← All ivy-nodes
IVYXSTUDIO · IVY NODE
K
Knowledge Base Index
ivy.node.kb-index · v0.1.0
ivyx✓
Embeds text chunks and stores them in a knowledge-base file that kb-search reads, returning how many chunks it holds. Uses a local embedding model (fastembed), so no service and no API key are needed.
#rag#embedding#index#knowledge-base#vector
Inputs
| Field | Type | Description |
|---|---|---|
| chunksrequired | array | The text chunks to index, such as text-chunker's chunks. |
| index_pathrequired | string | The knowledge-base file to write. |
| model | string | The embedding model; English by default, the multilingual one for other languages. |
| replace | boolean | Start the file afresh rather than adding to it. |
Outputs
| Field | Type | Description |
|---|---|---|
| countrequired | integer | How many chunks the file holds. |
| dimrequired | integer | The embedding size. |
| index_pathrequired | string | The file's absolute path. |
Source
python
inp = __ivy_ctx__["nodes"][__ivy_node_id__]["input"]
_MODEL_CACHE = globals().setdefault("_ivy_embed_models", {})
def embed(texts, model_name):
"""fastembed (ONNX, no torch). The model is loaded once per kernel."""
from fastembed import TextEmbedding
model = _MODEL_CACHE.get(model_name)
if model is None:
model = _MODEL_CACHE[model_name] = TextEmbedding(model_name)
return [[float(x) for x in v] for v in model.embed(list(texts))]
import json, os
chunks = [c for c in inp["chunks"] if isinstance(c, str) and c.strip()]
index_path = inp["index_path"]
model_name = inp.get("model", "BAAI/bge-small-en-v1.5")
replace = bool(inp.get("replace", True))
if not chunks:
raise ValueError("There are no chunks to index.")
index = {"model": model_name, "items": []}
if not replace and os.path.exists(index_path):
with open(index_path) as handle:
index = json.load(handle)
if index.get("model") != model_name:
raise ValueError(f"{index_path} was built with {index.get('model')}; use that model or replace it.")
vectors = embed(chunks, model_name)
index["items"].extend({"text": t, "vector": v} for t, v in zip(chunks, vectors))
os.makedirs(os.path.dirname(os.path.abspath(index_path)), exist_ok=True)
with open(index_path, "w") as handle:
json.dump(index, handle)
out = __ivy_ctx__["nodes"][__ivy_node_id__]["output"]
out["count"] = len(index["items"])
out["dim"] = len(vectors[0])
out["index_path"] = os.path.abspath(index_path)Tests
Requires: python:3.9, fastembed, model download on first use
- indexes
Three chunks are embedded at 384 dimensions.
- nothing
Empty chunks are refused.