Document Q&A
ivy.agent.document-qa · v1.0.0
ivyx
Indexes every PDF in a folder and has a model answer a question from the passages closest to it.
Steps
In the order the agent's file lists them, each with where its inputs come from and the values the file fixes. A branch or a loop is a step too.
- 01Ivy NodeRead the PDFsivy.node.pdf-folder-text
Extracts the text of every PDF in a folder, returning each document's text with its file name and all of it as one text.
- takes
folderfrom the agent's input- returns
documents
- 02Ivy NodeClean the textivy.node.text-normalizer
Cleans text before chunking: Unicode composition, control characters removed, runs of whitespace collapsed, optional lowercasing.
- takes
textfrom step 1
- 03Ivy NodeSplit into passagesivy.node.text-chunker
Splits a long text document into overlapping chunks suitable for embedding and retrieval-augmented generation (RAG).
- takes
textfrom step 2- set
chunk_size500overlap100
- 04Ivy NodeIndex the passagesivy.node.kb-index
Embeds text chunks and stores them in a knowledge-base file that kb-search reads, returning how many chunks it holds.
- takes
chunksfrom step 3- set
index_pathkb/documents.jsonreplacetrue- returns
chunks,index_path
- 05Ivy NodeFind the closest passagesivy.node.kb-search
Finds the chunks of a knowledge-base file closest in meaning to a question and returns them as context, best first, with their scores.
- takes
index_pathfrom step 4questionfrom the agent's input- set
top_k3- returns
sources
- 06Model turnAnswer from the passages
Answer the question using only the passages. If the passages do not contain the answer, say that the documents do not say. Answer in one or two sentences.
- takes
context.questionfrom the agent's input,questioncontext.passagesfrom step 5,context- returns
answer
Nodes it brings
Adding this agent in IVYX Studio adds these nodes with it. A node your workspace already has is kept as it is, even at another version.
What it touches
Collected from what each of its nodes declares, plus the model call when a step is a model turn. A declaration is the author's statement, and it is what policy rules select on.
Inputs
| Field | Type | Description |
|---|---|---|
| folderrequired | string | The folder to read PDFs from. |
| questionrequired | string | The question to answer from the documents. |
Outputs
| Field | Type | Description |
|---|---|---|
| answer | string | The model's answer, from the passages only. |
| sources | array | The passages the answer was given, as {text, score}, best first. |
| documents | integer | How many PDFs were read. |
| chunks | integer | How many passages the index holds. |
| index_path | string | The knowledge-base file. |
Tests
2 of 2 test cases passed on Oct 4, 2026, in the publisher's own environment, before this version was published. The registry keeps that record; it does not run the cases again.
Requires: python:3.9, a model core, the embedding model, downloaded on first use
- refund-five-days
Three policy PDFs: the refund period is read from the refund policy.
- given
folderdata/docsquestionHow long does it take to get a refund?- expects
documentsequals 3answermatches (5|[Ff]ive) business daysindex_pathmatches kb/documents\.json$
- refund-ten-days
Another folder states another period, so the answer comes from the folder given.
- given
folderdata/docs-bquestionHow long does it take to get a refund?- expects
answermatches (10|[Tt]en) business days