← All ivy-nodes
IVYXSTUDIO · IVY NODE
P
PDF Folder Text
ivy.node.pdf-folder-text · v0.1.0
ivyx✓
Extracts the text of every PDF in a folder, returning each document's text with its file name and all of it as one text. Files that are not PDFs are skipped and a PDF that cannot be read is reported rather than failing the step.
#pdf#document#folder#text#rag
Inputs
| Field | Type | Description |
|---|---|---|
| folderrequired | string | The folder to read PDFs from. |
Outputs
| Field | Type | Description |
|---|---|---|
| documentsrequired | array | Each PDF as {name, text, page_count}, by file name. |
| textrequired | string | Every document's text, each under a '# name' heading. |
| countrequired | integer | How many PDFs were read. |
| unreadablerequired | array | PDFs that could not be read, as {name, error}. |
Source
python
inp = __ivy_ctx__["nodes"][__ivy_node_id__]["input"]
from pypdf import PdfReader
def pdf_pages(path):
reader = PdfReader(path)
return [(page.extract_text() or "").strip() for page in reader.pages]
import os
folder = inp["folder"]
if not os.path.isdir(folder):
raise NotADirectoryError(f"{folder} is not a folder.")
documents, unreadable = [], []
for name in sorted(os.listdir(folder)):
if not name.lower().endswith(".pdf"):
continue
try:
pages = pdf_pages(os.path.join(folder, name))
except Exception as error:
unreadable.append({"name": name, "error": str(error)[:200]})
continue
documents.append({"name": name, "text": "\n\n".join(p for p in pages if p), "page_count": len(pages)})
out = __ivy_ctx__["nodes"][__ivy_node_id__]["output"]
out["documents"] = documents
out["text"] = "\n\n".join(f"# {d['name']}\n\n{d['text']}" for d in documents)
out["count"] = len(documents)
out["unreadable"] = unreadableTests
Requires: python:3.9, pypdf
- folder
Two PDFs are read in name order and the text file is skipped.
- no-folder
A path that is not a folder is refused.