← All ivy-nodes
IVYXSTUDIO · IVY NODE
P

PDF Folder Text

ivy.node.pdf-folder-text · v0.1.0

ivyx✓

Extracts the text of every PDF in a folder, returning each document's text with its file name and all of it as one text. Files that are not PDFs are skipped and a PDF that cannot be read is reported rather than failing the step.

#pdf#document#folder#text#rag

Inputs

FieldTypeDescription
folderrequiredstringThe folder to read PDFs from.

Outputs

FieldTypeDescription
documentsrequiredarrayEach PDF as {name, text, page_count}, by file name.
textrequiredstringEvery document's text, each under a '# name' heading.
countrequiredintegerHow many PDFs were read.
unreadablerequiredarrayPDFs that could not be read, as {name, error}.

Source

python

inp = __ivy_ctx__["nodes"][__ivy_node_id__]["input"]

from pypdf import PdfReader


def pdf_pages(path):
    reader = PdfReader(path)
    return [(page.extract_text() or "").strip() for page in reader.pages]

import os

folder = inp["folder"]
if not os.path.isdir(folder):
    raise NotADirectoryError(f"{folder} is not a folder.")
documents, unreadable = [], []
for name in sorted(os.listdir(folder)):
    if not name.lower().endswith(".pdf"):
        continue
    try:
        pages = pdf_pages(os.path.join(folder, name))
    except Exception as error:
        unreadable.append({"name": name, "error": str(error)[:200]})
        continue
    documents.append({"name": name, "text": "\n\n".join(p for p in pages if p), "page_count": len(pages)})

out = __ivy_ctx__["nodes"][__ivy_node_id__]["output"]
out["documents"] = documents
out["text"] = "\n\n".join(f"# {d['name']}\n\n{d['text']}" for d in documents)
out["count"] = len(documents)
out["unreadable"] = unreadable

Tests

Requires: python:3.9, pypdf

  • folder

    Two PDFs are read in name order and the text file is skipped.

  • no-folder

    A path that is not a folder is refused.