← All ivy-nodes
IVYXSTUDIO · IVY NODE
P

PDF Text

ivy.node.pdf-text · v0.1.0

ivyx✓

Extracts the text of a PDF file, page by page and as one text, with the page count. A scanned PDF with no text layer returns empty text.

#pdf#document#text#extract

Inputs

FieldTypeDescription
pathrequiredstringPath to the PDF file.

Outputs

FieldTypeDescription
textrequiredstringAll pages' text, separated by blank lines.
pagesrequiredarrayEach page's text.
page_countrequiredintegerHow many pages the PDF has.

Source

python

inp = __ivy_ctx__["nodes"][__ivy_node_id__]["input"]

from pypdf import PdfReader


def pdf_pages(path):
    reader = PdfReader(path)
    return [(page.extract_text() or "").strip() for page in reader.pages]

pages = pdf_pages(inp["path"])

out = __ivy_ctx__["nodes"][__ivy_node_id__]["output"]
out["text"] = "\n\n".join(p for p in pages if p)
out["pages"] = pages
out["page_count"] = len(pages)

Tests

Requires: python:3.9, pypdf

  • two-pages

    A two-page contract reads page by page.

  • not-a-pdf

    A file that is not a PDF is refused.