← All ivy-nodes
IVYXSTUDIO · IVY NODE
P
PDF Text
ivy.node.pdf-text · v0.1.0
ivyx✓
Extracts the text of a PDF file, page by page and as one text, with the page count. A scanned PDF with no text layer returns empty text.
#pdf#document#text#extract
Inputs
| Field | Type | Description |
|---|---|---|
| pathrequired | string | Path to the PDF file. |
Outputs
| Field | Type | Description |
|---|---|---|
| textrequired | string | All pages' text, separated by blank lines. |
| pagesrequired | array | Each page's text. |
| page_countrequired | integer | How many pages the PDF has. |
Source
python
inp = __ivy_ctx__["nodes"][__ivy_node_id__]["input"]
from pypdf import PdfReader
def pdf_pages(path):
reader = PdfReader(path)
return [(page.extract_text() or "").strip() for page in reader.pages]
pages = pdf_pages(inp["path"])
out = __ivy_ctx__["nodes"][__ivy_node_id__]["output"]
out["text"] = "\n\n".join(p for p in pages if p)
out["pages"] = pages
out["page_count"] = len(pages)Tests
Requires: python:3.9, pypdf
- two-pages
A two-page contract reads page by page.
- not-a-pdf
A file that is not a PDF is refused.