← All ivy-nodes
IVYXSTUDIO · IVY NODE
W

Web Page Text

ivy.node.web-page-text · v0.1.0

ivyx✓

Fetches a web page and returns its readable text and title, without scripts, styles or markup. A page that answers with an error status fails the step.

#web#http#fetch#page#text

Inputs

FieldTypeDescription
urlrequiredstringThe page's address.
timeout_secondsnumberGive up after this many seconds.

Outputs

FieldTypeDescription
textrequiredstringThe page's readable text, one block per line.
titlerequiredstringThe page title.
statusrequiredintegerThe HTTP status.
urlrequiredstringThe address after redirects.

Source

python

inp = __ivy_ctx__["nodes"][__ivy_node_id__]["input"]

import requests
from bs4 import BeautifulSoup

url = inp["url"]
timeout = float(inp.get("timeout_seconds", 20))
response = requests.get(url, timeout=timeout, headers={"User-Agent": "IVY node web-page-text"})
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
for tag in soup(["script", "style", "noscript", "template"]):
    tag.decompose()
title = soup.title.get_text(strip=True) if soup.title else ""
if soup.title:
    soup.title.decompose()
text = "\n".join(line for line in (l.strip() for l in soup.get_text("\n").splitlines()) if line)

out = __ivy_ctx__["nodes"][__ivy_node_id__]["output"]
out["text"] = text
out["title"] = title
out["status"] = response.status_code
out["url"] = response.url

Tests

Requires: python:3.9, requests, beautifulsoup4

  • reads

    A page's text comes back without its script or style.

  • missing

    A 404 fails the step.