← All ivy-nodes
IVYXSTUDIO · IVY NODE
W
Web Page Text
ivy.node.web-page-text · v0.1.0
ivyx✓
Fetches a web page and returns its readable text and title, without scripts, styles or markup. A page that answers with an error status fails the step.
#web#http#fetch#page#text
Inputs
| Field | Type | Description |
|---|---|---|
| urlrequired | string | The page's address. |
| timeout_seconds | number | Give up after this many seconds. |
Outputs
| Field | Type | Description |
|---|---|---|
| textrequired | string | The page's readable text, one block per line. |
| titlerequired | string | The page title. |
| statusrequired | integer | The HTTP status. |
| urlrequired | string | The address after redirects. |
Source
python
inp = __ivy_ctx__["nodes"][__ivy_node_id__]["input"]
import requests
from bs4 import BeautifulSoup
url = inp["url"]
timeout = float(inp.get("timeout_seconds", 20))
response = requests.get(url, timeout=timeout, headers={"User-Agent": "IVY node web-page-text"})
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
for tag in soup(["script", "style", "noscript", "template"]):
tag.decompose()
title = soup.title.get_text(strip=True) if soup.title else ""
if soup.title:
soup.title.decompose()
text = "\n".join(line for line in (l.strip() for l in soup.get_text("\n").splitlines()) if line)
out = __ivy_ctx__["nodes"][__ivy_node_id__]["output"]
out["text"] = text
out["title"] = title
out["status"] = response.status_code
out["url"] = response.urlTests
Requires: python:3.9, requests, beautifulsoup4
- reads
A page's text comes back without its script or style.
- missing
A 404 fails the step.