Text Normalizer
ivy.node.text-normalizer · v0.1.0
ivyx✓
Cleans text before chunking: Unicode composition, control characters removed, runs of whitespace collapsed, optional lowercasing. Returns the cleaned text and how many characters it dropped, so a pipeline can see whether its source was as tidy as it looked. Sits directly in front of text-chunker; pure compute, no network and no file access.
Inputs
| Field | Type | Description |
|---|---|---|
| textrequired | string | The text to clean. |
| collapse_whitespace | boolean | Collapse runs of spaces and tabs to one space and trim. Default true. |
| strip_control | boolean | Remove control characters, keeping newlines and tabs. Default true. |
| lowercase | boolean | Lowercase the result. Default false. |
| unicode_form | string | Unicode normalization form: NFC, NFD, NFKC or NFKD. |
Outputs
| Field | Type | Description |
|---|---|---|
| textrequired | string | The cleaned text. |
| removedrequired | integer | Characters dropped: source length minus result length. |
Source
python
# Input preparation
inp = __ivy_ctx__["nodes"][__ivy_node_id__]["input"]
text = inp["text"]
collapse_whitespace = bool(inp.get("collapse_whitespace", True))
strip_control = bool(inp.get("strip_control", True))
lowercase = bool(inp.get("lowercase", False))
form = inp.get("unicode_form", "NFC")
# Compute
import re
import unicodedata
def normalize(text, *, collapse_whitespace=True, strip_control=True,
lowercase=False, form="NFC"):
"""
Clean text before it is chunked and embedded.
The order is not arbitrary. Unicode composition runs first, so a combining
accent is one character before anything measures length; then control
characters; then whitespace, so a control character removed from the
middle of a word does not leave a double space behind it.
A non-breaking space counts as whitespace. It reads as a space to a
person, and a chunker that treats it as a letter produces chunks that are
subtly longer than the text looks.
"""
if not isinstance(text, str):
raise ValueError("text must be a string")
if form not in ("NFC", "NFD", "NFKC", "NFKD"):
raise ValueError("unicode_form must be one of NFC, NFD, NFKC, NFKD")
result = unicodedata.normalize(form, text)
if strip_control:
result = "".join(
ch for ch in result
if ch in ("\n", "\t") or unicodedata.category(ch)[0] != "C"
)
if collapse_whitespace:
# Horizontal whitespace only, so paragraph structure survives.
result = re.sub(r"[^\S\n]+", " ", result)
result = re.sub(r"[ \t]*\n[ \t]*", "\n", result)
# A blank line is a paragraph boundary and it SURVIVES: collapsing it
# would erase the structure a chunker is trying to respect. Three or
# more newlines are not a bigger boundary than one blank line, so they
# become exactly one.
result = re.sub(r"\n{3,}", "\n\n", result)
result = result.strip()
if lowercase:
result = result.lower()
return result
cleaned = normalize(text, collapse_whitespace=collapse_whitespace,
strip_control=strip_control, lowercase=lowercase,
form=form)
# Output collection (runner reads __ivy_ctx__)
out = __ivy_ctx__["nodes"][__ivy_node_id__]["output"]
out["text"] = cleaned
out["removed"] = len(text) - len(cleaned)Tests
Requires: python:3.9
- collapse-and-trim
Runs of spaces, tabs and a non-breaking space collapse to single spaces, and the result is trimmed.
- paragraph-break-survives
A blank line is a paragraph boundary and is kept; only the spaces around it go.
- runs-of-newlines-collapse-to-one-break
Four newlines are not a bigger boundary than one blank line.
- control-chars
A zero-width space and a bell character are removed, and the count says two went.
- unknown-unicode-form
A Unicode form the node does not support is named as an error, with the list.