← All ivy-nodes
IVYXSTUDIO · IVY NODE
T

Text Normalizer

ivy.node.text-normalizer · v0.1.0

ivyx✓

Cleans text before chunking: Unicode composition, control characters removed, runs of whitespace collapsed, optional lowercasing. Returns the cleaned text and how many characters it dropped, so a pipeline can see whether its source was as tidy as it looked. Sits directly in front of text-chunker; pure compute, no network and no file access.

#text#rag#cleaning

Inputs

FieldTypeDescription
textrequiredstringThe text to clean.
collapse_whitespacebooleanCollapse runs of spaces and tabs to one space and trim. Default true.
strip_controlbooleanRemove control characters, keeping newlines and tabs. Default true.
lowercasebooleanLowercase the result. Default false.
unicode_formstringUnicode normalization form: NFC, NFD, NFKC or NFKD.

Outputs

FieldTypeDescription
textrequiredstringThe cleaned text.
removedrequiredintegerCharacters dropped: source length minus result length.

Source

python

# Input preparation
inp = __ivy_ctx__["nodes"][__ivy_node_id__]["input"]
text = inp["text"]
collapse_whitespace = bool(inp.get("collapse_whitespace", True))
strip_control = bool(inp.get("strip_control", True))
lowercase = bool(inp.get("lowercase", False))
form = inp.get("unicode_form", "NFC")

# Compute
import re
import unicodedata


def normalize(text, *, collapse_whitespace=True, strip_control=True,
              lowercase=False, form="NFC"):
    """
    Clean text before it is chunked and embedded.

    The order is not arbitrary. Unicode composition runs first, so a combining
    accent is one character before anything measures length; then control
    characters; then whitespace, so a control character removed from the
    middle of a word does not leave a double space behind it.

    A non-breaking space counts as whitespace. It reads as a space to a
    person, and a chunker that treats it as a letter produces chunks that are
    subtly longer than the text looks.
    """
    if not isinstance(text, str):
        raise ValueError("text must be a string")
    if form not in ("NFC", "NFD", "NFKC", "NFKD"):
        raise ValueError("unicode_form must be one of NFC, NFD, NFKC, NFKD")
    result = unicodedata.normalize(form, text)
    if strip_control:
        result = "".join(
            ch for ch in result
            if ch in ("\n", "\t") or unicodedata.category(ch)[0] != "C"
        )
    if collapse_whitespace:
        # Horizontal whitespace only, so paragraph structure survives.
        result = re.sub(r"[^\S\n]+", " ", result)
        result = re.sub(r"[ \t]*\n[ \t]*", "\n", result)
        # A blank line is a paragraph boundary and it SURVIVES: collapsing it
        # would erase the structure a chunker is trying to respect. Three or
        # more newlines are not a bigger boundary than one blank line, so they
        # become exactly one.
        result = re.sub(r"\n{3,}", "\n\n", result)
        result = result.strip()
    if lowercase:
        result = result.lower()
    return result


cleaned = normalize(text, collapse_whitespace=collapse_whitespace,
                    strip_control=strip_control, lowercase=lowercase,
                    form=form)

# Output collection (runner reads __ivy_ctx__)
out = __ivy_ctx__["nodes"][__ivy_node_id__]["output"]
out["text"] = cleaned
out["removed"] = len(text) - len(cleaned)

Tests

Requires: python:3.9

  • collapse-and-trim

    Runs of spaces, tabs and a non-breaking space collapse to single spaces, and the result is trimmed.

  • paragraph-break-survives

    A blank line is a paragraph boundary and is kept; only the spaces around it go.

  • runs-of-newlines-collapse-to-one-break

    Four newlines are not a bigger boundary than one blank line.

  • control-chars

    A zero-width space and a bell character are removed, and the count says two went.

  • unknown-unicode-form

    A Unicode form the node does not support is named as an error, with the list.