← All ivy-nodes
IVYXSTUDIO · IVY NODE
D

Data Frame Filter

ivy.node.data-frame-filter · v0.1.0

ivyx✓

Filters, projects and limits a list of row objects, returning the kept rows with how many were dropped. Conditions are ANDed and support eq, ne, gt, gte, lt, lte, contains and in; a row that lacks the column simply does not match rather than raising, so ragged records are safe to filter. Makes data-frame-creator useful without writing code; pure compute.

#data#filter#table

Inputs

FieldTypeDescription
rowsrequiredarrayThe rows to filter.
wherearrayConditions, ANDed. Each is {column, op, value}.
selectarrayKeep only these columns. All of them when omitted.
limitintegerKeep at most this many rows.

Outputs

FieldTypeDescription
rowsrequiredarrayThe kept rows.
countrequiredintegerHow many rows were kept.
droppedrequiredintegerHow many rows were removed.

Source

python

# Input preparation
inp = __ivy_ctx__["nodes"][__ivy_node_id__]["input"]
rows = inp["rows"]
where = inp.get("where") or []
select = inp.get("select")
limit = inp.get("limit")

# Compute
OPS = {
    "eq": lambda a, b: a == b,
    "ne": lambda a, b: a != b,
    "gt": lambda a, b: a is not None and a > b,
    "gte": lambda a, b: a is not None and a >= b,
    "lt": lambda a, b: a is not None and a < b,
    "lte": lambda a, b: a is not None and a <= b,
    "contains": lambda a, b: b in a if isinstance(a, (str, list)) else False,
    "in": lambda a, b: a in b if isinstance(b, list) else False,
}


def matches(row, condition):
    """One condition against one row. An absent column never matches."""
    column = condition.get("column")
    op = condition.get("op", "eq")
    if column is None:
        raise ValueError("every condition needs a column")
    if op not in OPS:
        raise ValueError(f"unknown op {op!r}; one of " + ", ".join(sorted(OPS)))
    if column not in row:
        return False
    return OPS[op](row[column], condition.get("value"))


def filter_rows(rows, where=(), select=None, limit=None):
    """
    Filter, project and limit a list of row dicts.

    Rows rather than a pandas frame on purpose: this node's neighbours pass
    JSON, and a frame would make every caller convert twice. `where` is a list
    of conditions and they are ANDed, which is the case a pipeline actually
    writes; anything richer belongs in code, not in a config object.

    A row missing the column simply does not match, rather than raising.
    Filtering heterogeneous records is the ordinary reason to reach for this,
    and a filter that crashes on a ragged row is no use there.
    """
    if not isinstance(rows, list):
        raise ValueError("rows must be a list of objects")
    kept = []
    for row in rows:
        if not isinstance(row, dict):
            raise ValueError("every row must be an object")
        if all(matches(row, c) for c in where):
            kept.append(row)
    if select is not None:
        if not isinstance(select, list):
            raise ValueError("select must be a list of column names")
        kept = [{k: r[k] for k in select if k in r} for r in kept]
    if limit is not None:
        if not isinstance(limit, int) or limit < 0:
            raise ValueError("limit must be a non-negative integer")
        kept = kept[:limit]
    return kept


result = filter_rows(rows, where=where, select=select, limit=limit)

# Output collection (runner reads __ivy_ctx__)
out = __ivy_ctx__["nodes"][__ivy_node_id__]["output"]
out["rows"] = result
out["count"] = len(result)
out["dropped"] = len(rows) - len(result)

Tests

Requires: python:3.9

  • filters-and-counts

    A greater-than condition keeps the matching rows and counts the dropped ones.

  • projects-and-limits

    select keeps named columns and limit caps the rows.

  • ragged-rows-do-not-crash

    A row missing the filtered column is dropped rather than raising.

  • unknown-op

    An operator the node does not have is named as an error, with the list.