Skip to content

Handle nulls and errors

A null input gives a null result; an invalid parameter raises. That one rule governs the whole surface, and this page shows how to work with it. For the exhaustive table of cases, see Reference / Parameters and contracts.

Let nulls flow through

A null in any input on a row produces null on that row, never a stand-in constant, so you can score a frame with gaps and deal with them downstream:

import polars as pl
import polars_stats as ps

df = pl.DataFrame({"x": [0.0, None, 1.0]}, schema={"x": pl.Float64})

print(df.with_columns(density=ps.Normal().pdf("x")))
shape: (3, 2)
┌──────┬──────────┐
│ x    ┆ density  │
│ ---  ┆ ---      │
│ f64  ┆ f64      │
╞══════╪══════════╡
│ 0.0  ┆ 0.398942 │
│ null ┆ null     │
│ 1.0  ┆ 0.241971 │
└──────┴──────────┘

A null parameter behaves the same way: that row's result is null, and no error is raised. This matters when parameters are estimated, because an under-sized group yields a null sigma. The rule has no exceptions, so a null parameter nulls the row even where the evaluation point alone would have settled the answer; see Reference / Parameters and contracts.

Read an undefined moment

A moment the distribution does not have is null on every row whose parameters are valid. Cauchy has no moments of any order, so mean(), variance() and std() are null while median() and entropy() stay ordinary values:

heavy_tailed = ps.Cauchy(loc=0.0, scale=2.0)

print(
    pl.DataFrame({"_": [0]}).select(
        mean=heavy_tailed.mean(), median=heavy_tailed.median(), entropy=heavy_tailed.entropy()
    )
)
shape: (1, 3)
┌──────┬────────┬──────────┐
│ mean ┆ median ┆ entropy  │
│ ---  ┆ ---    ┆ ---      │
│ f64  ┆ f64    ┆ f64      │
╞══════╪════════╪══════════╡
│ null ┆ 0.0    ┆ 3.224171 │
└──────┴────────┴──────────┘

That null is an answer rather than a failure, so a downstream .mean(), .sum() or .drop_nulls() absorbs it silently. Check for it where a moment feeds an aggregate. An invalid scale on the row still raises, so a null never stands in for a bad parameter.

A moment whose integral diverges is +inf, as in scipy. Pareto(scale, shape).mean() is +inf for shape <= 1 and variance() / std() for shape <= 2, and each is finite above its threshold, so a column of shapes straddling 1 yields finite means and infinities rather than nulls, and an aggregate over it is inf rather than silently shortened.

Find the rows that would raise

An invalid parameter value (sigma <= 0, max <= min, p outside [0, 1], a NaN or an infinity) fails the whole evaluation with a ComputeError, whatever else is on the row. Locate the offending rows with an ordinary filter before scoring:

frame = pl.DataFrame(
    {
        "reading": [1.0, 2.0, 3.0],
        "mu": [0.0, 0.0, 0.0],
        "sigma": [1.0, -1.0, 0.0],
    }
)

print(frame.filter(pl.col("sigma") <= 0))
shape: (2, 3)
┌─────────┬─────┬───────┐
│ reading ┆ mu  ┆ sigma │
│ ---     ┆ --- ┆ ---   │
│ f64     ┆ f64 ┆ f64   │
╞═════════╪═════╪═══════╡
│ 2.0     ┆ 0.0 ┆ -1.0  │
│ 3.0     ┆ 0.0 ┆ 0.0   │
└─────────┴─────┴───────┘

Score anyway, quarantining the bad rows

Filter the invalid rows out of the scored branch. Nulls need no special handling; they propagate:

dist = ps.Normal(mu="mu", sigma="sigma")

scored = frame.filter(pl.col("sigma") > 0).with_columns(upper_tail=dist.sf("reading"))
quarantined = frame.filter(pl.col("sigma") <= 0)

print(scored)
print(quarantined)
shape: (1, 4)
┌─────────┬─────┬───────┬────────────┐
│ reading ┆ mu  ┆ sigma ┆ upper_tail │
│ ---     ┆ --- ┆ ---   ┆ ---        │
│ f64     ┆ f64 ┆ f64   ┆ f64        │
╞═════════╪═════╪═══════╪════════════╡
│ 1.0     ┆ 0.0 ┆ 1.0   ┆ 0.158655   │
└─────────┴─────┴───────┴────────────┘
shape: (2, 3)
┌─────────┬─────┬───────┐
│ reading ┆ mu  ┆ sigma │
│ ---     ┆ --- ┆ ---   │
│ f64     ┆ f64 ┆ f64   │
╞═════════╪═════╪═══════╡
│ 2.0     ┆ 0.0 ┆ -1.0  │
│ 3.0     ┆ 0.0 ┆ 0.0   │
└─────────┴─────┴───────┘

If you would rather keep every row, replace the invalid parameters with null and let the result null out:

guarded = frame.with_columns(sigma=pl.when(pl.col("sigma") > 0).then("sigma").otherwise(None)).with_columns(
    upper_tail=dist.sf("reading")
)
print(guarded)
shape: (3, 4)
┌─────────┬─────┬───────┬────────────┐
│ reading ┆ mu  ┆ sigma ┆ upper_tail │
│ ---     ┆ --- ┆ ---   ┆ ---        │
│ f64     ┆ f64 ┆ f64   ┆ f64        │
╞═════════╪═════╪═══════╪════════════╡
│ 1.0     ┆ 0.0 ┆ 1.0   ┆ 0.158655   │
│ 2.0     ┆ 0.0 ┆ null  ┆ null       │
│ 3.0     ┆ 0.0 ┆ null  ┆ null       │
└─────────┴─────┴───────┴────────────┘

Doing this converts a loud failure into a silent one, so do it only where a null result is genuinely the answer you want.

Catch the error instead

When failing the query is acceptable and you only need to report it, catch pl.exceptions.ComputeError. The message names the parameter and the offending values:

try:
    frame.with_columns(dist.sf("reading"))
except pl.exceptions.ComputeError as exc:
    print(str(exc).splitlines()[0])
the plugin failed with message: sigma must be finite and strictly positive, got -1

A scalar parameter is checked once and a column parameter over the whole column before any row computes, and a bad scalar and a bad column row surface the same way:

try:
    pl.DataFrame({"x": [0.5]}).with_columns(ps.Normal(mu=0.0, sigma=-1.0).pdf("x"))
except pl.exceptions.ComputeError as exc:
    print(str(exc).splitlines()[0])
the plugin failed with message: sigma must be finite and strictly positive, got -1

A wrong parameter type is caught earlier, at construction, with a TypeError: no query runs.