Handle nulls and errors¶
A null input gives a null result; an invalid parameter raises. That one rule governs the whole surface, and this page shows how to work with it. For the exhaustive table of cases, see Reference / Parameters and contracts.
Let nulls flow through¶
A null in any input on a row produces null on that row, never a stand-in constant, so you can score a frame with
gaps and deal with them downstream:
import polars as pl
import polars_stats as ps
df = pl.DataFrame({"x": [0.0, None, 1.0]}, schema={"x": pl.Float64})
print(df.with_columns(density=ps.Normal().pdf("x")))
shape: (3, 2)
┌──────┬──────────┐
│ x ┆ density │
│ --- ┆ --- │
│ f64 ┆ f64 │
╞══════╪══════════╡
│ 0.0 ┆ 0.398942 │
│ null ┆ null │
│ 1.0 ┆ 0.241971 │
└──────┴──────────┘
A null parameter behaves the same way: that row's result is null, and no error is raised. This matters when
parameters are estimated, because an under-sized group yields a null sigma. The rule has no exceptions, so a
null parameter nulls the row even where the evaluation point alone would have settled the answer; see
Reference / Parameters and contracts.
Read an undefined moment¶
A moment the distribution does not have is null on every row whose parameters are valid. Cauchy has no
moments of any order, so mean(), variance() and std() are null while median() and entropy() stay
ordinary values:
heavy_tailed = ps.Cauchy(loc=0.0, scale=2.0)
print(
pl.DataFrame({"_": [0]}).select(
mean=heavy_tailed.mean(), median=heavy_tailed.median(), entropy=heavy_tailed.entropy()
)
)
shape: (1, 3)
┌──────┬────────┬──────────┐
│ mean ┆ median ┆ entropy │
│ --- ┆ --- ┆ --- │
│ f64 ┆ f64 ┆ f64 │
╞══════╪════════╪══════════╡
│ null ┆ 0.0 ┆ 3.224171 │
└──────┴────────┴──────────┘
That null is an answer rather than a failure, so a downstream .mean(), .sum() or .drop_nulls() absorbs
it silently. Check for it where a moment feeds an aggregate. An invalid scale on the row still raises, so a
null never stands in for a bad parameter.
A moment whose integral diverges is +inf, as in scipy.
Pareto(scale, shape).mean() is +inf for shape <= 1 and variance() / std() for shape <= 2, and each is
finite above its threshold, so a column of shapes straddling 1 yields finite means and infinities rather than
nulls, and an aggregate over it is inf rather than silently shortened.
Find the rows that would raise¶
An invalid parameter value (sigma <= 0, max <= min, p outside [0, 1], a NaN or an infinity) fails the
whole evaluation with a ComputeError, whatever else is on the row. Locate the offending rows with an ordinary filter
before scoring:
frame = pl.DataFrame(
{
"reading": [1.0, 2.0, 3.0],
"mu": [0.0, 0.0, 0.0],
"sigma": [1.0, -1.0, 0.0],
}
)
print(frame.filter(pl.col("sigma") <= 0))
shape: (2, 3)
┌─────────┬─────┬───────┐
│ reading ┆ mu ┆ sigma │
│ --- ┆ --- ┆ --- │
│ f64 ┆ f64 ┆ f64 │
╞═════════╪═════╪═══════╡
│ 2.0 ┆ 0.0 ┆ -1.0 │
│ 3.0 ┆ 0.0 ┆ 0.0 │
└─────────┴─────┴───────┘
Score anyway, quarantining the bad rows¶
Filter the invalid rows out of the scored branch. Nulls need no special handling; they propagate:
dist = ps.Normal(mu="mu", sigma="sigma")
scored = frame.filter(pl.col("sigma") > 0).with_columns(upper_tail=dist.sf("reading"))
quarantined = frame.filter(pl.col("sigma") <= 0)
print(scored)
print(quarantined)
shape: (1, 4)
┌─────────┬─────┬───────┬────────────┐
│ reading ┆ mu ┆ sigma ┆ upper_tail │
│ --- ┆ --- ┆ --- ┆ --- │
│ f64 ┆ f64 ┆ f64 ┆ f64 │
╞═════════╪═════╪═══════╪════════════╡
│ 1.0 ┆ 0.0 ┆ 1.0 ┆ 0.158655 │
└─────────┴─────┴───────┴────────────┘
shape: (2, 3)
┌─────────┬─────┬───────┐
│ reading ┆ mu ┆ sigma │
│ --- ┆ --- ┆ --- │
│ f64 ┆ f64 ┆ f64 │
╞═════════╪═════╪═══════╡
│ 2.0 ┆ 0.0 ┆ -1.0 │
│ 3.0 ┆ 0.0 ┆ 0.0 │
└─────────┴─────┴───────┘
If you would rather keep every row, replace the invalid parameters with null and let the result null out:
guarded = frame.with_columns(sigma=pl.when(pl.col("sigma") > 0).then("sigma").otherwise(None)).with_columns(
upper_tail=dist.sf("reading")
)
print(guarded)
shape: (3, 4)
┌─────────┬─────┬───────┬────────────┐
│ reading ┆ mu ┆ sigma ┆ upper_tail │
│ --- ┆ --- ┆ --- ┆ --- │
│ f64 ┆ f64 ┆ f64 ┆ f64 │
╞═════════╪═════╪═══════╪════════════╡
│ 1.0 ┆ 0.0 ┆ 1.0 ┆ 0.158655 │
│ 2.0 ┆ 0.0 ┆ null ┆ null │
│ 3.0 ┆ 0.0 ┆ null ┆ null │
└─────────┴─────┴───────┴────────────┘
Doing this converts a loud failure into a silent one, so do it only where a null result is genuinely the answer you want.
Catch the error instead¶
When failing the query is acceptable and you only need to report it, catch pl.exceptions.ComputeError. The message
names the parameter and the offending values:
try:
frame.with_columns(dist.sf("reading"))
except pl.exceptions.ComputeError as exc:
print(str(exc).splitlines()[0])
A scalar parameter is checked once and a column parameter over the whole column before any row computes, and a bad scalar and a bad column row surface the same way:
try:
pl.DataFrame({"x": [0.5]}).with_columns(ps.Normal(mu=0.0, sigma=-1.0).pdf("x"))
except pl.exceptions.ComputeError as exc:
print(str(exc).splitlines()[0])
A wrong parameter type is caught earlier, at construction, with a TypeError: no query runs.
Related¶
- Reference / Parameters and contracts: the full
table, including
NaN, out-of-support, and out-of-range quantiles. - Explanation / Design notes: why an invalid parameter raises rather than nulling.
- Explanation / Design notes:
why an undefined moment is
nulland a divergent one is+inf.