Parameters and contracts¶
What every distribution accepts, what it returns, and what happens when an input is missing or invalid. This page is the lookup table; the how-to guides show what to do about it.
Accepted inputs¶
Constructor parameters and value-keyed method arguments follow the same coercion rules, with one asymmetry: a float
parameter rejects an int and an integer bound rejects a float, while a method argument accepts either.
The bounds of Uniform and DiscreteUniform sit on opposite sides of the float rule despite sharing the names min
and max.
| Input | Float parameter (mu, sigma, loc, scale, p, rate, a, b, Uniform's min / max) |
Count parameter (n) |
Integer bound (DiscreteUniform's min / max) |
Method argument (x, q) |
|---|---|---|---|---|
float |
accepted | TypeError |
TypeError |
accepted |
int |
TypeError |
accepted (0 to 2**63 - 1) |
accepted (any Int64) |
accepted |
bool |
TypeError |
TypeError |
TypeError |
TypeError |
str |
read as pl.col(name) |
read as pl.col(name) |
read as pl.col(name) |
read as pl.col(name) |
pl.Expr |
passed through | passed through | passed through | passed through |
pl.Series |
wrapped as pl.lit(series) |
wrapped as pl.lit(series) |
wrapped as pl.lit(series) |
wrapped as pl.lit(series) |
| anything else | TypeError |
TypeError |
TypeError |
TypeError |
A str is always a column reference, never a literal value. A Python scalar becomes pl.lit(value), a length-1
scalar column that the Rust plugin broadcasts to the call's row count
(see Architecture).
Polars' own scalar semantics then apply. An expression whose inputs are all constant is a scalar column:
df.select(Normal(0.0, 1.0).mean()) returns one row, group_by().agg() returns a scalar rather than a list per
group, and a 0-row frame still returns one row. Any column-valued input sets the length instead.
sample() and samples() are the exception. They pass a per-row index as a hidden full-length input, so they are
full height whatever the parameters are, and a 0-row frame returns no rows.
A length-1 expression is accepted wherever a column is and broadcasts rather than truncating, so
Normal(mu=pl.col("mu").mean(), sigma=1.0).pdf("x") is full height. Lengths that are neither equal nor 1 raise.
Length-1 inputs inside over() and group_by().agg() need polars 1.34
Before polars 1.34 a length-1 input is mishandled by polars itself once the expression is evaluated inside a
partition context. Depending on the distribution and method you get either a PanicException or, below 1.33,
values returned in group order rather than scattered back to row order. Both defects are fixed upstream and
need no change here.
Nothing else is affected. select and with_columns broadcast correctly on every supported version, and so
does the pattern of aggregating first and applying the distribution to the result:
(
frame.group_by("g")
.agg(mu=pl.col("x").mean(), sigma=pl.col("x").std())
.with_columns(p=ps.Normal(mu="mu", sigma="sigma").cdf(1.0))
)
Computing a parameter with .over("g") and using it in a plain select is fine too; it is only putting the
distribution expression itself inside over() or agg() that is affected.
A column in either position must be numeric: any Int* or UInt* width up to Int128 / UInt128, Float16 /
Float32 / Float64, or Decimal. A Null-typed column counts as numeric and propagates nulls. Numeric columns
are cast to Float64 at evaluation, so an integer column works wherever a float is expected.
That cast is the plugin's. The closed-form moments are polars arithmetic on the parameter itself, so there polars
decides: a moment that is the parameter (Normal(mu="mu").mean()) keeps the column's dtype, and a Decimal or
Null-typed sigma raises InvalidOperationError where polars has no kernel for it on the bare column (pow,
log, exp: Normal's variance and entropy, LogNormal's mean, variance, std and entropy, Pareto's
variance and entropy, Weibull's entropy).
The integer parameters are the exception to the cast: the count n and DiscreteUniform's bounds must already hold
integers, of any integer dtype, because casting a float one would silently truncate. n widens to UInt64 and the
bounds to Int64, so a UInt64 bound value above i64::MAX raises rather than wrapping. The rule is judged on the
dtype, so a float column raises even when every value in it is null; a Null-dtype column
propagates nulls like any other parameter.
In either position the Rust plugin enforces that rule on every method: a Boolean, String, Categorical,
Enum, Struct, Object or temporal column raises ComputeError naming the column and the dtype the plugin
received (Object arrives as binary), and no row computes. Nothing is parsed and nothing is read as 0 / 1.
The one variation is the exception type: where a closed-form moment's own polars arithmetic meets a parameter
column before the plugin does (n * p on a String p), polars raises InvalidOperationError instead. Cast a
column to Float64 when you mean a number. (On polars older than 1.25, an Object column in either position raises
PanicException from polars' own arrow export instead.)
Parameter validity¶
Values are validated at evaluation, not at construction, and identically for scalar and column-valued parameters.
The count parameter n is the exception: a Python int outside [0, 2**63 - 1] raises ValueError at
construction, since it is coerced to a UInt64 literal and passed to the fast paths as a kwarg. An n column
may hold any count its dtype can, up to UInt64.
| Distribution | Required | Also required |
|---|---|---|
Beta(a, b) |
a > 0, b > 0 |
both finite |
Cauchy(loc, scale) |
scale > 0 |
both finite |
Exponential(rate) |
rate > 0 |
finite |
LogNormal(mu, sigma) |
sigma > 0 |
both finite |
Normal(mu, sigma) |
sigma > 0 |
both finite |
Pareto(scale, shape) |
scale > 0, shape > 0 |
both finite |
Uniform(min, max) |
max > min |
max - min finite |
Weibull(shape, scale) |
shape > 0, scale > 0 |
both finite |
Bernoulli(p) |
0 <= p <= 1 |
|
Binomial(n, p) |
n >= 0, 0 <= p <= 1 |
n integral |
DiscreteUniform(min, max) |
min <= max, both inclusive |
max - min + 1 fits Int64 |
Geometric(p) |
0 < p <= 1 |
p = 0 rejected, unlike Bernoulli |
A violation raises ComputeError and fails the whole evaluation. The check runs over each parameter column before
any row computes, so an invalid value raises even on a row whose other parameter is null or whose evaluation point
is null or NaN. See nulls, NaNs and errors below.
Sampling¶
sample(seed) returns one variate per row; samples(size, seed) returns Array(inner=<element dtype>, shape=size).
Both arguments are checked at call time. A size <= 0 raises ValueError, and so does a seed outside
[0, 2**63). The seed crosses FFI as a pickled kwarg that Rust decodes as i64, one bit short of the u64 it
is read into, so polars' own Expr.sample takes seeds this library refuses. A seed or size that is not an
int raises TypeError, bool and numpy.int64 included.
size has no maximum. A call allocates rows * size elements up front: a product that does not fit a usize,
or an allocation the allocator refuses, raises ComputeError. When the request reaches the allocator, the message
names the size, the row count and the byte count; a size too large for the plugin's kwargs decoder to read
raises before that with a generic message. A request the allocator accepts but the machine cannot back is still
killed by the OS, with no exception to catch.
Element dtype is per distribution and is not normalised to Float64:
| Distribution | Sample dtype |
|---|---|
Bernoulli |
Boolean |
Binomial, Geometric |
UInt64 |
DiscreteUniform |
Int64 |
Beta, Cauchy, Exponential, LogNormal, Normal, Pareto, Uniform, Weibull |
Float64 |
| Aspect | Behaviour |
|---|---|
| Output length | the surrounding context: frame length under select / with_columns, partition length under over / group_by |
seed=<int> |
deterministic across OS, architecture, chunking, thread count, and engine (in-memory or streaming) |
seed=None |
non-reproducible by design (OS entropy), resolved once per call |
| Output name | "sample" / "samples" when every parameter is a scalar, otherwise the first parameter expression's root name |
samples(1, seed) |
bit-identical to sample(seed); growing size extends each row's array without changing existing draws |
| Null parameter on a row | sample yields null; samples yields a null array, not an array of nulls |
Nulls, NaNs and errors¶
null is reserved for missing inputs; an invalid parameter raises. A silent null from a bad parameter would be
indistinguishable from a legitimately missing input, and would propagate wrong answers downstream.
The rows are in precedence order: per row, the first that applies decides, so a raise outranks a null and a null
parameter outranks a NaN evaluation point. "Present" means non-null.
| Situation | When detected | Behaviour |
|---|---|---|
Wrong parameter type (a list, an int for a float parameter, a bool) |
Python __init__ |
TypeError, no query runs |
Non-numeric column in either position (Boolean, String, Categorical, Enum, Struct, Object, temporal) |
Rust evaluation | ComputeError naming the column and its dtype, no row computes; polars' InvalidOperationError where a closed-form moment's arithmetic meets a parameter column first (see Accepted inputs) |
Invalid present parameter value: outside its domain, NaN, +inf or -inf, as a scalar or on any column row |
Rust evaluation | ComputeError, fails the whole evaluation, never silently nulls, whatever else is on the row: a null sibling parameter, a null or NaN evaluation point |
| The same, on a 0-row frame | Rust evaluation | A constant parameterisation still raises; a parameter column returns an empty result (see below) |
rows * size in samples too large to address or allocate |
Rust evaluation | ComputeError, fails the whole evaluation (see Sampling) |
null parameter on a row |
per row | null on that row, on every method and at every evaluation point, NaN and off-support included (see below) |
null value or quantile argument on a row |
per row | null on that row |
A moment the distribution does not have (Cauchy.mean(), variance(), std()) |
per row | null on every valid row; an invalid parameter on the row still raises |
A moment whose integral diverges (Pareto.mean() at shape <= 1, variance() / std() at shape <= 2) |
per row | +inf, as scipy returns; a null is reserved for a moment with no value at all |
NaN value or quantile argument on a row |
per row | NaN on that row, ppf / isf included (matches scipy) |
q outside [0, 1] in ppf / isf |
per row | null, guaranteed for every distribution and both parameter regimes (pinned by tests/property/ppf_domain_test.py). q exactly 0 or 1 is in range and maps to a support bound |
x outside the support (e.g. pdf below a Uniform's min) |
per row | 0.0 (matches scipy) |
pmf(3.5) for a discrete distribution |
per row | 0.0 (matches scipy) |
Deep-tail underflow (sf of a 60-sigma event) |
per row | 0.0; log_sf keeps resolution where it has a stable form, see Numerical accuracy |
A null parameter nulls the answer, on the support and off it. A missing parameter is a missing answer, as
null * 0 is null in polars, so no off-support constant outranks it: on a row whose p is null,
Bernoulli(p="p").pmf(2) is null rather than 0.0, and so is Exponential(rate="rate").cdf(-1). The same holds
per bound for Uniform, whose two bounds might otherwise have settled the answer between them: on a row whose min
is null and whose max is 1.0, pdf(5.0) and cdf(5.0) are both null, as is every method at 0.5. Both
inverses null at every quantile under either null bound, in range and out.
On a 0-row frame, a constant parameterisation is still checked and a parameter column is not. A Python scalar and
a length-1 expression (pl.lit(5.0), pl.col("x").min()) are one parameterisation, validated once per call before
any row is read, so Uniform(5.0, 2.0) and Uniform(pl.lit(5.0), pl.lit(2.0)) both raise on an empty frame. A
parameter column is validated over its values, and an empty column has none, so Uniform(pl.col("lo"), pl.col("hi"))
returns an empty result instead. This applies to a value the strict cast refuses (Binomial(n=pl.lit(-5))) as much as
to one outside its domain. The same once-per-call check runs before the value column's dtype gate, so when both are
invalid a constant parameterisation reports the parameter and a parameter column reports the value column.
Cauchy is the only shipped distribution with a moment that has no value, and Pareto the only one with a moment
that diverges on part of its parameter range; every other one has finite moments wherever its parameters are valid.
The policy behind the two answers is in
Design notes.
Related¶
- Handle nulls and errors: what to do about each of these cases.
- Design notes: why raising, not nulling.