Discrete distributions¶
Every discrete distribution inherits the shared method surface from DiscreteDistribution (shown first); the concrete
classes follow, alphabetically, and list only what they define or override.
Base class¶
DiscreteDistribution
¶
Abstract base class for discrete univariate distributions.
pmf
¶
pmf(value: float | IntoExprColumn) -> Expr
Probability mass function, P(X = value). Nulls and NaNs in value are propagated.
log_pmf
¶
log_pmf(value: float | IntoExprColumn) -> Expr
Natural logarithm of the pmf. Nulls and NaNs in value are propagated.
sample
¶
sample(seed: int | None = None) -> Expr
Draw one random variate per row.
Output length follows the surrounding context (frame length under select / with_columns,
partition length under over / group_by). Each row's draw derives from a sub-seed mixed from
seed and the row's position, so the result is independent of chunking and thread scheduling.
A row with an invalid parameter raises; a row with a null parameter yields null. The output is
named "sample" when every parameter is constant; otherwise it follows the first parameter
expression (polars root-name semantics, so .name.* modifiers keep working).
Source code in polars_stats/distributions/_base.py
samples
¶
Draw size random variates per row, as Array(inner=<element dtype>, shape=size).
Each row's draws are consecutive values from one per-row stream keyed by seed and the row's
position, so samples(size=1) matches sample for the same seed and growing size extends
each row's array without changing the existing draws. A null-parameter row yields a null array
(outer validity), an invalid parameterisation raises. Named like sample, as "samples".
Source code in polars_stats/distributions/_base.py
cdf
¶
cdf(value: float | IntoExprColumn) -> Expr
Cumulative distribution function, P(X <= value). Nulls and NaNs in value are propagated.
log_cdf
¶
log_cdf(value: float | IntoExprColumn) -> Expr
Natural logarithm of the cdf. Nulls and NaNs in value are propagated.
sf
¶
sf(value: float | IntoExprColumn) -> Expr
Survival function, P(X > value) = 1 - cdf(value). Nulls and NaNs in value are propagated.
log_sf
¶
log_sf(value: float | IntoExprColumn) -> Expr
Natural logarithm of the survival function. Nulls and NaNs in value are propagated.
ppf
¶
ppf(quantile: float | IntoExprColumn) -> Expr
Percent point function (inverse cdf).
A quantile outside [0, 1] yields null. Nulls are propagated and a NaN quantile yields
NaN, matching scipy.
Source code in polars_stats/distributions/_base.py
isf
¶
isf(quantile: float | IntoExprColumn) -> Expr
Inverse survival function, the value x with sf(x) == quantile.
Same domain contract as ppf, with the endpoints reversed: quantile outside [0, 1] yields
null, nulls propagate, NaN yields NaN.
Source code in polars_stats/distributions/_base.py
mean
abstractmethod
¶
variance
abstractmethod
¶
std
¶
median
¶
Distributions¶
Bernoulli
¶
Bernoulli(p: float | IntoExprColumn)
Bases: DiscreteDistribution
Bernoulli distribution with success probability p.
Equivalent to scipy.stats.bernoulli(p).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
p
|
float | IntoExprColumn
|
Success probability, with |
required |
An invalid p (p < 0, p > 1 or NaN) is not checked at construction; it raises
InvalidOperation (a ComputeError) when any method is evaluated. A null p nulls every method,
on the support and off it.
Source code in polars_stats/distributions/_bernoulli.py
mean
¶
variance
¶
entropy
¶
Shannon entropy, -p * log(p) - (1 - p) * log1p(-p), with 0 * log 0 = 0 at p in {0, 1}.
log1p(-p) rather than log(1 - p), which collapses to 0.0 below p ~ 1.1e-16.
Source code in polars_stats/distributions/_bernoulli.py
Binomial
¶
Bases: DiscreteDistribution
Binomial distribution: number of successes in n trials, each with success probability p.
Equivalent to scipy.stats.binom(n, p). The argument order differs from statrs
(Binomial(p, n)); this class follows scipy's (n, p).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n
|
int | IntoExprColumn
|
Number of trials, an integer |
required |
p
|
float | IntoExprColumn
|
Success probability in |
required |
Parameters are validated at evaluation, where a negative n column, a non-integer n dtype or a
p outside [0, 1] raises InvalidOperation (a ComputeError). A scalar n is the one
exception: it is coerced to a UInt64 literal and passed to the fast paths as a kwarg, neither of
which can carry an out-of-range count, so it raises ValueError at construction. Construction
otherwise rejects only wrong types (TypeError). Null parameters propagate to null, a
Null-dtype n column included; the dtype rule is judged first, so a float n column
raises even when every value in it is null.
Source code in polars_stats/distributions/_binomial.py
DiscreteUniform
¶
Bases: DiscreteDistribution
Discrete uniform distribution over the integers {min, ..., max}, both bounds inclusive.
Equivalent to scipy.stats.randint(low=min, high=max + 1). The max argument is
inclusive, unlike scipy's exclusive high: the support is {min, ..., max} and
cdf(max) == 1.
Two further divergences from scipy: median is the midpoint (min + max) / 2, not the
support point scipy's ppf(0.5) reports, and ppf(0) / isf(1) clamp to the support
rather than answering its below-support sentinel min - 1.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
min
|
int | IntoExprColumn
|
Inclusive lower bound, an integer anywhere in |
required |
max
|
int | IntoExprColumn
|
Inclusive upper bound, with |
required |
An invalid parameterisation (max < min, or a width max - min + 1 overflowing Int64)
is not checked at construction; it raises InvalidOperation (a ComputeError) when a method is
evaluated. Null bounds propagate to null.
Source code in polars_stats/distributions/_discrete_uniform.py
support_size
property
¶
Support count N = max - min + 1, as Float64, validated in Rust; rounded above 2**53.
mean
¶
variance
¶
median
¶
Median, the midpoint (min + max) / 2, which for an even support size is not a support point.
Diverges from scipy, which reports that support point: scipy.stats.randint(low=1, high=7).median()
is 3.0 against this library's 3.5.
Source code in polars_stats/distributions/_discrete_uniform.py
Geometric
¶
Geometric(p: float | IntoExprColumn)
Bases: DiscreteDistribution
Geometric distribution: the number of trials up to and including the first success.
Equivalent to scipy.stats.geom(p). The support is the positive integers: a draw of k
means trials 1 .. k - 1 failed and trial k succeeded, so pmf(1) = p and the mass
decays geometrically above it. Textbooks that count failures before the first success start the
support at 0 instead; that is a different parameterisation and not what this class computes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
p
|
float | IntoExprColumn
|
Success probability of each trial, with |
required |
An invalid p (p <= 0, p > 1 or NaN) is not checked at construction; it raises
InvalidOperation (a ComputeError) when any method is evaluated. A null p nulls every method,
on the support and off it. Samples are UInt64 trial counts, so unlike Bernoulli the degenerate
p = 0 point mass is not representable.