Skip to content

Discrete distributions

Every discrete distribution inherits the shared method surface from DiscreteDistribution (shown first); the concrete classes follow, alphabetically, and list only what they define or override.

Base class

DiscreteDistribution

Abstract base class for discrete univariate distributions.

pmf

pmf(value: float | IntoExprColumn) -> Expr

Probability mass function, P(X = value). Nulls and NaNs in value are propagated.

Source code in polars_stats/distributions/_base.py
def pmf(self, value: float | IntoExprColumn) -> pl.Expr:
    """Probability mass function, `P(X = value)`. Nulls and NaNs in `value` are propagated."""
    return self._pmf(as_expr(value))

log_pmf

log_pmf(value: float | IntoExprColumn) -> Expr

Natural logarithm of the pmf. Nulls and NaNs in value are propagated.

Source code in polars_stats/distributions/_base.py
def log_pmf(self, value: float | IntoExprColumn) -> pl.Expr:
    """Natural logarithm of the pmf. Nulls and NaNs in `value` are propagated."""
    return self._log_pmf(as_expr(value))

sample

sample(seed: int | None = None) -> Expr

Draw one random variate per row.

Output length follows the surrounding context (frame length under select / with_columns, partition length under over / group_by). Each row's draw derives from a sub-seed mixed from seed and the row's position, so the result is independent of chunking and thread scheduling.

A row with an invalid parameter raises; a row with a null parameter yields null. The output is named "sample" when every parameter is constant; otherwise it follows the first parameter expression (polars root-name semantics, so .name.* modifiers keep working).

Source code in polars_stats/distributions/_base.py
def sample(self, seed: int | None = None) -> pl.Expr:
    """Draw one random variate per row.

    Output length follows the surrounding context (frame length under `select` / `with_columns`,
    partition length under `over` / `group_by`). Each row's draw derives from a sub-seed mixed from
    `seed` and the row's position, so the result is independent of chunking and thread scheduling.

    A row with an invalid parameter raises; a row with a null parameter yields null. The output is
    named `"sample"` when every parameter is constant; otherwise it follows the first parameter
    expression (polars root-name semantics, so `.name.*` modifiers keep working).
    """
    return self._draw("sample", seed=_checked_seed(seed))

samples

samples(size: int, seed: int | None = None) -> Expr

Draw size random variates per row, as Array(inner=<element dtype>, shape=size).

Each row's draws are consecutive values from one per-row stream keyed by seed and the row's position, so samples(size=1) matches sample for the same seed and growing size extends each row's array without changing the existing draws. A null-parameter row yields a null array (outer validity), an invalid parameterisation raises. Named like sample, as "samples".

Source code in polars_stats/distributions/_base.py
def samples(self, size: int, seed: int | None = None) -> pl.Expr:
    """Draw `size` random variates per row, as `Array(inner=<element dtype>, shape=size)`.

    Each row's draws are consecutive values from one per-row stream keyed by `seed` and the row's
    position, so `samples(size=1)` matches `sample` for the same seed and growing `size` extends
    each row's array without changing the existing draws. A null-parameter row yields a null array
    (outer validity), an invalid parameterisation raises. Named like `sample`, as `"samples"`.
    """
    return self._draw("samples", size=_checked_size(size), seed=_checked_seed(seed))

cdf

cdf(value: float | IntoExprColumn) -> Expr

Cumulative distribution function, P(X <= value). Nulls and NaNs in value are propagated.

Source code in polars_stats/distributions/_base.py
def cdf(self, value: float | IntoExprColumn) -> pl.Expr:
    """Cumulative distribution function, `P(X <= value)`. Nulls and NaNs in `value` are propagated."""
    return self._cdf(as_expr(value))

log_cdf

log_cdf(value: float | IntoExprColumn) -> Expr

Natural logarithm of the cdf. Nulls and NaNs in value are propagated.

Source code in polars_stats/distributions/_base.py
def log_cdf(self, value: float | IntoExprColumn) -> pl.Expr:
    """Natural logarithm of the cdf. Nulls and NaNs in `value` are propagated."""
    return self._log_cdf(as_expr(value))

sf

sf(value: float | IntoExprColumn) -> Expr

Survival function, P(X > value) = 1 - cdf(value). Nulls and NaNs in value are propagated.

Source code in polars_stats/distributions/_base.py
def sf(self, value: float | IntoExprColumn) -> pl.Expr:
    """Survival function, `P(X > value) = 1 - cdf(value)`. Nulls and NaNs in `value` are propagated."""
    return self._sf(as_expr(value))

log_sf

log_sf(value: float | IntoExprColumn) -> Expr

Natural logarithm of the survival function. Nulls and NaNs in value are propagated.

Source code in polars_stats/distributions/_base.py
def log_sf(self, value: float | IntoExprColumn) -> pl.Expr:
    """Natural logarithm of the survival function. Nulls and NaNs in `value` are propagated."""
    return self._log_sf(as_expr(value))

ppf

ppf(quantile: float | IntoExprColumn) -> Expr

Percent point function (inverse cdf).

A quantile outside [0, 1] yields null. Nulls are propagated and a NaN quantile yields NaN, matching scipy.

Source code in polars_stats/distributions/_base.py
def ppf(self, quantile: float | IntoExprColumn) -> pl.Expr:
    """Percent point function (inverse cdf).

    A `quantile` outside `[0, 1]` yields **null**. Nulls are propagated and a `NaN` quantile yields
    `NaN`, matching scipy.
    """
    return self._ppf(as_expr(quantile))

isf

isf(quantile: float | IntoExprColumn) -> Expr

Inverse survival function, the value x with sf(x) == quantile.

Same domain contract as ppf, with the endpoints reversed: quantile outside [0, 1] yields null, nulls propagate, NaN yields NaN.

Source code in polars_stats/distributions/_base.py
def isf(self, quantile: float | IntoExprColumn) -> pl.Expr:
    """Inverse survival function, the value `x` with `sf(x) == quantile`.

    Same domain contract as `ppf`, with the endpoints reversed: `quantile` outside `[0, 1]` yields
    null, nulls propagate, `NaN` yields `NaN`.
    """
    return self._isf(as_expr(quantile))

mean abstractmethod

mean() -> Expr

Expected value E[X].

Source code in polars_stats/distributions/_base.py
@abstractmethod
def mean(self) -> pl.Expr:
    """Expected value `E[X]`."""

variance abstractmethod

variance() -> Expr

Variance Var[X] = E[(X - E[X])^2].

Source code in polars_stats/distributions/_base.py
@abstractmethod
def variance(self) -> pl.Expr:
    """Variance `Var[X] = E[(X - E[X])^2]`."""

std

std() -> Expr

Standard deviation, sqrt(variance).

Source code in polars_stats/distributions/_base.py
def std(self) -> pl.Expr:
    """Standard deviation, `sqrt(variance)`."""
    return self.variance().sqrt()

median

median() -> Expr

Median, ppf(0.5).

Source code in polars_stats/distributions/_base.py
def median(self) -> pl.Expr:
    """Median, `ppf(0.5)`."""
    return self._ppf(as_expr(0.5))

entropy abstractmethod

entropy() -> Expr

Differential or Shannon entropy, in nats.

Source code in polars_stats/distributions/_base.py
@abstractmethod
def entropy(self) -> pl.Expr:
    """Differential or Shannon entropy, in nats."""

Distributions

Bernoulli

Bernoulli(p: float | IntoExprColumn)

Bases: DiscreteDistribution

Bernoulli distribution with success probability p.

Equivalent to scipy.stats.bernoulli(p).

Parameters:

Name Type Description Default
p float | IntoExprColumn

Success probability, with 0 <= p <= 1. Either a Python float or an IntoExprColumn (pl.Expr, pl.Series or column name str) carrying one probability per row.

required

An invalid p (p < 0, p > 1 or NaN) is not checked at construction; it raises InvalidOperation (a ComputeError) when any method is evaluated. A null p nulls every method, on the support and off it.

Source code in polars_stats/distributions/_bernoulli.py
def __init__(self, p: float | IntoExprColumn) -> None:
    self._p = coerce_param(p, name="p")
    self._scalar_kwargs = scalar_kwargs(p=scalar_float(p))

mean

mean() -> Expr

Expected value, p.

Source code in polars_stats/distributions/_bernoulli.py
def mean(self) -> pl.Expr:
    """Expected value, ``p``."""
    return self._validated_params

variance

variance() -> Expr

Variance, p * (1 - p).

Source code in polars_stats/distributions/_bernoulli.py
def variance(self) -> pl.Expr:
    """Variance, ``p * (1 - p)``."""
    p = self._raw_p
    return self._moment(p * (1 - p))

entropy

entropy() -> Expr

Shannon entropy, -p * log(p) - (1 - p) * log1p(-p), with 0 * log 0 = 0 at p in {0, 1}.

log1p(-p) rather than log(1 - p), which collapses to 0.0 below p ~ 1.1e-16.

Source code in polars_stats/distributions/_bernoulli.py
def entropy(self) -> pl.Expr:
    """Shannon entropy, ``-p * log(p) - (1 - p) * log1p(-p)``, with ``0 * log 0 = 0`` at ``p in {0, 1}``.

    ``log1p(-p)`` rather than ``log(1 - p)``, which collapses to ``0.0`` below ``p ~ 1.1e-16``.
    """
    p = self._raw_p
    shannon = pl.when((p == 0) | (p == 1)).then(0.0).otherwise(-p * p.log() - (1 - p) * (-p).log1p())
    return self._moment(shannon)

Binomial

Binomial(n: int | IntoExprColumn, p: float | IntoExprColumn)

Bases: DiscreteDistribution

Binomial distribution: number of successes in n trials, each with success probability p.

Equivalent to scipy.stats.binom(n, p). The argument order differs from statrs (Binomial(p, n)); this class follows scipy's (n, p).

Parameters:

Name Type Description Default
n int | IntoExprColumn

Number of trials, an integer >= 0. Either a Python int (at most 2**63 - 1) or an IntoExprColumn (pl.Expr, pl.Series or column name str) carrying one count per row. A column must have an integer dtype, of any width up to UInt64, and may hold any count the dtype can; cast a float one yourself (pl.col("n").cast(pl.Int64)).

required
p float | IntoExprColumn

Success probability in [0, 1]. Either a Python float or an IntoExprColumn.

required

Parameters are validated at evaluation, where a negative n column, a non-integer n dtype or a p outside [0, 1] raises InvalidOperation (a ComputeError). A scalar n is the one exception: it is coerced to a UInt64 literal and passed to the fast paths as a kwarg, neither of which can carry an out-of-range count, so it raises ValueError at construction. Construction otherwise rejects only wrong types (TypeError). Null parameters propagate to null, a Null-dtype n column included; the dtype rule is judged first, so a float n column raises even when every value in it is null.

Source code in polars_stats/distributions/_binomial.py
def __init__(self, n: int | IntoExprColumn, p: float | IntoExprColumn) -> None:
    self._n = coerce_n(n, name="n")
    self._p = coerce_param(p, name="p")
    self._scalar_kwargs = scalar_kwargs(n=scalar_int(n), p=scalar_float(p))

mean

mean() -> Expr

Expected value, n * p.

Source code in polars_stats/distributions/_binomial.py
def mean(self) -> pl.Expr:
    """Expected value, ``n * p``."""
    return self._moment(self._n * self._p)

variance

variance() -> Expr

Variance, n * p * (1 - p).

Source code in polars_stats/distributions/_binomial.py
def variance(self) -> pl.Expr:
    """Variance, ``n * p * (1 - p)``."""
    return self._moment(self._n * self._p * (1 - self._p))

entropy

entropy() -> Expr

Shannon entropy in nats, the exact support sum -sum_k pmf(k) log pmf(k).

0 at the degenerate endpoints p in {0, 1}. There is no closed form, so binomial_entropy evaluates the sum in Rust.

Source code in polars_stats/distributions/_binomial.py
def entropy(self) -> pl.Expr:
    """Shannon entropy in nats, the exact support sum ``-sum_k pmf(k) log pmf(k)``.

    ``0`` at the degenerate endpoints ``p in {0, 1}``. There is no closed form, so ``binomial_entropy``
    evaluates the sum in Rust.
    """
    return self._param_plugin("entropy")

DiscreteUniform

DiscreteUniform(min: int | IntoExprColumn, max: int | IntoExprColumn)

Bases: DiscreteDistribution

Discrete uniform distribution over the integers {min, ..., max}, both bounds inclusive.

Equivalent to scipy.stats.randint(low=min, high=max + 1). The max argument is inclusive, unlike scipy's exclusive high: the support is {min, ..., max} and cdf(max) == 1.

Two further divergences from scipy: median is the midpoint (min + max) / 2, not the support point scipy's ppf(0.5) reports, and ppf(0) / isf(1) clamp to the support rather than answering its below-support sentinel min - 1.

Parameters:

Name Type Description Default
min int | IntoExprColumn

Inclusive lower bound, an integer anywhere in Int64: [-2**63, 2**63 - 1]. Either a Python int (rejected at construction outside that range) or an IntoExprColumn (pl.Expr, pl.Series or column name str) carrying one bound per row; a column may be any integer dtype, judged by its values fitting Int64.

required
max int | IntoExprColumn

Inclusive upper bound, with max >= min (min == max is a one-point mass) and the width max - min + 1 fitting Int64. Same accepted types and range as min.

required

An invalid parameterisation (max < min, or a width max - min + 1 overflowing Int64) is not checked at construction; it raises InvalidOperation (a ComputeError) when a method is evaluated. Null bounds propagate to null.

Source code in polars_stats/distributions/_discrete_uniform.py
def __init__(self, min: int | IntoExprColumn, max: int | IntoExprColumn) -> None:  # noqa: A002
    self._min = coerce_int(min, name="min")
    self._max = coerce_int(max, name="max")
    self._scalar_kwargs = scalar_kwargs(min=scalar_int(min), max=scalar_int(max))

support_size property

support_size: Expr

Support count N = max - min + 1, as Float64, validated in Rust; rounded above 2**53.

mean

mean() -> Expr

Expected value, (min + max) / 2.

Source code in polars_stats/distributions/_discrete_uniform.py
def mean(self) -> pl.Expr:
    """Expected value, ``(min + max) / 2``."""
    return self._moment(self._midpoint)

variance

variance() -> Expr

Variance, (N**2 - 1) / 12.

Source code in polars_stats/distributions/_discrete_uniform.py
def variance(self) -> pl.Expr:
    """Variance, ``(N**2 - 1) / 12``."""
    return (self.support_size**2 - 1) * ONE_TWELFTH

median

median() -> Expr

Median, the midpoint (min + max) / 2, which for an even support size is not a support point.

Diverges from scipy, which reports that support point: scipy.stats.randint(low=1, high=7).median() is 3.0 against this library's 3.5.

Source code in polars_stats/distributions/_discrete_uniform.py
def median(self) -> pl.Expr:
    """Median, the midpoint ``(min + max) / 2``, which for an even support size is not a support point.

    **Diverges from scipy**, which reports that support point: ``scipy.stats.randint(low=1, high=7).median()``
    is ``3.0`` against this library's ``3.5``.
    """
    return self._moment(self._midpoint)

entropy

entropy() -> Expr

Shannon entropy in nats, log(N).

Source code in polars_stats/distributions/_discrete_uniform.py
def entropy(self) -> pl.Expr:
    """Shannon entropy in nats, ``log(N)``."""
    return self.support_size.log()

Geometric

Geometric(p: float | IntoExprColumn)

Bases: DiscreteDistribution

Geometric distribution: the number of trials up to and including the first success.

Equivalent to scipy.stats.geom(p). The support is the positive integers: a draw of k means trials 1 .. k - 1 failed and trial k succeeded, so pmf(1) = p and the mass decays geometrically above it. Textbooks that count failures before the first success start the support at 0 instead; that is a different parameterisation and not what this class computes.

Parameters:

Name Type Description Default
p float | IntoExprColumn

Success probability of each trial, with 0 < p <= 1. Either a Python float or an IntoExprColumn (pl.Expr, pl.Series or column name str) carrying one probability per row.

required

An invalid p (p <= 0, p > 1 or NaN) is not checked at construction; it raises InvalidOperation (a ComputeError) when any method is evaluated. A null p nulls every method, on the support and off it. Samples are UInt64 trial counts, so unlike Bernoulli the degenerate p = 0 point mass is not representable.

Source code in polars_stats/distributions/_geometric.py
def __init__(self, p: float | IntoExprColumn) -> None:
    self._p = coerce_param(p, name="p")
    self._scalar_kwargs = scalar_kwargs(p=scalar_float(p))

mean

mean() -> Expr

Expected value, 1 / p.

Source code in polars_stats/distributions/_geometric.py
def mean(self) -> pl.Expr:
    """Expected value, ``1 / p``."""
    return 1 / self._validated_params

variance

variance() -> Expr

Variance, (1 - p) / p**2.

Source code in polars_stats/distributions/_geometric.py
def variance(self) -> pl.Expr:
    """Variance, ``(1 - p) / p**2``."""
    p = self._raw_p
    return self._moment((1 - p) / p**2)

std

std() -> Expr

Standard deviation, sqrt(1 - p) / p.

Not variance().sqrt(): squaring and unsquaring p overflows about 150 decades earlier.

Source code in polars_stats/distributions/_geometric.py
def std(self) -> pl.Expr:
    """Standard deviation, ``sqrt(1 - p) / p``.

    Not ``variance().sqrt()``: squaring and unsquaring ``p`` overflows about 150 decades earlier.
    """
    p = self._raw_p
    return self._moment((1 - p).sqrt() / p)

entropy

entropy() -> Expr

Shannon entropy, (-(1 - p) * log1p(-p) - p * log(p)) / p, with 0 * log 0 = 0 at p = 1.

Source code in polars_stats/distributions/_geometric.py
def entropy(self) -> pl.Expr:
    """Shannon entropy, ``(-(1 - p) * log1p(-p) - p * log(p)) / p``, with ``0 * log 0 = 0`` at ``p = 1``."""
    p = self._raw_p
    q = 1 - p
    shannon = pl.when(p == 1).then(0.0).otherwise((-q * (-p).log1p() - p * p.log()) / p)
    return self._moment(shannon)