Skip to content

Sample reproducibly

Draw random variates as part of a query, with results you can reproduce on another machine.

Draw one variate per row

sample(seed) draws from each row's own distribution:

import polars as pl
import polars_stats as ps

dist = ps.Normal(mu="mu", sigma="sigma")
df = pl.DataFrame({"mu": [0.0, 10.0], "sigma": [1.0, 2.0]})

print(df.with_columns(draw=dist.sample(seed=42)))
shape: (2, 3)
┌──────┬───────┬───────────┐
│ mu   ┆ sigma ┆ draw      │
│ ---  ┆ ---   ┆ ---       │
│ f64  ┆ f64   ┆ f64       │
╞══════╪═══════╪═══════════╡
│ 0.0  ┆ 1.0   ┆ -0.473562 │
│ 10.0 ┆ 2.0   ┆ 13.301167 │
└──────┴───────┴───────────┘

Draw several variates per row

samples(size, seed) returns an Array of width size per row:

print(df.with_columns(draws=dist.samples(3, seed=42)))
shape: (2, 3)
┌──────┬───────┬─────────────────────────────────┐
│ mu   ┆ sigma ┆ draws                           │
│ ---  ┆ ---   ┆ ---                             │
│ f64  ┆ f64   ┆ array[f64, 3]                   │
╞══════╪═══════╪═════════════════════════════════╡
│ 0.0  ┆ 1.0   ┆ [-0.473562, 0.525173, -0.03833… │
│ 10.0 ┆ 2.0   ┆ [13.301167, 10.87918, 8.6149]   │
└──────┴───────┴─────────────────────────────────┘

samples(size=1) matches sample bit for bit for the same seed, and increasing size extends each row's array without changing the draws already there.

size has no maximum, and the whole rows * size result is materialised at once: a million rows at size=10_000 asks for 74.5 GiB. Too large to allocate raises ComputeError; too large for the machine to back is an OS kill, as the equivalent numpy array would be. Draw in batches of rows instead.

Reduce or expand the draws

Stay in the array to summarise per row:

print(
    df.with_columns(draws=dist.samples(4, seed=7)).with_columns(
        mc_mean=pl.col("draws").arr.mean(),
        mc_max=pl.col("draws").arr.max(),
    )
)
shape: (2, 5)
┌──────┬───────┬─────────────────────────────────┬───────────┬───────────┐
│ mu   ┆ sigma ┆ draws                           ┆ mc_mean   ┆ mc_max    │
│ ---  ┆ ---   ┆ ---                             ┆ ---       ┆ ---       │
│ f64  ┆ f64   ┆ array[f64, 4]                   ┆ f64       ┆ f64       │
╞══════╪═══════╪═════════════════════════════════╪═══════════╪═══════════╡
│ 0.0  ┆ 1.0   ┆ [-0.433693, -0.386766, … 0.202… ┆ 0.094614  ┆ 0.996389  │
│ 10.0 ┆ 2.0   ┆ [11.258921, 13.898999, … 9.976… ┆ 11.020545 ┆ 13.898999 │
└──────┴───────┴─────────────────────────────────┴───────────┴───────────┘

Or explode to get one row per draw:

print(df.with_columns(draw=dist.samples(2, seed=7)).explode("draw"))
shape: (4, 3)
┌──────┬───────┬───────────┐
│ mu   ┆ sigma ┆ draw      │
│ ---  ┆ ---   ┆ ---       │
│ f64  ┆ f64   ┆ f64       │
╞══════╪═══════╪═══════════╡
│ 0.0  ┆ 1.0   ┆ -0.433693 │
│ 0.0  ┆ 1.0   ┆ -0.386766 │
│ 10.0 ┆ 2.0   ┆ 11.258921 │
│ 10.0 ┆ 2.0   ┆ 13.898999 │
└──────┴───────┴───────────┘

Make results reproducible

Pass an integer seed. The output is then deterministic across operating systems, architectures, chunk layouts, thread counts, and both the in-memory and streaming engines:

lazy = pl.LazyFrame({"mu": [0.0, 10.0], "sigma": [1.0, 2.0]})
expr = ps.Normal(mu="mu", sigma="sigma").sample(seed=42)

print(lazy.with_columns(draw=expr).collect())
print(lazy.with_columns(draw=expr).collect(engine="streaming"))
shape: (2, 3)
┌──────┬───────┬───────────┐
│ mu   ┆ sigma ┆ draw      │
│ ---  ┆ ---   ┆ ---       │
│ f64  ┆ f64   ┆ f64       │
╞══════╪═══════╪═══════════╡
│ 0.0  ┆ 1.0   ┆ -0.473562 │
│ 10.0 ┆ 2.0   ┆ 13.301167 │
└──────┴───────┴───────────┘
shape: (2, 3)
┌──────┬───────┬───────────┐
│ mu   ┆ sigma ┆ draw      │
│ ---  ┆ ---   ┆ ---       │
│ f64  ┆ f64   ┆ f64       │
╞══════╪═══════╪═══════════╡
│ 0.0  ┆ 1.0   ┆ -0.473562 │
│ 10.0 ┆ 2.0   ┆ 13.301167 │
└──────┴───────┴───────────┘

Omit seed (or pass None) for non-reproducible draws seeded from OS entropy. There is no global seed to set: the seed argument is the only control, which means a sampled column is reproducible from the query alone.

Reproducibility holds across releases too, from 0.0.1 onward: for a given seed, row position and parameters, sample and samples return the same values on every later version. Changing a draw algorithm is therefore a breaking change, and it will be called out as one in that version's release notes rather than shipped alongside an unrelated fix.

Determinism does not depend on evaluation order. Each row derives its own generator from (seed, row_index), so a row's draw depends only on its position in the frame, never on how Polars chunked or threaded the data.