Skip to content
tbosierPublic

About

Bayesian models in Python. Inference in Rust.

Topics

Resources

Contributing

Stars

14 stars

Watchers

0 watching

Forks

Repository files navigation

rustmc

Bayesian models in Python. Inference in Rust.

rustmc focuses on small, structured models you need to fit repeatedly: regressions, group comparisons, calibration, and forecasts. Build a model once, fit new datasets, and keep the posterior draws for prediction and diagnostics.

The project is alpha. The Python package is supported; the Rust API is still changing. Check convergence and model fit on your own data.

pip install rustmc

NumPy is the only required Python dependency. Install rustmc[viz] for ArviZ and Matplotlib. Python 3.9–3.14 are covered by install tests.

Fit a regression

This example estimates an instrument's offset, gain, and measurement noise.

import numpy as np
import rustmc as rmc

rng = np.random.default_rng(42)
x = np.linspace(-2, 2, 100)
y = 0.3 + 1.2 * x + rng.normal(0, 0.2, x.size)

model = rmc.ModelBuilder()
offset = model.normal_prior("offset", 0.0, 1.0)
gain = model.normal_prior("gain", 1.0, 0.5)
noise = model.half_normal_prior("noise", 0.5)
model.normal_likelihood("reading", offset + gain * "x", noise, "y")
compiled = model.compile()

fit = compiled.sample(
    {"x": x, "y": y}, chains=4, warmup=1000, draws=1000, seed=42,
    show_progress=False,
)
print(fit.summary())

future = fit.predict({"x": np.array([-1.0, 0.0, 1.0])}, seed=43)
print(np.quantile(future["reading"], [0.025, 0.975], axis=(0, 1)))

predict keeps the (chain, draw, observation) axes. Use expected=True for the conditional mean without new observation noise. Priors above are chosen for this example's units.

Reuse the model

compiled.sample() accepts another dataset with the same columns and a different number of rows. compiled.sample_batch() fits independent datasets with stable IDs:

batch = compiled.sample_batch(
    [{"x": x, "y": y}, {"x": x[:50], "y": y[:50]}],
    ids=["instrument-a", "instrument-b"],
    chains=4, warmup=1000, draws=1000, threads=2, errors="collect",
    show_progress=False,
)
for instrument in batch.ids:
    if instrument not in batch.errors:
        print(instrument, batch.get(instrument).summary())
print(batch.errors)

Independent fits do not share information. For related groups, build one partial-pooling model.

What's included

  • NUTS and HMC with autodiff, constrained parameters, and parallel chains.
  • Scalar and vector regressions, group indexing, nonlinear expressions, and custom log-density terms. See custom models.
  • Prior and posterior prediction, pointwise log likelihood, R-hat, effective sample size, Monte Carlo error, and ArviZ export.
  • Exact Gaussian AR regression, Gaussian hierarchical models, and Kalman/FFBS algorithms for state-space models.
  • Forecasting workflows for structural, count, hurdle, and runoff models, with joint predictive paths and backtests.
  • Versioned model and fit artifacts. Compiled model artifacts omit training data; fitted artifacts include it. Neither resumes sampler adaptation or RNG state.

Two engines, one result surface

Models you write with ModelBuilder compile to a differentiable graph and are fitted by NUTS or HMC. The forecasting models are different: structural, seasonal, AR, dynamic GLM, hurdle, runoff, and the Gaussian hierarchy are hand-written samplers that do not use that graph, its autodiff, or its samplers. Each exploits structure the general sampler cannot, and they are not all the same kind: Gibbs with FFBS for the Gaussian state-space models, exact independent conjugate draws for AR, block elliptical slice sampling for dynamic GLMs, and — for runoff — either exact conjugate draws or latent-count Gibbs, depending on whether every ultimate total is known. sampler_stats on a fit reports which one ran, and whether warmup applied.

What they share is narrower than "one engine" suggests. diagnostics is genuinely common: R-hat, ESS and MCSE are computed by the same code for every fit. The batch executor is shared inside Rust, not just at the Python edge. state_space and forecast_diagnostics are shared among the forecasting models but are not used by the graph sampler at all. What every model does share is the Python surface: summary() and diagnostics() mean the same thing wherever you find them. A change to the NUTS sampler does not change a forecast, and vice versa.

The modeling language is deliberately small. PyMC and Stan offer broader model support. rustmc aims to earn its place through repeated fitting and a few well-tested specialized algorithms. Performance depends on the workload; see the benchmark protocol.

Start here

The roadmap tracks five priorities: statistical release gates, representative benchmarks, native model artifacts, bounded batches, and consistent results and diagnostics. Most of that work sits in the shared layer, so it reaches the forecasting models and the graph models alike.

For source builds and checks, see Contributing. MIT licensed.

About

Bayesian models in Python. Inference in Rust.

Topics

Resources

Contributing

Stars

14 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages