Towards Data Science iconTowards Data ScienceSep 25, 2026 ~10 min source read

Autoregressive Rollout and Uncertainty Propagation: How to Turn One-Step Probabilities into Honest Multi-Step Forecasts

Part II explains why repeating a one-step probabilistic predictor deterministically destroys uncertainty, and shows how to assemble a joint forecast over a horizon by sampling the model’s conditional distributions step by step.

Share this story

Send the public story page.

Useful takeaways from this story.

A model trained with Gaussian NLL learns a correct per-step conditional mean and variance, but that alone does not produce a correct joint forecast over multiple steps.

Naively feeding the model’s predicted mean forward (deterministic rollout) discards aleatoric uncertainty and produces overconfident sequence forecasts.

The useful part

If your model is probabilistic, you want that forecast to carry honest uncertainty all the way to the end of the horizon. Fixing this requires thinking carefully about what a joint distribution over future values actually is, and how to approximate it with Monte Carlo sampling. The fix is barely more code than the naive version, but it gives a fundamentally different and fundamentally better forecast.

How it works

  • Everything here is built and checked in the companion notebook: the two rollout approaches, applied to the same trained model, on the same signal, with coverage and PIT diagnostics (all explained in this post).
  • There are two strategies that look similar in code but are conceptually and statistically very different.
  • Deterministic rollout feeds only μ \mu μ forward, so later σ \sigma σ predictions assume the previous value was known exactly.
  • The deterministic rollout: feed μ \mu μ back The simplest thing you could do: at each step, predict (μ, σ) (\mu, \sigma) (μ, σ), take only the μ \mu μ, plug it into the context as if it were an observed...
  • By feeding μ \mu μ back, you're pretending the previous prediction was perfect.

What to take from it

The problem starts the moment someone asks: okay, but what happens over the next 50 steps? ··· The multi-step inference problem In any real application, you don't just need the next value. The future values are correlated with each other because each one feeds into the next.

Example or evidence

  • So in principle, you can build the joint distribution one step at a time: predict step T + 1 T+1 T + 1 given the observed history, then predict step T + 2 T+2 T + 2 given the observed history plus step T +...
  • It's one distribution over H H H -dimensional space, a distribution over entire sequences.
  • The value at step T + 3 T+3 T + 3 depends on what happened at steps T + 1 T+1 T + 1 and T + 2 T+2 T + 2.
  • The σ \sigma σ values you get capture only the local noise at each step in isolation, the irreducible randomness if the input were known exactly.

Details worth keeping

At inference time, in any real application, this is rarely enough. You want to know what happens over the next 10, 100, or 1000 steps. You want a forecast, not a single prediction.

Related coverage

  • Towards Data Science: First in a series on probabilistic forecasting for physical signals. Next: what happens when you roll the forecast forward more than one step.
  • Plainenglish: A model that scores 94 percent on your eval set and then falls apart on real traffic does not have a model problem. It has an eval set… Continue reading on Python in Plain English »
  • Columbia: It's time for another one of Bob's favorite sort of post, which is when I use simulation to demonstrate a statistical principle.
  • Medium: Four gaps that sit between an offline number and what people actually experience — and the one test that exposes all of them. Continue reading on Medium »
  • Medium: You. Continue reading on Data Science Collective »

More context around this story.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app