R Bloggers iconR BloggersSep 14, 2026 ~5 min source read

Model-agnostic prediction intervals in Python and R: does nnetsauce’s QuantileRegressor hold up?

A hands-on look at nnetsauce’s QuantileRegressor: how it converts any sklearn-compatible regressor into quantile predictors, available in Python and R via reticulate, and how it performs in an out-of-the-box benchmark across many estimators and datasets.

Model-agnostic prediction intervals in Python and R: does nnetsauce’s `QuantileRegressor` hold up?

Share this story

Send the public story page.

Useful takeaways from this story.

In a quick example on the diabetes dataset, QuantileRegressor (scoring="residuals", level=95) produced empirical coverage of 94.7% on the test set (target 95%).

The author benchmarked 38 regressors, 6 datasets, two coverage targets, and multiple scoring strategies to evaluate out-of-the-box behavior without tuning base estimators.

# What QuantileRegressor does

# Python and R parity

# Quick reproducible example The article runs a small, runnable quickstart using the diabetes dataset with BayesianRidge as the base regressor. With level=95 and scoring="residuals", the code prints five intervals and computes empirical coverage on the test set. The reported empirical coverage was 94.7% against the 95% target for that run.

# Benchmark design and goals The author deliberately tested the wrapper "out of the box," omitting hyperparameter tuning for base estimators. The grid included:

  • 38 scikit-learn regressors (all_estimators(type_filter='regressor') minus certain meta-estimators)
  • 6 datasets: diabetes, linnerud, two synthetic sets (linear and mildly nonlinear), an anonymized Boston Housing, and a 600-row subsample of California Housing
  • Two coverage targets: 80% and 95%
  • Five QuantileRegressor scoring strategies plus nnetsauce's PredictionInterval (method="splitconformal") as a structurally different baseline
  • Two native quantile baselines: scikit-learn's linear QuantileRegressor and GradientBoostingRegressor(loss="quantile")

That setup produced 2,736 wrapped-estimator fits for the main grid, plus 24 fits for the native baselines.

# What the wrapper changes and what it doesn't QuantileRegressor does not replace or retrain the base regressor to produce quantiles. Instead, it optimizes an offset around the existing model's predictions. This design means the wrapper's behavior depends on the base model's point predictions and residual structure. Because the R binding calls the Python implementation directly, R users get identical behavior and limitations.

# Practical implications for practitioners

  • Expect results to reflect how well the base model's point predictions and residuals capture conditional error structure. No hyperparameter tuning was applied in the benchmark, so results represent typical first attempts.
  • Use the scoring strategy that matches your assumptions: residuals or conformal approaches will produce different intervals.

# Example outputs and verification The quick example prints intervals for individual test rows and computes empirical coverage. For the diabetes example with BayesianRidge and scoring="residuals", empirical coverage was 94.7% with a 95% target. That shows the method can approach target coverage in practice, at least on that dataset and configuration.

# What the benchmark can tell you The benchmark's breadth—many regressors, datasets, and scoring variants—helps reveal how stable QuantileRegressor is across typical off-the-shelf workflows. Because the author did not tune base models, the results illustrate what many practitioners will observe on first use.

# How to follow up Try QuantileRegressor on your own data with the base regressor you intend to use, compare scoring strategies, and consider modest tuning of the base estimator if intervals look overly wide or miscalibrated. If you use R, remember the implementation is the same Python object accessed through reticulate.

More context around this story.

ROC, Paper, Scissor, Shoe
R Bloggers iconR BloggersSep 2, 2026

ROC, Paper, Scissor, Shoe

📊 Working through ROC-AUC from scratch, then poking at its blind spots — low prevalence, calibration, and finally Decision Curve Analysis. Mostly notes to myself on what I learned (and got confused by) along the way. 🤔📈 Motivations We see ROC-AUC so often with classification models, we know the higher the be

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app