discretization_impact#
- superglm.discretization_impact(
- model: SuperGLM,
- X: FrameLike,
- y: NDArray,
- sample_weight: NDArray | None = None,
- *,
- offset: NDArray | None = None,
- n_bins: int = 100,
- bin_strategy: str = 'exposure_quantile',
- features: list[str] | None = None,
- band_se: float = 1.0,
- band_max_error: float = 0.1,
Analyse the impact of discretizing the smooth terms of a fit.
For each spline/polynomial feature, the smooth per-observation log-relativity is replaced with a family-appropriate bin average. For each continuous-by-continuous interaction it is replaced with the value at the nearest node of the
n_bins-per-axis grid the rating-table export ships – a SAMPLING rather than an averaging, because that block is keyed on axis values and a consumer’s only available lookup is the nearest one. Predictions are recomputed and compared to the originals.Both are covered because both are approximations the exported workbook carries, and reporting one without the other understates how far the table sits from the model (issue #287). The returned
metricsare joint over everything discretized in the call, which is the error a consumer applying the whole table actually gets.Under
weight_semantics="frequency"sample_weightis replication mass: bin geometry, bin averages, mean prediction change, and prediction correlation match literal integer row replication. Zero-frequency rows retain predictions and physicaln_obsentries but cannot change bin geometry or summary metrics. Under"prior"the weights are precisions: they weight deviance and the displayedsample_weighttotals, while bin geometry, bin averages, and pure prediction-comparison summaries use physical rows. A Tweedie fit under"prior"additionally requires them finite and strictly positive.- Parameters:
- modelSuperGLM
A fitted SuperGLM model.
- Xpandas or eager Polars DataFrame
Data used for analysis (typically training data).
- yNDArray
Response variable.
- sample_weightNDArray, optional
Nonnegative observation weights, read under the model’s declared
weight_semantics; a Tweedie fit under"prior"requires them strictly positive. Defaults to ones.- offsetNDArray, optional
Link-scale offset aligned to
X. Used when comparing original and discretized predictions for offset-fitted models.- n_binsint
Number of bins per feature (default 100).
- bin_strategystr
Binning strategy:
"exposure_quantile"(the retained public name) places edges at equal geometry-weight mass;"uniform"uses equal-width bins;"winsorized"uses geometry-weight quantiles on the interior [p5, p95] with dedicated tail bins;"exact"chooses the fewest bands that keep every value’s band average withinmin(band_se * SE, log(1 + band_max_error))of the curve, then the least weighted squared error, withn_binsas the maximum band count. When that is too few, every tolerance is widened by the least factor that fits andband_diagnosticsreports it. Geometry weight means replication mass under the frequency contract and unit physical-row mass for Tweedie.- featureslist[str], optional
Subset of names to discretize: spline/polynomial features, and continuous-by-continuous interaction names as they appear in
model._interaction_order. None means every one of both.- band_sefloat
Under
"exact": the tolerance in pointwise standard errors of the term’s log relativity (native centering). Default 1.0.- band_max_errorfloat
Under
"exact": the largest relative error of a band factor against the curve. Default 0.10.
- Returns:
- DiscretizationResult