discretization_impact#

superglm.discretization_impact(
model: SuperGLM,
X: FrameLike,
y: NDArray,
sample_weight: NDArray | None = None,
*,
offset: NDArray | None = None,
n_bins: int = 100,
bin_strategy: str = 'exposure_quantile',
features: list[str] | None = None,
band_se: float = 1.0,
band_max_error: float = 0.1,
) → DiscretizationResult#

Analyse the impact of discretizing the smooth terms of a fit.

For each spline/polynomial feature, the smooth per-observation log-relativity is replaced with a family-appropriate bin average. For each continuous-by-continuous interaction it is replaced with the value at the nearest node of the n_bins-per-axis grid the rating-table export ships – a SAMPLING rather than an averaging, because that block is keyed on axis values and a consumer’s only available lookup is the nearest one. Predictions are recomputed and compared to the originals.

Both are covered because both are approximations the exported workbook carries, and reporting one without the other understates how far the table sits from the model (issue #287). The returned metrics are joint over everything discretized in the call, which is the error a consumer applying the whole table actually gets.

Under weight_semantics="frequency" sample_weight is replication mass: bin geometry, bin averages, mean prediction change, and prediction correlation match literal integer row replication. Zero-frequency rows retain predictions and physical n_obs entries but cannot change bin geometry or summary metrics. Under "prior" the weights are precisions: they weight deviance and the displayed sample_weight totals, while bin geometry, bin averages, and pure prediction-comparison summaries use physical rows. A Tweedie fit under "prior" additionally requires them finite and strictly positive.

Parameters:
modelSuperGLM

A fitted SuperGLM model.

Xpandas or eager Polars DataFrame

Data used for analysis (typically training data).

yNDArray

Response variable.

sample_weightNDArray, optional

Nonnegative observation weights, read under the model’s declared weight_semantics; a Tweedie fit under "prior" requires them strictly positive. Defaults to ones.

offsetNDArray, optional

Link-scale offset aligned to X. Used when comparing original and discretized predictions for offset-fitted models.

n_binsint

Number of bins per feature (default 100).

bin_strategystr

Binning strategy: "exposure_quantile" (the retained public name) places edges at equal geometry-weight mass; "uniform" uses equal-width bins; "winsorized" uses geometry-weight quantiles on the interior [p5, p95] with dedicated tail bins; "exact" chooses the fewest bands that keep every value’s band average within min(band_se * SE, log(1 + band_max_error)) of the curve, then the least weighted squared error, with n_bins as the maximum band count. When that is too few, every tolerance is widened by the least factor that fits and band_diagnostics reports it. Geometry weight means replication mass under the frequency contract and unit physical-row mass for Tweedie.

featureslist[str], optional

Subset of names to discretize: spline/polynomial features, and continuous-by-continuous interaction names as they appear in model._interaction_order. None means every one of both.

band_sefloat

Under "exact": the tolerance in pointwise standard errors of the term’s log relativity (native centering). Default 1.0.

band_max_errorfloat

Under "exact": the largest relative error of a band factor against the curve. Default 0.10.

Returns:
DiscretizationResult