Validation And Model Comparison#
Validation should be part of the default pricing workflow, not something added after the model is already chosen.
Recommended Workflow#
split data into training and holdout
compare candidate models with
cross_validate()on training datainspect fold-level deltas, not just one mean metric
refit the chosen candidates on all training data
evaluate holdout Lorenz and double-lift charts
use diagnostics to check residual behavior and calibration
Cross-Validation#
cross_validate() is the main comparison tool. For REML pricing models, call
it with fit_mode="fit_reml" so each fold uses the same fitting story as the
final model.
The fitting examples on this page assume a Poisson rate response
(y = claim_count / exposure) fitted with weight_semantics="frequency", for
which exposure is a replication weight. The contract is declared per model and
the default is "prior", an EDM precision. See
Families & Dispersion.
from sklearn.model_selection import KFold
from superglm.model_selection import cross_validate
result = cross_validate(
model,
X,
y,
cv=KFold(n_splits=5, shuffle=True, random_state=42),
sample_weight=exposure,
fit_mode="fit_reml",
scoring=("deviance", "nll", "gini"),
return_oof=True,
)
Key outputs:
result.fold_scores: per-fold metrics, fit time, convergence, and EDFresult.mean_scores/result.std_scores: summary comparisonsresult.oof_predictions: out-of-fold predictions for every training row
Built-in deviance and NLL averages use the fitted model’s declared likelihood
size: sum(sample_weight) under weight_semantics="frequency" and the count
of positive-weight rows under "prior".
Typical metric interpretation:
Metric |
Measures |
Better |
|---|---|---|
|
probabilistic fit |
lower |
|
negative log-likelihood |
lower |
|
ranking / segmentation power |
higher |
Out-of-fold predictions are especially useful for challenger analysis because they let you compare models on the training portfolio without leaking each row into its own fitted mean.
Holdout Business Evidence#
After cross-validation, refit the chosen candidates on all training data and evaluate them on holdout.
The exposure= argument to the business-validation helpers below is a
portfolio aggregation weight. It does not change or infer the fitted model’s
declared sample_weight contract.
Lorenz Curve And Gini#
from superglm.validation import lorenz_curve
result = lorenz_curve(y_obs, y_pred, exposure=exposure)
print(f"Gini ratio: {result.gini_ratio:.4f}")
The Gini ratio measures ranking power relative to perfect foresight. It is the standard quick view of segmentation quality.
Do not score models on gini and balance alone: a fit with separated interaction cells can move out-of-sample deviance by orders of magnitude while both stay healthy. Keep out-of-sample deviance in every comparison.
Double Lift Chart#
Use the CAS-style double-lift chart when you need business-facing evidence that a challenger is improving on the current model rather than just fitting well in the abstract.
from superglm.validation import double_lift_chart
result = double_lift_chart(
y_obs=y_holdout,
y_pred_model=mu_new,
y_pred_current=mu_baseline,
exposure=exposure_holdout,
n_bins=20,
)
The chart sorts by the challenger/current relativity ratio, bins to equal exposure, and shows whether the challenger tracks Actual more closely than the baseline.
Diagnostic Plots#
Use diagnostics after you have a candidate worth defending.
model.plot_diagnostics(X, y, sample_weight=exposure)
The four-panel diagnostic figure includes:
Q-Q with simulation envelope
sample-weighted calibration
residuals versus linear predictor
residual histogram with normal overlay
Diagnostic weighting follows the declared contract too: a replication weight
repeats a row’s contribution without changing its response distribution, while
a prior weight gives row \(i\) observation-specific dispersion
\(\phi / w_i\). For exact discrete Poisson quantile residuals, diagnose raw
claim counts with log(exposure) as an offset; a rate plus replication weight
cannot reconstruct the corresponding count CDF.
For large datasets, the plotting code automatically switches to more efficient rendering paths such as hexbin density summaries.
Practical Comparison Advice#
use the same CV folds for all candidate models
compare fold-level deltas, not just headline means
keep holdout truly untouched until the challenger set is stable
use both probabilistic metrics and business ranking metrics
treat double-lift as the business communication chart
See the Plotting & Diagnostics Demo for a worked example on French MTPL2.