ModelMetrics#
- class superglm.ModelMetrics(
- model: SuperGLM,
- X=None,
- y=None,
- sample_weight=None,
- offset=None,
- *,
- _fit_data_matches: bool | None = None,
- _mu: NDArray | None = None,
- _null_mu: NDArray | None = None,
- _fit_stats=None,
Bases:
objectPost-fit diagnostics for a SuperGLM model.
- Parameters:
- modelSuperGLM
A fitted model.
- Xpandas or eager Polars DataFrame
Feature matrix used for fitting (or evaluation).
- yarray-like
Response variable.
- sample_weightarray-like, optional
Observation weights / sample_weight.
- offsetarray-like, optional
Offset term.
Notes
On the frame a model was fitted on, leverage and effective degrees of freedom come from the fit itself. Equal copies of that frame, or a pickled model, are recognised as the training rows only when the model keeps a structured or shape-constrained covariance. Otherwise the metrics are re-formed on the rows passed, which for a discrete fit can differ slightly from its binned fit.
- property eta: NDArray#
Linear predictor (link-scale fitted values).
- residuals( ) NDArray#
Compute residuals of the specified type.
Deviance and Pearson residuals use a compressed-row aggregate representation: a weighted row is multiplied by
sqrt(sample_weight)so its squared residual equals the sum of the squared residuals from literal replicated rows. Response, working, and quantile residuals retain per-row semantics and are not multiplied by the weight. Underweight_semantics="prior"the weight is part of the row’s own distribution and therefore stays in every diagnostic that depends on it. Influence diagnostics combine the aggregate residual with compressed leverage and therefore describe deletion of the whole weighted cell, not deletion of one literal expanded copy.- Parameters:
- kindstr
One of “deviance”, “pearson”, “response”, “working”, “quantile”.
- seedint or None
Random seed for quantile residuals (discrete families only). Default 42 for reproducibility. Ignored for non-quantile types.
- property leverage: NDArray#
Compressed-row hat diagonal under the supplied evaluation weights.
This is the influence of one weighted cell as represented in the compressed design, not the leverage of each literal expanded copy.
sum(h)is approximately effective_df - 1 (excluding the intercept).
- property cooks_distance: NDArray#
Compressed-row aggregate Cook’s distance.
Both the Pearson residual contribution and leverage correspond to deletion of the supplied weighted cell, rather than deletion of one literal expanded copy.
- property std_deviance_residuals: NDArray#
Standardized compressed-row deviance residuals.
Equal to
r_dev / sqrt(phi * (1 - h))using aggregate weighted residual and compressed-row leverage, so it represents deletion of the whole weighted cell rather than one literal expanded copy.
- property std_pearson_residuals: NDArray#
Standardized compressed-row Pearson residuals.
Equal to
r_pear / sqrt(phi * (1 - h))using aggregate weighted residual and compressed-row leverage, so it represents deletion of the whole weighted cell rather than one literal expanded copy.
- property coefficient_se: dict[str, NDArray]#
Per-group quasi-likelihood coefficient standard errors.
For known-scale families such as Poisson and Binomial, the covariance is scaled by the Pearson dispersion estimate
pearson_chi2 / residual_df, whereresidual_dffollows the family’s sample-weight semantics. For estimated-scale families such as Gaussian, Gamma, and Tweedie, it uses the fitted dispersionphi. Usecoefficient_se_rawfor the unscaled, exact-family covariance.This accessor is an explicit quasi-likelihood diagnostic. Rendered coefficient tables retain the fitted-family scale (and hence
phi=1for known-scale families) unless a quasi-summary option is requested by a future API.Inactive groups get all-zero SEs.
Note: These are conditional-on-the-selected-model SEs from the penalized estimate. They do not account for model selection uncertainty (same convention as glmnet / mgcv).
- property coefficient_se_raw: dict[str, NDArray]#
Per-group coefficient standard errors assuming phi=1.
For known-scale families these are the exact-family standard errors, without the Pearson quasi-likelihood correction applied by
coefficient_se. For estimated-scale families these omit the fitted dispersion and therefore differ fromcoefficient_sewheneverphi != 1.Inactive groups get all-zero SEs.
- property intercept_se: float#
Quasi-likelihood standard error of the intercept.
Computed from the [0,0] element of the augmented Fisher information inverse, which accounts for covariance between the intercept and all other coefficients. Uses the same dispersion convention as
coefficient_se, including Pearson dispersion for known-scale families. Rendered coefficient tables retain fitted-family scale.
- property intercept_se_raw: float#
Exact-family intercept SE for known scale, otherwise assuming phi=1.
- feature_se(name: str, n_points: int = 200) dict[str, Any]#
Quasi-likelihood SE of a feature curve, levels, or coefficient.
Uses the same dispersion convention as
coefficient_se, including Pearson dispersion for known-scale families. Rendered coefficient tables retain fitted-family scale.Refuses for a term carrying a post-fit shape repair: the band would be propagated from the unconstrained fit’s covariance around constrained coefficients, which is the quantity
summary()withholds for that term. A per-feature accessor can say so exactly, so it does – unlikecoefficient_se, which is one dict over every group and already spends the all-zeros array on “not selected”.
- summary( ) ModelSummary#
Formatted model summary with coefficient table.
The rendered coefficient table uses fitted-family covariance: known-scale families retain
phi=1and estimated-scale families use their fittedphi. It deliberately does not reuse the explicit Pearson quasi-likelihood accessorscoefficient_se,intercept_se, orfeature_se.The dict-like
standard_errorspayload retains the historicalcoefficient_sekey for the explicit quasi-likelihood accessor and labels its scale separately from the rendered fitted-family table.- Parameters:
- alphafloat
Significance level for confidence intervals (default 0.05 → 95% CI).
- detailstr
Level of detail for spline terms.
"compact"(default) shows one row per spline group."full"adds per-coefficient detail rows (ASCII: printed inline; HTML: pre-expanded<details>disclosure). Default"compact"still shows closed disclosures in HTML.- level_displaystr
Categorical level presentation.
"expanded"(default) shows exact original levels;"grouped"shows one row per fitted group with a membership legend.
- Returns:
- ModelSummary
Object with
__str__(ASCII),_repr_html_(HTML), and dict-like access for backward compatibility.