superglm.SuperGLM#
- class superglm.SuperGLM(
- family: str | Distribution = 'poisson',
- link: str | Link | None = None,
- penalty: Penalty | str | None = None,
- selection_penalty: float | Literal['auto'] | None = None,
- spline_penalty: float | None = None,
- penalty_features: str | list[str] | None = None,
- features: Mapping[Hashable, FeatureSpec] | None = None,
- splines: list[str] | None = None,
- n_knots: int | list[int] = 10,
- degree: int = 3,
- categorical_base: str = 'most_exposed',
- interactions: list[tuple[str, str] | object] | None = None,
- active_set: bool = False,
- direct_solve: str = 'auto',
- discrete: bool = False,
- n_bins: int | dict[str, int] = 256,
- tol: float = 1e-06,
- max_iter: int = 100,
- convergence: str = 'deviance',
- retain_fit_state: bool = True,
- separation: str = 'warn',
- group_pricing: Literal['rank', 'spanned'] = 'rank',
- weight_semantics: Literal['prior', 'frequency'] = 'prior',
Bases:
objectPenalised generalised linear model with splines, group penalties, and REML.
Supports Poisson, Gaussian, Gamma, NB2, Tweedie, and Binomial families with group lasso, sparse group lasso, or ridge penalties. Smoothing parameters can be estimated via REML (
fit_reml) or cross-validation (cross_validate).- Parameters:
- familystr or Distribution
Response distribution. Strings
"poisson","gaussian","gamma","binomial"are accepted for parameter-free families. For parameterized families use Distribution objects:Tweedie(p=1.5),NegativeBinomial(theta=1.0), or thefamiliesmodule (e.g.families.tweedie(p=1.5)). For"binomial", y must be in {0, 1} andpredict()returns probabilities.- linkstr or Link, optional
Link function. Defaults to the family’s configured default link.
- penaltystr or Penalty, optional
Penalty type. One of
"group_lasso","sparse_group_lasso","group_elastic_net","ridge", or a Penalty object. Defaults toGroupLasso.- selection_penaltyfloat, {“auto”}, or None, optional
Regularisation strength for the group penalty (feature selection).
None(default) and0.0disable selection."auto"explicitly requests calibration to 10% of lambda_max at fit time.- spline_penaltyfloat, optional
Within-group ridge shrinkage for spline smoothing. Defaults to 0.1.
- penalty_featuresstr or list[str], optional
Restrict the selection penalty to specific feature or group names.
None(default) applies to all penalizable groups.- featuresmapping[hashable, FeatureSpec], optional
Explicit feature specifications mapping hashable column labels to feature objects (
Spline,Categorical,Numeric,Polynomial). Mutually exclusive with splines.- splineslist[str], optional
Deprecated auto-detection shorthand. Column names in this list are treated as splines and all other columns are inferred as categorical or numeric. Use explicit
features={"age": Spline(...)}for new code. Mutually exclusive with features.- n_knotsint or list[int]
Number of interior knots for auto-detect splines.
- degreeint
B-spline degree for auto-detect splines.
- categorical_basestr
Base level strategy for auto-detected categoricals.
- interactionslist[tuple[str, str] or interaction spec], optional
Pairs of feature names to interact, or explicit interaction specifications such as
FactorSmooth. Tuple interaction types are auto-detected from their parent feature specs.- active_setbool
Use active-set cycling in the BCD solver.
- direct_solve{“auto”, “gram”, “qr”, “structured”}
Strategy for the direct IRLS solver (lambda1=0).
"auto"selects compact structured elimination for eligible random-effect and factor-smooth terms above the measured crossover, otherwise using Gram. A globally unidentifiable SZ system also retries on Gram with an explicit recorded reason."gram"forces the gram path."qr"uses QR on the materialised weighted design matrix — backward-stable but O(n·p²) per iteration. Intended for smaller datasets."structured"forces structured elimination for an eligible random-effect, FS, or SZ block.- discretebool
Use discretized basis matrices for large-n REML (fREML-style).
- n_binsint or dict[str, int]
Number of discretization bins per feature when
discrete=True.- tolfloat
Convergence tolerance for IRLS / PIRLS. Default
1e-6. Can also be set per-call viafit(tol=...)orfit_reml(pirls_tol=...). Fit-time values take precedence. Larger values (e.g.1e-6) converge faster but may stop before near-separated coefficients have stabilised.- max_iterint
Maximum IRLS / PIRLS outer iterations. Default
100.- convergence{“deviance”, “coefficients”}
Convergence criterion.
"deviance"(default) stops when relative deviance change drops below tol — fast, since well-identified coefficients lock in early."coefficients"(experimental) stops when the maximum relative coefficient change drops below tol. May not converge for near-separated levels where the MLE is at −∞.- retain_fit_statebool
If True (default), keep training-scale fit state such as the fitted design matrix for later diagnostics. If False, eagerly computes compact inference state after fitting, then releases row-scale training caches while preserving prediction, summaries, and term confidence intervals.
- separation{“warn”, “error”, “ignore”}
Build-time check for separated categorical cells: levels or crossed-interaction cells that carry exposure but whose responses all sit on the response boundary (e.g. no positive response under a log-link Tweedie/Poisson fit). Such cells have no finite maximum-likelihood effect, IRLS drifts until the objective stagnates, and the affected predictions collapse to the boundary while rank and aggregate metrics (gini, balance) still look healthy – only out-of-sample likelihood/deviance exposes the damage.
"warn"(default) emits aSeparationWarningnaming the offending cells and the remedies before fitting;"error"refuses the design with aSeparationError;"ignore"disables the check. Terms bounded by an active selection penalty are exempt, as their penalised optima are finite. The same mode governs the in-solver backstop that fires when an exhausted, stagnant IRLS run shows the extreme-working-weight signature of separation the build scan cannot see.- group_pricing{“rank”, “spanned”}
Dimension
p_gat which the selection penalty and the fallback df ledger price a group whose spec emits fewer columns than the term spans (a categorical interaction with empty or nested cells)."rank"(default) prices the emitted, identifiable width, following the group-lasso literature’s derivation ofsqrt(p_g)from the df of the group’s score statistic."spanned"prices the width the term spans – the historical behaviour – so cell pruning is a pure reparametrisation of the fit. The choice moves fitted results only for models that combine such an interaction with an active group penalty –selection_penalty > 0or an explicitpenalty=whoselambda1 > 0– and moves the reportedeffective_df/phi/AIC/BIC of any fit whose df falls back to the Breheny-Huang allocation.- weight_semantics{“prior”, “frequency”}
What
sample_weightsays about a row."prior"(default) reads it as an EDM prior weight – a statement of precision,Var(Y_i) = phi V(mu_i) / w_i– which is what you have when the response is an average:incurred / exposureweighted by exposure, or an average severity weighted by claim count. This is the reading R’sglmand glum give their single weight argument, and the one statsmodels callsvar_weights."frequency"reads it as a replication count, so an integer weight is exactly equivalent to repeating the row; that is statsmodels’freq_weightsand Stata’sfweight.The two agree only at
w == 1– integer weights do not make them coincide – and only the prior reading is a likelihood at fractional ones. They share a score equation, sobetais unchanged; what moves isphi, every Wald standard error and interval, residual degrees of freedom, the effectivenin AIC/BIC, and – through the REML criterion – the smoothing parameters, the effective degrees of freedom and hence the fitted surface. Spline knot placement moves too: frequency mass shapes the same support as replicated rows, while prior weights leave learned geometry a function of physical rows. Unweighted fits and fits withw == 1are identical under both.
See also
The methods and attributes of this class are grouped by task on the SuperGLM reference pages; each has its own page.