superglm.SuperGLM#

class superglm.SuperGLM(
family: str | Distribution = 'poisson',
link: str | Link | None = None,
penalty: Penalty | str | None = None,
selection_penalty: float | Literal['auto'] | None = None,
spline_penalty: float | None = None,
penalty_features: str | list[str] | None = None,
features: Mapping[Hashable, FeatureSpec] | None = None,
splines: list[str] | None = None,
n_knots: int | list[int] = 10,
degree: int = 3,
categorical_base: str = 'most_exposed',
interactions: list[tuple[str, str] | object] | None = None,
active_set: bool = False,
direct_solve: str = 'auto',
discrete: bool = False,
n_bins: int | dict[str, int] = 256,
tol: float = 1e-06,
max_iter: int = 100,
convergence: str = 'deviance',
retain_fit_state: bool = True,
separation: str = 'warn',
group_pricing: Literal['rank', 'spanned'] = 'rank',
weight_semantics: Literal['prior', 'frequency'] = 'prior',
)#

Bases: object

Penalised generalised linear model with splines, group penalties, and REML.

Supports Poisson, Gaussian, Gamma, NB2, Tweedie, and Binomial families with group lasso, sparse group lasso, or ridge penalties. Smoothing parameters can be estimated via REML (fit_reml) or cross-validation (cross_validate).

Parameters:
familystr or Distribution

Response distribution. Strings "poisson", "gaussian", "gamma", "binomial" are accepted for parameter-free families. For parameterized families use Distribution objects: Tweedie(p=1.5), NegativeBinomial(theta=1.0), or the families module (e.g. families.tweedie(p=1.5)). For "binomial", y must be in {0, 1} and predict() returns probabilities.

linkstr or Link, optional

Link function. Defaults to the family’s configured default link.

penaltystr or Penalty, optional

Penalty type. One of "group_lasso", "sparse_group_lasso", "group_elastic_net", "ridge", or a Penalty object. Defaults to GroupLasso.

selection_penaltyfloat, {“auto”}, or None, optional

Regularisation strength for the group penalty (feature selection). None (default) and 0.0 disable selection. "auto" explicitly requests calibration to 10% of lambda_max at fit time.

spline_penaltyfloat, optional

Within-group ridge shrinkage for spline smoothing. Defaults to 0.1.

penalty_featuresstr or list[str], optional

Restrict the selection penalty to specific feature or group names. None (default) applies to all penalizable groups.

featuresmapping[hashable, FeatureSpec], optional

Explicit feature specifications mapping hashable column labels to feature objects (Spline, Categorical, Numeric, Polynomial). Mutually exclusive with splines.

splineslist[str], optional

Deprecated auto-detection shorthand. Column names in this list are treated as splines and all other columns are inferred as categorical or numeric. Use explicit features={"age": Spline(...)} for new code. Mutually exclusive with features.

n_knotsint or list[int]

Number of interior knots for auto-detect splines.

degreeint

B-spline degree for auto-detect splines.

categorical_basestr

Base level strategy for auto-detected categoricals.

interactionslist[tuple[str, str] or interaction spec], optional

Pairs of feature names to interact, or explicit interaction specifications such as FactorSmooth. Tuple interaction types are auto-detected from their parent feature specs.

active_setbool

Use active-set cycling in the BCD solver.

direct_solve{“auto”, “gram”, “qr”, “structured”}

Strategy for the direct IRLS solver (lambda1=0). "auto" selects compact structured elimination for eligible random-effect and factor-smooth terms above the measured crossover, otherwise using Gram. A globally unidentifiable SZ system also retries on Gram with an explicit recorded reason. "gram" forces the gram path. "qr" uses QR on the materialised weighted design matrix — backward-stable but O(n·p²) per iteration. Intended for smaller datasets. "structured" forces structured elimination for an eligible random-effect, FS, or SZ block.

discretebool

Use discretized basis matrices for large-n REML (fREML-style).

n_binsint or dict[str, int]

Number of discretization bins per feature when discrete=True.

tolfloat

Convergence tolerance for IRLS / PIRLS. Default 1e-6. Can also be set per-call via fit(tol=...) or fit_reml(pirls_tol=...). Fit-time values take precedence. Larger values (e.g. 1e-6) converge faster but may stop before near-separated coefficients have stabilised.

max_iterint

Maximum IRLS / PIRLS outer iterations. Default 100.

convergence{“deviance”, “coefficients”}

Convergence criterion. "deviance" (default) stops when relative deviance change drops below tol — fast, since well-identified coefficients lock in early. "coefficients" (experimental) stops when the maximum relative coefficient change drops below tol. May not converge for near-separated levels where the MLE is at −∞.

retain_fit_statebool

If True (default), keep training-scale fit state such as the fitted design matrix for later diagnostics. If False, eagerly computes compact inference state after fitting, then releases row-scale training caches while preserving prediction, summaries, and term confidence intervals.

separation{“warn”, “error”, “ignore”}

Build-time check for separated categorical cells: levels or crossed-interaction cells that carry exposure but whose responses all sit on the response boundary (e.g. no positive response under a log-link Tweedie/Poisson fit). Such cells have no finite maximum-likelihood effect, IRLS drifts until the objective stagnates, and the affected predictions collapse to the boundary while rank and aggregate metrics (gini, balance) still look healthy – only out-of-sample likelihood/deviance exposes the damage. "warn" (default) emits a SeparationWarning naming the offending cells and the remedies before fitting; "error" refuses the design with a SeparationError; "ignore" disables the check. Terms bounded by an active selection penalty are exempt, as their penalised optima are finite. The same mode governs the in-solver backstop that fires when an exhausted, stagnant IRLS run shows the extreme-working-weight signature of separation the build scan cannot see.

group_pricing{“rank”, “spanned”}

Dimension p_g at which the selection penalty and the fallback df ledger price a group whose spec emits fewer columns than the term spans (a categorical interaction with empty or nested cells). "rank" (default) prices the emitted, identifiable width, following the group-lasso literature’s derivation of sqrt(p_g) from the df of the group’s score statistic. "spanned" prices the width the term spans – the historical behaviour – so cell pruning is a pure reparametrisation of the fit. The choice moves fitted results only for models that combine such an interaction with an active group penalty – selection_penalty > 0 or an explicit penalty= whose lambda1 > 0 – and moves the reported effective_df/phi/AIC/BIC of any fit whose df falls back to the Breheny-Huang allocation.

weight_semantics{“prior”, “frequency”}

What sample_weight says about a row. "prior" (default) reads it as an EDM prior weight – a statement of precision, Var(Y_i) = phi V(mu_i) / w_i – which is what you have when the response is an average: incurred / exposure weighted by exposure, or an average severity weighted by claim count. This is the reading R’s glm and glum give their single weight argument, and the one statsmodels calls var_weights. "frequency" reads it as a replication count, so an integer weight is exactly equivalent to repeating the row; that is statsmodels’ freq_weights and Stata’s fweight.

The two agree only at w == 1 – integer weights do not make them coincide – and only the prior reading is a likelihood at fractional ones. They share a score equation, so beta is unchanged; what moves is phi, every Wald standard error and interval, residual degrees of freedom, the effective n in AIC/BIC, and – through the REML criterion – the smoothing parameters, the effective degrees of freedom and hence the fitted surface. Spline knot placement moves too: frequency mass shapes the same support as replicated rows, while prior weights leave learned geometry a function of physical rows. Unweighted fits and fits with w == 1 are identical under both.

See also

The methods and attributes of this class are grouped by task on the SuperGLM reference pages; each has its own page.