superglm.SuperLSS#
- class superglm.SuperLSS(
- family: DistributionalFamily,
- /,
- *predictors: BoundPredictor,
- weight_semantics: Literal['prior', 'frequency'] = 'prior',
- discrete: bool = False,
- n_bins: int | Mapping[str, int] = 256,
- separation: Literal['warn', 'error', 'ignore'] = 'warn',
- coefficient_curvature: Literal['observed', 'fisher'] = 'observed',
Bases:
objectFit several parameters of a response distribution jointly.
Create a family, then pass it followed by one declaration for each of its parameters. For example, a Gaussian model has a
locationpredictor for its conditional mean and ascalepredictor for its standard deviation. Each predictor can use its own linear terms, smooths and categorical effects.- Parameters:
- familyDistributionalFamily
Response family, passed as the first positional argument. Create all predictor declarations with helpers on this same family instance.
- *predictorsBoundPredictor
One declaration per family parameter. A bare column name declares a numeric linear term; use
sfor a spline,catfor categories andrefor a random effect. An empty helper call declares an intercept-only predictor. Every parameter must be declared exactly once, in any order. Unpack a sequence of declarations with*.- weight_semantics{“prior”, “frequency”}, default=”prior”
Meaning of
sample_weightduring fitting and scoring. Prior weights describe precision under the family’s observation law; frequency weights count repeated observations. Families without a prior-weight law require unit prior weights or frequency weights.- discretebool, default=False
Use grouped marginal designs and row chunks during fitting. This can reduce design memory, but it does not make fitting out of core.
- n_binsint or mapping of str to int, default=256
Bin count for discrete fitting, shared by all features or supplied by feature name. Check sensitivity to the grid for your data.
- separation{“warn”, “error”, “ignore”}, default=”warn”
How to handle categorical cells whose response values imply no finite coefficient for a predictor. The family defines these boundaries.
- coefficient_curvature{“observed”, “fisher”}, default=”observed”
Curvature used for coefficient updates.
"fisher"requires a family that supplies expected information.
See also
GaussianLSGaussian location and standard deviation.
GammaLSPositive responses with mean and coefficient of variation.
TweedieLSSNonnegative responses with mean, dispersion and power.
bind_predictorDeclare predictors for a custom family.
Notes
Construction copies the family configuration and predictor declarations. Fitting one model does not fit the family or another model constructed from it. Use
fit_remlto estimate coefficients and smoothing parameters, orfitto hold smoothing parameters fixed.predictreturns the conditional mean for the built-in families.predict_parametersreturns every fitted distribution parameter, andpredict_linkreturns their linear predictors. Result columns use family parameter names. In particular, Tweedie’smu,phiandphelpers producemean,dispersionandpowercolumns.Examples
Declare a Gaussian model whose mean varies with age and region, with constant standard deviation:
>>> from superglm import GaussianLS, SuperLSS, cat, s >>> family = GaussianLS() >>> model = SuperLSS( ... family, ... family.location(s("age", kind="cr", k=8), cat("region")), ... family.scale(), ... ) >>> tuple(p.name for p in model.predictors) ('location', 'scale')
The model is ready for
model.fit_reml(X, y). The fit and prediction frames must contain the declared columns.
See also
The methods and attributes of this class are grouped by task on the SuperLSS reference pages; each has its own page.