superglm.SuperLSS#

class superglm.SuperLSS(
family: DistributionalFamily,
/,
*predictors: BoundPredictor,
weight_semantics: Literal['prior', 'frequency'] = 'prior',
discrete: bool = False,
n_bins: int | Mapping[str, int] = 256,
separation: Literal['warn', 'error', 'ignore'] = 'warn',
coefficient_curvature: Literal['observed', 'fisher'] = 'observed',
)#

Bases: object

Fit several parameters of a response distribution jointly.

Create a family, then pass it followed by one declaration for each of its parameters. For example, a Gaussian model has a location predictor for its conditional mean and a scale predictor for its standard deviation. Each predictor can use its own linear terms, smooths and categorical effects.

Parameters:
familyDistributionalFamily

Response family, passed as the first positional argument. Create all predictor declarations with helpers on this same family instance.

*predictorsBoundPredictor

One declaration per family parameter. A bare column name declares a numeric linear term; use s for a spline, cat for categories and re for a random effect. An empty helper call declares an intercept-only predictor. Every parameter must be declared exactly once, in any order. Unpack a sequence of declarations with *.

weight_semantics{“prior”, “frequency”}, default=”prior”

Meaning of sample_weight during fitting and scoring. Prior weights describe precision under the family’s observation law; frequency weights count repeated observations. Families without a prior-weight law require unit prior weights or frequency weights.

discretebool, default=False

Use grouped marginal designs and row chunks during fitting. This can reduce design memory, but it does not make fitting out of core.

n_binsint or mapping of str to int, default=256

Bin count for discrete fitting, shared by all features or supplied by feature name. Check sensitivity to the grid for your data.

separation{“warn”, “error”, “ignore”}, default=”warn”

How to handle categorical cells whose response values imply no finite coefficient for a predictor. The family defines these boundaries.

coefficient_curvature{“observed”, “fisher”}, default=”observed”

Curvature used for coefficient updates. "fisher" requires a family that supplies expected information.

See also

GaussianLS

Gaussian location and standard deviation.

GammaLS

Positive responses with mean and coefficient of variation.

TweedieLSS

Nonnegative responses with mean, dispersion and power.

bind_predictor

Declare predictors for a custom family.

Notes

Construction copies the family configuration and predictor declarations. Fitting one model does not fit the family or another model constructed from it. Use fit_reml to estimate coefficients and smoothing parameters, or fit to hold smoothing parameters fixed.

predict returns the conditional mean for the built-in families. predict_parameters returns every fitted distribution parameter, and predict_link returns their linear predictors. Result columns use family parameter names. In particular, Tweedie’s mu, phi and p helpers produce mean, dispersion and power columns.

Examples

Declare a Gaussian model whose mean varies with age and region, with constant standard deviation:

>>> from superglm import GaussianLS, SuperLSS, cat, s
>>> family = GaussianLS()
>>> model = SuperLSS(
...     family,
...     family.location(s("age", kind="cr", k=8), cat("region")),
...     family.scale(),
... )
>>> tuple(p.name for p in model.predictors)
('location', 'scale')

The model is ready for model.fit_reml(X, y). The fit and prediction frames must contain the declared columns.

See also

The methods and attributes of this class are grouped by task on the SuperLSS reference pages; each has its own page.