Quick Start#

This page covers scalar-response SuperGLM models. To model several parameters of a response distribution, follow Your first distributional model with SuperLSS.

Large-n REML#

If the dataset is large, keep the same workflow but turn on discrete REML:

model = SuperGLM(
    family="poisson",
    selection_penalty=0.0,
    discrete=True,
    n_bins=256,
    features=features,
)
model.fit_reml(df, y, sample_weight=exposure, max_reml_iter=30)

Legacy Auto-Detection#

Auto-detection remains available for compatibility, but explicit feature specs are the canonical API. Passing the legacy splines keyword emits a FutureWarning; calls continue to run during the 0.15 compatibility window, including with n_knots, degree, and categorical_base. Warning-as-error environments should migrate named smooths to features={"column": Spline(kind="ps")} before upgrading.

from superglm import Spline, SuperGLM

model = SuperGLM(
    family="poisson",
    selection_penalty=0.0,
    features={
        "DrivAge": Spline(n_knots=10),
        "VehAge": Spline(n_knots=10),
        "BonusMalus": Spline(n_knots=10),
    },
)
model.fit_reml(df, y, sample_weight=exposure)

Eager Polars Input#

The same named-frame API accepts eager Polars frames. Install Polars separately and collect a LazyFrame before passing it to SuperGLM.

import polars as pl

from superglm import Spline, SuperGLM

X = pl.DataFrame({"age": [20.0, 30.0, 40.0], "region": ["N", "S", "N"]})
y = [0.0, 1.0, 2.0]

model = SuperGLM(
    family="poisson",
    features={"age": Spline()},
).fit(X, y)
predictions = model.predict(X)

Input storage remains Polars through fitting and prediction. Public tabular results such as inference tables, diagnostics, plot data, and rating-table payloads remain pandas DataFrames.

Sparse Screening Or Fixed-Penalty Fitting#

If your goal is sparse screening or compression rather than GAM-style REML inference, use fit() with selection_penalty > 0 instead:

model = SuperGLM(
    family="poisson",
    penalty="group_elastic_net",
    selection_penalty=0.01,
    spline_penalty=0.1,
    features=features,
)
model.fit(df, y, sample_weight=exposure)

Weights And Offsets#

For the Poisson frequency examples on this page, sample_weight=exposure weights a per-exposure rate response by exposure. It does not enter the linear predictor or automatically multiply the mean. Weight semantics are family-specific. In particular, Tweedie uses finite, strictly positive EDM prior weights with Var(Y_i | x_i) = phi * mu_i**p / w_i, not replication counts. The exposure= fit alias has been removed.

Two common patterns for count models:

import numpy as np

# Raw count target: offset absorbs exposure, model estimates a rate
model.fit(df, claim_counts, offset=np.log(exposure))

# Rate target (count / exposure): sample_weight carries exposure
model.fit(df, claim_rate, sample_weight=exposure)

Next steps: