fit#

SuperGLM.fit(
X: object,
y: NDArray,
sample_weight: NDArray | None = None,
offset: NDArray | None = None,
*,
tol: float | None = None,
max_iter: int | None = None,
convergence: str | None = None,
record_diagnostics: bool = False,
) → SuperGLM#

Fit the model to data.

Parameters:
Xpandas or eager Polars DataFrame

Feature matrix with columns matching registered features. Lazy frames must be collected before fitting.

yarray-like

Response variable.

sample_weightarray-like, optional

Observation weights, read under the model’s declared weight_semantics. Defaults to 1 for all observations.

Under "prior" (the default) they are EDM prior weights: Y_i ~ ED(mu_i, phi / w_i), so Var(Y_i | x_i) = phi * V(mu_i) / w_i, and estimated dispersion uses the count of positive-weight rows minus edf. Under "frequency" they are replication counts: once feature geometry is fixed, integer weights are likelihood-equivalent to row replication, and estimated dispersion uses sum(w) - edf. A Tweedie fit under "prior" additionally requires finite, strictly positive weights, because its compound-Poisson normalizer carries log w.

Weights affect fitting but do not enter the linear predictor or automatically scale the conditional mean. The model mean is mu_i = g**-1(x_i.T @ beta + offset_i).

offsetarray-like, optional

Offset added to the linear predictor. sample_weight does not supply an offset. To make a raw-count mean scale with exposure, pass offset=np.log(exposure). For a per-exposure response, pass exposure as sample_weight without adding it again to the linear predictor.

record_diagnosticsbool

If True, record per-iteration IRLS diagnostics (W range, mu/eta range, step halvings, worst-observation indices) on result.iteration_log. Useful for debugging convergence.

Returns:
SuperGLM

The fitted model (self).