Categorical#

class superglm.Categorical(
base: str = 'most_exposed',
grouping=None,
*,
levels=None,
unseen: str = 'error',
)#

Bases: object

One-hot encoded categorical feature.

Parameters:
basestr

How to choose the reference level.

  • 'most_exposed' - level with highest total sample_weight (default, best for insurance)

  • 'first' - first level in the level universe (alphabetical when inferred, as declared when levels= or a categorical dtype bounds it)

Or pass a specific level name as a string.

groupingLevelGrouping, optional

Collapse original levels into groups before fitting.

levelslist | tuple | Series | ndarray | CategoricalDtype, optional

The level universe to bind to (spec 2026-08-11, §3.1). With a grouping this declares the RAW, pre-collapse universe. Levels with no training rows are pinned to base rather than dropped; training rows outside the universe are an error.

unseen{‘error’, ‘base’}

Predict-time policy for levels outside the universe. ‘error’ (default) is the historical behavior; ‘base’ routes those rows to the base level with one warning per call.

adopt_dtype_categories(categories: list) → None#

Adopt a dtype-declared universe unless one is already declared.

apply_level_binding(binding) → None#

Adopt a full-frame binding: universe if unset, base pin if unpinned.

resolve_binding(
values: NDArray,
sample_weight=None,
)#

Compute this spec’s full-frame binding without mutating the spec.

build(
x: NDArray,
sample_weight: NDArray[floating] | None = None,
) → GroupInfo#

Build sparse one-hot design columns, choosing the base level from x.

transform(
x: NDArray,
) → NDArray#

One-hot encode using levels learned during build().

score(
x: NDArray,
beta: NDArray[floating],
) → NDArray[floating]#

Score the fitted categorical contribution directly on new data.

reconstruct(
beta: NDArray[floating],
) → dict[str, Any]#

Coefficients -> relativity table.