cat#

superglm.cat(
column: str,
*,
base: str = 'most_exposed',
grouping: LevelGrouping | None = None,
levels: Any = None,
unseen: Literal['error', 'base'] = 'error',
) → BoundTerm#

Declare a categorical effect with a reference level.

Use cat("region") for categories, including categories stored as numbers. A bare string in a predictor always declares a numeric linear term and does not infer categorical encoding from the column’s dtype.

Parameters:
columnstr

Name of the categorical input column.

basestr, default=”most_exposed”

Reference level. Use the level with the greatest total sample weight, "first" for the first level, or a specific level name.

groupingLevelGrouping, optional

Combine input levels into groups before encoding.

levelssequence, data column or categorical dtype, optional

Declare the allowed input levels. With grouping, these are the original levels before grouping.

unseen{“error”, “base”}, default=”error”

Prediction policy for levels outside the fitted level universe. "base" uses the reference level and emits a warning.

Returns:
BoundTerm

A categorical declaration for the named column.

See also

Categorical

Encoding, grouping and level-universe rules.

re

A penalized effect with a coefficient for every level.