Feature Types¶
Choose the simplest feature type that matches the shape you want in the final pricing model.
Splines¶
Spline(kind, k) is the main public spline API. k is the public basis size
in the mgcv sense; the fitted smooth then absorbs the identifiability
constraint.
Spline(kind="ps", k=14) # default P-spline choice
Spline(kind="bs", k=14) # integrated-derivative B-spline smooth
Spline(kind="cr", k=10) # cubic regression spline
Spline(kind="ns", k=10) # natural spline
Spline(kind="ps", k=14, select=True) # REML + double-penalty shrinkage
Spline(kind="cr", k=12, m=(1, 2)) # multi-order penalty
Which spline kind to choose¶
| Kind | Use when | Notes |
|---|---|---|
"ps" |
default pricing spline | P-spline with difference penalty |
"bs" |
you want a proper B-spline smooth / mgcv-style bs basis |
integrated-derivative penalty on the same raw B-spline geometry |
"cr" |
you want a cubic regression spline / mgcv-style cr basis |
natural boundary constraints plus identifiability |
"ns" |
you want a natural spline with fixed natural boundaries | does not support monotone fitting |
Knot strategies¶
| Strategy | Description |
|---|---|
"uniform" |
evenly spaced interior knots |
"quantile_rows" |
more knots where the training data is dense |
"quantile_tempered" |
blend between uniform and quantile placement |
quantile_tempered with a small knot_alpha is often a good pricing default
for skewed variables like Bonus-Malus.
select=True¶
select=True adds mgcv-style double-penalty shrinkage to the spline term. This
is the REML-native way to let a smooth shrink toward linear or zero while
staying in the fit_reml() workflow.
Multi-order penalties with m=¶
m can be a single integer or a tuple. With a tuple, the spline emits
multiple penalty components on the same coefficient block, each with its own
REML smoothing parameter.
Current limitations:
select=True + m=(...)is not yet supported- tensor interactions with multi-order spline parents are not yet supported
kind="cr_cardinal"currently supports onlym=2
Monotone Splines¶
If monotonicity is part of the model specification, prefer solver-backed monotone fitting rather than post-fit repair.
from superglm import BSplineSmooth, Constraint, CubicRegressionSpline, PSpline
BSplineSmooth(n_knots=8, constraint=Constraint.fit.increasing) # QP
CubicRegressionSpline(n_knots=8, constraint=Constraint.fit.decreasing) # QP
PSpline(n_knots=10, constraint=Constraint.fit.increasing) # SCOP
See Monotone Splines for the full decision guide.
Polynomial¶
Orthogonal polynomials are a good option when the shape is simple and stable.
Categorical¶
Categoricals are one-hot encoded with a reference level. The entire factor is treated as one group for selection and inference.
Collapsing Sparse Levels¶
collapse_levels(...) lets you merge sparse levels for fitting while keeping
the mapping back to original levels for inference and plotting.
from superglm import Categorical, collapse_levels
grouping = collapse_levels(df["Area"], groups={"Rural": ["E", "F"]})
area = Categorical(base="most_exposed", grouping=grouping)
This is useful when a tariff factor has many thin levels but you still want a single grouped factor inside the model.
OrderedCategorical¶
Use OrderedCategorical(...) when a factor has a real order and you want a
smooth effect across levels. Level positions are equally spaced unless you
provide an explicit values={level: position} mapping.
The legacy basis="spline" and spline shortcut arguments such as kind= and
n_knots= are deprecated; configure them on basis=Spline(...). Step smoothing
with basis="step" is also deprecated and will be removed. Use Spline(...)
for smoothing or Categorical(...) for independent level effects.
Inference follows the spline model, not a saturated categorical model. The summary reports one Wood-style whole-smooth p-value for the ordered term; its null hypothesis is that the smooth contributes no variation after centering. Per-level rows are base-relative effect estimates with standard errors and confidence intervals, but deliberately have no p-values or significance codes. Changing the reporting base therefore changes the displayed level contrasts, not the whole-smooth p-value.
This interpretation depends on the numeric positions assigned to the levels.
order=[...] uses equal spacing on [0, 1]; use values={...} when the real
distances are unequal. If spacing and smoothness are not defensible assumptions,
fit Categorical(...) and use a whole-term comparison instead.
Numeric¶
Numeric() is a simple passthrough for continuous variables that should enter
linearly.
Interactions¶
Interactions are declared separately via interactions=[(...)], and the type
is inferred from the parent specs.
model = SuperGLM(
features={"age": Spline(kind="ps", k=14), "region": Categorical()},
interactions=[("age", "region")],
selection_penalty=0.01,
)
See Interactions for the full interaction map.