Piecewise#

class superglm.Piecewise(
breaks: Sequence[float | str] | int,
*,
base: float | str = 'most_exposed',
strategy: str = 'quantile',
lower: float | None = None,
upper: float | None = None,
extrapolation: str = 'clip',
degrees: Sequence[int] | None = None,
)#

Bases: object

Continuous piecewise-linear feature on stated breakpoints.

Given J breakpoints the knot vector is [lower, *breaks, upper] (J + 2 knots) and the basis is the degree-1 B-spline (hat) basis on those knots, evaluated directly rather than through scipy so that the outer segments extend linearly past the boundary knots. The hat at a base knot t_r is dropped for identifiability against the model intercept, so the group has J + 1 columns and each retained coefficient is exactly

v_j = f(t_j) - f(t_r) – the log relativity at knot j, against base.

Parameters:
breakssequence of float, or int

The breakpoints. A sequence is the primary mode and the defensible one: you state where the kinks are. An int places that many breakpoints at exposure-weighted quantiles of x, snapped to observed values – a convenience for exploration, not for a filed tariff. Heaped data (ages ending 0/5, whole-year tenures) routinely collapses tied quantiles, so int mode can realise fewer breakpoints than requested; it warns rather than raising, because that outcome is the library’s doing and not the caller’s.

As the basis= of an OrderedCategorical the breaks may also be stated as BAND NAMES (breaks=["Mi060", "Mi066"]), resolved to level positions when the ordered term is constructed; integer positions on the level axis are the escape hatch. On the numeric axis a name has nothing to resolve against, so string breaks refuse at build().

degreessequence of int, optional

One polynomial degree per segment (len(breaks) + 1 entries), default all 1 – the plain kinked line. 0 states a flat segment (the grouped/plateau tail); 2 and above add within-segment curvature. Seams stay value-continuous by construction (the grafted- polynomial device: Gallant & Fuller 1973; plateau tails: Anderson & Nelson 1975). Legal only when this Piecewise is the basis= of an OrderedCategorical: there the export contract is one row per band, so the table is exact at any degree, while on the numeric axis the exported workbook is exact under linear interpolation only at degree 1 – a numeric-axis build with any degree != 1 refuses loudly. Requires stated breaks (int-mode placement has no stated segments).

basefloat or {‘most_exposed’, ‘first’}

Reference knot, mirroring Categorical. 'most_exposed' picks the knot carrying the largest hat-carried mass under the fit weights; with no sample_weight the weights are ones, so it is the knot carrying the most rows. A float must equal exactly one knot. The resolved base is sticky: _base_knot persists across build() calls, so after a refit whose knot set changed, base='first' can legitimately remain on a surviving interior knot rather than move to the new first knot.

strategy{‘quantile’}

Placement rule, consulted only when breaks is an int.

lower, upperfloat, optional

Pin the outermost knots. Default min(x) / max(x). Pinning wider than the data states a rated range the tariff must cover. Pinning narrower is allowed too, and what happens to the rows outside [lower, upper] is the extrapolation parameter’s call: under "clip" they are grouped onto the boundary knots, so Piecewise(breaks, upper=u) fits identically to precomputing x.clip(max=u) – the tail-grouping idiom stated as a term parameter instead of a preprocessing step; under "extend" they load the linear tails, and their leverage lands entirely on the two boundary segments’ slopes; under "error" the build refuses.

extrapolation{‘clip’, ‘extend’, ‘error’}

Behaviour outside [lower, upper], mirroring Spline. "clip" (default) holds the boundary knot’s value. "extend" continues the boundary segments at their fitted slopes. "error" raises on out-of-range values. The policy binds at build() as well as at prediction, so the design that is fitted is the design the tariff states.

Notes

Outside [t_0, t_{J+1}] the term follows extrapolation, and the default mirrors this library’s splines: hold the boundary knot’s value flat. Under "extend" the boundary segments’ slopes continue past the boundary knots – and the slope is stated: lower / upper pin where the boundary segments start and the exported table prints the two boundary slopes, so the rule outside the tabulated range is reproducible by hand. The exported workbook states whichever rule is in force.

In the editor the term gets one control handle per knot – the handle is the coefficient, because the raw basis evaluated at the knots is the identity. The editor’s hard cap of 24 handles still applies, so a term with more than 24 knots displays a subsampled set of them and the knots without a handle cannot be dragged. State fewer breakpoints if every knot has to be editable.

build(
x: NDArray[floating],
sample_weight: NDArray[floating] | None = None,
) → GroupInfo#

Resolve knots and base from x, validate, and return the J+1 columns.

Validation runs in a fixed order, and the order is part of the contract because several rules fire together on a degenerate input: (1) finite x, (2) non-empty breaks, (3) known strategy, (4) increasing sequence breaks, (6a) lower < upper, then the extrapolation policy binds (clip groups the out-of-range rows onto the boundary knots, error refuses them), (5) int-mode placement, (6b) breaks strictly inside, (7) base resolves to one knot, (8) no zero-mass knot, (9) no in-range empty segment, (10) full column rank against the intercept, (11) thin-segment warning.

Level-axis-only declarations refuse before any of that: band-name breaks and per-segment degrees= exist only where an OrderedCategorical hosts this spec as its basis=.

transform(
x: NDArray,
) → NDArray[float64]#

Evaluate the identifiable retained columns on new data.

Out-of-range x follows the term’s extrapolation policy. On the legacy (all-degree-1) path these are the J+1 non-base hats; on the segmented path, the non-base knot-value groups then the curvature columns.

score(
x: NDArray,
beta: NDArray[floating],
) → NDArray[floating]#

Score the fitted piecewise contribution directly on new data.

reconstruct(
beta: NDArray[floating],
) → dict[str, Any]#

Coefficients -> per-knot relativities and derived segment slopes.

ordered_structural_rows() → list[StructuralContrastRow]#

Structural contrasts of a fitted piecewise term, for summaries.

One slope-change contrast per stated break and one curvature family per segment of degree >= 2. This is the fixed-knot truncated-power inference vocabulary: with the knots stated as inputs the design is linear in every parameter, the slope change at a stated join is the coefficient of its plus-function, and both hypotheses are ordinary Wald/F tests (Smith 1979, “Splines as a useful and convenient statistical tool”, Am. Statist. 33(2):57-62; Sprent 1961 for the known-changeover two-phase case). Deliberately NOT per-segment per-power z-rows: under C0 seams the segments share their joint values, so within-segment orthogonal components are not free parameters and that clean-z geometry does not exist here.

Contrast vectors are stated over the retained columns of transform and are invariant to the base-column choice.