Polynomial#

class superglm.Polynomial(
degree: int | None = None,
powers: Sequence[int] | None = None,
)#

Bases: object

Orthogonal polynomial feature (orthonormal in the training-weight measure).

Scales x to [-1, 1] using training-data min/max, seeds a Legendre basis (including the constant), and orthonormalizes it against the training sample_weight by weighted thin QR. The constant column is dropped from the emitted block — the model intercept carries it.

Normalization convention: with training weights w, the emitted components phi_j (one per stated power) satisfy

(1 / sum(w)) * sum_i w_i * phi_j(x_i) * phi_k(x_i) = delta_jk

i.e. they are orthonormal in the mean empirical inner product of the training sample_weight, and each is orthogonal to the constant in the same inner product. When exposure enters through an offset (the documented count workflow) sample_weight stays at ones, so the basis is orthonormalized against the row-count measure. The feature itself is weight-contract neutral: it follows exactly the stream its caller supplies. Scalar SuperGLM preserves likelihood-weight standardization, while distributional compilation supplies resolved data-derived geometry (physical rows for prior weights and replication mass for frequency weights). Under Gaussian/fixed-weight fitting where the supplied stream is also the fitting measure, this makes the per-power coefficient estimates exactly uncorrelated and gives the group penalty the within-group orthonormal geometry that the group lasso assumes (Yuan & Lin 2006, JRSS-B 68:49-67; Simon & Tibshirani 2012, Statistica Sinica 22(3):983-1001 — orthonormalizing within the group is exactly equivalent to their standardized group lasso).

The triangular factor of the weighted QR is stored as fitted state beside the min/max scaling. transform/score push new x through the same seed basis and the stored factor — they never re-orthogonalize against new data or new weights. Out-of-range x is plain polynomial evaluation on the scaled seed basis (unbounded growth); every orthogonality and uncorrelatedness statement is a property of the training measure only.

Honest caveats:

  • Exact uncorrelatedness of the per-power estimates holds when the supplied QR stream is also the fixed fitting measure (the scalar fixed-weight/Gaussian world). In a GLM it is approximate at the IRLS working weights; no published result quantifies that gap, and published group-penalty practice makes the same fixed spherical approximation of the working Hessian (Simon & Tibshirani 2012, section 5.3).

  • Dropping powers by their z-statistics is response-driven selection. Validate out-of-fold, or state powers from the plan.

Parameters:
degreeint, optional

Maximum polynomial degree; sugar for powers=range(1, degree+1). 2 (quadratic) or 3 (cubic) are the standard insurance choices. Defaults to 3 when neither degree nor powers is given. Mutually exclusive with powers.

powerssequence of int, optional

Distinct integers >= 1 naming the orthogonal components to keep, e.g. powers=[1, 2, 4]. The orthogonal basis is built up to max(powers) and the stated components are selected, so when the QR and fitting measures coincide, dropping a middle power leaves the retained components’ fitted coefficients unchanged. Excluding “power 3” excludes the degree-3 orthogonal component, not the raw x**3 monomial — on asymmetric exposure the degree-4 orthogonal polynomial carries x**3 monomial content. API precedent for a degree list: numpy.polynomial.Polynomial.fit(deg=[...]).

Notes

Group size = len(powers). Columns are emitted in ascending power order and summary rows are labelled by the stated power.

The QR-on-Legendre-seed build is the published standardized-group- lasso algorithm and is well conditioned at degree <= 8 on min/max- scaled data. If the degree ceiling ever rises, the upgrade path is the three-term recurrence for weighted discrete measures (Forsythe 1957; Gautschi 2004, Orthogonal Polynomials: Computation and Approximation): store the recurrence coefficients instead of the triangular factor and evaluate new x by re-running the recurrence.

The rank guard admits pivots down to 1e-10 of the largest (which is exactly 1 — the constant column under normalized weights), so a build just clearing it has cond(R) near 1e10 and its computed columns are orthonormal only to roughly eps * cond(R) ~ 1e-6: near the guard boundary the orthonormality and uncorrelatedness statements hold to that reduced precision, not machine precision.

build(
x: NDArray[floating],
sample_weight: NDArray[floating] | None = None,
) → GroupInfo#

Learn min/max and the weighted orthonormalization from x.

transform(
x: NDArray,
) → NDArray#

Evaluate the fitted orthonormal components at new x.

Pushes x through the stored min/max scaling, the Legendre seed, and the stored triangular factor.

score(
x: NDArray,
beta: NDArray,
) → NDArray#

Score the fitted polynomial contribution directly on new data.

reconstruct(
beta: NDArray,
n_points: int = 200,
) → dict[str, Any]#

Evaluate the fitted polynomial on a grid and return relativities.