estimate_p#

SuperGLM.estimate_p(
X: object,
y: NDArray,
sample_weight: NDArray | None = None,
offset: NDArray | None = None,
*,
fit_mode: str = 'fit',
search_fit_mode: str | None = None,
p_bounds: tuple[float, float] = (1.05, 1.95),
xatol: float = 0.001,
ci_alpha: float | None = None,
max_reml_iter: int | None = None,
progress_callback: Callable[[...], None] | None = None,
)#

Estimate Tweedie p via profile likelihood, refit, and return result.

At each candidate p the mean is refitted and phi is its maximum-likelihood value at that mean; p minimises the resulting profile negative log-likelihood by bounded Brent search (Dunn & Smyth 2005), and the published fit is refitted at the selected p. The search is local: if the profile has more than one interior minimum in p_bounds, it can return one that is not the global minimum, with no warning. Only an estimate on a bound is flagged.

Parameters:
Xpandas or eager Polars DataFrame

Feature matrix. Lazy frames must be collected before fitting.

yarray-like

Response variable.

sample_weightarray-like, optional

Finite, strictly positive EDM prior weights, not replication or frequency weights. The Tweedie variance convention is Var(Y_i | x_i) = phi * mu_i**p / w_i. Remove zero-weight rows consistently from X, y, sample_weight, and offset before calling this method.

offsetarray-like, optional

Offset added to the linear predictor.

fit_mode{“fit”, “reml”, “inherit”}

Fitting regime for the published final fit, and by default for each candidate p evaluation too.

search_fit_mode{“fit”, “reml”, “inherit”}, optional

Fitting regime for the candidate p evaluations alone, when it should differ from the regime the final fit is published under. The default None searches under fit_mode.

fit_mode="reml", search_fit_mode="fit" selects p under ordinary ML and then publishes one REML fit at the selected p, which is the cost of a single smoothing-parameter search rather than one per candidate. In either coupling the published fit is publication-grade: candidate fits only rank powers and run at the loose search tolerance, while the published refit runs the tight publication default and re-profiles dispersion against its own fitted mean, so the returned estimates describe the model you get back. Selecting p under a different objective than the one that publishes it is an approximation – on the synthetic benchmark fixture it moved p_hat by 1.1e-6 while running about 1.8x faster – so measure it on your own data before relying on it. A decoupled search also never meets the certifiable-region boundary a coupled REML search must route around (a coupled run warns when its optimum is pinned against that boundary), at the price that the publication refit may instead fail at the selected p with a typed superglm.PublicationModeError naming the ways out. Likelihood-ratio confidence intervals remain available either way; they invert the searched profile, so they describe the regime named by search_fit_mode.

p_boundstuple of float

Search interval for p, strictly inside (1, 2). An estimate on a bound is warned about and recorded in result.warnings: a maximum as p -> 1 can be an artefact of rounded responses. Near p = 1 the dispersion profile is multimodal and its global search grows like 1 / (p - 1); a power whose search would exceed its bound (superglm.NearPoissonDispersionError) is skipped as infeasible and recorded.

xatolfloat

Absolute resolution of the Brent search in p.

ci_alphafloat, optional

Significance level for an explicitly requested likelihood-ratio profile confidence interval. For example, 0.05 computes a 95% interval and caches it for model.summary(alpha=0.05). The default None performs no confidence-interval evaluations. If the searched winner’s fit, or a fit the interval evaluates, did not converge, the interval is still computed, and the cause is warned and recorded in result.warnings.

max_reml_iterint, optional

Outer-iteration budget for the REML publication refit alone; candidate search fits keep their own budget. Requires fit_mode="reml" – a pure-ML publication has no REML iteration to budget and refuses the parameter. The default None uses the fit_reml default of 20.

progress_callbackcallable, optional

Called as progress_callback(phase, payload): "profiling" with {"profile_trace": [row]} for each feasible search candidate, then "best_found" and "final_refit" with {"profile_estimate": ...}.