Predict#
predict() returns the conditional mean per row on
the response scale the model was fitted to, for a built-in family; a custom
family defines its own default prediction quantity.
predict_parameters() returns every fitted
parameter on its natural scale, one column per name in
parameter_names_, and
predict_link() the same columns on the link scale,
offsets included. predict_cdf() and
predict_quantile() evaluate the fitted law per
row. For uncertainty by simulation,
posterior_draws() draws coefficients from the
fit’s Bayesian posterior, posterior_bounds()
pushes them through a parameter, a predictive quantile, an exceedance
probability or an expected shortfall to give per-row intervals, and
posterior_predictive() simulates responses.
Return the conditional mean for each row with a built-in family. |
|
Return every fitted distribution parameter as a DataFrame. |
|
Return each parameter's linear predictor as a DataFrame. |
|
Return |
|
Return the |
|
Simulate responses for the rows of |
|
Return coefficient draws from the Bayesian posterior of the fit. |
|
Return per-row bounds on a quantity derived from the fitted parameters. |
Example#
The same simulated severity book as the inference page: 3,000 claim amounts whose log-normal location peaks near age 30 and shifts by region, and whose scale widens with age.
import numpy as np
import pandas as pd
from superglm import LogNormalLS, SuperLSS, cat, s
rng = np.random.default_rng(1)
n = 3000
book = pd.DataFrame(
{
"age": rng.uniform(18, 80, n),
"region": rng.choice(["North", "South", "East", "West"], n),
}
)
age = book["age"].to_numpy()
by_region = book["region"].map(
{"North": 0.0, "South": 0.25, "East": -0.15, "West": 0.1}
).to_numpy()
location = 7.0 + 0.9 * np.exp(-(((age - 30) / 14) ** 2)) + by_region
scale = 0.35 + 0.006 * (age - 18)
amount = np.exp(rng.normal(location, scale))
family = LogNormalLS(parametrisation="location")
model = SuperLSS(
family,
family.location(s("age", kind="cr", k=8), cat("region")),
family.scale(s("age", kind="cr", k=6)),
).fit_reml(book, amount, outer="efs+newton")
model.smoothing_certified_
True
predict_parameters gives the fitted law row by row, one column per name in
parameter_names_; predict_link gives the same two columns on their link
scales.
parameters = model.predict_parameters(book)
links = model.predict_link(book)
parameters.head(3).round(3).join(links.head(3).round(3), rsuffix="_link")
| location | scale | location_link | scale_link | |
|---|---|---|---|---|
| 0 | 6.993 | 0.533 | 6.993 | -0.648 |
| 1 | 6.955 | 0.663 | 6.955 | -0.426 |
| 2 | 7.944 | 0.415 | 7.944 | -0.904 |
Under this parametrisation location is the mean of log amount and scale
its standard deviation. The location link is the identity, so location_link
repeats it; the scale link is log(scale - 0.01), which is why 0.533 appears
as -0.648. Both columns move from row to row, which is the point of the
distributional fit: rows differ in spread as well as in level.
predict, predict_quantile and predict_cdf on the first five rows: the
conditional mean, the 90th percentile of the claim, and the observed claim
read back through its own fitted law.
first = book.head(5)
pd.DataFrame(
{
"age": first["age"].round(1),
"region": first["region"],
"mean": model.predict(first),
"q90": model.predict_quantile(first, 0.9),
"observed": amount[:5],
"pit": model.predict_cdf(first, amount[:5]),
}
).round(3)
| age | region | mean | q90 | observed | pit | |
|---|---|---|---|---|---|---|
| 0 | 49.7 | East | 1255.804 | 2157.383 | 609.814 | 0.138 |
| 1 | 76.9 | North | 1306.096 | 2452.503 | 577.417 | 0.184 |
| 2 | 26.9 | West | 3071.324 | 4796.758 | 3853.757 | 0.775 |
| 3 | 76.8 | West | 1465.575 | 2750.967 | 875.653 | 0.328 |
| 4 | 37.3 | South | 3080.704 | 5045.277 | 3200.296 | 0.624 |
predict is the conditional mean of the claim amount, not of its logarithm,
so it sits above the median of a right-skewed law. predict_quantile answers
the price question instead: the 90th percentile of what this policy would
claim, between 1.55 and 1.9 times the mean on these rows. predict_cdf reads
the observed claim back through its own fitted law — the five values are spread
across the unit interval, as they should be for rows that are neither
systematically over- nor under-predicted.
posterior_predictive simulates responses for those same rows, drawing
coefficients from the fit’s posterior and then a response from each drawn
law, so the interval carries both parameter uncertainty and the claim’s own
randomness.
draws = model.posterior_predictive(first, n_draws=500)
pd.DataFrame(
np.quantile(draws, [0.05, 0.5, 0.95], axis=0).T,
columns=["p05", "p50", "p95"],
).round(0)
| p05 | p50 | p95 | |
|---|---|---|---|
| 0 | 484.0 | 1073.0 | 2589.0 |
| 1 | 341.0 | 1050.0 | 3313.0 |
| 2 | 1498.0 | 2917.0 | 5829.0 |
| 3 | 401.0 | 1184.0 | 3653.0 |
| 4 | 1358.0 | 3025.0 | 6490.0 |
Each row’s 5% to 95% span covers a factor of roughly four to ten — the spread of a single claim dwarfs the uncertainty in where its law sits.