Predict#

predict() returns the conditional mean per row on the response scale the model was fitted to, for a built-in family; a custom family defines its own default prediction quantity. predict_parameters() returns every fitted parameter on its natural scale, one column per name in parameter_names_, and predict_link() the same columns on the link scale, offsets included. predict_cdf() and predict_quantile() evaluate the fitted law per row. For uncertainty by simulation, posterior_draws() draws coefficients from the fit’s Bayesian posterior, posterior_bounds() pushes them through a parameter, a predictive quantile, an exceedance probability or an expected shortfall to give per-row intervals, and posterior_predictive() simulates responses.

predict

Return the conditional mean for each row with a built-in family.

predict_parameters

Return every fitted distribution parameter as a DataFrame.

predict_link

Return each parameter's linear predictor as a DataFrame.

predict_cdf

Return P(Y <= y) per row; y is a scalar or one value per row.

predict_quantile

Return the p-quantile per row; p is a scalar or one value per row in (0, 1).

posterior_predictive

Simulate responses for the rows of X through the family's quantile.

posterior_draws

Return coefficient draws from the Bayesian posterior of the fit.

posterior_bounds

Return per-row bounds on a quantity derived from the fitted parameters.

Example#

The same simulated severity book as the inference page: 3,000 claim amounts whose log-normal location peaks near age 30 and shifts by region, and whose scale widens with age.

import numpy as np
import pandas as pd

from superglm import LogNormalLS, SuperLSS, cat, s

rng = np.random.default_rng(1)
n = 3000
book = pd.DataFrame(
    {
        "age": rng.uniform(18, 80, n),
        "region": rng.choice(["North", "South", "East", "West"], n),
    }
)
age = book["age"].to_numpy()
by_region = book["region"].map(
    {"North": 0.0, "South": 0.25, "East": -0.15, "West": 0.1}
).to_numpy()
location = 7.0 + 0.9 * np.exp(-(((age - 30) / 14) ** 2)) + by_region
scale = 0.35 + 0.006 * (age - 18)
amount = np.exp(rng.normal(location, scale))

family = LogNormalLS(parametrisation="location")
model = SuperLSS(
    family,
    family.location(s("age", kind="cr", k=8), cat("region")),
    family.scale(s("age", kind="cr", k=6)),
).fit_reml(book, amount, outer="efs+newton")
model.smoothing_certified_
True

predict_parameters gives the fitted law row by row, one column per name in parameter_names_; predict_link gives the same two columns on their link scales.

parameters = model.predict_parameters(book)
links = model.predict_link(book)
parameters.head(3).round(3).join(links.head(3).round(3), rsuffix="_link")
location scale location_link scale_link
0 6.993 0.533 6.993 -0.648
1 6.955 0.663 6.955 -0.426
2 7.944 0.415 7.944 -0.904

Under this parametrisation location is the mean of log amount and scale its standard deviation. The location link is the identity, so location_link repeats it; the scale link is log(scale - 0.01), which is why 0.533 appears as -0.648. Both columns move from row to row, which is the point of the distributional fit: rows differ in spread as well as in level.

predict, predict_quantile and predict_cdf on the first five rows: the conditional mean, the 90th percentile of the claim, and the observed claim read back through its own fitted law.

first = book.head(5)
pd.DataFrame(
    {
        "age": first["age"].round(1),
        "region": first["region"],
        "mean": model.predict(first),
        "q90": model.predict_quantile(first, 0.9),
        "observed": amount[:5],
        "pit": model.predict_cdf(first, amount[:5]),
    }
).round(3)
age region mean q90 observed pit
0 49.7 East 1255.804 2157.383 609.814 0.138
1 76.9 North 1306.096 2452.503 577.417 0.184
2 26.9 West 3071.324 4796.758 3853.757 0.775
3 76.8 West 1465.575 2750.967 875.653 0.328
4 37.3 South 3080.704 5045.277 3200.296 0.624

predict is the conditional mean of the claim amount, not of its logarithm, so it sits above the median of a right-skewed law. predict_quantile answers the price question instead: the 90th percentile of what this policy would claim, between 1.55 and 1.9 times the mean on these rows. predict_cdf reads the observed claim back through its own fitted law — the five values are spread across the unit interval, as they should be for rows that are neither systematically over- nor under-predicted.

posterior_predictive simulates responses for those same rows, drawing coefficients from the fit’s posterior and then a response from each drawn law, so the interval carries both parameter uncertainty and the claim’s own randomness.

draws = model.posterior_predictive(first, n_draws=500)
pd.DataFrame(
    np.quantile(draws, [0.05, 0.5, 0.95], axis=0).T,
    columns=["p05", "p50", "p95"],
).round(0)
p05 p50 p95
0 484.0 1073.0 2589.0
1 341.0 1050.0 3313.0
2 1498.0 2917.0 5829.0
3 401.0 1184.0 3653.0
4 1358.0 3025.0 6490.0

Each row’s 5% to 95% span covers a factor of roughly four to ten — the spread of a single claim dwarfs the uncertainty in where its law sits.