Language: English
Last updated: 2026-08-02 This page: Model documentation
Switch: Chinese
Language switch: Chinese
GeneralizedLinearModel provides the common GLM entry point for Gaussian, binomial, and Poisson families. PenalizedGeneralizedLinearModel and its typed wrappers add L1, L2, ElasticNet, group, and adaptive-penalty hooks while keeping the public API explicit.
Use typed penalized estimators for regularized models:
from statgpu.linear_model import (
GeneralizedLinearModel,
PoissonRegression,
PenalizedLinearRegression,
PenalizedLogisticRegression,
PenalizedPoissonRegression,
)Ridge, Lasso, and ElasticNet are sklearn-style thin wrappers over penalized Gaussian regression.
statgpu.linear_model.GeneralizedLinearModelstatgpu.linear_model.PoissonRegressionstatgpu.linear_model.PenalizedGeneralizedLinearModelstatgpu.linear_model.PenalizedLinearRegressionstatgpu.linear_model.PenalizedLogisticRegressionstatgpu.linear_model.PenalizedPoissonRegressionstatgpu.linear_model.Ridgestatgpu.linear_model.Lassostatgpu.linear_model.ElasticNet- Internal GLM core:
statgpu.glm_core
Ordinary GLM fits minimize the average negative log-likelihood for the selected family:
Penalized GLM adds a penalty term:
The intercept is not penalized. statgpu.glm_core is intentionally GLM-specific; Cox partial likelihood, panel objectives, time-series likelihoods, and zero-inflated composite likelihoods should use future objective layers rather than being forced into glm_core.
Smooth GLMs solve score equations through IRLS/Newton/L-BFGS-style updates when available. Non-smooth penalized objectives use proximal/KKT-style optimization through FISTA.
Current solver behavior:
| Setting | solver="auto" behavior |
|---|---|
PenalizedLinearRegression(penalty="l2") |
solver="exact" closed-form L2 path |
| `PenalizedLinearRegression(penalty="l1" | "elasticnet")` |
PenalizedLogisticRegression(penalty="l2") on NumPy/CPU |
IRLS |
PenalizedPoissonRegression(penalty="l2") on NumPy/CPU |
IRLS |
PenalizedLogisticRegression(penalty="l2") on CuPy/Torch GPU |
FISTA |
PenalizedPoissonRegression(penalty="l2") on CuPy/Torch GPU |
FISTA |
explicit solver="irls" |
backend-native IRLS on NumPy/CuPy/Torch |
explicit solver="newton" |
backend-native Newton on smooth objectives |
explicit solver="lbfgs" |
backend-native L-BFGS on smooth objectives (all GLM families + L2/ElasticNet) |
non-smooth penalty with solver="newton" or solver="lbfgs" |
raises ValueError |
Important device rule: explicit device="cuda" stays on CuPy, explicit device="torch" stays on Torch CUDA, and explicit solver choices do not silently fall back to CPU. Formula parsing may run on CPU, but the core fit/predict path is converted to the selected backend.
This page covers estimation-first GLM and penalized GLM APIs. Full strict inference parity is not yet exposed for the new penalized GLM layer. Existing inference-rich estimators remain documented separately:
LinearRegression: classical, HC0-HC3, HAC.Ridge: classical, HC0-HC3, HAC.LogisticRegression: classical, HC0-HC3, HAC.Lasso: OLS-style and bootstrap inference paths.
Future GLM inference work should align with the project-wide strict inference gate before release.
| Parameter | Default | Description |
|---|---|---|
family |
model-specific | GLM family, for example "gaussian", "binomial", or "poisson" |
penalty |
"l2" or model-specific |
none, l1, l2, elasticnet, and reserved structured penalties |
alpha |
1.0 or model-specific |
Penalty strength in statgpu objective scale |
l1_ratio |
None |
ElasticNet mixing parameter |
fit_intercept |
True |
Whether to fit an intercept |
solver |
"auto" |
Solver dispatch; see estimating-equation section |
device |
"auto" |
cpu, cuda, torch, or auto depending on estimator support |
max_iter |
model-specific | Maximum optimizer iterations |
tol |
model-specific | Convergence tolerance |
formula |
None |
Optional patsy-style formula used with data |
data |
None |
DataFrame used with formula |
Alpha scaling is explicit. Do not compare same-named parameters across frameworks without conversion:
- Ridge:
sklearn_alpha = n_samples * statgpu_alpha - Logistic L2:
sklearn_C = 1 / (n_samples * statgpu_alpha) - Poisson L2: align against sklearn
PoissonRegressor(alpha=...) - Poisson L1/ElasticNet: align against statsmodels
fit_regularized
from statgpu.linear_model import GeneralizedLinearModel, PenalizedLogisticRegression
# Ordinary Poisson GLM on GPU when the selected path supports it.
glm = GeneralizedLinearModel(family="poisson", device="cuda")
glm.fit(X, y_count)
# CPU L2 logistic path: auto selects IRLS.
logit_cpu = PenalizedLogisticRegression(
penalty="l2",
alpha=0.01,
solver="auto",
device="cpu",
)
logit_cpu.fit(X, y_binary)
# GPU L2 logistic path: auto selects GPU-capable FISTA.
logit_gpu = PenalizedLogisticRegression(
penalty="l2",
alpha=0.01,
solver="auto",
device="cuda",
)
logit_gpu.fit(X, y_binary)Formula support is optional:
pip install statgpu[formula]from statgpu.linear_model import LinearRegression, PenalizedPoissonRegression
lm = LinearRegression()
lm.fit(formula="y ~ x1 + x2 + C(group)", data=df)
pred = lm.predict(df_new)
pois = PenalizedPoissonRegression(penalty="l2", alpha=0.01)
pois.fit(formula="count ~ exposure + x1", data=df)Formula parsing runs on CPU and is intended as a convenience layer. For very large data, pass explicit X, y arrays.
For the current GLM refactor, strict numerical validation is performed through remote CPU/GPU accuracy and external-framework comparison scripts. The new penalized GLM layer does not yet expose strict inference outputs such as robust standard errors and confidence intervals.
solver="auto" is device-aware for penalized GLMs. It picks exact Ridge for Gaussian L2, IRLS for smooth CPU logistic/poisson L2, and FISTA for CuPy/Torch GPU logistic/poisson L2. Explicit irls, newton, and lbfgs run on the selected backend when mathematically valid.
PenalizedGLM_CV defaults to cv_strategy="strict". In strict mode every fold/alpha is evaluated with the requested max_iter and tol, and GPU optimizations are limited to caching, fused kernels, and batched validation-score transfers. The optional cv_strategy="two_stage" mode first screens the alpha grid with relaxed CV solves, then strictly refines the candidate alphas and performs a strict final refit. Because the screening step can change alpha ranking on close CV curves, two-stage mode emits ApproximateCVWarning unless acknowledge_approx=True is passed.
from statgpu.linear_model import PenalizedGLM_CV
# Default: strict CV.
strict_cv = PenalizedGLM_CV(
loss="poisson",
penalty="elasticnet",
cv_strategy="strict",
device="cuda",
)
# Opt-in approximate screening, strict candidate refinement and final refit.
fast_cv = PenalizedGLM_CV(
loss="poisson",
penalty="elasticnet",
cv_strategy="two_stage",
acknowledge_approx=True,
refine_top_k=3,
device="cuda",
)PenalizedGLM_CV(loss="cox_ph") uses a separate survival path rather than the
scalar-response GLM scorer. Pass y as an (n_samples, 2) array with columns
[time, event]. L1, L2, ElasticNet, SCAD, and MCP are supported on NumPy,
CuPy CUDA, and Torch CUDA. The path:
- preserves the two-column target and never fits an intercept;
- scores each held-out fold with unpenalized negative Cox partial likelihood per row;
- selects an alpha only when every evaluable fold supplies finite evidence;
- hard-fails without publishing fitted state when no alpha is supported; and
- refits
PenalizedCoxPHModelwithcompute_inference=False.
survival_y = np.column_stack([time, event])
cox_cv = PenalizedGLM_CV(
loss="cox_ph",
penalty="scad", # l1, l2, elasticnet, scad, or mcp
alpha_grid=[0.1, 0.03, 0.01],
cv=5,
cv_strategy="strict",
loss_kwargs={"ties": "efron"},
device="cpu", # or "cuda" / "torch"
).fit(X, survival_y)cv_strategy="two_stage", sample_weight, dictionary targets, and
post-selection coefficient inference are not supported for this Cox branch.
cv_results_ records per-fold losses, valid-evidence counts, event counts,
failure reasons, the tie method, and the final-refit class.
Common fitted attributes and methods include:
coef_intercept_n_iter_when exposed by the selected solverfitpredictpredict_probafor logistic modelsscorewhere implementedcv_results_forPenalizedGLM_CV, includingcv_strategy_,cv_selected_device_,refined_mask, and stage-1 scores when two-stage screening is enabled
Future unified result objects are reserved for later work and are not part of this page's public contract.
- Solver × Penalty Compatibility Matrix — full dispatch table for loss × penalty × solver combinations, CV fast paths, and inference support status.
- Why is
statgpu.lossesnot kept as a compatibility namespace? The uncommittedlosseslayer was GLM-specific, so it was renamed toglm_coreto avoid implying a project-wide objective system. - Does
device="cuda"force GPU for every GLM solver? Yes for supported GLM solver paths: CuPy is used for the core computation, or a clear error is raised. There is no silent CPU fallback for explicit CUDA/Torch requests. - Should I use formula on large GPU workloads? Usually no. Formula parsing is CPU-side convenience; use explicit arrays for large-scale GPU jobs.
- Are
Ridge,Lasso, andElasticNetaliases? No. They are thin wrappers so sklearn-style constructor behavior can remain clear.
Local checks cover imports and smoke tests only. Accuracy, runtime, GPU behavior, and external-framework comparisons run on the remote myconda environment.
v23c full matrix benchmark (2026-05-20): 1043/1043 ALL PASS across 7 families x 10 penalties x 3 scales x 3 backends, validated against sklearn and statsmodels. See dev/tests/_bench_v23c_report.md and dev/tests/_bench_full_matrix.py.
Validation coverage includes:
- CPU/CuPy/Torch coefficient and intercept differences.
- Objective gap and KKT residual checks for penalized paths.
- Gaussian penalized comparison against sklearn Ridge/Lasso/ElasticNet.
- Logistic comparison against sklearn.
- Poisson L2 comparison against sklearn.
- Poisson L1/ElasticNet comparison against statsmodels
fit_regularized. - Runtime benchmarks with warm-up and GPU synchronization.
Remote credentials must be supplied through environment variables and must not be committed.
- McCullagh, P., & Nelder, J. A. (1989). Generalized Linear Models (2nd ed.). Chapman & Hall/CRC.
- Hastie, T., Tibshirani, R., & Friedman, J. (2009). The Elements of Statistical Learning (2nd ed.). Springer.
- Friedman, J., Hastie, T., & Tibshirani, R. (2010). Regularization paths for generalized linear models via coordinate descent. Journal of Statistical Software, 33(1), 1-22. https://doi.org/10.18637/jss.v033.i01
- scikit-learn linear models documentation: https://scikit-learn.org/stable/modules/linear_model.html
- statsmodels GLM documentation: https://www.statsmodels.org/stable/glm.html