语言:中文
最后更新:2026-08-08
页面定位:模型文档
切换:English
statgpu.panel 提供六类面板数据估计器:
PanelOLS:个体和/或时间固定效应;RandomEffects:Swamy-Arora 可行 GLS 随机效应;PooledOLS:不去均值的堆叠 OLS;BetweenOLS:在个体均值上回归;FirstDifferenceOLS:个体内一阶差分;FamaMacBeth:逐期横截面回归后对系数取平均。
数组输入的数值路径支持 NumPy、CuPy CUDA 与 Torch CUDA。formula 构造以及字符串/分类 entity、time、cluster 标签属于明确的 CPU 元数据边界,只会把对齐后的紧凑编码传入数值后端。显式 GPU device 不会静默回退 CPU。
Tier-1 Panel 路线的 Stage B 在不改变 Stage-A 系数、预测、协方差归一化和 legacy inference 契约的前提下,新增参数型 fit statistics 和三类结构化 specification test。
from statgpu.panel import (
PanelOLS,
RandomEffects,
PooledOLS,
BetweenOLS,
FirstDifferenceOLS,
FamaMacBeth,
PanelTestResult,
PanelFitStatistics,
hausman_test,
pooling_f_test,
breusch_pagan_lm_test,
clustered_covariance,
two_way_clustered_covariance,
hac_covariance,
)诊断结果类和三类检验函数也从顶层 statgpu 导出。
| 模型 | 变换 | 主要推断选项 | Stage-B fit statistics |
|---|---|---|---|
PanelOLS |
个体/时间组内变换 | nonrobust、HC1 robust、clustered | within/between/overall R²、adjusted R²、classical model F、pooling F |
RandomEffects |
Swamy-Arora 可行 GLS | nonrobust | within/between/overall R²、adjusted R²、classical model F、Hausman 输入 |
PooledOLS |
堆叠 OLS | nonrobust、robust、clustered、HAC | overall R² 始终可用;提供 entity_ids 后增加 within/between R² 和 BP-LM;另有 adjusted R² 与 classical model F |
BetweenOLS |
个体均值 | nonrobust、robust | within/between/overall R²、adjusted R²、classical model F |
FirstDifferenceOLS |
个体内一阶差分 | nonrobust、robust | within/between/overall R²、差分拟合空间的 adjusted R²、classical model F |
FamaMacBeth |
逐期横截面回归 | nonrobust、Newey-West | 参数型 within/between/overall R²;不定义 residual-OLS adjusted R² 或 model F |
PanelOLS 在移除指定固定效应后拟合 OLS。仅使用个体效应时,
个体和时间双向固定效应会进一步移除时间均值并加回总体均值。
PooledOLS 拟合
其中 (X^+) 在需要时表示 Moore-Penrose 伪逆。BetweenOLS 对个体均值执行 OLS,FirstDifferenceOLS 对 (\Delta X) 和 (\Delta y) 执行 OLS,FamaMacBeth 对逐期系数向量求平均。
cov_type |
行为 |
|---|---|
"nonrobust" |
经典 OLS 协方差和 t 推断 |
"robust" |
HC1 sandwich 协方差和渐近正态推断 |
"clustered" |
在相应 estimator 支持范围内使用聚类稳健协方差 |
"hac" |
PooledOLS 使用 Bartlett/Newey-West HAC |
"newey-west" |
对 FamaMacBeth 系数时间序列应用 HAC |
Stage B 不修改原有 bse_、tvalues_、pvalues_、conf_int_ 或 estimator-specific covariance 定义。RandomEffects robust covariance、HC0/HC2/HC3、Driscoll-Kraay 和扩展 cluster correction 仍属于后续 Stage C。
对 PooledOLS(cov_type="hac"),应向 fit 传入 time_index=。实现采用稳定时间排序,并让 X、y 和 Stage-B entity diagnostic metadata 使用同一个 permutation。因此 BP-LM 与参数型 R² 不会因为 HAC 排序后仍使用旧 entity metadata 而发生错位。
秩亏设计需要区分拟合空间与系数空间:
- 拟合、预测、残差、RSS、有效 rank 与拟合空间比较仍然有效;
df_resid使用nobs - rank(X),而不是nobs - n_columns;- 精确共线时单个系数不唯一;
- 因而系数级 covariance、BSE、test statistic、p-value 和 confidence interval 不应被解释为唯一识别的系数推断。
Stage-B model-F 的 restriction rank 使用有效数值 rank,而不是直接使用原始列数。
支持的拟合对象提供结构化 PanelFitStatistics:
stats = model.fit_statistics_
stats.rsquared_within
stats.rsquared_between
stats.rsquared_overall
stats.rsquared_adj
stats.f_statistic
stats.f_pvalue
stats.f_df
stats.metadatawithin/between/overall R² 采用与 linearmodels 对齐的 parameter-based 定义:各统计量直接评价同一个拟合系数向量,而不是使用 fitted-value correlation 的平方。
给定 (\hat\beta):
- overall R² 在 level panel 上评价 (y-X\hat\beta);
- between R² 在个体均值上评价 (\bar y_i-\bar X_i\hat\beta);
- within R² 在 entity-demeaned 的 y 和 X 上评价同一系数向量。
只有 level regressor design 中存在实际可识别的常数项时,overall 和 between total sum of squares 才中心化。固定效应本身不会自动改变这一规则。RandomEffects 会检测传入 level design 中显式的非零常数列,并在 adjusted R² 与 restricted model-F 中保留其 quasi-demeaned transformed column。常数检测的容差相对于该列自身量级定义,因此仅改变单位不会把一个非零常数误判成 slope,反之亦然。若 total sum of squares 为 0,Stage-B 标准字段返回 0.0,并在 metadata["degenerate_total_ss"] 中标记。
PanelOLS.rsquared_within 为兼容 Stage A 保持原样。双向 FE 下,它描述历史上的 entity+time 完整 transformed fit,因此可能与标准化的 entity-within fit_statistics_.rsquared_within 不同。Stage B 不覆盖旧属性,并在 fit-statistics metadata 中保留兼容值。
新的标准化 diagnostics 使用完整 fixed-effect nuisance space 的 rank。当前 PanelOLS 不保留 exogenous intercept,因此标准 nuisance-effect rank 为:
- 仅 entity effects:(N);
- 仅 time effects:(T);
- entity + time effects:(N+T-C),其中 (C) 是观测到的 entity-time incidence graph 的 connected-component 数。
因此常见的 (N+T-1) 只是连通面板 (C=1) 的特例;若 incomplete panel 的 incidence graph 不连通,diagnostic df 使用实际 dummy-space rank,而不会硬编码连通情形。
若 transformed slope design 的数值 rank 为 (r_X),Stage B 使用
以及
来计算标准 FE model F 与 adjusted R²。
这套 diagnostic df 与历史公开 PanelOLS.df_resid 明确分离;旧 df 继续服务已有 covariance、t statistic、p-value、confidence interval 与 summary。Hausman 使用一份仅供 diagnostic 的小型 FE covariance,并按标准 nuisance-rank denominator 重标度;Stage-A 公共 BSE/CI 不受影响。
对 OLS-style panel estimator,Stage B 报告主拟合空间上的 classical homoskedastic joint-slope F:
其中 (q) 是有效 restriction rank。即使 estimator 的 coefficient covariance 使用 robust 或 clustered 选项,这一字段也不会静默变成 robust Wald test。当 unrestricted regression 精确拟合而 restricted regression 的 RSS 为正时,标准化结果采用经典极限值 F=inf、p=0,而不是把统计量标成 unavailable。RSS 的 zero/nesting tolerance 按 RSS 自身量级缩放,因此将 response 与 fitted coefficient 同时乘一个单位变换常数不会改变无量纲 F statistic 或 applicability。
FamaMacBeth 不定义 residual-OLS model F 或 adjusted R²,因为其 covariance 来自逐期横截面系数时间序列;Stage B 不会把 beta-series inference 重命名成 residual OLS inference。
所有 specification test 都返回 PanelTestResult。在计量意义上不适用的情况返回 applicable=False 并给出 reason;真正的 API 编程错误仍正常抛出异常。
对至少包含一个固定效应的已拟合 PanelOLS:
result = fe.pooling_f_test()
# 或
result = pooling_f_test(fe)经典 poolability 原假设是所有纳入的固定效应联合为 0。统计量使用同一个对齐后的估计样本比较 fixed-effect model 与其 nested pooled regression。
当 FE level design 没有显式常数时,pooled null 会先从 y 和 X 中投影掉共同常数,再拟合 slope,避免把共同均值误计入被检验的 fixed effects。分子和分母自由度由有效 nested-model rank 推导,而不是硬编码 effect count。
如果 fixed-effects fit 精确拟合(RSS_FE 在数值上为 0),而 restricted pooled RSS 实质为正,则经典极限结果为 F=inf、p=0;若 pooled 与 FE RSS 都为 0,则比值不定,返回结构化 inapplicable。以上 zero/nesting 判定使用相对于 RSS 量级的 tolerance,不使用固定的 unit-size floor。
即使 FE 对象的系数推断采用 robust 或 clustered covariance,这个 pooling F 仍是 classical/homoskedastic test。
向 PooledOLS.fit() 提供 entity_ids:
pooled = PooledOLS().fit(X, y, entity_ids=entity_ids)
result = pooled.breusch_pagan_lm_test()
# 或
result = breusch_pagan_lm_test(pooled)这里是 panel error-components Breusch-Pagan LM test,不是横截面异方差的 Breusch-Pagan test。Stage B 实现 one-way entity 版本,并包含 plm::plmtest(type="bp", effect="individual") 使用的 Baltagi-Li incomplete/unbalanced-panel 公式。
原假设是 entity random-effect variance 为 0。至少需要两个 entity、正的 pooled RSS,并且至少一个 entity 有重复观测。没有 entity_ids 时返回结构化 inapplicable 结果,不会根据行顺序猜测 panel structure。
fe = PanelOLS(entity_effects=True, cov_type="nonrobust").fit(
X, y, entity_ids=entity_ids
)
re = RandomEffects().fit(X, y, entity_ids=entity_ids)
result = fe.hausman_test(re)
# 或
result = hausman_test(fe, re)Stage B 实现 one-way entity FE-versus-RE 的经典二次型 Hausman:
适用性规则是显式的:
- FE 必须只有 one-way entity effects;
- FE coefficient covariance 必须是 classical/nonrobust;
- FE 与 RE 必须来自同一个对齐后的 y/entity 样本和同一个 canonical slope design;
- RE 可以比 FE 多一个显式常数,因为 entity FE 已吸收共同 intercept;该常数不会进入 Hausman coefficient vector;
- 行/样本一致性使用所有对齐 float64 slope-X/y 值的 collision-resistant SHA-256 digest,并结合 entity-code signature 与 canonical feature metadata,而不是只比较 shape 或低阶 moments;
- canonical slope position 会保留到各模型原始 coefficient/covariance position 的映射,因此 RE 的第 0 列 intercept 不会把
x1、x2错配到错误系数; - covariance difference 若实质上 indefinite,则返回 inapplicable,不通过 eigenvalue clipping 强行制造统计量。
数组输入下,在移除 RE-only constant 后 slope 会重新按 x1、x2、... 做 canonical 编号;named/formula design 则保留 slope 名称。Hausman covariance rank 与 identified-range tolerance 按 covariance/coefficient 自身量级缩放,所以改变 outcome 单位不会改变同一数学问题的 applicability。
GPU 拟合下,canonical slope full-content digest 仅为了 hashing 而通过有界 chunk 分批复制到 host。只有可进入 Stage-B Hausman 的 one-way entity nonrobust FE 与 RandomEffects 保留该 identity;robust/clustered、time-only、two-way FE 在 identity compare 前就会被判不适用,因此不承担完整 X/y hashing 开销。拟合对象只保存 digest/index metadata,不保留第二份完整 CPU design copy。统计估计、covariance 构造与 fit-statistic reduction 仍在所选数值 backend 上完成。
若 covariance difference 为 positive semidefinite 但 rank deficient,statgpu 提供显式记录的 generalized-inverse extension:只在 coefficient difference 位于 identified range 内时计算,并使用 numerical rank 作为 chi-square df。metadata 会记录 used_pinv=True 和 singular PSD generalized-inverse Hausman。Stage B 不实现 robust auxiliary-regression Hausman。
PanelOLS(
entity_effects=False,
time_effects=False,
cov_type="nonrobust",
device="auto",
)model.fit(X, y, entity_ids=entity_ids, time_ids=time_ids, cluster=cluster)formula 输入也可使用已有 pipe syntax,例如 "y ~ x1 + x2 | entity"。formula 的 missing-row filtering 会同步对齐 observation-level side arrays。
PooledOLS(
cov_type="nonrobust",
alpha=0.05,
bandwidth=None,
kernel="bartlett",
device="auto",
)model.fit(
X,
y,
cluster=None,
time_index=None,
entity_ids=None,
)clustered inference 需要 cluster。time_index 定义 HAC 的稳定时间排序。entity_ids 是可选项,不改变系数;提供后会启用标准化 within/between R² 与 panel BP-LM。
RandomEffects(device="auto")
BetweenOLS(cov_type="nonrobust", alpha=0.05, device="auto")
FirstDifferenceOLS(cov_type="nonrobust", alpha=0.05, device="auto")
FamaMacBeth(
cov_type="newey-west",
bandwidth=None,
alpha=0.05,
min_obs_per_period=1,
device="auto",
)FamaMacBeth.fit(..., entity_ids=None) 中的可选 entity IDs 只用于 Stage-B within/between R²;beta-series estimation 与 covariance path 保持不变。
import numpy as np
from statgpu.panel import PanelOLS, PooledOLS, RandomEffects
n_entities, n_times = 50, 10
n = n_entities * n_times
entity_ids = np.repeat(np.arange(n_entities), n_times)
time_ids = np.tile(np.arange(n_times), n_entities)
X = np.random.default_rng(0).normal(size=(n, 3))
y = X @ np.array([1.0, -0.5, 0.3]) + np.random.default_rng(1).normal(size=n) * 0.1
fe = PanelOLS(entity_effects=True, device="cpu").fit(
X, y, entity_ids=entity_ids
)
print(fe.fit_statistics_.rsquared_within)
print(fe.pooling_f_test())
pooled = PooledOLS(device="cpu").fit(X, y, entity_ids=entity_ids)
print(pooled.breusch_pagan_lm_test())
re = RandomEffects(device="cpu").fit(X, y, entity_ids=entity_ids)
print(fe.hausman_test(re))CuPy CUDA 使用 device="cuda";Torch CUDA 使用 CUDA tensor 并设置 device="torch"。Stage-B 的统计变换与 sufficient-statistic accumulation 跟随所选数值 backend;formula/label metadata、最终 scalar 与小型 covariance matrix 使用 CPU metadata boundary;Hausman-compatible one-way FE/RE 拟合还会仅为 collision-resistant identity hashing 对 canonical slope X/y 做有界分块 host copy。
常见拟合属性包括:
coef_;- 在系数空间可识别时的
bse_、tvalues_、pvalues_、conf_int_; - 历史兼容的
rsquared或rsquared_within; - 标准化的
fit_statistics_; nobs、df_resid与有效 rank;FamaMacBeth的betas_、cov_params_与n_periods。
PanelTestResult 包含 statistic、pvalue、distribution、df、null、alternative、applicable、reason 和 metadata。
formula evaluation 可能因为 missing value 删除行。entity、time、cluster 等 side array 会与保留行同步对齐。字符串和分类标签在 CPU 上 factorize;数值变换与 sufficient-statistic calculation 继续留在所选 backend。对 Hausman-compatible 拟合,会先排除 RE-only 显式常数,再把 canonical slope X/y 以有界 chunk 复制到 host 计算 SHA-256 identity;随后只保存 digest、原始 coefficient-index 映射、feature/entity metadata 与小型 covariance matrix,不保存第二份完整 CPU design copy。
Stage A / PR #119 建立共享 Panel framework,并在 Tesla P100 上通过 10 个 CuPy + 10 个 Torch exact-head physical cases。
Stage B 增加 maintained analytic/fitted-model regression tests、formula/missing-row alignment、Python 3.9 + Torch 2.0 CPU parity,以及可执行的 linearmodels==7.0 external-definition gate。physical runner 每个 backend 包含 17 个 estimator cases 与 4 个 Hausman diagnostic cases,其中包括 balanced/unbalanced 显式常数 RandomEffects,以及 FE 吸收 intercept、RE 显式估计 intercept 的 Hausman 参数化。最终 promotion 要求 dev/benchmarks/validate_panel_stage_b_gpu.py 在 clean exact commit 上同时通过 CuPy 与 Torch CUDA;该 runner 是 correctness/provenance gate,而不是性能 benchmark。另有独立 physical benchmark 用于测量 Hausman-compatible FE/RE 上剩余 full-content identity 开销。
- Hausman, J. A. (1978). Specification Tests in Econometrics.
- Breusch, T. S., & Pagan, A. R. (1980). The Lagrange Multiplier Test and its Applications to Model Specification in Econometrics.
- Baltagi, B. H., & Li, Q. (1990). A Lagrange Multiplier Test for the Error Components Model with Incomplete Panels.
- White, H. (1980). A heteroskedasticity-consistent covariance matrix estimator.
- Newey, W. K., & West, K. D. (1987). A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix.
- Fama, E. F., & MacBeth, J. D. (1973). Risk, return, and equilibrium.
- Cameron, A. C., Gelbach, J. B., & Miller, D. L. (2011). Robust inference with multiway clustering.