Skip to content

Commit 084b0b5

Browse files
committed
feat(genrm): GenRM 弹性扩缩容与 Autoscaler 接入
关联 #351。官方验收标准五项全部达成(证据矩阵见 PR 描述)。 单 commit 汇总全部开发提交;完整历史在 feat/genrm-elastic-scaling-history 分支(证据文件中的 hash 引用指向该分支)。 扩缩 API 与状态机(relax/utils/genrm_scale_registry.py,纯标准库): - POST /genrm/scale_out|scale_in,num_replicas 为目标绝对总数 - 同键幂等重放(verbatim NOOP replay);按模型互斥,未决清理同样阻塞 - 逐操作状态查询与 reconcile 端点 Manager 生命周期(relax/distributed/ray/genrm.py、multi_engine_manager.py): - 每副本独占 placement group,创建即登记 ownership,所有失败出口 (就绪超时、InfoActor 创建/探测/kill、abort)落到确认 REMOVED 或 pending cleanup,扩缩互斥保持到 reconcile 确认 - 健康检查通过后才发布路由;超时后迟到的结果不再发布 - 缩容最新优先选 victim,初始引擎受保护;drain 完成才销毁 - PG 释放以 Ray 确认 REMOVED 为准(remove_placement_group() 返回不算) - 物理完成栅栏:终态仅在 lifecycle 线程停止后计入 - multi_engine_manager +55 行:REMOVED 确认轮询与 pending 清理重试, 落在 GenRM/Teacher 共享骨架(Rollout 侧零改动,Task 3 融合点) Autoscaler 服务隔离(relax/utils/autoscaler/): - per-service runtime:独立采集器/决策引擎/策略/历史 - ServiceScalingPolicy 支持 GenRM 独立阈值,rollout 配置向后兼容 - /conditions、/scale_history、/metrics_history 支持 service= 过滤 - 指标逐字段有效性校验;unknown_engines 冻结 scale-in(健康路径与 旧实现一致,异常路径更保守,见 PR 描述"需维护者确认的共享语义变更") 可观测性: - /metrics 上报 Manager 实时容量 - TUI monitor --service genrm 视图与无头 SVG 截图 测试与 GPU 证据(demos/task4_genrm/EVIDENCE.md,全部 hash pin 可复算): - 96 个任务相关 CPU 测试(含 PG ownership 故障注入,修复前验证为红) - 手动 1→2→1:4,163 请求 0 失败,贪心前缀跨阶段一致 - Autoscaler 全周期:3,202 请求 0 失败;冻结阈值轮 r2 12/14(A7/B3 定性为验收口径缺陷:Ray 2.58 保留 REMOVED 墓碑)→ r3 以 v2 预注册 口径(非终态 PG 计数,阈值逐字冻结)14/14 全绿 - 奖励一致性:greedy 跨引擎逐输入一致;官方采样 4% 不稳定率定性为 逐引擎种子配置属性 - 故障注入:607.3s fail-closed 停等后 reconcile 清理确认;SIGKILL victim 后在飞请求全部重试到初始引擎 - 最小训练回路(dapo-genrm):step 训练、权重同步 200 OK ×9、judge 调用真实发生(judge_response 落盘) - 训练连续性(官方验收第 5 项):真实配方一次运行内 scale_out→ACTIVE (55s)→scale_in→COMPLETED(1s 排空),双窗口内训练事件无 >120s 停滞、弹性引擎真实服务奖励、8/8 rollouts 跑完、零错误 - B2 冒烟与连续性配方 scripts/training/genrm/run-qwen3-0.6B-4xgpu-*.sh
1 parent 6d78410 commit 084b0b5

68 files changed

Lines changed: 17562 additions & 151 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎demos/task4_genrm/EVIDENCE.md‎

Lines changed: 106 additions & 0 deletions
Large diffs are not rendered by default.

‎demos/task4_genrm/README.md‎

Lines changed: 20 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,20 @@
1+
# Task 4 contract demo
2+
3+
Run from this directory:
4+
5+
```bash
6+
python -m unittest -v test_contract_demo.py
7+
python contract_demo.py
8+
python render_demo.py results/contract-demo.json --output results/contract-demo.html
9+
```
10+
11+
Open `results/contract-demo.html` in a browser (GitHub's `blob` view shows the source; download the file or use a local checkout to run it). The page replays four deterministic scenarios: normal 1→2→1, health-check failure, backend still busy during drain, and cleanup failure followed by explicit reconciliation. Scrub the timeline to inspect route membership, admitted requests, PG ownership and the recovery boundary. The JSON is the exact event trace used by the page.
12+
13+
This exercises the proposed request and lifecycle contract, including unknown-request `404`, absolute-target validation, idempotency-key replay (a keyed `NOOP` is recorded and replayed verbatim after capacity changes, never executing a new operation), `409` on key/body mismatch or unresolved cleanup, and explicit retry/reconcile paths. Engines, workers, health checks, admission, backend idleness and PGs are in-memory fakes; the `workers` field is a contract signal, not a real worker process. The demo has no Ray, SGLang, GPU, Autoscaler or real scoring. It does not establish that Task 3 supplies the required non-cancelling drain or full-worker cleanup, or that text training continues during scaling. Those remain integration and acceptance work for [RFC #351](https://github.com/redai-studio/Relax/issues/351).
14+
15+
## GPU E2E drivers
16+
17+
- `e2e_genrm_scale.py` — manual `1→2→1` under continuous scoring: probes physical GPUs for the scale-out PG, records per-phase scores, per-engine served counts, GPU snapshots and machine verdicts.
18+
- `e2e_autoscaler_load.py` — full autoscaler cycle under a `LOW→HIGH→STEADY→LOW'` load curve with per-service GenRM thresholds; samples capacity, decisions and history at ~1 Hz.
19+
20+
Both drivers need a single node with free GPUs and the SGLang runtime; they deploy only Ray Serve apps they own and are expected to clean up in an outer `finally`. Machine-verdict summaries of the recorded runs live under `results/`; see [`EVIDENCE.md`](EVIDENCE.md) for the run-to-commit mapping, known limits and hash pins. Full per-request logs and the failed intermediate runs are on the `evidence/task4-genrm` branch.

‎demos/task4_genrm/contract_demo.py‎

Lines changed: 383 additions & 0 deletions
Large diffs are not rendered by default.

0 commit comments

Comments
 (0)