Skip to content

【No.4】GenRM 弹性扩缩容与自动伸缩 - #370

Open
shanyulu wants to merge 44 commits into
redai-studio:mainfrom
shanyulu:feat/genrm-elastic-scaling
Open

shanyulu wants to merge 44 commits into
redai-studio:mainfrom
shanyulu:feat/genrm-elastic-scaling

Conversation

@shanyulu

@shanyulu shanyulu commented Sep 24, 2026 •

Copy link
Copy Markdown

改动目的

为冻结奖励模型 GenRM 增加弹性副本生命周期,并接入现有 Autoscaler。负载升高时,新引擎通过健康检查后接流量;负载回落时,排空弹性副本并回收资源,保留初始引擎。

设计与取舍见 RFC #351,字段与错误码见 中文 API / OpenAPI。本期范围为单 Gateway 排空、单 GPU 弹性副本。

当前状态 结果
PR / 产品版本 PR 头 da4acbb;产品 0481701,relax/ 产品代码保持不变
CI 前一头 a55c85f 8/8 通过:pre-commit、三版 Python、H20 4-GPU 单测、三项训练集成(含一项 NPU);da4acbb 仅文档变更,CI 重跑中
专用验收 生命周期、评分对照、自动扩缩、故障收尾及训练运行证据见下表;各自保留运行版本
尚待确认 当前头 approval;artifact 归属、验收规模、内容确定性采样三项裁决

产品头之后仅变更 API / 证据文档、OpenAPI 生成脚本及两处测试 lint 注释,运行时代码未变。

实现结构

flowchart TB
    U["手动扩缩请求"] --> C["GenRM 控制 API<br/>绝对目标 · 幂等 · 互斥"]
    C --> M["GenRMManager<br/>PG 所有权 · 生命周期"]
    M --> E["初始 / 弹性引擎<br/>健康检查后发布"]
    M -.->|"可路由集合"| G["GenRM Gateway<br/>admission · 在途请求计数"]
    T["训练评分请求"] --> G
    G --> E
    G -.->|"指定 generation 排空证明"| C
    E -.->|"有效负载指标"| A["分服务 Autoscaler<br/>持续条件 · 冷却 · history"]
    A --> C
    classDef control fill:#e8f0fe,stroke:#4169a1,color:#172b4d;
    classDef serving fill:#e5f4ee,stroke:#328268,color:#173d31;
    class U,C,M,A control;
    class T,G,E serving;
Loading

实线为调用,虚线为状态、发现或观测关系。GenRM 扩缩不引入训练权重同步;actor→rollout 原有同步仍然保留。

建议审查顺序

审查入口 关键改动 重点检查
Manager 创建 PG 时即登记 owner;候选就绪后发布;固定 victim 排空;弹性槽位永久退役 迟到初始化不发布;PG 未确认回收时不放行;recover() 不重建已缩掉的副本
组件 / Registry 绝对目标、幂等指纹、同模型互斥、generation 排空、状态查询与 reconcile 超时只是 abort;逻辑终态与 cleanup_required 分离;无法取得权威容量时拒绝提交
Autoscaler / 指标 Rollout 与 GenRM 各自维护 discovery、collector、decision engine、pending 和历史;应用 service_policies 缺失/过期不当作空闲;不同条件使用自己的有效分母;冷却状态不串服务
评分参数 默认采样 seed 由模型、token 输入和有效参数派生;显式 seed 优先 改变重复请求随机性,是待确认的产品契约,不只是内部修复
TUI / API 文档 --service genrm;按服务展示状态、条件、历史;动态容量 保留 Rollout 默认视图;降级快照不能冒充资源释放证明

失败路径的行为

  • 扩容终态为 ACTIVE / PARTIAL / FAILED;PARTIAL 保留已发布副本,不代表达到目标。
  • 缩容先关闭 admission,再等指定 generation 的在途请求清零,最后退役 worker 并确认 owned PG 为 REMOVED。
  • 超时后保持原操作互斥;原物理线程未结束时 reconcile 返回可重试冲突。清理成功只清除 cleanup_required,不把原 FAILED 改成成功。
  • 容量分别报告 current / ready / occupied / pending_cleanup:未发布候选不增加服务容量;失败缩容 victim 未释放前仍计入已发布容量与资源占用。

正常链路、超时收尾时序和容量示例见 RFC 架构说明。

验证

原始证据固定于 fdf288d,各实验的运行版本见下表。评分复确认与任务相关 CPU 回归在最终产品 0481701 上完成。

验收项 运行版本 结果 原始证据
手动扩缩 a2ca6cb 1→2→1;4,163 请求零失败;初始引擎保留、资源回归 生命周期运行
评分一致性 0481701 50 输入、800 条引擎归属回复;完整解析、零截断;greedy 一致,采样 0/50 翻转 最终代码复确认
不同请求历史 0481701 初始引擎先额外处理 300 次采样;再比较 50 输入、400 条归属回复,0/50 翻转 历史扰动对照
控制契约 0481701 278 项任务相关 CPU 回归通过;幂等、409、非法目标与初始保护 CPU 回归记录
自动扩缩 5c1e2e7 冻结断言 14/14;1,568 请求零失败;弹性引擎处理 181 请求;真空闲缩容与资源回收通过 最终协议轮次
故障收尾 e7224af 长请求排空、deadline-abort 后保持互斥并 reconcile、kill-victim 三场景通过 故障注入
训练继续完成 945741e 8/8 步执行完成,8/8 rollout 完成;日志记录 64 个保存样本;扩容窗有 rollout 进展 日志与再分析

训练结论采用再分析 v2.1。仓库旧索引的“全程零错误”和“全程停顿不超过 120 秒”表述已撤回:日志有三条非良性带时间戳错误记录,judge 错误行数与双副本稳定窗口内的错误行数均为 0;事件间隔不能作为全程停顿上界。

其余限制:

  • 训练完成不等于零性能影响。 最大进度间隔跨过扩缩边界;没有同配置、不扩缩的对照。缩容窗约 1 秒,无窗内训练事件,不声称该窗内持续更新。
  • 64 个样本来自保存日志;rollout ID 无缺漏/重复,但逐样本 JSONL 未保留,未做内容级完整性与去重核验。该训练轮的弹性副本贡献仅有一个计数器归属请求;跨副本评分由独立归属实验验证。
  • 0/50 只说明固定输入与配置下未观察到翻转,不证明模型判题能力或普遍确定性。PG 资源计数排除 REMOVED 墓碑,不能用原始表长度判泄漏。
真实验收截图:GenRM 的状态、扩缩条件与操作历史

冻结协议最终轮的 GenRM Autoscaler TUI

截图来自自动扩缩最终轮,固定于上述 evidence commit;用于检查观测界面,不替代整轮事件与 verdict。Rollout 对照截图与原始记录保留在同一目录。

兼容性与未覆盖范围

旧 Rollout 顶层字段和默认 TUI 行为保留。两服务的冷却计时状态独立,但冷却时长配置仍由 deployment 提供。GenRM 自动扩缩目标目前只支持单模型实例,发现多个活跃实例时拒绝混合处理。

单 Gateway 计数不覆盖 direct client / 跨 Gateway 的完整排空;组件重启后的持久幂等、Manager 重启后的在途恢复、多 GPU 弹性副本,以及弹性操作与运行时 onload/offload 的协调,均不在本期范围。

请维护者确认

  1. 仓库范围:driver、manifest、原始 artifacts 的保留与迁出边界。
  2. 验收规模:4×RTX 4090、Qwen3-0.6B 是否满足本期生命周期与路由验收。
  3. 评分语义:是否接受未显式指定 seed 时的内容确定性采样;若需要重复调用的随机多样性,应另行定义请求身份契约。

当前头仍需正式 review。不可变证据目录保留全部已归档轮次及各自判定。

@shanyulu
shanyulu force-pushed the feat/genrm-elastic-scaling branch 5 times, most recently from 084b0b5 to 8dcde5e Compare September 25, 2026 08:42
关联 redai-studio#351。官方验收标准五项全部达成(证据矩阵见 PR 描述)。
单 commit 汇总全部开发提交;完整历史在
feat/genrm-elastic-scaling-history 分支(证据文件中的 hash 引用指向该分支)。

扩缩 API 与状态机(relax/utils/genrm_scale_registry.py,纯标准库):
- POST /genrm/scale_out|scale_in,num_replicas 为目标绝对总数
- 同键幂等重放(verbatim NOOP replay);按模型互斥,未决清理同样阻塞
- 逐操作状态查询与 reconcile 端点

Manager 生命周期(relax/distributed/ray/genrm.py、multi_engine_manager.py):
- 每副本独占 placement group,创建即登记 ownership,所有失败出口
  (就绪超时、InfoActor 创建/探测/kill、abort)落到确认 REMOVED 或
  pending cleanup,扩缩互斥保持到 reconcile 确认
- 健康检查通过后才发布路由;超时后迟到的结果不再发布
- 缩容最新优先选 victim,初始引擎受保护;drain 完成才销毁
- PG 释放以 Ray 确认 REMOVED 为准(remove_placement_group() 返回不算)
- 物理完成栅栏:终态仅在 lifecycle 线程停止后计入
- multi_engine_manager +55 行:REMOVED 确认轮询与 pending 清理重试,
  落在 GenRM/Teacher 共享骨架(Rollout 侧零改动,Task 3 融合点)

Autoscaler 服务隔离(relax/utils/autoscaler/):
- per-service runtime:独立采集器/决策引擎/策略/历史
- ServiceScalingPolicy 支持 GenRM 独立阈值,rollout 配置向后兼容
- /conditions、/scale_history、/metrics_history 支持 service= 过滤
- 指标逐字段有效性校验;unknown_engines 冻结 scale-in(健康路径与
  旧实现一致,异常路径更保守,见 PR 描述"需维护者确认的共享语义变更")

可观测性:
- /metrics 上报 Manager 实时容量
- TUI monitor --service genrm 视图与无头 SVG 截图

测试与 GPU 证据(demos/task4_genrm/EVIDENCE.md,全部 hash pin 可复算):
- 96 个任务相关 CPU 测试(含 PG ownership 故障注入,修复前验证为红)
- 手动 1→2→1:4,163 请求 0 失败,贪心前缀跨阶段一致
- Autoscaler 全周期:3,202 请求 0 失败;冻结阈值轮 r2 12/14(A7/B3
  定性为验收口径缺陷:Ray 2.58 保留 REMOVED 墓碑)→ r3 以 v2 预注册
  口径(非终态 PG 计数,阈值逐字冻结)14/14 全绿
- 奖励一致性:greedy 跨引擎逐输入一致;官方采样 4% 不稳定率定性为
  逐引擎种子配置属性
- 故障注入:607.3s fail-closed 停等后 reconcile 清理确认;SIGKILL
  victim 后在飞请求全部重试到初始引擎
- 最小训练回路(dapo-genrm):step 训练、权重同步 200 OK ×9、judge
  调用真实发生(judge_response 落盘)
- 训练连续性(官方验收第 5 项):真实配方一次运行内 scale_out→ACTIVE
  (55s)→scale_in→COMPLETED(1s 排空),双窗口内训练事件无 >120s
  停滞、弹性引擎真实服务奖励、8/8 rollouts 跑完、零错误
- B2 冒烟与连续性配方 scripts/training/genrm/run-qwen3-0.6B-4xgpu-*.sh
@shanyulu
shanyulu force-pushed the feat/genrm-elastic-scaling branch from 8dcde5e to fa3ca97 Compare September 25, 2026 09:09
@shanyulu
shanyulu marked this pull request as ready for review September 25, 2026 10:04
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, you can upgrade your account or add credits to your account and enable them for code reviews in your settings.

@rai-studio-bot

rai-studio-bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Nyanpasu 审查看板

审查状态: 💬 已完成 · 有补充意见

审查版本: da4acbb

第 18 轮(head da4acbb):F30 关闭——EVIDENCE.md 的 ⑤ 历史行补内联撤回标注(依据 v2.1 日志分类,经原始日志独立核实)、表前段落改 run-continuity 口径、证据链接重钉 fdf288d;B2 行 zero errors 经 verdicts.json 核实为逐运行事实。至此 30 项 findings 全部 resolved/superseded。结论不变:三项维护者决策(产物归属、验收规模、采样语义)待定,裁定前不作整体通过。

审查阶段进度范围与结果
常规审查 ✅ 已完成 范围:a4f55d7→da4acbba 九个增量提交全量复核;F22/F24/F26/F27/F28/F29/F30 验证修复,F23/F25 已解决;本机 249 项通过,a55c85f8 CI 8/8。EVIDENCE.md 已按再分析 v2.1 对齐(撤回标注 + 链接重钉 fdf288d,B2 逐运行事实核实保留);da4acbba 为纯文档增量,CI 待批准。
深度审查 ✅ 已完成 必要性审计完成:attention_backend 注入建议条件注入、后端用例保留;F22 复现证据沿用(真实 AutoscalerService 直连实测:声明后终态收尾、字段缺失仍 fail-closed)。真实 GPU 与 Blackwell 未验证。

审查发现

待处理
编号 严重性 问题状态规则来源
暂无待处理的记录。
已解决或已取代
编号 严重性 问题状态规则来源
F1 Medium severity reconcile 后仍误拒已完成操作的恢复 ✅ 已解决 —
F2 Medium severity 副本数缺少严格整数校验 ✅ 已解决 —
F3 Medium severity 连续性断言未分别覆盖扩缩容窗口 ✅ 已解决 —
F4 High severity 排空确认未绑定具体 victim ✅ 已解决 —
F5 High severity PG 删除首次失败后不再提交重试 ✅ 已解决 —
F6 High severity 热新增服务未启动指标采集器 ✅ 已解决 —
F7 Medium severity 启动覆盖致旧版 rollout URL 热更新失效 ✅ 已解决 —
F8 Medium severity 缺 throughput 时丢弃有效扩容压力 ✅ 已解决 —
F9 Medium severity GenRM 面板误显示 rollout 最近操作 ✅ 已解决 —
F10 Medium severity 历史淘汰后幂等记录永久遗留 ✅ 已解决 —
F11 Medium severity 慢服务阻塞其他服务后续评估轮次 ✅ 已解决 —
F12 Medium severity 后台评估循环不应用 PATCH 后的新配置 ✅ 已解决 —
F13 Medium severity 采样验证默认切片不足,实验无法执行 ✅ 已解决 —
F14 Medium severity 证据归档链接使用不存在的提交 SHA ✅ 已解决 —
F15 Medium severity 配置交接或删除 target 丢失等待中的操作 ✅ 已解决 —
F16 Medium severity 默认 SGLang 模式未消费请求种子 ✅ 已解决 —
F17 Medium severity 确定性模式下 min-p 请求触发 sampler 断言 ✅ 已解决 —
F18 Medium severity 四个验收摘要的清单 hash 不符 ✅ 已解决 —
F19 Low severity 完整 Manager 协议精简(转后续工作)
精简建议(非阻塞)
➖ 已被后续变更取代 —
F20 Medium severity 重放终态未查询清理状态便结束跟踪 ✅ 已解决 —
F21 Medium severity 后台恢复未更新冷却时间和最近动作 ✅ 已解决 —
F22 High severity 三态清理闸门使未声明契约的终态操作永不收尾 ✅ 已解决 —
F23 Low severity PR 描述状态小节与审查状态不符 ✅ 已解决 —
F24 Medium severity 确定性后端注入静默覆盖全局 attention-backend 开关 ✅ 已解决 —
F25 Low severity radix 可用性结论未覆盖 Blackwell+ ✅ 已解决 —
F26 Low severity pre-commit 失败:守卫后导入触发 E402 ✅ 已解决 —
F27 Low severity ② 重确认未记入 EVIDENCE.md,PR 描述证据链接指向不存在的 SHA ✅ 已解决 —
F28 Low severity 扩缩端点未入用户侧 API 文档,描述指向不覆盖新契约的指南 ✅ 已解决 —
F29 Low severity rollout 侧 cleanup_required 未同步到 OpenAPI 与 API 文档 ✅ 已解决 —
F30 Low severity EVIDENCE.md 历史行仍保留已撤回的零错误/停顿上界断言且未标注 ✅ 已解决 —
提交范围 · 接收 46 · 建议移出 1 · 待确认 15

接收 46 个文件 · 建议移出 1 个文件 · 待确认 15 个文件。移出与待确认部分暂停深审,不代表审查通过。

文件结论仓库维护必要性依据替代去向或方案
relax/backends/sglang/sglang_engine.py
relax/components/genrm.py
relax/distributed/ray/genrm.py
relax/distributed/ray/multi_engine_manager.py
relax/utils/autoscaler/autoscaler_service.py
relax/utils/autoscaler/config.py
relax/utils/autoscaler/metrics_collector.py
relax/utils/autoscaler/monitor.py
relax/utils/autoscaler/scaling_decision.py
relax/utils/genrm_scale_registry.py
接收 GenRM 弹性扩缩与按服务 Autoscaler 生产实现,RFC #351 要求须随本 PR 发布。 RFC #351 与现有调用方。 仅存外部代码无法实现服务契约。
tests/components/test_genrm_scale_endpoints.py
tests/components/test_genrm_scale_watcher.py
tests/distributed/ray/test_genrm_scale_placement.py
tests/utils/_dep_stubs.py
tests/utils/autoscaler/test_autoscaler_service.py
tests/utils/autoscaler/test_metrics_validity.py
tests/utils/autoscaler/test_monitor_service_view.py
tests/utils/autoscaler/test_scaling_decision.py
tests/utils/autoscaler/test_service_targets.py
tests/utils/test_genrm_scale_registry.py
接收 上述行为的持久回归,多数免 GPU;防 v0.0.1 复发主防线。 被测生产文件与 tests/ 布局。 单条 GPU smoke 无法覆盖故障模型。
scripts/training/genrm/run-qwen3-0.6B-4xgpu-genrm-continuity.sh
scripts/training/genrm/run-qwen3-0.6B-4xgpu-genrm-smoke.sh
接收 验收配方,位于 scripts/training/ 既有布局并被 README 指向。 AGENTS.md 启动约定与 README 引用。 移出会破坏布局与复现。
demos/task4_genrm/README.md
demos/task4_genrm/README_zh.md
demos/task4_genrm/contract_demo.py
demos/task4_genrm/render_demo.py
demos/task4_genrm/test_contract_demo.py
demos/task4_genrm/e2e_genrm_scale.py
demos/task4_genrm/e2e_autoscaler_load.py
demos/task4_genrm/e2e_autoscaler_preregistered.py
demos/task4_genrm/e2e_failure_injection.py
demos/task4_genrm/e2e_reward_consistency.py
demos/task4_genrm/e2e_sampling_divergence.py
demos/task4_genrm/e2e_train_continuity.py
接收 可复现验收机制:CPU 契约演示与六个 GPU E2E 驱动。 RFC #351 验收要求与 README 职责说明。 仅留截图/结论不可复算。
demos/task4_genrm/probe_sglang_serve.py
建议移出 一次性预检探针,仓库无引用与调用方。 文件自述 "Pre-flight probe" 与引用清单(未列入)。 移入 evidence/task4-genrm 分支。
demos/task4_genrm/EVIDENCE.md
demos/task4_genrm/results/autoscaler_preregistration_20260925.md
demos/task4_genrm/results/autoscaler_preregistration_v2_20260925.md
demos/task4_genrm/results/autoscaler_prereg_v2_20260925_r5/verdicts.json
demos/task4_genrm/results/autoscaler_run_20260924_v3/verdicts.json
demos/task4_genrm/results/b2_train_smoke_20260925/verdicts.json
demos/task4_genrm/results/b2_train_smoke_20260925_r2/verdicts.json
demos/task4_genrm/results/b2_train_smoke_20260925_r3/verdicts.json
demos/task4_genrm/results/e2e_run_20260924/verdicts.json
demos/task4_genrm/results/failure_injection_20260925_r2/verdicts.json
demos/task4_genrm/results/reward_consistency_20260925_r2/verdicts.json
demos/task4_genrm/results/reward_consistency_20260925_r3/verdicts.json
demos/task4_genrm/results/sampling_divergence_20260925_r3/verdicts.json
demos/task4_genrm/results/train_continuity_20260925_r4/verdicts.json
demos/task4_genrm/results/train_continuity_20260925_r5/verdicts.json
待确认 任务专用验收记录是否留仓、是否新建顶层 demos/ 需维护者定。 仓库现状(main 无 demos/)与 README 自述。 留 demos/、移 examples/、或仅留链接。
.gitignore
接收 忽略 demo 大体积产物与失败运行目录。 EVIDENCE.md 归档说明与输出路径。 不写规则则依赖手工清理。
tests/backends/sglang/test_sglang_engine.py
接收 确定性注意力后端修复的唯一持久回归。 被测 sglang_engine.py 与既有测试布局。 仅一次性 A/B 记录无法防回归。
relax/components/rollout.py
接收 F22 修复:rollout 状态模型显式声明 cleanup_required。 调用方:状态端点与 autoscaler 轮询。 按契约限定闸门会弱化 fail-closed。
tests/components/test_scale_response_models.py
接收 F22 线缆契约回归:真实 Pydantic 模型断言序列化。 被测 rollout.py 状态模型。 固定 payload 无法证明契约声明。
tests/utils/data/test_identity_window_sampler.py
接收 随 transfer_queue stub 所需的属性守卫。 _dep_stubs.py stub 与 tests/core 先例。 stub 补齐符号会伪造被测依赖。
docs/en/api/genrm.md
docs/zh/api/genrm.md
docs/public/openapi/genrm.json
scripts/tools/generate_openapi.py
docs/en/api/rollout.md
docs/zh/api/rollout.md
docs/public/openapi/rollout.json
接收 新增扩缩契约的用户侧 API 文档(en/zh)与 OpenAPI 规范及生成器维护:GenRM 全量契约与 rollout 侧 cleanup_required 三态契约,与 docs/ 既有按服务布局一致。 被文档化的生产契约(relax/components/genrm.py 端点、registry 状态机)与 docs/en/api/、docs/public/openapi/ 既有布局及 scripts/tools/generate_openapi.py 生成器。 仅在 demos README 记载:对不跑 demo 的 API 使用者不可发现,且仓库已有正式 API 文档布局。
精简审查与验证依据
审查范围进度结论
生产代码 ✅ 已完成 复核 0481701 契约修复(保留,见 F22);其余沿用既有结论。
测试 ✅ 已完成 复核本轮新增用例与 stub 守卫(保留);其余沿用既有结论。

生产代码的必要性与替代方案

范围必须保留的契约更简单的方案结论依据与限制
genrm_scale_registry.py:公开操作状态、互斥及 replay Issue #351;GenRM submit/status/reconcile。 用 Manager 当前列表/阶段替代。 保留 逐 victim 阶段会回退,公开状态与 replay/互斥须保存;F1/F10 探针验证。
components/genrm.py:严格 HTTP 校验、容量视图、watcher/reconcile 与 inflight 公开扩缩 API 与排空/晚到完成约束。 宽松容量回退或合并 watcher/reconcile。 保留 类型边界、只读降级、容量证明不同;F1/F2/F4 已验证。
components/genrm.py:可选 begin/execute/progress dispatch 分支 构造器仅建完整 GenRMManager。 用完整协议并让替身同构。 替换 正常路径替换后仍通过;异常路径未验证,仅非阻塞建议。
distributed/ray/genrm.py、multi_engine_manager.py:victim fence、候选/退出状态和 PG ownership Manager 生命周期、恢复与 drain。 单 Event/空槽/cleanup 集合。 保留 F4 旧 A 确认不能释放 B;F5 区分未提交与未完成。
components/genrm.py seed helper;sglang_engine.py GenRM 模式 冻结 judge 跨副本采样契约。 仅 server seed 或仅请求 seed。 保留 300 次发散实验与 batch 探针否定两替代(F16/F17)。
autoscaler_service.py:独立 runtime、worker、兼容视图;config.py、monitor.py rollout/GenRM 独立决策与既有调用。 共享状态或每轮 gather。 保留 gather 阻止健康服务后续轮次;兼容视图为既有调用所需。
metrics_collector.py、scaling_decision.py:逐字段有效性及 unknown gating 缺数据不得证明闲置。 零填充或整引擎丢弃。 保留 前者错误缩容,后者丢压力;F8 探针验证。
autoscaler 提交占位与 worker 交接 PATCH 与远端操作所有权。 仅响应后登记。 保留 3127bad HTTP 验证双向占位与交接。
GenRM min-p 请求边界校验 generate API 与确定性 sampler 限制。 仅依赖后端校验。 保留 ASGI 验证随机参数均 422 且零派发。
_call_engine_tracked 新增 HTTPException handler 前置校验先于 retry。 移除该 handler。 删除 3127bad 已删除,41 项测试通过。
操作记录 service_url 与旧记录 fallback status/reconcile 属提交服务。 禁止换址或按当前配置查询。 保留 逐操作地址经变异证实;fallback 不覆盖新地址。
共享派发与SUBMIT_UNKNOWN记录 POST 可接受后丢响应。 两份派发或异常删记录。 保留 共享派发减重复;接管清理收敛为权威查询。
内联与后台接管记账 status/reconcile 消费操作状态。 两路径分别解析。 合并 已统一到 4c13780 的同一状态处理。
PATCH候选配置 失败无部分写入。 逐字段回滚或改 live 对象。 保留 deepcopy 一次替换更简单;409 测试通过。
autoscaler_service.py:1017 三态闸门与 rollout 状态响应契约 GenRM 终态须给 cleanup_required;rollout 模型未声明。 对所有契约统一三态闸门。 替换 实探:rollout 契约下 3 轮仍 pending;补字段即收尾;GenRM 侧须保留。
sglang_engine.py:确定性模式的 attention_backend 注入 确定性推理需架构适配后端保 radix。 无条件固定或仅用全局。 替换 继承循环仅回填缺失键;显式注入须以全局未设置为条件(F24)。
components/rollout.py:状态契约的 cleanup_required 声明 rollout 终态即已清理;闸门要求显式 False。 按契约限定闸门。 保留 字段声明即恢复收尾且保留 fail-closed;9755486 补回归。

测试的必要性与替代方案

范围必须保留的契约更简单的方案结论依据与限制
test_genrm_scale_registry.py:协议与有界保留 状态/互斥/replay 与有界内存。 仅 happy path。 保留 NOOP/孤儿/TTL/cap 涉不同入口;4!=0 断言检出变异。
test_genrm_scale_endpoints.py;tests/utils/_dep_stubs.py HTTP 类型/路由/容量边界。 改 registry 单测。 保留 真实 FastAPI 曾检出强转;shim 不证明集成。
endpoint 的 partial-manager fixture/missing-hook PENDING 预期 构造器仅建完整 Manager。 替身共用完整协议。 替换 无 hook 契约不存在;替换后正常派发通过。
test_genrm_scale_watcher.py、reconcile/crash tests;test_genrm_scale_placement.py 晚到完成、fence 与 PG 边界。 单条成功 GPU 用例。 保留 成功路径不会制造 victim 确认或提交异常。
TestRequestSamplingSeed 及固定版本 backend 边界探针 参数合并、seed 与后端契约。 hash 自比较或单 golden。 保留 helper 诊断参数语义;GPU golden 不验证优先级。
autoscaler service_targets、monitor、metrics_validity、scaling_decision 测试 启动/PATCH、隔离与决策边界。 仅决策引擎单测。 保留 HTTP collector 与投影保护不同调用链。
demos/task4_genrm:manual、reward/divergence、failure、autoscaler、training drivers/证据 GPU 生命周期与训练连续性验收。 单条 smoke。 保留 各驱动独立 oracle;GPU 未亲跑。
新增提交/交接测试 POST 响应期间的所有权。 占位删除或 POST 前停顿。 替换 3127bad 换为响应窗口停顿替身。
随机/greedy min-p helper 回归 随机拒绝、greedy 兼容。 删 greedy 对照。 保留 41 项通过;对照避免过度拒绝。
丢响应恢复替身 真实操作身份与 cooldown。 永久 PENDING 替身。 替换 真实 registry 生成结果覆盖 FAILED 重放。
PATCH成功/拒绝用例 多字段成功、拒绝无部分写入。 仅成功路径。 保留 两结果不同;失败用例需独立旧值。
content-deterministic seed用例 同内容一致、输入变化影响派生。 仅 hash 检查。 保留 11 项通过;重复断言可合并未验证。
tests/utils/autoscaler:清理闸门与 decision 的替身契约覆盖 真实状态契约;rollout 模型无该字段。 替身固定返回该字段。 替换 全通过也发现不了 rollout 永不收尾路径;建议按契约生成。
test_sglang_engine.py:确定性后端 4 条用例与 stub 跳过防护 F22/F24 回归:序列化、收尾、优先级、守卫。 合并为单条端到端。 保留 各用例边界不同;CI 实际执行;缺口已由 F24 用例补齐。
新增用例:响应模型序列化、RolloutWire 生命周期×4、全局优先级×2、stub 属性守卫 F22/F24 回归:序列化、收尾、优先级、守卫。 合并为单条端到端。 保留 各用例边界不同;CI 实际执行;缺口已由 F24 用例补齐。
Powered by Nyanpasu with glm-5.3[1m] xhigh, please check the suggestions carefully.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

常规审查发现缩容 reconcile 在进度查询失败时会绕过排空证明,需要修复;详见行内讨论。独立设计、资源生命周期与 Autoscaler 深度检查仍在进行,本次为阶段性结论。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

Comment thread relax/components/genrm.py Outdated
Comment thread relax/components/genrm.py Outdated
Comment thread demos/task4_genrm/e2e_train_continuity.py Outdated

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

深度审查已完成,仍需修改;新增的排空确认、PG 清理重试和 Autoscaler 生命周期问题见行内讨论。相关 CPU 测试共 189 项通过,但未覆盖已复现的边界;未在本机验证 Ray/SGLang/GPU 集成运行。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

Comment thread relax/distributed/ray/genrm.py
Comment thread relax/distributed/ray/multi_engine_manager.py Outdated
Comment thread relax/utils/autoscaler/autoscaler_service.py Outdated
Comment thread relax/utils/autoscaler/config.py
Comment thread relax/utils/autoscaler/metrics_collector.py Outdated
Comment thread relax/utils/autoscaler/autoscaler_service.py
…le closed

Two review findings on the scale-in drain contract:

1. Reconcile no longer falls through when the manager progress query
   fails. Previously a timeout/error returned {} and reconcile proceeded
   with victim=None, so the manager retired its recorded victim on the
   caller's (missing) drain proof alone, terminating requests still in
   flight inside the single Gateway. Reconcile now fails closed with 503
   until the victim is positively identified and proven at zero
   in-flight (regression tests: progress-query failure and unknown
   victim both refuse; identified drained victim still reconciles).

2. Drain confirmation is bound to the victim the watcher observed. A
   multi-victim scale-in reuses the request id and swaps the drain event
   per victim, so a delayed duplicate confirm for the previous victim
   could release the next one without its in-flight count ever reaching
   zero. confirm_scale_drained now takes the observed victim address and
   rank and ignores stale identities under the scale lock; the watcher
   passes them (regression tests: stale confirm between victims and
   after retirement both ignored; legacy no-arg confirm still works).
…tors

Two review findings on resource and control-plane safety:

1. _remove_owned_pg previously used membership in _pending_pg_cleanup to
   skip resubmitting remove_placement_group(). When the first submission
   itself raised (before Ray accepted the deletion), the rank entered
   _pending_pg_cleanup, every later retry only polled the table, and the
   PG stayed CREATED forever -- GPU leak plus a permanently held model
   scale mutex. Deletion-submitted state is now tracked separately
   (_pg_remove_submitted): a failed submission is retried, an accepted
   one is only polled for REMOVED confirmation (regression tests: first
   submit failure resubmits then confirms REMOVED; accepted submission
   is never resubmitted).

2. AutoscalerService._rebuild_service_runtimes constructed a fresh
   MetricsCollector for a service target added via PATCH /config but
   never started it -- without an HTTP session the new target collected
   nothing and never autoscaled; a removed target kept its session
   alive. The rebuild now records collectors to start/stop and the
   async PATCH path starts/stops them before publishing the new runtime
   set; start() consumes the pending list (regression tests: PATCH adds
   a started collector, PATCH removing the target stops it).
…apshots, bounded history

Review findings on the scale control plane:

- num_replicas is now a strict int: JSON true / "2" / 2.0 are rejected
  422 at the HTTP boundary instead of being coerced by pydantic before
  the registry's isinstance checks (boundary test covers all three).
- Scale submission fails closed when the authoritative capacity cannot
  be queried: the mutating path uses strict capacity lookups and returns
  503 instead of submitting against a degraded snapshot where ready
  impersonates current; read-only /engines keeps its degraded view.
- A watcher crash no longer zero-fills capacity: the terminal snapshot
  falls back to the registry's last known values (previous terminal
  snapshot, else the submit-time observation, now exposed in status)
  instead of inventing current=0/ready=0, and an entirely empty
  progress payload keeps cleanup_required set (fail closed) while a
  non-empty legacy snapshot without the key keeps its old meaning.
- The registry is a long-running control plane: clean terminal
  operations are bounded to max_history (oldest evicted, counted in
  history_truncated), idempotency records whose clean outcome aged out
  of the replay window expire to fresh-request semantics, and dirty
  terminals / live operations are never evicted.
Review finding: the GenRM server args derived random_seed from the
engine rank, so the elastic replica sampled differently from the initial
one and the two replicas disagreed on borderline inputs under the
official sampling config (measured 2/50 verdict flips in the reward
consistency run). GenRM is a frozen scoring service: every replica must
sample identically for an identical request, so the seed is now the base
args.seed on every rank. The rollout builder keeps its rank-derived seed
(sampling diversity is wanted there, not here).
Review findings on the autoscaler control plane:

- Per-service evaluation now runs concurrently (asyncio.gather with
  per-service exception isolation): a stalled GenRM metrics endpoint can
  no longer delay the rollout service's evaluation cadence (regression
  test parks the GenRM evaluation and asserts rollout completes).
- A legacy PATCH of rollout_service_url is no longer shadowed by the
  startup service_targets entry: the two spellings are kept in sync in
  both directions, so the patched URL actually takes effect (tests for
  both directions).
- /status reports per-service last_scale_action / last_scale_time and
  the monitor's service view overlays them, so --service genrm shows
  GenRM's own scale history instead of the rollout runtime's.
- Metrics aggregation is per-field: an engine whose scrape is missing
  only throughput still contributes its observed token_usage and queue
  depth to scale-out signals, while remaining an unknown engine for the
  conservative scale-in gate (cross-layer regression test).
- Together with the collector hot start/stop from the previous commit,
  PATCHed service targets are fully live at runtime.
Review finding: the continuity verdict's single in-window check used the
scale-out COMPLETION time as its lower bound, so the whole scale-out
process was excluded and one train event between the two operations
satisfied the assertion even when the 1 s scale-in window contained
nothing. The verdict now asserts progress separately before the
operations, between them, and after scale-in -- the honest, observable
windows -- and the run-6 evidence is re-checked against them by the
final-head rerun (the 1 s drain window is shorter than one training
iteration; progress during it is not an observable claim).
@shanyulu

This comment has been minimized.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已完成 af82e6a 复查:原先的 P1 安全问题均已修复;八项反馈已解决,reconcile 仍有一项 P2 恢复问题,已在原讨论补充并降低优先级。另有历史保留和跨服务评估节奏两项 P2,见新增行内意见。本轮仅留非阻塞意见。

69 项组件/状态机测试与 137 项 Autoscaler 测试通过;真实 Ray/SGLang/GPU 集成未重跑。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

Comment thread relax/utils/genrm_scale_registry.py Outdated
Comment thread relax/utils/autoscaler/autoscaler_service.py Outdated
Evidence-provenance cleanup (review feedback):

- Remove all relative references ('final head' / 'final PR head' /
  'this PR head' / 'final-head rerun') from EVIDENCE.md; every run now
  names the commit that produced it: the seed-contract chain (score
  consistency r3, adversarial divergence r3, training smoke r3, training
  continuity r5) at 945741e; the round-1 fix batch (failure injection r2,
  autoscaler prereg r4/r5, smoke r2) at their exact heads (e7224af /
  8b0c6fe / fddb093); the preregistered v2 rounds at fa3ca97.
- State the pinning policy accordingly: the newest code exercised by any
  recorded acceptance run is 945741e; earlier verdicts are pinned to
  their own producing commits.
- Judge ground-truth correctness is now explicitly informational and
  reported per retained run (95%/95% in _r2, 97%/97% in _r3); replica
  agreement, not judge capability, is the acceptance criterion.
- Terminology: 'request-level sampling-seed contract' -> 'content-derived
  sampling-seed contract' (same content reproduces the same seed across
  replicas and repeated calls).

Artifact blobs are untouched; the SHA256 pin manifest still matches the
committed verdict files.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

审查结论:请求修改后再合并

本轮复审了上一轮两条 P1(4c137803)的修复,并复查弹性扩缩与 GenRM 扩缩注册表相关子系统;最新提交 16d7148 仅改动 demos/task4_genrm/EVIDENCE.md(证据行改为按产出提交固定),代码未变,行内意见锚点在该提交上依然成立。

上一轮意见(已修复,原线程已由作者关闭)

  • GenRM 终态重放不再被当作清理完成:接管后同一轮继续以状态接口为唯一权威判据,只有明确干净才进入 history;恢复矩阵 A–F 的回归覆盖到位。
  • 正常受理、内联同 key 重试、后台补投三条路径共用同一接管逻辑,冷却起点与操作账目一致,冷却一致性用例通过。

新发现(阻塞项)

  • P1:三态清理闸门把「状态响应缺少 cleanup_required」判为未确认,但该字段只存在于 GenRM 的状态契约;autoscaler 的默认目标 rollout 的状态响应模型未声明它,导致 rollout 的扩缩操作到达终态后无法收尾,进而永久冻结后续决策,并让 target 删除守卫持续返回 409。详见 relax/utils/autoscaler/autoscaler_service.py:1017 的行内意见(含实证与本机复现,另有两条修复方向可选)。

验证范围与缺口

  • 本机运行 Autoscaler 全部用例:149 项通过;另有 11 项因沙箱缺少 aiohttp 未能运行(环境限制,与本次改动无关)。
  • CI 检查 8/8 通过,但现有新增用例的替身均按 GenRM 契约返回该字段,未覆盖非 GenRM 服务契约路径,故 CI 绿灯不能证明上述回归不存在。
  • 生产/测试必要性与简化审计结论已记录在看板,本轮未提出新的简化类意见。
Powered by Nyanpasu with deepseek-v4.1-flash-ali medium, please check the suggestions carefully.

Comment thread relax/utils/autoscaler/autoscaler_service.py
Comment thread relax/utils/autoscaler/autoscaler_service.py Outdated
@rai-studio-bot

Copy link
Copy Markdown

P3 优先级:P3(非行级:PR 描述)

描述中新增的状态小节与当前审查状态不一致,建议同步更新,避免维护者据此判断合并条件:

  • 「Review and CI status」称 “All correctness findings through the current review round have dedicated fixes and regressions”,并称两条 round-8 问题仍 “awaiting bot re-verification”。实际状态:那两条已在 4c13780 复验通过(已在原线程回复),但本轮在 16d7148 上发现了一条未解决的 P1——三态清理闸门对未声明 cleanup_required 的服务契约(autoscaler 的默认目标 rollout)会把已完成的扩缩操作永久留在 pending_requests,冻结后续决策并让删除 target 长期返回 409。详见行内意见 discussion_r4110171098 与本轮审查 pullrequestreview-5324561491。
  • 「Post-freeze validation」把 4c13780 一行的暴露面记为 “none — requires response loss or failed operations”。这一判断低估了范围:三态闸门同时改变 rollout 的成功路径——状态响应缺少该字段即保持 UNKNOWN,操作到达终态后永不收尾(本机实探:真实 AutoscalerService + 按 rollout 契约返回状态的服务替身,3 轮后仍为 pending、history 为空、下一决策为 NONE;同一替身补上 cleanup_required: False 即正常收尾)。因此任何会触发 rollout 扩缩的场景都会覆盖该分支,修复 P1 时建议一并复核该行的结论。

修复顺序建议:先把 rollout 契约路径修好或明确作用域,再更新这两处描述(例如:两条 round-8 问题已复验通过;新增 P1 待修复;4c13780 的暴露面包含 rollout 成功路径)。

Powered by Nyanpasu with deepseek-v4.1-flash-ali xhigh, please check the suggestions carefully.

…-Blackwell

Final-code Autoscaler acceptance reconfirmation (Criterion 4, frozen
protocol) exposed a product regression introduced by the deterministic
sampling activation (945741e):

Reproduction (identical frozen driver, thresholds, load curve and
machine, same day):
- final code (deterministic ON): the engine wedged mid-load -- the
  scheduler stopped completing requests (served frozen at 843, 48
  requests hung, /health no answer, GPU util 0%), the elastic replica
  served 0 requests, and scale-in never became possible; the run failed.
- pre-deterministic code (8b0c6fe, r5's product code): 13/14 functional
  assertions passed, elastic replica served 461 requests in STEADY and
  automatic scale-in completed.

Root cause: SGLang applies the generic flashinfer attention-backend
default BEFORE its deterministic-inference handler. On pre-Blackwell
GPUs the handler's own deterministic fallback (fa3) is therefore
preempted, flashinfer is kept, and the radix cache is force-disabled
(flashinfer is not in SGLang's radix-supported deterministic set).
Without prefix reuse, a 48-way concurrent scoring load re-prefills every
long shared prompt until the scheduler wedges.

Fix: select the attention backend SGLang itself recommends for
deterministic inference on the detected architecture (fa3 on
pre-Blackwell, flashinfer on Blackwell+), so the radix cache stays
available. The per-request sampling-seed contract is unchanged
(deterministic mode stays on; the pytorch sampling backend still
consumes per-request seeds). Per-instance --genrm-engine-config
overrides keep the highest priority.

Regressions: deterministic attention-backend selection (fa3 pre-
Blackwell, flashinfer Blackwell+, probe failure defers to the SGLang
default), engine-config override precedence, and the server-args
pipeline carrying attention_backend alongside
enable_deterministic_inference.
FINAL_EVIDENCE_SHA = 6b501a2 (raw artifacts: the final-code confirmation
r3, the regression reproduction r2, the environment-invalid r1, the A/B
discrimination run at 8b0c6fe, and the four machine-blocked
score-consistency attempts).

EVIDENCE.md: the final-code Autoscaler confirmation row (frozen protocol,
14/14 at 5c1e2e7, deterministic sampling under the radix-preserving fa3
attention backend); the regression and environment failure records in
the failed-runs table; the pinning policy now names 5c1e2e7 as the
newest code exercised by any recorded acceptance run; the determinism
known-limit notes that fa3 keeps the radix cache available on
pre-Blackwell GPUs; the SHA256 manifest gains the final verdict blob.
…module

Suite-wide dependency-stub tests can replace sglang.srt.server_args
while relax.backends.sglang.sglang_engine is already imported, leaving
its module-level ServerArgs bound to a non-dataclass stub that
_compute_genrm_server_args cannot iterate. The two new deterministic
attention-backend tests now skip on what the engine module actually
sees (2 skips in the stubbed context, 9 passed standalone); the full
suite's failure set is byte-identical to the pre-change baseline.
@shanyulu

This comment has been minimized.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

本轮结论:仍需修改——此前提出的 P1(三态清理闸门与 rollout 状态契约不匹配,而 rollout 是 autoscaler 的默认目标)在最新 head a4f55d7 复核后依然存在,合并条件不变;已在原线程补充复核证据(discussion_r4110171098)。

本轮增量(16d7148f → a4f55d7:5c1e2e7 的确定性注意力后端修复、3127b63 的证据文档、a4f55d7 的测试跳过防护)另有两处非阻塞意见:

  • P2:显式写入的 "attention_backend" 会静默覆盖全局 --sglang-attention-backend(含探测失败的分支),非 CUDA 平台会被确定性写成 fa3;建议改为「全局未设置时才注入」或限定在 CUDA 设备上,并补一条优先级用例。
  • P3:注释中「让 radix 缓存保持可用」在 Blackwell+ 上不成立(仍走 flashinfer → disable_radix_cache),建议在注释与 known limits 中写明该边界。

核对信息:本轮审查 head 为 a4f55d7(目标 head 与 PR head 一致,两条行内意见均标注在该 head 上),该 head 的 8 项 CI 检查全部通过。范围决策未变:接收 36 / 建议移出 1 / 待确认 15,移出与待确认部分不计入本轮结论。

Powered by Nyanpasu with glm-5.3[1m] xhigh, please check the suggestions carefully.

Comment thread relax/backends/sglang/sglang_engine.py Outdated
Comment thread relax/backends/sglang/sglang_engine.py Outdated
Re-review P1: the three-state cleanup gate (UNKNOWN/False/True, missing
!= clean) treats every service status response without an explicit
cleanup_required as UNKNOWN, but that flag only existed in GenRM's
status contract. Rollout's ScaleOutStatusResponse and
ScaleInStatusResponse never declared it, so FastAPI's response_model
stripped the key from the wire and the shared autoscaler -- whose
default target is rollout -- read UNKNOWN for every rollout terminal
operation:

- terminal scale-out (ACTIVE/PARTIAL/FAILED/CANCELLED) and scale-in
  (COMPLETED/FAILED) stayed pending_requests forever;
- the next scaling evaluation stayed frozen ("in progress or awaiting
  cleanup") with no path to unblock;
- a removable rollout-contract target kept failing PATCH removal with
  409 (stale unresolved pending);
- history never recorded the finished operation.

Rollout has no deferred-cleanup lifecycle: a FAILED scale-out rolls its
engines back before reporting and a COMPLETED scale-in reports the
engines removed, so terminal status is cleanup-complete by contract.
The fix therefore makes the rollout status models state that explicitly
(cleanup_required: bool = False) instead of leaving the flag absent.
GenRM's semantics are unchanged: a GenRM POST reply still proves
nothing, a missing/malformed flag still reads UNKNOWN, terminal-dirty
still requires status/reconcile, and finalization still requires an
authoritative cleanup_required is False.

Coverage:
- response-model serialization tests pin model_dump()["cleanup_required"]
  is False for both scale-out and scale-in wire shapes (the response-
  model tests now import the component through tests.utils._dep_stubs
  so they actually run where transfer_queue is absent);
- shared-autoscaler regressions drive a real AutoscalerService against
  a fake speaking the rollout HTTP contract through the actual Pydantic
  response models: scale-out and scale-in terminal requests finalize
  into history, the next decision is no longer blocked, and a finalized
  target no longer 409s on removal;
- the fail-closed guard stays pinned: a rollout-contract response with
  the flag dropped (pre-fix wire shape) keeps the request pending and
  the decision frozen -- the fix lives in the contract, never in the
  autoscaler defaulting missing to clean.
…ic probe declines

Re-review P2: _compute_genrm_server_args pinned attention_backend
unconditionally -- including the architecture probe's None -- so the
sglang_* inheritance loop (which backfills only fields absent from
kwargs) could never apply the user's global --sglang-attention-backend
to a GenRM engine again. The docstring's "falls back to SGLang default"
only held for SGLang's internal default; the relax-level global flag
was silently disabled by 5c1e2e7.

The deterministic backend is now pinned only when the user has not
chosen one globally:

- explicit --sglang-attention-backend wins over the probe (restores the
  pre-5c1e2e7 inheritance behaviour);
- a failed probe defers to SGLang's default instead of pinning None
  over it (and over the user's global choice);
- the arch-appropriate deterministic backend (fa3 on pre-Blackwell,
  flashinfer on Blackwell+, radix cache preserved) is still pinned when
  the user set nothing, keeping the final-code autoscaler acceptance
  behaviour on this machine;
- per-instance --genrm-engine-config overrides still take priority over
  both (applied after the loop).

Regressions:
- an explicit global choice beats the probe result;
- a failed probe pins nothing (key absent, SGLang default applies) while
  the user's global choice still flows through the inheritance loop.
@shanyulu

This comment has been minimized.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

复审结论:上一轮的两条意见已在 04817012 验证修复,此前唯一的 P1(rollout 状态契约缺 cleanup_required,默认目标的已终态操作永不收尾)已消除,本轮增量无新增缺陷;因 demos/task4_genrm 证据文件的去向仍待维护者裁定(15 个待确认 + 1 个建议移出,见看板),本轮不给出整体通过。

  • P1 已修复(rollout 清理契约):两个状态模型显式声明 cleanup_required: bool = False,与终态语义一致(scale-out FAILED 前已回滚、scale-in COMPLETED 前已移除引擎)。本机以真实 AutoscalerService 直连复测:声明后终态操作 pending→history、后续决策解冻;字段缺失(修复前形状)仍保持 fail-closed 永久 pending。序列化与生命周期回归齐备,含把修复前形状钉死为「保持 pending + 决策冻结」的守卫用例(详见 原线程)。
  • P2 已修复(attention_backend 优先级):仅在全局 --sglang-attention-backend 未设置时注入,探测失败不再写入 None;两条优先级用例补齐,NPU/KLX 配方不再被覆盖(详见 原线程)。
  • P3 已按已知限制记录(radix / Blackwell+):线程回复与 EVIDENCE.md 的 pre-Blackwell 边界表述覆盖该提醒,后续按 upstream-dependent 跟进。
  • 验证环境说明:本机任务内测试 249 项通过(rollout 契约与 sglang 后端用例因沙箱缺 torch/transfer_queue/sglang 跳过,需 CI 全量环境执行);该 head 的 required workflows 仍在等待维护者批准运行。

范围决策:接收 39 / 建议移出 1 / 待确认 15;demos/task4_genrm 证据文件的仓库去向(留 demos/、移 examples/、或仓库仅留 README/EVIDENCE 链接指向证据分支)需要维护者拍板,裁定前不作整体通过。

Powered by Nyanpasu with glm-5.3[1m] xhigh, please check the suggestions carefully.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

补充意见:该 head 的 CI 中 Pre-commit Checks 失败——9755486 在 tests/utils/data/test_identity_window_sampler.py 的模块级守卫之后留下两处导入,触发 ruff E402×2(行内已附最小修复建议,P3 非阻塞,修复合入前请重跑 pre-commit)。Unit Tests (H20, 4 GPUs) 仍在运行,其中包含上轮本机沙箱无法执行的 rollout 契约与 sglang 优先级用例。其余结论同上一轮:两条复审意见已验证修复、无新增功能缺陷;范围决策(接收 39 / 建议移出 1 / 待确认 15)仍待维护者裁定。

Powered by Nyanpasu with glm-5.3[1m] xhigh, please check the suggestions carefully.

Comment thread tests/utils/data/test_identity_window_sampler.py Outdated
The module-level transfer_queue attribute guard (added with the rollout
contract stub support) places two imports after executable code; ruff's
E402 fires on the pushed tree and would fail the Pre-commit workflow.
Annotate both imports as guarded.
@shanyulu

shanyulu commented Sep 26, 2026 •

Copy link
Copy Markdown
Author

当前 PR 头 a55c85f,8/8 CI 通过;GenRM 产品代码保持在 0481701,仍待当前头正式 review。

请维护者确认三项范围与行为:

决策 当前方案 请确认
验收文件归属 driver、配方与小型摘要在 PR 中;原始运行另存不可变 evidence 分支 保留 demos/task4_genrm/、迁到指定目录,或主仓库仅保留必要入口与索引?
验收规模 4×RTX 4090、Qwen3-0.6B;单 Gateway、单 GPU 弹性副本,另有上游 GPU / 训练 CI 是否满足本期生命周期与路由验收;如不足,请指定模型或拓扑。
采样语义 未指定 seed 时按内容派生;显式 seed 优先。同内容重复调用也受确定性约束 接受该默认行为,还是要求重复调用保留独立随机性?

各实验的运行版本、结果与限制见 PR 正文,设计见 RFC #351。训练连续性以 再分析 v2.1为准:运行完成已验证,未证明零性能影响。

范围或语义裁决若要求改动,再提交最小增量并重新走 CI / review。

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

当前 head f507d3b7 的复审收口:全部 26 项审查问题已解决并验证——最后一条(pre-commit E402 注解)已在该 head 修复(ruff 复核通过,见原线程);此前 0481701 上的两条复审修复(rollout 清理契约、attention_backend 优先级)均经真实 AutoscalerService 直连实测与回归用例验证。该 head 的 8 项 CI 检查全部通过(含 H20 4-GPU 单测与三项集成测试),25 个 review 线程均已解决。

本轮不给出整体通过的唯一原因仍是范围决策:demos/task4_genrm 证据文件的去向(作者已在交接评论中列出 A/B/C 三个选项)需要维护者裁定,裁定前 15 个待确认与 1 个建议移出的路径保持可见(范围:接收 39 / 建议移出 1 / 待确认 15)。范围裁定(以及若需移动文件产生的新 head 复核)完成后即可给出最终结论。

Powered by Nyanpasu with glm-5.3[1m] xhigh, please check the suggestions carefully.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

补充意见(非阻塞):PR 描述把最终代码的 ② 重确认记为 PASS,证据确实存在——已在证据分支新顶端核对 reward_consistency_20260927_final_r6 与 sampling_divergence_20260927_final_r4 两个运行目录(旧顶端无、新顶端有)。但记录有两处待修:PR 描述中新加的证据分支链接使用了不存在的完整 SHA(1deb9d83f79b…,实际分支顶端为 1deb9d830710a923ba561c99a32106c00cfa325f),且 demos/task4_genrm/EVIDENCE.md 未收录这两个运行、文件头分支指针仍停在旧顶端 6b501a2(行内意见已附具体修改建议)。其余状态与上一轮一致:26 项审查问题全部解决并验证、CI 8/8、25 线程 resolved;范围裁定(demos 去向 A/B/C、验收规模)仍待维护者。

Powered by Nyanpasu with glm-5.3[1m] xhigh, please check the suggestions carefully.

Comment thread demos/task4_genrm/EVIDENCE.md Outdated
…tion PASS, r5 proxy incident, train-continuity reanalysis v2)

Evidence-branch pointer update only; no product change. Raw artifacts live on
the evidence branch @ d88d6c7 (r6/r4 verdict trees, the archived raw job log
with sha256, and the reanalysis v2 script + parser tests).
…ons table

Per review: the top of the file still pointed at evidence 6b501a2 and called
5c1e2e7 the newest acceptance code while the bottom appended 0481701
reconfirmations — an index that contradicted itself. The top is now the only
current-conclusions table: one row per official criterion, each bound to the
exact product SHA, the producing run and its raw inputs on evidence 66f1f78,
with an explicit superseded-by column. Reading rules state the caliber of
each claim (0/50 = these 50 inputs under this config; GenRM-path-only weight
sync; criterion 5 is continuity, no performance claim absent a control run;
content-deterministic semantics pending maintainer acceptance). A scope
commitments table separates delivered / not-implemented / to-be-negotiated
(cross-Gateway wording now 'not implemented this phase; interface and
responsibility to be negotiated'). The per-run history below is marked as
historical detail and keeps its own producing commits.
@shanyulu

This comment has been minimized.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

状态更新:5b512d0 的 8 项 CI 检查已全部通过(含 pre-commit、Python 3.10–3.12、H20 GPU 单测与三项训练集成)——此前记录的最后一个 CI 门槛清零;描述重写(What/Why/How/Testing)与新增的连续性可视化更正均与已核实记录一致(诚实限缩:不证明步延迟不变、撤回请求数归因)。全部 27 项审查问题维持 resolved/superseded。

本轮新增一条非阻塞意见:新增的扩缩端点未进入用户侧 API 文档(docs/en/api/genrm.md 仍只覆盖打分服务),而描述指其「describes the public contract」——建议补文档或改指向(行内已附,可与范围裁定的文档去向一并处理)。整体通过仍待三项维护者决策(demos 去向、验收规模、采样语义)。

Powered by Nyanpasu with glm-5.3[1m] xhigh, please check the suggestions carefully.

Comment thread relax/components/genrm.py
- docs/en/api/genrm.md + docs/zh/api/genrm.md: endpoints table, request/
  status contract, NOOP/idempotency semantics, scale-out/in state chains
  with terminal sets, cleanup_required and capacity fields (current/ready/
  occupied/pending_cleanup), strict-capacity 503 and fallback semantics
- docs/public/openapi/genrm.json: regenerate via scripts/tools/
  generate_openapi.py; 10 paths, 9 schemas, descriptions from code
  docstrings
- scripts/tools/generate_openapi.py: write trailing newline so reruns are
  byte-stable

Verified against relax/components/genrm.py, relax/utils/genrm_scale_registry.py
and relax/distributed/ray/genrm.py: every documented route, state, terminal
set, capacity formula and error code matches the implementation.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

增量复核(5b512d0 → 9095c66,文档提交):GenRM 扩缩契约的 en/zh API 文档与 OpenAPI 规范已核对——端点、严格整数校验(422)、幂等键重放语义(有界内存史、非跨重启保证)、三态清理与 reconcile 行为、capacity 字段含义及单 Gateway 范围警示均与实现一致;genrm.json 与生成器输出同步、schema 含 cleanup_required。F28 以两种方式(指向修正 + 正式文档)落实,已关闭。

一条新的非阻塞意见(行内):rollout 侧由 F22 修复引入的 cleanup_required 字段未同步到 docs/public/openapi/rollout.json 与 docs/en|zh/api/rollout.md——GenRM 侧已补齐,建议 rollout 侧一并补上(注意重跑生成器会带上 main 既有的 /predict 无关 drift)。范围决策随文档组更新为:接收 43 / 建议移出 1 / 待确认 15;该 head 的 CI 待批准运行。整体结论不变:三项维护者决策(demos 去向、验收规模、采样语义)待定。

Powered by Nyanpasu with glm-5.3[1m] xhigh, please check the suggestions carefully.

Comment thread relax/components/rollout.py
- docs/public/openapi/rollout.json: add cleanup_required (bool) to
  ScaleOutStatusResponse and ScaleInStatusResponse properties — manual
  patch in generator-faithful style (type + title, not in required),
  avoiding the unrelated /predict drift a full regeneration would bring
- docs/en/api/rollout.md + docs/zh/api/rollout.md: new 'Scale status
  cleanup contract' section (bilingual semantic parity) documenting the
  three-state contract shared with the Autoscaler (true = cleanup still
  pending; false = authoritative complete; absent = unknown, missing !=
  clean) and why rollout terminal states report false (scale-out engines
  serve or were rolled back; scale-in COMPLETED removed / FAILED rolled
  back)

Docs-only; no product code touched.
@shanyulu shanyulu changed the title 【No.4】feat(genrm): add elastic scaling and autoscaler support 【No.4】GenRM 弹性扩缩容与自动伸缩 Sep 27, 2026
# 📝 Documentation

- Point the two evidence index links to the reviewed fdf288d snapshot.
- Leave product code, historical verdicts and experiment versions unchanged.
- Keep this commit local; the public PR head and its CI remain unchanged.
# 📝 Documentation

- Withdraw whole-run zero-error and stall-bound claims using reanalysis v2.1.
- Distinguish log-reported sample counts from file-level sample verification.
- Keep original verdicts and producing versions; mark superseded interpretations.
- Correct the long-request drain description and completed reconfirmation status.
- Preserve product code, GPU results and public PR history.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

本轮描述更新撤回了训练连续性运行的两个旧断言(「全程零错误」与「全程停顿不超过 120 秒」),撤回依据经独立核实成立:再分析 v2.1 的日志分类(三条非良性带时间戳错误记录、judge 错误行 0、双副本稳定窗口内错误行 0)与归档原始日志一致,且「事件间隔不能作为停顿上界」的说明正确。这是又一次有效的诚实性修正。

一条新的非阻塞意见(行内):demos/task4_genrm/EVIDENCE.md 的训练连续性历史行仍原样保留这两个已撤回的断言,未按该文档自身「superseded rows say so in-line」的约定做标注——读者经合并门槛指向查阅清单时会看到与描述矛盾的未标注声明;建议补内联撤回标注并顺带把再分析引用更新到 v2.1。其余状态不变:29 项审查问题全部解决、CI 8/8@当前 head、三项维护者决策待定。

Powered by Nyanpasu with glm-5.3[1m] xhigh, please check the suggestions carefully.

Comment thread demos/task4_genrm/EVIDENCE.md Outdated

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants