Conversation
084b0b5 to
8dcde5e
Compare
关联 redai-studio#351。官方验收标准五项全部达成(证据矩阵见 PR 描述)。 单 commit 汇总全部开发提交;完整历史在 feat/genrm-elastic-scaling-history 分支(证据文件中的 hash 引用指向该分支)。 扩缩 API 与状态机(relax/utils/genrm_scale_registry.py,纯标准库): - POST /genrm/scale_out|scale_in,num_replicas 为目标绝对总数 - 同键幂等重放(verbatim NOOP replay);按模型互斥,未决清理同样阻塞 - 逐操作状态查询与 reconcile 端点 Manager 生命周期(relax/distributed/ray/genrm.py、multi_engine_manager.py): - 每副本独占 placement group,创建即登记 ownership,所有失败出口 (就绪超时、InfoActor 创建/探测/kill、abort)落到确认 REMOVED 或 pending cleanup,扩缩互斥保持到 reconcile 确认 - 健康检查通过后才发布路由;超时后迟到的结果不再发布 - 缩容最新优先选 victim,初始引擎受保护;drain 完成才销毁 - PG 释放以 Ray 确认 REMOVED 为准(remove_placement_group() 返回不算) - 物理完成栅栏:终态仅在 lifecycle 线程停止后计入 - multi_engine_manager +55 行:REMOVED 确认轮询与 pending 清理重试, 落在 GenRM/Teacher 共享骨架(Rollout 侧零改动,Task 3 融合点) Autoscaler 服务隔离(relax/utils/autoscaler/): - per-service runtime:独立采集器/决策引擎/策略/历史 - ServiceScalingPolicy 支持 GenRM 独立阈值,rollout 配置向后兼容 - /conditions、/scale_history、/metrics_history 支持 service= 过滤 - 指标逐字段有效性校验;unknown_engines 冻结 scale-in(健康路径与 旧实现一致,异常路径更保守,见 PR 描述"需维护者确认的共享语义变更") 可观测性: - /metrics 上报 Manager 实时容量 - TUI monitor --service genrm 视图与无头 SVG 截图 测试与 GPU 证据(demos/task4_genrm/EVIDENCE.md,全部 hash pin 可复算): - 96 个任务相关 CPU 测试(含 PG ownership 故障注入,修复前验证为红) - 手动 1→2→1:4,163 请求 0 失败,贪心前缀跨阶段一致 - Autoscaler 全周期:3,202 请求 0 失败;冻结阈值轮 r2 12/14(A7/B3 定性为验收口径缺陷:Ray 2.58 保留 REMOVED 墓碑)→ r3 以 v2 预注册 口径(非终态 PG 计数,阈值逐字冻结)14/14 全绿 - 奖励一致性:greedy 跨引擎逐输入一致;官方采样 4% 不稳定率定性为 逐引擎种子配置属性 - 故障注入:607.3s fail-closed 停等后 reconcile 清理确认;SIGKILL victim 后在飞请求全部重试到初始引擎 - 最小训练回路(dapo-genrm):step 训练、权重同步 200 OK ×9、judge 调用真实发生(judge_response 落盘) - 训练连续性(官方验收第 5 项):真实配方一次运行内 scale_out→ACTIVE (55s)→scale_in→COMPLETED(1s 排空),双窗口内训练事件无 >120s 停滞、弹性引擎真实服务奖励、8/8 rollouts 跑完、零错误 - B2 冒烟与连续性配方 scripts/training/genrm/run-qwen3-0.6B-4xgpu-*.sh
8dcde5e to
fa3ca97
Compare
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Nyanpasu 审查看板审查状态: 💬 已完成 · 有补充意见 审查版本: da4acbb 第 18 轮(head da4acbb):F30 关闭——EVIDENCE.md 的 ⑤ 历史行补内联撤回标注(依据 v2.1 日志分类,经原始日志独立核实)、表前段落改 run-continuity 口径、证据链接重钉 fdf288d;B2 行 zero errors 经 verdicts.json 核实为逐运行事实。至此 30 项 findings 全部 resolved/superseded。结论不变:三项维护者决策(产物归属、验收规模、采样语义)待定,裁定前不作整体通过。
审查发现待处理
已解决或已取代提交范围 · 接收 46 · 建议移出 1 · 待确认 15接收 46 个文件 · 建议移出 1 个文件 · 待确认 15 个文件。移出与待确认部分暂停深审,不代表审查通过。
精简审查与验证依据
生产代码的必要性与替代方案
测试的必要性与替代方案
Powered by Nyanpasu with glm-5.3[1m] xhigh, please check the suggestions carefully.
|
|||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
rai-studio-bot
left a comment
There was a problem hiding this comment.
常规审查发现缩容 reconcile 在进度查询失败时会绕过排空证明,需要修复;详见行内讨论。独立设计、资源生命周期与 Autoscaler 深度检查仍在进行,本次为阶段性结论。
rai-studio-bot
left a comment
There was a problem hiding this comment.
深度审查已完成,仍需修改;新增的排空确认、PG 清理重试和 Autoscaler 生命周期问题见行内讨论。相关 CPU 测试共 189 项通过,但未覆盖已复现的边界;未在本机验证 Ray/SGLang/GPU 集成运行。
…le closed
Two review findings on the scale-in drain contract:
1. Reconcile no longer falls through when the manager progress query
fails. Previously a timeout/error returned {} and reconcile proceeded
with victim=None, so the manager retired its recorded victim on the
caller's (missing) drain proof alone, terminating requests still in
flight inside the single Gateway. Reconcile now fails closed with 503
until the victim is positively identified and proven at zero
in-flight (regression tests: progress-query failure and unknown
victim both refuse; identified drained victim still reconciles).
2. Drain confirmation is bound to the victim the watcher observed. A
multi-victim scale-in reuses the request id and swaps the drain event
per victim, so a delayed duplicate confirm for the previous victim
could release the next one without its in-flight count ever reaching
zero. confirm_scale_drained now takes the observed victim address and
rank and ignores stale identities under the scale lock; the watcher
passes them (regression tests: stale confirm between victims and
after retirement both ignored; legacy no-arg confirm still works).
…tors Two review findings on resource and control-plane safety: 1. _remove_owned_pg previously used membership in _pending_pg_cleanup to skip resubmitting remove_placement_group(). When the first submission itself raised (before Ray accepted the deletion), the rank entered _pending_pg_cleanup, every later retry only polled the table, and the PG stayed CREATED forever -- GPU leak plus a permanently held model scale mutex. Deletion-submitted state is now tracked separately (_pg_remove_submitted): a failed submission is retried, an accepted one is only polled for REMOVED confirmation (regression tests: first submit failure resubmits then confirms REMOVED; accepted submission is never resubmitted). 2. AutoscalerService._rebuild_service_runtimes constructed a fresh MetricsCollector for a service target added via PATCH /config but never started it -- without an HTTP session the new target collected nothing and never autoscaled; a removed target kept its session alive. The rebuild now records collectors to start/stop and the async PATCH path starts/stops them before publishing the new runtime set; start() consumes the pending list (regression tests: PATCH adds a started collector, PATCH removing the target stops it).
…apshots, bounded history Review findings on the scale control plane: - num_replicas is now a strict int: JSON true / "2" / 2.0 are rejected 422 at the HTTP boundary instead of being coerced by pydantic before the registry's isinstance checks (boundary test covers all three). - Scale submission fails closed when the authoritative capacity cannot be queried: the mutating path uses strict capacity lookups and returns 503 instead of submitting against a degraded snapshot where ready impersonates current; read-only /engines keeps its degraded view. - A watcher crash no longer zero-fills capacity: the terminal snapshot falls back to the registry's last known values (previous terminal snapshot, else the submit-time observation, now exposed in status) instead of inventing current=0/ready=0, and an entirely empty progress payload keeps cleanup_required set (fail closed) while a non-empty legacy snapshot without the key keeps its old meaning. - The registry is a long-running control plane: clean terminal operations are bounded to max_history (oldest evicted, counted in history_truncated), idempotency records whose clean outcome aged out of the replay window expire to fresh-request semantics, and dirty terminals / live operations are never evicted.
Review finding: the GenRM server args derived random_seed from the engine rank, so the elastic replica sampled differently from the initial one and the two replicas disagreed on borderline inputs under the official sampling config (measured 2/50 verdict flips in the reward consistency run). GenRM is a frozen scoring service: every replica must sample identically for an identical request, so the seed is now the base args.seed on every rank. The rollout builder keeps its rank-derived seed (sampling diversity is wanted there, not here).
Review findings on the autoscaler control plane: - Per-service evaluation now runs concurrently (asyncio.gather with per-service exception isolation): a stalled GenRM metrics endpoint can no longer delay the rollout service's evaluation cadence (regression test parks the GenRM evaluation and asserts rollout completes). - A legacy PATCH of rollout_service_url is no longer shadowed by the startup service_targets entry: the two spellings are kept in sync in both directions, so the patched URL actually takes effect (tests for both directions). - /status reports per-service last_scale_action / last_scale_time and the monitor's service view overlays them, so --service genrm shows GenRM's own scale history instead of the rollout runtime's. - Metrics aggregation is per-field: an engine whose scrape is missing only throughput still contributes its observed token_usage and queue depth to scale-out signals, while remaining an unknown engine for the conservative scale-in gate (cross-layer regression test). - Together with the collector hot start/stop from the previous commit, PATCHed service targets are fully live at runtime.
Review finding: the continuity verdict's single in-window check used the scale-out COMPLETION time as its lower bound, so the whole scale-out process was excluded and one train event between the two operations satisfied the assertion even when the 1 s scale-in window contained nothing. The verdict now asserts progress separately before the operations, between them, and after scale-in -- the honest, observable windows -- and the run-6 evidence is re-checked against them by the final-head rerun (the 1 s drain window is shorter than one training iteration; progress during it is not an observable claim).
…ction, autoscaler r4/r5, B2 smoke
…ertions, 8/8 rollouts
This comment has been minimized.
This comment has been minimized.
rai-studio-bot
left a comment
There was a problem hiding this comment.
已完成 af82e6a 复查:原先的 P1 安全问题均已修复;八项反馈已解决,reconcile 仍有一项 P2 恢复问题,已在原讨论补充并降低优先级。另有历史保留和跨服务评估节奏两项 P2,见新增行内意见。本轮仅留非阻塞意见。
69 项组件/状态机测试与 137 项 Autoscaler 测试通过;真实 Ray/SGLang/GPU 集成未重跑。
Evidence-provenance cleanup (review feedback):
- Remove all relative references ('final head' / 'final PR head' /
'this PR head' / 'final-head rerun') from EVIDENCE.md; every run now
names the commit that produced it: the seed-contract chain (score
consistency r3, adversarial divergence r3, training smoke r3, training
continuity r5) at 945741e; the round-1 fix batch (failure injection r2,
autoscaler prereg r4/r5, smoke r2) at their exact heads (e7224af /
8b0c6fe / fddb093); the preregistered v2 rounds at fa3ca97.
- State the pinning policy accordingly: the newest code exercised by any
recorded acceptance run is 945741e; earlier verdicts are pinned to
their own producing commits.
- Judge ground-truth correctness is now explicitly informational and
reported per retained run (95%/95% in _r2, 97%/97% in _r3); replica
agreement, not judge capability, is the acceptance criterion.
- Terminology: 'request-level sampling-seed contract' -> 'content-derived
sampling-seed contract' (same content reproduces the same seed across
replicas and repeated calls).
Artifact blobs are untouched; the SHA256 pin manifest still matches the
committed verdict files.
rai-studio-bot
left a comment
There was a problem hiding this comment.
审查结论:请求修改后再合并
本轮复审了上一轮两条 P1(4c137803)的修复,并复查弹性扩缩与 GenRM 扩缩注册表相关子系统;最新提交 16d7148 仅改动 demos/task4_genrm/EVIDENCE.md(证据行改为按产出提交固定),代码未变,行内意见锚点在该提交上依然成立。
上一轮意见(已修复,原线程已由作者关闭)
- GenRM 终态重放不再被当作清理完成:接管后同一轮继续以状态接口为唯一权威判据,只有明确干净才进入 history;恢复矩阵 A–F 的回归覆盖到位。
- 正常受理、内联同 key 重试、后台补投三条路径共用同一接管逻辑,冷却起点与操作账目一致,冷却一致性用例通过。
新发现(阻塞项)
- P1:三态清理闸门把「状态响应缺少
cleanup_required」判为未确认,但该字段只存在于 GenRM 的状态契约;autoscaler 的默认目标rollout的状态响应模型未声明它,导致 rollout 的扩缩操作到达终态后无法收尾,进而永久冻结后续决策,并让 target 删除守卫持续返回 409。详见relax/utils/autoscaler/autoscaler_service.py:1017的行内意见(含实证与本机复现,另有两条修复方向可选)。
验证范围与缺口
- 本机运行 Autoscaler 全部用例:149 项通过;另有 11 项因沙箱缺少
aiohttp未能运行(环境限制,与本次改动无关)。 - CI 检查 8/8 通过,但现有新增用例的替身均按 GenRM 契约返回该字段,未覆盖非 GenRM 服务契约路径,故 CI 绿灯不能证明上述回归不存在。
- 生产/测试必要性与简化审计结论已记录在看板,本轮未提出新的简化类意见。
|
描述中新增的状态小节与当前审查状态不一致,建议同步更新,避免维护者据此判断合并条件:
修复顺序建议:先把 rollout 契约路径修好或明确作用域,再更新这两处描述(例如:两条 round-8 问题已复验通过;新增 P1 待修复;
Powered by Nyanpasu with deepseek-v4.1-flash-ali xhigh, please check the suggestions carefully.
|
…-Blackwell Final-code Autoscaler acceptance reconfirmation (Criterion 4, frozen protocol) exposed a product regression introduced by the deterministic sampling activation (945741e): Reproduction (identical frozen driver, thresholds, load curve and machine, same day): - final code (deterministic ON): the engine wedged mid-load -- the scheduler stopped completing requests (served frozen at 843, 48 requests hung, /health no answer, GPU util 0%), the elastic replica served 0 requests, and scale-in never became possible; the run failed. - pre-deterministic code (8b0c6fe, r5's product code): 13/14 functional assertions passed, elastic replica served 461 requests in STEADY and automatic scale-in completed. Root cause: SGLang applies the generic flashinfer attention-backend default BEFORE its deterministic-inference handler. On pre-Blackwell GPUs the handler's own deterministic fallback (fa3) is therefore preempted, flashinfer is kept, and the radix cache is force-disabled (flashinfer is not in SGLang's radix-supported deterministic set). Without prefix reuse, a 48-way concurrent scoring load re-prefills every long shared prompt until the scheduler wedges. Fix: select the attention backend SGLang itself recommends for deterministic inference on the detected architecture (fa3 on pre-Blackwell, flashinfer on Blackwell+), so the radix cache stays available. The per-request sampling-seed contract is unchanged (deterministic mode stays on; the pytorch sampling backend still consumes per-request seeds). Per-instance --genrm-engine-config overrides keep the highest priority. Regressions: deterministic attention-backend selection (fa3 pre- Blackwell, flashinfer Blackwell+, probe failure defers to the SGLang default), engine-config override precedence, and the server-args pipeline carrying attention_backend alongside enable_deterministic_inference.
FINAL_EVIDENCE_SHA = 6b501a2 (raw artifacts: the final-code confirmation r3, the regression reproduction r2, the environment-invalid r1, the A/B discrimination run at 8b0c6fe, and the four machine-blocked score-consistency attempts). EVIDENCE.md: the final-code Autoscaler confirmation row (frozen protocol, 14/14 at 5c1e2e7, deterministic sampling under the radix-preserving fa3 attention backend); the regression and environment failure records in the failed-runs table; the pinning policy now names 5c1e2e7 as the newest code exercised by any recorded acceptance run; the determinism known-limit notes that fa3 keeps the radix cache available on pre-Blackwell GPUs; the SHA256 manifest gains the final verdict blob.
…module Suite-wide dependency-stub tests can replace sglang.srt.server_args while relax.backends.sglang.sglang_engine is already imported, leaving its module-level ServerArgs bound to a non-dataclass stub that _compute_genrm_server_args cannot iterate. The two new deterministic attention-backend tests now skip on what the engine module actually sees (2 skips in the stubbed context, 9 passed standalone); the full suite's failure set is byte-identical to the pre-change baseline.
This comment has been minimized.
This comment has been minimized.
rai-studio-bot
left a comment
There was a problem hiding this comment.
本轮结论:仍需修改——此前提出的 P1(三态清理闸门与 rollout 状态契约不匹配,而 rollout 是 autoscaler 的默认目标)在最新 head a4f55d7 复核后依然存在,合并条件不变;已在原线程补充复核证据(discussion_r4110171098)。
本轮增量(16d7148f → a4f55d7:5c1e2e7 的确定性注意力后端修复、3127b63 的证据文档、a4f55d7 的测试跳过防护)另有两处非阻塞意见:
- P2:显式写入的
"attention_backend"会静默覆盖全局--sglang-attention-backend(含探测失败的分支),非 CUDA 平台会被确定性写成fa3;建议改为「全局未设置时才注入」或限定在 CUDA 设备上,并补一条优先级用例。 - P3:注释中「让 radix 缓存保持可用」在 Blackwell+ 上不成立(仍走
flashinfer→disable_radix_cache),建议在注释与 known limits 中写明该边界。
核对信息:本轮审查 head 为 a4f55d7(目标 head 与 PR head 一致,两条行内意见均标注在该 head 上),该 head 的 8 项 CI 检查全部通过。范围决策未变:接收 36 / 建议移出 1 / 待确认 15,移出与待确认部分不计入本轮结论。
Re-review P1: the three-state cleanup gate (UNKNOWN/False/True, missing
!= clean) treats every service status response without an explicit
cleanup_required as UNKNOWN, but that flag only existed in GenRM's
status contract. Rollout's ScaleOutStatusResponse and
ScaleInStatusResponse never declared it, so FastAPI's response_model
stripped the key from the wire and the shared autoscaler -- whose
default target is rollout -- read UNKNOWN for every rollout terminal
operation:
- terminal scale-out (ACTIVE/PARTIAL/FAILED/CANCELLED) and scale-in
(COMPLETED/FAILED) stayed pending_requests forever;
- the next scaling evaluation stayed frozen ("in progress or awaiting
cleanup") with no path to unblock;
- a removable rollout-contract target kept failing PATCH removal with
409 (stale unresolved pending);
- history never recorded the finished operation.
Rollout has no deferred-cleanup lifecycle: a FAILED scale-out rolls its
engines back before reporting and a COMPLETED scale-in reports the
engines removed, so terminal status is cleanup-complete by contract.
The fix therefore makes the rollout status models state that explicitly
(cleanup_required: bool = False) instead of leaving the flag absent.
GenRM's semantics are unchanged: a GenRM POST reply still proves
nothing, a missing/malformed flag still reads UNKNOWN, terminal-dirty
still requires status/reconcile, and finalization still requires an
authoritative cleanup_required is False.
Coverage:
- response-model serialization tests pin model_dump()["cleanup_required"]
is False for both scale-out and scale-in wire shapes (the response-
model tests now import the component through tests.utils._dep_stubs
so they actually run where transfer_queue is absent);
- shared-autoscaler regressions drive a real AutoscalerService against
a fake speaking the rollout HTTP contract through the actual Pydantic
response models: scale-out and scale-in terminal requests finalize
into history, the next decision is no longer blocked, and a finalized
target no longer 409s on removal;
- the fail-closed guard stays pinned: a rollout-contract response with
the flag dropped (pre-fix wire shape) keeps the request pending and
the decision frozen -- the fix lives in the contract, never in the
autoscaler defaulting missing to clean.
…ic probe declines Re-review P2: _compute_genrm_server_args pinned attention_backend unconditionally -- including the architecture probe's None -- so the sglang_* inheritance loop (which backfills only fields absent from kwargs) could never apply the user's global --sglang-attention-backend to a GenRM engine again. The docstring's "falls back to SGLang default" only held for SGLang's internal default; the relax-level global flag was silently disabled by 5c1e2e7. The deterministic backend is now pinned only when the user has not chosen one globally: - explicit --sglang-attention-backend wins over the probe (restores the pre-5c1e2e7 inheritance behaviour); - a failed probe defers to SGLang's default instead of pinning None over it (and over the user's global choice); - the arch-appropriate deterministic backend (fa3 on pre-Blackwell, flashinfer on Blackwell+, radix cache preserved) is still pinned when the user set nothing, keeping the final-code autoscaler acceptance behaviour on this machine; - per-instance --genrm-engine-config overrides still take priority over both (applied after the loop). Regressions: - an explicit global choice beats the probe result; - a failed probe pins nothing (key absent, SGLang default applies) while the user's global choice still flows through the inheritance loop.
This comment has been minimized.
This comment has been minimized.
rai-studio-bot
left a comment
There was a problem hiding this comment.
复审结论:上一轮的两条意见已在 04817012 验证修复,此前唯一的 P1(rollout 状态契约缺 cleanup_required,默认目标的已终态操作永不收尾)已消除,本轮增量无新增缺陷;因 demos/task4_genrm 证据文件的去向仍待维护者裁定(15 个待确认 + 1 个建议移出,见看板),本轮不给出整体通过。
- P1 已修复(rollout 清理契约):两个状态模型显式声明
cleanup_required: bool = False,与终态语义一致(scale-out FAILED 前已回滚、scale-in COMPLETED 前已移除引擎)。本机以真实AutoscalerService直连复测:声明后终态操作 pending→history、后续决策解冻;字段缺失(修复前形状)仍保持 fail-closed 永久 pending。序列化与生命周期回归齐备,含把修复前形状钉死为「保持 pending + 决策冻结」的守卫用例(详见 原线程)。 - P2 已修复(attention_backend 优先级):仅在全局
--sglang-attention-backend未设置时注入,探测失败不再写入None;两条优先级用例补齐,NPU/KLX 配方不再被覆盖(详见 原线程)。 - P3 已按已知限制记录(radix / Blackwell+):线程回复与
EVIDENCE.md的 pre-Blackwell 边界表述覆盖该提醒,后续按 upstream-dependent 跟进。 - 验证环境说明:本机任务内测试 249 项通过(rollout 契约与 sglang 后端用例因沙箱缺 torch/transfer_queue/sglang 跳过,需 CI 全量环境执行);该 head 的 required workflows 仍在等待维护者批准运行。
范围决策:接收 39 / 建议移出 1 / 待确认 15;demos/task4_genrm 证据文件的仓库去向(留 demos/、移 examples/、或仓库仅留 README/EVIDENCE 链接指向证据分支)需要维护者拍板,裁定前不作整体通过。
rai-studio-bot
left a comment
There was a problem hiding this comment.
补充意见:该 head 的 CI 中 Pre-commit Checks 失败——9755486 在 tests/utils/data/test_identity_window_sampler.py 的模块级守卫之后留下两处导入,触发 ruff E402×2(行内已附最小修复建议,P3 非阻塞,修复合入前请重跑 pre-commit)。Unit Tests (H20, 4 GPUs) 仍在运行,其中包含上轮本机沙箱无法执行的 rollout 契约与 sglang 优先级用例。其余结论同上一轮:两条复审意见已验证修复、无新增功能缺陷;范围决策(接收 39 / 建议移出 1 / 待确认 15)仍待维护者裁定。
The module-level transfer_queue attribute guard (added with the rollout contract stub support) places two imports after executable code; ruff's E402 fires on the pushed tree and would fail the Pre-commit workflow. Annotate both imports as guarded.
|
当前 PR 头 请维护者确认三项范围与行为:
各实验的运行版本、结果与限制见 PR 正文,设计见 RFC #351。训练连续性以 再分析 v2.1为准:运行完成已验证,未证明零性能影响。 范围或语义裁决若要求改动,再提交最小增量并重新走 CI / review。 |
rai-studio-bot
left a comment
There was a problem hiding this comment.
当前 head f507d3b7 的复审收口:全部 26 项审查问题已解决并验证——最后一条(pre-commit E402 注解)已在该 head 修复(ruff 复核通过,见原线程);此前 0481701 上的两条复审修复(rollout 清理契约、attention_backend 优先级)均经真实 AutoscalerService 直连实测与回归用例验证。该 head 的 8 项 CI 检查全部通过(含 H20 4-GPU 单测与三项集成测试),25 个 review 线程均已解决。
本轮不给出整体通过的唯一原因仍是范围决策:demos/task4_genrm 证据文件的去向(作者已在交接评论中列出 A/B/C 三个选项)需要维护者裁定,裁定前 15 个待确认与 1 个建议移出的路径保持可见(范围:接收 39 / 建议移出 1 / 待确认 15)。范围裁定(以及若需移动文件产生的新 head 复核)完成后即可给出最终结论。
rai-studio-bot
left a comment
There was a problem hiding this comment.
补充意见(非阻塞):PR 描述把最终代码的 ② 重确认记为 PASS,证据确实存在——已在证据分支新顶端核对 reward_consistency_20260927_final_r6 与 sampling_divergence_20260927_final_r4 两个运行目录(旧顶端无、新顶端有)。但记录有两处待修:PR 描述中新加的证据分支链接使用了不存在的完整 SHA(1deb9d83f79b…,实际分支顶端为 1deb9d830710a923ba561c99a32106c00cfa325f),且 demos/task4_genrm/EVIDENCE.md 未收录这两个运行、文件头分支指针仍停在旧顶端 6b501a2(行内意见已附具体修改建议)。其余状态与上一轮一致:26 项审查问题全部解决并验证、CI 8/8、25 线程 resolved;范围裁定(demos 去向 A/B/C、验收规模)仍待维护者。
…tion PASS, r5 proxy incident, train-continuity reanalysis v2) Evidence-branch pointer update only; no product change. Raw artifacts live on the evidence branch @ d88d6c7 (r6/r4 verdict trees, the archived raw job log with sha256, and the reanalysis v2 script + parser tests).
…ons table Per review: the top of the file still pointed at evidence 6b501a2 and called 5c1e2e7 the newest acceptance code while the bottom appended 0481701 reconfirmations — an index that contradicted itself. The top is now the only current-conclusions table: one row per official criterion, each bound to the exact product SHA, the producing run and its raw inputs on evidence 66f1f78, with an explicit superseded-by column. Reading rules state the caliber of each claim (0/50 = these 50 inputs under this config; GenRM-path-only weight sync; criterion 5 is continuity, no performance claim absent a control run; content-deterministic semantics pending maintainer acceptance). A scope commitments table separates delivered / not-implemented / to-be-negotiated (cross-Gateway wording now 'not implemented this phase; interface and responsibility to be negotiated'). The per-run history below is marked as historical detail and keeps its own producing commits.
This comment has been minimized.
This comment has been minimized.
rai-studio-bot
left a comment
There was a problem hiding this comment.
状态更新:5b512d0 的 8 项 CI 检查已全部通过(含 pre-commit、Python 3.10–3.12、H20 GPU 单测与三项训练集成)——此前记录的最后一个 CI 门槛清零;描述重写(What/Why/How/Testing)与新增的连续性可视化更正均与已核实记录一致(诚实限缩:不证明步延迟不变、撤回请求数归因)。全部 27 项审查问题维持 resolved/superseded。
本轮新增一条非阻塞意见:新增的扩缩端点未进入用户侧 API 文档(docs/en/api/genrm.md 仍只覆盖打分服务),而描述指其「describes the public contract」——建议补文档或改指向(行内已附,可与范围裁定的文档去向一并处理)。整体通过仍待三项维护者决策(demos 去向、验收规模、采样语义)。
- docs/en/api/genrm.md + docs/zh/api/genrm.md: endpoints table, request/ status contract, NOOP/idempotency semantics, scale-out/in state chains with terminal sets, cleanup_required and capacity fields (current/ready/ occupied/pending_cleanup), strict-capacity 503 and fallback semantics - docs/public/openapi/genrm.json: regenerate via scripts/tools/ generate_openapi.py; 10 paths, 9 schemas, descriptions from code docstrings - scripts/tools/generate_openapi.py: write trailing newline so reruns are byte-stable Verified against relax/components/genrm.py, relax/utils/genrm_scale_registry.py and relax/distributed/ray/genrm.py: every documented route, state, terminal set, capacity formula and error code matches the implementation.
rai-studio-bot
left a comment
There was a problem hiding this comment.
增量复核(5b512d0 → 9095c66,文档提交):GenRM 扩缩契约的 en/zh API 文档与 OpenAPI 规范已核对——端点、严格整数校验(422)、幂等键重放语义(有界内存史、非跨重启保证)、三态清理与 reconcile 行为、capacity 字段含义及单 Gateway 范围警示均与实现一致;genrm.json 与生成器输出同步、schema 含 cleanup_required。F28 以两种方式(指向修正 + 正式文档)落实,已关闭。
一条新的非阻塞意见(行内):rollout 侧由 F22 修复引入的 cleanup_required 字段未同步到 docs/public/openapi/rollout.json 与 docs/en|zh/api/rollout.md——GenRM 侧已补齐,建议 rollout 侧一并补上(注意重跑生成器会带上 main 既有的 /predict 无关 drift)。范围决策随文档组更新为:接收 43 / 建议移出 1 / 待确认 15;该 head 的 CI 待批准运行。整体结论不变:三项维护者决策(demos 去向、验收规模、采样语义)待定。
- docs/public/openapi/rollout.json: add cleanup_required (bool) to ScaleOutStatusResponse and ScaleInStatusResponse properties — manual patch in generator-faithful style (type + title, not in required), avoiding the unrelated /predict drift a full regeneration would bring - docs/en/api/rollout.md + docs/zh/api/rollout.md: new 'Scale status cleanup contract' section (bilingual semantic parity) documenting the three-state contract shared with the Autoscaler (true = cleanup still pending; false = authoritative complete; absent = unknown, missing != clean) and why rollout terminal states report false (scale-out engines serve or were rolled back; scale-in COMPLETED removed / FAILED rolled back) Docs-only; no product code touched.
# 📝 Documentation - Point the two evidence index links to the reviewed fdf288d snapshot. - Leave product code, historical verdicts and experiment versions unchanged. - Keep this commit local; the public PR head and its CI remain unchanged.
# 📝 Documentation - Withdraw whole-run zero-error and stall-bound claims using reanalysis v2.1. - Distinguish log-reported sample counts from file-level sample verification. - Keep original verdicts and producing versions; mark superseded interpretations. - Correct the long-request drain description and completed reconfirmation status. - Preserve product code, GPU results and public PR history.
rai-studio-bot
left a comment
There was a problem hiding this comment.
本轮描述更新撤回了训练连续性运行的两个旧断言(「全程零错误」与「全程停顿不超过 120 秒」),撤回依据经独立核实成立:再分析 v2.1 的日志分类(三条非良性带时间戳错误记录、judge 错误行 0、双副本稳定窗口内错误行 0)与归档原始日志一致,且「事件间隔不能作为停顿上界」的说明正确。这是又一次有效的诚实性修正。
一条新的非阻塞意见(行内):demos/task4_genrm/EVIDENCE.md 的训练连续性历史行仍原样保留这两个已撤回的断言,未按该文档自身「superseded rows say so in-line」的约定做标注——读者经合并门槛指向查阅清单时会看到与描述矛盾的未标注声明;建议补内联撤回标注并顺带把再分析引用更新到 v2.1。其余状态不变:29 项审查问题全部解决、CI 8/8@当前 head、三项维护者决策待定。



改动目的
为冻结奖励模型 GenRM 增加弹性副本生命周期,并接入现有 Autoscaler。负载升高时,新引擎通过健康检查后接流量;负载回落时,排空弹性副本并回收资源,保留初始引擎。
设计与取舍见 RFC #351,字段与错误码见 中文 API / OpenAPI。本期范围为单 Gateway 排空、单 GPU 弹性副本。
da4acbb;产品0481701,relax/产品代码保持不变a55c85f8/8 通过:pre-commit、三版 Python、H20 4-GPU 单测、三项训练集成(含一项 NPU);da4acbb仅文档变更,CI 重跑中产品头之后仅变更 API / 证据文档、OpenAPI 生成脚本及两处测试 lint 注释,运行时代码未变。
实现结构
flowchart TB U["手动扩缩请求"] --> C["GenRM 控制 API<br/>绝对目标 · 幂等 · 互斥"] C --> M["GenRMManager<br/>PG 所有权 · 生命周期"] M --> E["初始 / 弹性引擎<br/>健康检查后发布"] M -.->|"可路由集合"| G["GenRM Gateway<br/>admission · 在途请求计数"] T["训练评分请求"] --> G G --> E G -.->|"指定 generation 排空证明"| C E -.->|"有效负载指标"| A["分服务 Autoscaler<br/>持续条件 · 冷却 · history"] A --> C classDef control fill:#e8f0fe,stroke:#4169a1,color:#172b4d; classDef serving fill:#e5f4ee,stroke:#328268,color:#173d31; class U,C,M,A control; class T,G,E serving;实线为调用,虚线为状态、发现或观测关系。GenRM 扩缩不引入训练权重同步;actor→rollout 原有同步仍然保留。
建议审查顺序
recover()不重建已缩掉的副本cleanup_required分离;无法取得权威容量时拒绝提交service_policies--service genrm;按服务展示状态、条件、历史;动态容量失败路径的行为
ACTIVE / PARTIAL / FAILED;PARTIAL保留已发布副本,不代表达到目标。REMOVED。cleanup_required,不把原FAILED改成成功。current / ready / occupied / pending_cleanup:未发布候选不增加服务容量;失败缩容 victim 未释放前仍计入已发布容量与资源占用。正常链路、超时收尾时序和容量示例见 RFC 架构说明。
验证
原始证据固定于
fdf288d,各实验的运行版本见下表。评分复确认与任务相关 CPU 回归在最终产品0481701上完成。a2ca6cb0481701048170104817015c1e2e7e7224af945741e训练结论采用再分析 v2.1。仓库旧索引的“全程零错误”和“全程停顿不超过 120 秒”表述已撤回:日志有三条非良性带时间戳错误记录,judge 错误行数与双副本稳定窗口内的错误行数均为 0;事件间隔不能作为全程停顿上界。
其余限制:
REMOVED墓碑,不能用原始表长度判泄漏。真实验收截图:GenRM 的状态、扩缩条件与操作历史
截图来自自动扩缩最终轮,固定于上述 evidence commit;用于检查观测界面,不替代整轮事件与 verdict。Rollout 对照截图与原始记录保留在同一目录。
兼容性与未覆盖范围
旧 Rollout 顶层字段和默认 TUI 行为保留。两服务的冷却计时状态独立,但冷却时长配置仍由 deployment 提供。GenRM 自动扩缩目标目前只支持单模型实例,发现多个活跃实例时拒绝混合处理。
单 Gateway 计数不覆盖 direct client / 跨 Gateway 的完整排空;组件重启后的持久幂等、Manager 重启后的在途恢复、多 GPU 弹性副本,以及弹性操作与运行时 onload/offload 的协调,均不在本期范围。
请维护者确认
当前头仍需正式 review。不可变证据目录保留全部已归档轮次及各自判定。