Skip to content

Commit 46e5160

Browse files
committed
fix(studio): take the migration analysis as the model wrote it
The analysis turn moved from `codex exec --output-schema` to a Sandbox app-server dynamic tool, where the JSON Schema is only a declaration: a turn that answered with a progress note, a Markdown fence or a probe-shaped payload reached Studio's verifier anyway, and every rejection collapsed into one sentence that named neither the field nor the fix. The first submission that passed also became final, so the failing production run turned a plain Dify export into `MIGRATION_ANALYSIS_UNSUPPORTED` while the deterministic detector the image ships had silently found nothing (PyYAML lives on the 3.12 interpreter, the detector asked `/usr/bin/python3`). Studio now reads the archive itself before the turn: `migration/detection.py` walks the upload, parses whatever DSL is inside and records the files, the framework candidates and the evidence lines it actually saw, with a degraded report instead of a silent empty one. The turn gets one tool per verdict (`reportRecommendation`, `reportNeedsInput`, `reportUnsupported`) and `migration/analysis_contract.py` assembles the stored document, so the model's side of the contract is flat and forgiving: summary is the only required field, unknown keys, unknown frameworks and odd types are folded into notes, scalar values are normalised, and the detected candidates are always merged in. The bar that stays is the one that protects the user: `unsupported` needs two evidence lines that exist in the detected file list, and it fails open when detection itself is degraded. When a turn delivers nothing usable, Studio answers from the detection report (a conservative recommendation, the detected framework first) instead of spending the scripted driver on a second analysis, and the outcome is written to `diagnostics/analysis/model-turn.json` so the reason is readable on the page. The two real Baidu Qianfan and Dify exports that failed in production now reach `analysis_ready` with model-authored recommendations, and the seven payloads from the failed rollout are kept as regression cases.
1 parent 0e89974 commit 46e5160

11 files changed

Lines changed: 2381 additions & 357 deletions

‎frontend/server/migration/analysis_contract.py‎

Lines changed: 892 additions & 0 deletions
Large diffs are not rendered by default.

‎frontend/server/migration/app_server.py‎

Lines changed: 154 additions & 42 deletions
Original file line numberDiff line numberDiff line change
@@ -14,11 +14,13 @@
1414

1515
"""Codex app-server route analysis for Studio project migration.
1616
17-
The migration analysis result is a strict contract object. Delivering it through a
18-
registered dynamic tool keeps the payload in typed JSON-RPC arguments, so a progress
19-
update, a commentary message, or a Markdown-fenced reply can no longer be mistaken for
20-
the result. Rejections are returned to Codex as ``success: false`` so the same turn can
21-
correct itself instead of failing the whole migration.
17+
The analysis result arrives through registered dynamic tools, so a progress update, a
18+
commentary message, or a Markdown-fenced reply can no longer be mistaken for the result.
19+
Acceptance is deliberately liberal: Studio owns the state-file format, fills in the
20+
protocol bookkeeping itself, and defaults or drops whatever the model could not shape,
21+
so a model that writes imperfect arguments still lands a usable result. Only the
22+
destructive ``unsupported`` verdict is held to a deterministic bar, and a refusal is
23+
returned to Codex as ``success: false`` so the same turn corrects itself.
2224
2325
The same turn also carries ``askUser``. When the project cannot answer a question that
2426
changes the migration, Codex asks the user inside the turn and keeps analysing with the
@@ -42,21 +44,49 @@
4244
Answers,
4345
normalize_questions,
4446
)
47+
from .analysis_contract import (
48+
NEEDS_INPUT_KIND,
49+
RECOMMENDATION_KIND,
50+
UNSUPPORTED_KIND,
51+
AnalysisAcceptanceError,
52+
AnalysisAssemblyError,
53+
acceptance_feedback,
54+
analysis_tool_schema,
55+
build_analysis_result,
56+
)
4557
from .codex_tool_turn import DynamicTool, ToolTurnUnavailable, run_tool_turn
46-
from .contracts import MigrationContractError, validate_analysis_result
4758

4859
logger = logging.getLogger(__name__)
4960

50-
ROUTE_TOOL_NAME = "reportRoute"
51-
ROUTE_TOOL_DESCRIPTION = (
52-
"提交只读项目分析的最终结果。必须在完成分析后调用一次,"
53-
"参数严格遵循给定的 JSON Schema;被拒绝时按返回的错误修正后重新调用。"
54-
)
61+
RECOMMENDATION_TOOL_NAME = "reportRecommendation"
62+
NEEDS_INPUT_TOOL_NAME = "reportNeedsInput"
63+
UNSUPPORTED_TOOL_NAME = "reportUnsupported"
64+
TOOL_NAME_BY_KIND = {
65+
RECOMMENDATION_KIND: RECOMMENDATION_TOOL_NAME,
66+
NEEDS_INPUT_KIND: NEEDS_INPUT_TOOL_NAME,
67+
UNSUPPORTED_KIND: UNSUPPORTED_TOOL_NAME,
68+
}
69+
TOOL_DESCRIPTION_BY_KIND = {
70+
RECOMMENDATION_KIND: (
71+
"提交只读项目分析的最终结论:推荐一种可执行的迁移方式。"
72+
"在完成分析后调用一次;只有 summary 是必填的,其余字段能给多少给多少,"
73+
"缺失的字段会被 Studio 用已核实的事实补齐,不会被拒绝。"
74+
),
75+
NEEDS_INPUT_KIND: (
76+
"提交只读项目分析的结论:必须先由用户补充信息才能决定迁移方式。"
77+
"把要向用户提出的问题写入 questions(每项含 prompt)。"
78+
),
79+
UNSUPPORTED_KIND: (
80+
"提交「该项目无法迁移」的结论。这是最后一个手段,只用于材料不足或"
81+
"证据完整的高风险行为链;必须给出 summary 和至少两条指向项目真实文件的"
82+
"证据,证据无法核实会被拒绝。"
83+
),
84+
}
5585
# What Codex is told when nobody answered the questions in time.
5686
UNANSWERED_HINT = (
5787
"用户没有在时限内回答这些问题。请立即调用 "
58-
f"{ROUTE_TOOL_NAME} 并返回 status=needs_input,"
59-
"把原始问题写入 questions(每项包含 id、prompt、required),不要重复提问。"
88+
f"{NEEDS_INPUT_TOOL_NAME},"
89+
"把原始问题写入 questions(每项包含 prompt),不要重复提问。"
6090
)
6191

6292
_APP_SERVER_ENV = "AGENTKIT_MIGRATION_APP_SERVER"
@@ -86,34 +116,79 @@ class MigrationAnalysisUnavailable(RuntimeError):
86116
"""The Sandbox app-server could not produce a usable analysis result."""
87117

88118

89-
class RouteRecorder:
90-
"""Validate and retain one analysis result delivered by a dynamic tool call."""
119+
class AnalysisRecorder:
120+
"""Accept one terminal analysis submission and retain the assembled result.
91121
92-
def __init__(self, *, attempt: int, input_sha256: str) -> None:
122+
The record carries the outcome facts a caller may want to persist: which tool
123+
landed, what acceptance had to default or drop, and why earlier submissions were
124+
refused. None of that is a verdict about the project.
125+
"""
126+
127+
def __init__(
128+
self,
129+
*,
130+
attempt: int,
131+
input_sha256: str,
132+
detection: dict[str, object] | None = None,
133+
) -> None:
93134
self.attempt = attempt
94135
self.input_sha256 = input_sha256
136+
self.detection = detection
95137
self.result: dict[str, object] | None = None
96-
self.rejections: list[str] = []
97-
98-
def submit(self, arguments: dict[str, object]) -> CodexDynamicToolResult:
99-
candidate = {
100-
**arguments,
101-
"attempt": self.attempt,
102-
"input_sha256": self.input_sha256,
103-
}
138+
self.kind = ""
139+
self.notes: list[str] = []
140+
self.refusals: list[str] = []
141+
142+
def handler(
143+
self,
144+
kind: str,
145+
) -> Callable[[dict[str, object]], CodexDynamicToolResult]:
146+
def handle(arguments: dict[str, object]) -> CodexDynamicToolResult:
147+
return self._accept(kind, arguments)
148+
149+
return handle
150+
151+
def _accept(
152+
self,
153+
kind: str,
154+
arguments: dict[str, object],
155+
) -> CodexDynamicToolResult:
156+
if self.result is not None:
157+
return CodexDynamicToolResult(
158+
True,
159+
"分析结果已经提交,请直接给出简短的简体中文总结。",
160+
)
104161
try:
105-
validated = validate_analysis_result(candidate)
106-
except MigrationContractError as error:
107-
self.rejections.append(str(error))
162+
document, notes = build_analysis_result(
163+
kind,
164+
arguments,
165+
attempt=self.attempt,
166+
input_sha256=self.input_sha256,
167+
detection=self.detection,
168+
)
169+
except AnalysisAcceptanceError as error:
170+
self.refusals.append(
171+
"; ".join(f"{item.path}:{item.actual}" for item in error.issues)
172+
)
173+
return CodexDynamicToolResult(False, acceptance_feedback(kind, error))
174+
except AnalysisAssemblyError:
175+
logger.exception("Studio migration analysis assembly failed kind=%s", kind)
108176
return CodexDynamicToolResult(
109177
False,
110-
f"分析结果不符合协议({error})。请修正后重新调用 {ROUTE_TOOL_NAME}。",
178+
"Studio 暂时无法保存这次结论,请稍后重新调用同一个工具。",
179+
)
180+
self.result = document
181+
self.kind = kind
182+
self.notes = notes
183+
if notes:
184+
logger.info(
185+
"Studio migration analysis accepted with defaults kind=%s notes=%s",
186+
kind,
187+
notes,
111188
)
112-
if self.result is None:
113-
self.result = validated
114189
return CodexDynamicToolResult(
115190
True,
116-
"分析结果已接收。请用简体中文给出简短的用户可见总结。",
191+
"结论已接收。请用简体中文给出简短的用户可见总结,不要重复分析过程。",
117192
)
118193

119194

@@ -168,7 +243,6 @@ async def run_route_analysis(
168243
*,
169244
endpoint: str,
170245
prompt: str,
171-
schema: dict[str, object],
172246
cwd: str,
173247
attempt: int,
174248
input_sha256: str,
@@ -178,13 +252,22 @@ async def run_route_analysis(
178252
questioner: AnalysisQuestioner | None = None,
179253
idle_timeout_seconds: float | None = None,
180254
host_wait_seconds: Callable[[], float] | None = None,
255+
detection: dict[str, object] | None = None,
256+
diagnostics: dict[str, object] | None = None,
181257
) -> dict[str, object] | None:
182-
"""Run one analysis turn and return the validated route contract, if any.
258+
"""Run one analysis turn and return the accepted analysis document, if any.
183259
184260
Passing ``questioner`` registers ``askUser`` for this turn; the caller also owns
185261
the matching window through ``idle_timeout_seconds`` and ``host_wait_seconds``.
262+
``detection`` carries Studio's own verified facts, which acceptance uses to fill
263+
defaults and to check the one verdict that must cite real files. ``diagnostics``,
264+
when given, is filled in place with what happened, so a caller can persist it.
186265
"""
187-
recorder = RouteRecorder(attempt=attempt, input_sha256=input_sha256)
266+
recorder = AnalysisRecorder(
267+
attempt=attempt,
268+
input_sha256=input_sha256,
269+
detection=detection,
270+
)
188271
extra_tools = (
189272
(
190273
DynamicTool(
@@ -197,34 +280,63 @@ async def run_route_analysis(
197280
if questioner is not None
198281
else ()
199282
)
283+
# Counting Codex' own output is what lets a caller tell "the turn never reached
284+
# the model" (an infrastructure fallback) from "the model worked and delivered
285+
# nothing acceptable" (a conclusion Studio must fall back for itself).
286+
seen = {"events": 0}
287+
288+
def sink(event: object) -> None:
289+
seen["events"] += 1
290+
if event_sink is not None:
291+
event_sink(event)
292+
293+
tools = tuple(
294+
DynamicTool(
295+
name=TOOL_NAME_BY_KIND[kind],
296+
description=TOOL_DESCRIPTION_BY_KIND[kind],
297+
schema=analysis_tool_schema(kind),
298+
handler=recorder.handler(kind),
299+
)
300+
for kind in (RECOMMENDATION_KIND, NEEDS_INPUT_KIND, UNSUPPORTED_KIND)
301+
)
200302
try:
201303
await run_tool_turn(
202304
endpoint=endpoint,
203305
prompt=prompt,
204306
cwd=cwd,
205-
tool_name=ROUTE_TOOL_NAME,
206-
tool_description=ROUTE_TOOL_DESCRIPTION,
207-
tool_schema=schema,
208-
handler=recorder.submit,
307+
tools=tools,
209308
has_result=lambda: recorder.result is not None,
210309
model=model,
211310
timeout_seconds=timeout_seconds,
212-
event_sink=event_sink,
311+
event_sink=sink,
213312
extra_tools=extra_tools,
214313
idle_timeout_seconds=idle_timeout_seconds,
215314
host_wait_seconds=host_wait_seconds,
216315
)
217316
except ToolTurnUnavailable as error:
218317
raise MigrationAnalysisUnavailable(str(error)) from error
318+
if diagnostics is not None:
319+
diagnostics.update(
320+
{
321+
"accepted": recorder.result is not None,
322+
"kind": recorder.kind,
323+
"notes": list(recorder.notes),
324+
"refusals": list(recorder.refusals),
325+
"events": seen["events"],
326+
}
327+
)
219328
return recorder.result
220329

221330

222331
__all__ = [
223332
"AnalysisQuestioner",
333+
"AnalysisRecorder",
224334
"MigrationAnalysisUnavailable",
225-
"ROUTE_TOOL_DESCRIPTION",
226-
"ROUTE_TOOL_NAME",
227-
"RouteRecorder",
335+
"NEEDS_INPUT_TOOL_NAME",
336+
"RECOMMENDATION_TOOL_NAME",
337+
"TOOL_DESCRIPTION_BY_KIND",
338+
"TOOL_NAME_BY_KIND",
339+
"UNSUPPORTED_TOOL_NAME",
228340
"app_server_analysis_enabled",
229341
"ask_tool_handler",
230342
"run_route_analysis",

‎frontend/server/migration/codex_tool_turn.py‎

Lines changed: 17 additions & 13 deletions
Original file line numberDiff line numberDiff line change
@@ -85,10 +85,11 @@ async def run_tool_turn(
8585
endpoint: str,
8686
prompt: str,
8787
cwd: str,
88-
tool_name: str,
89-
tool_description: str,
90-
tool_schema: dict[str, object],
91-
handler: ToolHandler,
88+
tool_name: str = "",
89+
tool_description: str = "",
90+
tool_schema: dict[str, object] | None = None,
91+
handler: ToolHandler | None = None,
92+
tools: tuple[DynamicTool, ...] | None = None,
9293
has_result: Callable[[], bool],
9394
thread_id: str = "",
9495
model: str = "",
@@ -120,15 +121,18 @@ async def run_tool_turn(
120121
session.cwd = cwd
121122
if model:
122123
session.model = model
123-
for tool in (
124-
DynamicTool(
125-
name=tool_name,
126-
description=tool_description,
127-
schema=tool_schema,
128-
handler=handler,
129-
),
130-
*extra_tools,
131-
):
124+
if tools is None:
125+
if not tool_name or tool_schema is None or handler is None:
126+
raise ValueError("run_tool_turn needs either tools or one primary tool")
127+
tools = (
128+
DynamicTool(
129+
name=tool_name,
130+
description=tool_description,
131+
schema=tool_schema,
132+
handler=handler,
133+
),
134+
)
135+
for tool in (*tools, *extra_tools):
132136
session.register_dynamic_tool(
133137
tool.name,
134138
tool.description,

0 commit comments

Comments
 (0)