version1 是在 version_mvp 基础上的一次完整功能扩展。MVP 版本主要证明了“读取 GitHub PR -> 解析 patch -> 构建上下文 -> 调用 LLM -> 输出 JSON/Markdown 报告”这条最小链路;version1 则把它推进成一个可以覆盖更多审查入口、可以在 GitHub Actions 中自动运行、并带有评测集与示例输出的工程化版本。
Compared with origin/codex/version_mvp, version1 adds 44 changed files and roughly 2.6k lines of implementation, tests, examples, and docs. The main change is that the agent is no longer PR-only: it now treats PR, commit, compare range, and local git diff as different sources of the same ChangeSet.
与 version_mvp 相比,version1 的重点变化如下:
| Area | MVP | version1 |
|---|---|---|
| Review target / 审查目标 | 只支持 GitHub PR URL | 支持 GitHub PR、GitHub commit、GitHub compare range、本地 git diff |
| Pipeline abstraction / 管线抽象 | PR 专用流程 | 新增 ReviewTargetInfo、ReviewTargetRef、ChangeSet,统一多来源变更 |
| CLI commands / 命令行 | fetch、review 面向 PR |
fetch <target>、review <target> 自动识别目标,新增 review-action、eval-dataset |
| Local mode / 本地模式 | 不支持 | 支持 local 审查未提交的工作区变更,包含完整 unified diff 拆分 |
| GitHub Actions / 自动化 | 不发布 GitHub 评论 | 新增 .github/workflows/ai-review.yml,支持 PR 自动摘要评论和 push commit 评论 |
| Summary comments / 摘要评论 | 只生成本地报告 | 新增可更新的 PR summary comment,使用 marker 避免重复刷屏 |
| Evaluation / 评测 | 无评测集 | 新增 50 条 JSONL case,覆盖 target parser、diff parser、filter、Action event、schema、issue detection 等 |
| Example outputs / 示例输出 | 单个 demo 输出 | 新增 PR、commit、compare、local diff 四类示例输出 |
| LLM robustness / 模型鲁棒性 | 直接解析 JSON | 支持 fenced JSON/前后解释文本解析、JSON repair、timeout 配置、token usage 合并 |
| Environment loading / 环境变量 | 默认当前目录 .env |
自动从项目根目录定位 .env,并支持 OPENAI_TIMEOUT_SECONDS |
| Test coverage / 测试覆盖 | MVP 单测 | 当前公开主分支可复测 120 passed,覆盖多目标、GitHub Actions、评论、评测集、本地 diff、验证工具等测试 |
Version1 的新增代码主要落在这些模块:
src/pr_agent/targets/: 统一解析 PR、commit、compare、local target。src/pr_agent/review/runner.py: 把 CLI 与 GitHub Actions 复用的 review pipeline 独立出来。src/pr_agent/diff/full_parser.py: 将完整git diff拆成文件级 patch,支持本地模式。src/pr_agent/github/actions.py: 将 GitHub Actions event 转换成 review target。src/pr_agent/github/comments.py: 生成 PR/commit summary comment。src/pr_agent/evaluation/dataset.py: 加载、统计和评分评测 JSONL。examples/与evaluation/: 提供可展示的审查结果与评测数据。
version2 turns the reviewer from a single broad pass into a coordinated review system:
- Multi-Agent Reviewer: Bug Reviewer, Test Reviewer, Security Reviewer, and Performance Reviewer run specialized passes. A deterministic Coordinator deduplicates and ranks findings.
- Patch Suggestion / Test Suggestion: Findings can include structured fix plans, optional patch snippets, validation commands, test scenarios, assertions, and optional test code.
- Evaluation Report:
pr-agent eval-reportscores 20-30 PR-level cases withvalid_finding_rate,line_hit_rate,false_positive_rate,fixability_rate, latency, token usage, and optional cost. - Design reports: See
docs/version2-design.mdand the agent harness note indocs/agent-harness-design.md.
v2.1 adds evidence-verified review. The model still proposes candidate findings, but findings can now pass through a controlled verification layer before publication:
flowchart TD
A["Multi-Agent Reviewer"] --> B["Coordinator"]
B --> C["Verification Planner"]
C --> D["Policy Gate"]
D --> E["Static Tools / Sandbox Tools"]
E --> F["Evidence Adjudicator"]
F --> G["Final Review Report"]
Why this matters: LLM-only review can produce plausible false positives. v2.1 records whether each finding is supported, contradicted, or inconclusive, which tools ran, what evidence they produced, and whether that evidence changed the confidence or publication decision.
New review options:
pr-agent review local `
--out outputs/local-review `
--verify static `
--workspace . `
--verification-budget 3 `
--verification-timeout 45Verification modes:
--verify off: default v2 behavior; no new tools run.--verify static: read-only repository tools only: repository search, file reads, test discovery, and dependency inspection.--verify sandbox: static tools plus allowlisted Docker checks such aspython -m pytest -q <approved_test_path>,ruff check <approved_paths>, andmypy <approved_paths>.
Standalone verification:
pr-agent verify outputs/local-review/review_result.json --workspace . --mode static --out outputs/verified-reviewv2.1 output adds:
verification_report.jsonartifacts/verification/<finding-id>/search_result.jsonartifacts/verification/<finding-id>/test_discovery.json- sandbox logs such as
pytest.logwhen sandbox mode is eligible - finding-level
verificationobjects insidereview_result.json
Security boundary:
- The LLM cannot provide raw shell commands.
- Every tool request is filtered through a deterministic allowlist policy.
- Sandbox execution copies the workspace into a cleaned temporary directory, drops
.git,.env, keys, virtualenvs,node_modules, and build outputs, then runs Docker with no network, no forwarded secrets, read-only mount, dropped capabilities, memory/CPU/pid limits, and a timeout. - GitHub Actions automatically downgrades fork PRs from sandbox verification to static verification.
Evaluation now includes verification metrics:
| Metric | v2 | v2.1 |
|---|---|---|
| valid_finding_rate | yes | yes |
| false_positive_rate | yes | yes, expected lower |
| line_hit_rate | yes | yes |
| fixability_rate | yes | yes |
| verification_coverage | no | yes |
| supported_finding_rate | no | yes |
| contradicted_suppression_rate | no | yes |
| average latency | yes | yes, expected higher |
| p95 latency | yes | yes |
| token cost | yes | yes |
Expanded live E2E manual run:
Command:
python -m pr_agent.main run-live-e2e --cases evaluation/live_e2e_cases.jsonl --out outputs/live-e2e-expanded --verify static --verification-budget 3 --verification-timeout 45Summary:
- Dataset: 20 cases, Python + C++.
- Main run: 19 completed, 1 provider disconnect (
LIVE017). - Retry:
LIVE017completed separately. - Verification mode:
static; most verification results are expected to stayinconclusive. - Detailed report:
outputs/live-e2e-expanded/manual_judgement_report.md.
| Case | Type | Expected | Manual result |
|---|---|---|---|
| LIVE001 | Python bug | None dereference before guard | Found; extra related test noise |
| LIVE002 | Python security | SQL injection via f-string | Found; duplicate security finding suppressed |
| LIVE003 | Python security | Authorization token logging | Found; extra test noise |
| LIVE004 | Python mixed | Unsafe YAML + N+1 lookup | Found both; avoided SQL false positive |
| LIVE005 | Python clean | Parameterized SQL should be clean | Clean result |
| LIVE006 | Python clean | Guarded refactor should avoid None false positive | Avoided target false positive, but noisy test/contract findings |
| LIVE007 | Python clean/tests | Existing tests should suppress broad test complaint | Clean result; evidence gate suppressed candidates |
| LIVE008 | Python performance | Cache disable is benchmark-dependent | Cautious performance warning; static evidence inconclusive |
| LIVE009 | Python blocker | SyntaxError/import blocker | Found |
| LIVE010 | Python conditional crash | Divide-by-zero when total == 0 |
Found; extra test noise |
| LIVE011 | Python logic/security | Authorization ownership check inverted | Found; duplicate bug/security/test noise |
| LIVE012 | Python concurrency | Lock removed, race/lost update | Found; static evidence inconclusive |
| LIVE013 | Python resource leak | File handle leak | Found; category noise across bug/security/test |
| LIVE014 | Python clean concurrency | Lock retained, should be clean | Clean result; low-value test candidate suppressed |
| LIVE015 | C++ blocker | Missing semicolon compile failure | Found |
| LIVE016 | C++ conditional crash | Null pointer dereference before guard | Found; extra test noise |
| LIVE017 | C++ lifetime | Dangling c_str() pointer |
Partial; provider retry succeeded, root cause mentioned but public finding was only a test finding |
| LIVE018 | C++ memory | Early-return memory leak | Partial; root cause recognized, but public finding was only a test finding |
| LIVE019 | C++ concurrency | Mutex removed, data race | Found; static evidence inconclusive |
| LIVE020 | C++ clean RAII | RAII ownership should avoid memory false positive | Avoided memory/pointer false positive, but published test noise |
Manual takeaway: the agent catches obvious crash, build, security, logic, and lock-removal issues well in both Python and C++. The main weakness is publication quality: duplicate root causes, test-review noise, and C++ memory/lifetime bugs sometimes being surfaced only as missing-test findings.
See docs/version2_1_design.md, docs/sandbox-security.md, and docs/verification-evaluation.md.
AI PR Review Agent is a local-first code review agent for GitHub Pull Requests, GitHub commits, GitHub compare ranges, and local git diffs. It focuses on repository-aware context, structured LLM output, conservative review, result validation, reproducible reports, and evaluation.
AI PR Review Agent 是一个本地优先的 AI 代码审查 Agent,支持 GitHub PR、GitHub commit、GitHub compare range 和本地 git diff。它不是简单把 diff 丢给大模型,而是围绕 diff 解析、上下文构建、结构化输出、保守审查策略、结果校验、报告生成和评测数据做成的工程化 Agent。
At the system level, it can also be viewed as a read-only agent harness for code review: the harness controls the review target, context, reviewer orchestration, validation, tracing, reporting, and evaluation around the LLM reviewer.
从系统层看,它也可以理解为一个面向代码审查的只读 Agent Harness:由 harness 管理审查目标、上下文、reviewer 编排、结果校验、trace、报告和评测,而不是让模型直接自由操作仓库。
Given one review target, the same command can automatically detect the target type:
输入一个 review target 后,同一个命令会自动识别目标类型:
https://github.com/owner/repo/pull/123for GitHub PR review.https://github.com/owner/repo/commit/<sha>for single-commit review.https://github.com/owner/repo/compare/<base>...<head>for compare/range review.localfor uncommitted local git diff review againstHEAD.
Core workflow:
核心流程:
- Fetch PR, commit, or compare metadata from the GitHub REST API, or read local
git diff. - Parse patch text into structured
DiffHunkandDiffLineobjects with old/new line numbers. - Normalize every source into one shared
ChangeSetmodel. - Filter generated files, lock files, binary files, large patches, and removed files.
- Retrieve lightweight repository context, including surrounding code, README excerpts, and related test file candidates.
- Call an OpenAI-compatible LLM to generate conservative code review findings.
- Repair or validate JSON model output, then validate findings with Pydantic schemas, confidence thresholds, file-path checks, and line-number checks.
- Generate
review_result.json,review_report.md, andtrace.jsonl. - Optionally publish a GitHub summary comment in GitHub Actions.
Most toy AI code review demos stop at "send the diff to an LLM and ask for comments." That is not enough for real engineering use. Code review quality depends on whether the system understands changed lines, nearby code, file type, repository hints, and whether model output can be validated and traced.
很多 AI Code Review 玩具项目只是“把 diff 发给大模型,让它吐几条建议”。这在工程上远远不够。真实代码审查的难点不只是调用模型,而是:
- How to parse PR/commit/range/local diffs into reliable changed-line structures.
- How to provide enough repository context without flooding the prompt.
- How to make model output machine-checkable.
- How to reduce false positives with conservative prompts and validators.
- How to preserve trace, latency, token usage, and reproducible reports.
- How to run the same review pipeline locally, from a PR URL, and inside CI.
- How to evaluate parser, schema, target detection, and issue-detection behavior over labeled cases.
- One review command / 一个 review 指令:
pr-agent review <target>supports PR, commit, compare, and local diff targets. - Target auto-detection / 自动识别目标: Parses GitHub PR URLs, commit URLs, compare URLs, and
local. - Shared ChangeSet abstraction / 统一 ChangeSet 抽象: Normalizes target metadata, changed files, and parsed hunks before review.
- Unified diff parser / Diff 解析: Converts patch text into hunks and lines with old/new line numbers.
- Full local diff parser / 完整本地 diff 解析: Splits full
git diffoutput into file-level patches for local working tree review. - Repository context retrieval / 仓库上下文检索: Adds surrounding code, README snippets, and related test file candidates from GitHub or local files.
- Multi-agent reviewer / 多 Agent 审查: Runs Bug, Test, Security, and Performance reviewer passes, then coordinates duplicate findings.
- Structured findings / 结构化建议: Uses
ReviewFindingandReviewResultschemas for reliable downstream processing. - Patch and test suggestions / 修复与测试建议: Emits structured fix plans, optional patch snippets, validation commands, test scenarios, assertions, and optional test code.
- JSON recovery / JSON 输出恢复: Handles plain JSON, fenced JSON, JSON embedded in text, and one-shot JSON repair.
- Quality gates / 质量过滤: Filters low-confidence findings, invalid file paths, weak evidence, hallucinated line numbers, deterministic false positives, LLM-verifier rejections, and low-value style nits.
- Markdown report / Markdown 报告: Renders readable review reports for demos and portfolio presentation.
- GitHub Actions review / GitHub Actions 自动审查: Resolves
pull_requestandpushevents, runs review, uploads artifacts, and publishes summary comments. - Evaluation dataset / 评测数据集: Provides 50 labeled JSONL cases, 25 PR-level cases,
eval-dataset, andeval-reportcommands for coverage and scoring.
Review Target
- GitHub PR URL
- GitHub commit URL
- GitHub compare URL
- local git diff
↓
Target Parser
- detect pull_request / commit / compare / local_diff
↓
ChangeSet Loader
- fetch GitHub metadata and changed files
- or read local git diff
- normalize target + files + hunks_by_file
↓
Diff Parser
- parse unified diff hunk headers
- split full local git diff into file patches
- track old/new line numbers
- produce DiffHunk and DiffLine
↓
File Filter
- skip lock/generated/binary/large/removed files
↓
Context Retriever
- target patch
- surrounding code around changed lines
- README excerpt
- related test file candidates
↓
General Reviewer
- OpenAI-compatible LLM call
- conservative review prompt
- JSON-only output requirement
- JSON repair fallback
↓
Validator
- Pydantic schema validation
- confidence threshold
- file path and line number checks
- evidence and suggestion checks
↓
Renderer / Publisher
- review_result.json
- review_report.md
- trace.jsonl
- optional GitHub summary comment
这个设计把“变更来源”和“审查流程”解耦。PR、commit、compare、local diff 都先变成统一的 ChangeSet,后面的 context retrieval、LLM review、validator、renderer 和 GitHub Actions 发布流程都可以复用同一套逻辑。
For the agent harness boundary, component mapping, and data flow, see docs/agent-harness-design.md.
Create and install a virtual environment:
创建并安装虚拟环境:
py -3.12 -m venv .venv-win
.\.venv-win\Scripts\python.exe -m pip install -e ".[dev]"
copy .env.example .envThis README uses .venv-win for Windows PowerShell commands to avoid confusion with MSYS/Mingw-created virtual environments, which expose executables under bin/ instead of the standard Windows Scripts/ directory. If you create a normal Windows CPython environment named .venv, replace .venv-win with .venv in the commands.
本 README 在 Windows PowerShell 命令中使用 .venv-win,是为了避免和 MSYS/Mingw Python 创建的 .venv 混淆;后者通常使用 bin/,而不是标准 Windows 的 Scripts/。如果你用标准 Windows CPython 创建名为 .venv 的环境,只需要把命令里的 .venv-win 换成 .venv。
Fill .env:
填写 .env:
GITHUB_TOKEN=your_github_token
OPENAI_API_KEY=your_openai_or_compatible_api_key
OPENAI_BASE_URL=https://api.openai.com/v1
OPENAI_MODEL=gpt-4.1-mini
OPENAI_TIMEOUT_SECONDS=120
VERIFIER_OPENAI_API_KEY=your_verifier_openai_or_compatible_api_key
VERIFIER_OPENAI_BASE_URL=https://api.openai.com/v1
VERIFIER_OPENAI_MODEL=gpt-4.1-mini
VERIFIER_OPENAI_TIMEOUT_SECONDS=60Fetch metadata and structured diff only:
只获取元数据和结构化 diff:
# GitHub PR
.\.venv-win\Scripts\pr-agent.exe fetch https://github.com/owner/repo/pull/123 --out outputs/pr-fetch
# GitHub commit
.\.venv-win\Scripts\pr-agent.exe fetch https://github.com/owner/repo/commit/<sha> --out outputs/commit-fetch
# GitHub compare/range
.\.venv-win\Scripts\pr-agent.exe fetch https://github.com/owner/repo/compare/main...feature --out outputs/compare-fetch
# Local uncommitted git diff against HEAD
.\.venv-win\Scripts\pr-agent.exe fetch local --out outputs/local-fetchRun full review with the same command shape:
使用同一个 review <target> 指令运行完整审查:
# GitHub PR
.\.venv-win\Scripts\pr-agent.exe review https://github.com/owner/repo/pull/123 --out outputs/pr-review
# GitHub commit
.\.venv-win\Scripts\pr-agent.exe review https://github.com/owner/repo/commit/<sha> --out outputs/commit-review
# GitHub compare/range
.\.venv-win\Scripts\pr-agent.exe review https://github.com/owner/repo/compare/main...feature --out outputs/compare-review
# Local uncommitted git diff against HEAD
.\.venv-win\Scripts\pr-agent.exe review local --out outputs/local-reviewRun GitHub Actions mode locally:
本地模拟 GitHub Actions 模式:
.\.venv-win\Scripts\pr-agent.exe review-action --event-path path/to/event.json --event-name pull_request --dry-runRun evaluation dataset validation:
运行评测数据集校验:
.\.venv-win\Scripts\pr-agent.exe eval-dataset --dataset evaluation/cases.jsonl --out outputs/eval_report.jsonRun PR-level evaluation report:
.\.venv-win\Scripts\pr-agent.exe eval-run --cases evaluation/runnable_pr_cases.jsonl --out examples/evaluation/run --llm-mode deterministic
.\.venv-win\Scripts\pr-agent.exe eval-report --cases evaluation/pr_cases.jsonl --predictions evaluation/pr_predictions.example.jsonl --out outputs/pr_evaluation_report.jsonRun tests:
运行测试:
.\.venv-win\Scripts\python.exe -m pytestCurrent test status:
当前测试状态:
120 passed
本项目可以作为一个独立的 AI Review 工具,被其他 GitHub 仓库通过 GitHub Actions 自动调用。其他仓库不需要复制 src/ 代码,只需要复制一个 workflow yml。
在本仓库中找到示例文件:
examples/external_repo_github_action/ai-review.yml
把它复制到目标仓库的这个位置:
.github/workflows/ai-review.yml
复制完成后,目标仓库结构应类似:
target-repo/
.github/
workflows/
ai-review.yml
当前示例会在目标仓库发生以下事件时自动运行:
on:
push:
branches:
- "**"
pull_request:
types:
- opened
- synchronize
- reopened
- ready_for_review含义:
push: 任意分支 push 后触发,review 本次 push 的before...after变更范围,并把 summary comment 发到本次 push 的 head commit。pull_request.opened: 新建 PR 后触发,review 这个 PR。pull_request.synchronize: PR 分支新增 commit 后触发,重新 review 这个 PR。pull_request.reopened: PR 重新打开后触发。pull_request.ready_for_review: draft PR 变成 ready 后触发。
示例 yml 会从本工具仓库安装 AI Review Agent:
run: python -m pip install "git+https://github.com/Atirian-Chen/ai-pr-review-agent.git@version2.1"因此,示例 yml 的可用前提是:
Atirian-Chen/ai-pr-review-agent对目标仓库的 GitHub Actions runner 可访问。version2.1branch or tag 已经推送到 GitHub。
如果你后续改用新的 release/tag,例如 v2.1.0,把安装行改成:
run: python -m pip install "git+https://github.com/Atirian-Chen/ai-pr-review-agent.git@v2.1.0"如果本工具仓库是 private,目标仓库还需要额外配置读取本工具仓库的 token;public 仓库不需要。
每个使用该工具的目标仓库都必须配置:
OPENAI_API_KEY
VERIFIER_OPENAI_API_KEY
配置位置:
Target repository -> Settings -> Secrets and variables -> Actions -> New repository secret
Secret 名称填:
OPENAI_API_KEY
Secret 值填模型服务商提供的 API key。
如果启用 LLM verifier,再新增一个 secret:
VERIFIER_OPENAI_API_KEY
这个 key 专门给第二遍验证模型使用。没有配置时,工具仍会运行确定性 verifier,但会在结果统计里把 LLM verifier 标记为 skipped。
GITHUB_TOKEN 不需要手动创建,GitHub Actions 会自动提供:
GITHUB_TOKEN: ${{ github.token }}示例 yml 中模型相关参数是明文写在 workflow 里的:
OPENAI_BASE_URL: https://api.deepseek.com
OPENAI_MODEL: deepseek-v4-pro
OPENAI_TIMEOUT_SECONDS: "500"
VERIFIER_OPENAI_BASE_URL: https://api.deepseek.com
VERIFIER_OPENAI_MODEL: deepseek-v4-flash
VERIFIER_OPENAI_TIMEOUT_SECONDS: "120"含义:
OPENAI_BASE_URL: OpenAI-compatible API 地址。OPENAI_MODEL: 主 reviewer 使用的模型名,建议使用更强的推理模型。OPENAI_TIMEOUT_SECONDS: 主 reviewer 单次 LLM 请求超时时间,单位是秒。VERIFIER_OPENAI_BASE_URL: LLM verifier 使用的 OpenAI-compatible API 地址。VERIFIER_OPENAI_MODEL: 第二遍验证模型,建议使用 mini、flash、chat 等更便宜更快的模型。VERIFIER_OPENAI_TIMEOUT_SECONDS: LLM verifier 单次请求超时时间,单位是秒。
如果使用官方 OpenAI API,可以改成:
OPENAI_BASE_URL: https://api.openai.com/v1
OPENAI_MODEL: gpt-4.1-mini
OPENAI_TIMEOUT_SECONDS: "120"
VERIFIER_OPENAI_BASE_URL: https://api.openai.com/v1
VERIFIER_OPENAI_MODEL: gpt-4.1-mini
VERIFIER_OPENAI_TIMEOUT_SECONDS: "60"如果使用 DeepSeek 或其他兼容服务,保持服务商要求的 URL 和模型名即可。
示例 yml 需要这些权限:
permissions:
contents: write
issues: write
pull-requests: write含义:
contents: write: 允许在 push 场景给 head commit 写 commit comment。issues: write: 允许在 PR conversation 中创建或更新 summary comment。pull-requests: write: 允许读取 PR 元数据,并在 PR 场景下发布或更新 summary comment。
配置完成后:
- 目标仓库 push 后,会自动 review 本次 push 的变更,并评论到最后一次 commit。
- 目标仓库 PR 创建或更新后,会自动 review PR,并在 PR conversation 中创建或更新一条 AI Review Summary。
- 每次运行都会上传
outputs/github-action作为 GitHub Actions artifact,里面包含review_result.json、review_report.md、trace.jsonl和summary_comment.md。
Troubleshooting:
故障排查:
If dependency installation fails on Windows MSYS/Git Bash with a pydantic-core build error, create the virtual environment with a standard CPython 3.11+ interpreter, for example the Python Launcher (py -3.11 -m venv .venv-win) or the installer from python.org.
If review fails with ReadTimeout, LLM API request timed out, or a provider-specific timeout, the LLM provider did not finish sending the response before the timeout. Increase OPENAI_TIMEOUT_SECONDS / VERIFIER_OPENAI_TIMEOUT_SECONDS in .env, or the matching timeout_seconds value in configs/default.yml, then retry.
如果在 Windows MSYS/Git Bash 环境中安装依赖时遇到 pydantic-core 构建错误,建议改用标准 CPython 3.11+ 创建虚拟环境,例如 Python Launcher(py -3.11 -m venv .venv-win)或 python.org 安装版。
Example targets:
示例目标:
PR: https://github.com/octocat/Hello-World/pull/1
Commit: https://github.com/octocat/Hello-World/commit/7044a8a032e85b6ab611033b2ac8af7ce85805b2
Compare: https://github.com/octocat/Hello-World/compare/553c2077f0edc3d5dc5d17262f6aa498e69d6f8e...7044a8a032e85b6ab611033b2ac8af7ce85805b2
Local: local
Generated example files:
生成的示例文件:
- PR example:
examples/octocat_hello_world/review_result.json,examples/octocat_hello_world/review_report.md - Commit example:
examples/octocat_commit/review_result.json,examples/octocat_commit/review_report.md - Compare example:
examples/octocat_compare/review_result.json,examples/octocat_compare/review_report.md - Local diff example:
examples/local_git_diff/review_result.json,examples/local_git_diff/review_report.md
Example finding:
示例 finding:
### 1. [Minor][maintainability] Concatenated command and description without spacing
- File: `README:2`
- Confidence: 0.95
- Evidence: +$ mkdir ~/Hello-WorldCreates a directory for your project called "Hello-World" in your user directory
- Why it matters: The added lines incorrectly combine the shell command and its explanation into a single string, making the README confusing and unreadable.
- Suggestion: Separate the command and its description, e.g., using a code block for the command and plain text for the explanation.Example GitHub summary comment:
示例 GitHub 摘要评论:
<!-- ai-pr-review-agent:summary-comment -->
## AI Review Summary
- Target: `pull_request` `#123`
- Risk: Low
- Files reviewed: 3 / 5
- Findings: 1
- Model: gpt-4.1-mini
### Findings
1. **Minor / maintainability** `README.md:2`: Concatenated command and description without spacing (0.95)Example metrics:
示例指标:
Findings: 1
Critical: 0
Major: 0
Minor: 1
Nit: 0
Model: deepseek-v4-pro
Latency seconds: 23.73
Estimated tokens: 1563
- One command, multiple sources / 一个命令,多种来源: Users keep using
review <target>while the target parser handles PR, commit, compare, and local diff. - ChangeSet abstraction / ChangeSet 抽象: Different source types are normalized before review, avoiding duplicated reviewer logic.
- PR remains the main story / PR 仍是主线: PR is still the best fit for GitHub code review workflows, while commit/compare/local support pre-PR and incremental scanning.
- CI integration without coupling / CI 集成但不绑死 CI: GitHub Actions mode reuses the same runner as local CLI, so CI is an entry point rather than a separate review implementation.
- Controlled tool execution / 受控工具执行: By default the agent only reads metadata, diffs, and file content. v2.1 can run static verification tools, and can run minimal allowlisted checks in a Docker sandbox when explicitly enabled.
- Structured output first / 优先结构化输出: Findings must match Pydantic schemas, so they can be filtered, sorted, evaluated, and rendered.
- Conservative review strategy / 保守审查策略: The prompt asks the LLM to report only issues supported by diff or context.
- Validation after generation / 生成后校验: The validator filters weak findings by confidence, file path, line number, evidence, and suggestion quality.
- Repair before failure / 先修复再失败: LLM output may contain markdown fences or extra text, so version1 attempts robust JSON extraction and repair before failing the run.
- Lightweight context before vector RAG / 先做轻量上下文: Version1 uses diff, surrounding code, README excerpts, and related test candidates before introducing embeddings.
- Evaluation as a first-class artifact / 评测作为一等产物: The project includes labeled cases so future iterations can measure parser, schema, and issue-detection changes.
Completed in version1:
version1 已完成:
- Multi-target review / 多目标审查: GitHub PR, commit, compare, and local diff now share one review pipeline.
- GitHub workflow / GitHub 工作流: Added GitHub Action triggers for
pull_requestandpush. - PR summary comment / PR 摘要评论: Added bot-generated summary comments with update markers.
- Evaluation dataset / 评测集: Added 50 labeled cases plus dataset validation and scoring helpers.
Next steps:
后续计划:
- Specialized reviewers / 多维度 reviewer: Split the current GeneralReviewer into Bug, Security, Performance, Test, and Maintainability reviewers.
- Aggregator / 聚合器: Deduplicate findings, sort by severity and confidence, and enforce per-category limits.
- Repository config / 仓库级配置: Support repository-level
ai-review.ymlfor filters, thresholds, model, and reviewer switches. - Richer local mode / 更完整的本地模式: Add staged-only, unstaged-only, and explicit
main..featurelocal range support. - Metrics / 指标统计: Track JSON valid rate, false positive rate, valid suggestion rate, line accuracy, latency, and cost.
- Repository-level RAG / 仓库级 RAG: Add keyword retrieval, AST/symbol extraction, and optional vector retrieval.
- Review history / 审查历史: Save repeated runs and compare finding changes over time.
This project is an open-source-style abstraction of my AICR-related internship experience. During the internship, I worked with AI-assisted code analysis, incremental code scanning, and AI-generated code comments. This project turns that industrial experience into a runnable and demonstrable LLM engineering system.
这个项目是我对淘天 AICR 相关实习经历的开源化抽象。在实习中,我接触过 AI 代码分析、存量/增量代码扫描和 AI 注释生成。这个项目把这些经历进一步沉淀成一个可运行、可测试、可展示的 LLM 应用工程项目。
The main story is:
核心叙事是:
AI code review is not just an LLM prompt.
It is an engineering pipeline around diff parsing, repository context,
structured model output, conservative review policy, validation, reporting,
CI integration, and eventually evaluation.
AI 代码审查不是简单写一个 prompt。
它是围绕 diff 解析、仓库上下文、结构化模型输出、保守审查策略、
结果校验、报告生成、CI 集成以及后续评测构建的一套工程化流程。