Databend welcomes AI-assisted development.
We already ship agent-oriented product features and maintain repository guidance for coding agents (AGENTS.md, agents/). This policy is about how humans and AI collaborate on contributions so maintainers can keep the warehouse correct, reviewable, and safe.
This policy covers contributions to this repository: code, tests, docs, issues, PR descriptions, and review discussion.
- AI is a normal engineering tool. Using models, agents, or IDE assistants is allowed and encouraged when it improves quality or speed.
- A human remains accountable. Whoever opens the PR owns the change: correctness, security, compatibility, tests, and follow-up fixes.
- Understanding is mandatory. If you cannot explain the change in your own words, it is not ready to submit.
- Reviewer time is scarce. Low-effort AI output that shifts the real work onto maintainers will usually be closed.
- Database semantics need evidence. Planner, executor, storage, meta, and SQL behavior changes need tests that can falsify the implementation — not vibes.
You may use AI to:
- Explore the codebase and form a root-cause hypothesis
- Draft implementations, refactors, tests, docs, and benchmarks
- Generate alternative designs for a human to compare
- Help with formatting, boilerplate, and mechanical edits
- Assist local review before you open a PR
When working with coding agents in this repository, follow AGENTS.md and the linked docs under agents/. Prefer the smallest relevant build/test loop, keep diffs reviewable, and treat tests as part of the change.
The following will usually result in the issue/PR being closed, possibly without extended discussion:
-
Ownerless autonomous contribution
- Agents may open PRs, update branches, and use
ghtooling — seeagents/commit-and-pr.md. What is banned is a PR with no accountable human: every PR must name an author-side person who has read the diff, can explain the changes, will answer review questions, and owns follow-up fixes. This person is not the reviewer — reviewers are assigned separately via CODEOWNERS - Bulk drive-by PRs that show no local validation or understanding of Databend module boundaries
- Agents may open PRs, update branches, and use
-
AI-generated conversation as a substitute for thinking
- Pasting model replies as your response to maintainer questions
- Long, generic, high-confidence commentary that does not address the concrete code or failure mode
- Using AI to argue for a change you cannot defend yourself
- Exception: using AI to translate or polish your own reasoning (for example, writing in Chinese and posting an AI-assisted English version) is fine and encouraged — the ideas must be yours
-
Slop submissions
- PRs that do not compile, fail basic lint, or clearly were not read by the author
- Fake or tautological tests that only mirror the implementation
- Huge unrelated diffs, noisy renames, or “cleanup” mixed into a behavior change without need
- Secret redaction failures, license-incompatible pasted code, or unexplained dependency churn
-
Bypassing collaboration norms
- Skipping issue/RFC discussion for large semantic or architectural changes
- Ignoring CLA, PR template, CODEOWNERS review paths, or requested test evidence
Agent-opened PRs are welcome. Every PR — however it was produced — must have a responsible human (the PR author, or the person named in the PR body for bot-authored PRs) who has:
- Read every line of the final diff. Unread code does not enter review. If the diff is too large to read, it is too large to submit — split it
and who can:
- State the user-visible or developer-visible problem in plain language
- Explain why this approach is correct for Databend’s architecture
- Point to the tests or manual validation that would fail if the fix were wrong
- Answer review questions without outsourcing the reply to a model transcript
- Own production risk: compatibility, upgrade/migration, performance, and data safety
If no responsible human engages when maintainers ask, the PR may be closed regardless of code quality.
AI assistance does not lower the quality bar. It raises the author’s duty to filter bad output before review.
Generated code fails in patterns that hand-written code rarely does. Before requesting review, the responsible human should walk the diff specifically looking for:
- Hallucinated interfaces: calls to functions, settings, or SQL behaviors that do not exist in this codebase or behave differently than the model assumed
- Plausible-but-wrong edge cases: NULL handling, overflow, empty input, timezone/precision, non-UTF-8 — verify against actual Databend behavior, not the model’s memory of “how databases work”
- Tests that mirror the implementation: a test that re-derives expected output from the same logic proves nothing; expected values should come from an independent source (MySQL/other engines, a spec, or hand computation)
- Silently swallowed errors:
unwrap_or_default, broadmatch _ =>, or droppedResults inserted to make code compile - Unnecessary abstraction and dead code: traits, generics, helpers, and config knobs that no second caller needs; delete them before review
- Comments and names that don’t match behavior: generated comments often describe intent, not what the code does
- Divergence from neighboring code: if the surrounding module solves the same problem differently, follow it or explain why not
Fixing these before review is the author’s job. Finding them in review is a signal the diff was not read.
Every PR must complete the ## AI assistance section. This declaration is a review-entry requirement:
- If AI materially helped with code, tests, docs, or the PR summary, describe the type of assistance and the affected scope
- If no AI was involved beyond ordinary autocomplete, spell check, or linter suggestions, write
None - Describe the work, not the vendor: do not include product, model, or provider names
A complete, valid example:
## AI assistance
- AI usage: An AI coding agent drafted the storage iterator patch; I reworked the error paths and added logic tests
- Responsible human: @your-actual-github-id
- [x] The responsible human has read every line of this diff and can explain each changeFor a PR without material AI assistance:
## AI assistance
- AI usage: None
- Responsible human: @your-actual-github-id
- [x] The responsible human has read every line of this diff and can explain each changeThe declaration identifies how the change was produced; it does not reduce the responsible human’s obligations. Maintainers may also request a walkthrough for high-risk changes (catalog/meta, storage correctness, transactions, planner semantics, security boundaries).
AI-assisted changes must still satisfy normal Databend contribution expectations:
- Prefer small, reviewable PRs over agent-produced megadiffs
- Split mechanical regeneration from semantic changes
- Do not mix drive-by refactors into bug fixes
- Bug fixes: include a regression test whenever practical
- Planner / executor / storage behavior: add or update logic tests (or the relevant suite) with expected output when deterministic
- Performance claims: include before/after measurements, not anecdotes
- “No Test — Explain why” in the PR template requires a real reason, not convenience
- Run the smallest relevant checks first, then stronger validation before handoff
- For Rust changes intended to land, do not submit known clippy/format failures
- If full workspace validation is expensive, say what you ran and what remains uncovered
- Write PR summaries for humans: motivation, behavior change, risks, validation
- Quote model output only when needed, inside blockquotes, with your own interpretation
- Keep discussion concrete and tied to files, tests, and failure modes
Review the change, not the tool.
- At least one human must have read the diff before merge. Approvals from AI review bots do not count toward this; they are assistants, not reviewers
- Spend human attention on correctness, data safety, compatibility, performance, and test adequacy
- Prefer questions that check author understanding over style nits already handled by CI. For high-risk areas, one probing question (“why is this branch safe when the snapshot is stale?”) reveals more than ten nits
- Read the tests first: verify they can fail, and that expected outputs come from an independent source rather than the implementation itself
- If a PR looks AI-generated and low-effort, ask for a tighter summary, tests, or a reduced diff; close it if the author cannot engage
- Domain CODEOWNERS remain the authority for their areas
- It is fine to use AI to help review, but merge decisions stay with humans
- You are responsible for ensuring contributed material is license-compatible, whether typed or generated
- Do not paste proprietary code, private customer data, credentials, or internal secrets into prompts in a way that causes them to land in the repository
- Treat security-sensitive areas (auth, authorization, multi-tenant isolation, sandbox UDF boundaries) as high scrutiny regardless of how the code was produced
| Doc | Role |
|---|---|
AI_POLICY.md (this file) |
Community rules for AI-assisted contribution and review |
AGENTS.md + agents/ |
How coding agents should work inside this repository |
.github/PULL_REQUEST_TEMPLATE.md |
PR checklist (CLA, summary, tests, change type, AI declaration) |
.github/CODEOWNERS |
Domain reviewers for merge paths |
If maintainer guidance for a specific PR conflicts with this document, follow the maintainer’s explicit guidance for that PR, then propose an update here if the exception should become policy.
- Maintainers may request changes, ask for human clarification, require more tests, or close PRs that violate this policy
- Repeated low-effort or unattended agent spam may result in blocked contributions
- Good-faith AI-assisted work that is well explained, tested, and reviewable is welcome
Use AI freely.
Read every line before you submit.
Submit only what you understand, can test, and can defend.
Do not waste reviewer time with unattended slop.