An evidence-first bond analysis agent for Chinese market data
Numbers are code-calculated.
Narratives are LLM-assisted.
Every output is provenance-tracked.
BondLens is a lightweight AI agent platform for Chinese bond market analysis. It unifies deterministic analytics, optional LLM narration, and a reviewer-facing Trust Layer — so every number can be audited, every answer can be replayed, and every limitation is stated honestly.
Non-investment advice. For learning, research, portfolio demonstration, and interview discussion only.
Project page: https://phoenix0531-sudo.github.io/BondLens/
Static demo packs (no API key): docs/demo_runs/
A core design principle of BondLens (shared with platforms such as FinRobot) is the strict separation between deterministic financial computation and LLM-based narration.
| Layer | What produces it | Can invent numbers? |
|---|---|---|
| Yield / volume / percentiles / rankings | Deterministic tools (bond_agent/tools.py) |
No |
| Taxonomy / maturity buckets / peer spread | Rule-based classifiers + pure Python stats | No |
| Data lineage (live / snapshot / static) | Data resolver | No |
| Evidence ledger claims | Built from tool outputs | No |
| Final narrative text | Deterministic report, or LLM only if guardrail + judge pass | Text only; numbers must match evidence |
In short: tools compute, models narrate, trust decides.
BondLens turns a natural-language bond question into an auditable analysis run:
- Resolve data (live AkShare → cached snapshot → local Excel)
- Plan intent (overview / search / ranking / outliers / monitor / composite / bond report)
- Run deterministic tools
- Build structured evidence (market, peer, monitor, quality, maturity coverage)
- Compose a report with risk notes and mandatory limitations
- Optionally polish with an LLM under numeric + language guardrails
- Score trust, export a Bond Evidence Pack, and store a replay summary
- Answer Snapshot: 3-sentence headline + key metrics; full body collapsed by default
- SSE stream + soft final render: tool-step progress, token preview, final summary card without forced full-page reload; share/full board still via
result_url - Bilingual UI (zh default): query/cookie language memory, explicit zh/en switch, bilingual provenance lines
- Bond type mix + maturity buckets: conservative name-rule taxonomy (no rating inference)
- Peer comparison: same type + maturity bucket spread vs peers
- Cross-section monitor board: high yield / low volume / yield outliers / missing maturity
- Maturity / residual board: live coverage, cashflow teaching duration/DV01, perpetual dual scenarios (first finite leg + theoretical consol)
- Trust score + stress view + audit folds: guardrail / judge / risk / ledger behind details
Data Ops live / snapshot / static sample + lineage + maturity enrichment
Agent Core Planner → Tools → Evidence → Report
Trust Layer Guardrail + Judge + Risk Profile + Trust Score + Replay + Evals
flowchart TD
A[User Question] --> B[Data Source Resolver]
B --> C[Planner multi-intent]
C --> D[Deterministic Tools]
D --> E[Structured Evidence]
E --> F[Report + Limitations]
F --> G{Optional LLM}
G -->|guardrail pass| H[Narrated answer]
G -->|fail or disabled| I[Deterministic answer]
H --> J[Judge + Trust + Pack + Replay]
I --> J
Inspired by deterministic compute, LLM narration research platforms, BondLens specializes the idea for Chinese bonds with claim-level evidence, answer judging, red-team evals, and reviewer-facing evidence packs — not a multi-role equity research desktop.
UI screenshots will be recaptured after the product surface is fully stabilized. Current images under
docs/screenshots/reflect an earlier workbench revision.
Agent Workbench |
Answer, Tool Trace, and Evidence |
Risk Profile and Answer Judge |
Replay Dashboard |
pip install -r requirements.txt
# preferred local demo bind
export FLASK_RUN_HOST=127.0.0.1
export PORT=8765
export BOND_DATA_MODE=static
export SECRET_KEY=local-dev
python app.py
# open http://127.0.0.1:8765/agent
# try: 当前样本收益率分布是什么样?Or open a pre-generated pack with no server:
export FLASK_RUN_HOST=127.0.0.1
export PORT=8765
export BOND_DATA_MODE=auto # live first, then snapshot, then static
python app.py
# force live: BOND_DATA_MODE=live
# watch data_source.runtime_mode, maturity board, and trust score when live degradesOptional LLM polish (never required):
export OPENAI_API_KEY=...
export OPENAI_BASE_URL=http://127.0.0.1:31876/v1 # example: local OpenAI-compatible gateway
export OPENAI_MODEL=grok-4.5
export OPENAI_API_STYLE=chat
# Keys stay in process env only. Do not commit secrets.| Tool | Inputs | Deterministic outputs |
|---|---|---|
search_bonds |
name / type / maturity / yield filters | match_count, records |
describe_market |
active frame | yield/volume summaries, segments, data quality |
rank_bonds |
by, top_n, order | ranked records |
detect_yield_outliers |
method, threshold | outlier_count, scores |
compare_bond_to_market |
bond / record | percentiles, peer comparison |
build_market_monitor |
top_n | high-yield / low-volume / outliers / missing maturity |
generate_bond_report |
tool outputs + plan | analysis, risk notes, limitations |
Numbers in the final answer must come from these tools (or be rejected by the guardrail).
Each answer includes trust_score (0–100) built from evidence quality, data freshness/degradation,
ledger coverage, guardrail outcome, judge outcome, and a forced non-advisory penalty.
Every run can export a portable Bond Evidence Pack (JSON + static HTML):
- question / intent / tools
- data lineage + maturity coverage
- trust score + adjustments
- guardrail + judge + risk profile
- evidence ledger + final answer
- mandatory limitations
python scripts/generate_demo_packs.pyRuntime packs: .tmp/evidence_packs/
Committed demos: docs/demo_runs/
Live/snapshot feeds have no native maturity field. BondLens enriches matched names from the local security master and exposes:
GET/POST /api/maturity/unmatched?format=csv&data_mode=static
GET/POST /api/maturity/unmatched?format=json&data_mode=static
The UI maturity board links to the same export for unmatched bond names.
当前样本收益率分布是什么样?
搜索23附息国债26并给出收益率分析
打开今日市场监控面板:高收益、低成交与异常
按收益率列出最高的前5只债券
有没有收益率异常的债券?
筛选国债收益率大于 2.5 的债券
POST /api/agent/query
Content-Type: application/json
{
"question": "搜索23附息国债26并给出收益率分析",
"data_mode": "auto"
}Streaming (SSE):
POST /api/agent/stream
Content-Type: application/json
{
"question": "当前债券市场样本概览如何?",
"data_mode": "static"
}Events include status (tool steps), token (partial text), and final (soft-render view + result_url).
Operational endpoints:
GET /healthz
GET /api/agent/schema
GET /replay
GET /packs/<pack_id>.html
GET /packs/<pack_id>.json
GET /api/maturity/unmatched
Deployment notes: docs/deployment.md
Primary: ChinaMoney / AkShare-style spot deal fetch (direct preferred)
Snapshot: .tmp/bond_spot_deal_snapshot.csv
Final fallback: data/testdata.xlsx
- Live fields used include bond name, clean price, yield, BP change, weighted yield, volume, and native residual maturity when present (
termToMaturity) - Residual maturity may still be incomplete → coverage board + trust penalty on weak coverage
- Cashflow duration / DV01 are teaching-level level-coupon estimates, not OAS / full call-tree valuation
- Perpetual-style residuals expose dual scenarios (first finite leg + theoretical perpetual), not a multi-century fake tenor
- No issuer ratings, financial statements, guarantees, or credit events
- Yield is a risk signal, not a trade instruction
Modes:
auto -> live first, cached snapshot second, local fallback third
live -> live source requested; fallback reason shown if it degrades
static -> local Excel only
- Data resolver chooses live / snapshot / static with honest lineage
- Planner classifies multi-intent and selects tools
- Tools run pure Python analytics over the active frame
- Evidence is structured and ledger-backed
- Report is composed from evidence with risks and limitations
- Optional LLM may narrate only after local evidence exists
- Guardrail + judge accept or reject model text
- Trust score + Evidence Pack + replay make the run reviewable without dumping raw JSON
If OPENAI_API_KEY is not set, the project still runs with deterministic fallback output.
Recorded against local new-api (http://127.0.0.1:31876/v1) with model
deepseek-v4-flash-search, BOND_DATA_MODE=static, temperature=0, one-shot
numeric repair when guardrail fails.
Stable first bond for report questions: 06国开24 (bond-name ascending, mergesort).
Full table: docs/demo_runs/llm_matrix_deepseek_v4.md · raw: llm_matrix_deepseek_v4.json
| Scenario | Lang | Threshold | Result | Notes |
|---|---|---|---|---|
| overview | zh | 3/3 final LLM | 3/3 | direct pass |
| bond report | zh | 3/3 final LLM | 3/3 | direct pass |
| overview | en | >=2/3 | 2/3 | 1 residual invented 5%; repair may recover |
| bond report | en | >=2/3 | 3/3 | residual duration%/percentile invents caught by guardrail + repair |
Historical grok-4.5 matrix (same thresholds, earlier channel):
docs/demo_runs/llm_matrix_grok45.md.
On this host, gpt-5.4* / grok-4.5 channels are often unavailable or timeout;
deepseek-v4-flash-search is the currently verified chat model.
Honest residuals:
- Provider channel churn still forces deterministic fallback when models vanish
- Guardrails stay on; unsupported numbers never become final
- English overviews may invent bare shares like
5%and fall back once in three - One-shot repair rewrites only after a failed numeric check; not a free pass
- End-to-end latency often 8–25s on deepseek; repair path can exceed 40s
- Soft-render shows a summary card; full dashboard tables remain on
result_url - Not implemented: WebSocket tick feed, true OAS / full call-tree perpetual pricing, desktop GUI/CLI
This matrix is evidence of a working path, not a zero-bug claim.
This project started as a 2024 undergraduate thesis: a Flask-based bond data analysis system. The original thesis version is preserved and should not be rewritten:
- Original thesis branch:
undergraduate-thesis-2024 - Current branch:
main
MIT
BondLens AI is an engineering and research demonstration. It does not provide investment advice, does not claim complete market coverage, and does not replace professional fixed-income research tools.



