Bibliographic Intelligence for Thought Emergence
Let every idea have a source, and every judgment have an anchor.
🔥 BITE Community | 💬 WeChat / BITE WeChat Group
🔥 News: BITE's public evidence layer is published on HuggingFace dataset PaperBite-Assets, covering
L0-L3structured paper assets (Markdown analysis notes + figures + manifests). Incrementally sync withscripts/sync_assets_from_hf.py; if you work on AI-related research, it is a strong starting point for building your own evidence vault.🔥 News: The formal local analysis chain now uses the v06 four-section note format: MinerU parse/reuse, chunk anchor extraction, main analysis JSON, section writing, figure/table slot selection, and structural validation. The default keeps up to 6 core figures/tables to reduce redundant visual dumps and per-paper analysis cost. For downloaded queues with enough API quota, use
--jobs 50for high-throughput batch analysis.
What is BITE? BITE is a local-first workflow framework for structured paper analysis and research memory, purpose-built for knowledge-grounded research agents. It transforms paper analysis into structured notes and builds a persistent, reusable research memory.
Who is this for? Researchers building paper-grounded knowledge bases, agent-assisted literature workflows, or evidence-backed idea generation.
🧠 Knowledge first, not execution first. Many AI research tools focus on helping you run experiments or draft papers. BITE focuses on the upstream question: when an agent makes a research decision, does it have enough structured, searchable paper evidence in hand?
🧩 Turn structured paper analysis into reusable research memory. BITE organizes paper PDFs and paper lists into layered local assets: source literature, single-paper evidence units, domain knowledge surfaces, cross-domain evidence accumulation, and downstream idea or experiment records.
🪶 Local-first with low lock-in. The default workflow is local files only: PDFs, Markdown notes, JSONL indexes, and idea notes all live under
obsidian-vault/. Normal use does not require a server, database, or service deployment.
💡 BITE is a methodology and local knowledge workflow, not a closed platform. What matters is the layered research assets you keep accumulating.
BITE is not centered on idea generation in isolation. The core claim is that research directions should emerge from an accumulated, structured, and traceable evidence base, then be stress-tested before execution.
This diagram shows BITE's six-layer asset hierarchy: L0-L3 (knowledge building, powered by PaperBite), L4 (emergence), and L5 (validation).
The table below follows the diagram from bottom to top:
| Level | Output | Role |
|---|---|---|
L0 |
paper PDFs | preserve source literature |
L1 |
single-paper analysis | extract idea, design, and evidence |
L2 |
Domain Research Vault | support domain-level induction and deduction |
L3 |
Cross-Domain Research Vault | support transfer and idea emergence |
L4 |
Idea Vault | emergence layer |
L5 |
Experiment Vault | validation layer |
Give BITE a research direction, and it helps you build the knowledge base step by step:
collect candidate papers / import local PDFs
-> download when needed
-> integrated analysis chain
(MinerU parse/reuse -> structured analysis -> vault export)
-> optional index refresh
-> query / ideate / review / export
You can use it in four common modes:
| Mode | Purpose | Typical entry |
|---|---|---|
| Build | Collect candidates, download or import PDFs, run the integrated analysis chain, and refresh the index when needed | research-workflow |
| Query | Retrieve papers by topic, task, method, venue, year, title, or technique tags | papers-query-knowledge-base |
| Decision | Compare methods before choosing baselines, changing a design, or writing related work | papers-query-knowledge-base |
| Idea | Generate, focus, and stress-test research directions grounded in the local knowledge base | research-brainstorm-from-kb, idea-focus-coach, reviewer-stress-test |
git clone https://github.com/<your-username>/BITE.git
cd BITE
conda env create -f environment/environment.yml
conda activate biteCreate a repo-root .env when you need model keys, model names, or parser
overrides. Use environment/.env.example as a
reference.
MinerU is the PDF parsing component inside BITE's local analysis chain. You no
longer need a separate MinerU batch-preparation phase before analysis:
scripts/run_local_paper_analysis.py can call MinerU during analysis, or reuse
existing parse outputs when you already have them. Minimal verification:
mineru --help should run, or .env should set MINERU_CLI_PATH.
/research-workflow
I want to build a knowledge base for controllable motion generation from PDFs.
Please tell me the next step and the expected outputs.
To use BITE's pre-built structured paper assets, sync from HuggingFace by layer:
pip install huggingface_hub
# Text only: analysis notes + indexes (~43 MB)
python scripts/sync_assets_from_hf.py --mode text
# Assets only: figures and tables (~1.8 GB)
python scripts/sync_assets_from_hf.py --mode assets
# Everything (default)
python scripts/sync_assets_from_hf.py --mode all --dry-run # preview first
python scripts/sync_assets_from_hf.py # full sync
# Explicitly replace your local paper list with the public PaperBite list
python scripts/sync_assets_from_hf.py --mode paper-list --overwrite-paper-listPaperBite shards use vault-relative paths (analysis/, index/, and
assets/) and extract directly under obsidian-vault/. This makes them
suitable as a drop-in public evidence vault for BITE. paper_list.csv is synced
only when explicitly requested, so your local paper list is not overwritten by
default. The public assets do not include the full original PDF corpus; keep
downloading or importing paperPDFs/ locally when PDFs are needed.
For a single paper, start directly from the PDF. The runner performs MinerU
parse or cache reuse, chunk evidence extraction, main analysis JSON generation,
section writing, figure/table placement, vault export, and structural
validation. Pass --mineru-output or --mineru-output-root only when you
already have parse outputs to reuse.
python3 scripts/run_local_paper_analysis.py \
--pdf "obsidian-vault/paperPDFs/<Venue_Year>/<Paper>.pdf" \
--conf-year "<Venue_Year>" \
--export-vaultThe default note body has four sections: 概要, 核心方法与创新机理,
实验与关键发现, and 定位与知识库关联. --max-note-images 6 is the
default figure/table budget for retaining task definition, core pipeline, and
key result or ablation visuals without flooding the top-level note.
For batch analysis, the queue runner calls the same formal chain per row:
python3 scripts/run_paper_list_analysis.py \
--source obsidian-vault/paper_list.csv \
--state Downloaded \
--jobs 50 \
--export-vault \
--max-note-images 6If provider rate limits, MinerU I/O, or local memory become the bottleneck,
drop to --jobs 10 or --jobs 20 and rerun. The runner resumes completed
work by default.
Build a topic knowledge base from scratch
/research-workflow
I want to build a knowledge base for text-driven reactive motion generation.
Start by collecting candidate papers and tell me which skill to use at each stage.
Collect candidate papers from a GitHub paper list
/papers-collect-from-github-repo
Collect papers related to controllable human motion generation from this GitHub repository: <URL>
Keep only items related to diffusion, controllability, real-time generation, or long-form motion.
Write a candidate list suitable for the downstream download workflow.
Run the formal local analysis chain
Run the full chain directly from a PDF:
python3 scripts/run_local_paper_analysis.py \
--pdf "obsidian-vault/paperPDFs/<Venue_Year>/<Paper>.pdf" \
--conf-year "<Venue_Year>" \
--export-vault \
--reasoning-effort max \
--part-thinking disabled \
--writer-thinking disabledReuse existing MinerU output when available:
python3 scripts/run_local_paper_analysis.py \
--mineru-output "<mineru_output_dir>" \
--paper-pdf "obsidian-vault/paperPDFs/<Venue_Year>/<Paper>.pdf" \
--conf-year "<Venue_Year>" \
--export-vaultFor batch analysis, the queue runner calls the same formal chain per row:
python3 scripts/run_paper_list_analysis.py \
--source obsidian-vault/paper_list.csv \
--state Downloaded \
--jobs 50 \
--export-vault \
--max-note-images 6| Need | Skill |
|---|---|
| Decide the next pipeline step | research-workflow |
| Collect candidates from web pages | papers-collect-from-web |
| Collect candidates from GitHub paper lists | papers-collect-from-github-repo |
| Download PDFs from candidate rows | papers-download-from-list |
| Generate a structured single-paper analysis | scripts/run_local_paper_analysis.py |
| Rebuild the local index | papers-build-index |
| Query or compare papers from local notes | papers-query-knowledge-base |
| Generate grounded research ideas | research-brainstorm-from-kb |
| Focus an idea into an executable plan | idea-focus-coach |
| Run reviewer-style stress tests | reviewer-stress-test |
| Export share-ready Markdown | notes-export-share-version |
See .claude/skills/README.md for the full skill map.
BITE intentionally stays plain: folders, Markdown, JSONL, CSV, and
SKILL.md. The same research memory can therefore be shared by multiple agents:
- Claude Code / Cursor can read
.claude/skillsdirectly. - Codex CLI can use
scripts/setup_shared_skills.pyto generate local aliases. - Other agents can read
obsidian-vault/index/index.jsonlandobsidian-vault/analysis/directly.
Codex CLI compatibility
Claude Code / Cursor does not need this step. Codex CLI does.
python3 scripts/setup_shared_skills.py
python3 scripts/setup_shared_skills.py --checkObsidian setup
- Obsidian is optional but recommended as a visualization layer.
- Open
obsidian-vault/as an Obsidian vault if you want graph view, backlinks, and manual browsing. - Do not treat Obsidian pages as a separate source of truth.

