From an approved research plan to screened papers, page-level evidence, and an editable cited report.
A local-first research workspace for macOS.
Download · Workflow · Product tour · My role · Build from source
An editable report beside the exact source passage that supports it.
Spark Agent helps a researcher move from a question to a report without hiding the work inside a chat transcript. You approve the exact scope and plan; a durable, bounded research agent then executes one approved step at a time, preserves what happened, and stops when it needs a human decision.
The result is not a single generated answer. It is a reviewable chain of discovery records, screening decisions, local PDFs, exact-page quotations, an extraction matrix, analysis artifacts, and an editable report with citations.
Research work is usually split across search tabs, spreadsheets, PDF readers, notebooks, and writing tools. The researcher carries the context between them, while a chat history is a weak record of what was searched, excluded, quoted, or changed.
Spark makes the research artifacts the product record. The agent advances bounded work from real workspace state; the researcher remains responsible for scope, screening, interpretation, and conclusions.
Ask a bounded question
-> approve the exact providers, queries, budget, and stopping policy
-> review and screen the candidate papers
-> attach and read local PDFs
-> save verbatim evidence with the exact page
-> compare extracted findings
-> edit and export a cited report
Inside that approved envelope, the Research Agent can choose the next allowed query, observe novelty and evidence coverage, retry a known-safe failure, stop early, or ask for a revised plan. It cannot add a provider, widen a query, increase a budget, grant itself a permission, screen a paper, or make the final scientific conclusion.
Every meaningful step is preserved as structured state rather than left in a transcript:
understand -> plan -> act -> observe -> decide -> continue / ask / stop
|
v
verify and preserve
Approve the exact question, query set, provider set, result budget, and stopping policy. Spark runs bounded Crossref and OpenAlex discovery, normalizes and deduplicates candidates, and lets the researcher mark each one as Include, Exclude, or Awaiting before it can enter the evidence set.
Save a verbatim passage from a local PDF with its exact page. Carry that evidence through extraction and writing, then reopen the supporting passage directly from the report citation.
Approved plans, completed steps, artifacts, failures, and pending decisions are durable project state. A provider interruption or application restart does not turn the workflow into a guessing exercise.
I led the product definition and orchestration of Spark Agent: positioning, research-product analysis, the agent-native operating model, workflow scope, information architecture, interaction trade-offs, acceptance criteria, and coordination across multiple models.
Codex and other models produced a substantial share of the implementation under those constraints. I reviewed the resulting behavior against real workflows, redirected work that crossed the product boundary, and accepted features only when the UI, persistence, recovery, and evidence trail worked together.
This project represents my product judgment and ability to direct AI-assisted implementation. It is not a claim that I manually wrote every line of code.
- Exact plan and scope approval before remote literature discovery.
- Bounded Crossref and OpenAlex discovery with normalization, deduplication, coverage observations, durable operation identity, and restart recovery.
- Human Include, Exclude, or Awaiting screening decisions.
- Local PDF reading and verbatim evidence capture with exact-page reopening.
- A persistent extraction matrix and editable cited report with separate document and citation exports.
- Local CSV analysis through approved Python and Jupyter execution, with figures, tables, notebooks, logs, environment metadata, hashes, and artifact lineage.
- Project-scoped research memory and a reviewed, replay-tested lifecycle for project-local reusable procedures.
- The agent advances one approved research step at a time.
- It may reorder approved queries or stop from novelty and evidence coverage.
- It may not expand the provider, query, budget, method, permission, or disclosure boundary without a revised approval.
- Candidate metadata is not full-text evidence. A verified local PDF is required before a source can support a claim.
- Screening judgments, evidence interpretation, and scientific conclusions remain human decisions.
Download Spark Agent v0.2.0 from GitHub Releases. Spark Agent requires macOS 13 or newer and a running Docker Desktop or OrbStack installation. This build is distributed without Apple notarization; macOS may ask you to confirm that you trust the downloaded application.
Model credentials are optional. Crossref/OpenAlex discovery, local PDF evidence, deterministic CSV analysis, and approval-gated workflows work without a model key.
Requirements:
- macOS 13 or newer
- Node.js 20 and pnpm 9
- Docker Desktop or OrbStack
The packaged .dmg does not require Node.js or pnpm. It currently does require
Docker Desktop or OrbStack to be installed and running because Spark's bundled
local research services run in pinned containers. If the container engine is
not ready, Spark keeps the question editable in the current window and shows
the exact recovery action before any search is started.
git clone https://github.com/shawliu998/spark-agent.git
cd spark-agent
pnpm install
pnpm mvp:devThe first start builds two pinned local services and can take several minutes.
Later starts reuse the container cache. Stop the scoped local services with
Ctrl-C or:
pnpm science:downThe default deterministic literature and dataset workflows, PDF import, CSV analysis, and Jupyter execution work without a model key. Optional model-assisted planning, synthesis, and PaperQA use an OpenAI-compatible credential stored in macOS Keychain:
pnpm model-key:set
pnpm model-key:statusSelect the non-secret provider model before launch:
export SPARK_AGENT_LLM_MODEL='your-provider-model-id'
# Optional: export OPENAI_API_BASE='https://your-provider.example/v1'
# Optional: export SPARK_AGENT_EMBEDDING_MODEL='your-compatible-embedding-id'
pnpm mvp:devPublic and LAN endpoints must use HTTPS. Plain HTTP is accepted only for literal
localhost or a loopback IP address.
Spark Agent Desktop (Tauri + React)
|
+-- research workspaces
+-- @spark/research-sdk
+-- approval and evidence surfaces
|
+-- science-core (FastAPI + SQLite)
+-- projects, plans, jobs, approvals, reports, memory
+-- bounded discovery and Research Agent decisions
+-- evidence and artifact integrity
|
+-- Unix-domain socket
|
+-- science-runtime
+-- approved Python and Jupyter
+-- no network
Bundled OpenCode sidecar
+-- replaceable model runtime behind packages/sdk
+-- project-local Skills and MCP capabilities
The desktop UI never calls the model runtime directly. Product-owned domain contracts live behind the research SDK and Science Core; model providers, Skills, and MCP servers remain replaceable.
Quality
The Research Agent workflow is covered by:
- 3 focused backend acceptance tests;
- 138 focused frontend contract tests;
- 12 fixed v1.3 agent evaluation cases;
- the full desktop TypeScript check;
- targeted whitespace and repository checks; and
- a live persisted page-2 evidence capture.
Run the same focused evidence:
bash scripts/quality/validate-agent-interview-story.shRun the complete local quality gate with:
pnpm install --frozen-lockfile
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install \
'./services/science-core/vendor/paper-search-mcp/paper_search_mcp-0.1.4+spark.3-py3-none-any.whl'
python -m pip install -e './services/science-core[literature,dev]' \
-e './services/science-runtime[dev]'
pnpm qualityTrust boundaries
- Workspace-scoped file access and explicit approval for material scope and high-risk operations.
- Loopback-only authenticated local services and allowlisted CORS.
- No arbitrary direct shell mode in the product UI.
- Revision-checked, idempotent workflows with bounded recovery and fail-closed review.
- Python and Jupyter run in a no-network container with a read-only root filesystem, resource limits, and a Unix-domain socket.
- Trusted reads re-verify source containment and SHA-256 before accepting files or artifacts.
This is research software. Model answers and generated analyses still require review before publication or consequential use.
Brand, upstream, and attribution
The canonical application icon is
apps/desktop/src/assets/spark-app-icon.png;
the horizontal product wordmark is
apps/desktop/src/assets/spark-wordmark.png.
Platform icon derivatives are generated from the application-icon master.
The desktop shell reuses MIT-licensed portions of Open Science Desktop v0.1.9. Spark-specific research pages, domain contracts, science services, approval state, and isolated execution are maintained in this repository. Spark Agent is independent from Open Science Desktop and its maintainers.
See LICENSE and THIRD_PARTY_NOTICES.md.



