Skip to content

Latest commit

 

History

History
358 lines (278 loc) · 12.4 KB

File metadata and controls

358 lines (278 loc) · 12.4 KB

Demo script

A 7-step runbook that takes a reviewer from git clone to an end-to-end demonstration of every Sentinel capability — citation-grounded RAG, refusal, schema-constrained extraction, the human-in-the-loop review queue, and the dashboard — in roughly fifteen minutes on a developer laptop. An optional final section repeats the demo on AWS using the M10 Terraform stack.

All sample data is synthetic. The data/sample/ corpus is generated by scripts/gen_synthetic_corpus.py with a fixed seed (deterministic and reproducible); see data/sample/README.md.

Prerequisites

Tool Version Why
Docker + Docker Compose recent Postgres 16 + pgvector
uv 0.4+ Python toolchain (installs the 3.12 venv on first run)
Node 20 LTS Vite dev server for the frontend
Anthropic API key claude-sonnet-4-6 access for /query and /extract
OpenAI API key text-embedding-3-small access for embeddings at ingest time
or Google AI Studio key gemini-3.5-flash + gemini-embedding-2 drives both LLM and embeddings on one free key

Without API keys you can still run the test suite (it uses the deterministic fake LLM and embedder) but /query and /extract against the real synthetic corpus need real keys. The lowest-friction path is a single free Google AI Studio key with LLM_PROVIDER=gemini and EMBEDDINGS_PROVIDER=gemini (see the README's "Google-only quickstart").


Step 1 — Clone and start the stack

git clone https://github.com/div0rce/sentinel.git
cd sentinel

# Copy the env template and fill in your keys.
cp .env.example .env
$EDITOR .env   # set ANTHROPIC_API_KEY and OPENAI_API_KEY

# Bring up Postgres 16 + pgvector locally.
docker compose up -d db

# Install Python deps and run the dev API in one command.
make dev   # uv sync + uvicorn backend.app.main:app --reload, port 8000

In a second terminal, start the frontend:

cd frontend
npm ci
npm run dev    # Vite dev server on :5173, proxies /query|/extract|/review|/dashboard|/health to :8000

Open http://localhost:5173 in a browser. The SPA loads on the Query route.

screenshot: docs/screenshots/01-query-empty.png — empty Query page after first load.

Step 2 — Migrate and seed the synthetic corpus

Apply migrations and ingest the committed sample documents:

make migrate    # alembic upgrade head — creates tables + enables the vector extension
make seed       # python -m backend.app.ingest --path data/sample

The ingest pipeline is idempotent: re-running make seed is a no-op because the documents.hash check short-circuits identical content. The PII redaction pass runs before the chunk store, so the database never sees raw email addresses or phone numbers.

Sanity-check the database (should report ~15 documents and a few hundred chunks, depending on the latest synthetic corpus size):

psql postgres://sentinel:sentinel@localhost:5432/sentinel \
  -c "select count(*) as documents from documents;
      select count(*) as chunks from chunks;"

Step 3 — Query the corpus, get a cited answer

The flagship capability: ask a natural-language question, receive a source-cited answer.

In the SPA, paste this question on the Query page and click Ask:

What is the total amount due on the Initech Components invoice issued on 2026-01-22?

Or hit the API directly:

curl -s http://localhost:8000/query \
  -H 'Content-Type: application/json' \
  -d '{"query":"What is the total amount due on the Initech Components invoice issued on 2026-01-22?"}' \
  | jq

Expected response shape (the actual answer text is generated by Claude and may phrase the answer differently, but the citation invariants are deterministic):

{
  "status": "answered",
  "answer": "The total amount due on the Initech Components invoice issued on 2026-01-22 is $90,006.92 [chunk:42].",
  "citations": [{"chunk_id": 42, "document_id": 1, "score": 0.83, "text": "..."}],
  "reason": null
}

Two invariants you can verify by inspection:

  • Citation-or-refuse. Every claim is annotated [chunk:N]. The backend parses those markers and refuses the response if any cited id was not in the retrieved set (reason: "invalid_citation").
  • PII redaction is pre-LLM. The prompt sent to Claude contains [REDACTED:EMAIL] etc. in place of any matched PII; you can verify by setting SENTINEL_LOG_FORMAT=console and tailing the API logs.

screenshot: docs/screenshots/02-query-cited.png — Query page rendering an answered response with the citation chip showing source chunk text.

Step 4 — Query the corpus, observe a refusal

Ask a question with no support in the synthetic corpus:

When did the first humans land on the moon?

curl -s http://localhost:8000/query \
  -H 'Content-Type: application/json' \
  -d '{"query":"When did the first humans land on the moon?"}' | jq

Expected response:

{
  "status": "refused",
  "answer": "",
  "citations": [],
  "reason": "no_support"
}

The retrieval top-score is below RAG_SIMILARITY_THRESHOLD, so the system refuses before calling the LLM — no answer is hallucinated, no token spend. This is the citation-or-refuse policy doing its job.

screenshot: docs/screenshots/03-query-refusal.png — refusal banner on the Query page with reason: no_support.

Step 5 — Extract a structured record from a document

Pick the document id of an ingested invoice (a fresh make seed typically makes the first invoice id 1). Then call /extract with the registered invoice schema:

curl -s http://localhost:8000/extract \
  -H 'Content-Type: application/json' \
  -d '{"document_id":1, "schema_name":"invoice"}' | jq

Expected shape:

{
  "status": "ok",
  "document_id": 1,
  "schema_name": "invoice",
  "extraction_id": 1,
  "payload": {
    "vendor": "Initech Components",
    "invoice_number": "INV-2026000",
    "issue_date": "2026-01-22",
    "total_due": 90006.92
  },
  "field_confidence": {
    "vendor": 0.97,
    "invoice_number": 0.99,
    "issue_date": 0.94,
    "total_due": 0.71
  },
  "field_citations": {
    "vendor": [10],
    "invoice_number": [10],
    "issue_date": [10],
    "total_due": [12]
  },
  "requires_review": true,
  "low_confidence_fields": ["total_due"],
  "reason": null
}

Three things to point out to a reviewer:

  • Per-field provenance. field_citations maps each extracted field to the chunk id that supports it. The backend validates that each cited id was in the retrieval set (the same citation-validity rule as /query).
  • Per-field confidence. The LLM emits a self-reported confidence per field; we use it as a routing signal (M5 / M6) but do not interpret it as calibrated probability — see docs/evaluation.md.
  • Routing. requires_review is true because at least one field's confidence is below CONFIDENCE_REVIEW_THRESHOLD (default 0.75). The /extract handler has already routed the successful extraction through the workflow engine, which inserted a workflow_items row in needs_review and written one audit_events row tagged actor=system action=workflow.routed.

screenshot: docs/screenshots/04-extract-result.png — JSON viewer in the SPA Query/Extract panel showing the structured record with the low-confidence field highlighted.

Step 6 — Approve in the human-in-the-loop review queue

In the SPA, navigate to the Review tab. The queue lists every workflow item in needs_review state. The extraction from Step 5 should be at the top.

Enter your name in the actor field (any string; it's recorded verbatim as the audit event's actor), optionally add a note, and click Approve.

Behind the scenes:

POST /review/{id}/approve  body: {"actor": "Reviewer", "note": "verified against invoice PDF"}

The backend transitions the workflow item from needs_review to auto_approved (the terminal state) and writes one new audit_events row tagged actor=Reviewer action=review.approved — both in the same transaction.

Confirm the audit trail:

psql postgres://sentinel:sentinel@localhost:5432/sentinel <<'SQL'
WITH target_item AS (
  SELECT wi.id
  FROM workflow_items wi
  JOIN extractions e ON e.id = wi.extraction_id
  WHERE e.schema_name = 'invoice'
  ORDER BY wi.updated_at DESC, wi.id DESC
  LIMIT 1
)
SELECT ae.id, ae.actor, ae.action, ae.before, ae.after, ae.request_id, ae.ts
FROM audit_events ae
JOIN target_item ti
  ON ae.target_type = 'workflow_item'
 AND ae.target_id = ti.id
ORDER BY ae.id;
SQL

The CTE finds the most recently updated invoice workflow item, so the query does not depend on a particular local id sequence.

You should see two rows: one system / workflow.routed (from Step 5) and one Reviewer / review.approved (from this step). Replaying these events in order reproduces the workflow item's current state — the property tested in backend/tests/test_audit_events_append_only.py.

screenshot: docs/screenshots/05-review-queue.png — Review queue with one item highlighted, approve/reject buttons visible, actor field filled in.

Step 7 — Dashboard

Click the Dashboard tab. The lazy-loaded route renders four KPIs:

  • Volume — daily ingestion counts over the last 30 days (synthetic corpus → spike on the day of make seed).
  • Categories — extraction counts grouped by schema name (invoice, potentially others as you call /extract more).
  • Confidence histogram — distribution of per-field confidence across all extractions, bucketed.
  • SLA — count of needs_review items older than the configured threshold (default 24 h). Useful to demo what happens when reviewers fall behind.

Each panel is a Recharts component fed by a single typed API call from frontend/src/api.ts; the underlying endpoints are read-only and live in backend/app/routers/dashboard.py.

screenshot: docs/screenshots/06-dashboard.png — Dashboard rendering with the four panels populated. Capture this after Step 5 and Step 6 so the categories panel and the SLA panel both have data.

Teardown (local)

docker compose down -v   # removes the Postgres volume so a fresh demo starts clean

Optional — repeat the demo on AWS

The M10 Terraform stack provisions an ephemeral demo deployment in us-east-1. The full operator runbook (apply / write secrets / deploy / destroy) lives in infra/README.md. Short version:

cd infra
export TF_VAR_db_password="$(openssl rand -base64 24)"
export TF_VAR_github_repository="OWNER/sentinel"   # for the OIDC role; optional
terraform fmt -recursive -check
terraform init
terraform validate
terraform plan -out=plan.tfplan
terraform apply plan.tfplan

# Write the API keys out-of-band (not in tfstate)
aws ssm put-parameter --name /sentinel/anthropic_api_key \
  --type SecureString --value "$ANTHROPIC_API_KEY" --overwrite
aws ssm put-parameter --name /sentinel/openai_api_key \
  --type SecureString --value "$OPENAI_API_KEY" --overwrite
# Only if deploying with -var='llm_provider=gemini' / 'embeddings_provider=gemini':
aws ssm put-parameter --name /sentinel/gemini_api_key \
  --type SecureString --value "$GEMINI_API_KEY" --overwrite

# Force the backend to pick up the new secret values
aws ecs update-service --cluster sentinel-cluster \
  --service sentinel-backend --force-new-deployment --no-cli-pager

# The ALB DNS name is the demo URL.
terraform output alb_dns_name

Migrations against RDS run as a one-off Fargate task; recipe in infra/README.md. After capturing screenshots, immediately:

terraform destroy

The cost posture (~$45/month idle floor, dominated by the ALB + Fargate + RDS) is documented in infra/README.md. Leaving the stack running overnight is ~$1.50; leaving it for a month is ~$45.


Cross-references

All screenshots referenced above are placeholders. Capture them on a real demo run; commit to docs/screenshots/ (gitignored by default — flip the rule when capturing for a portfolio review).