PoliCast is a real-time, signal-driven prediction market platform. It grounds each forecast in the right kind of evidence β live market data (crypto / FX / commodity prices, mirrored prediction-market odds, polls) for quantitative questions, and entity-matched, LLM-vetted news for qualitative ones β then re-forecasts the moment those signals move and pushes the change live to connected users.
- No page refreshes. No manual polling. No noise. Quiet markets hold steady; markets with fresh evidence update on their own.
Markets dashboard. Live YES/NO odds, probability sparklines and the signal type behind each market (news vs. market data).
Market detail. Probability history for a quantitative market ("Will Bitcoin exceed $200K before 2027?"), anchored to live CoinGecko readings.
| Key drivers and market info | Live signal feed: readings the engine used |
|---|---|
![]() |
![]() |
- Correct signals, not keyword soup: Every question carries structured intent β entities (with aliases, so Beijing matches China), keywords, and the metric that resolves it. Signals are scored with an entity-aware relevance model, deduplicated, and then vetted by an LLM pass that judges each news item against the question's resolution criteria and tags its YES/NO direction.
- Grounded in real market data: Quantitative questions ("Will Bitcoin exceed $200K?", "Will oil top $120/bbl?") are anchored to live readings from CoinGecko, Yahoo Finance, and Polymarket β not guessed from headlines.
- Event-driven + tiered refresh: A fast lane polls quantitative markets every ~20 min; scheduled ingestion + a baseline sweep cover the rest. A question re-forecasts the instant meaningful signals or a price move land β and a question with no new signals does not move (no random-walk drift).
- Live probability streams: Only markets that actually moved are pushed to connected browsers over a persistent WebSocket channel.
- On-demand AI briefings: Progressive 3-paragraph briefings (Current Landscape, Key Catalysts, Outlook) streamed via Server-Sent Events (SSE).
-
Guardrails: Per-cycle movement is clamped (Β±8% news / Β±20% authoritative market signals) within a
$[2%, 98%]$ band; mutations and scheduler controls are gated by an admin API key.
PoliCast is a decoupled client-server system. Three APScheduler jobs feed a single signal β forecast pipeline; only markets that actually move are broadcast to clients.
flowchart TB
subgraph Client["Client Β· React 19 Β· Polymarket-style"]
UI[Dashboard & Detail]
end
subgraph API["FastAPI Service"]
REST[REST /api/v1]
WS[WebSocket /ws]
SSE[SSE /analysis]
end
subgraph Sched["APScheduler Jobs"]
QJOB[refresh_quant<br/>~20 min Β· quantitative only]
IJOB[ingest_signals<br/>~2 h Β· all questions]
UJOB[update_markets<br/>~4 h Β· baseline sweep]
end
subgraph Pipe["Signal β Forecast Pipeline"]
NORM[Normalizer<br/>entity relevance Β· dedup Β· staleness]
RELV[LLM Relevance + Direction<br/>vets news Β· tags YES/NO]
ENG[Probability Engine<br/>quantitative grounding Β· guardrails]
end
subgraph Sources["Data Sources"]
MKT[CoinGecko / Yahoo Finance]
ODDS[Polymarket odds / Polls]
NEWS[GDELT / RSS / NewsAPI]
end
LLM[Groq Β· Claude Β· OpenAI fallback]
DB[(SQLite / PostgreSQL)]
UI -->|initial load| REST
WS -->|live updates| UI
UI -->|request briefing| SSE
IJOB --> Sources
QJOB --> MKT
QJOB --> ODDS
Sources --> NORM
NORM --> RELV
RELV -. Groq .-> LLM
RELV --> DB
IJOB -->|significant signal / metric move| ENG
QJOB -->|metric move| ENG
UJOB --> ENG
ENG -. LLM .-> LLM
ENG --> DB
ENG -->|broadcast movers| WS
REST --> DB
SSE -. LLM .-> LLM
How a signal becomes a live probability update, end to end:
- Structured intent β each question stores
entities(+ aliases),keywords, asignal_type, and ametric_configdescribing the number that resolves it. - Ingestion β the normalizer runs the matching adapters per question: quantitative sources (CoinGecko / Yahoo Finance / Polymarket / polls) for numeric questions, plus news (GDELT / RSS / NewsAPI). Signals are entity-scored, deduplicated, and stale-filtered.
- LLM relevance + direction β surviving news candidates are batched to the LLM, which drops off-topic items and tags each with a YES/NO direction against the resolution criteria. Quantitative readings skip this step (already authoritative).
- Persistence β signals are stored with their
signal_kind(news / market / odds / poll), relevance, direction, and anymetric_value. - Refresh trigger β a question is re-forecast immediately when enough significant news arrives or its metric moves past a threshold; the fast lane does this for quantitative markets every ~20 min, and a slow sweep is the safety net. A question with no new signals is skipped.
- Forecast β the engine builds a grounded prompt (price-vs-threshold gap + direction-tagged news) and asks the LLM for a new probability + drivers. Movement is guardrail-clamped. If the LLM is unavailable, quantitative markets fall back to a deterministic metric nudge; everything else holds steady.
- Broadcast β only markets that actually moved get a new history point and a WebSocket push to the UI.
- Briefings β on demand, the SSE endpoint streams a cached-or-fresh 3-paragraph analysis.
sequenceDiagram
participant SRC as Data Sources
participant NRM as Normalizer
participant LLM as LLM Β· Groq
participant DB as Database
participant ENG as Probability Engine
participant WS as WebSocket
participant UI as React Client
SRC->>NRM: fetch (entities + keywords + metric config)
NRM->>NRM: relevance score Β· dedup Β· drop stale
NRM->>LLM: batch news candidates
LLM-->>NRM: relevant? + direction (YES/NO)
NRM->>DB: persist signals (news + metric readings)
alt significant news OR metric moved β₯ threshold
NRM->>ENG: event-driven re-forecast (this question)
ENG->>LLM: grounded prompt (price gap + direction-tagged news)
LLM-->>ENG: new probability + drivers
ENG->>DB: write probability + history (only if moved)
ENG->>WS: broadcast (movers only)
WS->>UI: live update
else no new signals
NRM-->>NRM: skip β market holds steady (no drift)
end
| Layer | Technologies |
|---|---|
| Frontend | React 19, React Router, Vite, Chart.js, Polymarket-style dark theme |
| Backend | FastAPI, SQLAlchemy ORM, Alembic Migrations, APScheduler |
| Database | SQLite (Local Development) / PostgreSQL (Production) |
| LLM Inference | Groq (llama-3.3-70b-versatile) |
| Market Data | CoinGecko (crypto), Yahoo Finance (commodities/FX/indices), FRED (economic data, optional key), Polymarket + Manifold (mirrored odds) |
| News Sources | GDELT 2.0 Doc API, RSS (BBC, Guardian, Al Jazeera, CNBC, oilprice), GNews + NewsAPI (optional keys) |
| Relevance | Entity-aware lexical scoring β LLM relevance & direction pass |
| Authentication | Header-based API key validation (X-API-Key) |
.
βββ backend/
β βββ main.py # FastAPI app setup, WebSocket route, and scheduler lifespan
β βββ routes.py # API routing for 22 endpoints (CRUD, resolve, history)
β βββ models.py # Database schema mappings (5 core tables)
β βββ database.py # Session and DB engine configurations
β βββ config.py # Pydantic Settings singleton parses .env
β βββ auth.py # X-API-Key validation helper
β βββ seed.py # Seeds 52 geopolitical prediction markets
β βββ ai/
β β βββ llm_client.py # Multi-provider client wrapper with automatic fallback
β β βββ guardrails.py # Single source of truth for probability clamping
β β βββ relevance_filter.py # LLM relevance + YES/NO direction pass over news
β β βββ probability_engine.py # Grounded prompts, structured JSON, delta guardrails
β βββ sources/
β β βββ base.py # QuestionContext + BaseSourceAdapter + RawSignalData
β β βββ relevance.py # Shared entity-aware relevance scoring
β β βββ gdelt.py # GDELT Doc API client (entity/keyword queries)
β β βββ rss.py # Multi-feed RSS scraper
β β βββ newsapi.py # NewsAPI.org adapter (optional key)
β β βββ gnews.py # GNews adapter β production-safe news (optional key)
β β βββ market_data.py # CoinGecko + Yahoo Finance quantitative readings
β β βββ fred.py # FRED economic indicators (optional key)
β β βββ prediction_markets.py # Polymarket + Manifold odds mirrors
β β βββ polls.py # Config-driven polling averages
β β βββ normalizer.py # Runs adapters, dedup + staleness, ranking
β βββ scheduler/
β β βββ jobs.py # 3 jobs: refresh_quant, ingest_signals, update_markets
β βββ alembic/ # Database migration scripts (Current Head: b2c3d4e5f6a1)
βββ frontend/
βββ src/
β βββ App.jsx # App root, handles bootstrap fetch and socket patches
β βββ hooks/
β β βββ useMarketSocket.js # WebSocket client with automatic exponential backoff
β βββ pages/
β β βββ Dashboard.jsx # Active prediction markets grid
β β βββ QuestionDetail.jsx # Focus detail page containing live chart and drivers
β βββ components/
β β βββ Navbar.jsx # Header bar showing connection status indicators
β β βββ MarketAnalysis.jsx # Progressively renders streamed SSE briefings
β β βββ SignalsPanel.jsx # Displays raw news signals backing the prediction
β βββ api/questions.js # Endpoint HTTP utility functions
βββ .env # API URL configurations
βββ package.json
git clone https://github.com/zacktiger/PoliCast.git
cd PoliCastCreate a .env file in the root directory:
# Database configuration (Defaults to SQLite for local development)
DATABASE_URL=sqlite:///./geopolitics.db
# LLM Providers (Configure at least one to enable AI scoring)
GROQ_API_KEY=your-groq-api-key-here # Primary (https://console.groq.com/keys)
ANTHROPIC_API_KEY= # Fallback (https://console.anthropic.com/)
OPENAI_API_KEY= # Fallback (https://platform.openai.com/)
# Control-plane Security (Protects mutations, deletions, and scheduler configuration)
# Required for DELETE /questions, PUT /settings, and POST /scheduler/pause|resume
ADMIN_API_KEY=your-secure-admin-api-key-hereNote
If ADMIN_API_KEY is not provided, endpoints performing data mutations or scheduler modifications will return 503 Service Unavailable for security.
# Setup virtual environment
python -m venv .venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
# Install dependencies & run migrations
pip install -r backend/requirements.txt
python -m alembic upgrade head
# Seed initial geopolitical markets (inserts 52 markets)
python -m backend.seed
# Start FastAPI development server
uvicorn backend.main:app --reload --host 0.0.0.0 --port 8000- API Documentation: Accessible at http://localhost:8000/docs
cd frontend
npm install
npm run dev- Dashboard App: Open your browser to http://localhost:5173
| Method | Path | Description | Authentication |
|---|---|---|---|
GET |
/health |
Service health & active LLM provider check | None |
GET |
/questions |
Fetch all prediction market summaries | None |
POST |
/questions |
Create a new prediction market | None |
GET |
/questions/{id} |
Get market details, histories, and active drivers | None |
PUT |
/questions/{id} |
Modify market metadata | None |
DELETE |
/questions/{id} |
Delete prediction market and associated rows | Admin Key Required |
POST |
/questions/{id}/tick |
Trigger a manual mock market fluctuation | None |
POST |
/questions/{id}/resolve |
Finalize prediction market outcomes | None |
POST |
/questions/{id}/analysis |
Stream AI analysis briefing via Server-Sent Events | None |
GET |
/questions/{id}/signals |
Retrieve raw ingested news feed signals | None |
GET |
/questions/{id}/history |
Extract historical probability records | None |
GET |
/questions/{id}/drivers |
View active forces/reasons pushing the target probability | None |
POST |
/questions/{id}/drivers |
Add a manual probability driver | None |
DELETE |
/questions/{id}/drivers/{d_id} |
Remove a probability driver | None |
| Method | Path | Description | Authentication |
|---|---|---|---|
GET |
/sources |
Get active news adapters and sync statistics | None |
POST |
/sources/fetch |
Manually run GDELT signal ingestion in background | None |
POST |
/markets/update |
Manually trigger AI forecasting loop in background | None |
GET |
/settings |
Get current parameters and guardrail thresholds | None |
PUT |
/settings |
Adjust guardrails or TTL criteria | Admin Key Required |
GET |
/scheduler/status |
View background jobs status and upcoming runs | None |
POST |
/scheduler/pause |
Pause scheduled background tasks | Admin Key Required |
POST |
/scheduler/resume |
Resume scheduled background tasks | Admin Key Required |
Provides read-only, real-time broadcasts. The client initiates a connection and receives push frames on target events:
// 1. Sent immediately on connection
{ "type": "connected", "message": "PoliCast WS ready" }
// 2. Sent every 30 seconds as server-side keep-alive
{ "type": "ping" }
// 3. Pushed when probability calculations update
{
"type": "market_update",
"data": [
{
"id": "uuid-here",
"probability": 78.4,
"last_updated": "2026-06-28T13:00:00"
}
]
}ingest_signals_job (every N hours) β βββ for each question (10 in parallel) β βββ adapters fetch articles + market/poll/odds data β β βββ relevance.py β score each article 0β1 (term matching) β βββ normalizer β drop score < 0.25, drop stale, dedup β βββ relevance_filter.py (LLM #1) β βββ drop off-topic news β βββ set relevance_score + direction β βββ save to DB as RawSignal (is_processed = False) β βββ important news? (β₯2 articles with rel β₯ 0.5, or a quantitative move β₯1%) βββ YES β forecast this question NOW ββββββββββ βββ NO β wait for the 4h sweep βββββββββββββββ€ β update_markets_job (every 4h) βββββββββββββββββββββ€ βΌ _forecast_question βββ load unprocessed signals β (last 24h, top 30 by relevance) βββ probability_engine.py (LLM #2) β βββ new probability + reasoning + drivers β βββ LLM failed? β βββ have market data β move toward it β βββ no β keep probability as is βββ guardrails.py β max Β±8%, keep within 2β98% βββ mark signals processed βββ moved? βββ YES β save history + drivers β β WebSocket β chart updates βββ NO β nothing
-
Tiered, event-driven refresh β not a fixed clock.
-
Fast lane (
refresh_quant, ~20 min): re-reads live prices/odds for quantitative markets and re-forecasts any whose metric moved past a threshold. -
Ingestion (
ingest_signals, ~2 h): refreshes news + quantitative signals for every question and event-triggers a re-forecast when meaningful signals arrive. -
Baseline sweep (
update_markets, ~4 h): safety net that forecasts any question still holding unprocessed signals. - Intervals are config-driven and can be changed at runtime via
PUT /settings(jobs are rescheduled in place).
-
Fast lane (
-
Grounded forecasting. For quantitative questions the engine surfaces the live reading and the gap to the resolution threshold ("68% below
$200K β YES less likely"), then blends in LLM-vetted, direction-tagged news. Movement is clamped per cycle ($ \pm 8%$ news,$\pm 20%$ authoritative market signals) inside a$[2%, 98%]$ band. - No random walk. A question with no new signals is left untouched β no history point, no drift. If the LLM is unavailable, quantitative markets fall back to a deterministic metric nudge; qualitative markets simply hold.
-
SSE briefings (on-demand). Cached by
prob_moved_atand regenerated only when the probability actually changed, streamed word-by-word into three sections (Current Landscape, Key Catalysts, Outlook).
- User accounts + mock trading: JWT authentication, mock portfolio balances, YES/NO contract purchases, and historical P&L analytics.
- Production deployment: Production-grade multi-container setups using PostgreSQL, Nginx reverse proxying, SSL setup, and CI/CD pipelines.
- Test coverage: Complete pytest configurations, LLM-response mocking, and edge-case guardrail testing.



