Skip to content
zacktigerPublic

About

A geopolitical prediction platform combining agentic AI, automated news signal extraction, and probabilistic forecasting models. The system continuously updates event probabilities and visualizes trends through an interactive React dashboard.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

26 Commits

Folders and files

Repository files navigation

PoliCast

Build Status License FastAPI React PostgreSQL Groq WebSocket

PoliCast is a real-time, signal-driven prediction market platform. It grounds each forecast in the right kind of evidence β€” live market data (crypto / FX / commodity prices, mirrored prediction-market odds, polls) for quantitative questions, and entity-matched, LLM-vetted news for qualitative ones β€” then re-forecasts the moment those signals move and pushes the change live to connected users.

  • No page refreshes. No manual polling. No noise. Quiet markets hold steady; markets with fresh evidence update on their own.

πŸ“Έ Screenshots

Markets dashboard. Live YES/NO odds, probability sparklines and the signal type behind each market (news vs. market data).

Markets dashboard

Market detail. Probability history for a quantitative market ("Will Bitcoin exceed $200K before 2027?"), anchored to live CoinGecko readings.

Market detail with probability history

Key drivers and market info Live signal feed: readings the engine used
Key drivers and market info Live signal feed

🌟 Key Features

  • Correct signals, not keyword soup: Every question carries structured intent β€” entities (with aliases, so Beijing matches China), keywords, and the metric that resolves it. Signals are scored with an entity-aware relevance model, deduplicated, and then vetted by an LLM pass that judges each news item against the question's resolution criteria and tags its YES/NO direction.
  • Grounded in real market data: Quantitative questions ("Will Bitcoin exceed $200K?", "Will oil top $120/bbl?") are anchored to live readings from CoinGecko, Yahoo Finance, and Polymarket β€” not guessed from headlines.
  • Event-driven + tiered refresh: A fast lane polls quantitative markets every ~20 min; scheduled ingestion + a baseline sweep cover the rest. A question re-forecasts the instant meaningful signals or a price move land β€” and a question with no new signals does not move (no random-walk drift).
  • Live probability streams: Only markets that actually moved are pushed to connected browsers over a persistent WebSocket channel.
  • On-demand AI briefings: Progressive 3-paragraph briefings (Current Landscape, Key Catalysts, Outlook) streamed via Server-Sent Events (SSE).
  • Guardrails: Per-cycle movement is clamped (Β±8% news / Β±20% authoritative market signals) within a $[2%, 98%]$ band; mutations and scheduler controls are gated by an admin API key.

πŸ“ Architecture

PoliCast is a decoupled client-server system. Three APScheduler jobs feed a single signal β†’ forecast pipeline; only markets that actually move are broadcast to clients.

flowchart TB
    subgraph Client["Client Β· React 19 Β· Polymarket-style"]
        UI[Dashboard & Detail]
    end

    subgraph API["FastAPI Service"]
        REST[REST /api/v1]
        WS[WebSocket /ws]
        SSE[SSE /analysis]
    end

    subgraph Sched["APScheduler Jobs"]
        QJOB[refresh_quant<br/>~20 min Β· quantitative only]
        IJOB[ingest_signals<br/>~2 h Β· all questions]
        UJOB[update_markets<br/>~4 h Β· baseline sweep]
    end

    subgraph Pipe["Signal β†’ Forecast Pipeline"]
        NORM[Normalizer<br/>entity relevance Β· dedup Β· staleness]
        RELV[LLM Relevance + Direction<br/>vets news Β· tags YES/NO]
        ENG[Probability Engine<br/>quantitative grounding Β· guardrails]
    end

    subgraph Sources["Data Sources"]
        MKT[CoinGecko / Yahoo Finance]
        ODDS[Polymarket odds / Polls]
        NEWS[GDELT / RSS / NewsAPI]
    end

    LLM[Groq Β· Claude Β· OpenAI fallback]
    DB[(SQLite / PostgreSQL)]

    UI -->|initial load| REST
    WS -->|live updates| UI
    UI -->|request briefing| SSE

    IJOB --> Sources
    QJOB --> MKT
    QJOB --> ODDS
    Sources --> NORM
    NORM --> RELV
    RELV -. Groq .-> LLM
    RELV --> DB

    IJOB -->|significant signal / metric move| ENG
    QJOB -->|metric move| ENG
    UJOB --> ENG
    ENG -. LLM .-> LLM
    ENG --> DB
    ENG -->|broadcast movers| WS

    REST --> DB
    SSE -. LLM .-> LLM
Loading

πŸ”„ Data Flow

How a signal becomes a live probability update, end to end:

  1. Structured intent β€” each question stores entities (+ aliases), keywords, a signal_type, and a metric_config describing the number that resolves it.
  2. Ingestion β€” the normalizer runs the matching adapters per question: quantitative sources (CoinGecko / Yahoo Finance / Polymarket / polls) for numeric questions, plus news (GDELT / RSS / NewsAPI). Signals are entity-scored, deduplicated, and stale-filtered.
  3. LLM relevance + direction β€” surviving news candidates are batched to the LLM, which drops off-topic items and tags each with a YES/NO direction against the resolution criteria. Quantitative readings skip this step (already authoritative).
  4. Persistence β€” signals are stored with their signal_kind (news / market / odds / poll), relevance, direction, and any metric_value.
  5. Refresh trigger β€” a question is re-forecast immediately when enough significant news arrives or its metric moves past a threshold; the fast lane does this for quantitative markets every ~20 min, and a slow sweep is the safety net. A question with no new signals is skipped.
  6. Forecast β€” the engine builds a grounded prompt (price-vs-threshold gap + direction-tagged news) and asks the LLM for a new probability + drivers. Movement is guardrail-clamped. If the LLM is unavailable, quantitative markets fall back to a deterministic metric nudge; everything else holds steady.
  7. Broadcast β€” only markets that actually moved get a new history point and a WebSocket push to the UI.
  8. Briefings β€” on demand, the SSE endpoint streams a cached-or-fresh 3-paragraph analysis.
sequenceDiagram
    participant SRC as Data Sources
    participant NRM as Normalizer
    participant LLM as LLM Β· Groq
    participant DB as Database
    participant ENG as Probability Engine
    participant WS as WebSocket
    participant UI as React Client

    SRC->>NRM: fetch (entities + keywords + metric config)
    NRM->>NRM: relevance score Β· dedup Β· drop stale
    NRM->>LLM: batch news candidates
    LLM-->>NRM: relevant? + direction (YES/NO)
    NRM->>DB: persist signals (news + metric readings)
    alt significant news OR metric moved β‰₯ threshold
        NRM->>ENG: event-driven re-forecast (this question)
        ENG->>LLM: grounded prompt (price gap + direction-tagged news)
        LLM-->>ENG: new probability + drivers
        ENG->>DB: write probability + history (only if moved)
        ENG->>WS: broadcast (movers only)
        WS->>UI: live update
    else no new signals
        NRM-->>NRM: skip β€” market holds steady (no drift)
    end
Loading

πŸ› οΈ Tech Stack

Layer Technologies
Frontend React 19, React Router, Vite, Chart.js, Polymarket-style dark theme
Backend FastAPI, SQLAlchemy ORM, Alembic Migrations, APScheduler
Database SQLite (Local Development) / PostgreSQL (Production)
LLM Inference Groq (llama-3.3-70b-versatile) $\rightarrow$ Anthropic Claude $\rightarrow$ OpenAI GPT-4o
Market Data CoinGecko (crypto), Yahoo Finance (commodities/FX/indices), FRED (economic data, optional key), Polymarket + Manifold (mirrored odds)
News Sources GDELT 2.0 Doc API, RSS (BBC, Guardian, Al Jazeera, CNBC, oilprice), GNews + NewsAPI (optional keys)
Relevance Entity-aware lexical scoring β†’ LLM relevance & direction pass
Authentication Header-based API key validation (X-API-Key)

πŸ“‚ Repository Structure

.
β”œβ”€β”€ backend/
β”‚   β”œβ”€β”€ main.py              # FastAPI app setup, WebSocket route, and scheduler lifespan
β”‚   β”œβ”€β”€ routes.py            # API routing for 22 endpoints (CRUD, resolve, history)
β”‚   β”œβ”€β”€ models.py            # Database schema mappings (5 core tables)
β”‚   β”œβ”€β”€ database.py          # Session and DB engine configurations
β”‚   β”œβ”€β”€ config.py            # Pydantic Settings singleton parses .env
β”‚   β”œβ”€β”€ auth.py              # X-API-Key validation helper
β”‚   β”œβ”€β”€ seed.py              # Seeds 52 geopolitical prediction markets
β”‚   β”œβ”€β”€ ai/
β”‚   β”‚   β”œβ”€β”€ llm_client.py         # Multi-provider client wrapper with automatic fallback
β”‚   β”‚   β”œβ”€β”€ guardrails.py         # Single source of truth for probability clamping
β”‚   β”‚   β”œβ”€β”€ relevance_filter.py   # LLM relevance + YES/NO direction pass over news
β”‚   β”‚   └── probability_engine.py # Grounded prompts, structured JSON, delta guardrails
β”‚   β”œβ”€β”€ sources/
β”‚   β”‚   β”œβ”€β”€ base.py               # QuestionContext + BaseSourceAdapter + RawSignalData
β”‚   β”‚   β”œβ”€β”€ relevance.py          # Shared entity-aware relevance scoring
β”‚   β”‚   β”œβ”€β”€ gdelt.py              # GDELT Doc API client (entity/keyword queries)
β”‚   β”‚   β”œβ”€β”€ rss.py                # Multi-feed RSS scraper
β”‚   β”‚   β”œβ”€β”€ newsapi.py            # NewsAPI.org adapter (optional key)
β”‚   β”‚   β”œβ”€β”€ gnews.py              # GNews adapter β€” production-safe news (optional key)
β”‚   β”‚   β”œβ”€β”€ market_data.py        # CoinGecko + Yahoo Finance quantitative readings
β”‚   β”‚   β”œβ”€β”€ fred.py               # FRED economic indicators (optional key)
β”‚   β”‚   β”œβ”€β”€ prediction_markets.py # Polymarket + Manifold odds mirrors
β”‚   β”‚   β”œβ”€β”€ polls.py              # Config-driven polling averages
β”‚   β”‚   └── normalizer.py         # Runs adapters, dedup + staleness, ranking
β”‚   β”œβ”€β”€ scheduler/
β”‚   β”‚   └── jobs.py               # 3 jobs: refresh_quant, ingest_signals, update_markets
β”‚   └── alembic/             # Database migration scripts (Current Head: b2c3d4e5f6a1)
└── frontend/
    β”œβ”€β”€ src/
    β”‚   β”œβ”€β”€ App.jsx          # App root, handles bootstrap fetch and socket patches
    β”‚   β”œβ”€β”€ hooks/
    β”‚   β”‚   └── useMarketSocket.js # WebSocket client with automatic exponential backoff
    β”‚   β”œβ”€β”€ pages/
    β”‚   β”‚   β”œβ”€β”€ Dashboard.jsx # Active prediction markets grid
    β”‚   β”‚   └── QuestionDetail.jsx # Focus detail page containing live chart and drivers
    β”‚   β”œβ”€β”€ components/
    β”‚   β”‚   β”œβ”€β”€ Navbar.jsx   # Header bar showing connection status indicators
    β”‚   β”‚   β”œβ”€β”€ MarketAnalysis.jsx # Progressively renders streamed SSE briefings
    β”‚   β”‚   └── SignalsPanel.jsx # Displays raw news signals backing the prediction
    β”‚   └── api/questions.js # Endpoint HTTP utility functions
    β”œβ”€β”€ .env                 # API URL configurations
    └── package.json

πŸš€ Local Quickstart

1. Clone the Repository

git clone https://github.com/zacktiger/PoliCast.git
cd PoliCast

2. Configure Environment Variables

Create a .env file in the root directory:

# Database configuration (Defaults to SQLite for local development)
DATABASE_URL=sqlite:///./geopolitics.db

# LLM Providers (Configure at least one to enable AI scoring)
GROQ_API_KEY=your-groq-api-key-here        # Primary (https://console.groq.com/keys)
ANTHROPIC_API_KEY=                          # Fallback (https://console.anthropic.com/)
OPENAI_API_KEY=                             # Fallback (https://platform.openai.com/)

# Control-plane Security (Protects mutations, deletions, and scheduler configuration)
# Required for DELETE /questions, PUT /settings, and POST /scheduler/pause|resume
ADMIN_API_KEY=your-secure-admin-api-key-here

Note

If ADMIN_API_KEY is not provided, endpoints performing data mutations or scheduler modifications will return 503 Service Unavailable for security.

3. Spin up the Backend

# Setup virtual environment
python -m venv .venv
source .venv/bin/activate  # On Windows: .venv\Scripts\activate

# Install dependencies & run migrations
pip install -r backend/requirements.txt
python -m alembic upgrade head

# Seed initial geopolitical markets (inserts 52 markets)
python -m backend.seed

# Start FastAPI development server
uvicorn backend.main:app --reload --host 0.0.0.0 --port 8000

4. Run the Frontend

cd frontend
npm install
npm run dev

πŸ”Œ API Reference

Core Market Actions

Method Path Description Authentication
GET /health Service health & active LLM provider check None
GET /questions Fetch all prediction market summaries None
POST /questions Create a new prediction market None
GET /questions/{id} Get market details, histories, and active drivers None
PUT /questions/{id} Modify market metadata None
DELETE /questions/{id} Delete prediction market and associated rows Admin Key Required
POST /questions/{id}/tick Trigger a manual mock market fluctuation None
POST /questions/{id}/resolve Finalize prediction market outcomes None
POST /questions/{id}/analysis Stream AI analysis briefing via Server-Sent Events None
GET /questions/{id}/signals Retrieve raw ingested news feed signals None
GET /questions/{id}/history Extract historical probability records None
GET /questions/{id}/drivers View active forces/reasons pushing the target probability None
POST /questions/{id}/drivers Add a manual probability driver None
DELETE /questions/{id}/drivers/{d_id} Remove a probability driver None

Operations & Scheduler

Method Path Description Authentication
GET /sources Get active news adapters and sync statistics None
POST /sources/fetch Manually run GDELT signal ingestion in background None
POST /markets/update Manually trigger AI forecasting loop in background None
GET /settings Get current parameters and guardrail thresholds None
PUT /settings Adjust guardrails or TTL criteria Admin Key Required
GET /scheduler/status View background jobs status and upcoming runs None
POST /scheduler/pause Pause scheduled background tasks Admin Key Required
POST /scheduler/resume Resume scheduled background tasks Admin Key Required

πŸ›œ WebSocket API (/ws)

Provides read-only, real-time broadcasts. The client initiates a connection and receives push frames on target events:

// 1. Sent immediately on connection
{ "type": "connected", "message": "PoliCast WS ready" }

// 2. Sent every 30 seconds as server-side keep-alive
{ "type": "ping" }

// 3. Pushed when probability calculations update
{
  "type": "market_update",
  "data": [
    {
      "id": "uuid-here",
      "probability": 78.4,
      "last_updated": "2026-06-28T13:00:00"
    }
  ]
}

ingest_signals_job (every N hours) β”‚ β”œβ”€β”€ for each question (10 in parallel) β”‚ β”œβ”€β”€ adapters fetch articles + market/poll/odds data β”‚ β”‚ └── relevance.py β†’ score each article 0–1 (term matching) β”‚ β”œβ”€β”€ normalizer β†’ drop score < 0.25, drop stale, dedup β”‚ └── relevance_filter.py (LLM #1) β”‚ β”œβ”€β”€ drop off-topic news β”‚ └── set relevance_score + direction β”‚ β”œβ”€β”€ save to DB as RawSignal (is_processed = False) β”‚ └── important news? (β‰₯2 articles with rel β‰₯ 0.5, or a quantitative move β‰₯1%) β”œβ”€β”€ YES β†’ forecast this question NOW ─────────┐ └── NO β†’ wait for the 4h sweep ─────────────── β”‚ update_markets_job (every 4h) ───────────────────── β–Ό _forecast_question β”œβ”€β”€ load unprocessed signals β”‚ (last 24h, top 30 by relevance) β”œβ”€β”€ probability_engine.py (LLM #2) β”‚ └── new probability + reasoning + drivers β”‚ └── LLM failed? β”‚ β”œβ”€β”€ have market data β†’ move toward it β”‚ └── no β†’ keep probability as is β”œβ”€β”€ guardrails.py β†’ max Β±8%, keep within 2–98% β”œβ”€β”€ mark signals processed └── moved? β”œβ”€β”€ YES β†’ save history + drivers β”‚ β†’ WebSocket β†’ chart updates └── NO β†’ nothing


πŸ’‘ How the Engine Works

  1. Tiered, event-driven refresh β€” not a fixed clock.
    • Fast lane (refresh_quant, ~20 min): re-reads live prices/odds for quantitative markets and re-forecasts any whose metric moved past a threshold.
    • Ingestion (ingest_signals, ~2 h): refreshes news + quantitative signals for every question and event-triggers a re-forecast when meaningful signals arrive.
    • Baseline sweep (update_markets, ~4 h): safety net that forecasts any question still holding unprocessed signals.
    • Intervals are config-driven and can be changed at runtime via PUT /settings (jobs are rescheduled in place).
  2. Grounded forecasting. For quantitative questions the engine surfaces the live reading and the gap to the resolution threshold ("68% below $200K β†’ YES less likely"), then blends in LLM-vetted, direction-tagged news. Movement is clamped per cycle ($\pm 8%$ news, $\pm 20%$ authoritative market signals) inside a $[2%, 98%]$ band.
  3. No random walk. A question with no new signals is left untouched β€” no history point, no drift. If the LLM is unavailable, quantitative markets fall back to a deterministic metric nudge; qualitative markets simply hold.
  4. SSE briefings (on-demand). Cached by prob_moved_at and regenerated only when the probability actually changed, streamed word-by-word into three sections (Current Landscape, Key Catalysts, Outlook).

πŸ—ΊοΈ Roadmap

  • User accounts + mock trading: JWT authentication, mock portfolio balances, YES/NO contract purchases, and historical P&L analytics.
  • Production deployment: Production-grade multi-container setups using PostgreSQL, Nginx reverse proxying, SSL setup, and CI/CD pipelines.
  • Test coverage: Complete pytest configurations, LLM-response mocking, and edge-case guardrail testing.

About

A geopolitical prediction platform combining agentic AI, automated news signal extraction, and probabilistic forecasting models. The system continuously updates event probabilities and visualizes trends through an interactive React dashboard.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages