Parses, sanitizes, and visualizes crowd-sourced Rivian R2 pre-order data into an interactive Plotly dashboard. It pulls two live Google Sheets (an orders/deliveries tracker and a separate reservations-only tracker) via their CSV export endpoints, cleans them (dedup, VIN recovery, date normalization, geo enrichment), removes reservation-holders who have already ordered, and produces a tidy CSV plus a 10-chart HTML dashboard.
r2_dashboard run-in-place launcher (./r2_dashboard); also `python3 src/pipeline.py`
requirements.txt pandas, numpy, plotly, PyYAML, beautifulsoup4
src/
config.py paths, run timestamps + loaders for the conf/ YAML files
pipeline.py main() orchestration + report printing (fetch -> clean -> render)
ingest/ get + clean the data
fetch.py live-sheet fetch with caching + change detection
parsing.py pure parsing / VIN / date / geo helpers
schema_check.py locates columns by name; verifies them against schema.yaml
loaders.py load_and_clean, load_reservations
render/ build the webpage
colors.py color-transform functions + derived display palettes
charts.py the ten fig_* chart builders + helpers
page.py BeautifulSoup DOM population, HTML helpers, SECTIONS, build_dashboard
templates/ valid standalone page shell + assets, filled at render time
page.html valid HTML shell (empty id'd slots, populated via the DOM)
styles.css page stylesheet (its own <style> slot)
head.js pre-paint theme set (no flash)
theme.js re-tint chart chrome on light/dark toggle
nav.js sidebar hamburger + scroll-spy
conf/ data/config YAML (loaded by config.py)
dimensions.yaml category vocabulary: each column's label, order, blank handling, caveat, and
per-category label/color/marker (paints, wheels, interiors, regions, ...);
published as r2_dimensions.json
palette.yaml chart fills that don't name a category (take-rate, timeline, accents)
theme.yaml page & chart chrome for light/dark (CSS variables + chart retint colors)
schema.yaml sheet sources, column maps, sanitize bounds, option vocab
geo.yaml state/province -> region + coordinates, factory, province aliases
delivery.yaml delivery-estimate normalization (tokens, overrides, month names)
overrides.yaml manual curation: overrides (edit existing rows) + additions (forum-only orders)
data/
raw/ timestamped live caches (committed as dated fetch history)
processed/ cleaned CSV output
output/ dashboard HTML output
tests/
test_parsing.py unit tests (run via pytest OR plain python3)
From the project root:
./r2_dashboard # or: python3 src/pipeline.py
./r2_dashboard --offline # skip the live fetch: newest known cache, no new cache writtenIt's a run-in-place project (no install step). Dependencies are listed in
requirements.txt (pandas, numpy, plotly, PyYAML, beautifulsoup4).
data/processed/r2_orders_clean.csv— the cleaned, tidy dataset.data/processed/r2_dimensions.json—src/conf/dimensions.yamlas JSON: each CSV column's label, category order, blank handling, small-n rule, caveat/note text, and per-category colors and markers.data/processed/r2_series.json— daily counts of what was true by each date: orders, final VINs assigned, final delivery dates set, deliveries, outstanding and converted reservations, each counted on its own event date from today's data, so late reports revise past points. Aggregates only; no per-order history.output/r2_orders_dashboard.html— the interactive dashboard.data/raw/r2_orders_live_*.csv,data/raw/r2_reservations_live_*.csv— timestamped live caches. A new cache is written only when the fetched content differs from the newest known cache, on disk or committed onorigin/main(change detection), so a cache's timestamp marks when the data last changed. If a live fetch fails, or with--offline, the newest known cache is used and nothing is written. Caches from local builds are worth committing too: each one can only add a change the scheduled deploy missed, and a duplicate is harmless. Why the history is kept as separate snapshots and not a single tracked file or a database is recorded indocs/data-layer.md(Snapshot storage) and #89.
Two live Google Sheets, pulled on demand via their CSV export endpoint:
- Orders & Deliveries tracker (one row per person; VIN, config, delivery estimate).
- Reservations tracker (reservation-only holders — no order/VIN/config).
The data is self-reported and noisy; treat all figures as indicative. It is
always pulled live and cached under data/raw/; there are no hand-maintained
snapshots — the raw caches are committed, so data/raw/ is a dated,
change-detected history of the sheets (useful for trend analysis).
Because both sheets are hand-maintained forms, columns are located by name
rather than position, and only the columns actually used are read. So the sheets
can be reordered, or grow new questions anywhere, with no effect. What is checked
on every run is that each column named in src/conf/schema.yaml is present
exactly once: a mapped column that has been renamed, removed, or duplicated
stops the pipeline, since it would otherwise read as empty (or ambiguously) for
every row and quietly skew every figure. Failing means the deployed dashboard
stays on its last good build until schema.yaml is updated to match. A merely
new column can't affect anything, so it's listed in the dashboard's
data-quality panel instead, as a nudge that new data is available.
The dashboard is published at https://emroch.com/r2-dashboard on Cloudflare's free tier, refreshed automatically:
- GitHub Actions (
.github/workflows/deploy.yml) runs the pipeline daily (and on demand / on push), deploys the static output to a Cloudflare Pages project via Wrangler, and commits any refresheddata/raw/caches back so the fetch history accrues. - A small Cloudflare Worker (
worker/) routesemroch.com/r2-dashboard*to that Pages project (the HTML is self-contained, so no asset rewriting is needed)..github/workflows/worker.ymldeploys it whenever anything underworker/lands onmain— validating the bundle on PRs first, and smoke-testing the live route afterwards, since the Worker has no preview environment.
The Python build runs only in Actions — Cloudflare serves and routes but can't run
pandas/plotly. One-time setup (API token, secrets, Pages project) is noted in the
workflow files. The API token needs Workers Scripts · Edit and Workers
Routes · Edit on the emroch.com zone in addition to Pages · Edit, since the
same token deploys both.
python3 tests/test_parsing.py # no pytest required
# or
pytest tests/The dashboard's ⚑ Report issue menu offers three routes, because they lead different places:
- Your own order data is wrong → the forum thread the tracker is compiled
from, which carries the form for updating your entry. This pipeline only
reads those sheets, so fixing the data at the source is what reaches everyone
using it — not just this page. Each sheet's thread is set as
thread_urlinsrc/conf/schema.yamland also linked beside that sheet in the page header. - Something's wrong with this page → the
dashboard-report issue form,
prefilled with the build it was opened from (as-of date, each sheet's
last-updated time, and the deployed commit), so a report is reproducible
without the reader having to describe their build. Lands labelled
from-dashboard. - Message @emroch on the forum → for anyone without a GitHub account, since filing an issue requires signing in.
Planned improvements are tracked as GitHub issues.
MIT © 2026 Eric Roch