A comprehensive guide to building discrete event simulation engines for modeling and validating high-level system designs.
This repository provides everything needed to understand and build a Discrete Event Simulation (DES) engine for distributed system architecture. It walks through the theory, data structures, and implementation details required to simulate system behavior — letting you discover bottlenecks, test scaling strategies, and validate designs before writing production code.
hld-simulator-docs/
├── docs/ # Theory, teaching curriculum & system reference
│ ├── README.md # Curriculum index (5-part learning guide)
│ ├── SYSTEM_OVERVIEW.md # How the simulator works end-to-end (UI, CLI, engine)
│ ├── theoretical-foundations.md # Academic theory: queueing, DES, probability, reliability
│ ├── 01-system-diagrams.md # Part 1: Nodes, edges, graph patterns
│ ├── 02-simulation-fundamentals.md # Part 2: Events, time, the event loop
│ ├── 03-data-structures-and-mechanics.md # Part 3: Min-heap, PRNG, distributions, G/G/c/K
│ ├── 04-distributed-systems-and-failures.md # Part 4: Network physics, failure propagation
│ ├── 05-devs-chaos-and-analysis.md # Part 5: DEVS formalism, chaos engineering, output analysis
│ └── question-platform-hardening/ # Implementation guide: the question & grading platform (PRs #212–#214 + follow-ups)
├── schema/
│ ├── complete_simulator_schema.ts # Full TypeScript type definitions (2300+ lines)
│ └── README.md # Schema documentation
├── canonical-catalogue/
│ ├── *.csv # 17 reference catalogue files
│ └── README.md # Catalogue documentation
├── planning/ # Implementation roadmap
│ ├── IMPLEMENTATION_PLAN.md # Phased build plan (10 phases)
│ └── TICKETS.md # 46 engineering tickets
└── design-decisions/ # Architecture decision records
├── adr-internal-modularity-over-plugin-system.md
└── adr-no-custom-change-detection.md
The single source of truth for understanding the simulator end-to-end. Covers how the simulation engine works, the three user phases (BUILD → SIMULATE → ANALYSE), UI representation (screen layout, canvas states, inspector, JSON topology viewer, results tray), CLI commands, component inventory, feature-to-ticket map, and design foundations.
Maps academic theory to simulator features — queueing theory (G/G/c/K, Little's Law), DEVS formalism, probability distributions, reliability theory, graph theory, and control theory.
Covers the building blocks of any system diagram:
- Nodes — source, processing, storage, routing, sink, and composite nodes with their properties (capacity, processing speed, availability)
- Edges — synchronous, asynchronous, streaming, conditional, and weighted connections with failure modes
- Patterns — sequence, fork, join, branch, loop, and parallel composition
- Real-world examples across domains (hospitals, factories, e-commerce, web systems)
Introduces core simulation concepts:
- The three ingredients: model (structure), engine (time progression), observer (measurements)
- Events, state, and the event loop
- Key parameters: arrival rate (lambda), capacity (K, c), service rate (mu), utilization (rho)
- Queues, overflow strategies, and Little's Law
- Randomness, distributions (exponential, log-normal, Poisson), and deterministic replay via seeded PRNGs
Covers implementation in depth with working code:
- Min-heap — O(log n) event queue with full JavaScript implementation
- Precision & determinism — BigInt timestamps, SFC32 PRNG, distribution generators
- G/G/c/K queueing model — formalizing node behavior with Kendall's notation
- Workload generation — constant, Poisson, bursty, diurnal, and spike traffic patterns
- Simulation engine — complete implementation with event handlers, latency percentiles (P50/P90/P95/P99), and Little's Law verification
Models real-world distributed system complexity:
- Distributed systems — dependency graphs, critical vs optional dependencies, fallback behavior
- Network physics — latency decomposition (propagation + transmission + processing + queuing + jitter), congestion modeling
- Failure modes — crash, omission, timing, response, Byzantine; resource exhaustion taxonomy
- Failure propagation — timeout cascades, retry amplification, resource starvation, thundering herd, cache stampede
- Resilience patterns — circuit breaker, bulkhead, retry with backoff, backpressure, rate limiting, load shedding
Formalizes the simulator and adds validation:
- DEVS formalism — atomic and coupled DEVS models, time advance functions, hierarchical composition
- Chaos engineering — structured experiment workflow, fault injection catalog, pre-built experiments
- Output analysis — metrics collection, waterfall traces, heatmaps, causal failure graphs, bottleneck identification
A PR-by-PR, concept-first implementation guide to how the simulator became a
deterministic question and grading platform. It is the learning companion to
specs/rubric-engine-and-question-platform-architecture.md:
the spec is the blueprint; this guide is the field notes on what was built and
why. Each document follows What → How → Why → Trade-offs.
- 01 — Question Platform Foundation —
QuestionPackage,AttemptStatelifecycle, structural & invariant checks, iframe-embed security - 02 — Evaluation Contract & CLI — versioned contract, parser invariants, exit-code taxonomy, Game Playground adapter
- 03 — Rubric Engine Hardening — check kinds & statuses, execution rows, short-circuit skip semantics, hashed test IDs
- 04 — Rebase & Contract Reconciliation — restacking overlapping PRs and regenerating frozen fixtures safely
- 05 — Design Decisions & Trade-offs — a consolidated decision log (D1–D19) with the criteria behind each choice
- 06 — Grading-Safe Persistence & the Evaluation Envelope — the immutable, tamper-evident submission record: integrity checksum, replay digest, append-only archive
- 07 — Production Embed Runtime & Origin Security — hardening the iframe seam: trusted-origin handshake, configured-allowlist-vs-TOFU, no wildcard for sensitive messages
- 08 — The Presentation Layer: EnvironmentProfile — the visibility + capability lens (AUTHOR/ASSIGNMENT/PRACTICE) over one QuestionPackage: presets, safe resolver, applied gates
The canonical-catalogue/ directory contains 17 CSV reference files covering:
- Component taxonomy — ~110+ component types across 13 categories (compute, network, storage, messaging, orchestration, security, observability, DevOps, data infra, real-time, integration, consensus, DNS)
- Component specification — YAML schema template and uniform attributes every component must support
- Simulation primitives — event types, workload profiles (spike, steady-state, diurnal, bursty), and fault injection modes
- Failure modes — cascading failures, backpressure, split-brain, thundering herd, resource exhaustion, and propagation rules
- Patterns & anti-patterns — 23 architectural patterns (CQRS, Saga, Circuit Breaker, etc.) and 8 anti-patterns to detect
- Metrics & SLIs — latency percentiles, throughput, availability, saturation, cost, and recovery time
- Policies & invariants — idempotency, causal ordering, consistency, and security checks
- Pre-built scenarios — 7 deterministic test scenarios (cache stampede, DB failover, network partition, auth outage, traffic spike, and more)
- Provider mapping — AWS / GCP / Azure equivalents for multi-cloud simulation
- Implementation guidance — architecture recommendations, utility components, and a completeness checklist
See canonical-catalogue/README.md for detailed descriptions of each file.
schema/complete_simulator_schema.ts consolidates the full type system for the simulator (2300+ lines), incorporating definitions from both the documentation and the canonical catalogue. It is organized into 17 parts:
- Component types — union types for all ~110+ component types plus a unified
ComponentType - Component specification —
ComponentDefinitionwith identity, resources, lifecycle, dependencies, health checks, telemetry, SLOs, fault injection, scaling, and security - Simulation events —
SimulationEventwith 50+ event types and typedEventDatavariants - Failure propagation —
FailurePropagationwith conditions and cascading effects - Workload profiles — 8 traffic pattern types (steady-state, spike, diurnal, sawtooth, bursty, long-tail, replay, custom)
- Fault injection —
FaultInjectionwith 14 fault types and deterministic/probabilistic/conditional timing - Metrics & outputs —
MetricsDefinition,SimulationOutputwith traces, heatmaps, causal graphs, and reproducibility specs - Scaling & invariants — horizontal/vertical scaling simulation, shard rebalancing, and invariant checks
- Provider configs — cloud-specific latency, quotas, and cost profiles (includes pre-built
AWS_PROFILE) - Utilities —
ScenarioComposer,CostCalculator,ImpactCalculator,ReplayEngine,DesignComparator - Built-in scenarios — 5 pre-configured
BUILT_IN_SCENARIOS(cache stampede, DB crash, network partition, auth outage, traffic spike) - Distribution configs — 12 statistical distributions (normal, log-normal, exponential, Poisson, Weibull, gamma, beta, Pareto, empirical, mixture, etc.)
- Component configs — type-specific configurations for APIs, databases, caches, queues, streams, serverless functions, CDNs, SFUs, and gateways
See schema/README.md for the full breakdown.
The planning/ directory contains the implementation roadmap:
-
Implementation Plan — 10-phase build plan covering topology JSON format, core primitives, simulation engine, network modeling, failure injection, resilience patterns, metrics/output, chaos scenarios, UI integration, and advanced features. Includes dependency graph, file structure, and the critical path for an MVP.
-
Tickets — 46 self-contained engineering tickets with detailed specs, acceptance criteria, dependency chains, and size estimates. Organized into 12 phases covering the core engine, UI components, topology state management, CLI, and more.
The design-decisions/ directory contains architecture decision records (ADRs):
-
Internal Modularity Over Plugin System — Why the engine uses internal module boundaries instead of a runtime plugin system. The core DES loop is domain-agnostic; domain logic (queueing, network, failures) is structured as modules with clean interfaces, but ships as one package.
-
No Custom Change Detection — Why no mutation observer or custom reactivity is needed. BUILD-phase state uses Zustand selector subscriptions. SIMULATE-phase data uses Web Worker
postMessage. Both feed into React's standard re-render cycle. -
Canonical Node Architecture Refactor — Engine-first node modeling plus production-grade semantic naming and domain-first folder structure standards (Component vs Node vocabulary, typed boundaries, workspace persistence contract, and migration path from legacy React Flow saves).