An evidence-first research platform for deterministic, reproducible Windows driver reliability experiments.
KCrashLab studies the parts of low-level failure research that are easy to overlook: deterministic input generation, recovery after interruption, exact failure identity, trigger minimization, cold replay, provenance, and independently verifiable evidence.
The default workflow runs entirely in user mode against a synthetic state machine. A separately gated Windows-lab track exists for the repository-owned target, but it is never selected by ordinary campaign commands.
Important
All evidence currently committed to this repository is simulated. The presence of Windows-lab source code does not prove that a driver was loaded, a system failure occurred, a checkpoint was restored, or a real vulnerability was found. KCrashLab does not target third-party drivers and does not generate exploits.
| Area | Status | What the status means |
|---|---|---|
| Deterministic simulation | ✅ | Implemented and exercised by repository tests and recorded synthetic artifacts |
| Resumable campaign controller | ✅ | Append-only journal, recovery, terminal-state handling, and idempotent evidence production |
| Case mutation and corpus scheduling | ✅ | Deterministic policies, stateful sequences, semantic feedback, and explicit scheduler telemetry |
| Triage, replay, and minimization | ✅ | Exact signatures, replay voting, hierarchical reduction, and lineage-preserving cases |
| Evidence verification | ✅ | Content hashes plus semantic consistency checks across summaries, cases, reports, and raw tables |
| Controlled Windows-lab source | 🧪 | Gated implementation is present; no runtime result is committed |
| Third-party target support | — | Intentionally out of scope |
✅ means implemented in the public simulation path. 🧪 means source is available for review but runtime evidence is absent. — means intentionally unsupported.
flowchart LR
caseIr["1 · Case IR<br/>canonical identity"]
simulate["2 · Deterministic simulation<br/>scheduler + semantic feedback"]
finding["3 · Finding closure<br/>exact signature + minimize + replay"]
evidence["4 · Evidence closure<br/>provenance + manifest + verifier"]
caseIr --> simulate --> finding --> evidence
windowsLab["Separate Track B<br/>explicitly gated Windows lab"]
caseIr -. manual gate · never automatic .-> windowsLab
classDef input fill:#e8f5e9,stroke:#2e7d32,color:#102a13;
classDef process fill:#e3f2fd,stroke:#1565c0,color:#0d2740;
classDef result fill:#fff3e0,stroke:#ef6c00,color:#3d2200;
classDef evidence fill:#ede7f6,stroke:#6a1b9a,color:#25102f;
classDef gated fill:#f5f5f5,stroke:#616161,color:#212121,stroke-dasharray:5 5;
class caseIr input;
class simulate process;
class finding result;
class evidence evidence;
class windowsLab gated;
Solid arrows show the ordinary user-mode workflow. The dashed branch is separately gated and is never selected automatically. See the recorded G3 discovery, M1 minimization and 3/3 replay, and the research claims ledger for the evidence behind this flow.
Finding a failure is only the beginning. A credible research workflow must also answer:
- Can the exact input be identified independently of file formatting?
- Can an interrupted campaign resume without inventing or losing state?
- Can the same failure be reproduced from a clean baseline?
- Can the trigger be reduced while preserving its exact signature?
- Can another reviewer verify the result without trusting the producer?
- Are simulated observations clearly separated from real-machine observations?
KCrashLab turns those questions into explicit contracts, tests, artifacts, and failure-closed gates.
The controller communicates through ILabBackend. The simulation backend is the only backend available to ordinary campaigns. The Windows-lab path has a separate entry point, private profile, explicit confirmation boundary, fixed repository target, and clean-checkpoint replay policy.
Inputs are normalized into a versioned JSON representation. Semantic identity is a SHA-256 digest of canonical content, while lineage metadata remains available for audit. Equivalent inputs therefore deduplicate even when they were reached through different mutation paths.
Parent, operator, and candidate decisions use independent seed-derived decision lanes. Candidate caps use HASH_RANKED_V1: the complete valid candidate set is ranked from the campaign seed, operator identifier, and semantic case identifier before truncation. This avoids favoring whichever fields happen to be enumerated first.
Campaign summaries expose the termination reason, scheduling iterations and limit, duplicate-candidate skips, empty polls, candidate-selection rule, and per-operator cap. A scheduler-limited run cannot be silently reported as a fully consumed execution budget or proof of global search-space exhaustion.
Findings are grouped by a versioned signature derived from normalized triage fields rather than by a generic failure label. Replay uses explicit voting, and minimization is accepted only while the target signature remains unchanged.
Campaign transitions are recorded in an append-only SQLite journal. Restarting the controller reconstructs state and resumes the unfinished stage. Evidence publication is idempotent, so recovery does not create multiple authoritative bundles for one campaign.
Every bundle includes a SHA-256 manifest, but verification does not stop there. KCrashLab also checks cross-file identities, summary bounds, trial counts, paired seeds, claims, canonical cases, CSV rows, reports, and provenance fields. A self-consistent hash over inconsistent research data is rejected.
- Windows, Linux, or macOS for the simulation path
- .NET SDK
8.0.100exactly, as pinned byglobal.json - No administrator rights, WDK, Hyper-V, or VM
dotnet --version
dotnet restore KCrashLab.sln
dotnet build KCrashLab.sln --configuration Release --no-restore
dotnet test KCrashLab.sln --configuration Release --no-builddotnet run --project src/KCrashLab.Cli --configuration Release -- `
campaign run `
--scenario dump-ready `
--case samples/cases/state-original.case.json `
--output artifacts/demo
dotnet run --project src/KCrashLab.Cli --configuration Release -- `
evidence verify artifacts/demodotnet run --project src/KCrashLab.Cli --configuration Release -- `
fuzz run `
--seed samples/cases/state-safe-seed.case.json `
--budget 256 `
--campaign-seed 20260831 `
--output artifacts/fuzz
dotnet run --project src/KCrashLab.Cli --configuration Release -- `
fuzz verify artifacts/fuzzdotnet run --project src/KCrashLab.Cli --configuration Release -- `
experiment e1 `
--seed samples/cases/state-safe-seed.case.json `
--budget 256 `
--trials 20 `
--base-seed 20260831 `
--output artifacts/e1
dotnet run --project src/KCrashLab.Cli --configuration Release -- `
experiment e2 `
--seed samples/cases/state-reset-seed.case.json `
--budget 512 `
--trials 20 `
--base-seed 20260831 `
--output artifacts/e2
dotnet run --project src/KCrashLab.Cli --configuration Release -- `
experiment verify artifacts/e1
dotnet run --project src/KCrashLab.Cli --configuration Release -- `
experiment verify artifacts/e2E1 is a paired 2×2 ablation of corpus admission and parent selection. E2 changes only the maximum sequence length, making it an expressiveness check for a known multi-step synthetic condition—not a general performance claim. See the E1 and E2 experiment records for controls and interpretation limits.
Reviewer-facing synthetic artifacts are stored under:
results/recorded/g3— deterministic discovery;results/recorded/e1— policy ablation;results/recorded/e2— stateful versus single-call experiment;results/recorded/minimization-replay— exact-signature minimization and 3/3 simulated replay evidence (record).
Local artifacts/* directories are disposable outputs and are not authoritative. A recorded result is accepted only after generation from a clean commit, semantic verification, and an independent deterministic rerun with a matching manifest. Provenance records the source revision, GIT_TREE_SHA256_V1 digest algorithm and digest, Case IR version, engine version, experiment-definition digest, and timestamp policy.
For experiment commands, --recorded-at SOURCE_COMMIT_TIME resolves the full checked-out HEAD when --git-commit is omitted. Canonical provenance fails closed unless the working tree is clean and the requested commit equals HEAD; source identity is calculated from the pinned commit's Git tree and blob objects rather than platform-dependent checkout bytes.
The v0.2 evidence freeze uses two Git layers. Source commit 55c31b25902157cddc9014bb9fdaa598bede40a2 contains the engine, provenance contracts, and tests used to generate the records. The child revision containing results/recorded/* is the evidence/release layer, and the release tag should point to that revision. Exact artifact reproduction therefore requires checking out the source commit and writing the rerun below ignored artifacts/; running the same command from the evidence commit intentionally records a different HEAD and cannot produce the checked-in manifest.
Historical artifacts do not retroactively inherit stronger guarantees from newer code. Read docs/status.md before citing a result.
The repository also contains a separately gated implementation for the project-owned test target:
- a safe-default KMDF build and a distinct opt-in lab build;
- a fixed device interface and allowlisted request contract;
- a Windows guest dispatcher with a durable attempt journal;
- a versioned private profile and fail-closed validation;
- exact VM/checkpoint/private-network checks;
- one discovery attempt followed by three clean-checkpoint replays;
- dump-readiness checks, WinDbg parsing, exact signatures, and sanitized public evidence construction.
This track is not run by normal CI and is not runtime-verified by the checked-in repository. Its entry point requires an explicitly prepared, disposable, isolated owner-controlled environment. The full prerequisites and publication boundary are documented in the controlled-lab runbook and publication checklist.
| Path | Responsibility |
|---|---|
src/KCrashLab.Contracts |
Versioned backend, case, experiment, and evidence contracts |
src/KCrashLab.Domain |
Canonicalization, scheduling, mutation, signatures, replay, and minimization |
src/KCrashLab.Storage |
SQLite journal and content-addressed storage |
src/KCrashLab.Simulation |
Virtual clock, scripted backend, synthetic target, and environment gates |
src/KCrashLab.Controller |
Resumable orchestration and artifact production |
src/KCrashLab.GuestAgent |
Windows-only dispatcher for the fixed project-owned contract |
drivers/KCrashLab.Target |
Repository-owned KMDF target source |
schemas |
JSON Schema 2020-12 contracts |
tests |
Unit, component, contract, recovery, and golden-fixture tests |
docs |
Architecture, methods, safety boundaries, claims, and experiment records |
SIMULATEDandREAL_LABare separate evidence classes and may not be conflated.- No checked-in artifact currently establishes a real-machine result.
- Synthetic discovery latency is not a benchmark of real target performance.
- E2 is a controlled sanity experiment, not a universal claim about stateful methods.
- Manifests establish integrity after production; they do not prove that the producing machine was honest.
- Raw memory captures and unreviewed diagnostic output are private by default.
- Third-party targeting, exploitability assessment, and exploit generation are outside project scope.
The mapping from every public claim to its supporting artifact is maintained in the research claims and evidence ledger.
- Architecture
- Implementation status
- Threat model
- Lab safety policy
- Reviewer packaging guide
- Security policy
- Citation metadata
Licensed under the Apache License 2.0.