Simulation evidence establishes a result within a declared model, numerical regime, parameter domain, and diagnostic. It is not automatically evidence about nature.
The agent can choose an underpowered design, an imperfect observable, or an overstrong interpretation. Provenance makes such mistakes inspectable; it does not prevent every scientifically plausible error.
Hard problems may require a validated guided starting point. This lowers startup cost but constrains the explored model family and must not be confused with evidence for the active hypothesis.
An agent can optimize a declared metric rather than the intended physics. Complementary observables, fresh data, adversarial cases, and independent implementations remain necessary.
Passing conservation and interface checks is not a convergence proof. Particle, grid, timestep, domain, seed, boundary, and model-form sensitivity must be matched to the scope of each claim.
The startup literature search is bounded and cannot establish novelty. A result intended for scientific publication still requires broader literature review and expert comparison.
Version 0.1 demonstrates evidence-governed autonomous computation and real failure discovery. The next step is to apply the same harness to other simulation-gated fields and to hunt independently confirmed new results in the origin domain and beyond.