Safety and usefulness of tool-using agent policies under synthetic tool, observation, memory, and constraint shift.
It does not measure real production safety, real trading performance, real calendar reliability, exploit capability, or private-data handling.
Research comparisons of agent policies, monitors, and fault mitigations.
No real APIs, no operational side effects, no dangerous procedures, no financial advice.
All data and environments are synthetic. Faults never mutate hidden ground truth. Agents and monitors receive a visible-only ObservationContext; hidden ground truth and evaluator execution methods are used only after the policy decision for scoring and replay.
Simple policies, compact environments, and synthetic assumptions limit external validity.
Use deterministic seeds. Run bash scripts/run_repro.sh.
Agents and monitors should not be able to access hidden evaluator state. The current implementation enforces this by passing ObservationContext instead of the full environment object, by redacting evaluator-only fault labels before policy decisions, and by testing that hidden methods such as hidden_ground_truth, tool_response, and execute are absent from that context.