Skip to content

Repository files navigation

Clarification-Guided Reward Learning

CI Python License

A correction tells a robot which state a person preferred. It does not tell the robot why. This simulation tests whether one short follow-up question can separate reward hypotheses that all fit the same corrected state.

System overview of a state correction, clarification question, posterior update, and next action

The overview follows the first two objects in the committed default trace. The yellow cup correction leaves several reward hypotheses compatible; an entropy-gated feature question increases the designated true hypothesis from 0.294 to 0.624. On the next red cup, correction-only selects Q4 and is corrected to Q1, while the clarification condition selects Q1 directly. These are deterministic consequences of one toy configuration, not participant or real-robot results. Vector PDF · figure contract and provenance

Research Status

I started this simulation during the 2024 Carnegie Mellon Robotics Institute Summer Scholars program. The working paper had no user-study results, so the claim here stays narrow: the code implements the proposed inference mechanism and tests it inside a fixed toy hypothesis space.

It does not establish that clarification improves real-robot performance, task completion time, user satisfaction, or generalization.

Research Question

A state correction can be overspecified. If a person moves a red glass cup from one dishwasher quadrant to another, the robot observes the preferred state but not the reason: color, material, object type, or a conjunction of features.

This reference implementation compares two conditions on the same deterministic sequence:

  1. Correction only: update reward-hypothesis beliefs from the corrected state.
  2. Correction + clarification: apply the same correction update, then ask which object features motivated the correction when posterior entropy remains high.

Method

For a human-corrected state S_h, robot state S_r, reward hypothesis theta, and rationality parameter beta, the correction likelihood uses a Bradley-Terry model:

P(S_h > S_r | theta) = exp(beta R_theta(S_h))
                         -----------------------------------------
                         exp(beta R_theta(S_h)) + exp(beta R_theta(S_r))

The posterior is the normalized product of prior and likelihood. A feature answer uses the paper's illustrative noise model: likelihood 0.8 for an exact feature-structure match and 0.2 otherwise. Clarification is gated by normalized entropy rather than asked unconditionally.

The code keeps each assumption explicit in inference.py and simulation.py.

Illustrative Output

With the default six hand-authored hypotheses, three objects, beta=2.0, and clarification likelihood 0.8:

Condition Final posterior on designated true hypothesis Final normalized entropy
Correction only 0.7009 0.5701
Correction + clarification 0.9036 0.2510

These values are a deterministic code-path check, not an empirical result. Change the hypothesis space, feature noise, threshold, objects, or true model and the trace changes.

Correction-only and clarification posterior and entropy trajectories

Reproduce It

git clone https://github.com/ethanvillalovoz/clarification-guided-reward-learning.git
cd clarification-guided-reward-learning

python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"

clarification-reward-demo --output-dir artifacts/latest

The command writes:

  • trace.json: complete priors, likelihood-driven posteriors, actions, answers, and entropy.
  • comparison.png: correction-only versus clarification trajectories.
  • reasoning-snapshot.png: the first correction, clarification answer, and posterior update in one figure.
  • belief-update.webp: stage-by-stage belief animation.
  • system-overview.{svg,pdf,png}: wide method overview and next-object consequence.

A committed reference trace makes the default configuration inspectable without running Python.

Verification

ruff check src tests scripts
pytest -q
MPLBACKEND=Agg clarification-reward-demo --output-dir artifacts/check

Tests cover normalization, numerical stability, Bradley-Terry directionality, correction updates, feature clarification, entropy reduction, deterministic replay, and invalid configurations.

Package Layout

src/clarification_reward_learning/
  models.py          objects, reward hypotheses, and default toy problem
  inference.py       likelihoods, Bayesian updates, entropy, and action selection
  simulation.py      correction-only and clarification experiment traces
  visualization.py   static, animated, and vector system-overview artifacts
  cli.py             reproducible command-line entry point
tests/               focused inference and simulation regression tests
examples/            committed default trace
docs/                RISS paper, poster, slides, video, and method notes
scripts/             committed public-figure regeneration

Research Artifacts

The original presentation-oriented implementation and prototype history remain available in the v1.0.0 tag. Version 1.1 replaces that path with a smaller, testable reference implementation; it does not rewrite the historical paper.

Limitations

  • Fixed discrete hypothesis space and three synthetic objects.
  • Hand-authored rewards, clarification accuracy, and entropy threshold.
  • No participant study, real robot, language understanding, or learned question policy.
  • No robustness analysis for misspecified or incomplete hypothesis spaces.
  • The expected-reward action policy is deterministic and omits dynamics beyond quadrant placement.
  • The committed comparison is descriptive for one configuration and has no statistical uncertainty.

These limitations are the next research work, not hidden implementation details.

Acknowledgments

Developed during CMU RISS 2024 with Michelle Zhao and mentorship from Dr. Henny Admoni and Dr. Reid Simmons. Thanks to Rachel Burcin and Dr. John Dolan for leading the RISS program.

License

Released under the MIT License.

About

Tested simulation of clarification-guided reward learning from human state corrections, developed during CMU RISS.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages