A unified toolkit for researchers and engineers working on AI physical reasoning. PRKit provides a shared foundation for representing physics problems, running inference with multiple model providers, evaluating outputs with physics-aware comparators, and building structured annotation workflows.
PRKit applies a βunified interfaceβ idea to the full physical-reasoning loop (data β annotation β inference β evaluation), rather than focusing on datasets alone.
PRKit centers on core components that define the physical reasoning ontology. Three integrated subpackages build on this foundation:
- Core components:
PhysicsDomain,PhysicsProblem,PhysicsAnswer,PhysicsDataset,PhysicsSolution,BaseModelClient,create_model_client,PRKitLoggerβthe shared abstractions used across the toolkit. prkit.datasets: A Datasets-like hub that downloads/loads benchmarks into the unified schema (PhysicsProblem,PhysicsDataset).prkit.annotation: Workflow-oriented tools for structured, lower-level labels (e.g., domain/subdomain, theorem usage).prkit.evaluation: Evaluate-like components for physics-oriented scoring and comparison (e.g., symbolic/numerical answer matching).
from prkit.datasets import DatasetHub
from prkit.core.model_clients import create_model_client
# Load any benchmark into the unified schema (PhysicsProblem, PhysicsDataset)
dataset = DatasetHub.load("physreason", variant="full", split="test")
# Run inference with the unified model client (core component)
client = create_model_client("gpt-4.1-mini")
for problem in dataset[:3]:
print(client.solve_physics_problem(problem)[:200])The same pattern works across different datasets and model providersβswap the dataset name or model identifier.
For the standalone "is this physics answer right?" use case, use the light-import
verifierβa math-verify-shaped API that, unlike math-verify, is unit- and
symbolic-aware and imports no model clients, dataset hub, or provider SDKs:
from prkit.verify import verify
v = verify("9.8 m/s^2", "9.8 m/sΒ²") # verify(gold, pred) -> Verdict
v.correct # True β the unit suffix normalizes (math-verify strips units)
v.units_ok # True
v.symbolic_equiv # None (numeric case); True for e.g. verify("v = a t", "v = t a")
v.scorer_version # stamped so a stored score is attributable to its scorerUnderneath verify is the physics-semantics layer. It models a question's contract
q and an answer's typed semantics a, and judges equivalence as a question-conditioned
relation Eq(a_pred, a_ref ; q) β deterministically, not by string match. It exposes three
build actions and two judge entry points, all importable from prkit.semantics:
extract_prediction_answer_semantics(answer_text)β deterministically type a prediction;create_reference_semantics(problem, model_client=None)β build a reference(q_ref, a_ref)(deterministic whenmodel_clientis omitted, LLM-assisted otherwise);generate_prediction_semantics(problem, solver_client, ...)β solve, then type the answer;compare_protocol_answers(pred, ref, ...)β reference-based judgement;compare_predictions(a_i, a_j, ...)β reference-free (symmetric) judgement.
See PHYSICS_SEMANTICS.md for the full story and doc map.
Quick Links:
- π§ CORE.md - Core components: domain model, model client, logger, and definitions
- π DATASETS.md - Complete guide to supported datasets and benchmarks
- π§ͺ PHYSICS_SEMANTICS.md - Physics-semantics layer:
q/a, the five build/judge steps, and the doc map - π EVALUATION.md - The deterministic physics-semantics scorer (
verify/SemanticsScorerβVerdict) - π·οΈ ANNOTATION.md - Human annotation tasks (gold, correctness)
- π CHANGELOG.md - Version history and release notes
- Python 3.10+ (required)
# Install the latest stable version
pip install physical-reasoning-toolkit
# Verify installation
python -c "import prkit; print(prkit.__version__)"Step 1: Clone the Repository
git clone https://github.com/sherryzyh/physical_reasoning_toolkit.git
cd physical_reasoning_toolkitStep 2: Install
# Install the package (regular install for end users)
pip install .
# Verify installation
python -c "import prkit; print('β
Toolkit installed successfully!')"Option 1: Export as environmental variable
# For model provider integration (optional)
export OPENAI_API_KEY="your-openai-api-key"
export GEMINI_API_KEY="your-gemini-api-key"
export DEEPSEEK_API_KEY="your-deepseek-api-key"
export XAI_API_KEY="your-xai-api-key"
export DASHSCOPE_API_KEY="your-dashscope-api-key"
# For logging configuration (optional)
export PRKIT_LOG_LEVEL=INFO
export PRKIT_LOG_FILE=/var/log/prkit.log # Optional: defaults to {cwd}/prkit_logs/prkit.log if not setOption 2: Create a .env file at your project root
π See CORE.md (Model Client section) for supported providers and usage.
python -c "
import prkit
from prkit.datasets import DatasetHub
from prkit.annotation import AnnotationOrchestrator
print('β
All packages imported successfully!')
print(f'PRKit version: {prkit.__version__}')
"Installing the package provides a prkit console command for dataset workflows:
prkit --version # Print the installed version
prkit list # List available datasets
prkit info ugphysics # Show dataset metadata (JSON)
prkit download ugphysics # Download a dataset into the cache dir
prkit download seephys --split test # Download a specific split
prkit download phyx --data-dir ./data # Download into a custom directory
# Annotation tasks
prkit annotate gold domain seephys # Expert labels gold domains (terminal)
prkit annotate correctness PATH/TO/seephys_gpt-5 # Judge model answers vs gold (Streamlit UI)The dataset commands are thin wrappers over DatasetHub, so the cache directory,
variants, and splits behave exactly as they do in the Python API. prkit annotate
routes to a human annotation task via the AnnotationOrchestrator.
physical_reasoning_toolkit/
βββ src/prkit/ # Main package (modern src-layout)
β βββ core/ # Core components (domain models, model clients, logging)
β βββ datasets/ # Dataset loading and management
β βββ annotation/ # Human annotation tasks (gold, correctness)
β βββ evaluation/ # Evaluation metrics and benchmarks
β βββ semantics/ # Physics-aware answer normalization and comparison
βββ docs/ # User guides and reference documentation
βββ tests/ # Unit tests
βββ pyproject.toml # Package configuration
βββ LICENSE # MIT License
βββ README.md # This file
Note: The actual dataset files are stored externally (see Environment Setup section). This repository contains only the toolkit code, examples, and documentation.
In Repository (Code & Documentation):
- β src/prkit/: Complete toolkit with core components and 3 subpackages
- β tests/: Unit tests (for contributors)
External (Data & Runtime):
- π Data Directory: Dataset files (set via
DATASET_CACHE_DIR) - π API Keys: Model provider credentials (if applicable)
- π Log Files: Runtime logs (default:
{cwd}/prkit_logs/prkit.log, can be overridden viaPRKIT_LOG_FILE)
The toolkit is organized around core components and three subpackages that use them. Subpackages depend only on prkit.core; there are no direct dependencies between prkit.datasets, prkit.annotation, and prkit.evaluation.
| Component | Purpose |
|---|---|
prkit.core |
Core components, see below |
prkit.datasets |
Dataset hub: loaders, downloaders, unified schema |
prkit.evaluation |
Comparators and accuracy metrics |
prkit.annotation |
Workflow pipelines for domain/theorem annotation |
The essential building blocks of the physical-reasoning-toolkit. All datasets, inference, evaluation, and annotation workflows use these components.
- PhysicsDomain β Enumeration of physics subfields (mechanics, thermodynamics, quantum mechanics, optics, etc.) for problem classification. Aligned with UGPhysics, PHYBench, TPBench. Use
PhysicsDomain.from_string()for flexible parsing. - PhysicsProblem β The canonical representation of a physics problem. Required:
problem_id,question. Optional:answer(PhysicsAnswer),solution,domain,image_path,problem_type(MC/OE),options,correct_option. Supports dictionary-like access andload_images()for visual problems. - PhysicsAnswer β Thin observation record:
value(str, verbatim), optionalunit(observed unit string), optionalsource_type(dataset-native type tag, verbatim), andmetadatadict. The canonical answer kind (AnswerObjectKind, 9 object kinds) is derived on demand by theprkit.semanticslayer β it is not stored onPhysicsAnswer. - PhysicsDataset β Collection of
PhysicsProbleminstances. Indexing, slicing,get_by_id(),filter_by_domain(),take(),sample(),save_to_json()/from_json(). Providesget_statistics()for domain and problem-type distribution. - PhysicsSolution β Bundles a
PhysicsProblem, modelagent_answer, and optionalintermediate_steps. Captures the full solution trace for evaluation and analysis. - BaseModelClient β Abstract base for model clients. Subclasses implement
chat(user_prompt, image_paths=None). - PRKitLogger β Centralized logging with colored output, file logging, and env config (
PRKIT_LOG_LEVEL,PRKIT_LOG_FILE, etc.).
π See CORE.md for the full domain model, entity relationships, subpackage dependency diagram, and import reference.
The deterministic physics-semantics scorer: prkit.verify.verify (light-import, one-call)
and the prkit.scoring family β SemanticsScorer (binary), the EedScorer/SeedScorer
edit-distance baselines, the graded SemanticsEedScorer/SemanticsSeedScorer, and the
model-graded LLMJudgeScorer β all returning the canonical Verdict. Wraps the
prkit.semantics.comparison engine. (The legacy prkit.evaluation comparator/evaluator
stack is deprecated; prkit.evaluation.llm_judge stays.)
π EVALUATION.md Β· PHYSICS_SEMANTICS.md
Dataset hub with a Datasets-like interface: DatasetHub.load() for PHYBench, PhysReason, UGPhysics, SeePhys, PhyX (plus JEEBench, TPBench loaders). Auto-download, variant selection, and reproducible sampling.
π DATASETS.md
Human-in-the-loop annotation tasks dispatched by one AnnotationOrchestrator: gold (an expert authors the gold label for an attribute such as domain) and correctness (a human judges model answers against the gold reference via a Streamlit UI). Run with prkit annotate <task> ....
π ANNOTATION.md
# Check Python version
python --version # Should be 3.10+
# If using wrong version
python -m venv venv
source venv/bin/activate# Reinstall in development mode
pip install -e .
# Check installation
pip show physical-reasoning-toolkit# Set data directory (external to repository)
export DATASET_CACHE_DIR=/path/to/your/data
# Check directory structure
ls -la $DATASET_CACHE_DIR
# Verify dataset files exist
ls -la $DATASET_CACHE_DIR/ugphysics/
ls -la $DATASET_CACHE_DIR/PhysReason/- Review logs: Check logging output for detailed error information
- Verify setup: Run the testing commands above
- Check data: Ensure datasets are properly downloaded and accessible
- Check documentation: Start with the root docs linked below
- GitHub Issues: Report bugs or request features
- Discussions: Share ideas and get help
# Clone and install in development mode
git clone https://github.com/sherryzyh/physical_reasoning_toolkit.git
cd physical_reasoning_toolkit
pip install -e ".[dev]"
# Run code quality tools
black src/
isort src/
mypy src/
# Run tests
pytest tests/- Follow existing patterns: Use consistent logging and error handling
- Add tests: Include tests for new functionality
- Update documentation: Add examples and update README files
- Maintain compatibility: Ensure changes don't break existing functionality
- Fork the repository
- Create a feature branch
- Make your changes with tests
- Ensure all tests pass
- Submit a pull request with clear description
If you use PRKit in your research, please cite it as follows:
BibTeX:
@software{zhang2026physicalreasoningtoolkit,
author = {Zhang, Yinghuan},
title = {Physical Reasoning Toolkit},
year = {2026},
license = {MIT},
url = {https://github.com/sherryzyh/physical_reasoning_toolkit},
abstract = {A unified toolkit for researchers and engineers working on AI physical reasoning. PRKit provides a shared foundation for representing physics problems, running inference with multiple model providers, evaluating outputs with physics-aware comparators, and building structured annotation workflows.}
}For citation files, see CITATION.cff and CITATION.bib in the repository root.
PRKit integrates and builds upon several excellent physics reasoning benchmarks and datasets. We thank the creators of:
- PhysReason, PHYBench, UGPhysics, SeePhys, PhyX, and other benchmark datasets
- The open-source community for their valuable contributions and feedback
Note: For detailed citations and references to the original dataset papers, please see the Citations section in DATASETS.md.
This project is licensed under the MIT License - see the LICENSE file for details.
Ready to advance physics reasoning research! πβ¨
Quick Links: pip install physical-reasoning-toolkit | GitHub | Documentation | Issues