- Branch from
main:git checkout -b feature/your-change - Keep changes scoped to one concern (one feature, one bugfix, one refactor per PR)
- Run the full test suite before opening a PR:
pytest tests/ -v - Run the full pipeline once locally to confirm nothing broke:
python main.py - Open a PR against
mainwith a clear description of what changed and why
- Follow the existing module boundaries: ingestion only touches
data/raw/, features only read fromdata/processed/cleaned_panel.csv, training only reads the feature panel, etc. Don't reach across layers. - Every public function needs a docstring explaining why it exists, not just what it does -
see
src/features/feature_engineering.pyfor the expected level of detail. - No placeholder code (
# TODO,pass # implement later) in anything merged tomain. - New features go through
config/config.yaml, not hardcoded constants, wherever a value might reasonably change between environments or experiments.
- Add its hyperparameters under
models:inconfig/config.yaml - Register it in
build_models()insrc/training/train_model.py - If it needs tuning, add a search space + estimator factory in
src/training/hyperparameter_tuning.py - Add a unit test in
tests/test_training.pyconfirming it fits on toy data
- Add a
build_featuresstep insrc/features/feature_engineering.py - Document why the feature should help the model in the function docstring
- Add a unit test in
tests/test_features.pyconfirming correctness (no leakage across countries, no unexpected NaNs, expected columns present)
If the real datasets (country_metadata.csv, country_year_indicators.csv,
economic_stress_score.csv, indicator_dictionary.csv) change shape, update
EXPECTED_SCHEMA in src/validation/schema_validation.py first - it is the single source
of truth for what "valid input" means, and the rest of the pipeline assumes it's accurate.
Open a GitHub issue with: what you ran, what you expected, what happened instead, and the
relevant lines from logs/pipeline.log.