Skip to content

Latest commit

 

History

History
41 lines (35 loc) · 2.17 KB

File metadata and controls

41 lines (35 loc) · 2.17 KB

Contributing

Workflow

  1. Branch from main: git checkout -b feature/your-change
  2. Keep changes scoped to one concern (one feature, one bugfix, one refactor per PR)
  3. Run the full test suite before opening a PR: pytest tests/ -v
  4. Run the full pipeline once locally to confirm nothing broke: python main.py
  5. Open a PR against main with a clear description of what changed and why

Code standards

  • Follow the existing module boundaries: ingestion only touches data/raw/, features only read from data/processed/cleaned_panel.csv, training only reads the feature panel, etc. Don't reach across layers.
  • Every public function needs a docstring explaining why it exists, not just what it does - see src/features/feature_engineering.py for the expected level of detail.
  • No placeholder code (# TODO, pass # implement later) in anything merged to main.
  • New features go through config/config.yaml, not hardcoded constants, wherever a value might reasonably change between environments or experiments.

Adding a new model

  1. Add its hyperparameters under models: in config/config.yaml
  2. Register it in build_models() in src/training/train_model.py
  3. If it needs tuning, add a search space + estimator factory in src/training/hyperparameter_tuning.py
  4. Add a unit test in tests/test_training.py confirming it fits on toy data

Adding a new feature family

  1. Add a build_features step in src/features/feature_engineering.py
  2. Document why the feature should help the model in the function docstring
  3. Add a unit test in tests/test_features.py confirming correctness (no leakage across countries, no unexpected NaNs, expected columns present)

Data changes

If the real datasets (country_metadata.csv, country_year_indicators.csv, economic_stress_score.csv, indicator_dictionary.csv) change shape, update EXPECTED_SCHEMA in src/validation/schema_validation.py first - it is the single source of truth for what "valid input" means, and the rest of the pipeline assumes it's accurate.

Reporting issues

Open a GitHub issue with: what you ran, what you expected, what happened instead, and the relevant lines from logs/pipeline.log.