The scaled experiments adapt ExPLAIND to LLM checkpoints by explaining one optimizer update at a time with StepExplainer. The EuroLLM scripts are in euro_llm/ and assume access to converted Hugging Face checkpoints, original optimizer checkpoints, recovered training batches, and evaluation samples.
Run commands from the repository root.
Install the base package dependencies and the optional scaled-experiment dependencies before running these scripts:
pip install -r requirements.txt
pip install -r requirements-scaled.txtThe scaled requirements cover the Hugging Face, Accelerate, SafeTensors, TensorDict, and nnsight packages used by the EuroLLM scripts and accelerated StepExplainer implementations.
python scaled_experiments/euro_llm/sample_lang_batches.pyThis samples language-balanced JSONL batches into results/lang_batches/. The script streams from Hugging Face datasets and may take a long time depending on network and dataset cache state.
python scaled_experiments/euro_llm/verify_blimp_scores.py \
--conv_checkpoint_dir /path/to/converted/eurollm/1b \
--checkpt_dir /path/to/original/megatron/checkpoints \
--score_dir results/euro_llm_scores/verify_blimp \
--blimp_path results/blimp_scores/blimp_samples.json \
--model_config_path scaled_experiments/llama1b/training/configs/models/llama1b.json \
--optimizer_config_path scaled_experiments/llama1b/training/configs/optimizer/adamw.json \
--split_langs_trainUse compute_hypothetical_blimp_scores.py for the hypothetical-BLiMP variant and compute_scores.py if you want to call the lower-level compute_scores(...) function directly from another script.
python scaled_experiments/euro_llm/check_verification.pyUpdate the checkpoint root in that script or adapt it into a one-off analysis script for your score directory.
Use explaind.accelerated.explainer.StepExplainer when full training-history tracking is too expensive and you can reconstruct a single training update. The sample Explainer class at the bottom of explaind/accelerated/explainer.py documents the required subclass interface.
To integrate a custom model:
- Subclass
StepExplainer. - In
load_checkpoint, load the pre-update model, optimizer, scheduler if used, training batch or grouped training instances, test instances, learning rate, optimizer-state mapping, and update-step id. - In
compute_step_model, reproduce the exact optimizer update being explained, keep a copy of the pre-update model and optimizer state, then call_wrap_model("", model_before, model_after). - Set
gradient_scaling,train_grad_scaling,test_grad_scaling, andtest_loss_scalingto match the training recipe, especially if gradient accumulation or clipping was used. - Provide
train_instancesandtest_instancesas dictionaries mapping readable type names to tensor batches. Those type names become the score keys. - Use
expl_type="loss"for loss decomposition. Output decomposition is intentionally not implemented in the current accelerated path. - Call
compute_influence_scores(),save_scores_to_disk(), and optionallyverify_loss_decomposition().
The LLaMA and EuroLLM implementations in explaind/accelerated/llama_explainer.py and explaind/accelerated/eurollm_explainer.py are concrete references for optimizer restoration, recovered-batch loading, language splitting, and verification.
These files or directories are referenced by the scaled scripts but are not expected to be committed to this repository:
- converted model checkpoints, e.g.
/path/to/converted/eurollm/1b/...; - original Megatron optimizer checkpoints, e.g.
mp_rank_00/model_optim_rng.pt; - recovered training batches such as
results/lang_batches/eurollm_phase1_train_batch.jsonl; - BLiMP sample files under
results/blimp_scores/; - model and optimizer config files under
scaled_experiments/llama1b/training/configs/...if the external training code is not included.
The required Python packages for these scripts are listed in requirements-scaled.txt.