This repo fine-tunes GPT-2 Small (gpt2, ~124M params) on AG News for text classification and evaluates:
- Baseline full-precision (FP32) fine-tuned model
- INT8 Post-Training Quantization (from-scratch, no bitsandbytes)
- bitsandbytes: load in 8-bit
- bitsandbytes: load in 4-bit NF4
It computes Accuracy/Precision/Recall/F1, confusion matrices, memory footprint, and latency, then builds report_1.pdf and report_2.pdf.
python -m venv .venv && source .venv/bin/activate
pip install --upgrade pip
# Choose CUDA or CPU wheels appropriately
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
pip install transformers datasets evaluate scikit-learn matplotlib bitsandbytes reportlab accelerate tqdmpython -m src.train_baseline --epochs 3 --batch_size 16 --eval_bs 32 --max_length 256python -m src.eval_model --checkpoint checkpoints/gpt2small-agnews --tag baselinepython -m src.ptq_scratch --checkpoint checkpoints/gpt2small-agnews --out checkpoints/gpt2small-agnews-int8
python -m src.eval_model --checkpoint checkpoints/gpt2small-agnews-int8 --tag int8_scratchpython -m src.quant_bnb --checkpoint checkpoints/gpt2small-agnews --mode 8bit
python -m src.quant_bnb --checkpoint checkpoints/gpt2small-agnews --mode 4bitEach command evaluates and writes metrics & confusion plots:
results/<tag>_metrics.jsonoutputs/confusion_<tag>.png
python -m src.generate_reportsOutputs:
report_1.pdf(Evaluation metrics answers)report_2.pdf(Methods + auto-inserted tables/figures fromresults/andoutputs/)
- Default model is GPT-2 Small (
gpt2) — not GPT-2 Medium/Large. - CPU works, but bitsandbytes is best on CUDA; scripts detect device and degrade gracefully.
- All paths are created automatically under
checkpoints/,results/,outputs/.