Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Advanced NLP (M25) - Assignment 2 — GPT-2 Small Solution

This repo fine-tunes GPT-2 Small (gpt2, ~124M params) on AG News for text classification and evaluates:

  • Baseline full-precision (FP32) fine-tuned model
  • INT8 Post-Training Quantization (from-scratch, no bitsandbytes)
  • bitsandbytes: load in 8-bit
  • bitsandbytes: load in 4-bit NF4

It computes Accuracy/Precision/Recall/F1, confusion matrices, memory footprint, and latency, then builds report_1.pdf and report_2.pdf.

0) Environment

python -m venv .venv && source .venv/bin/activate
pip install --upgrade pip
# Choose CUDA or CPU wheels appropriately
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
pip install transformers datasets evaluate scikit-learn matplotlib bitsandbytes reportlab accelerate tqdm

1) Train baseline (FP32)

python -m src.train_baseline --epochs 3 --batch_size 16 --eval_bs 32 --max_length 256

2) Evaluate baseline

python -m src.eval_model --checkpoint checkpoints/gpt2small-agnews --tag baseline

3) From-scratch INT8 PTQ (weights-only, per-channel)

python -m src.ptq_scratch --checkpoint checkpoints/gpt2small-agnews --out checkpoints/gpt2small-agnews-int8
python -m src.eval_model --checkpoint checkpoints/gpt2small-agnews-int8 --tag int8_scratch

4) bitsandbytes 8-bit + 4-bit NF4 (no further training)

python -m src.quant_bnb --checkpoint checkpoints/gpt2small-agnews --mode 8bit
python -m src.quant_bnb --checkpoint checkpoints/gpt2small-agnews --mode 4bit

Each command evaluates and writes metrics & confusion plots:

  • results/<tag>_metrics.json
  • outputs/confusion_<tag>.png

5) Build reports (PDF)

python -m src.generate_reports

Outputs:

  • report_1.pdf (Evaluation metrics answers)
  • report_2.pdf (Methods + auto-inserted tables/figures from results/ and outputs/)

Notes

  • Default model is GPT-2 Small (gpt2) — not GPT-2 Medium/Large.
  • CPU works, but bitsandbytes is best on CUDA; scripts detect device and degrade gracefully.
  • All paths are created automatically under checkpoints/, results/, outputs/.

About

Fine-tuned GPT-2 on AG News for news classification, with reproducible preprocessing, training, evaluation, and FP32 vs INT8 / 4-bit inference comparison.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages