Custom PyTorch data pipeline and tokenization engine built from scratch for pre-training an Urdu Large Language Model.