Bakes a system prompt into a LoRA adapter, so a chat model behaves as if prompted without the prompt in its context. It trains on the KL divergence to the prompted model, after Prompt Baking (arXiv 2409.13697). PyTorch, Transformers and PEFT, Qwen3.5 by default, and it fits an 8 GB GPU.
nlp machine-learning deep-learning transformers pytorch lora knowledge-distillation kl-divergence fine-tuning peft huggingface large-language-models llm prompt-engineering qwen system-prompt context-distillation prompt-baking
-
Updated
Oct 1, 2026 - Python