🔥 Learning When to Trust via Selective Context Preference Optimization
benchmark post-training opd on-policy hallucination sft dpo llm llms on-police llm-evaluation llm-agents on-policy-learning on-policy-distillation
-
Updated
Aug 12, 2026 - Python