Hi,
Thanks for maintaining this repo and survey!
We have recently released two technical reports on rubric-based post-training:
https://arxiv.org/abs/2604.13029
The first part focuses on DPO training, utilizing instance-specific rubrics for offline data filtering and on-policy preference pair construction.
https://arxiv.org/abs/2605.30244
The second part focuses on GRPO training, introducing a dual-path mechanism for criterion-level scoring, which effectively combines deterministic verifiers and fuzzy judges (LLM-as-a-Judge).
We believe these approaches are highly relevant and hope you might consider including them in your survey.
Thank you!
Hi,
Thanks for maintaining this repo and survey!
We have recently released two technical reports on rubric-based post-training:
https://arxiv.org/abs/2604.13029
The first part focuses on DPO training, utilizing instance-specific rubrics for offline data filtering and on-policy preference pair construction.
https://arxiv.org/abs/2605.30244
The second part focuses on GRPO training, introducing a dual-path mechanism for criterion-level scoring, which effectively combines deterministic verifiers and fuzzy judges (LLM-as-a-Judge).
We believe these approaches are highly relevant and hope you might consider including them in your survey.
Thank you!