Implementation of the training framework proposed in Self-Rewarding Language Model, from MetaAI
-
Updated
Apr 11, 2024 - Python
Implementation of the training framework proposed in Self-Rewarding Language Model, from MetaAI
SQL-o1: A Self-Reward Heuristic Dynamic Search Method for Text-to-SQL
Reinforcement Learning of Vision Language Models with Self Visual Perception Reward
SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
An official implementation of "SPARK: Synergistic Policy And Reward Co-Evolving Framework"
GenPark AI Agent Skill - Iterative self-rewarding LLM judge scoring multi-turn agent outputs and synthesizing contrastive hard-negative prompt variations.
GenPark AI Agent Skill - Direct Preference Optimization (DPO) implicit reward calculation, reference policy log-ratio tracking, and pairwise preference loss evaluation.
GenPark AI Agent Skill - Prioritized experience replay buffer for agent reinforcement fine-tuning (RFT / GRPO) with advantage estimation and importance sampling weights.
GenPark AI Agent Skill - Pairwise Bradley-Terry and Elo rating tournament engine for ranking multi-agent strategies, prompt mutations, and generated tool solutions.
GenPark AI Agent Skill - Direct Preference Optimization (DPO) implicit reward calculation, reference policy log-ratio tracking, and pairwise preference loss evaluation.
GenPark AI Agent Skill - Pairwise Bradley-Terry and Elo rating tournament engine for ranking multi-agent strategies, prompt mutations, and generated tool solutions.
GenPark AI Agent Skill - Kahneman-Tversky Optimization (KTO) loss evaluator aligning agents directly on unpaired binary feedback (thumbs-up / thumbs-down) using prospect theory loss aversion.
GenPark AI Agent Skill - Iterative self-rewarding LLM judge scoring multi-turn agent outputs and synthesizing contrastive hard-negative prompt variations.
GenPark AI Agent Skill - Prioritized experience replay buffer for agent reinforcement fine-tuning (RFT / GRPO) with advantage estimation and importance sampling weights.
GenPark AI Agent Skill - Kahneman-Tversky Optimization (KTO) loss evaluator aligning agents directly on unpaired binary feedback (thumbs-up / thumbs-down) using prospect theory loss aversion.
Class-Conditional self-reward mechanism for improved Text-to-Image models
Does a model training on its own judgment collapse? On verifiable math: the reward gets hacked (a brevity reward halves answer length) but capability does not collapse. An honest failure analysis with a judge-quality probe and per-round trajectory datasets.
To associate your repository with the self-rewarding topic, visit your repo's landing page and select "manage topics."