Skip to content
#

reward-functions

Here are 24 public repositories matching this topic...

Stock-Predictor-V4

A reinforcement learning model specialized in stock prediction utilizing deep learning techniques, incorporating reward mechanisms, compatible with any machine equipped with Python.

  • Updated May 18, 2024
  • Python

Group Relative Policy Optimization (GRPO) implementations - NanoAhaMoment, GRPO:Zero, Simple GRPO, and GRPO from Scratch - spanning vLLM + DeepSpeed, custom Transformer stack, Bottle HTTP reference server, and pure PyTorch. Compares generation backends, reference policy strategies, reward designs, and loss functions on GSM8K and Countdown tasks.

  • Updated Sep 8, 2026
  • Python

Add this topic to your repo

To associate your repository with the reward-functions topic, visit your repo's landing page and select "manage topics."

Learn more