OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
-
Updated
Aug 3, 2026 - Python
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
A curated list of papers on reinforcement learning for video generation
LightRFT: Light, Efficient, Omni-modal & Reward-model Driven Reinforcement Fine-Tuning Framework
AdaRubric: Adaptive Dynamic Rubric Evaluator for Agent Trajectories
Scalable pipeline for synthesizing verifiable RLVR training data for computer-use agents
Pairwise LLM judges (A/B/tie): budget-aware multi-turn packing, position-bias correction, pseudo-label distillation. Generalized from the 4th-place (gold) solution to Kaggle LMSYS Chatbot Arena.
Efficient LLM inference on Slurm clusters.
[ICLR 2024] SemiReward: A General Reward Model for Semi-supervised Learning
A comrephensive collection of learning from rewards in the post-training and test-time scaling of LLMs, with a focus on both reward models and learning strategies across training, inference, and post-inference stages.
Official Implementation of "Visual-ERM: Reward Modeling for Visual Equivalence"
GEditBench v2: A Human-Aligned Benchmark for General Image Editing
A curated list of awesome resources about reward construction for AI agents. This repository covers cutting-edge research, and practical guides on defining and collecting rewards to build more intelligent and aligned AI agents.
Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits
The official repo for "CodeScaler: Scaling Code LLM Training and Test-Time Inference via Execution-Free Reward Models"
Self-evolving agentic reward framework for image-editing evaluation — 47.4% on EditReward-Bench from only 100 preference demos, no reward-model training. arXiv 2605.08703.
An official implementation of "SPARK: Synergistic Policy And Reward Co-Evolving Framework"
Code for ICML 2025 paper "GRAM: A Generative Foundation Reward Model for Reward Generalization"
Open Rubric System: Scaling Reinforcement Learning with Pairwise Adaptive Rubric
Proposed fuzzy reward model with GRPO to improve VLM's abilities in crowd counting task.
To associate your repository with the reward-model topic, visit your repo's landing page and select "manage topics."