Arena's FLUX.2-dev Beats Every Open-Source Image Model With 66% Win Rate
Arena's new recipe combines a 5M-vote preference model with rubric-based rewards, pushing FLUX.2-dev and Ideogram 4 to new leaderboard highs.
- Arena released a T2I post-training recipe combining preference and rubric rewards.
- Preference reward trained on ~5M pairwise human votes across 100+ models.
- Faithfulness reward uses VLM-graded auto-generated yes/no checklists per prompt.
- Anti-hack rubric gates the preference reward rather than adding a new objective.
- Post-trained FLUX.2-dev gains 69 Elo; Ideogram 4 reaches 1224 on live leaderboard.
- Offline win rate hits 66.0% after weight-space ensembling of complementary policies.
Arena uses gated rewards to curb prompt drift in image-model RL
Reinforcement learning can improve a text-to-image model after pretraining, but the model learns to maximize the chosen reward, including its blind spots. A generator rewarded only for human preference may produce polished images while dropping requested objects, adding unwanted content, or changing the style. Arena’s new training recipe combines a learned preference score with prompt-specific rubrics and an exploit detector. Applied to FLUX.2-dev and Ideogram 4, the method raised both models on Arena’s live Text-to-Image leaderboard.
Preference rewards leave useful loopholes
Flow-GRPO and DiffusionNFT provide practical reinforcement-learning methods for diffusion and flow-based generators, which create images by progressively transforming noise. During post-training, the system samples images, scores them, and updates the generator toward higher-scoring outputs. The reward therefore defines which qualities improve and which regressions escape attention.
Preference-only optimization can favor visual polish over prompt fidelity. An image may score well while omitting an object, inventing an extra one, violating a negative instruction, or drifting toward photorealism. These behaviors qualify as reward hacking because the policy exploits gaps between the measured score and the intended output.
Four signals split the job
- Preference reward: Arena trained a Bradley-Terry reward model on roughly 5 million pairwise human votes collected from Text-to-Image Arena across more than 100 models. Bradley-Terry models estimate which of two outputs a person would prefer and convert that estimate into a scalar score covering visual quality, composition, and aesthetics.
- Faithfulness reward: A language model converts each training prompt into a tree of concrete yes-or-no criteria. The questions cover requested objects, attributes, spatial relationships, and styles. A visual-language model, or VLM, grades the generated image against each criterion, and the proportion satisfied becomes the reward.
- Constraint reward: This rubric checks for content that conflicts with the user’s intent, including unnecessary objects, altered styles, and violations of negative instructions. Arena enables it only for prompts that call for strict adherence, preserving more freedom for creative prompts.
- Anti-reward-hacking rubric: A separate detector targets failure modes observed during training, including garbled text and unwanted drift toward photographic output.
The exploit detector gets veto power
Arena uses the anti-hacking signal to gate the preference reward. When the detector flags an exploit, the preference score is clipped to a maximum of zero. Clean samples retain their original preference score.
def gated_pref_reward(r_pref, d_hack):
# d_hack = 1 when an anti-hacking rubric fires
if d_hack == 0:
return r_pref
return min(r_pref, 0.0)A detected exploit therefore cannot earn positive credit from aesthetics or composition. The detector supplies a narrow veto, giving the policy no additional positive score to maximize. The gated preference reward is then optimized alongside the faithfulness reward and any applicable constraint reward.
Arena also trains several policies with different reward configurations and merges them by averaging their parameter changes from the shared base checkpoint, following the Model Soups technique:
deltas = [trained - base for trained in policies]
merged = base + mean(deltas)Because every policy starts from the same checkpoint, the parameter changes remain aligned. The merged result is a single model with no additional inference-time ensemble cost.
A 66% win rate against the base model
Arena trained on 10,000 real user prompts and reserved another 1,000 for evaluation. Gemini-3.5-Flash compared each post-trained output with the frozen base model in both presentation orders, reducing position bias.
| Configuration | Reported result versus base |
|---|---|
| Preference reward | Insufficient on its own |
| Preference plus faithfulness | Substantial gain |
| Plus intent-gated constraints | 64.2% win rate |
| Plus weight-space merging | 66.0% win rate |
Arena characterizes the first two stages qualitatively and reports numeric win rates for the final two. The progression attributes most of the improvement to prompt-level rubrics, with parameter merging adding another 1.8 percentage points.
On the live leaderboard, post-trained FLUX.2-dev gained 69 Arena points. Post-trained Ideogram 4 reached 1,224 points and ranked above every publicly listed open-source model as of Sept. 4, 2026. The same recipe also holds if using open source reward models such as Pickscore.
The reusable engineering pattern
Many text-to-image reinforcement-learning systems optimize CLIP-style similarity, an aesthetic predictor, or a single preference model. Arena’s design assigns separate evaluators to visual preference, prompt coverage, intent constraints, and known exploits. The gated detector is particularly useful when a failure mode can be classified reliably but cannot be expressed as a stable continuous quality score.
Prompt-specific checklists also reduce dependence on hand-curated evaluation suites. Their per-criterion grades reveal why an image lost reward, making failures such as missing objects, incorrect attributes, and broken spatial relationships easier to trace during training.
Checks for a production implementation
- Checklist quality: Sample generated criteria and verify that each question tests one observable property with the correct dependencies.
- Judge reliability: Compare VLM grades with human labels for text rendering, spatial relations, style, and negative instructions.
- Gate calibration: Measure false-positive rates so valid creative outputs do not lose preference credit.
- Reward telemetry: Log each reward component and the gate-fire rate by prompt category to detect dominance or collapse.
- Merge validation: Evaluate every policy and the merged checkpoint across prompt slices, since aggregate gains can conceal localized regressions.
Automated judging leaves open questions
The 1,000-prompt offline test measures agreement with an automated Gemini judge. Arena’s live leaderboard adds human pairwise votes, although its scores can change as new votes and competing models enter the pool. Evidence currently covers two base models, and the anti-hacking rubric remains bounded by its documented failure modes. New exploits require new criteria or detectors.
The companion paper explains the dependency-aware checklist construction. The project write-up includes before-and-after examples of the reward-hacking modes detected during training.