Apple's RLTL;DR Teaches AI to Learn From Its Own Failures
Apple researchers show a 9B model can crack tasks it failed 128 times in a row by backpropagating short self-written notes instead of full solutions.
- Apple proposes RLTL;DR, letting a policy write one-sentence hints after failures and backprop on them.
- On Pass@128=0 tasks, Qwen 3.5 9B goes from 0-1% to 12-13% Pass@1 without hints at eval.
- Hints are tiny: averaging 17 tokens each, inserted as user messages between sequential attempts.
- Stripped-down variant SFTL;DR trains only on 4k (task, insight) tuples and nearly matches full method.
- Short TL;DR hints beat detailed diagnostics; episode-specific details like IDs block generalization.
- Method helps only on frontier-difficulty tasks where GRPO flatlines; neutral on normal-difficulty training.
Apple’s RLTL;DR Trains Models on Their Own Failure Notes
Reinforcement learning with verifiable rewards stalls when a model cannot produce a single correct answer. Every rollout receives the same zero reward, leaving the optimizer without a useful gradient. Apple’s machine learning group addresses that failure mode in the RLTL;DR paper by having the model write a one-sentence lesson after each failed attempt and training it to recover that lesson directly from the task.
On tool-calling and coding tasks where Qwen 3.5 9B Thinking failed all 128 sampled attempts, standard Group Relative Policy Optimization, or GRPO, remained at 0% to 1% Pass@1. RLTL;DR reached 12% to 13% Pass@1 during evaluation, when no hints were supplied in context, according to the reported results.
Why GRPO Flatlines
GRPO samples several answers for the same prompt, scores them, and updates the policy according to each answer’s reward relative to the group average. If every answer earns the same reward, each relative advantage becomes zero. That commonly happens on frontier-difficulty tasks because every sampled answer fails.
Earlier approaches supplied hints from a stronger teacher or a reference solution, creating guidance that may be unavailable during deployment. RLTL;DR uses the policy’s failed attempts and verifier error messages, without requiring an external teacher or successful reference trajectory.
Failures Become Compact Training Targets
RLTL;DR modifies sampling and optimization through two connected steps:
-
Sequential attempts with accumulated hints. The policy attempts the same task repeatedly. After each failure, it receives the failed rollout and verifier errors, then writes a one-sentence lesson such as “Remember to paginate search results.” Later attempts receive all lessons generated for that task. Hint injection activates when the task’s running success rate is 50% or lower.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.