Most reinforcement learning methods train an AI to maximize its average reward. That sounds sensible—but averages can hide something important. https://arxiv.org/abs/2609.
Think of two AI systems that both score an average of 70. One almost always produces results around 70. The other usually scores 60, but occasionally discovers an exceptional solution scoring 100. If we are allowed to sample the AI many times and choose the best result, the second system could be far more valuable.
This is the problem addressed by Tail-Likelihood Reinforcement Learning (TailRL).
Instead of asking only, “What reward does the AI achieve on average?”, TailRL asks a different question across the entire range of outcomes:
“How likely is the AI to produce something better than a particular reward threshold?”
The method effectively turns continuous rewards into a collection of success-or-failure questions. Importantly, it gives greater learning weight to rare, high-reward outcomes rather than allowing them to disappear into the average.
There is an interesting connection to Best-of-k sampling: when we generate multiple attempts and keep the best one, performance depends on whether the model has retained enough probability mass around those unusually good solutions. TailRL explicitly trains the model to preserve that potential.
The practical advantage is significant. The researchers report that TailRL helps models avoid getting trapped in mediocre solutions across tasks including object localization, maze navigation, GUI grounding, and code optimization. More importantly, the resulting models become better at exploiting additional sampling at inference time—the more attempts you give them, the more useful those attempts become.







