Claim: we’ve solved the AI slop problem (!) 💩🧹✨
Blog post: https://facebookresearch.github.io/RAM/blogs/unslop/ by: Jason Weston.
Key idea: take *expert* human writing and learn rubrics that find the gap between experts and models. Train with those rubrics.
We train with RL-XAR (RL with eXpert Aligned Rubrics) & see large performance gains on writing scientific paper sections, Pulitzer prize novel continuations and high quality Wikipedia pages.
First: The failure of standard LLM Judgements 💀
On paper writing tasks, strong judges (GPT-5.6 or Opus-4.8) think current ‘slop’ models are better than humans on selected high quality papers (using either pairwise, or using standard rubrics).
Our method can learn rubrics where the human is considered better by the grader (right in fig) – the key to training.
RL-XAR (RL with eXpert Aligned Rubrics) recipe 👩🍳:
Collect examples of high quality human-written texts learn LLM judgments via rubrics that score those expert texts higher than model generations RL on the learnt rubrics.
This procedure is iterated until meta-optimization of the rubrics can no longer find a discernible gap.
Optimization of the rubrics changes them from preferring current ‘slop’ models to preferring humans.
We perform RL-XAR across 3 tasks (writing scientific paper sections, story continuations and writing high quality Wikipedia pages) using Qwen3.5-27B, rewarding each generation by the learned rubric as scored by a cross-family Qwen3.8–2.4T-A95B judge.
Final results are shown with a further judge, GPT5-6. We find strong gains across all 3 tasks, particularly the first two.
We also performed a small human eval on papers and stories, with large wins for RL-XAR (89% and 95% win rates).
For papers, training tends to fix over-scoping and sectional focus. For stories it tends to remove clichéd writing.
#artificialintelligence #LLM #AIjudge #ReinforcementLearning
Reinforcement learning from eXpert-Aligned Rubrics for expert-level text generation.
