A model cannot grade its own homework
Self-correction fails without external feedback, and training on your own output degrades the model. What the collapse and verifier research says you should do instead
· 7 min read
Research · Pillar
Models that train on their own output.
Bootstrapped reasoning, self-rewarding training, and the hard limits the research keeps finding: model collapse and the myth of unaided self-correction.