Governing an agent that rewrites itself
An agent that edits its own code needs a gate, not a hope. What the safety literature says about reward hacking, sandboxes, oversight, and typed contracts
· 8 min read
Research · Pillar
Agents that rewrite their own playbook.
Skill libraries, workflow memory, automated agent design, and code that modifies itself. What the evidence shows, and where it breaks.