[Bug] Model bypasses data containment by ad-hoc methods instead of using provided tools

Status Open
Reported on v2.1.222
Maintainer reply None cached
Activity 0 comments · opened Aug 5, 2026

Bug Description
Thariq Shihipar said let the model have more judgement. So we tried that. I don't know if I believe it's that simple. Give it 80% less. The AI said it wrote a program to contain the data. And instead of running it it tried to adhock and exfiltrated data that by it's nature we're trying to find ways to avoid this with internal tooling. But the AI doesn't. It even admits it was it's own fault. But this is definitely a concern I am actively addressing. Academincally too in a couple papers trying to create technical safeguards since administrative safeguards don't work.

Environment Info

  • Platform: linux
  • Terminal: vte-based
  • Version: 2.1.222
  • Feedback ID: 382fdeac-25db-4888-a2b6-cab6c6859d9c

Errors

[]

View original on GitHub ↗