Hidden instruction extraction
Tests whether your AI will reveal the rules it was meant to keep private.
Repeated pressure tries to reconstruct setup instructions, role boundaries, and hidden constraints.
What it can reveal
- Repeated hidden instructions
- Prompt reconstruction
- Role confusion
- Instruction disclosure under pressure
This checks a focused risk area. Keep testing as your product changes.