Petri
Alignment auditing agent for probing language model behavior
Petri is an open-source alignment auditing agent/toolbox for probing language model behavior. It generates realistic audit scenarios, orchestrates multi-turn interactions between auditor and target models, simulates tools and rollbacks, and scores transcripts with judge models to detect alignment issues such as deception, sycophancy, harmful cooperation, and evaluation awareness.

Recent stories
0 linked stories
No linked stories yet.