Testing AI-based security systems can be tough amid growing fears about agents cheating, hallucinating, and escaping containment, but a group of researchers believe they've found a way to better ...