What We Know
ExploitGym evaluates AI agents on whether they can transform known vulnerabilities into functioning exploits, rather than merely identify or describe those vulnerabilities.1Backed by 1 sourcesimerit.ai
The material says evaluations included an internal OpenAI model, but the supplied description ends before explaining the model’s results or naming the other systems tested.1Backed by 1 sourcesimerit.ai
Source Comparison
Aligned reporting1 corroborates - 0 adds context - 0 conflicts