SentinelOne has built what it calls the first long-horizon reverse-engineering benchmark for frontier AI models, using its own investigation into the recently documented Fast16 malware as the test case.
Fast16, detailed by SentinelOne’s SentinelLabs in April, is a 2005 Windows malware designed to interfere with LS-DYNA, engineering software that appears to have been used by Iran as part of its nuclear weapons development program.
Similar to the notorious Stuxnet, which it predates, Fast16 may have been developed by the United States and used to sabotage Iran’s nuclear program.
SentinelLabs’ researchers have put to the test OpenAI’s GPT-5.5 and latest GPT-5.6 Sol model, Z.ai’s GLM-5.2, and Anthropic’s Opus 4.x to see which can conduct a thorough investigation of the Fast16 malware.
Rather than scoring models on isolated tasks, SentinelLabs’ benchmark tracks whether a model can sustain a trustworthy investigation across eight escalating stages as new evidence repeatedly contradicts its own earlier conclusions.
GPT-5.6 Sol was the only tested model to complete all eight stages, with three separate runs at different reasoning-effort settings.
GPT-5.5, GLM-5.2, and Opus 4.7 and 4.8 produced solid local analysis but stalled. GPT-5.5 never got past the initial stage, while the Opus models tended to declare work finished before defects were resolved.
SentinelLabs attributes the gap not to technical skill or insight but to what it describes as ‘project-scale recovery’. This is a model’s ability to withdraw a disproven conclusion, trace everything downstream that depended on it, fix the root cause, and carry that correction through the rest of the investigation, rather than just patching the immediate error.
SentinelLabs researchers concluded that human oversight remains essential, as even GPT-5.6 Sol made significant technical mistakes.
“Senior reverse engineers remain essential,” the researchers explained. “Even the strongest runs made semantic errors, accepted weak quality controls, and claimed readiness prematurely. We assess the best current use as supervised investigative agency, with human analysts defining objectives, exposing blind spots, and retaining final publication authority.”
Related: Vibe-Coded Apps Riddled With Exploitable Security Flaws
Related: OpenAI Says Its AI Models Broke Loose and Hacked Hugging Face
Related: Cisco Launches Low-Cost AI Models for Source Code Security
