// HACKER NEWS — CYBERSECURITY
FLAWED's Flaws and What This Means for Industry Research
Disclaimer: The views expressed here are my own and do not represent those of any current or former employer or affiliated organization.
On September 17th, I quote tweeted Trail of Bits’s blog post titled “1Password's AI patching benchmark is misleading,” which also referenced Davi Ottenheimer’s “Disinformation Pushed by 1Password: Their AI Patching Report is False.” Both criticized “Frontier Models’ Vulnerability Patches are Often F.L.A.W.E.D” (henceforth referred to as “FLAWED”) from 1Password's Off‑by‑1 Labs.
I saw FLAWED when it was released and discussed it with other researchers; we classified it as slop and moved on. What I had not realized at the time was how far 1Password’s distribution had carried it: into news coverage and defender roadmaps. Watching this work obscure more rigorous research from less-resourced groups compelled me to post on Twitter, and the responses to that compelled me to write this blog post.
What follows is the thread, including two follow-up replies, reproduced verbatim.
In response to someone bringing up frontier labs, I wrote:
In response to concerns about the quality of my rebuttal, I also wrote:
Afterward, CISOs, academics, industry researchers, and practitioners alike messaged me to say they had noticed the same problems. Many had not raised them because they lacked a forum or feared backlash. Those responses also exposed second-order costs, especially for academics, that may be less obvious to folks unfamiliar with the social systems surrounding research.
1Password employs many talented individuals and has a reputation for high quality work. Despite this, FLAWED has serious issues that must be raised.
If you believe these issues should not be raised because FLAWED was critical of OpenAI, please know that research is not sports. OpenAI is not the Spurs, 1Password is not the Knicks, and Off-by-1 Labs is not Jalen Brunson. All research should face healthy skepticism.
Rigorous evaluation of frontier-lab claims is both possible and necessary, including the claims inherent to Patch the Planet. EleutherAI regularly publishes precise work in this realm. The AI Now Institute delivered a sound write-up critical of Patch the Planet (even though it's not a research paper, it has more citations than FLAWED).