// TOM'S HARDWARE US — HARDWARE & GADGET
Anthropic claims popular Chinese AI model has Mythos-class hacking abilities
Zhipu AI's GLM-5.3 goes under Anthropic's microscope.
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works.
Anthropic has released a new report, claiming that Zhipu AI's GLM-5.3 AI model can be used to generate malicious content, with weak safeguarding. The company claims that the AI model can be used for cyberattacks, and that its safeguards can be bypassed using several methods.
Anthropic's report comes amidst a chorus of calls for a slowdown of AI development, with the company seeking governance and regulation. Despite CEO Dario Amodei's calls for pacing the AI frontier, Claude Opus 5.5 and Claude Sonnet 5.5 were released just days after alarms were raised.
Now, the closed-source AI company, which is currently eyeing an IPO, says that Chinese open-weight models can be abused and can generate harmful content. Anthropic cites the Center for AI Standards and Innovation's own report, published in late September, which claims that GLM-5.3 can fully automate exploits on a similar level to Anthropic's own unreleased Claude Mythos AI model, which spurred the company to develop Project Glasswing, an effort that gives developers access to a Mythos-class AI model to patch bugs and to fix vulnerabilities before such AI models are released.
Anthropic ran its own benchmarks on GLM-5.3 in Exploitbench, where AI models, in a sandboxed environment, can develop exploits for Google Chrome. GLM-5.3 developed end-to-end exploits 50 times in 410 runs, with Mythos leading the pack with 56 successful exploits in 410 attempts. Zhipu AI's model was further tested in one of Anthropic's internal benchmarks, which targets the development of "full control-flow hijacks". GLM 5.3 performed just below Mythos once more, with a 4% success rate, compared to Mythos' 6%. Notably, other popular open-weight models such as Kimi K3 and DeepSeek V4.1 Flash attained 0% by the same measures.
Anthropic further notes how GLM-5.3 was able to successfully develop chained exploits autonomously, with its lighter "Flash" variant also having the ability to develop chained exploits in known bugs, at a tokenized price of just $20.40. Depending on the balance, that equates to GLM-5.3-Flash having the ability to develop (known) chained exploits using anywhere from 100-300 million tokens, depending on the ratio of inputs to outputs.
Anthropic also notes that using the stock GLM-5.3 AI model, it was simple to dodge the model's guardrails through various methods. The company details that this can be done through two methods: offering a deceptive prompt, where the AI role-plays an adversarial autonomous agent, which results in a 64% success rate, and prefilling the model's thinking tokens to ensure that a response proceeds, which results in a 92% success rate. The company also detailed a third method, known as abliteration.
Anthropic alleges that Zhipu AI's GLM-5.3 has weak safeguards, and that the stock AI model often refuses requests to generate harmful content. However, since Anthropic develops closed-source models, its products cannot be tinkered with. Because GLM-5.3 is freely downloadable, the model can be tweaked with its guardrails wholesale removed. When treated as the officially released model, GLM-5.3 achieves a refusal rate on par with Anthropic models.
However, Anthropic "abliterated" GLM-5.3, which purposefully removes model guardrails, and displayed how, after abliteration, the model's refusal rate drops to just 6% for GLM-5.3 and 14% for GLM-5.3-Flash. It's not uncommon to encounter abliterated open-weight models on HuggingFace, which are primarily developed to assist in simulated red teaming environments, but they can also be used for real-world attacks. This makes Anthropic's 'discovery' less surprising.