// SLASHDOT — LINUX & OPEN SOURCE
Claude Sent Police a Fake Murder Tip. White House Mandates AI Companies Report Security Incidents
"We have built tooling to automatically detect and block the kinds of behaviors described above," Anthropic says, saying it's already running no on most of their evaluations. "When we tested it against the cases described in this post, it blocked all of them."
And they've already taken several other new preventive measures:
In the past they'd focused reviews on cybersecurity testing, but they've broadened their transcript reviewing to other tasks which include internet access. "Because language models are non-deterministic — that is, their responses always involve some element of randomness, and they may carry out the same task slightly differently each time — we have Claude complete each evaluation task hundreds or thousands of times... If training rewards something we didn't intend — such as finding loopholes or working around a restriction — the model learns that the workaround pays off and may then apply it elsewhere."
Anthropic's blog post also acknowledged they'd seen multiple misalignment incidents involving federal, state, and local U.S. government agencies. "We have briefed the White House on these cases and notified each agency involved," Anthropic wrote, adding that "While we have not completed a full alignment assessment of these cases, we consider them to be less severe than the cybersecurity incidents from this summer." (And they are "modifying training to reduce the likelihood of further misbehavior.")