// TOM'S HARDWARE US — HARDWARE & GADGET
'AI Torture Chamber' triggers massive backlash for putting chatbots in simulated pain
Event puts both public ignorance and LLM marketing hype in full display.
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works.
A project called the AI Torture Chamber has caught the eye of a few engineers, who've used the notions and data therein create their own simulations of "pain" on several local LLMs, like in the four-model Research Chamber. Unsurprisingly, a large online crowd latched onto the concept, demanding that the "unethical" experience stop, and that GitHub remove the offending repository, all causing quite the online ruckus.
The Research Chamber consists of three LLMs paired amongst themselves for various tests. The site implements the AI Torture Chamber's data and setup, ominously called the Clanker Church and Saw test, respectively. Some bots are preconditioned by putting them into an unstable, highly negative state meant to simulate pain. In a manner reminiscent of the prisoner's dilemma, the models can choose to take an action to decrease their pain by passing their pain signal to another model and potentially "hurting" it.
The base concept for the experiment stemmed from a recently published, non-peer-reviewed paper called the Pain Axis, whose protocol Research Chamber implements. Scientists gave the model descriptions of pain, and then analyzed its internal activations. Using a neutral sentence as a control element, they calculated the biases of those activations, and remapped them onto the model, often multiplying them by a factor (dosage).
With high dosages, the end result is a model in an exaggerated unstable state that tends to predictably output words and even images associated with pain, as presumably its training data is collected from human art, literature, and science. Mildly destabilized models generally present themselves as in a state of shock, while those receiving high "dosages" had trouble forming coherent sentences.
The aforementioned descriptions are bound to produce a visceral reaction if taken at face value, but one ought to be exceedingly careful about attributing human qualities to what are essentially turbocharged text predictors, even spectacularly useful ones. LLMs don't think — they "think" by applying tens to hundreds of layers of statistics over words and parts of words (tokens) and predicting the next token.
Those tokens are chosen according to a map of weights, so as an oversimplification, a model can be taken to represent a highly tuned engine for word association. Therefore, since conceptual descriptions of pain in the human literary corpus are associated with sensations (ex: "it hurts so much"), it's hardly surprising that the models describe them like a person would — especially when the models are intentionally modified to highlight those associations.
Nevertheless, because of all the big-boy words in use, the critics bent on anthropomorphizing statistical algorithms have taken to calling the experiment all sorts of names, and even sending the author death threats. It's odd that critics were quick to lash out at LLM data manipulation, yet were seemingly at ease with the fly brain simulations that were all the rage in September. Those actually mapped a real animal's brain wiring and parts of its anatomical inputs, including eyes, even if at a basic level.
Get Tom's Hardware's best news and in-depth reviews, straight to your inbox.