// NATURE NEWS — SPAZIO & SCIENZA
The Einstein test: what happens when AI tries to rediscover relativity?
Search author on:
PubMed
Google Scholar
In 1915, Albert Einstein unveiled his general theory of relativity and transformed our view of the fabric of the physical world. The theory, which explains gravitation as a deformation of space-time by mass, is a pinnacle of modern physics that underpins cosmology, from work on black holes to measurements of gravitational waves, and is used routinely to guide space missions and GPS satellites.
It has also become a yardstick for leaders in the field of artificial-intelligence technology, who are asking whether their creations could ever make a breakthrough on that level. At the India AI Summit in New Delhi this February, Demis Hassabis, the co-founder of Google DeepMind in London, proposed training a large language model (LLM) on all that was known before a particular cut-off date — he suggested the year 1911 — to see whether it could reproduce general relativity. “That would be a good test for AGI,” Hassabis said, referring to the nebulous concept of artificial general intelligence that is a goal for many in the AI industry.
Such a test needn’t specifically involve general relativity. In December 2024, Owain Evans, a researcher at the non-profit organization Truthful AI in Berkeley, California, gave a talk about ‘vintage’ or ‘historical’ LLMs, which would be trained only on historical data up to a certain date. Evans asked what such models might be able to rediscover.
This year, several teams have built instances of vintage models, including an effort at Hassabis’s test. But their early attempts reveal more about the limitations of current AI than about its strengths.
A relativity-like breakthrough is not inherently out of reach for AI, says Ido Kaminer, a specialist in quantum optics at Technion —Israel Institute of Technology in Haifa, who co-authored a preprint titled ‘Can AI follow in Einstein’s footsteps?’, posted in July1.
But, he and his colleagues argue, it won’t happen without rethinking some of the principles on which today’s models are built.
Hassabis had mentioned the Einstein test analogy in media interviews last year, but another Google DeepMind researcher, Tom Zahavy, described it in detail in a position paper posted on his website in January. Titled ‘LLMs can’t jump’ (see go.nature.com/3ykartg), the paper highlights the current inability of LLMs to make jumps of reasoning like Einstein’s.
Zahavy wrote, as philosophers of science have long recognized, that such advances require not inductive reasoning that derives a general rule from the accumulation of data or examples, but abductive reasoning: “a creative leap that invents a cause for a singular phenomenon”.
That isn’t the forte of current AI models, which seem better suited to doing the grunt work of science than to making transformative discoveries. The models look for correlations in vast data sets, through being trained on known examples and then guided by prompts to supply the most statistically likely answers for unknown cases.