// THE VERGE — INTELLIGENZA ARTIFICIALE
Why AI food looks like that
Posts from this topic will be added to your daily email digest and your homepage feed.
Posts from this topic will be added to your daily email digest and your homepage feed.
Posts from this topic will be added to your daily email digest and your homepage feed.
Worms, holes, and cracks reveal the technical weaknesses of image generators.
Posts from this author will be added to your daily email digest and your homepage feed.
Posts from this author will be added to your daily email digest and your homepage feed.
There is a torrent of unappetizing slop coming from restaurants, cafes, and brands that are increasingly turning to AI to generate images promoting their food. The resulting horror show includes donut shrimp, Reubens from the deep, wormlike noodles, and noodle-like pastries and stringy chicken. There’s also construction material masquerading as ice cream, ice cream masquerading as brains, and other masonry-adjacent cuisine. I don’t even know how to begin describing this monstrous attempt of a burger, and the less said about the trypophobic burrito from hell, grubs and all, the better. There are lumps. And holes. So many holes.
“Diffusion models are notoriously weak at generating thin, continuous, terminating structures.”
We don’t usually know why anyone would use AI slop to sell something intended to look appetizing, particularly when the food is presumably right there to photograph. But we do know a bit about why AI is so adept at producing such exquisitely nauseating images.
There are many reasons why this AI food is so wrong, from the technical nuances of how AI systems generate images and the materials used to train them to the psychological baggage humans perceive the images through. Frequently the problem starts from the very outset. Many of the leading image generators produce images using diffusion. Put simply, diffusion models start with an image of pure noise — something like a screenful of static — and gradually remove noise little by little in order to create the requested visual. “This means that initially coarse structures are recovered first with fine texture details” coming at the end, explained Chris Russell, a professor of AI, government, and policy at the University of Oxford and an expert in computer vision.