Why AI-generated food images look so wrong

The Verge surveyed restaurants, cafes and brands that have started using AI image generators to produce photos of their food, and the results are often unappetizing: donut shrimp, wormlike noodles, noodle-like pastries, stringy chicken, ice cream that looks like cracked concrete or brains, and burgers and burritos riddled with holes and lumps. The piece asked several academics to explain why AI-generated food so often looks wrong, and the answers point to technical weaknesses in how the images are made, gaps in what the underlying models actually understand, and quirks of human perception.
Most leading image generators work by diffusion: they start from an image of pure noise, similar to television static, and remove that noise step by step until the requested picture emerges. Chris Russell, a professor of AI, government and policy at the University of Oxford and a computer vision expert, explained that this means coarse structure is recovered first and fine texture is added at the end. If the model gets the basic structure of an object wrong early on, it then piles vivid texture on top of that flawed structure, the same category of failure that produces a person with six fingers instead of five, which Russell said could also explain the donut shrimp.
Even when the underlying structure is sound, the finer details can still go wrong. Giovanbattista Califano, a behavioral scientist who studies responses to AI-generated imagery at the University of Naples Federico II, said diffusion models are notoriously weak at generating thin, continuous, terminating structures. Noodles, strands and tendrils fall exactly into that category, so the models produce spaghetti-like artifacts that bleed into places with no anatomical or culinary logic, and struggle to decide where such shapes should stop or what they should be attached to. Other repeating textures, such as bubbles and seeds, are similarly hard for the models to contain within sensible boundaries and often spill where they should not, which Califano said helps explain why so many AI food images look noodly, oddly patterned and dotted with the kind of clustered holes that can trigger trypophobia.
A deeper problem is that the models have no real understanding of what a sandwich, a noodle or a burrito actually is. Roland Meyer, a professor of digital cultures and arts at the University of Zurich, said AI image generation reproduces looks without proper knowledge about the world: the systems have learned, on a statistical level, what these foods tend to look like, but not why they look that way or how they behave physically. That gap can also make otherwise coherent AI images read as wrong to viewers. Michael Cook, a senior lecturer in computer science at King's College London, pointed to ice cream that resembles cracked concrete or burgers that look built from rocks: those textures might look normal in an architectural context, but become wrong once imagined as edible food, and the models have no built-in sense of that distinction.
The training data compounds the problem. Cook noted that because so little is known about how these systems are trained, it is not clear what mix of content they receive or what associations they form. Food photography is often highly stylized, with sharp contrasts, intense colors and glossy lighting, and Meyer said models pick up on those surface qualities without understanding the professional aesthetic strategies behind them, which he called the source of much of their unsettling quality. Simon Colton, a professor of computational creativity, games and artificial intelligence at Queen Mary University of London, added that ordinary images are underrepresented online because nobody wants to post a picture of a boring apple, while stranger images spread widely on platforms like Reddit or as memes, skewing the models' learned associations around food. Cook also said it is well known that AI systems are now increasingly trained on AI-generated material, including a recent trend of AI-generated videos of people jumping into piles of food, and research suggests training models on other models' outputs can cause a form of "model collapse," leading to visual degeneration and growing sameness between images.
How the images are produced and used adds further problems. Both user prompts and the system-level instructions companies build into their products may not yield good results: a prompt might be as vague as simply asking for a sandwich, or use language such as "be precise," which makes sense for text but not for image generation. Low-resolution images can also be scaled up well beyond their intended size, magnifying imperfections or forcing the model to fill in gaps it cannot resolve accurately.
The result lands squarely in the uncanny valley, and humans are especially well equipped to notice it. Scientists believe disgust evolved partly to protect people from parasites, pathogens and toxins, making people particularly attuned to signs that food may be unsafe. Califano said this makes the uncanny valley for food feel even more visceral than the one people experience with almost-human faces. Noodly tendrils resemble worms or parasites, clusters of holes suggest infestation, and off colors and textures signal contamination or spoilage, so AI-generated food triggers a primal sense that something is wrong even when it superficially resembles a meal.
Key facts
- Diffusion models generate images by removing noise step by step, recovering coarse structure first and adding fine texture last, so an early structural mistake gets compounded by texture piled on top of it, the same failure class as six-fingered hands.
- Diffusion models are notoriously weak at rendering thin, continuous, terminating shapes such as noodles, strands and tendrils, which is why AI food images are so often noodly and dotted with clustered holes.
- The models have no real understanding of what food actually is; they reproduce statistical appearances learned from images without grasping the physical logic behind them, per Roland Meyer of the University of Zurich.
- Training data skews toward unusual, highly stylized food photography since ordinary images are rarely posted online, and models are increasingly trained on other AI-generated material, which research links to a degrading 'model collapse' effect.
- Human disgust evolved to flag parasites, pathogens and spoilage, which makes people unusually sensitive to spotting the wormlike tendrils, holes and off textures that AI food images tend to produce.
Why it matters
Restaurants, cafes and brands are increasingly turning to AI image generators instead of photographing their actual food, and the results frequently look nauseating rather than appetizing. This piece is useful less as a story about any single failure than as a plain explanation of why generative AI keeps stumbling on this particular category of image: it connects a visible, widely mocked phenomenon to concrete mechanics of how diffusion models work, giving readers a grounded way to understand a broader class of AI image failures beyond just food.
Who it affects
Directly, the restaurants, cafes and brands generating these images and the customers who see them online or in marketing material. More broadly, anyone relying on AI image generators for photorealistic product or food photography, since the same structural and training-data weaknesses described here apply beyond food to any subject involving thin, irregular shapes or strong physical expectations.
How to use it
The article points to two practical failure points worth avoiding: vague prompts, such as simply asking for 'a sandwich,' tend to leave too much for the model to guess, and text-oriented instructions like 'be precise' do not translate into better image results. It also flags that low-resolution AI images blown up well past their intended size will magnify artifacts or force the model to fill in gaps it cannot resolve well, so upscaling should be treated as a risk rather than a free enhancement.
How solid is it
The explanation draws on five academics from different institutions and disciplines, computer vision, behavioral science, digital culture studies, computer science and computational creativity, at Oxford, the University of Naples Federico II, the University of Zurich, King's College London and Queen Mary University of London, and their accounts converge on complementary mechanisms (diffusion structure, training data composition, human perception) rather than contradicting one another. No specific AI model or product is named or tested, and no statistics are given for how often these artifacts occur or how long the trend has been visible, so the piece is a qualitative, expert-sourced explainer rather than a measured study.
Risks and caveats
The 'model collapse' claim about training on AI-generated material is attributed to unnamed research rather than a specific cited paper, so its scope and strength cannot be checked from this piece alone. No timeframe is given for when AI food imagery became common or how widespread it now is, and no single generator or company is identified as responsible, so the explanations describe diffusion-based image generation broadly rather than any one product's behavior.
“Diffusion models are notoriously weak at generating thin, continuous, terminating structures. Noodles, strands, and tendrils are exactly the kind of geometry that trips this up, so you get spaghetti-like artifacts bleeding into places with no anatomical or culinary logic.”
— Giovanbattista Califano, behavioral scientist, University of Naples Federico II