AI-generated food images are increasingly appearing across restaurants, cafes and brand promotions, yet the results are frequently unappetising. The output includes shrimp resembling donuts, wormlike noodles, stringy chicken, and ice cream that looks more like construction material or brains than something edible. Many of these images are riddled with lumps, holes and unsettling patterns.
The reasons behind these distorted visuals span the technical workings of AI systems, the material used to train them, and the way humans perceive the results. Understanding these factors helps explain why so many promotional food images generated by artificial intelligence miss the mark so dramatically.
How Diffusion Models Create Distorted Food
Many leading image generators rely on a technique known as diffusion. These models begin with an image of pure noise, similar to a screen full of static, and gradually remove the noise to build the requested visual. As Chris Russell, a professor of AI, government and policy at the University of Oxford and an expert in computer vision, explained, this means coarse structures are recovered first, with fine texture details arriving at the end.
According to Russell, problems often emerge before those finer details are even added. A model may get the basic structure of an object wrong at an early stage, then apply vivid texture details on top of that flawed foundation. He compared this to the same kind of failure that produces a person with six fingers instead of five, a mechanism that could also account for shrimp shaped like donuts.
Why Noodles, Bubbles and Holes Go Wrong
Even when the underlying structure is sound, finer details can still fail. Giovanbattista Califano, a behavioural scientist who studies responses to AI-generated imagery at the University of Naples Federico II in Italy, noted that diffusion models are notoriously weak at generating thin, continuous, terminating structures. Noodles, strands and tendrils are exactly the kind of geometry that trips these systems up, producing spaghetti-like artefacts that bleed into places with no anatomical or culinary logic.
Once a model begins generating such shapes, it can struggle to determine where they should stop or what they should be attached to. Califano added that other repeating textures, such as bubbles and seeds, are similarly difficult for diffusion models to contain within sensible boundaries, often spilling into areas where they do not belong. This explains why so many AI food images appear relentlessly noodly, unsettlingly patterned, and clustered with holes that can trigger trypophobia.
A further complication is that the systems have no genuine understanding of the objects they create. Artificial intelligence does not know what a sandwich, a noodle or a burrito actually is, nor does it comprehend the objects it is generating. This lack of conceptual understanding sits at the root of many of the failures, compounding the technical limitations of how these images are assembled.
Source
Image: theverge.com