Generative AI Goes 'MAD' When Trained on AI-Created Data Over Five Times

MAD
(Image credit: Rice University)

A new study on AI has found an inherent limitation on current-generation networks such as the ones employed by ChatGPT and Midjourney. It seems that AI networks trained on AI outputs (like the text created by ChatGPT or the image output created by a Stable Diffusion model) tend to go "MAD" after five training cycles with AI-generated data. As you can see in the above images, the result is oddly mutated outputs that aren't reflective of reality.

MAD - short for Model Autophagy Disorder — is the acronym used by the Rice and Stanford University researchers involved in the study to describe how AI models, and their output quality, collapses when repeatedly trained on AI-generated data. As the name implies, the model essentially "eats itself," not unlike the Ouroboros of myth. It loses information on the tails (the extremes) of the original data distribution, and starts outputting results that are more aligned with the mean representation of data, much like the snake devouring its own tail.

Latest Videos FromTom's Hardware
Francisco Pires
Freelance News Writer

Francisco Pires is a freelance news writer for Tom's Hardware with a soft side for quantum computing.

  • The Historical Fidelity
    Idk if this is inappropriate for comments so if it is please let me know.

    But the first thing that came to mind is that we found out that AI to AI intimate contact can lead to contracting Artificial Syphilis (Syphilis is the STD that makes humans go mad in the brain if left untreated)
    Reply
  • sepuko
    That means you can poison internet content with image previews and data that appears has the traits of AI generated one so AIs get MAD when crawling. So now AI owners will need to start validating the data the AI trains on. What worries me is that we hold them to no standard when it comes to the selection they feed the AI when training it and that trustworthiness of the output is very dubious, image and code generation aside.
    Reply
  • peachpuff
    Did someone hurt ai's feelings? Boohoo...
    Reply
  • HonkingAntelope
    This is by far the number one issue all future LLMs will run into. And that is the lack of "clean" data for training. More so, when you have an arms race, it's very possible for an opponent to introduce subtle poison into the data in order to manipulate the training outcome. Especially if an LLM is being trained essentially by just crawling the internet on a 10Gbps link
    Reply
  • Giroro
    " ...starts outputting results that are more aligned with the mean representation of data, much like the snake devouring its own tail."

    ... I never knew that the ouroboros was known for outputting results that align with the mean representation of data. I thought it represented infiinity.
    Reply
  • Giroro
    The Historical Fidelity said:
    Idk if this is inappropriate for comments so if it is please let me know.

    But the first thing that came to mind is that we found out that AI to AI intimate contact can lead to contracting Artificial Syphilis (Syphilis is the STD that makes humans go mad in the brain if left untreated)

    More like inbreeding.

    And we are already on *at least* the second generation of AI training AI. The overwhelming majority of online text content has been AI generated for at least a few years. Pretty much since content farms realized the only way to SEO for google listings was to grab "common knowledge" information (which is always unsourced and often incorrect) then repeatedly restate and repackage those bullet points into different ways people might ask full questions into a smart speaker or digital assistant. This destroys the nuance of the information, to the point it quickly loses meaning and eventually becomes wrong and contradictory.
    It's why you can't find the info you're looking for on google anymore. It's also why Bard is essentially unusable, and already showing signs of getting worse.
    Reply
  • Kamen Rider Blade
    So would this be a case of "Mad AI" Disease?
    Reply
  • domih
    Duh, the scientists at Hollywood predicted the issue back in 1996: https://en.wikipedia.org/wiki/Multiplicity_(film).
    Reply
  • bit_user
    It's not surprising, really. Your model is only as good as your data.

    I do find it interesting that, in the sample images, it seems to have latched onto JPEG-style image compression artifacts and treated them as if they're part of the underlying object.

    But, one thing about the article @Francisco Alexandre Pires : the primary focus seems to be on image generators, and yet you seem to conflate them with LLMs:
    "In essence, training an LLM on its own (or anothers') outputs creates a convergence effect on the data that composes the LLM itself. "
    Yes, the phenomenon was found to apply to LLMs, but let's be clear: image generators aren't LLMs. For instance, Stable Diffusion is a latent diffusion model, as its name implies. The paper actually used an image generator called StyleGAN-2, which is a generative adversarial network.

    Please take care not to treat LLM as a short-hand for advanced neural networks, except when it actually applies.
    Reply
  • bit_user
    Kamen Rider Blade said:
    So would this be a case of "Mad AI" Disease?
    Mad Cow Disease is caused by a specific defective protein (AKA prion), which interferes with protein synthesis in a way that causes yet more of them to be produced. It got propagated through the cattle industry feeding slaughterhouse scraps to other cows. Even cooking the scraps isn't enough to breakdown enough of the prions.

    So, in a way, it's not completely off the mark. However, mere cow cannibalism isn't enough. You need to introduce the defective protein into the cycle, at some point.
    Reply