AI Lie: Machines Don’t Learn Like Humans (And Don’t Have the Right To)

AI
(Image credit: Shutterstock)

Some AI doomers want you to believe that sentient computers are close to destroying humanity. But the most dangerous aspect of chatbots such as ChatGPT and Google SGE today is not their ability to produce robot assassins with German accents, scary though that sounds. It's their unlicensed use of copyrighted text and images as “training data,” masquerading as "learning" -- which leaves human writers and artists competing against computers using their own words and ideas to put them out of business. And it risks breaking the open web marketplace of ideas that’s existed for nearly 30 years.

There’s nothing wrong with LLMs as a technology. We’re testing a chatbot on Tom’s Hardware right now that draws training data directly from our original articles; it uses that content to answer reader questions based exclusively on our expertise.

Latest Videos FromTom's Hardware
Avram Piltch
Managing Editor: Special Projects

Avram Piltch is Managing Editor: Special Projects. When he's not playing with the latest gadgets at work or putting on VR helmets at trade shows, you'll find him rooting his phone, taking apart his PC, or coding plugins. With his technical knowledge and passion for testing, Avram developed many real-world benchmarks, including our laptop battery test.

  • BX4096
    No offense to the author, but this indiscriminate assault on LLMs (what do math errors have to do with copyright, anyway?) reads less like an valid viewpoint and more like irrational neo-luddism from someone too concerned about losing their job to a new tech to stay objective. Not that I can't relate – my job is under threat as well – or there aren't any valid issues to be had with the technology, but I'm not going to abandon all logic and reason simply because my emotions tell me to.

    The article is too long to go point by point, but in a nutshell, the author doesn't seem to realize (or rather, willing to admit) that the more content these LLMs consume, the less recognizable any of it will be in regard to particular sources. And man, if you're going to put random quotes around stuff you don't like, you should at least put them around the "unlicensed" part in your "'training data' masquerading as 'learning'" bit. Aside from the fact that the training LLMs seems to fall under fair use under US copyright law, there's nothing there to copyright since these models do not "copy" work in the traditional sense and merely digest it to predict outcomes. Using a large number of patterns to predict the best next word in a sequence is a long cry from reproducing a copyrighted work without permission, which is why none of the major corporations seem concerned about the wave of legal challenges this produced.

    If anything, you should put blame on human masses for being so thoroughly unoriginal that you can train a mindless machine to believably mimic their verbal diarrhea on the fly. A smaller web with less redundancy and regurgitation and more accessible answers and solutions? The horror! Now, where do I sign up?
    Reply
  • setx
    BX4096 said:
    Using a large number of patterns to predict the best next word in a sequence is a long cry from reproducing a copyrighted work without permission
    Indeed, for example compression algorithms do exactly that and no one has any problem with that.

    BX4096 said:
    If anything, you should put blame on human masses for being so thoroughly unoriginal that you can train a mindless machine to believably mimic their verbal diarrhea on the fly.
    It's not that humans are that unoriginal, it's that some humans are used to put minimal effort in their job but now the same (minimal) quality can be produced by machines cheaper.
    Reply
  • JarredWaltonGPU
    BX4096 said:
    A smaller web with less redundancy and regurgitation and more accessible answers and solutions? The horror! Now, where do I sign up?
    This is assuredly NOT where this will all lead. If anything, it's going to be more content, generated by AI, creating billions of meaningless drivel web pages that are grammatically correct but factually highly questionable. Sure, you can go straight to whatever LLM tool you want (ChatGPT, Bard, Sydney, etc.) and generate that content yourself. It won't be any better, and potentially not any worse, than the myriad web pages that did the same thing but created a website from it. 🤷‍♂️

    I'm not personally that worried for my job, because so far no one is actually figuring out a way to have AI swap graphics cards, run hundreds of benchmarks, create graphs from the data, and analyze the results. But there will be various "best graphics cards" pages that 100% crib from my work and the work of other people to create a page that offers similar information with zero testing. Smaller sites will inevitably suffer a lot.
    Reply
  • hotaru251
    JarredWaltonGPU said:
    because so far no one is actually figuring out a way to have AI swap graphics cards, run hundreds of benchmarks, create graphs from the data, and analyze the results.
    i mean swapping gpu from a system can be done with an assembly style arm & positioning data of where it goes & rest can be done with scripts for the ai to activate.

    not really "not figured out" and more nobody cares enough to as cost too high.

    and on context of copyright to train LLM....music & film industry made that pretty clear decades ago.

    You ask for permission to use it and if you dont get a yes you can't. (and yes ppl can use content by modifying it, but the llm/ai is learning off the raw content which it cant change so couldnt use that bypass)
    Reply
  • apiltch
    BX4096 said:
    No offense to the author, but this indiscriminate assault on LLMs (what do math errors have to do with copyright, anyway?) reads less like an valid viewpoint and more like irrational neo-luddism from someone too concerned about losing their job to a new tech to stay objective. Not that I can't relate – my job is under threat as well – or there aren't any valid issues to be had with the technology, but I'm not going to abandon all logic and reason simply because my emotions tell me to.
    The bots cannot exist without the training data; ergo, one might say that the entire bot is a derivative work. These issues are still being litigated in court and we don't know exactly what the decisions will be in terms of law and the decision might be different in one country than another.

    The IP and tech worlds have never seen anything like LLMs so what they are doing is really unprecedented. The premise that these bots have the legal right to hoover up the data is based on the philosophical view that they are "learning like a person does." In the article, I quoted a couple of people who have literally said that the bots should be thought of as "students" or a writer who is "inspired by" the work they read.

    What I'm trying to argue -- and we'll see how the legal cases pan out -- is that morally, philosophically and technically, LLMs do not learn like people and do not think like people. We should not think of the training process the same way we think of a person learning, because that's not what is happening. It's more akin to one service scraping data from another.

    Is scraping all publicly-accessible data from a website legal? That has definitely not been established. LinkedIn has won a lawsuit against a company that scraped its user profiles, which are all visible on the public web (https://www.socialmediatoday.com/news/LinkedIn-Wins-Latest-Court-Battle-Against-Data-Scraping/635938/). As a society, we recognize that just because a human can view it that doesn't mean a machine has the right to record and copy it at scale.

    Now I know someone will say "this isn't copying." In order to create the LLM, they are taking the content and ingesting it and using it to create a database of tokens. The company that scraped LinkedIn wasn't word-for-word reproducing LinkedIn content either; it was taking their asset and repurposing it for profit. What about Clearview AI, which scraped billions of photos from Facebook (again, pages that were visible on the public Internet)?
    Reply
  • BX4096
    apiltch said:
    What I'm trying to argue -- and we'll see how the legal cases pan out -- is that morally, philosophically and technically, LLMs do not learn like people and do not think like people. We should not think of the training process the same way we think of a person learning, because that's not what is happening. It's more akin to one service scraping data from another...
    I mean, just because a search engine doesn't process and index information the way a human does, doesn't mean that we should limit their access to all types of searchable data and force them to ask explicit permission for every single public result found. We had the same types of arguments when Google rose to power, and we all know how that particular battle played out.

    So yes, we can all mourn the countless librarian jobs (not to mention, our privacy) lost to the likes of Google and Amazon, but let's not pretend that any of us want to go back to the reality of spending an entire afternoon on simply looking up a fact or finding a book in a library a few miles away from your house. Machine training, or "learning", is the inevitable future, and that future is already here. Just take one look at the recent Photoshop 2024 release to realize that things are never going back to what we considered normal. The way I see it, content creators only have two viable choices: adapt or perish, which is the harsh reality of all things obsolete. As for new low-effort, regurgitated content popping up all over the web mentioned by Jarred, we already had that for the past decade or two. Same s***, different toolset, if you ask me.
    Reply
  • JarredWaltonGPU
    BX4096 said:
    I mean, just because a search engine doesn't process and index information the way a human does, doesn't mean that we should limit their access to all types of searchable data and force them to ask explicit permission for every single public result found. We had the same types of arguments when Google rose to power, and we all know how that particular battle played out.
    AI / LLM generation of content with no attribution is nothing like search. It's more akin to Google/Bing/etc. taking the best search results, ripping out all of the content, and presenting a summary page that never links to anyone else (except maybe Amazon and such with affiliate links). Search works because it drives traffic to the websites, and at most there's a small snippet from the linked article. That's "fair use" and it benefits both parties. LLM training only benefits one side, for profit, and that's a huge issue.
    Reply
  • BX4096
    JarredWaltonGPU said:
    AI / LLM generation of content with no attribution is nothing like search. It's more akin to Google/Bing/etc. taking the best search results, ripping out all of the content, and presenting a summary page that never links to anyone else (except maybe Amazon and such with affiliate links). Search works because it drives traffic to the websites, and at most there's a small snippet from the linked article. That's "fair use" and it benefits both parties. LLM training only benefits one side, for profit, and that's a huge issue.
    Not quite true, either. It may be a small part of it that relates to quoting very specific facts, like individual benchmark numbers that may be linked to a particular source, but the rest of it is too generic to be viewed this way. Writing a generic paragraph in the style of Steven King – loosely speaking –is not ripping off Steven King, and neither is quoting 99.9% facts known to humans. To whom should it attribute the idea that Paris is the capital of France, for example? Because the vast majority of the queries are just like that. Moreover, considering the way LLMs digest information to produce their content, I frankly can't even imagine how they could properly source it all. You say they're nothing like humans, but it's a lot like trying to attribute an idea that just popped up in your head that is based on multiple sources and events in your life going decades back. It seems inconceivable.

    It appears to me see that you're looking at this from the niche position of your website and not broadly enough to see the big picture. Even when asking a related question like "what's the best gamer GPU on the market right now?", you won't get a rephrased summary from Tom's latest GPU so-and-so. What you'll get is likely an aggregate reply that stems from dozens of websites like yours (well, that including more inferior ones), and probably an old one at that – "RTX 3090!" – since the training data for these things can lag quite a bit. The doctrine of fair use also includes the idea of "transformative purpose", which is exactly what is going on with probably at least 99.99% of ML generated content already.

    Again, it's not like I don't empathize with any of your concerns, but considering the mind-blowing opportunities presented by this new technology, I can't help but see these objections as beating a dead horse at best and standing in the way of inevitable progress at worst. I wish we could all focus on the truly amazing stuff it can and will soon do, rather than fixate on what seems to be futile doomsayer negativity that - at least in my view - will almost certainly lose in court and therefore ultimately cause the very people affected even more stress and financial troubles.I could be wrong, but that's my current take on this.
    Reply
  • Co BIY
    A legal way of flagging a websight to specifically exclude AI scaping seems reasonable to me.

    More importantly in my opinion is requiring all AI generated content to be marked as AI generated.
    Reply
  • coolitic
    While I consider most AI-generated content to be "grammatically-correct garbage", and definitely don't see AI as thinking as people do, I also don't really care for scraping to be outlawed, and I don't care if "search engines are ok because they benefit the website". Avram is certainly correct in analyzing how AI objectively works and what it is capable of producing (like a person who understands the worth of their tool); but also yeah, with all these articles, I do feel a bit of protectionism and defensiveness for his field coming out from them.

    Frankly, even if it wasn't outlawed, the fact that there's a lot of people willing to "poison" AI content (which is easy to do) for either serious reasons or the lulz, means that the legality won't be the be-all and end-all. And perhaps more importantly, it seems that LLMs are running out of "fresh" data and collapsing on themselves, both over the excessive data that they do have, and the cannibalization that stems from "Model Autophagy Disorder". I do genuinely think that LLMs are being pushed to do broader things than they are really capable of, and both corporations and researchers in the field refuse to admit the limitations because they'd lose funding. Just remember folks, we've had big AI hype before, as early as in the 70s IIRC, and it moreorless went down in the same cycle.
    Reply