Google Wants AI Scraping to Be 'Fair Use.' Will That Fly in Court?

AI
(Image credit: Shutterstock)

What do you think would happen if I tried this? I stroll into a bank and see a wad of cash within arm’s reach behind an unoccupied teller window. I grab the dough and start walking out the door with it when a police officer, very rudely, stops me. “I’m entitled to take this money,” I say. “Because nobody at the bank told me not to.”

If you think my defense is implausible, then you don’t work for Google. This week, the search giant said that it wants to change copyright laws so that it can grab any content it wants from the Internet, use it as training data for its AI products, and argue “fair use” if anyone objects to the plagiarism stew Google’s cooking up. Google’s figleaf to copyright holders: they’ll find a way to let you opt-out.

Latest Videos FromTom's Hardware
Avram Piltch
Managing Editor: Special Projects

Avram Piltch is Managing Editor: Special Projects. When he's not playing with the latest gadgets at work or putting on VR helmets at trade shows, you'll find him rooting his phone, taking apart his PC, or coding plugins. With his technical knowledge and passion for testing, Avram developed many real-world benchmarks, including our laptop battery test.

  • kano1337
    I think gigantic companies like them intend to pay a relatively small sum if they get fined over AI-training-related things, but they do not inted to stop this kind of training. Maybe unless managers will be "collected" like they were in Full Tilt Poker's case (Poker's Black Friday), so some hard governmental actions will not be taken, unless the regulations will not take a more strict approach, the direction of the corporate world will not really change.
    So they have very good law specialists, likely their profit coming from the training will exceed the fines big time.

    Also, regulations-wise pulling the plug on AI, before small entities, or individuals get access to very powerful AI tools maybe even would be beneficial to the mega and giga companies, as real breakthroughs still would most likely be achieved "only" by them. But maybe it is already late, and too much of the genie is out of the bottle.

    All in all, I would like to see AI improving and going forward, as it is a very exciting field, but I would like to see more emphasis being put on more grounded researches, with less controversies like this, and less showboating type of usecases of AI.
    Reply
  • hotaru251
    Opt in should be optional, but opt out should be DEFAULT.

    scraping data is theft if you are for profit company.
    Reply
  • DSzymborski
    Scraping copyrighted content may be theft. A lot of this is largely untested.

    Data itself is trickier to argue as theft. It really depends. Facts themselves are not copyrightable under US law. Specific *presentations* of them may be.
    Reply
  • peachpuff
    It would be a shame if us plebs would trick ai with fake/wrong data just for the hell of it...
    Reply
  • vanadiel007
    This has already been tested in courts. Sony lawyers made the argument that possessing MP3 files is equal to stealing. Consumers were arguing that a copy is not equal to stealing because it's a copy. This is the exact same argument, but this time from a Company rather than a consumer, trying to make the argument that copying data into an AI database is not stealing but merely "scraping" ie replicating the original data.
    Reply
  • kjfatl
    The solution is simple. If the input is generated substantially using data acquired through 'fair use", all generated outputs must be labeled as available for 'fair use' by others and made available for others to use, not hidden behind a firewall.

    If we aren't careful, Google will effectively own everything.
    Reply
  • thisisaname

    What do you think would happen if I tried this? I stroll into a bank and see a wad of cash within arm’s reach behind an unoccupied teller window. I grab the dough and start walking out the door with it when a police officer, very rudely, stops me. “I’m entitled to take this money,” I say. “Because nobody at the bank told me not to.”

    More like going into a newsagents and taking photographs of the newspapers.
    Reply
  • bigdragon
    Google should have to live by the same fair use standards they put on their content creators. Given how ridiculous YouTube and search can be with filtering out content or complying with obviously bogus DMCA take-down requests, Google should also have to comb through their datasets to aggressively apply every complaint. It's only fair!
    Reply
  • DavidLejdar
    If i.e. Google would be smart, they would realize that they are digging a hole under their feet, if they seriously consider putting professional content creators on the side-line. The "AI-output" would quickly become repetitive, and out-dated, possibly even ending up digesting just what some propaganda network puts out.

    "AI as librarian" would certainly make more sense. I'd even go for "personal assistant". In example, I would ask PAI to compile a table containing links to all published works by Avram Pitch, containing "AI". Right now, I would have to do a search manually, including specific commands to not have the results be flooded with stuff I am not looking for specifically right now. And having PAI, which uses some personal data from me to actually improve my experience - personal data, such as what my preference for file format of the table is, which I can tell-it/customize - that would seem way cooler. And it would arguably be the actual next logical step in the development of the "web-search experience".

    And such experience may perhaps not have as much broad appeal, as a "mysterious oracle" does have, or not sound as much of a financial venue as creating a corporate environment, which users are tied to, where they have their daily schedule determined by an algorithm, and where they are eventually told to crush another billionaire's fiefdom. But it would feel more like Web 3.0, opening the door to more quality.

    In example, with the mentioned table, and some additions to it, it would save some time towards writing an article in an non-English language, which uses linked quotes, and which may help to create interest in web-stuff in overall. And such could create more traffic for various sites and services, who are then more likely to be able to afford full-time positions, which in turn could mean e.g. more news articles.

    Meanwhile, PAIs could help readers to filter what every reader is individually interested in, instead of the quest for an ultimate algorithm, which caters to everyone, while it doesn't necessarily go even beyond being driven by the web-usage of a relatively small group of power-users, who are not really doing anything but to sit in their basement day and night, while flooding the net with their output. Which is good for them, but there are also poorer citizen, who may not even have a smartphone to type on. So an algorithm assuming that the power-users are what the majority says and likes, that makes the online experience in some cases a quite estranged one.
    Reply
  • Leptir
    Google is beyond ruthless is preventing bots from scraping anything off google. But they give themselves the right to scrape whatever they want. Rank hypocrites!
    Reply