Your support is crucial to Anoraker. If you buy something through our links, we might earn an affiliate commission at no extra cost to you. Learn More

How Apple is Stealing YouTube Video Content to Train AI, and Why It Does Matter

How Apple is Stealing YouTube Video Content to Train AI, and Why It Does Matter

What This Video Covers

  • What exactly was taken from more than 170,000 YouTube videos, and how was that data used for AI training?
  • Why creator-written scripts, subtitles, and transcripts raise different questions than simply scraping random public webpages.
  • How the idea that anything publicly accessible online is automatically fair game starts to break down when creative work is involved.
  • What this kind of AI training means for creators who spend real time and money producing the material behind their videos.

Advertisement

Recently, Wired reported that Apple, Salesforce, Nvidia, and more are using "stolen" data from YouTube to train AI. The data consists of subtitle transcripts scraped from over 170,000 YouTube videos across some of the biggest names.

In some cases, like MKBHD (https://www.youtube.com/shorts/xiJMjTnlxg4) those transcriptions are hand-created at a cost. But even when they aren't, they're formed off of scripts written by creators.

But if it's on the internet, it's freely available for any use right? Let's get into that thought, because I disagree.