, , , , , , , , ,

AI video startup Runway reportedly trained on ‘thousands’ of YouTube videos without permission

AI company Runway reportedly scraped “thousands” of YouTube videos and pirated versions of copyrighted movies without permission. 404 Media obtained alleged internal spreadsheets suggesting the AI video-generating startup trained its Gen-3 model using YouTube content from channels like Disney, Netflix, Pixar and popular media outlets.

An alleged former Runway employee told the publication the company used the spreadsheet to flag lists of videos it wanted in its database. It would then download them without detection using open-source proxy software to cover its tracks. One sheet lists simple keywords like astronaut, fairy and rainbow, with footnotes indicating whether the company had found corresponding high-quality videos to train on. For example, the term “superhero” includes a note reading, “Lots of movie clips.” (Indeed.)

Other notes show Runway flagged YouTube channels for Unreal Engine, filmmaker Josh Neuman and a Call of Duty fan page as good sources for “high movement” training videos.

“The channels in that spreadsheet were a company-wide effort to find good quality videos to build the model with,” the former employee told 404 Media. “This was then used as input to a massive web crawler which downloaded all the videos from all those channels, using proxies to avoid getting blocked by Google.”

Screnshot of the Runway AI homepad.
Runway

A list of nearly 4,000 YouTube channels, compiled in one of the spreadsheets, flagged “recommended channels” from CBS New York, AMC Theaters, Pixar, Disney Plus, Disney CD and the Monterey Bay Aquarium. (Because no AI model is complete without otters.)

In addition, Runway reportedly compiled a separate list of videos from piracy sites. A spreadsheet titled “Non-YouTube Source” includes 14 links to sources like an unauthorized online archive of Studio Ghibli films, anime and movie piracy sites, a fan site displaying Xbox game videos and the animated streaming site kisscartoon.sh.

In what could be viewed as a damning confirmation that the company used the training data, 404 Media found that prompting the video generator with the names of popular YouTubers listed in the spreadsheet spit out results bearing an uncanny resemblance. Crucially, entering the same names in Runway’s older Gen-2 model — trained before the alleged data in the spreadsheets — generated “unrelated” results like generic men in suits. Additionally, after the publication contacted Runway asking about the YouTubers’ likenesses appearing in results, the AI tool stopped generating them altogether.

“I hope that by sharing this information, people will have a better understanding of the scale of these companies and what they’re doing to make ‘cool’ videos,” the former employee told 404 Media.

When contacted for comment, a YouTube representative pointed Engadget to an interview its CEO Neal Mohan gave to Bloomberg in April. In that interview, Mohan described training on its videos as a “clear violation” of its terms. “Our previous comments on this still stand,” YouTube spokesperson Jack Mason wrote to Engadget.

Runway did not respond to a request for commeInt by the time of publication.

At least some AI companies appear to be in a race to normalize their tools and establish market leadership before users — and courts — catch onto how their sausage was made. Training with permission through licensed deals is one thing, and that’s another tactic companies like OpenAI have recently adopted. But it’s a much sketchier (if not illegal) proposition to treat the entire internet — copyrighted material and all — as up for grabs in a breakneck race for profit and dominance.

404 Media’s excellent reporting is worth a read.

This article originally appeared on Engadget at https://www.engadget.com/ai-video-startup-runway-reportedly-trained-on-thousands-of-youtube-videos-without-permission-182314160.html?src=rss

https://www.engadget.com/ai-video-startup-runway-reportedly-trained-on-thousands-of-youtube-videos-without-permission-182314160.html?src=rss


Leave a Reply

Your email address will not be published. Required fields are marked *

July 2024
M T W T F S S
1234567
891011121314
15161718192021
22232425262728
293031  

About Us

Welcome to encircle News! We are a cutting-edge technology news company that is dedicated to bringing you the latest and greatest in everything tech. From automobiles to drones, software to hardware, we’ve got you covered.

At encircle News, we believe that technology is more than just a tool, it’s a way of life. And we’re here to help you stay on top of all the latest trends and developments in this ever-evolving field. We know that technology is constantly changing, and that can be overwhelming, but we’re here to make it easy for you to keep up.

We’re a team of tech enthusiasts who are passionate about everything tech and love to share our knowledge with others. We believe that technology should be accessible to everyone, and we’re here to make sure it is. Our mission is to provide you with fun, engaging, and informative content that helps you to understand and embrace the latest technologies.

From the newest cars on the road to the latest drones taking to the skies, we’ve got you covered. We also dive deep into the world of software and hardware, bringing you the latest updates on everything from operating systems to processors.

So whether you’re a tech enthusiast, a business professional, or just someone who wants to stay up-to-date on the latest advancements in technology, encircle News is the place for you. Join us on this exciting journey and be a part of shaping the future.

Podcasts

TWiT 989: Executive Laundry Folding Disorder – Crowdstrike, Prime Day, Stremaing the Olympics This Week in Tech (Audio)

Biden drops out of the race via tweet "CrowdStrike update that caused global outage likely skipped checks, experts say Amazon's Prime Day causes worker injuries, Senate probe finds Kaspersky Lab Closing U.S. Division; Laying Off Workers Researchers: Weak Security Defaults Enabled Squarespace Domains Hijacks The Paris Olympics Will Show Us the Future of Sports on TV Judge dismisses much of SEC suit against SolarWinds over cybersecurity disclosures FBI Used New Cellebrite Software to Access Trump Shooter's Phone Hundreds more Californians get housing with Apple's help Microsoft's new AI system 'SpreadsheetLLM' unlocks insights from spreadsheets, boosting enterprise productivity Bethesda Game Studios workers unionize under Communications Workers of America Host: Leo Laporte Guests: Lisa Schmeiser, Ashley Esqueda, and Anthony Ha Download or subscribe to this show at https://twit.tv/shows/this-week-in-tech Get episodes ad-free with Club TWiT at https://twit.tv/clubtwit Sponsors: canary.tools/twit – use code: TWIT lookout.com mintmobile.com/twit NetSuite.com/TWIT motific.ai
  1. TWiT 989: Executive Laundry Folding Disorder – Crowdstrike, Prime Day, Stremaing the Olympics
  2. TWiT 988: Flaming Corn Maze – AT&T Breach, Galaxy Z Fold6, Olympic Disinfo
  3. TWiT 987: Often Plagiarized, Never Equalled – Sapce Junk, Threads Hits 175M Users, AIndependence
  4. TWiT 986: Our Dope GPS! – Supreme Court Decisions, Snapdragon X Elite Tests
  5. TWiT 985: TikTok With Wings – AT&T Landlines, US Bans Kaspersky and DJI