- By JeffkomStory Team
- Published on
Adobe Faces Class-Action Lawsuit Over Alleged Use of Pirated Books in AI Training
Adobe, like many major tech companies, has aggressively embraced artificial intelligence in recent years. From AI-powered design tools to generative media platforms, the company has positioned itself as a leader in creative AI. But that push may now come at a legal cost.
A newly proposed class-action lawsuit accuses Adobe of using pirated books to train one of its AI language models. And raising fresh concerns about how AI systems are built and whose work is being used behind the scenes.
What the Lawsuit Claims
According to the lawsuit, Adobe trained its SlimLM language model using copyrighted books, including Lyon’s own works.
SlimLM is described by Adobe as a lightweight language model designed for document assistance tasks, particularly on mobile devices. According to Adobe, the model was pre-trained using SlimPajama-627B, an open-source dataset released by AI chipmaker Cerebras in June 2023.
However, Lyon’s lawsuit argues that SlimPajama itself is derived from another dataset called RedPajama. Which allegedly includes a controversial collection of pirated books known as Books3.
The Problem With Books3 and RedPajama
Books3 is a massive dataset containing around 191,000 books. It has become a recurring point of legal conflict in the AI industry, as many authors claim their copyrighted works were included without consent, credit, or compensation.
According to the lawsuit, SlimPajama was created by copying and modifying RedPajama, which in turn includes Books3. Because of this chain, the lawsuit argues that Adobe indirectly used copyrighted material when training SlimLM.
“The SlimPajama dataset was created by copying and manipulating the RedPajama dataset,” the complaint states, “and therefore contains the Books3 dataset, including the copyrighted works of the Plaintiff and Class members.”
Adobe Is Not Alone
Adobe is far from the only company facing these accusations.
-
Apple was sued in September over claims that its Apple Intelligence models were trained on copyrighted material without permission.
-
Salesforce faced a similar lawsuit in October, also tied to the RedPajama dataset.
-
Anthropic, the company behind Claude AI, agreed to pay $1.5 billion to authors earlier this year to settle claims that it used pirated books for training.
These cases highlight a growing legal backlash against how generative AI models are trained.
A Bigger Issue for the AI Industry
At the heart of these lawsuits is a fundamental question:
Can AI companies legally train models on copyrighted material without permission?
AI systems require enormous amounts of data to function effectively. But as authors, artists, and publishers push back, courts are increasingly being asked to define the boundaries between innovation and intellectual property rights.
The Anthropic settlement was seen by many as a potential turning point—suggesting that using pirated or unauthorized content may no longer be legally or financially sustainable.
What Comes Next for Adobe and AI Companies
Adobe has not yet publicly resolved the claims, and the lawsuit is still in its early stages. But the case adds to mounting pressure on AI developers to be more transparent about training data and to establish clearer licensing practices.
As generative AI continues to expand, these legal battles may shape the future of how models are built—and who gets paid for the knowledge that powers them.
One thing is clear: the era of “train first, ask later” may be coming to an end.
Here are some related articles you may find interesting:
OpenAI Adds AI Doomer Paul Christiano to Its Board of Directors
OpenAI has added AI researcher Paul Christiano to the board of the OpenAI Foundation. His appointment...
Travis Kalanick’s Atoms May Enter the Robotaxi Business
Travis Kalanick’s Atoms is reportedly exploring the robotaxi and autonomous vehicle market, signaling...
Startup ARR Is Less Secure Than Ever as AI Changes Enterprise Buying
Startup ARR Faces a New Challenge
Annual recurring revenue (ARR) has long been one of the most important...
Fashion Startup Atorie Raises $9.5M to Make Luxury Goods More Affordable
Fashion is changing fast. Consumers want better quality without paying extremely high luxury markups....
Generalist Robotics Startup Reaches $3 Billion Valuation After Fresh Funding
The robotics industry is gaining momentum as investors continue to back startups building AI systems...
How AI Accounting Startup Rillet Raised $100M and Became a Unicorn in 48 Hours
The AI accounting industry is gaining momentum as businesses look for smarter ways to manage finance...
Every Fusion Startup That Has Raised Over $100 Million
Fusion energy is moving from a long-term scientific ambition toward a potentially transformative source...
Cognition Eyes $40 Billion Valuation as AI Coding Agent Devin Gains Enterprise Traction
Cognition Reportedly Targets $40 Billion Valuation
AI coding startup Cognition, the company behind AI...
Naïve Raises $28.5M to Automate Business Setup and AI Operations
Naïve, an AI infrastructure startup, has raised $28.5 million in Series A funding to help developers...
Trump DOJ to Oversee OpenAI Green Card Hiring Practices After $3.2 Million Settlement
OpenAI Agrees to DOJ Oversight Over PERM Hiring Process
OpenAI and its former subsidiary, Statsig, have...
Popular Posts

OpenAI Adds AI Doomer Paul Christiano to Its Board of Directors
JeffkomStory Team
OpenAI has added AI researcher

Travis Kalanick’s Atoms May Enter the Robotaxi Business
JeffkomStory Team
Travis Kalanick’s Atoms is reportedly
Startup ARR Is Less Secure Than Ever as AI Changes Enterprise Buying
JeffkomStory Team
Startup ARR Faces a New

Fashion Startup Atorie Raises $9.5M to Make Luxury Goods More Affordable
JeffkomStory Team
Fashion is changing fast. Consumers
Join Our Newsletter
Start your day with impactful startup stories and concise news! All delivered in a quick five-minute read in your inbox.