The legal fight over books used to train large language models is splitting into two questions: whether learning from copyrighted text is lawful, and whether the company acquired its training copies legally. That distinction may shape both future lawsuits and how AI developers build datasets.
TechCrunch points to a ruling involving Anthropic in which Judge William Alsup treated model training as a transformative use comparable to a writer studying existing literature. The court nevertheless penalized the company for obtaining books from unauthorized “shadow libraries.” The result does not create a simple rule that all AI training is legal or illegal; the source and handling of each copy can matter independently from the training process.
For authors, that distinction is uncomfortable. A model may have learned from their work without permission and may compete with some forms of writing, yet copyright law does not automatically grant control over every form of analysis or learning. AI companies, meanwhile, cannot assume that a fair-use argument for training excuses piracy used to assemble a dataset.
More cases and appeals are likely to refine the boundaries. For now, the safest practical lesson for model developers is straightforward: document where training material came from, license it when required, and do not treat an argument about transformative use as permission to acquire unlawful copies.