A company that publicly offered to scan physical books for AI training data has scrubbed that service from its website, claiming the offering never actually existed.
The unnamed firm had advertised the capability to convert printed books into digital formats for machine learning purposes. This sparked backlash from publishers and authors concerned about unauthorized use of copyrighted material in AI model training. The company's sudden reversal suggests the service faced legal or commercial pressure.
Book scanning for AI training has become a flashpoint in the broader debate over AI data sourcing. Major language models like ChatGPT relied on internet text and potentially copyrighted works for training. Authors including John Grisham and others have filed lawsuits against OpenAI and other AI companies for alleged copyright infringement. Publishing houses fear widespread digital scanning of their catalogs could undermine sales and author compensation.
The company's statement that "no such service was ever brought to life" attempts damage control but raises questions about why the offering appeared on their site in the first place. This mirrors similar backpedaling from AI firms facing public outcry over data practices. The move reflects growing legal and reputational risks companies face when handling copyrighted material for machine learning.
The incident underscores mounting pressure on AI developers to address copyright concerns. Publishers and authors increasingly demand transparency about training data sources and compensation frameworks. Regulators worldwide are also scrutinizing AI companies' data harvesting practices. Whether through formal licensing agreements or policy changes, the industry faces mounting expectations to respect intellectual property rights rather than exploit them for training efficiency.
This episode demonstrates that public perception and legal exposure can force AI companies to abandon controversial practices quickly, even after announcing them publicly.
