NEW YORK / RankWire.AI / – Hachette Book Group, Cengage Learning, and Elsevier have initiated legal action against Google concerning its Gemini artificial intelligence platform. Author Scott Turow and his organization, S.C.R.I.B.E., have joined the proposed class action. The lawsuit was filed on July 10 in the U.S. District Court for the Southern District of New York. The plaintiffs accuse Google of copying millions of copyrighted books and journal articles without authorization during the development and training of Gemini. As of July 15, the court had yet to rule on the allegations or certify the class.

The complaint states that Google sourced material via Google Books, Google Play Books, and Google Scholar. Publishers and authors provided works for specific purposes, including search functions, sales, and research access. The plaintiffs argue that these arrangements did not permit broader commercial AI training. They also allege Google downloaded extensive web-scraped datasets containing copyrighted content. The filing claims some of this material originated from known pirate sources and paywalled services.
The 57-page complaint outlines four claims under federal law. Three relate to alleged reproduction through Google services, web scraping, and Gemini’s development or training. The fourth invokes the Digital Millennium Copyright Act. The plaintiffs allege Google removed or altered copyright management information from training data. The filing also references internal discussions about utilizing publisher-supplied books. One assessment estimated potential fines between $10 billion and $100 billion. The court has not yet evaluated these claims.
Proposed class includes registered works
The proposed class encompasses owners of registered U.S. copyrights for qualifying books and journal articles. Eligible books must have an International Standard Book Number (ISBN). Eligible articles require a Digital Object Identifier or International Standard Serial Number. The class definition covers works allegedly copied from Google services or downloaded through web scraping. It also includes works supposedly reproduced during Gemini’s development or training.
Registration timing also dictates class eligibility. One criterion requires registration within five years of publication and prior to Google’s alleged reproduction or distribution. Another mandates registration within three months of publication. The lawsuit excludes government agencies, Google affiliates, certain court participants, and individuals who properly opt out. The court must approve the class designation before the case proceeds to represent the broader group.
Complaint requests damages and an accounting
The plaintiffs seek statutory damages or actual damages for proven infringements. They also request Google’s profits attributable to any confirmed copyright violations. Their remedies include an injunction, legal costs, and a jury trial. The complaint does not specify a total damages amount but asks Google to disclose Gemini training data, collection methods, and known capabilities through a court-ordered accounting.
This accounting would identify copyrighted works used for Gemini’s training and detail how Google collected, copied, processed, and encoded these materials. The plaintiffs also request the court to oversee the destruction of unauthorized copies under Google’s control. Earlier, Hachette and Cengage sought to join separate AI litigation against Google in California. This New York case expands the scope to include Elsevier, Turow, and S.C.R.I.B.E., with claims related to Google services, web scraping, and Gemini training.
