Startups & Business

The Perplexity lawsuit and the future of AI licensing

Legal battles involving Perplexity, Anthropic, and the New York Times highlight the growing tension between AI developers and publishers. A $1.5 billion Anthropic settlement sets a major financial benchmark for how companies must handle unlicensed datasets.

The Perplexity lawsuit and the future of AI licensing

Dow Jones and the New York Post filed an amended complaint against Perplexity in the Southern District of New York. The plaintiffs, which include the parent company News Corporation, allege that Perplexity engages in copyright infringement and trademark dilution. They claim the AI scours the internet to gather information and compiles it into answers that allow users to "Skip the Links." This architecture, known as retrieval-augmented generation, interacts with large language models to optimize accuracy. The plaintiffs argue this process allows users to bypass publisher websites, which harms advertising and subscription revenue. Perplexity responded to these claims by stating that news organizations want to control publicly reported facts. The lawsuit seeks an injunction to stop the alleged scraping and the destruction of databases that contain the copyrighted works.

The Anthropic Settlement and the Piracy Penalty

The Anthropic settlement changed how companies view the risk of using unlicensed datasets. In September 2025, a court granted preliminary approval to a $1.5 billion settlement between Anthropic and authors. This case involved allegations that Anthropic used pirated books to train Claude. The judge in the Bartz v. Anthropic case ruled that training on lawfully acquired books is a transformative fair use. However, the court found that using pirated copies is not fair use. The settlement of $1.5 billion for authors after the court ruled that training on pirated books is not fair use provides a clear financial benchmark for how much companies must pay to avoid massive litigation costs regarding stolen datasets. This distinction between lawful and pirated material is a main factor in current litigation.

The legal distinction between training on lawful and pirated data creates a clear divide in risk profiles. Companies training models on materials acquired through lawful channels face different challenges than those using shadow libraries. In the Bartz case, the court separated the acquisition of the data from the training process itself. This means a company might win on the transformative nature of its training but lose on the method of data collection. This reality forces AI developers to evaluate their training pipelines with extreme care.

The New York Times and the Regurgitation Problem

The New York Times case against OpenAI and Microsoft focuses on the issue of content regurgitation. The Times alleges that OpenAI trained ChatGPT on its copyrighted articles and that the model reproduces verbatim excerpts. This differs from the way search engines show snippets of information. In January 2026, Judge Sidney Stein affirmed an order for OpenAI to produce 20 million anonymized ChatGPT logs. The Times also seeks to prove that the model memorizes substantial portions of its articles. The company has also sought revenue through other channels to protect its subscription business. In May 2025, the Times signed a licensing deal with Amazon worth an estimated $20 million to $25 million per year. This agreement covers the Times, NYT Cooking, and The Athletic for use in Amazon products like Alexa and Rufus.

The litigation moves through discovery and tests whether the output of an AI constitutes an infringing work. The Times argues that the ability of a chatbot to reproduce its content removes the incentive for users to visit the original source. This conflict between AI output and publisher traffic is the core of the legal battle. You should observe the distinction between a search engine that directs traffic and an answer engine that replaces it. Will the New York Times find a way to reclaim its visibility in answer engines without abandoning its legal fight against OpenAI?

Non-Generative AI and the Thomson Reuters Precedent

The Thomson Reuters v. Ross Intelligence case provides a different outcome for non-generative AI tools. In February 2025, a court in Delaware rejected the fair use defense for Ross Intelligence. The court found that using Westlaw headnotes to train an AI legal research tool was not transformative. It also determined that the tool served as a market substitute and caused market harm. The court found that the 2,243 headnotes at issue matched the original content so closely that no reasonable jury could find otherwise. This case is currently on appeal before the Third Circuit.

The ruling in the Thomson Reuters case emphasizes that market harm is a significant factor in fair use analysis. The court did not view the copying as a technological input that is removed from the commercial value of the original work. Instead, the court saw the AI tool as a direct threat to an existing product and to a licensing market. This decision shows that the legal treatment of AI depends heavily on whether the tool is generative or a specialized search tool. The distinction between these two types of technology changes the fair use equation.

The Market Dilution Theory in the Meta Case

The Kadrey v. Meta decision introduced the theory of market dilution into the fair use analysis. The judge in this case noted that LLM training has the potential to flood the market with competing works. This differs from the "direct substitution" arguments seen in other cases. However, the court still granted summary judgment for Meta. The judge stated that the plaintiffs failed to provide meaningful evidence regarding market dilution. This ruling shows that while the theory exists, it requires a strong factual record to succeed in court.

The judge in the Kadrey case also addressed the analogy of training children to write. The court argued that using books to teach children to write is not like using books to create a product that a single individual can use to generate countless competing works. This distinction attempts to separate human learning from machine training. The decision suggests that the unprecedented nature of AI technology is a factor in how courts weigh the impact on creators.

The Economics of Perplexity and the Revenue Gap

Perplexity’s financial trajectory is rapid. The company reached an estimated $500 million in annualized revenue by April 2026. This is an increase from $100 million just one year earlier. The company has a valuation of $20 billion from its September 2025 funding round. Perplexity uses a revenue-sharing program with over 30 publishers. This program was launched in July 2024. The program includes a $42.5 million pool for its partners. Perplexity also offers subscription tiers like Pro for $20 per month and Max for $200 per month. The company provides usage-based credits for the Computer agent launched in February 2026.

The revenue for publishers through Perplexity is often small. One publishing executive said the money from Perplexity is useful but is much less than what OpenAI offers. Perplexity’s revenue per employee is $200 million for a team of 250 people. The company manages its growth through consumer subscriptions and enterprise contracts. It also uses commerce integrations with partners like PayPal and Venmo.

AI Buyer Number of Partners Reported Value
OpenAI ~24 $250 million (News Corp)
Meta 8 $50 million (News Corp)
Microsoft 11 Undisclosed
Perplexity 30+ $42.5 million (Pool)

The Three Tiers of the AI Licensing Market

The AI content licensing market operates across three tiers. The first tier involves bilateral deals between major AI companies and select publishers with global brand recognition. The second tier is an intermediary layer of companies including TollBit, Sphere AI, ScalePost, Created by Humans, ProRata, and Miso.ai. This tier also includes big tech firms like Cloudflare and Microsoft. The third tier is the long tail of media producers like local newspapers and regional broadcasters. Most publishers belong to the third tier. These publishers are often absent from the licensing market. This creates a "publisher double bind" where the firms causing traffic erosion also control the licensing infrastructure.

The structure of these tiers creates a massive imbalance in bargaining power. Large publishers can negotiate direct deals, but local outlets have no such access. The intermediary layer tries to provide bot detection and pay-per-use pricing, but these startups are vulnerable to acquisition by large tech firms. This dynamic mirrors the shift from search to social media, where platforms eventually controlled the economic logic of the content.

The Verdict on AI Copyright Litigation

The current legal landscape favors a distinction between how data is acquired and how it is used. The $1.5 billion Anthropic settlement shows that piracy carries a high price. The Bartz ruling shows that training on lawfully acquired books may be protected as transformative. The New York Times continues to fight for its archive’s visibility against the risk of regurgitation. The legal battle has moved from theoretical arguments to a contest over the actual merits of fair use and market harm.

The decision in the Thomson Reuters case shows that non-generative tools face a higher bar for fair use. The Kadrey case shows that the theory of market dilution is a difficult hurdle to clear without specific evidence. I recommend watching the outcome of the New York Times case to see if the court finds that verbatim reproduction of content overrides the transformative nature of the model. The outcome of these cases will determine if publishers can maintain their economic viability in an era of answer engines.