TL;DR:
A Southern District of New York magistrate judge issued an August 26, 2026 order in Encyclopaedia Britannica, Inc. v. Perplexity AI, Inc., shaping AI related discovery and data handling. The order establishes concrete deadlines for obtaining the trial transcript, requires production of Perplexity’s source-code samples and other AI data handling details, and lays out a timeline for producing clickthrough data, Discord data, Slack evidence, and related materials. A follow-on decision on RAG and user activity log data is anticipated, with a status conference set for September 8, 2026. For trial teams, the ruling signals intensified scrutiny of AI retrieval and data provenance in litigation and underscores specific data-gathering protocols that may govern other AI-fueled disputes. Objection practice teams should anticipate challenges around authenticating AI outputs, exposing retrieval pipelines, and cross-examining machine-generated content.
Background
Encyclopaedia Britannica, Inc. and Merriam-Webster, Inc. filed suit against Perplexity AI, Inc. in the Southern District of New York over concerns that Perplexity’s real-time AI responses reproduce or closely resemble the plaintiffs’ reference works. The August 2026 order arises in the ongoing discovery phase of that matter and reflects the court’s active involvement in defining how AI-driven evidence is retrieved, analyzed, and produced for litigation. The case is identified as Encyclopaedia Britannica, Inc. v. Perplexity AI, Inc., No. 25 Civ. 7546 (JLR) (SLC). The docket confirms the court’s focus on the data that underpins Perplexity’s answers, including source code aspects, data pipelines, and activity logs. (docs.justia.com)
What the August 26 Order Does
- Transcript procurement: The court requires the parties to obtain a copy of the conference transcript by August 28, 2026. This creates a documented record of the discussions about discovery parameters, including sensitive data issues. (docs.justia.com)
- Source-code samples and explanation: The plaintiffs must identify and provide Perplexity examples from the retrieval pipeline showing how Perplexity evaluates Britannica and Merriam-Webster works, along with an explanation of any additional information sought about that process. The parties must meet and confer on these issues. (docs.justia.com)
- Clickthrough data production: By September 3, 2026, the plaintiffs will produce “Clickthrough Data,” and Perplexity will produce documents it agreed to produce in response to a specific discovery request. A follow-up conference is scheduled to review progress. (docs.justia.com)
- Discord and Slack The order contemplates discovery regarding responsive communications from Discord, and it notes that Slack data may be subject to production; the court directs continued meet-and-confer as discovery evolves. (docs.justia.com)
- Dow Jones data and related filings: The court references Dow Jones filings concerning retrieval-augmented generation (RAG) and user activity data (UAL), and directs production or sharing of particular Dow Jones materials relevant to the case. This signals the court’s intent to compare and contrast multiple AI systems and their training or output provenance. (docs.justia.com)
- ESI search date ranges: The order sets specific date ranges for various custodians’ data, with timelines including June 1, 2021 through December 31, 2025 for several business custodians, January 1, 2022 through December 31, 2025 for editorial custodians, and a separate window for Merriam-Webster custodial data. These custodial scoping decisions demonstrate a careful balance between broad data collection and practical limits on production. (docs.justia.com)
- Follow-up status conference: A further telephone conference is scheduled for September 8, 2026 at 4:30 p.m. ET to address discovery progress and any ripe issues. (docs.justia.com)
- Separate decision on RAG/UAL data forthcoming: The court indicates that a separate ruling will be issued on RAG and UAL data issues, signaling that this area remains a live dispute with substantial strategic significance for how AI tools are evaluated in court. (docs.justia.com)
Practical implications for trial teams
- AI provenance becomes a central battleground: The order highlights that courts are increasingly demanding visibility into the provenance of AI outputs. For trial teams, this means preparing to defend or challenge the reliability of AI-generated content by pointing to underlying data sources, training inputs, and the retrieval process that produced a given answer. Firms should plan to present or challenge source-code snippets and retrieval logic during cross-examination or in pretrial motions as appropriate.
- Structured data requests as a standard tool: The explicit deadlines for transcripts, code samples, and data categories demonstrate a framework that practitioners can mimic in other AI-related disputes. When litigating AI or data-driven claims, counsel should tailor discovery requests to retrieve “how” and “why” an AI tool produced a given response, not only “what” the output is.
- Data governance and preservation play a critical role: The order’s focus on Clickthrough Data, Discord, and Slack data underscores the importance of preserving user interactions, system prompts, and communications around AI outputs. Litigation teams should coordinate early with IT and eDiscovery to identify and preserve relevant data stores and communications channels that could be used to reconstruct the tool’s behavior.
- Cross-referencing multiple AI platforms: By referencing Dow Jones and other sources, the court signals that comparative analyses of multiple AI systems may become routine in future cases. Practitioners should be prepared to discuss differences in retrieval pipelines, content governance, and output diversity across systems, and to map those differences to potential evidentiary objections or admissibility considerations.
- Case strategy implications for objections and authentication: The emergence of detailed AI data discovery influences how objections might be framed for machine-generated content. Counsel should anticipate arguments about authentication, chain of custody, and the reliability of machine-generated conclusions, and consider leveraging practice tools that train on objection to AI outputs, such as AI-driven mock voir dire or cross-examination drills.
- Practical steps for counsel now:
- Coordinate with the IT and eDiscovery teams to identify potential custodians and data sources for Discord, Slack, source code, and clickstream data.
- Develop a concrete plan to request and review Perplexity’s retrieval pipeline artifacts and any associated documentation.
- Prepare for intense scrutiny of RAG and UAL data once the separate ruling issues are resolved.
- Use trial-prep resources to rehearse challenges to AI outputs and to drill objections anchored in evidence rules and data provenance.
Lookahead and implications for the broader practice
The August 26, 2026 order illustrates a continuing federal trend toward rigorous, data-driven scrutiny of AI-generated evidence. As courts grapple with proprietary retrieval systems, training data, and the governance of machine-generated outputs, trial teams should expect more formalized discovery playbooks for AI disputes. This development also has implications for how opposing counsel will prepare to present AI-derived content in court, potentially affecting pretrial motions, Daubert-like challenges, and the admissibility framework for machine-generated material. For practitioners seeking practical training in handling AI-driven evidentiary issues, Objection Academy offers targeted drills on objections to AI-generated content, data provenance, and real-time cross-examinations of technology-based evidence, providing a useful companion to the evolving standards reflected in cases like Encyclopaedia Britannica v. Perplexity AI.
Sources
- Encyclopaedia Britannica, Inc. v. Perplexity AI, Inc., No. 25 Civ. 7546 (JLR) (SLC) – Order dated August 26, 2026; transcript deadline and data-handling directives, Discord/Slack data considerations, and future RAG/UAL data decision. Justia Dockets & Filings, 1:25-cv-07546. August 26, 2026. (docs.justia.com)
- Encyclopaedia Britannica, Inc. v. Perplexity AI, Inc. – Case details and docket activity (SDNY) – Docket entries and status updates. Justia, 25 Civ. 7546. August 2026. (dockets.justia.com)
Note: For practitioners tracking AI‑related discovery developments, this SDNY order is a concrete, actionable milestone that informs how AI data and provenance must be handled in court. Objection Academy remains a resource for sharpening courtroom technique around AI evidence, including how to frame and sustain objections to machine-generated content.