For two full years, OpenAI reportedly told a federal court it couldn't search its own ChatGPT logs. Turns out, it already had — on a sample of 78 million of them.
That's the bombshell at the center of a new sanctions motion filed by The New York Times and other news organizations suing OpenAI for copyright infringement. If the allegation holds up, it means OpenAI didn't just drag its feet on discovery. It allegedly looked the court in the eye and described a technical limitation that didn't exist.
The whole thing unraveled during a re-deposition of OpenAI privacy engineer Vincent Monaco in April. He was brought back to the stand after a previous appearance left the court unsatisfied, and his testimony apparently opened a can of worms that two years of careful legal positioning had kept sealed. According to the NYT's filing, Monaco revealed that OpenAI had already conducted searches across large anonymized samples of ChatGPT logs — including datasets of 10 million and 78 million entries — well before the lawsuit even began.
The news plaintiffs argue this matters enormously. Those logs are potentially the most consequential evidence in the entire case. They could show whether real users were regularly prompting ChatGPT to reproduce paywalled articles verbatim — exactly the kind of behavior that would support the infringement claims. Or they could show the opposite, potentially helping OpenAI's fair use defense. Either way, keeping them hidden for two years wasn't a minor procedural hiccup.
The sanctions request argues that OpenAI's concealment inflated legal costs, stretched out discovery, and wasted the court's time. That's the kind of conduct judges tend to take seriously, regardless of which side they sympathize with on the underlying claims.
OpenAI, for its part, framed the sanctions motion as a privacy grab dressed up in legal language. A spokesperson argued that the NYT's recent decision to drop certain claims was a sign its case was falling apart, and accused news plaintiffs of trying to dig through logs belonging to users who have nothing to do with the lawsuit. The company says it's defending both fair use principles and ordinary people's privacy.
The NYT pushed back hard on that characterization. A spokesperson said dropping those claims was a strategic move to streamline the case and add Microsoft as a defendant — not a concession. The core argument, they insist, hasn't changed: that OpenAI and Microsoft used millions of copyrighted works without permission to build a product that competes directly with the outlets that produced them.
What makes this moment particularly awkward for OpenAI is the timing. The company has spent considerable energy positioning itself as a trustworthy, transparency-minded actor in the AI space. Getting caught — by one of its own witnesses, no less — allegedly misrepresenting its technical capabilities to a federal court is not a great look for that narrative.
The sanctions motion is still heavily redacted, so the full picture isn't public yet. But what's visible is enough to suggest this case is about to get significantly messier for OpenAI before it gets any cleaner.