• Sources: TechCrunch report, HN discussion
  • Summary: Newly unredacted material in the New York Times case quotes internal remarks from Microsoft and OpenAI staff about how training corpora were assembled. The Times cites more than 91,692 copies of plaintiff works in OpenAI mid-training datasets, and points to Microsoft's own data showing Copilot cut click-through rates for the nytimes.com domain by as much as 93 percent compared to traditional Bing search. TechCrunch carries an explicit provenance caveat, that much of the new information comes from the Times' own brief rather than the underlying exhibits, which remain sealed, and that the quotes appear without their original context.
  • Why it matters: The figures bear on whether fair use holds for models trained on paywalled text, while resting on a plaintiff's brief whose supporting exhibits are still sealed.

send feedback on this story