• Sources: primary, report, earlier accusation, discussion, HN thread, earlier HN thread
  • Summary: Andreas Thom states that after OpenAI's non-sofic-group announcement he emailed Mark Sellke and Sebastien Bubeck with two separate questions, whether months of ChatGPT conversations about the expander matching problem, his own and a Dresden colleague's, entered training data, and whether they were accessible to the reasoning process. He states Sellke's complete answer was that his conversations with ChatGPT did not happen, and that he now reads that categorical answer as addressing only the second question. He connects it to OpenAI's position in the Buckmaster and Alpoge case, where OpenAI states no specific user data was accessed but cannot rule out that de-identified data derived from their product usage helped improve its models, and OpenAI has not responded to his post.
  • Why it matters: The unresolved question is whether discussing unpublished work with a vendor's product leaves that work available to the vendor, and the two accounts now on record do not answer it the same way.
  • Follow-up: Whether OpenAI responds to Thom's post, and whether it states what its mathematics corpus contains.

send feedback on this story