Blog

Draft in an AI Chat: Who Else Can See Yours?

A mathematician asked if his drafts trained the model and got no answer. The five questions your company should ask first.

September 09, 2026 · Agência Primeira Página

Draft in an AI Chat: Who Else Can See Yours?

A mathematician at New York University asked whether the drafts he typed into an artificial intelligence tool had been used to train the model. He never received an answer.

The case is under public dispute and nothing has been proven. But the question left hanging is the same one any company should ask before pasting a budget, a contract, or a client list into a chat.

What happened, in order

  • August 15, 2026: Tristan Buckmaster, of the Courant Institute, and Levent Alpoge, a researcher at Anthropic, prove a result about fluid equations. The work was not yet public.
  • The following week: news of the result reaches OpenAI.
  • Days later: a team led by Sebastien Bubeck, who heads OpenAI's math division, produces a proof that extends the argument.
  • September 6: Buckmaster is told about OpenAI's result and asks whether the model had access to his sessions in an AI-assisted coding tool, where he kept the project's drafts.
  • The answer: he was told the model did not consult user data. On training, he says he received no answer.
  • September 8: the pair publishes three computer-verified proofs. Fields Medalist Terence Tao called the result a remarkable achievement.

OpenAI denies it. Bubeck called the allegations false and inflammatory, and stated that neither the researchers nor the agents saw the pair's work before publication, and that their prompts were not used. The company also states that an internal model solved the problem using 10,000 simultaneous agents over 88 hours.

The question that went unanswered

Notice the difference between these two things, because it's the whole lesson of the episode:

  • Using your text as a prompt — someone takes what you wrote and feeds it to the model at that moment.
  • Using your text as training — what you wrote enters the material that shapes the model, and goes on to influence answers for anyone, afterward, without anyone needing to open your file.

The public denial covered the first. The second is the one that should change what you paste into a chat, and it's exactly the one that went unanswered.

Why this matters to a small business

No one is going to steal your mathematical proof. But what typically passes through an everyday company's chat is: a priced proposal, a cost spreadsheet, an entire contract, process documentation, system code, a client list with names and phone numbers.

Two practical consequences, no drama:

  • Competitive advantage. Your cost table is what sets your proposal apart from a competitor's.
  • Legal obligation. A client's personal data pasted into a third-party service is data processing, and Brazilian law holds whoever controls the data responsible, not the tool. We've already written about this in what the TikTok fine teaches.

A personal account and a business account aren't the same thing

The general market rule today: corporate plans and API access typically don't use your content for training by default; free and personal accounts frequently do, with an opt-out available in some menu.

Two honest caveats. Policy changes over time, and each provider writes its own. That's why the only answer that counts is the one you read and saved yourself, with a date — not the one someone told you.

Five questions to ask your provider

  1. Is what I type used to train the model? By default, or only if I authorize it?
  2. How long is the content kept, and where?
  3. Who, within your company, can read what I send, and under what circumstances?
  4. What happens when I delete a conversation: does it also disappear from training, or just from my screen?
  5. Can I turn off history and keep using the service normally?

Ask in writing. A salesperson's answer in a meeting is worthless six months later.

What not to paste into a chat you don't control

  • Identifiable personal data of a client, employee, or supplier.
  • Passwords, access keys, or credentials of any kind.
  • An entire contract, especially one with a confidentiality clause.
  • Cost and margin tables.
  • Anything that, if published tomorrow, would cost you your advantage.

For everything else, which is most of the work, use it freely. The practical approach is to swap the real name for a nickname before pasting — the model helps just the same, and the sensitive data never leaves your desk.

When the data is too sensitive

There's the option of running your own model, with the data kept in-house. It has become far cheaper than it used to be, and we've already covered the case of a company that swapped rented AI for an open model. It's worth it when the data is the asset; not for writing emails.

And if the choice is to take the discount in exchange for data, which some providers openly offer, make it a conscious decision: that's what we analyze in the discount you pay with your data.

Setting up this policy and deciding where each type of data can circulate is part of what we do in AI implementation for businesses. A written policy solves it, and it needs to exist before the whole team starts using it.

Sources

Coverage of the episode published on September 8, 2026 by Fortune, Engadget, and Scientific American, including Tristan Buckmaster's claims, Sebastien Bubeck's public denial on behalf of OpenAI, and the publication of the proofs verified by Buckmaster and Levent Alpoge. None of the claims had been substantiated as of this writing.

Frequently asked questions

What did OpenAI get accused of doing in the Navier–Stokes case?

Mathematician Tristan Buckmaster alleges that the company began pursuing the same problem after learning of his then-unpublished work, and says he never got a straight answer on whether his private drafts in an AI-assisted coding tool were used to train the model. OpenAI denies it, calls the allegations false, and says it never used the prompts or saw the work before publication.

What's the difference between using my text as a prompt and using it for training?

Using it as a prompt means handing the model what you wrote in that moment so it can generate a reply. Using it for training means folding your content into the material that shapes the model itself, so it can influence other people's answers later on, without anyone ever opening your file. These are two different things and deserve separate questions.

Is what I type into ChatGPT used to train the model?

It depends on the plan and the settings. As a general rule of thumb, enterprise plans and API access typically don't use your content for training by default, while free and personal accounts often do, with an option to opt out. Policies change, so it's worth reading your provider's current version and saving a dated copy.

What should I never paste into an AI chat?

Personally identifiable data about a client, employee, or vendor; passwords or access keys; a full contract, especially one with a confidentiality clause; cost and margin spreadsheets; and anything that, if it leaked, would cost you your competitive edge. For everything else, swapping real names for aliases before pasting handles most cases.

Does using AI with client data carry legal risk in Brazil?

Yes. Pasting a client's personal data into a third-party service counts as processing personal data, and Brazilian law holds the data controller responsible, not the tool being used. It's worth putting in writing what can and can't be shared before rolling this out to the team.

Is it worth running your own model to protect your data?

It's worth it when the data is the business's core asset, like a client database, pricing, or information under contractual confidentiality. The cost has dropped considerably in recent years. For everyday tasks like drafting an email or summarizing a document, it's not worth the effort: anonymizing before you paste already does the job.