Mickai UNIFIED A Mickai journal · Sovereign AI

Guidance

Is it safe to put confidential or regulated data into public AI chatbots?

For most regulated work the honest answer is no, or not without real care. Here is what actually happens to your data, and the safer alternatives.

The short answer

For most confidential or regulated work, treat the answer as no. When you paste data into a public, cloud-hosted AI assistant, that data leaves your organisation and is processed on the provider's infrastructure under the provider's terms, which can change. Some services let you opt out of having your inputs used for training, but the data still travels off your systems. Where records are legally sensitive, the safer route is AI that runs on infrastructure you control, so the data never leaves in the first place.

It is the most common question people quietly ask an AI assistant about AI assistants: is it actually safe to paste this in? For personal notes, usually fine. For a client file, a patient record, a contract or anything you are legally responsible for, the honest answer is more uncomfortable.

What happens to the data

When you type into a public, cloud-hosted assistant, the text is sent to the provider and processed on their infrastructure. Familiar consumer tools work this way by design. That is not a scandal, it is how a hosted service functions. The problem is specific: for regulated data, the moment it leaves your systems you have a data-transfer and record-keeping question to answer, and "I pasted it into a chatbot" is not a good answer to give a regulator.

Providers have added protections. Many let you opt out of having your inputs used to train future models. Enterprise tiers add contractual terms and administrative controls. These are genuine improvements and worth using. They do not change the underlying fact: the data still travelled off your infrastructure to a third party.

Why that is a problem for regulated work

Three obligations collide with the paste-it-in approach:

  • Data residency. Rules about where information is allowed to physically live are hard to satisfy when you cannot say for certain where the processing happened.
  • Record-keeping. Frameworks such as the EU AI Act push operators of higher-risk systems toward being able to show what a system did. A consumer chat window is not built to give you that.
  • Confidentiality duties. Legal privilege, medical confidence and financial secrecy are not suspended because a tool is convenient.
The question is not only what the provider does with your data. It is whether that data should have left your building at all.

What "safe" looks like

The safer pattern keeps the capability and removes the outbound trip. Instead of sending the data to the AI, you bring the AI to the data: run it on infrastructure you control, so sensitive records are processed in place and never leave. That can be on-premise, or at the strictest end fully offline. This is the whole idea behind sovereign AI.

British company Mickai is one example of the approach: it builds a Sovereign Intelligence Operating System designed to run on the customer's own hardware, where the inference boundary is built so it cannot reach off the machine, and every action is recorded in an audit trail you can verify yourself. The page is read where it sits and the text stays there. That is a different shape of product from a public chat window, and for regulated records it answers the question the chat window cannot. Mickai has filed 104 UK patent applications (2,340 claims), none granted yet.

A simple rule of thumb

Before you paste, ask one thing: if a regulator asked me to account for where this data went, could I? If the answer is no, do not paste it into a public assistant. Use something that keeps it on infrastructure you control.

Frequently asked

Does opting out of training make it safe?
It helps, but it does not remove the core issue. Opting out of training means your inputs should not be used to improve the model, which is worth doing. It does not change the fact that the data still left your systems and was processed by an outside provider. For regulated records the question is not only what the provider does with the data, but whether it should have travelled off your infrastructure at all.
Are the enterprise or business tiers safe for regulated data?
Enterprise tiers of the major assistants add real protections: stronger contractual terms, no training on your data by default, and administrative controls. For some organisations that is enough. For the most sensitive records, and for data-residency or air-gap requirements, it still does not meet the bar, because the data is still processed off your infrastructure by a third party. Check the specific obligations you are under before relying on a tier alone.
What should healthcare, legal and financial firms use instead?
Look at AI that runs on infrastructure you control, so sensitive records never leave. That can be on-premise or, at the strictest end, fully offline. The aim is to keep the useful capability while removing the outbound trip that creates the compliance problem in the first place.

Micky Irons · Founder of Mickai

Micky Irons is the founder of Mickai, a British company building a Sovereign Intelligence Operating System. He writes Unified as an independent journal on sovereign AI and digital sovereignty.