For weeks in January and early February 2026, Microsoft 365 Copilot read and summarised confidential emails despite sensitivity labels and data loss prevention (DLP) policies being correctly configured to block such behaviour. This was confirmed by Microsoft in Service Advisory CW1226324. The bug affected emails in users’ Sent Items and Drafts folders, meaning legal communications, business agreements, or sensitive health information could all be processed by an AI that explicit organizational policies said should never touch.
Microsoft claimed that users only accessed information they were already authorised to see. This may be technically accurate, as Copilot operates within the user’s mailbox context, but sensitivity labels weren’t in place to stop users from reading their own email; they were there to stop the AI from processing confidential content. For several weeks, the AI processed it anyway.
A single point of failure
The incident made visible the fact that every control designed to keep Copilot away from confidential data – sensitivity labels, data loss prevention (DLP) policies, and access restrictions – lived inside the same platform as Copilot itself. There was no independent layer, no secondary check, no safety net.
Nobody builds a vault where the door lock, the alarm, and the surveillance cameras all run through a single circuit breaker. But that’s what happened here. Microsoft was the AI provider, the security control provider, and the only entity with visibility into whether those controls were working. Therefore, when the guardrails failed, organizations had no independent way to detect the failure.
An industrywide problem
It’s important to note that Copilot is a powerful tool, and code bugs happen to every vendor. The team identified the issue and rolled out a fix. For that they deserve credit. The problem isn’t that Microsoft had a bug. The problem is that the architecture led to a single bug causing a governance failure that had no independent detection for weeks.
This pattern isn’t unique to Microsoft. Whether it’s Copilot, Google Gemini for Workspace, Salesforce Einstein, or any other enterprise AI tool, the AI platform typically provides the governance controls and organizations trust those controls to work. When they don’t, there’s nothing underneath.
The World Economic Forum’s 2026 Global Cyber Security Outlook found that data leaks through generative AI are now the top cyber security concern of CEOs. Yet roughly one-third of organizations still have no process to validate AI security before deployment.
The WEF report also warned that without strong governance, AI agents can accumulate excessive privileges or propagate errors at scale. To combat this, they recommend continuous verification, audit trails, and zero-trust principles that treat every AI interaction as untrusted by default. The Copilot incident demonstrates why those recommendations exist.
Compliance exposure
If Copilot processed emails containing protected health information, organizations may need to assess whether this constitutes a reportable breach under the Data Protection Act 2018. The question isn’t whether the user was authorised, it’s whether the AI’s processing was authorised under the business associate agreement. Microsoft’s public statement doesn’t resolve that analysis.
Under GDPR, Article 32 requires appropriate technical measures for security of processing. If an organization’s sole measure was a vendor’s sensitivity labels that failed for weeks, that’s a difficult argument to take to regulators. The EU AI Act’s Article 12 adds a further layer. If the only records of what the AI accessed come from the vendor that had the failure, organizations lack the independent documentation the regulation demands.
What the fix looks like
Of course, the answer isn’t to stop using AI, as such tools deliver real productivity gains. The answer is to stop trusting AI platforms to govern themselves. Defence in depth is not a new concept. We’ve applied it to network security for decades through firewalls, intrusion detection, endpoint protection, and network segmentation. We all know that multiple independent layers, each capable of catching what the others miss, are a good thing. Yet, for AI governance, we’ve been operating with a single layer for too long.
Defence in depth for AI governance requires an independent data layer between AI platforms and sensitive content. AI doesn’t get direct access to repositories. It authenticates through an external governance layer that enforces policies independently. Purpose binding that restricts which data classifications AI can access, least-privilege controls, continuous verification, and audit trails that the organization controls.
Make sure this doesn’t happen to you
Every major technology shift creates a moment where organizations decide whether to bolt security on after the fact or build it into the architecture from the start. We saw it with cloud migration and remote work. Now we’re seeing it with AI.
The organizations that treat this bug as a wake-up call to build independent AI governance at the data layer will be able to scale AI adoption with confidence. They’ll satisfy regulators with independent evidence and be able to protect sensitive data through architecture that doesn’t depend on trust.
The labels were in place and the policies were configured. The AI read the confidential emails anyway. Make sure this doesn’t happen to you. Ensure you have an independent governance layer that will catch it.
The author
Tim Freestone is Chief Strategy Officer, Kiteworks






