The gap
A policy document does not change what an assistant can retrieve
A classification policy describes intent. An access model produces behaviour. The distinction is easy to state and easy to lose, because a signed policy feels like a completed control and generates the same paperwork an implemented control would generate. Microsoft says the quiet part directly in its own service assurance material: data classification levels by themselves are simply labels or tags that indicate the value or sensitivity of the content. The protection comes from what reads the label.
When a company connects an enterprise assistant to its content stores, the assistant inherits the effective permission map rather than the policy. Microsoft states its own position without hedging: Copilot only surfaces organizational data to which individual users have at least view permissions, and the grounding layer honours that boundary. Read as reassurance, that sentence closes the conversation. Read correctly, it opens it, because the sentence is a statement about permissions and says nothing at all about whether those permissions are right. Microsoft makes the corollary explicit elsewhere, noting that because of the power and speed with which AI can proactively surface content, generative AI amplifies the problem and risk of oversharing.
The mechanism is worth stating plainly, because it is not an attack and there is no vulnerability involved. A permission granted to an individual for one project in 2021, on a folder that later became a parent of other folders, is still granted. A group created for a study that closed still has members. A departing employee whose account was disabled leaves behind the group memberships that survived them. A guest account created for a CRO transition was never removed. Each of these is individually defensible and collectively they form a permission map nobody has read end to end. Human search never surfaced most of it because humans navigate by folder and by memory. A retrieval layer navigates by index.
At a clinical-stage biotech client, during a first engagement, connected AI reached employee-information files in Box through existing user permissions. Nothing was misconfigured in the AI tool. No control failed. The files were reachable by those users before the assistant existed, and the assistant simply made that reachability legible for the first time. That is the shape of the finding in almost every environment we look at, and it is why the remediation is a content and identity project rather than an AI project.
The consequence for scoping is that a classification engagement cannot begin with the label set. It begins with the current state of permissions, because the taxonomy you design has to be applied to content whose access map you understand. Designing five beautiful tiers and then discovering that half the estate has broken permission inheritance in ways nobody can explain is a common and avoidable sequencing error.
What a policy produces
A shared vocabulary, an audit artefact, and an agreed statement of who should see what.
What a policy does not produce
A label on a file, a group in the directory, or a rule in the platform that reads either one.
What AI changes
Nothing about the permissions. Only the speed and completeness with which they are exercised.
Where to start
The current access map, not the target taxonomy. The taxonomy has to land on real content.
The AI did not create the exposure. It read the permission map faster and more completely than any person ever had.
Related evidence and next steps
- Microsoft: data classification and labels— Microsoft service assurance on what a classification level is and is not.
- Microsoft 365 Copilot privacy and permissions— The permissions statement, the grounding boundary and the training exclusions.
- Microsoft Purview data security for AI— Microsoft on how generative AI amplifies existing oversharing.
- AI Access Exposure Review— The review that produces the backlog this build works through.




