Copilot Readiness for Life Sciences: What Purview Protects, and What It Doesn’t
Microsoft Purview is a real control plane. It is also not the finished answer to “is our AI deployment safe?” This page sets out what labels and DLP genuinely do, what Microsoft’s own documentation says they do not cover, and how both map to 21 CFR Part 11 and EU GMP Annex 11.
Four assumptions worth testing before you switch Copilot on
Each of these is a statement people make in kickoff meetings. Each is contradicted, in whole or in part, by Microsoft product documentation. Verified against Microsoft Learn on 27 August 2026.
01
It only shows people what they can already see
True, and that is the problem statement rather than the reassurance. Copilot inherits your permission model at machine speed. If access has drifted for a decade, the assistant is the first tool fast enough to make that visible.
We labelled the sites, so the documents are covered
Container labels govern the container. Microsoft documents that items do not inherit them, get no content marking or encryption, and cannot display a container label in Copilot or support label inheritance.
Manual labelling is E3-class. Automatic labelling, adaptive scopes, records settings, insider risk and the DLP rule that stops Copilot reading a labelled file are all E5-class. A Plan 1 licence must sit alongside Plan 2 or labelling fails silently.
Restricted SharePoint Search is retiring and its replacement is explicitly not a security boundary. Neither changes permissions, and neither removes content from the search index.
Purview is not one product. It is a set of loosely coupled control planes — sensitivity labels, data loss prevention, data lifecycle and records management, and insider risk — that share an administration portal and very little else. Understanding where each one enforces, and where it merely annotates, is the difference between a control narrative that survives an inspection and one that does not.
Definitions
A sensitivity label is metadata, not a permission
Microsoft describes a sensitivity label as information stored in clear text in the metadata for files and emails, which is precisely why third-party apps and services can read it. That single design decision explains most of what Purview can and cannot do. A label travels with the content wherever it is saved or stored, and each item that supports labels can carry exactly one.
Because the label is metadata rather than an access decision, the label by itself enforces nothing. Enforcement arrives from one of four other places: encryption applied by the label through the Azure Rights Management service, content markings applied by Office clients, container settings applied to a site or group, or some separate policy — data loss prevention, Copilot DLP, retention — that uses the label as a condition. Microsoft supports marking-only labels explicitly, describing the mode as labelling content without using any protection settings. A marking-only label is a clear-text tag. It is useful for reporting and for driving other policies, and it stops nobody.
Label scope is chosen when the label is created and determines both which settings you can configure and whether users see the label in a given application. The scopes are files and other data assets, emails, meetings, and groups and sites. The last of these only becomes available once container labelling has been enabled in the tenant, which is a separate one-time step that many organisations never complete.
Label priority is positional and matters in two places: it drives the justification prompt when a user downgrades a label, and it resolves auto-labelling conflicts, where Microsoft documents that the last sensitive label is selected and then, if applicable, the last sublabel. It matters a third time in Copilot, where inheritance onto generated content also follows the highest-priority rule.
There are hard tenant limits worth knowing before a taxonomy workshop. More than a thousand labels are technically supported, but a maximum of 500 applies when the label configures encryption specifying users and permissions. Microsoft’s field guidance is far tighter than the technical ceiling: real-world deployments show effectiveness is noticeably reduced when users have more than five main labels or more than five sublabels per main label, and some applications cannot display all labels when too many are published to the same user.
One behaviour surprises almost every administrator. If you delete a sensitivity label from the Purview portal, the label is not automatically removed from content and any protection settings continue to be enforced on content that had that label applied. Deleting a label is a portal cleanup action, not a decryption action.
What a label always does
Writes a clear-text tag into the item that travels with it and that other systems and policies can read.
What a label sometimes does
Applies encryption, content markings, or container settings — but only when those settings are configured on it.
What a label never does
Change who has permission to the item in SharePoint, OneDrive, Box or Egnyte. That remains the permission model’s job.
Design consequence
Keep the sensitivity taxonomy small and separate from your content taxonomy. Four or five tiers, applied across the content model, not mirroring it.
A label is an assertion about content. Only encryption, a container setting, or a separate policy turns that assertion into a control.
Manual, default and mandatory labelling are three different levers
These are label policy settings, not label settings, and conflating them produces rollouts that either annoy everyone or protect nothing. A default label applies to unlabeled documents, Loop items, emails, meeting invitations, new containers and Power BI content. Mandatory labelling blocks save, send or create until the user chooses. Justification on downgrade prompts the user and writes their reason into Activity Explorer.
Microsoft attaches an explicit warning to the first of these: it is usually not a good idea to select a label that applies encryption as a default label to documents. The reason is downstream breakage rather than user irritation, and it is covered in detail later on this page. The warning attached to mandatory labelling is about behaviour: without user training, these settings can result in inaccurate labelling, and they can frustrate users with frequent prompts. Both warnings come from the product documentation, not from a critic.
Justification on downgrade is on by default for files, emails and meetings, and is not available for groups and sites. In Office applications the prompt fires once per app session; in the Information Protection client it fires per file. This is a prompt, not a block. It creates an audit record and a friction point. It does not prevent a determined user from moving a document down a tier, which is the correct design — but it means the compliance value depends entirely on somebody owning the review of those events.
Label policies themselves have priority, with the highest order number winning on a setting conflict, and Microsoft asks you to allow up to 24 hours for label and policy changes to replicate. That replication window is the single most common cause of a control test that appears to fail. Test after the window, not during it.
Microsoft’s naming guidance is worth adopting verbatim because it removes an argument from the taxonomy workshop. It recommends label names that are self-descriptive and highlight relative sensitivity clearly, noting that Confidential and Restricted may leave users guessing which label is appropriate while Confidential and Highly Confidential are clearer about which is more sensitive. The company also documents its own pilot lesson: it originally used a label named Internal, found there were legitimate reasons for a document to be shared externally, and shifted to General.
Default label
Applies where nothing was chosen. Powerful for coverage, dangerous if it applies encryption.
Mandatory labelling
Blocks the action until a label is picked. Buys coverage at the cost of accuracy unless training precedes it.
Justification on downgrade
A prompt and an audit record, once per app session. Valuable only if someone reviews Activity Explorer.
Replication
Allow up to 24 hours for label and policy changes to reach clients before you conclude a control does not work.
Client-side and service-side automatic labelling are two different products
This is the single most misreported area in third-party write-ups about Purview, and getting it wrong changes both the design and the licence bill. Client-side automatic labelling runs in the Office client as the user edits or composes. Service-side automatic labelling runs in the service, over data at rest in SharePoint and OneDrive and data in transit through Exchange. They are configured in different places and they behave differently.
Only the client-side mechanism can recommend a label and let the user accept or reject it; Microsoft states that in both recommend cases the user decides whether to accept or reject the label. Only the service-side mechanism supports simulation, PDFs, images and optical character recognition, nested AND/OR/NOT rule logic, location restriction, labelling of incoming external mail, and the removal of an existing label in SharePoint or OneDrive. Client-side labelling depends on the Office application version; service-side labelling does not.
The Exchange caveat is the one most summaries omit and it changes the scope of a mailbox remediation plan entirely. Microsoft writes that for Exchange, service-side auto-labelling does not include emails at rest, meaning mailboxes. Turning on service-side automatic labelling does not retro-label the archive. Similarly, DLP does not scan or match previously existing email items stored in a mailbox or archive.
There is a second scoping limit with the same shape. Auto-labelling applies your selected sensitive information types only to content created or modified after those information types were created or modified. Static back-catalogue files are simply not seen. Reaching them requires on-demand classification, which is a separate operation and a separate planning conversation.
One accuracy limit is specific to life sciences. Intellectual property, CMC documentation and discovery data are exactly the content that needs trainable classifiers rather than pattern-matching sensitive information types, and Microsoft documents that automatic labelling is not supported for Office for the web when the conditions include trainable classifiers. If a meaningful share of your authoring happens in the browser, that gap is structural rather than a tuning problem.
Tuning happens through confidence levels and instance counts rather than through the choice of classifier. Microsoft’s own worked example escalates by count: recommend a lower Confidential sublabel at one to nine matches, a broader Confidential sublabel at ten or more, and escalate to Highly Confidential at three to nine high-confidence matches or ten or more low-confidence matches. Failures have a specific home in the portal — the policy’s labelled items tab, failed view — and a policy with no visible failures is not the same as a policy that matched what you intended.
Client-side only
Recommend-a-label with user override, and dependence on the Office application version.
Service-side only
Simulation, PDF and OCR support, nested rule logic, location scoping, external inbound mail, and label removal.
Neither
Retro-labelling an Exchange archive, or reaching files that have not been created or modified since the classifier changed.
Licence split
Both mechanisms are E5-class. Enterprise Mobility + Security E5 buys client-side automatic labelling only.
Encryption is the only label setting that controls access
Content markings — headers, footers, watermarks — are cosmetic and reversible, and they carry their own limits: watermarks cap at 255 characters and headers and footers at 1,024, with Excel capped at 255 including invisible formatting codes. Encryption is the setting that actually decides who can open a file. It also introduces every hard dependency in the system.
Microsoft is explicit that encryption prerequisites are not validated when you configure them: when you configure encryption settings, there is no check to validate that these prerequisites are met. Azure Rights Management must be activated. Network reachability must exist. Entra cross-tenant access settings and Conditional Access must not block the flow. Exchange must be configured for Azure Rights Management, and Microsoft spells out the consequence if it is not — users cannot view encrypted emails or encrypted meeting invitations on mobile phones or in Outlook on the web, encrypted emails cannot be indexed for search, and Exchange Online DLP cannot be configured for Rights Management protection.
There are two encryption modes and the difference matters enormously downstream. Admin-defined permissions assign permission levels or custom usage rights at label-configuration time. User-defined permissions let the user decide at the point of protection. The second mode looks flexible and breaks two things: Copilot cannot process unopened files in SharePoint or OneDrive protected that way, and label inheritance is not supported when the encryption is configured for user-defined permissions.
A further behaviour catches teams that route protected documents through email. When an encrypted mail or meeting invitation carries unencrypted Office attachments, those attachments automatically inherit the same encryption settings, and Microsoft states that you cannot turn off this encryption inheritance. That is usually what you want and occasionally a surprise, particularly for a submission package assembled from mixed sources.
The usage rights inside a label are not decoration either. The right that governs whether an AI assistant can summarise content is EXTRACT, shown in the portal as copy and extract content. It becomes the central mechanism in the Copilot section below.
DLP, lifecycle management, and insider risk do different jobs
Data loss prevention is a policy engine with conditions and actions applied to locations. The supported locations are Exchange email, SharePoint sites, OneDrive accounts, Teams chat and channel messages, Defender for Cloud Apps instances, Windows 10 and 11 and recent macOS devices, on-premises repositories, Fabric and Power BI workspaces, and Microsoft 365 Copilot in preview. Conditions can include a sensitivity label, a sensitive information type, or a retention label. Actions run from a policy tip through block with override and justification to a hard block or quarantine at rest.
One operational detail about DLP evidence catches investigations out. DLP alerts are available in the Microsoft Defender portal for six months, but only for 30 days in the Purview DLP alerts dashboard. If your incident process reads the Purview dashboard, your effective retention is a month.
Data Lifecycle Management and Records Management are two solutions sharing three mechanisms: retention policies, retention labels, and retention label policies, plus email archiving. Retention policies are the location-scoped cornerstone; retention labels are the exception mechanism for individual items. Microsoft directs regulated record-keeping specifically to the second solution: if you need to manage high-value items for business, legal or regulatory record-keeping requirements, use retention labels with records management. For a life-sciences company, that sentence decides where the trial master file conversation belongs.
The licensing line runs straight through the middle of this. Basic org-wide and location-wide retention policies are E3-class. Adaptive policy scopes — the query-driven scoping most enterprises actually want — are E5-class only, as are auto-applied retention labels, trainable-classifier-driven retention labels, and every setting that makes a retention label a record: event-based retention start, disposition review, mark as record or regulatory record, and automatic label change at the end of a period.
Insider Risk Management is detection and investigation, not prevention, and it is E5-class throughout. It is built pseudonymised by default with role-based access controls and its own audit log. Several policy templates are directly relevant to life sciences: data theft by departing users, data leaks including priority and risky users, risky AI usage, risky agent usage, risky browser usage, and the security policy violation set. There is also a patient data misuse template in preview, which requires the Microsoft 365 HR connector plus a healthcare-specific data connector and scores behaviour inside electronic health record systems. Forensic evidence is a separate opt-in capacity add-on sold in units of 100 GB per month.
Audit deserves its own note because it is frequently over-promised in control narratives. Audit (Standard) is E3-class and now explicitly includes audit for Microsoft Copilot interactions, though those records only generate when Copilot is licensed and in use. Audit (Premium), with one-year retention, custom retention policies, crucial events and a high-bandwidth Management Activity API, is E5-class, and ten-year retention requires a further add-on.
DLP
Conditions and actions across ten location types, including Microsoft 365 Copilot in preview. Alerts age out of the Purview dashboard in 30 days.
Data Lifecycle Management
Retention policies for scale, retention labels for exceptions. Adaptive scopes are E5-class only.
Records Management
Where regulated records belong. Record and regulatory-record settings are E5-class only.
Insider Risk Management
Detects and investigates. Does not prevent. E5-class, pseudonymised by default, with a metered forensic-evidence add-on.
What each licence tier actually buys, as documented
Rows summarise the Microsoft Purview service description as published on Microsoft Learn and read on 27 August 2026. SKU naming in this area changed materially during 2025 and 2026 as Microsoft Purview Suite branding displaced older Microsoft 365 E5 Compliance names. Confirm your own entitlement against the current service description before designing a control around it.
Capability
E3-class SKUs
E5-class SKUs
Notes and additional licences
Manual sensitivity labelling
Included
Included
Also OneDrive Plan 2, EMS E3/E5, Office 365 E3/E5, AIP Plan 1 and Plan 2
Client-side automatic labelling
Not included
Included
Enterprise Mobility + Security E5 buys client-side automatic labelling only
Service-side automatic labelling
Not included
Included
Also Purview Suite variants, Microsoft 365 E5 Information Protection and Governance, Office 365 E5
Scanner-based discovery
Included
Included
Discovery is E3-class; labelling the discovered content is not
DLP for Exchange, SharePoint, OneDrive
Included
Included
Also Business Premium and the relevant Plan 2 standalone SKUs
DLP for Teams and endpoint DLP
Not included
Included
Advanced Outlook policy tips sit on the same side of the line
DLP to restrict Copilot processing a labelled item
Not included
Included
Explicitly No for Business Premium, E3-class and Office 365 E3-class SKUs
DLP to safeguard Copilot prompts
Included
Included
Available to all users of Microsoft Copilot and Copilot Chat
Label inheritance from Copilot input to output
Included
Included
Requires a Microsoft 365 Copilot licence in addition
Retention policies, org-wide or location-wide
Included
Included
The cornerstone mechanism; scoping is where the tier split appears
Adaptive policy scopes
Not included
Included
Query-driven scoping is E5-class only
Auto-applied and trainable-classifier retention labels
Not included
Included
Includes default library label and adaptive-scope application
Record and regulatory-record settings
Not included
Included
Event-based start, disposition review, mark as record, end-of-period label change
Retention over Copilot interactions
Included
Included
Requires a Microsoft 365 Copilot licence in addition
Audit (Standard), including Copilot interactions
Included
Included
Records only generate when Copilot is licensed and in use
Audit (Premium)
Not included
Included
One-year retention; ten-year retention needs a further add-on
Insider Risk Management
Not included
Included
Forensic evidence sold separately in 100 GB per month units
Restricted Content Discovery
Not included
Not included
Requires a Copilot licence plus SharePoint Advanced Management
Entra dynamic membership groups
Depends
Depends
Entra ID P1 per unique user who is a member of a dynamic group
Entra access reviews
Depends
Depends
Requires Entra ID Governance or Entra Suite; some capability works on P2
What Microsoft documents Purview does not cover
This is the section that matters most, and none of it is our opinion. Every limitation below is published by Microsoft in its own product documentation. A company that deploys Copilot believing labels have solved access control has bought a false sense of safety, and the correction is available for free on Microsoft Learn.
The baseline
Copilot honours your permissions, which is the problem statement
Microsoft’s core claim about Copilot is a permissions claim, not a classification claim: Copilot only surfaces organizational data to which individual users have at least view permissions. The grounding layer holds the same line, with the semantic index honouring the user identity-based access boundary so that the grounding process only accesses content the current user is authorized to access. Prompts, responses and data accessed through Microsoft Graph are not used to train foundation models.
Read the corollary in the same documentation and the commercial reality of a readiness engagement becomes obvious. Microsoft asks customers to make sure they are using the permission models available in Microsoft 365 services, such as SharePoint, so the right users or groups have the right access to the right content. Elsewhere it states the consequence directly: because of the power and speed with which AI can proactively surface content, generative AI amplifies the problem and risk of oversharing or leaking data.
Copilot does not create an oversharing problem. It industrialises one that already exists. A permission that has been wrong since 2019 was previously protected by the practical difficulty of finding the document. An assistant removes that difficulty in one sentence. At a clinical-stage biotech client, during a first engagement, connected AI reached employee-information files in Box through existing user permissions — not through a flaw in the AI product, and not through a flaw in Box, but through permissions that had been in place long enough that nobody was reviewing them.
Microsoft also states plainly that the responses generative AI produces are not guaranteed to be 100% factual. That belongs in any control narrative written for a quality function, alongside the observation that all Copilot prompts run in the security context of the user who initiates them.
The question is never whether the assistant respects permissions. It is whether your permissions have ever been tested at the speed the assistant works.
Four categories of content are, by Microsoft’s own documentation, invisible to the labelling layer when Copilot looks at them. In each case the content may be perfectly well governed by some other mechanism; what is absent is the label-based protection an organisation thinks it has.
Container labels do not reach items. Microsoft writes that sensitivity labels applied to containers — groups and sites — are not inherited by items in those containers, and as a result the items will not display their container label in Copilot and cannot support sensitivity label inheritance. A Confidential-labelled Team looks governed in the admin portal and contributes nothing to Copilot’s handling of the documents inside it.
Teams meeting and chat labels are not recognised at all. Microsoft states that sensitivity labels which protect Teams meetings and chat are not currently recognised by Copilot and agents — no label display, no copy prevention, no inheritance. Invitations, responses and calendar events are exempt from that gap, but meeting-invitation metadata is unlabelled in a different way: the label attaches to the invitation body, so a question such as what meetings do I have on Monday returns date, time and recipient data carrying no label.
Graph connector and plugin data arrives unlabelled. Where content is indexed from an external source, Microsoft documents that sensitivity labels and encryption applied to that data from external sources are not recognised by Microsoft 365 Copilot Chat. Power BI is called out as the notable exception where this does bite. For a biotech that has connected a document management system, a CRO portal or a laboratory system through a connector, this is the gap with the largest surface area.
Double Key Encryption is invisible. Items protected by DKE are not returned by Copilot, and Copilot is disabled in the application while a DKE item is open. That is protective, and it is also a capability loss to plan for rather than discover.
Container labels
Govern the container. Contribute nothing to item-level handling in Copilot.
Teams meetings and chat
Labels not recognised by Copilot or agents: no display, no copy prevention, no inheritance.
Graph connector content
External labels and encryption are not recognised by Copilot Chat.
Double Key Encryption
Items are not returned, and Copilot is disabled in-app while one is open.
Where a control exists but does not stop what you expect
The second category is more subtle than a missing label and more dangerous, because a control is visibly configured and a reasonable person concludes the case is closed. Four documented behaviours belong here, and each has a different remedy.
Encryption-based gating works through the EXTRACT usage right rather than through the label name. If a label applies encryption, Copilot returns content only when the user holds EXTRACT — shown in the portal as copy and extract content — in addition to VIEW. Where a user has VIEW but not EXTRACT, Microsoft documents that Copilot will not summarise the content but can reference it with a link so the user can open and view it outside Copilot, and that when a user has that content open in an application they cannot use Copilot at all. Note that the Rights Management owner always holds EXTRACT on content they encrypted themselves.
Information Rights Management applied independently of a label can be bypassed. Where content is encrypted separately and grants VIEW but not EXTRACT while the label itself applies no encryption, Microsoft states the content can be returned by Copilot and therefore sent to a source item. The related SharePoint behaviour has the same shape: IRM library settings apply usage rights when files are downloaded, not when they are created or uploaded, so IRM alone will not stop Copilot summarising a file. The documented fix is a label applying encryption without the EXTRACT right.
Files uploaded directly into a prompt bypass DLP entirely. Microsoft writes that DLP cannot scan the contents of files uploaded directly into prompts, so evaluation of the uploaded file for sensitive data does not occur, and that DLP only checks the text typed into the prompt. For a regulated organisation this is the most operationally important sentence in the whole Copilot documentation set, because it means the control you licensed at E5 does not cover the easiest path a user has to a model.
Blocked items still appear in citations. Where a Copilot DLP rule prevents an item being processed, Microsoft documents that identified items still appear in the citations of the response while the content of the item is not used. The item name, and therefore often the subject matter, still travels. Two further configuration rules constrain the design: you cannot use both a sensitive information type condition and a sensitivity label condition in the same rule, and selecting the Microsoft 365 Copilot location disables every other location in that policy.
Timing is its own gap. Updates to a DLP policy can take up to four hours to reflect in the Copilot experience, and if a sensitivity label is applied mid-session the policy is enforced starting the next time the file is opened. A control that is correct on paper is not in force during that window.
One agent-specific behaviour deserves separate attention. Microsoft documents that for the Teams Channel Agent, permissions are not checked for all users in the channel, with the result that Channel Agent could summarise the content of items that one or more users in the channel do not have permission to open. Information barriers are not supported for it, and Copilot DLP cannot stop it summarising labelled files.
A configured control and an enforced control are different things. The gap between them is measured in usage rights, upload paths, citation metadata and four-hour propagation windows.
Audit is not usage reporting, and inheritance beats human judgement
Two closing items in this section change what a compliance function can honestly write down. The first is about evidence. The second is about who wins when a label is disputed.
Microsoft states that Copilot auditing is not intended to be used as the basis for Copilot usage reporting, and separately that auditing captures the Microsoft 365 Copilot activity of search but not the actual user prompt or response. Reconstructing an interaction therefore requires eDiscovery or the DSPM for AI activity explorer, which are different tools with different retention and different access controls. Any control narrative that promises a prompt-level audit trail must name the tool that will produce it and the retention that applies.
On inheritance, Copilot applies the highest-priority label among the items it referenced when it creates new content, and in Copilot Chat responses that reference multiple items users see the highest-priority label among them. The unusual part is the override rule: Microsoft documents that unlike other automatic labelling scenarios, an inherited label applied when you create new content will replace a lower-priority label that was manually applied. This is the one documented case where an automatic labelling decision overrides a deliberate human one, and it is generally the behaviour you want — but it needs to be in the training material, because a user who chose General and finds Highly Confidential will otherwise file a ticket.
For an EU-facing organisation there is one further published caveat worth recording in a data protection impact assessment. Microsoft states Copilot compliance with GDPR and the EU Data Boundary while noting that models provided by Anthropic as a subprocessor are currently excluded from the EU Data Boundary.
Restricted SharePoint Search is retiring. Its replacement is not a security boundary either
Restricted SharePoint Search was the emergency brake many organisations pulled before a Copilot pilot. Microsoft documents that it is retiring and that new enablement was blocked starting 31 July 2026. It was never a control in the first place: Microsoft states that it is not a security boundary and does not change any permissions on SharePoint sites, caps the allow list at 100 sites, and leaks anyway, because if a user recently accessed a site, or the site was shared with that user in Teams or Outlook, the site appears in the user’s results and responses even if it is not on the allow list.
The replacement, Restricted Content Discovery, requires both a Microsoft Copilot licence and SharePoint Advanced Management. Microsoft is equally direct about its limits: it does not change existing permissions and does not remove content from the Microsoft 365 search index, while Purview eDiscovery and automatic labelling keep working on the content. What it does add is the removal of AI entry points from the site interface, so users no longer see the Copilot button, AI action menus including agent creation, or the option to create pages with AI.
Plan for latency and plan for the cost to grounding quality. On sites with more than 500,000 items, Microsoft says an update to Restricted Content Discovery could take more than a week to fully process, and it cautions against wide use precisely because starving Copilot of grounding degrades the product you are paying for. Treat it as containment while the permission model is corrected, with a date on which it comes back off.
Not a permission change
Neither feature alters SharePoint permissions. The people who could open the document yesterday can still open it today.
Not an index removal
Content stays in the Microsoft 365 search index. eDiscovery and automatic labelling continue to operate on it.
Slow at real scale
Above 500,000 items on a site, Microsoft says an update can take more than a week to fully process.
Scale limits, simulation mode, and the over-encryption failure surface
The third category of surprise is not conceptual. It is operational: hard numbers that decide whether a policy can be activated at all, a simulation mode that behaves nothing like a dry run, and a documented list of ways encryption breaks the systems around it. These are the findings that turn a six-week rollout into a six-month one when they are discovered late.
Hard numbers
The limits that decide whether a policy can be turned on
Service-side automatic labelling carries published ceilings, and at biotech scale at least one of them usually binds. A tenant can automatically label a maximum of 100,000 files per day. A tenant can hold a maximum of 100 automatic labelling policies. In the portal, each policy can target up to 100 explicitly included or excluded locations, with the default of all locations exempt from that count.
The limit that stops projects is the simulation ceiling. Simulation supports a maximum of four million matched files, and Microsoft states that if more than this number of files are matched from an auto-labelling policy, you cannot turn on the policy to apply the labels. A broadly scoped first policy across a mature SharePoint estate will hit that ceiling, and the remedy is to narrow the rule or the location scope rather than to request an exception.
A second constraint is geographic rather than numerical: Microsoft notes that automatic labelling is not currently available in all regions because of a backend Azure dependency. Confirm availability for your tenant region before the design assumes it.
Identity-side limits belong in the same planning conversation because access model work usually accompanies labelling work. A single Entra tenant can have a maximum of 15,000 dynamic membership groups; the body of a membership rule cannot exceed 3,072 characters; you cannot mix users and devices in one dynamic group; and you cannot manually add or remove a member of a dynamic group, because the rule is the only source of truth. Microsoft also advises minimising the use of match or contains operators because they degrade processing time.
One identity risk is worth stating explicitly, because it is the mechanism by which a carefully designed access group quietly stops meaning anything. Microsoft writes that the security of a dynamic group’s membership depends on who can modify the attributes referenced in the rule, and notes that some on-premises attributes might be configured with permissions that allow users to modify their own values. If a user can edit the attribute, the user can edit themselves into the group. Microsoft’s structural answer is that role-assignable groups already prevent this risk by requiring assigned rather than dynamic membership.
100,000 files per day
The tenant ceiling for automatic labelling throughput.
4,000,000 matched files
Exceed this in simulation and the policy cannot be activated at all.
100 policies, 100 locations
Per tenant and per policy respectively, when locations are listed explicitly in the portal.
15,000 dynamic groups
The Entra tenant ceiling, with a 3,072-character cap on each membership rule body.
Teams reach for simulation expecting the behaviour of a database transaction that can be rolled back. It is closer to a live read-only pass with side effects, and Microsoft documents each of them.
It is not silent: simulation still generates activity alerts. If your alert policy is unscoped, the first simulation run will mail a large volume of notifications to the compliance team, and the second-order effect is that people start ignoring the channel. Scope or temporarily disable the alert policy for the simulation window, and record that you did.
A newly created policy cannot be managed for roughly 24 hours while the backend provisions, with turn on, edit and delete greyed out. A duplicated policy lands in simulation by default and does not tell you. Both behaviours produce the same support ticket, which is that the policy appears to be doing nothing.
Most importantly, simulation shows the result of a single policy, and the enforced result can differ from any single policy’s simulation because label priority resolves conflicts between policies at enforcement time. Exchange results are additionally not reproducible, because simulation evaluates mail sent and received during the run rather than a fixed corpus.
Practical parameters: a run targets 12 hours to complete, the portal displays a maximum of 100 items per site, and the CSV export ceiling is 50,000 records. Microsoft’s recommended pattern is to start with one SharePoint site or a single library, widen iteratively, and use the run to estimate production runtime rather than to prove correctness.
Simulation answers how many and roughly which. It does not answer what will actually happen when three policies and a label hierarchy interact.
The documented ways encryption breaks the systems around it
Encryption is the only label setting that controls access, so a team under pressure often reaches for it first. Microsoft’s SharePoint and OneDrive documentation contains the most complete published account of what that costs, and reading it before the design workshop is the cheapest hour available to a project.
Files labelled and encrypted before the tenant enabled sensitivity labels for SharePoint and OneDrive are never recognised; the documented remedy is to download those files and upload them again to their original location. Labels that set user access to content to expire on anything other than never, and labels using Double Key Encryption, cannot be processed at all — those documents are not returned in search results even if they are updated, and the label, along with all of its sublabels if it is a parent, does not display in Office for the web. With Hold Your Own Key and Double Key Encryption, coauthoring, eDiscovery, DLP and search all fail together.
Three breakages are specific enough to predict the affected documents in a biotech. An encrypted Office file larger than 12 MB that is copied or moved to a different site can no longer be processed by SharePoint. Files containing Power Query data, data stored by custom add-ins, custom XML parts, a bibliography, or a SharePoint Document ID cannot be processed when labelled and encrypted from Office desktop applications. In practice that reads as every Excel model with a data connection, every Word document produced with a reference manager, and every document uploaded into a Document-ID-enabled library.
Collaboration and automation degrade too. For encrypted documents, printing, downloading, exporting and creating a copy are not supported in Office for the web. Office desktop and mobile applications do not support coauthoring for files labelled with encryption by default; that requires a separate enablement step. An application or service using a service principal name that downloads an encrypted file and uploads it again with a label applying different encryption settings receives an upload failure — Microsoft’s own worked example is Defender for Cloud Apps changing a label from Confidential to Highly Confidential. External-tenant guests cannot open user-defined-permission files in Office for the web even when the label grants access to their home-tenant group.
Reversing the decision is not clean either. Microsoft warns that turning the feature off later produces a mixed protection state, and that this mixed protection state can result in higher administrative overheads with help desk incidents to investigate the inconsistent behaviour, and might also be a problem for compliance requirements.
Removal policies behave counter-intuitively for the same underlying reason as automatic labelling. Simulation matches every file meeting the conditions, but on activation the service only re-evaluates files whose state has recently changed, so untouched files are skipped and on-demand classification is required to force a re-scan. A removal policy cannot override an encrypting label, a conflicting apply-policy beats a remove-policy, and removing a label that applies encryption also removes the encryption — intentional, and frequently surprising to administrators expecting unlabelled but still protected. Cross-tenant migration has its own trap that matters for biotech acquisitions and CRO transitions: after cross-tenant migration, the source-tenant label often is not recognised in the target tenant.
Predictable casualties
Excel models with data connections, Word files with a bibliography, and anything in a Document-ID-enabled library.
The 12 MB rule
An encrypted Office file over 12 MB stops being processable once it is moved or copied to another site.
Search and discovery
Expiry-configured and Double Key Encryption labels make documents unsearchable, even after they are updated.
Automation
Service-principal re-upload with different encryption fails. Plan integrations before, not after.
Half of this project is not technical, and Microsoft says so
The most useful planning number in Microsoft’s documentation is not a limit. It is an observation about where the effort goes: Microsoft customers indicate that approximately 50% of an information protection project is business focused rather than technical, so end-user training and communication are critical to success.
The mechanics Microsoft recommends are modest and specific. Publish new labels to a few test users first, wait at least an hour, and verify the label behaviour before widening; wait at least a day before making a label available to more users. Create a working virtual team rather than assigning the rollout to one administrator. Test and tailor label names and tooltips with the people who will apply them, rather than with the people who designed them.
Ship a custom help page through the label policy learn-more URL. This is a small detail with a disproportionate effect: without it, users do not get the corresponding menu option in the Office sensitivity button at all, so the moment of confusion has no exit.
For pilots, Microsoft suggests using credit-card numbers as the training classifier even in organisations that do not process them, because the concept of these being sensitive items that need protection is easily understood by users. Production classifiers then move to what actually matters — trainable classifiers for intellectual property, exact data match for subject and employee data — once the mechanism itself is understood.
Microsoft’s current deployment model inverts the traditional advice and is worth reading in full before committing to a sequence. Its premise is that crawl-walk-run approaches stall on three specific things: defining the label taxonomy, concerns about encryption affecting users and line-of-business applications, and limited adoption through manual labelling. Its answer is to secure by default at the all-employees tier, derive file labels from site labels to reach scale quickly, train users to update labels for sharing exceptions instead of teaching them when to protect, use automatic labelling only for higher-sensitivity escalations, accelerate DLP, and use insider risk to spot suspicious labelling behaviour. Its milestone ladder puts a SharePoint default label without encryption first, extends the default to all files second, and adds encryption third.
21 CFR Part 11: §11.10(d) and §11.10(g) are two different requirements
Part 11 §11.10(d) requires limiting system access to authorized individuals. Paragraph (g) separately requires authority checks to ensure that only authorized individuals can use the system, electronically sign a record, access the operation or computer system input or output device, alter a record, or perform the operation at hand. A folder permission satisfies (d). Only rights that gate specific operations satisfy (g), which is why a permission model alone is not a Part 11 answer.
Paragraph (a) requires validation of systems to ensure accuracy, reliability, consistent intended performance, and the ability to discern invalid or altered records. Purview is not validated by Microsoft on your behalf, and no licence tier changes that.
Paragraph (b) requires the ability to generate accurate and complete copies of records in both human readable and electronic form suitable for inspection, review and copying by the agency. Test this against encrypted content specifically: an inspector-facing export path must survive whatever protection you applied.
Paragraph (c) requires protection of records to enable their accurate and ready retrieval throughout the records retention period. A label that applies encryption is a commitment to key custody for the whole of that period.
Paragraph (e) requires secure, computer-generated, time-stamped audit trails that independently record the date and time of operator entries and actions that create, modify or delete electronic records.
Paragraph (k) requires documentation controls including revision and change control procedures that maintain an audit trail documenting time-sequenced development and modification of systems documentation.
Read against the Copilot gaps above, the practical question becomes narrow and answerable: for each GxP record class, which system performs the authority check, and does an AI assistant sit inside or outside that system’s boundary?
EU GMP Annex 11: the clauses an AI deployment actually touches
Annex 11 §12.1 asks for physical or logical controls restricting access to a computerised system to authorised persons, and §12.2 makes the extent of those controls depend on the criticality of the system — the clause that justifies a tiered model rather than a uniform one. §12.3 is the sentence most access-review programmes are built on: creation, change and cancellation of access authorisations should be recorded.
Section 12.4 extends the same logic to the data itself: management systems should be designed to record the identity of operators entering, changing, confirming or deleting data, including date and time. An assistant that writes into a governed system inherits that requirement rather than escaping it.
Section 9 covers audit trails on a risk basis and requires that for change or deletion of GMP-relevant data the reason should be documented, and that audit trails need to be available and convertible to a generally intelligible form and regularly reviewed. Regular review is a resourcing commitment, not a configuration setting.
Section 7.1 requires that access to data should be ensured throughout the retention period, and Section 17 adds that archived data should be checked for accessibility, readability and integrity. Both bear directly on any decision to encrypt content that will be retained for decades.
Section 6 covers accuracy checks for data entered manually or transferred between systems, which is the clause an integration between a collaboration store and a validated system has to satisfy.
PIC/S PI 041-1 states the least-privilege expectation in terms a Purview or Entra designer can implement directly: systems should support different user access roles and assignment of a role should follow the least-privilege rule, assigning the minimum necessary access level for any job function. It also prohibits shared passwords, requires that administrator access rights not be given to normal users, and requires that systems be able to generate a list of users with actual access for use during periodic user reviews.
WHO Technical Report Series No. 1033, Annex 4 adds that there should be a documented system defining the access and privileges of users, that inactivated users should be retained in the system with a list of active and inactivated users maintained throughout the system life cycle, and that shared logins or generic user access should not be used for systems generating, amending or storing GxP data.
MHRA data-integrity guidance supplies the vocabulary that ties all of this together: ALCOA, meaning attributable, legible, contemporaneous, original and accurate, extended by ALCOA+ to complete, consistent, enduring and available.
HIPAA and GDPR: what is Required, what is Addressable, and where PHI must not be
Under 45 CFR 164.312(a)(1), the access-control standard requires technical policies and procedures allowing access only to those persons or software programs that have been granted access rights. Unique user identification and an emergency access procedure are Required. Automatic logoff and encryption and decryption are Addressable. Any page telling you HIPAA requires encryption is misreading the rule — and the phrase persons or software programs already reaches an AI agent without amendment.
Also Required, under 164.308(a)(1)(ii)(D): procedures to regularly review records of information system activity such as audit logs, access reports and security incident tracking reports. Access authorization and access establishment and modification, under 164.308(a)(4), are Addressable, as are the three workforce security specifications.
Addressable does not mean optional. It means a covered entity must assess whether the specification is reasonable and appropriate, implement it if so, and document the decision either way. It does mean that a control narrative should say which analysis was performed rather than asserting a requirement that is not in the text.
GDPR Article 5 supplies the classification argument in one sentence. Article 5(1)(c) requires data adequate, relevant and limited to what is necessary; 5(1)(f) requires processing in a manner that ensures appropriate security; and 5(2) makes the controller responsible for, and able to demonstrate compliance with, all of it. You cannot demonstrate compliance over content you have never categorised.
Article 32(1) requires appropriate technical and organisational measures appropriate to the risk, expressly including pseudonymisation and encryption, and a process for regularly testing, assessing and evaluating the effectiveness of those measures. Unlike HIPAA, GDPR has a built-in requirement for periodic control-effectiveness review, which is exactly what a retrieval test provides evidence for.
On protected health information the practical rule is blunt: identifiable subject data should not sit in a general-purpose collaboration store, because those are precisely the stores an AI grounding layer indexes. Where it must be there, de-identify first. HHS Office for Civil Rights recognises two routes, Expert Determination and Safe Harbor. GDPR does not recognise Safe Harbor: pseudonymised data remains personal data, so an EU-facing programme cannot treat a HIPAA-de-identified set as out of scope.
One configuration follows directly from the Copilot gaps documented above. A restricted tier that applies encryption without the EXTRACT usage right is the only Purview configuration that reliably stops Copilot summarising PHI, because label-based Copilot DLP is E5-only, still shows citations, and is bypassed entirely by direct upload into a prompt.
What to do instead of assuming: test what the assistant can actually retrieve
Every finding on this page points to the same method. Do not audit the configuration and infer the exposure. Ask the assistant, as a real person, for the things that would hurt, and record what comes back. Configuration review explains a result; it is a poor substitute for producing one.
Method
Ask the assistant, not the admin console
A configuration review tells you what should happen. A retrieval test tells you what does. The two disagree often enough — because of container labels, connector content, propagation windows, licence gaps and permissions nobody remembers granting — that only the second is worth putting in front of a board or a partner.
The test is built from personas rather than accounts. A persona is a role with a plausible permission set: a clinical operations associate, a contract CRA with guest access, a commercial analyst, a finance manager, a departing employee in their notice period, an external collaborator on one project. For each persona, the exercise runs a fixed set of retrieval attempts against the content classes that carry regulatory or commercial weight, and records four things: what was asked, what the assistant returned, what it cited without returning, and which control produced the result.
That fourth column is where the value is. A blocked result explained by a label with no EXTRACT right is a durable control. A blocked result explained by the fact that nobody has yet indexed the site is not a control at all. Recording the mechanism rather than the outcome is what makes the test repeatable after the next tenant change.
Include the paths the documentation tells you are open. Upload a document directly into a prompt and confirm that DLP did not evaluate it. Ask a question whose answer sits in a Teams chat and observe that no label is displayed. Query content reached through a Graph connector and check whether the source system’s classification survived the journey. These are not clever attacks; they are the scenarios Microsoft has already documented, tested against your tenant.
Then re-test after the propagation windows have elapsed. A four-hour Copilot DLP delay and a 24-hour label replication window will otherwise produce a false negative that is worse than no test, because it is written down.
Personas, not accounts
Six to ten role profiles that reflect real permission sets, including guests and departing staff.
Content classes, not folders
PHI, trial master file content, CMC and manufacturing, submissions, discovery data and intellectual property, HR, board material.
Record the mechanism
Which control produced each result. An accident is not a control, even when the outcome looks right.
Respect the clocks
Re-test after four hours for Copilot DLP and after 24 hours for label and policy replication.
Related evidence and next steps
AI access exposure review— The engagement that produces this evidence, across Microsoft 365, Box and Egnyte.
CISA Zero Trust Maturity Model v2.0 (PDF)— The Data Categorization and Data Access maturity ladders, useful as an external maturity anchor. The cover reads April 2023 while the internal revision table reads March 2022.
A remediation ladder ordered by exposure, not by product feature
Findings arrive unordered. Fixing them in the order the portal presents them is how a programme spends its first quarter on label taxonomy while an over-shared site index stays open. The ordering principle is simple: close what is reachable today, then reduce what is reachable in principle, then build the classification that keeps it closed.
First, remove the reachable exposure. Over-broad site permissions, orphaned guest access, everyone-except-external-users grants and abandoned sites are corrected directly, because no labelling design changes what they expose. Microsoft’s own access-review rationale is the cleanest published statement of why: excessive access rights can lead to compromises, and excessive access rights can also lead to audit findings as they indicate a lack of control over access.
Second, contain what cannot be fixed this quarter. Restricted Content Discovery, where the licence exists, keeps a site out of grounding while its permissions are repaired, with the explicit understanding that it is not a security boundary and comes back off on a named date. Retention and lifecycle policies retire content that has not been touched in years and should not be in an index at all.
Third, build the classification model — four or five tiers, named for relative sensitivity, applied across the content map rather than mirroring it. Start without encryption, following Microsoft’s own ladder, and add encryption only where the receiving systems have been tested against the failure surface documented above.
Fourth, install the periodic review that keeps the model true. PIC/S expects systems to generate a list of users with actual access for use during periodic user reviews; Annex 11 §12.3 expects creation, change and cancellation of access authorisations to be recorded; GDPR Article 32 expects regular testing of control effectiveness. Entra access reviews and entitlement management, licensed appropriately, are a reasonable implementation of all three, with recurrence set weekly, monthly, quarterly or annually.
Fifth, re-run the retrieval test. The same personas, the same questions, the same four columns. A remediation programme that cannot show a before-and-after on identical questions has produced activity rather than assurance.
Close what is reachable. Contain what is not yet fixable. Classify what remains. Review it on a schedule. Then prove it with the same test you started with.
Most of the life-sciences companies we speak to are not a pure Microsoft estate. Regulated content is split across SharePoint, a Box or Egnyte tenant, an eTMF, a QMS and a handful of CRO portals. Purview governs one of those. The readiness question is identical in each, and only the mechanism changes.
Box binds enforcement to classification through Shield access policies rather than through the label itself. A Box tenant supports up to 25 classification labels, files and folders carry one label at a time, only defined internal roles may change a classification, and external users can never modify one. Deleting a classification is permanent and strips it from every file and folder it was applied to. The enforceable controls attached to a label are external collaboration restriction, shared link restriction, download and print restriction separable for managed and external users across web, mobile and desktop, integration restriction, FTP restriction, watermarking, and signature-request restriction.
Two Box details matter for anyone auditing an older tenant. Classification originally shipped in Box Governance, and since October 2020 the security features in Box Governance — classification, shared link restriction and content security — have not been available to new Box Governance customers. Box further documents that the ability to create, edit and delete content security policies and shared link policies was sunset on 6 February 2026, with those policies disabled for all customers in May 2026, while security classification continues to be available. On a current tenant, classification-driven enforcement means Shield.
The integration restriction is the Box control most relevant to this page and the one most often left unset. It is what stops a connected application or agent downloading the content, which is precisely the mechanism by which an assistant reaches a file it should not summarise.
Egnyte works the other way around: it is a folder-ACL model in which classification informs and permissions enforce. Content classification lives in Egnyte Secure and Govern and is switched off by default for privacy reasons. Permissions inherit downward unless a more specific permission exists on a subfolder, inheritance can be broken per folder without cascading, and the default permission when adding a user is Viewer. Egnyte’s own guidance is a least-privilege statement: start with minimal access at the top-level folders and assign permissions or ownership as needed at lower levels.
Egnyte also documents the failure mode that produces an unexplainable access map three years later. Where a folder owner grants a permission on a parent but only owns part of the tree beneath it, the permissions change only as far as the topmost subfolders where the grantor loses owner permission. The grant silently under-applies, and nothing reports it as partial.
The regulatory texts themselves, rather than a summary of them
Every regulatory statement on this page comes from a primary text, and the texts are short enough to read. We link them here rather than paraphrasing them further, because the paraphrase is where errors enter — encryption becomes required, an addressable specification becomes mandatory, an archive period becomes a retention policy, and a control narrative inherits a mistake nobody can trace.
One sourcing caveat belongs on the record. The official Electronic Code of Federal Regulations restricts programmatic access, so the CFR text underlying this page was read through the Cornell Legal Information Institute mirror. Cornell is faithful and widely used, but it is not the official source and does not display a current-as-of date. Before a citation reaches a validation document or an inspection response, open the eCFR page and confirm both the wording and its currency.
On the same principle, this page makes no claim about pending rulemaking. Proposals to amend the HIPAA Security Rule have been discussed publicly, and their status was not verified here. Design against the rule as it currently reads, and track proposals separately rather than pre-emptively treating a proposal as the standard.
For content taxonomy in clinical development, the TMF Reference Model published by CDISC is the canonical structure and is worth linking for a different reason than the others: it is the model most often mistaken for a sensitivity taxonomy. It is a content taxonomy of zones, sections and artifacts, and it belongs in the eTMF system’s metadata. The sensitivity taxonomy stays at four or five tiers and is applied across it. Trying to make labels mirror the reference model is the most common design error in life-sciences classification projects.
The retention clock in that same domain is what makes an encryption decision consequential. EU Clinical Trials Regulation Article 58 requires the sponsor and investigator to archive the trial master file content for at least 25 years after the end of the trial, kept readily available and accessible to competent authorities, with any alteration traceable. A label applying encryption with a tenant key must remain decryptable for a quarter of a century, by people who do not work at the company yet. That is a key-custody and super-user commitment, and it deserves an explicit answer before the label is published, not after.
CDISC TMF Reference Model— Eleven zones, sections and artifacts. A content taxonomy, not a sensitivity taxonomy.
Scope and boundaries
What this engagement is, and what it is not
IntuitionLabs is an AI consultancy for pharmaceutical and biotechnology companies, founded in 2023. This work sits at the intersection of enterprise AI deployment, information architecture and regulated-industry expectations. It is deliberately narrow, and being explicit about the boundary is more useful than a broader claim.
We prepare clients for audits, certifications and security questionnaires. We are not an auditor and not a certifying body, and we never issue certifications. Nothing in a readiness review certifies a tenant as Part 11 compliant, HIPAA compliant or GDPR compliant, and no vendor including Microsoft can do that on your behalf, because compliance attaches to your intended use and your quality system rather than to a product configuration.
IntuitionLabs does not hold SOC 2 or ISO 27001. We say so on every security engagement, because a consultancy that is vague about its own posture while auditing yours has already answered the more interesting question. Our own security posture and the documents we can share are published on our trust pages rather than described here.
Where deeper security architecture review is warranted, the engagement is led by a named security architect on our expert bank: an independent consultant with seventeen years in information security, including a decade on the central security team of a major enterprise infrastructure vendor across design and architecture review, security requirements review, source code review, penetration testing and vulnerability response, and earlier client-facing consultancy work leading mobile penetration testing. The specialist is named to you in the engagement letter, is a specialist on our expert bank rather than a member of staff, and on a review of this kind did not build what they are reviewing.
Finally, a note on currency. Purview documentation and licensing change quickly, and the service description SKU names moved materially during 2025 and 2026. Every licensing statement on this page reflects the Microsoft Purview service description as read on 27 August 2026. Re-verify against the live page before making a purchasing decision, and treat any third-party summary — including this one — as a starting point rather than an authority.
The AI security service line— How exposure review, classification design, vendor assessment and questionnaire readiness fit together.
How a Copilot readiness review runs
Six weeks is typical for a mid-size tenant, and the shape does not change much with size. The output is evidence about your environment, not a maturity score. Every finding names the Microsoft or vendor documentation that explains it, so your own team can verify the conclusion without us.
Entitlement and configuration baseline
What your licences actually include, which Purview capabilities are enabled, whether container labelling and SharePoint label integration were ever turned on, and where a Plan 1 assignment is missing alongside Plan 2.
Where PHI, trial master file content, CMC records, submissions, discovery data, HR and board material actually live today, including the copies in personal OneDrive and the CRO portals nobody inventoried.
Fixed questions asked as six to ten role personas, recording what returned, what was cited without returning, and which control produced each result. Re-run after the documented propagation windows.
Each finding paired with the primary source that explains it, separating a genuine control from an accident of indexing, and separating an E5 gap from a configuration gap so budget conversations stay honest.
Ordered by exposure rather than by product feature: close what is reachable, contain what is not yet fixable with a removal date, classify what remains, then install the periodic review that keeps it true.
Findings phrased so they answer the questions a partner, investor or large pharma customer will actually ask, with the boundary between assessed and certified stated plainly rather than blurred.
It makes one safe if the underlying permission model is already correct, and it does not fix a permission model that is wrong. Microsoft states that Copilot only surfaces organizational data to which individual users have at least view permissions, and in the same documentation asks customers to make sure the permission models in SharePoint and other services already give the right users access to the right content. Purview adds labelling, data loss prevention, lifecycle management and insider risk detection on top of that permission model. It does not replace it, and Microsoft publishes a substantial list of Copilot scenarios in which labels and label-based DLP are not recognised at all.
Per the Microsoft Purview service description, E3-class licences include manual sensitivity labelling, DLP for Exchange Online, SharePoint Online and OneDrive, org-wide retention policies, and Audit (Standard). Every automation layer sits above that line: client-side and service-side automatic labelling, adaptive policy scopes, auto-applied retention labels, records and regulatory-record settings, Teams and endpoint DLP, Insider Risk Management, and the label-based DLP rule that stops Copilot processing an item. The service description also states plainly that scanner-based discovery is supported with a Microsoft 365 E3 license while sensitivity labeling, including automatic or policy-based labeling, requires a Microsoft 365 E5 license or Microsoft 365 Information Protection and Governance. SKU names in this area moved during 2025 and 2026, so confirm your own entitlement against the current service description rather than a summary.
No, and this is the most common misunderstanding we encounter. A container label applied to a Team, Microsoft 365 Group or SharePoint site sets privacy, guest access, external sharing, unmanaged-device access, authentication context and discoverability for the container. Microsoft documents that items in these containers do not inherit the labels and therefore do not apply any item-level label settings such as content markings and encryption. The consequence for an AI deployment is direct: because the items carry no label, they will not display their container label in Copilot and cannot support sensitivity label inheritance.
Only partly, and the gap is important. Microsoft documents that DLP cannot scan the contents of files that a user uploads directly into a prompt, so no evaluation of the uploaded file for sensitive data occurs; DLP only checks the text typed into the prompt itself. Prompt DLP is available to all users of Microsoft Copilot and Copilot Chat, but the rule that stops Copilot processing a labelled file or email is an E5-class capability. A company that relies on label-based DLP as its only control has left the direct-upload path open regardless of licence.
Microsoft documents that Restricted SharePoint Search is retiring and that new enablement was blocked starting 31 July 2026. Its replacement, Restricted Content Discovery, requires both a Copilot licence and SharePoint Advanced Management. Neither is a security boundary: Microsoft states that Restricted Content Discovery does not change existing permissions and does not remove content from the Microsoft 365 search index. It also warns that on sites with more than 500,000 items an update could take more than a week to fully process, so it is a containment measure while permissions are fixed, not a substitute for fixing them.
Almost certainly not as a first move, and Microsoft is unusually candid about why. Its own current deployment model puts encryption last: a default Confidential label without encryption, blocked from external sharing by DLP, then extension of that default to all files, and only then encryption on the default label. Microsoft also documents a long list of things that break when files are labelled with encryption in SharePoint and OneDrive, including encrypted Office files over 12 MB moved between sites, files containing Power Query data, custom XML parts, a bibliography or a SharePoint Document ID, and any label that configures content expiry or Double Key Encryption. For a company with a 25-year archive obligation, encryption is also a key-custody commitment, not just a usability question.
Part 11 §11.10(d) requires limiting system access to authorized individuals, and §11.10(g) separately requires authority checks to ensure that only authorized individuals can use the system, sign a record, access the input or output device, alter a record, or perform the operation at hand. Those are two distinct requirements, and a folder permission satisfies only the first. Purview labelling and DLP contribute evidence toward both, but the operation-level authority check normally lives in the validated application, not in the collaboration store. Purview also does not perform system validation under §11.10(a), and Microsoft does not certify a tenant as Part 11 compliant.
Not as written. In 45 CFR 164.312(a)(1), the technical access-control standard requires implementing technical policies and procedures to allow access only to those persons or software programs that have been granted access rights. Under that standard, unique user identification and an emergency access procedure are marked Required, while automatic logoff and encryption and decryption are marked Addressable. Addressable does not mean optional in practice, because a covered entity must assess and document its decision, but it does mean that a page telling you HIPAA requires encryption is misreading the rule. Note also the phrase persons or software programs: the text already reaches an AI agent without amendment.
Longer than most teams assume, which matters when a control is being tested or an incident is being contained. Microsoft documents that updates to a DLP policy can take up to four hours to reflect in the Microsoft 365 Copilot and Copilot Chat experience, and that if a sensitivity label is applied mid-session the policy is enforced starting the next time the file is opened. Sensitivity label and policy changes more generally can take up to 24 hours to replicate. Build these delays into the test plan rather than concluding that a control failed.
A tested account of what your assistant can retrieve, for a defined set of personas, against the content classes that carry regulatory weight. That means documented retrieval attempts and their results, the tenant configuration and licence entitlements that explain those results, the specific gaps between what the organisation believed was protected and what is, and a remediation sequence ordered by exposure rather than by product feature. It is an assessment, not a certification: IntuitionLabs prepares clients for audits, certifications and security questionnaires, and is not an auditor or a certifying body.
Purview is a Microsoft 365 control plane, so if your regulated content lives primarily in Box or Egnyte the equivalent question is which of those platforms enforce, and how. Box binds enforcement to classification through Shield access policies covering external collaboration, shared links, download and print, integrations and watermarking, and supports up to 25 classification labels with one label per item. Egnyte is a folder-ACL model in which classification informs and permissions enforce, and its content classification is switched off by default. The same readiness question applies in all three: what can the assistant actually retrieve as a given person, and which control stopped it.
Not to the depth most quality and legal teams expect. Microsoft states that auditing captures the Microsoft 365 Copilot activity of search but not the actual user prompt or response, and that audit is not intended to be used as the basis for Copilot usage reporting. Prompts and responses are reachable through eDiscovery or the DSPM for AI activity explorer instead. If your control narrative depends on reconstructing what an assistant told a specific user on a specific day, verify that the retrieval path exists and is retained for the period you need before you write the narrative.
Find Out What Your Assistant Can Actually Retrieve
Bring your licence mix, the stores where regulated content lives, and the personas you worry about. We will tell you what a retrieval test would cover, what Purview will and will not resolve in your environment, and what the remediation sequence looks like before anyone signs anything.