Claude

IntuitionLabs is now a member of the Claude Partner Network – AI training and upskilling with Claude for pharma and biotech. Book a call.

IntuitionLabs
Life sciences information security team reviewing a document classification and access model

Document Classification and Access Model Design

A four-tier policy in a Word document changes nothing an AI assistant can retrieve. We design the label taxonomy, the identity groups underneath it, the platform controls wired to both, and the migration of the content you already have.

Four layers, built in this order

Each layer is useless without the one before it. A taxonomy with no groups behind it cannot be enforced. Groups with no platform configuration behind them are documentation. Platform configuration applied to unmigrated content protects the files created after today and nothing else.

01
The taxonomy
Four or five tiers sized to what people will actually apply, named so the relative sensitivity is obvious, and mapped to the content classes a life-sciences company genuinely holds.
See the security practice
02
The identity groups
The Entra or Active Directory structure that each tier resolves to. Least privilege by design, membership that can be explained, and a review cadence that produces evidence.
Find the exposure first
03
Platform enforcement
Box Shield policies, Egnyte folder permissions, SharePoint container settings and Purview labels, configured against the group structure rather than against individual names.
Purview in detail
04
Content migration
The unglamorous part. Scoped pilots, simulation runs, auto-labeling where a rule can express the class, and manual triage where it cannot. Then a standing review.
The information layer

A classification policy is not a classification model

Almost every life-sciences company we speak to has a data classification policy. It is usually four tiers, it is usually two pages, and it is usually accurate. It also usually has no corresponding object anywhere in Box, Egnyte, SharePoint or Entra. That gap is invisible until something starts reading content at machine speed, at which point it becomes the whole problem.

The gap

A policy document does not change what an assistant can retrieve

A classification policy describes intent. An access model produces behaviour. The distinction is easy to state and easy to lose, because a signed policy feels like a completed control and generates the same paperwork an implemented control would generate. Microsoft says the quiet part directly in its own service assurance material: data classification levels by themselves are simply labels or tags that indicate the value or sensitivity of the content. The protection comes from what reads the label.

When a company connects an enterprise assistant to its content stores, the assistant inherits the effective permission map rather than the policy. Microsoft states its own position without hedging: Copilot only surfaces organizational data to which individual users have at least view permissions, and the grounding layer honours that boundary. Read as reassurance, that sentence closes the conversation. Read correctly, it opens it, because the sentence is a statement about permissions and says nothing at all about whether those permissions are right. Microsoft makes the corollary explicit elsewhere, noting that because of the power and speed with which AI can proactively surface content, generative AI amplifies the problem and risk of oversharing.

The mechanism is worth stating plainly, because it is not an attack and there is no vulnerability involved. A permission granted to an individual for one project in 2021, on a folder that later became a parent of other folders, is still granted. A group created for a study that closed still has members. A departing employee whose account was disabled leaves behind the group memberships that survived them. A guest account created for a CRO transition was never removed. Each of these is individually defensible and collectively they form a permission map nobody has read end to end. Human search never surfaced most of it because humans navigate by folder and by memory. A retrieval layer navigates by index.

At a clinical-stage biotech client, during a first engagement, connected AI reached employee-information files in Box through existing user permissions. Nothing was misconfigured in the AI tool. No control failed. The files were reachable by those users before the assistant existed, and the assistant simply made that reachability legible for the first time. That is the shape of the finding in almost every environment we look at, and it is why the remediation is a content and identity project rather than an AI project.

The consequence for scoping is that a classification engagement cannot begin with the label set. It begins with the current state of permissions, because the taxonomy you design has to be applied to content whose access map you understand. Designing five beautiful tiers and then discovering that half the estate has broken permission inheritance in ways nobody can explain is a common and avoidable sequencing error.

What a policy produces

A shared vocabulary, an audit artefact, and an agreed statement of who should see what.

What a policy does not produce

A label on a file, a group in the directory, or a rule in the platform that reads either one.

What AI changes

Nothing about the permissions. Only the speed and completeness with which they are exercised.

Where to start

The current access map, not the target taxonomy. The taxonomy has to land on real content.

The AI did not create the exposure. It read the permission map faster and more completely than any person ever had.

Related evidence and next steps

Sizing

Five labels, not fifty

The most common design failure in a classification project is a taxonomy that is technically elegant and operationally unusable. It happens because the people designing it are the people who understand the distinctions, and they are not the people who will apply them at eleven at night before a submission deadline. Three independent sources converge on a much smaller number than most first drafts propose, and the convergence is the strongest evidence-backed recommendation available on this subject.

Microsoft product documentation reports field observation rather than theory: real-world deployments show that effectiveness is noticeably reduced when users have more than five main labels or more than five sublabels per main label. It adds a practical constraint on top of the human one, noting that some applications cannot display all your labels when too many are published to the same user. Microsoft service assurance guidance arrives independently at the same shape, stating that a data classification framework is typically comprised of three to five classification levels and recommending no more than five top-level parent labels, each with five sublabels, to keep the interface manageable.

Box reaches a similar place from a different direction. Its platform permits up to twenty-five classification labels and assigns exactly one label per file or folder, and its own general recommendation is three: one for public content, one for content intended to stay inside the organisation, and one for content requiring specific authorisation. Box goes further than most vendors in the direction of restraint, suggesting that for genuinely public content some organisations keep the scheme simple by not classifying it at all. Two vendors with different architectures and different commercial incentives land within two labels of each other.

Naming matters as much as counting, and it is cheaper to get right at design time than to change later. Microsoft recommends label names that are self-descriptive and that highlight relative sensitivity clearly, giving the specific example that Confidential and Restricted may leave users guessing which is more sensitive, while Confidential and Highly Confidential are unambiguous. Microsoft also publishes a useful piece of its own deployment history: its corporate framework originally used a label named Internal during the pilot phase, found there were legitimate reasons for a document to be shared externally, and shifted to General. A label that quietly forbids a legitimate business action gets ignored, then resented, then bypassed.

The number that should govern the project plan, though, is not a label count. Microsoft states that customers indicate approximately fifty percent of an information protection project is business focused rather than technical, and that end-user training and communication is therefore critical to success. Half. If a proposal for this work allocates ninety percent of its effort to configuration, it is not a plan for a classification model; it is a plan for a set of labels that will be applied inconsistently and trusted accordingly.

Microsoft field limit

Effectiveness noticeably reduced above five main labels or five sublabels per label.

Microsoft framework guidance

Three to five levels; no more than five parents with five sublabels each.

Box platform shape

Twenty-five labels supported, one per item, three recommended in general practice.

The real budget line

Around half the project is business change, per Microsoft own customer reporting.

Related evidence and next steps

Content classes

What a life-sciences company actually holds

A generic four-tier scheme fails in life sciences for a specific reason: the organisation is not one records environment but several, with retention clocks that do not reconcile and audiences that do not overlap. A clinical-stage biotech is simultaneously a clinical research organisation, a manufacturer, a regulatory filer, a research laboratory and a public company. The taxonomy has to sit above all of those without pretending they are the same.

The classes worth naming explicitly are identifiable subject data, clinical study documentation, chemistry and manufacturing content, regulatory submissions, discovery data and intellectual property, commercial and medical affairs material, human resources records, and board and financial content. Each of these has a different regulatory anchor. Clinical study documentation answers to ICH E6 and, in the European Union, to the Clinical Trials Regulation. Manufacturing content answers to EU GMP Annex 11 and to PIC/S guidance on data management and integrity. Submissions answer to 21 CFR Part 11. Discovery data answers to no prescriptive regulation at all and to trade-secret law and patent strategy instead, which makes it the class most often left out of a scheme designed by a quality function.

The retention clocks are the part that constrains design most severely, because they outlast the technology. Under Article 58 of the EU Clinical Trials Regulation the sponsor and the investigator shall archive the content of the clinical trial master file for at least twenty-five years after the end of the clinical trial, kept readily available and accessible upon request to the competent authorities, with any alteration to the content traceable. In the United States, 21 CFR 312.62(c) requires an investigator to retain records for two years following approval of a marketing application for the indication, or two years after the investigation is discontinued and FDA is notified. Annex 11 section 7.1 adds that access to data should be ensured throughout the retention period, and section 17 that archived data should be checked for accessibility, readability and integrity.

There is a strong temptation to make the sensitivity taxonomy mirror the content taxonomy, and it should be resisted. The Trial Master File Reference Model, now published by CDISC, is the canonical content taxonomy for clinical development: eleven zones covering trial management, central trial documents, regulatory, approvals, site management, investigational product, safety reporting, testing, third parties, data management and statistics, subdivided into sections and then into artifacts with unique identifiers, each marked core or recommended, and existing at trial, country and site level. That is a document model and it belongs in the eTMF system metadata. The sensitivity taxonomy stays at four or five tiers and is applied across it. Conflating the two produces a label list with hundreds of entries and a workforce that stops labelling.

The mapping exercise, then, is not one-to-one. It is a matrix: each content class is assigned a default tier, a small number of documented exceptions where a subset escalates, and an owner who can adjudicate. Discovery chemistry might be Highly Confidential by default with no exceptions. Commercial content might be General by default with escalation for pre-launch pricing. Manufacturing content might be Confidential with an escalation for process detail that constitutes trade secret. Writing that matrix takes a week of workshops with people who own the content, and skipping it is the reason so many label sets go unused.

Clinical documentation

EU CTR Article 58 sets a twenty-five year archive clock that outlives most technology decisions.

Investigator records

21 CFR 312.62(c) sets a two-year clock keyed to approval or to discontinuation and FDA notification.

Discovery and IP

No prescriptive regulation, indefinite retention, and the class most often omitted from a quality-led scheme.

The separation rule

Content taxonomy lives in the system of record. Sensitivity taxonomy stays at four or five tiers above it.

Related evidence and next steps

PHI

The one class that should not be in the collaboration store

Every other class in the matrix is a question of which tier. Identifiable subject data is a question of whether it should be in a general-purpose collaboration platform at all. The answer, in almost every case we have examined, is no, and the reasoning is specific rather than reflexive.

Under 45 CFR 164.312(a)(1), a covered entity must implement technical policies and procedures for electronic information systems that maintain electronic protected health information to allow access only to those persons or software programs that have been granted access rights. The phrase software programs is doing real work in 2026. An AI agent is a software program. The Security Rule reaches it as written, without amendment. That means the question is not whether an assistant is permitted to see protected health information; it is whether the access rights granted to it were granted deliberately, which they generally were not, because they were inherited from a user account whose folder permissions predate the assistant.

Box arrives at the same conclusion from an operational direction and publishes it as a healthcare pattern: many organisations settle on a basic schema plus specific categorisation for content containing protected health information. In other words, PHI is not a level of confidentiality, it is a class of content that needs its own handling regardless of how confidential the surrounding material is. Microsoft own reference framework names it too, placing protected health information alongside sensitive personally identifiable information, cardholder data and bank account data at the Highly Confidential level. Two vendors, one giving PHI a dedicated category and one placing it at the top tier, and both treating it as a named exception rather than an ordinary case.

If identifiable subject data genuinely has to live in the collaboration store, de-identify it first. The HHS Office for Civil Rights documents exactly two methods that satisfy the Privacy Rule de-identification standard: Expert Determination and Safe Harbor. Safe Harbor requires removal of the enumerated identifiers together with the absence of actual knowledge that the residual information could identify an individual. One important caveat for any company with European operations: the GDPR does not recognise Safe Harbor, and pseudonymised data remains personal data under the GDPR. A dataset that is properly de-identified for HIPAA purposes is not thereby out of scope for European obligations, and treating it as such is a common and expensive mistake.

Where a Restricted or PHI tier is implemented in Purview, the configuration that reliably keeps an AI assistant away from the content is a label applying encryption that does not grant the EXTRACT usage right. Microsoft documents the resulting behaviour precisely: without EXTRACT, Copilot will not summarize the content but can reference it with a link, and when a user has that content open in an application they cannot use Copilot at all. The alternative levers are weaker. Label-based DLP for the Copilot location is E5-class, blocked items still appear in the citations of a response, and Microsoft states that DLP cannot scan the contents of files uploaded directly into a prompt, so a determined user with a copy on their desktop bypasses it entirely.

The statutory hook

HIPAA 164.312(a)(1) covers persons or software programs. An AI agent is already in scope.

The vendor consensus

Box gives PHI its own category; Microsoft places it at Highly Confidential. Both treat it as exceptional.

The de-identification routes

Expert Determination or Safe Harbor under 45 CFR 164.514(b). Nothing else satisfies the standard.

The GDPR caveat

Pseudonymised data is still personal data. HIPAA de-identification does not clear European obligations.

PHI is not a confidentiality level. It is a class of content that needs its own handling wherever it appears.

Related evidence and next steps

Three platforms, three different enforcement models

Vendor documentation as of 27 August 2026. These are architectural differences rather than quality rankings, and in practice the correct platform is usually the one already deployed. What changes between them is which design decisions are available to you.

Design questionBoxEgnyteSharePoint with Microsoft Purview
Primary modelClassification label plus Shield access policyFolder access-control list; classification is advisorySensitivity label plus container settings plus DLP
Label ceilingUp to 25 classification labels, one per file or folderNo label-driven enforcement; permissions decide1,000+ labels supported, 500 where the label specifies users and permissions
Is classification on by defaultLabels must be created; Shield policies must then be bound to themContent classification toggle is off by default, described by Egnyte as a privacy choiceLabels must be created and published; container labelling needs separate enablement
Who can change a labelA global setting selects which collaborator roles may change it; external users never canNot applicable; folder ownership and permission level governAny user with the label published to them, subject to justification on downgrade
Strongest enforceable controlShield integration restriction blocking all integrations from downloading contentLeast-privilege folder permissions with inheritance broken deliberatelyEncryption applied by the label, with usage rights including EXTRACT
Automated classificationBox Shield automated classification on upload, update, move and copyJurisdiction-scoped policies over included scan paths onlyClient-side and service-side auto-labeling, gated by E5-class licensing
Main licensing trapClassification moved from Box Governance to Box Shield; Governance security policies are sunsettingClassification sits in the Secure and Govern tier; confirm entitlement before designManual labelling is E3-class; almost every automation layer is E5-class
Documented failure to plan forDeleting a classification is permanent and strips it from every file and folderA partial-ownership grant applies only as far as the grantor owns the subfoldersOver-encryption breaking co-authoring, search, eDiscovery and integrations
Where it lands for AIIntegration restriction is the control that stops a connected app pulling bytesPermissions are the only real control; classification informs triageEXTRACT-less encryption is the control; label-based Copilot DLP is E5 and partial

The identity layer is the access model

Labels describe content. Groups decide who reads it. Every enforceable control on every one of these platforms resolves, eventually, to an identity object, which is why a classification project that does not touch the directory has not produced an access model. This is also the layer where regulated life sciences has clearer, older and more specific obligations than most industries realise.

Group design

The directory is where a taxonomy becomes enforceable

A tier called Confidential is a word until something resolves it to a set of people. In practice that resolution happens through a security group, and the quality of the group structure sets a hard ceiling on the quality of everything above it. A well-named group with unexplainable membership is worse than no group, because it creates the appearance of control while granting access nobody reviewed.

The design work has three parts. The first is deciding which groups exist and what each one means, expressed as a sentence a non-specialist can check: not GRP-SP-CLIN-RW-02 but the people who may edit clinical study documentation for active studies. The second is deciding how membership is determined, which is the choice between assigned membership that someone maintains and dynamic membership driven by directory attributes. The third is deciding what reviews that membership and how often, which is the part almost always deferred and the part regulators actually ask about.

Dynamic membership groups are the standard answer to least privilege at scale, and they carry real constraints worth knowing before the design rather than after. A Microsoft Entra tenant supports a maximum of fifteen thousand dynamic membership groups. The membership rule body cannot exceed 3,072 characters. Users and devices cannot be mixed in a single dynamic group, and a device rule cannot reference the owner user attributes. You cannot manually add or remove a member of a dynamic group; the rule is the only source of truth, which is a feature until the day you need an exception. Microsoft also advises minimising the use of match and contains operators because they degrade processing time.

The failure mode that matters most in a regulated organisation is one Microsoft states directly and that very few design documents account for: the security of a dynamic group membership depends on who can modify the attributes referenced in the rule. Attributes synchronised from on-premises Active Directory may carry self-write permissions, which means a user can in principle edit themselves into an access group by editing their own department or job title. Microsoft frames the consequence bluntly, noting that the security of that access is only as strong as the write controls on the attributes in the rule, and points to the structural fix: role-assignable groups already prevent this risk by requiring assigned rather than dynamic membership.

The practical output of this workstream is therefore not just a group list. It is a group list, a written statement of how each membership is determined, an attribute-permission audit for every attribute used in a dynamic rule at both the cloud and on-premises source, and an explicit decision about which groups are too sensitive to be dynamic at all. That last decision is usually short: administrative groups, anything gating Highly Confidential content, and anything gating identifiable subject data.

Tenant ceiling

A maximum of 15,000 dynamic membership groups per Microsoft Entra tenant.

Rule ceiling

The membership rule body cannot exceed 3,072 characters, and match or contains operators slow processing.

No manual override

Members cannot be added or removed by hand in a dynamic group. The rule is the only source of truth.

The real control

Who can write the attribute the rule reads. Audit that at both Entra and the on-premises source.

Related evidence and next steps

Access reviews

PIC/S PI 041-1 already requires the review you are not running

Periodic user access review is usually presented to a life-sciences client as an information security good practice, which makes it optional in the eyes of anyone budgeting. It is not optional. It is written into GxP data-management guidance in language specific enough to design against, and this is the strongest available bridge between a security programme and a quality system.

PIC/S PI 041-1, Good Practices for Data Management and Integrity in Regulated GMP and GDP Environments, dated 1 July 2021, is the most operationally specific access-control text in GxP. It states that user access controls shall be configured and enforced to prohibit unauthorised access to, changes to and deletion of data. It states that systems should support different user access roles and that assignment of a role should follow the least-privilege rule, assigning the minimum necessary access level for any job function. It states that administrator access rights should not be given to normal users, as a matter of segregation of duties, and that system administrators should normally be independent from users performing the task with no involvement or interest in the outcome of the data generated. And it states, in a single sentence that maps directly onto a Microsoft Entra access review, that systems should be able to generate a list of users with actual access including user identification and roles, and that the list should be used during periodic user reviews.

The World Health Organization reaches the same place in Technical Report Series No. 1033, Annex 4, its 2021 guideline on data integrity. It requires a documented system defining the access and privileges of users, requires that access and privileges be in accordance with the role and responsibility of the individual, requires that a limited number of personnel with no conflict of interest be appointed as system administrators, and requires that inactivated users be retained in the system with a list of active and inactivated users maintained throughout the system life cycle. It also prohibits shared logins or generic user access for systems generating, amending or storing GxP data, as does PIC/S, which prohibits shared passwords even for reasons of financial savings.

Annex 11 makes the evidence requirement explicit in one line at section 12.3: creation, change and cancellation of access authorisations should be recorded. That is not a request for a policy. It is a request for a record, produced each time access changes, that an inspector can read. Most environments we review can produce the current state and cannot produce the history, which is a gap that costs nothing to close going forward and a great deal to reconstruct backwards.

On the implementation side, Microsoft Entra access reviews cover security-group and Microsoft 365 group membership, enterprise-application assignment, Entra and Azure resource roles through Privileged Identity Management, and access-package assignments, with reviewers configurable as named people, group owners or the users themselves, on a weekly, monthly, quarterly or annual recurrence. Microsoft frames the rationale in terms an auditor recognises: excessive access rights can lead to compromises, and excessive access rights can also lead to audit findings as they indicate a lack of control over access. The capability requires a Microsoft Entra ID Governance or Entra Suite subscription, which is a budget line that has to be raised at design time rather than discovered at implementation.

PIC/S PI 041-1

Least privilege by role, segregation of administrator duties, and a user list used in periodic reviews.

WHO TRS 1033 Annex 4

Documented access system, retained inactivated users, and no shared or generic logins for GxP data.

Annex 11 section 12.3

Creation, change and cancellation of access authorisations should be recorded. A record, not a policy.

The tooling

Entra access reviews implement it, on an ID Governance or Entra Suite entitlement.

The regulator asked for periodic user access review before anyone connected an assistant to the file share. The AI deployment only made the omission visible.

Related evidence and next steps

Drift

Groups rot, and the rot has a predictable shape

An access model is not a static artefact. It degrades, and it degrades in ways specific enough to monitor for. Microsoft own rationale for access reviews lists the failure patterns, and each one describes an environment we have looked at.

The first pattern is accumulation in privileged roles, where too many users hold administrative rights because granting them was the fastest way to resolve an incident. The second is repurposing, where a group created for one thing is reused for another with higher stakes. Microsoft calls this out specifically, suggesting it is useful to ask a group owner to review dynamic membership before the group is used in a different risk context. The third is guests who were onboarded for a CRO transition, a diligence process or a co-development programme and never removed. The fourth is exception lists, which are created because a policy was too blunt and then outlive the policy that justified them.

On the file-store side the equivalent decay is inheritance. Egnyte documents a failure mode that is an unusually clean example of how an access map becomes unexplainable: when a folder owner grants a permission on a parent folder, the permissions would actually change only as far as the topmost subfolders where the grantor loses owner permission. The grant appears to have been made. Part of the tree receives it. The rest silently does not. Three years and several administrators later, nobody can say why one subfolder behaves differently from its sibling, and the honest answer is that a partial-ownership grant in 2023 stopped halfway down the tree.

Egnyte own design guidance is the correct antidote and is stated as plainly as any vendor states it: start with minimal access at the top-level folders and assign permissions and ownership to users as needed at the lower-level folders. The default permission when adding a new user is Viewer, which is the right default and one that administrators override upward reflexively under time pressure. Inheritance can be broken deliberately per folder, and when it is broken the administrator chooses between removing inherited permissions and keeping existing ones, a choice that determines whether the break is a tightening or merely a freeze.

Detecting drift requires an artefact rather than an intention. On the Egnyte side that is the Folder Permissions Report. On the Microsoft side it is a combination of access reviews, the SharePoint data access governance reports, and the weekly assessments produced by data security posture management for AI, which by default runs over the top one hundred SharePoint sites by usage. None of these is complete on its own, and none of them reviews itself. A named owner and a calendar entry are as load-bearing as the tooling.

Privileged accumulation

Administrative rights granted to resolve an incident are almost never handed back.

Group repurposing

A group built for one purpose reused for a higher-risk one, with membership nobody re-reviewed.

Orphaned guests

CRO, diligence and co-development guests that outlive the engagement that created them.

Partial grants

Egnyte documents a grant that reaches only as far as the grantor own subfolders. Half the tree changes.

Related evidence and next steps

What this engagement is and is not

We design and build the model, and we work through the migration with the people who own the content. We are not an auditor, not a certifying body, and we do not issue certifications. Where an independent view of the finished design is valuable, it can be reviewed by a named security architect on our expert bank who did not build what they are reviewing.

IntuitionLabs was founded in 2023 and does not hold SOC 2 or ISO 27001. Our own security posture is documented on our security page and in our Trust Center rather than implied here. Everything on this page describes work we do for clients, not accreditations we hold.

Talk Through Your Estate

Design and build

Taxonomy, group structure, platform configuration and the migration plan, produced with your IT, quality and legal owners.

Independent review

Optional review by a named independent security architect on the expert bank, separate from the team that produced the design.

Not certification

We prepare for audits, certifications and security questionnaires. We never audit or certify, and we issue no certificates.

Platform enforcement, compared honestly

The enforcement layer is where a taxonomy stops being a spreadsheet. Each of the three platforms a mid-size life-sciences company is likely to run enforces classification differently, has a different set of controls people forget to configure, and has at least one documented behaviour that will surprise the team implementing it. This section covers what each one can actually do. The Microsoft-specific detail has its own page and is not repeated here.

Box

Labels are cheap; the Shield policy behind them is the control

Box has the simplest model of the three and it lands, deliberately, in almost the same place as the more elaborate ones. Up to twenty-five classification labels exist per enterprise, each file or folder carries exactly one at a time, and a label consists of a name of up to forty characters, a colour and an advisory message. Which collaborator roles may change a classification is a single global setting, and external users can never modify a classification.

The gap in most Box tenants is not the labels. Organisations create the labels and stop, which produces a coloured tag with an advisory message and no enforcement anywhere. The enforcement lives in Box Shield access policies, which are bound to a classification and control external collaboration restriction, shared link restriction, download and print restriction separately for managed and external users across web, mobile and desktop, integration restriction, FTP restriction, watermarking, and restrictions on Box Sign signature requests.

Of that list, two settings matter disproportionately and are the two most often left off. Download and print restriction for managed users is the control that actually prevents an authorised insider walking out with a file, and it is unpopular precisely because it constrains people who are allowed to see the content. Integration restriction, which blocks all integrations from downloading content, is the one that matters most in the AI era, because it is the setting that determines whether a connected application or agent can pull the bytes rather than merely see that the file exists.

There is a live licensing history that anyone auditing an older Box tenant needs to know, and it changes what a correct design looks like. Classification originally shipped in Box Governance. Box documents that since October 2020, security features including classification, the shared link restriction policy and content security in Box Governance are no longer available to new Box Governance customers, and that the ability to create, edit and delete content security policies and shared link policies was sunset on 6 February 2026, with those policies to be disabled for all customers in May 2026 while security classification continues to be available. The practical consequence is that on a current tenant, classification-driven enforcement means Box Shield access policies, and a design that assumes Governance policies is designing against a control that is being withdrawn.

Two more Box behaviours belong in a design document. Deleting a classification is permanent, and Box removes it from every file and folder to which it was applied, so a taxonomy revision is not a rename operation. And Box Shield offers automated classification that can identify personally identifiable information as content is uploaded, updated, moved or copied into specified folders, which is the mechanism to reach for when the class can be expressed as a pattern and the volume is beyond manual triage.

The label

Name up to 40 characters, a colour and an advisory message. One per item, twenty-five per enterprise.

The control

A Box Shield access policy bound to the label. Without one, the label enforces nothing.

The AI-relevant setting

Integration restriction: block all integrations from downloading content.

The licensing shift

Classification enforcement now means Shield. Governance content and shared-link policies are being withdrawn.

Related evidence and next steps

Egnyte

A folder model where classification informs and permissions enforce

Egnyte is architecturally different from the other two and the difference has to be respected rather than papered over. There is no label that carries enforcement. Content classification produces findings; folder permissions produce access. A design that treats an Egnyte classification as if it were a Purview label will produce a document that does not describe the running system.

Content classification lives in Egnyte Secure and Govern and is switched off by default, which Egnyte describes as a privacy decision. Turning it on begins with selecting a regional jurisdiction, which surfaces the built-in regulatory compliance policies relevant to that jurisdiction, after which custom policies can be added from built-in patterns, keywords or custom keywords. Egnyte gives unusually blunt advice about scope, warning that only the policies needed for regulatory compliance should be selected because enabling all the policies may generate too many results to be useful. That is a sentence worth quoting at anyone who wants to switch everything on for completeness.

Scanning is bounded in a way that affects planning. Egnyte documents that content classification will only analyse files in the paths that have been included for scanning in each source, so an unincluded path is simply invisible to the findings, not classified as clean. Results appear in a Sensitive Content view grouped by containing folder, with actions to allow, move or delete. Egnyte publishes file-type and file-size constraints for classification, and those should be confirmed against the current Content Classification FAQs during design rather than assumed from a summary.

The permission model is four folder access levels: Owner, Full, Editor and Viewer, with a fifth viewer-only level available for project folders. Permissions inherit downward by default, so a permission added on a shared parent folder also applies to its subfolders unless a more specific permission exists for that user on the subfolder. Inheritance can be broken per folder by administrators and by power-user folder owners, and breaking it does not cascade; when inheritance is turned off the administrator chooses between removing inherited permissions and keeping existing ones.

The specific behaviour to design around is the partial-ownership grant described earlier: a grant made by an owner reaches only as far as the topmost subfolders where the grantor loses owner permission. Combined with an ad-hoc culture of breaking inheritance to solve individual problems, this is how an Egnyte tree acquires an access map that cannot be explained from the outside. The remediation is a documented ownership model, deliberate rather than reactive inheritance breaks, and a scheduled read of the Folder Permissions Report. For companies running Egnyte as their regulated content store, we cover the platform in more depth on the dedicated Egnyte service page.

Off by default

The content classification toggle is switched off by default, which Egnyte states is for privacy reasons.

Narrow the jurisdiction

Egnyte warns that enabling all policies may generate too many results to be useful.

Scope is literal

Only files in included scan paths are analysed. An unincluded path is invisible, not clean.

Enforcement is the ACL

Four access levels, downward inheritance, deliberate breaks, and a report someone must read.

Related evidence and next steps

SharePoint and Purview

The most capable model, and the most conditional

Microsoft offers the deepest control set of the three and the largest number of ways to build something that appears configured and is not. This page covers the design decisions that shape the access model. The full platform treatment, including the licensing matrix and the documented Copilot protection gaps, has its own page and is not duplicated here.

The first design decision is what a label is. Microsoft states that because the label is stored in clear text in the metadata for files and emails, third-party apps and services can read it, and that the label stays with the content no matter where it is saved or stored. Each item that supports sensitivity labels can have a single label applied. That clear-text design is why labels interoperate well and also why a label by itself enforces nothing: the enforcement comes from encryption applied by the label, content markings, container settings, or a separate policy such as DLP or retention that uses the label as a condition. Marking-only labelling, which Microsoft supports explicitly as labelling content without using any protection settings, is a signal rather than a control.

The second is the container question, and it is the most commonly misunderstood behaviour in the product. A sensitivity label applied to a Microsoft 365 group, Teams team or SharePoint site sets privacy, external user access, external sharing, access from unmanaged devices, an Entra Conditional Access authentication context, discoverability of private teams and shared-channel controls. It does not label the items inside. Microsoft is unambiguous: items in these containers do not inherit the labels and therefore do not apply any item-level label settings such as content markings and encryption. A Confidential site full of unlabelled documents is a normal and expected state, not a misconfiguration.

The third is dependency. Several container settings are inert without a separate configuration elsewhere, and Microsoft warns that there is no check in the sensitivity label configuration that the dependencies are in place. If the dependent Conditional Access policy for SharePoint is not configured, the unmanaged-devices option specified on the label will have no effect. Separately, sensitivity labels for Office files in SharePoint and OneDrive must be enabled at the tenant before those services can process encrypted files at all; until they are, Microsoft states that co-authoring, eDiscovery, data loss prevention, search and other collaborative features will not work for those files. A label that looks configured in the portal and does nothing in production is almost always one of these two.

The fourth is licensing, which shapes the design more than anyone wants it to. Manual labelling reaches down to E3-class subscriptions, while the automation layers on top of it, including client-side and service-side auto-labeling, sit at E5-class. That statement is taken from the consolidated Microsoft Purview service description and is accurate as of 27 August 2026; the SKU names in that document have changed materially over the last two years and should be re-read rather than quoted from memory. The practical consequence for a design is that an E3 tenant can have a taxonomy and manual application of it, and cannot have the automation that makes a large migration tractable, which is a budget conversation to have before the workshop rather than after the pilot.

The fifth is what happens at the seams. Uploading a document carrying a higher-priority label to a site with a lower one is not blocked; it generates an audit event and a mismatch email to the uploader and to site owners, which most tenants either never route or suppress outright, losing the signal. Clearing a container setting does not retract access already granted: Microsoft notes that guests who accessed a site because a setting was previously selected can still access it after the setting is cleared. And deleting a label does not undo it, because the label is not automatically removed from content and any protection settings continue to be enforced on content that had that label applied.

A label is metadata

Clear text in the file. Readable by third parties, and enforcing nothing on its own.

Containers do not cascade

Items in a labelled site do not inherit the label and get no marking or encryption from it.

Dependencies are unvalidated

No check confirms the Conditional Access policy behind an unmanaged-device setting exists.

Deletion is not reversal

Deleting a label leaves the protection it applied still enforced on the content.

Related evidence and next steps

Common ground

What every platform leaves unconfigured

Across the three platforms the same categories of setting go unconfigured, for the same reasons. They are the settings that constrain people who are allowed to see the content, the settings that require a second system to be configured first, and the settings that generate a signal nobody agreed to own.

The first category is the insider-facing control: download and print restriction for managed users in Box, encryption usage rights that withhold extraction in Purview, and viewer-only permission levels in Egnyte. These are unpopular because they inconvenience people with a legitimate need, and they are the only controls in the set that address the case where the person accessing the content is supposed to have access to it. In an AI context they matter more than they used to, because an integration acting as a user inherits that user rights and the distinction between viewing and extracting becomes the whole control.

The second is the cross-system dependency: a Conditional Access policy that has to exist for a label setting to mean anything, a tenant-level enablement that has to be run before encrypted files can be processed, a jurisdiction that has to be selected before classification policies appear. Each of these is a place where the portal shows a configured state and the runtime behaviour is unchanged, and none of them raises an error. The only defence is a verification step in the implementation plan that tests behaviour rather than reading configuration.

The third is the unowned signal. Purview generates a mismatch email when a higher-labelled document lands in a lower-labelled site, and generates justification text in Activity Explorer when a user downgrades a label. Egnyte generates a Sensitive Content view and a Folder Permissions Report. Box generates Shield alerts. Data loss prevention alerts have asymmetric lifetimes that catch people out: Microsoft documents that DLP alerts are available in the Defender portal for six months but only for thirty days in the Purview DLP alerts dashboard. Every one of these produces evidence that a control is working or failing, and every one of them is worthless without a named person and a cadence.

The fourth is scope creep on the review itself. Insider risk management in Purview is explicitly a detection and investigation capability rather than a preventive one, built pseudonymised by default with role-based access controls and its own audit log, and it includes a patient data misuse template that requires the human resources connector and a healthcare-specific data connector to score behaviour inside electronic health record systems. That is a genuine capability for a commercial-stage company, and it is also a project rather than a checkbox. It belongs on a roadmap, not in the first phase.

Insider-facing controls

The ones that constrain authorised people. Unpopular, and the only defence against authorised extraction.

Silent dependencies

Settings that require another system first, and that fail without raising any error.

Unowned signals

Mismatch emails, justification logs, permission reports. Evidence nobody reads is not evidence.

Alert lifetimes

DLP alerts persist six months in Defender and thirty days in the Purview DLP dashboard.

Related evidence and next steps

Migration is the project. The taxonomy is the easy part.

Designing five tiers takes a workshop. Applying them to content you already have takes months, and it is where the engagement either succeeds or quietly stops. The correct opening move is a scoped pilot: one library, a small group of users, a classifier everyone understands, and a stated intention to widen only once behaviour is verified. Microsoft advises publishing new labels to a few test users first, waiting at least an hour, verifying behaviour, then waiting a day before widening.
  • Pick a pilot classifier for teachability rather than for need. Microsoft suggests using credit card numbers even where they are not a real concern, because the concept of these being sensitive items that need protection is easily understood by users, and swapping to production classifiers such as trainable classifiers for intellectual property once the mechanics are proven.
  • Set the label policy Learn More URL to a custom internal help page. Microsoft notes that if it is not set, users do not get the corresponding menu option in the Office sensitivity button at all, which removes the one in-product route a confused user has.
  • Test names and tooltips with the people who will apply them, not with the people who wrote them. Microsoft advises testing and tailoring label names and tooltips with the people who need to apply them, and recommends a working virtual team rather than a single owner for the deployment.
  • Stage the widening deliberately: a library, then a site, then a department, then the tenant. Each stage should have a stated behaviour to verify and a stated reason to stop.
  • Expect the pilot to change the taxonomy. If it does not, the pilot was not honest about how the labels were being applied in practice.
Regulatory documentation being reviewed during a content classification migration pilot

Simulation mode is not a silent dry run

Service-side auto-labeling supports simulation, and it is the right thing to run before enforcing anything. It is also widely misunderstood in ways that produce avoidable incidents on the first attempt. Microsoft documents that simulation still generates activity alerts, so an unscoped alert policy will email the compliance function during the run. Policy management is greyed out for roughly twenty-four hours after a policy is created, and a duplicated policy lands in simulation without notice.
  • Simulation shows the result of a single policy. Once several policies interact, Microsoft notes the enforced result can differ from any single policy simulation, so a clean simulation is not a prediction of the enforced outcome.
  • A simulation run can take twelve hours to complete, with a display cap of one hundred items per site and an export ceiling of fifty thousand records, so the review process has to be designed around sampling rather than exhaustive reading.
  • Exchange results are not reproducible between runs, because simulation evaluates mail sent and received during the run window rather than a fixed corpus.
  • There is a hard activation ceiling: simulation supports a maximum of four million matched files, and above that number Microsoft states you cannot turn the policy on. Rule tuning to reduce matched volume is therefore sometimes a prerequisite to deployment rather than an optimisation.
  • Removal and downgrade policies behave counter-intuitively on activation. Microsoft documents that the service only re-evaluates files whose state has recently changed, so untouched files are skipped even though simulation matched them, and on-demand classification is needed to force a re-scan.
Engineers reviewing auto-labeling simulation results before enforcing a classification policy

Over-encryption is the failure mode that ends programmes

Encryption is the only label setting that genuinely controls access, and it has by far the largest documented failure surface. Microsoft own current deployment model defers it deliberately: the milestone ladder sets a SharePoint default label without encryption first, extends that default to all files second, and only then adds encryption. The published phased alternative agrees, with no encryption on any label in phase one. Going hard early is how this work gets discredited internally.
  • Files labelled and encrypted before sensitivity labels were enabled for SharePoint and OneDrive are never recognised by those services; the documented remedy is to download and re-upload them to their original location.
  • Labels configured with content expiry or Double Key Encryption cannot be processed at all, and Microsoft states those documents are not returned in search results even if they are updated.
  • An encrypted Office file larger than twelve megabytes stops being processable by SharePoint once it is copied or moved to a different site.
  • Files containing Power Query data, data stored by custom add-ins, custom XML parts, a bibliography or a SharePoint Document ID cannot be processed when labelled and encrypted from Office desktop applications. For a biotech that means every Excel model with a data connection and every Word document with a reference manager.
  • Printing, downloading, exporting and creating a copy are unsupported for encrypted documents in Office for the web, and Office desktop and mobile applications do not co-author encrypted files by default.
  • An application using a service principal that downloads an encrypted file and re-uploads it with different encryption settings gets an upload failure, which breaks automated pipelines that touch labelled content.
  • Turning the feature off later does not return you to the starting state. Microsoft warns that the resulting mixed protection state can result in higher administrative overheads with help desk incidents to investigate the inconsistent behaviour, and that it might also be a problem for compliance requirements.
  • For clinical documentation the twenty-five year archive clock makes encryption a records question rather than only a usability one. Any design that encrypts trial master file content needs an explicit key-custody and super-user answer that survives everyone currently employed.
Laboratory data workflows affected by encryption applied to classified documents

The regulatory frame, cited precisely

Classification and access design in life sciences is not governed by one regulation. It sits at the intersection of health-information privacy, electronic records rules, computerised systems guidance, data-protection law and data-integrity expectations, each of which asks for something slightly different. Getting the citations exactly right matters, because the imprecise version of each of these drives the wrong design.

HIPAA

Required and Addressable are not the same word

The HIPAA Security Rule is the most frequently paraphrased and most frequently mis-paraphrased authority in this area. The paraphrase that causes the most damage is that HIPAA requires encryption. It does not, and the specification structure is worth reading rather than summarising.

Under 45 CFR 164.312(a)(1), the access control standard requires technical policies and procedures for electronic information systems that maintain electronic protected health information to allow access only to those persons or software programs that have been granted access rights. Beneath that standard, unique user identification is Required and emergency access procedure is Required, while automatic logoff is Addressable and encryption and decryption is Addressable. Under 164.312(e), transmission security, both integrity controls and encryption are Addressable. Addressable does not mean optional in the colloquial sense, but it does mean that the implementation is assessed against reasonableness and appropriateness rather than mandated outright.

What is unambiguously Required, and much more relevant to this work, sits in the administrative safeguards. 45 CFR 164.308(a)(1)(ii)(D) requires procedures to regularly review records of information system activity, such as audit logs, access reports and security incident tracking reports. That is a Required specification for periodic review, and it is the HIPAA counterpart to the PIC/S and WHO access-review language discussed earlier. Under 164.308(a)(3), workforce security, all three specifications are Addressable: authorisation and supervision, workforce clearance procedure and termination procedures. Under 164.308(a)(4), information access management, isolating health care clearinghouse functions is Required while access authorisation and access establishment and modification are Addressable.

Section 164.312(b), audit controls, is a standard with no implementation specifications: implement hardware, software or procedural mechanisms that record and examine activity in information systems that contain or use electronic protected health information. Section 164.312(c) covers integrity with an Addressable specification to corroborate that electronic protected health information has not been altered or destroyed in an unauthorised manner, and 164.312(d) covers person or entity authentication as a standard.

One practical note on sourcing. The official Electronic Code of Federal Regulations blocks automated retrieval, so the text quoted in our design documents is verified against the Cornell Legal Information Institute mirror and confirmed by a person against eCFR before it is relied on. We cite both. We also do not make any assertion on this page about proposed changes to the Security Rule, because a proposal is not a rule and design decisions taken on the strength of a proposal are difficult to defend either way.

Required under 164.312(a)

Unique user identification and an emergency access procedure.

Addressable under 164.312(a)

Automatic logoff, and encryption and decryption.

Required under 164.308

Information system activity review: regularly review audit logs, access reports and incident tracking.

The AI-relevant phrase

Access only for persons or software programs granted access rights. Already written for agents.

Related evidence and next steps

Part 11 and Annex 11

Access limitation and authority checks are two requirements

For a GxP-relevant record, the electronic records rules add a requirement that a folder permission model does not satisfy on its own. The distinction is in the text of 21 CFR 11.10 and it is easy to read past.

Paragraph (d) requires limiting system access to authorized individuals. Paragraph (g) separately requires the use of authority checks to ensure that only authorized individuals can use the system, electronically sign a record, access the operation or computer system input or output device, alter a record, or perform the operation at hand. Those are two obligations, and they are not satisfied by the same control. A folder access-control list answers (d): this person may open this content. Only role-based rights that gate specific operations answer (g): this person may sign, this person may alter, this person may not. That is precisely why an Egnyte or Box folder-permission model, however well designed, is not by itself a Part 11 answer for a record that lives under Part 11.

The surrounding paragraphs shape the rest of the design. Paragraph (a) requires validation of systems to ensure accuracy, reliability, consistent intended performance, and the ability to discern invalid or altered records. Paragraph (b) requires the ability to generate accurate and complete copies of records in both human readable and electronic form suitable for inspection, review and copying by the agency. Paragraph (c) requires protection of records to enable their accurate and ready retrieval throughout the records retention period, which is the paragraph that collides directly with a twenty-five year encryption key custody question. Paragraph (e) requires secure, computer-generated, time-stamped audit trails independently recording the date and time of operator entries and actions that create, modify or delete electronic records. Paragraph (k) adds revision and change control procedures maintaining an audit trail of time-sequenced development and modification of systems documentation.

EU GMP Annex 11 covers the same ground for computerised systems in a GMP context and adds the evidence requirement explicitly. Section 12.1 requires physical or logical controls to restrict access to the computerised system to authorised persons. Section 12.2 states that the extent of security controls depends on the criticality of the computerised system, which is the clause that justifies a tiered model rather than a uniform one. Section 12.3 requires that creation, change and cancellation of access authorisations be recorded. Section 12.4 requires management systems to record the identity of operators entering, changing, confirming or deleting data including date and time. Section 9 requires risk-based audit trails, documented reasons for change or deletion of GMP-relevant data, and audit trails that are available, convertible to a generally intelligible form and regularly reviewed.

Where this work intersects with validated systems, it is validation work and should be treated as such rather than as an IT change. Our computer system validation and Part 11 software practices cover that boundary, and the sequencing question, whether a classification change to a validated environment requires revalidation, is a question for the client quality system rather than for a security design document.

11.10(d)

Limiting system access to authorized individuals. A folder ACL can satisfy this.

11.10(g)

Authority checks over specific operations. A folder ACL cannot satisfy this.

11.10(c)

Ready retrieval throughout the retention period, which constrains any encryption decision.

Annex 11 section 12.2

Security control extent depends on system criticality. The clause that justifies tiering.

Related evidence and next steps

GDPR

Accountability is the classification argument in one clause

For any company with European subjects, staff or trial sites, the General Data Protection Regulation supplies both a reason to classify and a requirement to keep testing whether the classification is working. It is a more demanding regime than HIPAA in one specific respect that matters here.

Article 5(1)(c) requires personal data to be adequate, relevant and limited to what is necessary in relation to the purposes for which they are processed, which is data minimisation and which is difficult to demonstrate over an uncategorised estate. Article 5(1)(e) adds storage limitation and Article 5(1)(f) adds integrity and confidentiality, requiring processing in a manner that ensures appropriate security of the personal data including protection against unauthorised or unlawful processing. Article 5(2) then supplies the sentence that is effectively the business case for this entire workstream: the controller shall be responsible for, and be able to demonstrate compliance with, paragraph 1. You cannot demonstrate compliance over content you have never categorised.

Article 32(1) requires appropriate technical and organisational measures to ensure a level of security appropriate to the risk, expressly including the pseudonymisation and encryption of personal data, and expressly including a process for regularly testing, assessing and evaluating the effectiveness of technical and organisational measures. That last clause is the one that distinguishes the GDPR from HIPAA in practice: it builds periodic control-effectiveness review into the regulation itself rather than leaving it to a separate administrative standard. Article 32(4) is the access-control clause proper, requiring that any person acting under the authority of the controller or processor who has access to personal data does not process them except on instructions from the controller.

For companies deploying Microsoft Copilot in Europe there is a live data-residency detail worth knowing at design time rather than at contract review. Microsoft states compliance with the GDPR and the EU Data Boundary for Copilot, with the documented caveat that models provided by Anthropic as a subprocessor are currently excluded from the EU Data Boundary. That is not a reason to avoid any particular vendor. It is a fact that belongs in the record of processing activities and in the conversation with the data protection officer before deployment rather than after.

The interaction with de-identified clinical data is the point most often missed. A dataset processed through HIPAA Safe Harbor is de-identified for the purposes of the US Privacy Rule. The GDPR does not recognise that mechanism, and pseudonymised data remains personal data under the GDPR. An EU-facing programme therefore cannot treat a HIPAA-de-identified set as out of scope, and a classification model that assigns such a dataset to a low tier on the strength of the HIPAA determination is making a European mistake for an American reason.

Article 5(2)

Demonstrate compliance. Not possible over an estate that has never been categorised.

Article 32(1)

Includes a requirement for a process regularly testing the effectiveness of the measures.

Article 32(4)

People with access must not process personal data except on the controller instructions.

The Safe Harbor trap

HIPAA de-identification does not remove GDPR scope. Pseudonymised data is still personal data.

Related evidence and next steps

Data integrity

ALCOA+, and the guidance documents worth naming

Three data-integrity documents carry most of the weight in GxP environments, and quoting them by identity rather than by reputation is what separates a design document an inspector accepts from one they question. All three are freely available, and all three say things about access that a security design should be built to satisfy.

The MHRA GXP Data Integrity Guidance and Definitions, revision 1 of March 2018, defines data integrity as the degree to which data are complete, consistent, accurate, trustworthy and reliable, and that these characteristics of the data are maintained throughout the data life cycle. Its data governance definition covers the arrangements ensuring that data, irrespective of format, are recorded, processed, retained and used to ensure the record throughout the data lifecycle, and it places accountability explicitly: senior management should be accountable for the implementation of systems and procedures to minimise the potential risk to data integrity. Its glossary is the standard reference for ALCOA, meaning attributable, legible, contemporaneous, original and accurate, with the plus adding complete, consistent, enduring and available.

PIC/S PI 041-1 of 1 July 2021 supplies the access specifics already covered: least privilege by role, prohibition of shared passwords even for reasons of financial savings, segregation of administrator rights from normal users, administrator independence from the users performing the task with no interest in the outcome of the data generated, and a generated list of users with actual access used during periodic user reviews. Read as a specification for an identity design it is remarkably close to a modern least-privilege architecture, written for pharmaceutical manufacturing.

WHO Technical Report Series No. 1033, Annex 4, the 2021 guideline on data integrity that replaced TRS 996 Annex 5, adds the governance framing. It requires a written policy on data integrity, requires data governance to ensure the application of ALCOA+ principles, requires governance to address roles, responsibilities, accountability and segregation of duties throughout the life cycle, requires a documented system defining access and privileges, requires that inactivated users be retained with a maintained list of active and inactivated users across the system life cycle, and prohibits shared logins or generic user access for systems generating, amending or storing GxP data.

The reason to hold all three rather than picking one is jurisdictional. A company with a UK inspection history, an EU manufacturing site and a WHO-prequalified product is answerable to all of them, and their access requirements are compatible but not identical. Designing to the union rather than to whichever one someone quoted in a meeting is both cheaper and more defensible.

MHRA revision 1, March 2018

The data integrity and data governance definitions, and the ALCOA+ glossary.

PIC/S PI 041-1, 1 July 2021

The most operationally specific access-control text in GxP guidance.

WHO TRS 1033 Annex 4, 2021

Written policy, documented privileges, retained inactivated users, no generic logins.

Design to the union

The three are compatible but not identical. Build to all of them, not to whichever was quoted.

Related evidence and next steps

Zero trust

A maturity ladder to place yourself on

Two public documents are worth using as external anchors when a board or an auditor asks how mature the model is. Neither is a life-sciences document, and that is exactly why they are useful: they are neutral, dated, freely available, and specific enough to place an organisation on a scale rather than assert a grade.

NIST Special Publication 800-207, Zero Trust Architecture, was finalised on 11 August 2020 by Scott Rose and Oliver Borchert of NIST, Stu Mitchell of Stu2Labs and Sean Connelly of the Department of Homeland Security. Its core proposition is that zero trust assumes there is no implicit trust granted to assets or user accounts based solely on their physical or network location, or based on asset ownership. That proposition is directly relevant to a content estate, because the traditional model of trust in a life-sciences file share is exactly locational: it is on the internal share, therefore employees may read it. There is no revision two of SP 800-207 as of this writing, and citations to one should be treated with suspicion.

The CISA Zero Trust Maturity Model, version 2.0, dated April 2023 on its cover, is the more practical of the two for this purpose. It is structured as five pillars, identity, devices, networks, applications and workloads, and data, with three cross-cutting capabilities, visibility and analytics, automation and orchestration, and governance, and four maturity stages: traditional, initial, advanced and optimal. A note on the document itself: the cover states April 2023 and version 2.0, while the internal revision table records version 2.0 as March 2022. We cite the cover date and flag the discrepancy rather than quietly picking one.

Version 2.0 added data categorization as a function under the data pillar, and its ladder is the single best external scale for a classification programme. Traditional is limited and ad hoc data categorization capabilities. Initial is a data categorization strategy with defined labels and manual enforcement mechanisms. Advanced automates some data categorization and labeling processes in a consistent, tiered, targeted manner with simple, structured formats and regular review. Optimal automates data categorization and labeling enterprise-wide with robust techniques, granular structured formats, and mechanisms to address all data types. Most of the companies we work with are at traditional and believe they are at initial, because the policy exists and the labels do not.

The data access ladder in the same document is the identity-side twin, running from static access controls at traditional to dynamic just-in-time and just-enough data access controls enterprise-wide with continuous review of permissions at optimal. Used together the two ladders give a board a defensible statement of position and a defensible statement of target, without anyone having to claim a certification or invent a score. The model aligns to OMB M-22-09 and Executive Order 14028, and its data pillar footnotes the NIST NCCoE data classification project as further reference.

NIST SP 800-207

Final 11 August 2020. No implicit trust from network location or asset ownership.

CISA ZTMM v2.0

Five pillars, three cross-cutting capabilities, four stages. Cover dated April 2023.

Data categorization ladder

Traditional, initial, advanced, optimal. The best external scale for a classification programme.

Where most companies sit

Traditional, while believing they are initial, because the policy exists and the labels do not.

A maturity ladder lets you state a position and a target without claiming a certification you do not hold.

Related evidence and next steps

Ongoing

This does not finish, and pretending otherwise is the mistake

Every proposal for this work is asked when it ends. The honest answer is that the build ends and the model does not. Content grows, people join and leave, projects close, contract organisations rotate, groups get reused, and a permission granted for a reason nobody recorded outlives the reason. Microsoft says it about its own product area: deploying an information protection solution is not a linear deployment but iterative, and often circular.

The recurring work has four components and they are not evenly distributed through the year. Access review is the largest and is calendar-driven, ideally quarterly for groups gating Confidential and above and annually for the rest, producing the list of users with actual access that PIC/S expects. New-content triage handles material entering the estate that the auto-labeling rules cannot classify, which is always more than the design predicted, and it is the component that decays first when nobody owns it. Drift detection reads the artefacts the platforms already generate: mismatch events, downgrade justifications, permission reports, posture assessments. Taxonomy maintenance is the smallest and the least frequent, and it is the one that should require a change-control conversation rather than an administrator with a portal open.

Two events reset a meaningful part of the model and should be planned for rather than absorbed. The first is corporate change: mergers, acquisitions, divestitures and CRO transitions. Microsoft documents that after cross-tenant migration the source-tenant label often is not recognised in the target tenant, which for a biotech in the middle of an acquisition means a labelled estate arriving with its labels functionally inert. The second is platform change, and it happens more often than anyone budgets for. The Box classification entitlement moved from Governance to Shield. Restricted SharePoint Search is retiring, with Microsoft blocking new enablement from 31 July 2026, and its replacement, Restricted Content Discovery, requires both a Copilot licence and SharePoint Advanced Management and is explicitly not a security boundary either.

Both of those SharePoint discovery features deserve a sentence of caution because they are frequently mistaken for access controls. Microsoft states that Restricted SharePoint Search is not a security boundary and does not change any permissions on SharePoint sites, that its allow list is capped at one hundred sites, and that a site a user recently accessed or that was shared with them appears in their results even if it is not on the allow list. Of Restricted Content Discovery, Microsoft states that it does not change existing permissions and does not remove content from the Microsoft 365 search index, and that for sites with more than five hundred thousand items an update could take more than a week to fully process. They are discovery brakes. They are not the access model, and a design that leans on either as a control is misreading the documentation.

We structure this as a standing engagement rather than as a project with a closing report, because the alternative is a model that is accurate on the day it is delivered and progressively less accurate every week after. The cadence is agreed with the client, the review artefacts are the ones the platforms already produce, and the output is a backlog that someone is accountable for closing. That is a less satisfying deliverable than a finished architecture diagram, and it is the one that keeps working.

Access review

Calendar-driven, quarterly for the higher tiers, producing the user list PIC/S expects.

New-content triage

Everything the rules cannot classify. Always more than predicted, and the first thing to decay.

Drift detection

Reading the mismatch events, downgrade justifications and permission reports already being generated.

Reset events

M&A and CRO transitions, and platform changes such as the Box and SharePoint shifts of 2026.

Related evidence and next steps

What the engagement actually produces

Six artefacts, each of which is a thing that exists in a system rather than a thing that exists in a document. The written deliverables describe them; they are not the deliverable.

Tier model and content matrix

Four or five named tiers, and a matrix assigning every content class a default tier, its documented escalations and a named owner who can adjudicate a disputed case.

See the practice

Group architecture

The security groups each tier resolves to, how each membership is determined, which groups must not be dynamic, and the attribute write-permission audit behind the ones that are.

Start with the review

Platform configuration

Box Shield policies, Egnyte permission and classification settings, or SharePoint container labels and Purview policies, built against groups rather than individual names.

Purview specifics

Migration plan and pilot

A scoped pilot, a simulation plan with alerting scoped in advance, auto-labeling rules where a rule can express the class, and a manual triage queue where it cannot.

The information layer

Review cadence and evidence

The access-review schedule, the reviewer assignments, and the record of creation, change and cancellation of access authorisations that Annex 11 section 12.3 asks for.

Validation boundary

Questionnaire-ready evidence pack

The same design, expressed in the form a partner security questionnaire or a customer audit asks for, so the work is answered once rather than reassembled each time.

Questionnaire readiness

Where this sits in the security practice

This page is one service in a cluster. The review finds the exposure, this engagement builds the model that closes it, and the remaining pages cover the vendors you connect, the questionnaires you answer and the testing that verifies what you built.

Before
AI Access Exposure Review
A read-only assessment of what a connected assistant can currently reach across your content stores. It produces the prioritised backlog that this engagement works through.
Review first
Alongside
AI Vendor Security Assessment
Evaluating the AI vendors and connectors you are about to grant access to, including data handling, tenancy, subprocessors and administrative controls.
Assess the vendor
After
Penetration Testing and Secure Code Review
Independent testing of the systems around the model, performed only under a signed authorization letter scoping exactly what may be touched. Testing is insured.
See testing scope

Questions about classification and access design

A policy states which categories exist and who may see each one. It does not apply a label to a file, it does not create the security group that the label maps to, and it does not configure the platform control that reads either one. Microsoft states the position plainly in its own service assurance documentation: classification levels by themselves are simply labels or tags that indicate the value or sensitivity of the content. The enforcement is separate work, in a different system, done by a different person. Most of the companies we meet have completed the first part and none of the second, which is why an AI assistant connected to the same content behaves as if the policy did not exist.
Four or five top-level tiers, with sublabels only where a real sharing distinction exists. Three independent Microsoft sources converge on this. Its service assurance guidance says a data classification framework is typically comprised of three to five classification levels and recommends no more than five top-level parent labels each with five sublabels. Its product documentation reports that real-world deployments show effectiveness is noticeably reduced when users have more than five main labels or more than five sublabels per main label. Box independently caps its own model at twenty-five classification labels and recommends three in general practice. A fourteen-tier taxonomy is not more precise; it is a taxonomy people guess at.
Only in specific configurations, and you should know which ones before you rely on any of them. Microsoft documents that Copilot only surfaces organizational data to which individual users have at least view permissions, so permissions are the primary control and the label is secondary. Where a label applies encryption, Microsoft requires the user to hold the EXTRACT usage right as well as VIEW for an AI app to return the data, and states that without EXTRACT, Copilot will not summarize the content but can reference it with a link. Label-based DLP for the Copilot location is a separate, E5-class capability, it still shows blocked items in citations, and Microsoft documents that it cannot scan the contents of files a user uploads directly into a prompt.
A label is metadata describing the content. A group is the object that actually decides who can open it. Every enforceable control on every platform resolves to an identity: a Box collaborator role, an Egnyte folder permission, an Entra security group behind a SharePoint site or a Purview encryption setting. The work is designing that group structure so it can be reasoned about, mapping each classification tier to specific groups, removing the ad-hoc individual grants that accumulated over years, and establishing a review cadence. PIC/S PI 041-1 requires that systems support different user access roles and that role assignment follow the least-privilege rule, and that a list of users with actual access be generated and used during periodic user reviews.
They are different models rather than different quality levels, and the right answer is usually the one you already run. Box is a label-plus-policy model: up to twenty-five classification labels, one per item, with Box Shield access policies bound to the label. Egnyte is a folder-ACL model where classification is advisory and permissions enforce, and its classification feature is switched off by default. SharePoint with Purview is the most capable and the most conditional, because much of what people assume is included sits behind E5-class licensing. We work with the platform in place and design the model that platform can actually enforce, rather than recommending a migration you did not ask for.
Encryption is the only label setting that genuinely controls access, and it is also the setting with the largest documented failure surface. Microsoft lists a long set of behaviours for SharePoint and OneDrive: labelled and encrypted Office files larger than twelve megabytes stop being processable when moved to a different site; files containing Power Query data, custom add-in data, custom XML parts or a SharePoint Document ID cannot be processed when labelled and encrypted from desktop apps; printing, downloading, exporting and copying are unsupported for encrypted documents in Office for the web; and desktop and mobile apps do not co-author encrypted files by default. Microsoft own current deployment model defers encryption to the third milestone for exactly this reason.
Partly, and not in the way most people expect. Microsoft documents that auto-labeling applies your selected sensitive information types only to content created or modified after those information types were created or modified, so a static back catalogue is not covered without on-demand classification. Service-side Exchange labelling covers mail in transit and explicitly does not include emails at rest in mailboxes. There are hard ceilings: a maximum of one hundred thousand automatically labelled files per tenant per day, a maximum of one hundred auto-labeling policies, and a simulation ceiling of four million matched files above which you cannot turn the policy on. Auto-labeling is a scaling tool for the classes you can express as a rule, not a substitute for the taxonomy decision.
It is the right step, but it is not the silent dry run people assume. Microsoft documents that simulation still generates activity alerts, so an unscoped alert policy will mail the compliance team during the run. Policy management is greyed out for roughly twenty-four hours after a policy is created. A duplicated policy lands in simulation by default without notifying you. Simulation shows the result of a single policy, so the enforced result can differ from any single simulation once several policies interact. A run can take twelve hours. We plan the alert scoping and the review window before the first simulation rather than discovering these behaviours during it.
Directly, and with more precision than most write-ups apply. 21 CFR 11.10(d) requires limiting system access to authorized individuals. 11.10(g) separately requires authority checks to ensure that only authorized individuals can use the system, electronically sign a record, access the operation or computer system input or output device, alter a record, or perform the operation at hand. Those are two different requirements: a folder ACL satisfies the first and not the second. EU GMP Annex 11 section 12 adds that logical controls should restrict access to authorised persons, that the extent of security controls depends on the criticality of the system, and that creation, change and cancellation of access authorisations should be recorded, which is an access-review evidence requirement stated in a single line.
No, and getting this right matters because the wrong version of it drives bad design. Under 45 CFR 164.312(a)(1), encryption and decryption is an Addressable implementation specification, as is automatic logoff. What is Required under that same standard is unique user identification and an emergency access procedure. Separately, 45 CFR 164.308(a)(1)(ii)(D) makes information system activity review Required: procedures to regularly review records of information system activity such as audit logs, access reports and security incident tracking reports. The standard itself requires access only for those persons or software programs that have been granted access rights, and the phrase software programs already reaches an AI agent without any amendment.
The taxonomy and identity design is a matter of weeks; the migration of existing content is a matter of months and depends entirely on how much content you have and how much of it has a clear owner. It is not finished, and we would rather say that up front. Content grows, people join and leave, projects end, CROs rotate, groups get repurposed, and a permission granted for one project outlives it. Microsoft describes deploying an information protection solution as not a linear deployment but iterative, and often circular. We structure the work as an initial build followed by a standing engagement covering review cadence, drift detection and the backlog the reviews produce.
No. We prepare clients for audits, certifications and security questionnaires, and we are not an auditor and not a certifying body. We do not issue certifications of any kind. Design and build work is performed by IntuitionLabs; where an independent review of the resulting design is valuable, the engagement can be reviewed by a named security architect on our expert bank who did not build what they are reviewing. IntuitionLabs itself does not hold SOC 2 or ISO 27001, and our own security posture is documented separately rather than implied by this page.
Start With What You Actually Have

Start With What You Actually Have

Bring the classification policy you already wrote, the content stores you already run, and the AI deployment you are planning or have already made. We will tell you which parts of the policy can be enforced today, which need a group structure built first, and what the migration realistically involves.

Book a Meeting

© 2026 IntuitionLabs. All rights reserved.