Claude

IntuitionLabs is now a member of the Claude Partner Network – AI training and upskilling with Claude for pharma and biotech. Book a call.

IntuitionLabs
Quality and IT reviewers examining supplier qualification evidence for an AI vendor in a life sciences setting

AI Vendor Security Assessment: GxP Supplier Qualification for AI Vendors

A short, fixed-price review of one AI vendor, run against the standards an inspector will use — not only the questionnaire the vendor filled in about itself.

What you are actually deciding

A vendor assessment is not a security opinion. It is a record of four decisions that your quality system, your data protection obligations and your contract will each depend on afterwards.

Decision 01
Whether this use is in scope at all
GxP or not, patient data or not, regulatory submission or not. The answer sets the depth of everything that follows and is the cheapest thing to get wrong.
Validation context
Decision 02
What evidence you must hold yourself
Not what the vendor holds. Draft Annex 11 requires documentation you can access and explain from your own facility, which is a different artefact from a report under NDA.
See the service line
Decision 03
What the contract has to say
Audit conditions, inspection support, sub-processor approval, an exit strategy, and the right to test a new version before it is released to you.
Part 11 context
Decision 04
What you own when it ends
Your records, your audit trails and your metadata, in a form you can read without the vendor. This is where most AI procurement is quietly weakest.
Information layer

A security review and a supplier qualification are not the same instrument

Both are legitimate, both are necessary, and they answer different questions. A generic security review asks whether the vendor runs a competent security programme. A life-science supplier qualification asks whether you, the regulated user, can still discharge a legal responsibility you are not permitted to delegate. Answering one does not answer the other, and this is the single most consequential misunderstanding in AI procurement inside pharma and biotech today.

The starting point

A questionnaire the vendor filled in is not an assessment

The standard procurement artefact is a spreadsheet the vendor completed about itself, returned to a buyer who has no practical way to test any individual answer. That is a useful structured disclosure. It is not diligence, and the market already knows it: research cited by Vanta reports that only 34% of third-party risk management professionals believe questionnaire responses are accurate, and that a single questionnaire takes somewhere between five and fifteen hours to complete.

Two things follow from that number, and they point in the same direction. The first is that adding more questions does not add more assurance. If the instrument is distrusted, a longer instrument is distrusted at greater cost to both sides. The second is that verifiable evidence beats volume of answers by a wide margin. A vendor that names its model provider and cites the specific contractual clause governing training on customer data has told you more in one line than a two-hundred-question spreadsheet of yes and no.

The failure mode is rarely dishonesty. It is that the person completing the questionnaire is answering in good faith about a general product while you are buying a specific deployment, and the general answer and the specific answer diverge in exactly the places that matter. Data residency is the clearest example: a company can be headquartered in the EU, hosted in the EU, and still process your prompts in whatever geography its chosen deployment type routes to. The questionnaire has no field for that distinction, so it is answered "yes" and nobody has lied.

An assessment is therefore not a longer questionnaire. It is a short list of high-consequence claims, each of which is checked against a primary source: a contract clause, a configuration state the vendor can demonstrate, a scope statement on a certificate, a report section, a published sub-processor list. Everything else can be taken on disclosure. This is what makes the work finishable in days rather than becoming an open-ended audit that nobody completes before the purchase order is signed.

It also means that the output has to record what could not be verified. An assessment that presents only what was confirmed is a sales document. The useful deliverable states plainly which claims were checked against evidence, which were accepted on the vendor's word, and which the vendor declined or was unable to answer. The third category is usually the most informative part of the file.

Check

Claims where a wrong answer changes the buy decision or creates a regulatory exposure. Verify these against primary sources.

Accept

Claims that are low consequence or independently unverifiable. Record them as disclosure, not as verified fact.

Escalate

Refusals, non-answers and contradictions between documents. These are findings in their own right.

Re-ask

Anything version-dependent. A model, a deployment type or a sub-processor list is a moving target, not a settled fact.

The point of an assessment is not to collect more answers. It is to decide which few answers you are going to insist on being able to prove.

Related evidence and next steps

In force today

EU GMP Annex 11 and Chapter 7 already require a supplier assessment

This is not a new or speculative obligation created by AI. The current Annex 11 to the EU GMP guide came into operation on 30 June 2011 and its section 3 covers suppliers and service providers directly. Chapter 7 on outsourced activities came into operation on 31 January 2013 and was revised specifically to align with ICH Q10, so that it covers outsourced GMP activities beyond contract manufacture and analysis. A software service is squarely within its scope.

Annex 11 section 3.1 requires that where third parties provide, install, configure, integrate, validate, maintain, modify or retain a computerised system, or process data, "formal agreements must exist" with "clear statements of the responsibilities of the third party" — and it adds that internal IT departments "should be considered analogous". Section 3.2 states that "the competence and reliability of a supplier are key factors when selecting a product or service provider" and that "the need for an audit should be based on a risk assessment". That sentence is the legal basis for the supplier audit request that lands, often unexpectedly, on a SaaS vendor's desk.

Section 3.4 is the clause that catches AI vendors out, and it deserves to be read slowly: "Quality system and audit information relating to suppliers or developers of software and implemented systems should be made available to inspectors on request." A regulator can ask to see the vendor's quality system evidence through the customer. A confidential report the customer has never been permitted to retain does not automatically satisfy that. Section 4.5 adds that the regulated user "should take all reasonable steps, to ensure that the system has been developed in accordance with an appropriate quality management system" and that "the supplier should be assessed appropriately".

Chapter 7 sharpens the test. Section 7.5 states that "prior to outsourcing activities, the Contract Giver is responsible for assessing the legality, suitability and the competence of the Contract Acceptor". Three separate tests, and a security attestation speaks to at most one of them. Section 7.4 places ultimate responsibility with the Contract Giver. Section 7.11 requires that the Contract Acceptor "should not subcontract to a third party any of the work entrusted to him under the Contract without the Contract Giver's prior evaluation and approval" — which, for an AI product whose model provider is a subcontractor, is a hard requirement rather than a courtesy notification. Section 7.13 states that outsourced activities "may be subject to inspection by the competent authorities": a software vendor selling into GMP can be inspected.

ICH Q10 is the parent text Chapter 7 was aligned to. Its section 2.7 states that "the pharmaceutical company is ultimately responsible to ensure processes are in place to assure the control of outsourced activities", and requires assessment of "the suitability and competence of the other party" before outsourcing, using audits, material evaluations and qualification. ICH Q9(R1), adopted at Step 4 on 18 January 2023, supplies the proportionality rule that keeps this sane: "the level of effort, formality and documentation of the quality risk management process should be commensurate with the level of risk", and the evaluation of risk "should be based on scientific knowledge and ultimately link to the protection of the patient".

Read together, these texts do something that a security framework never does. They locate accountability in a specific legal person — the regulated user — and then require that person to produce evidence about somebody else's organisation. That inversion is why a vendor assessment in life sciences is an obligation you own rather than a service you buy, and why the deliverable has to be something you can hold and explain, not something the vendor holds on your behalf.

Annex 11 §3.1

Formal agreements with clear statements of third-party responsibilities. Internal IT is treated the same way.

Annex 11 §3.4

Supplier quality and audit information available to inspectors on request, through you.

Chapter 7 §7.5

Legality, suitability and competence — three tests, assessed before you outsource.

Chapter 7 §7.11

No subcontracting without your prior evaluation and approval. The model provider is a subcontractor.

Related evidence and next steps

Draft, not final

Draft Annex 11 section 7 states the difference in plain language

On 7 July 2025 the European Commission opened a joint stakeholder consultation on three EudraLex Volume 4 documents: a revised Chapter 4 on documentation, a revised Annex 11 on computerised systems, and an entirely new Annex 22 on artificial intelligence. The consultation closed on 7 October 2025. As of 27 August 2026 no final text has been published, so everything in this section is draft language and must be read as such. It is nonetheless the clearest available statement of what regulators expect, and it is already shaping supplier questionnaires.

The draft Annex 11 grows to seventeen sections, and section 7 is retitled "Supplier and Service Management". Section 7.1 addresses reliance on a vendor's own qualification work directly: "When a regulated user is relying on a vendor's qualification of a system used in GMP activities … this does not change the requirements put forth in this document. The regulated user remains fully responsible for these activities based on the risk they constitute on product quality, patient safety and data integrity." There is no reading of that sentence in which buying a certified product transfers the obligation.

Section 7.2 sets the depth of work: the regulated user "should, according to risk and system criticality, conduct an audit or a thorough assessment to determine the adequacy of the vendor or service provider's implemented procedures, the documentation associated with the deliverables, and the potential to leverage these rather than repeating the activities." The phrase "rather than repeating the activities" is important and is often missed. The draft is not asking you to re-validate the vendor's software. It is asking you to determine, on evidence, how much of the vendor's work you are entitled to lean on.

Section 7.4 is the requirement that no attestation report satisfies on its own: documentation "is accessible and can be explained from their facility". Not held by the vendor. Not available on request under an NDA that forbids retention. Accessible, in your building, and explainable by your people to an inspector who is standing in front of them. If your only artefact is a confidential report your quality lead has not read, you do not meet this description, however good the report is.

Section 7.5 then sets a nine-point minimum for the contract itself: the activities and documentation to be provided; the company procedures and regulatory requirements to be met; regular, ad hoc and incident reporting with service levels, key performance indicators, answer and resolution times; "conditions for supplier audits"; "support during regulatory inspections, if so requested"; issue resolution; "requirements and processes for communication of quality and security related issues"; "an exit strategy by which the regulated user may retain control of system data"; and "the process for release of new system versions and on the regulated user's possibility to test these prior to release".

The last two items are the ones that collide hardest with how modern software is actually built and sold. An exit strategy that returns control of system data is not a data deletion clause; it is a data return clause with a format obligation attached. And a customer right to test new versions before release is structurally incompatible with continuous deployment and with silent model version changes. Neither is impossible — several vendors already offer pinned versions, release channels and advance notice windows — but they have to be negotiated before signature, because they are architecture decisions as much as contract terms.

Draft Annex 11 section 7.4 is the sentence to bring to a vendor call: the documentation must be "accessible and can be explained from their facility".

Related evidence and next steps

The consequence

Why an attestation report cannot close a GxP supplier assessment

Put the two documents side by side and the mismatch is not a matter of rigour. A SOC 2 Type 2 report can be an excellent, demanding piece of work and still leave every GxP question open, because it was built to answer a different question for a different reader.

Start with what a SOC 2 report is. It is an attestation examination performed under AICPA attestation standards by a licensed CPA firm, resulting in a report containing a practitioner's opinion against the Trust Services Criteria — formally TSP section 100, the 2017 criteria with revised points of focus issued in 2022. The five trust services categories are Security, Availability, Processing Integrity, Confidentiality and Privacy. Only Security, the common criteria, is mandatory; the other four are elective and appear only if the service organisation chose them. So the first question about any SOC 2 report is which categories are in scope, and a report covering Security alone says nothing about availability or privacy.

Now map that onto draft Annex 11 section 7.4. The report is normally provided under a non-disclosure agreement, frequently on view-only terms, and its contents are the vendor's confidential information. Your quality unit cannot usually retain it, cannot excerpt it into a supplier file, and often cannot explain its control set from your facility. Chapter 7 section 7.11 asks something the report does not attempt: prior evaluation and approval of the subcontractors, which for an AI product means the model provider by name. Section 7.13 asks whether the vendor understands it may be inspected by a competent authority, which is a contractual commitment, not an audit finding.

The gap is not filled by choosing a stricter framework either. ISO/IEC 27001 certifies a management system within a declared scope. ISO/IEC 42001 certifies an AI management system. HITRUST verifies a prescriptive control set. All three are third-party conformity assessments of the vendor's own programme. None of them produces the artefact that your inspector will ask you for, which is your assessment, in your file, with your risk rationale, referencing evidence you can access.

This is why the practical recommendation for a small vendor selling into GMP is often counter-intuitive. If your buyer is a GMP manufacturer, spending the certification budget on a validation support package, an audit-ready quality manual and a GxP-competent contract answers the question that is actually being asked, whereas an ISO certificate scoped to corporate IT does not. That is the same analysis run from the seller's side, and it is the subject of our security questionnaire readiness work, linked below.

None of this argues for ignoring security attestations. They are the fastest available signal that a vendor has an operating security programme, and their absence in a company of any size is itself informative. The argument is narrower and it is about sequencing: read the report properly, then run the supplier assessment, and do not let the first stand in for the second.

Different reader

A SOC 2 report is written for a service organisation's customers generally. A supplier assessment is written for your inspector.

Different question

Is the security programme competent, versus can the regulated user still discharge its responsibility.

Different custody

A report held under NDA is not documentation accessible and explainable from your facility.

Different scope lever

Trust services categories and a system description, versus GMP criticality and patient risk.

Related evidence and next steps

Two instruments, two different questions

The rows are not degrees of strictness. They are different objects of assessment, drawn from draft EU GMP Annex 11 section 7, EU GMP Chapter 7, ICH Q10 section 2.7 and ICH Q9(R1).

QuestionGeneric security reviewGxP supplier assessment
What is being protectedConfidentiality, integrity and availability of dataPatient safety, product quality and data integrity
Who is accountableShared, defined by contractThe regulated user, always and irreducibly (draft Annex 11 §7.1)
Test applied to the vendorIs the security programme competentLegality, suitability and competence (Chapter 7 §7.5)
What drives the depthData sensitivity and contract valueRisk to patient, product and data; GMP criticality (ICH Q9(R1))
Where the evidence livesVendor trust centre, usually under NDAAccessible and explainable from your facility (draft Annex 11 §7.4)
Sub-processorsNotify and allow objection (GDPR Article 28)Prior evaluation and approval, assessment passed through (Chapter 7 §7.11)
Release managementA documented change management policyYour right to test new versions before release (draft Annex 11 §7.5)
ExitA data deletion clauseAn exit strategy by which you retain control of system data (draft Annex 11 §7.5)
Regulator involvementNoneThe vendor may be inspected (Chapter 7 §7.13; Annex 11 §3.4)
Electronic recordsNot addressed21 CFR Part 11 and Annex 11: audit trails, signature meaning, readable export
AI specificallyData-handling questionsDraft Annex 22: static, deterministic models only in critical GMP use

What SOC 2, ISO 27001, ISO 42001 and HITRUST each prove

Each of these is a real instrument with a defined scope, a defined issuer and a defined limit. Most procurement failures come from treating them as interchangeable badges rather than reading what each one actually asserts. This section is the reading guide: what the document is, who issued it, what it covers, and the specific sentence you should look for before you accept it as evidence.

Attestation

SOC 2: read the opinion and the exceptions, not the badge

SOC 2 is not a certification. It is an attestation examination resulting in a report that contains a practitioner's opinion, and only a licensed CPA firm can issue one. A vendor that says it is "SOC 2 certified" has already told you something about how carefully it reads its own compliance material.

The Type 1 and Type 2 distinction is about time, not thoroughness. A Type 1 opinion covers the suitability of the design of controls as of a point in time. A Type 2 opinion covers design and operating effectiveness throughout a period, normally between three and twelve months. A first-year vendor presenting a three-month Type 2 window has demonstrably less evidence than one presenting twelve months, and the period end date matters as much as the opinion itself. This also means "SOC 2 in thirty days" is either a Type 1 or a misstatement, because a Type 2 cannot be compressed below its own observation window and evidence generated after the window does not count.

The report has five sections and only two of them repay close reading. Section 1 is management's assertion. Section 2 is the independent service auditor's report, containing the opinion — unqualified, qualified, adverse or a disclaimer. Section 3 is the system description written by management. Section 4 sets out the criteria, the controls, the tests performed and the results of tests, and this is where exceptions live. Section 5 is other information not covered by the auditor's opinion, and can contain management commentary that has not been tested at all. If you read only two things, read the opinion paragraph in section 2 and the exceptions in section 4.

Then check the bridge letter, if there is one. A bridge letter, sometimes called a gap letter, covers the interval between the end of the last Type 2 period and today. It is written and signed by the vendor's management, not by the CPA firm, and it carries no opinion. It typically does not cover more than three months. A vendor whose only current evidence is a bridge letter is offering you a self-assertion in the shape of an assurance document, and it should be treated accordingly.

Finally, note what the instrument does not attempt. It does not prove the absence of exceptions. It does not cover categories the vendor did not select. It does not evaluate the product's functional correctness or the quality of any model. HIPAA, GDPR and AI-specific obligations are not in the Trust Services Criteria at all, so a SOC 2 report is not evidence of compliance with any of them.

Related evidence and next steps

Certification

ISO/IEC 27001:2022: the scope statement and the Statement of Applicability

ISO/IEC 27001:2022 is Edition 3, published in October 2022, and unlike SOC 2 it genuinely is a certification: an accredited certification body issues a certificate, normally on a three-year cycle with annual surveillance audits, following a stage 1 documentation audit and a stage 2 implementation audit.

Clauses 4 to 10 are the mandatory management-system requirements: context, leadership, planning, support, operation, performance evaluation and improvement. Annex A is a reference set of controls, and this is the single most misunderstood point for buyers. Annex A is not a mandatory checklist. Applicability is decided by the organisation's own risk assessment and recorded in the Statement of Applicability, which must justify any exclusions.

Two ISO 27001 certificates are therefore not comparable until you have read two documents: the scope statement printed on the certificate, and the Statement of Applicability behind it. A certificate scoped to "the corporate IT function at the London office" tells you nothing whatsoever about the SaaS platform you are buying. Asking for the scope statement is not an aggressive request; it is the only way to make the certificate mean anything, and a vendor that resists it is telling you where the scope boundary probably sits.

As reported by Secureframe, the 2022 revision of Annex A contains 93 controls in four themes — organisational, people, physical and technological, at 37, 8, 14 and 34 controls respectively — reduced from 114 controls across 14 domains in the 2013 version, with 11 new controls including threat intelligence, information security for use of cloud services, and data masking. We attribute those counts rather than quote them: ISO publishes the standard behind a paywall, so the breakdown here comes from a consistent secondary source rather than from the standard text.

For an AI vendor specifically, the interesting Annex A control is the cloud services one, and the interesting question is whether the model provider was in scope of the certified ISMS at all. It very often was not, because the certificate predates the AI feature.

Related evidence and next steps

AI management systems

ISO/IEC 42001 and the arrival of accredited AI certification

ISO/IEC 42001:2023, Information technology — Artificial intelligence — Management system, is Edition 1, published in December 2023, 51 pages, developed by ISO/IEC JTC 1/SC 42. It is a management system standard using Plan-Do-Check-Act, structurally parallel to ISO 27001, and it is the first of its kind.

As reported by Vanta, its Annex A contains 38 controls organised into nine control areas from A.2 to A.10: policies related to AI, internal organisation, resources for AI systems, assessing impacts of AI systems, the AI system life cycle, data for AI systems, information for interested parties, use of AI systems, and third-party and customer relationships. The same paywall caveat applies as for ISO 27001; treat the counts as reported rather than quoted.

The standard is appearing in AI vendor diligence for three converging reasons. It is the only certifiable AI governance standard, so it is the natural box for a procurement team to tick. ISO/IEC 42006:2025, published on 7 July 2025, now sets the additional requirements for bodies that audit and certify AI management systems, supplementing ISO/IEC 17021-1 rather than replacing it — which is what makes an accredited certificate possible rather than a self-declared one. And ANAB names both 42001 and 42006 in its accreditation programme, so accredited certification bodies now exist in practice.

The third reason is procurement mechanics rather than standards development. The Shared Assessments SIG 2026 release added a comprehensive mapping to ISO 42001, which pushes AI governance questions into every SIG-based third-party risk programme in a single step. If your organisation uses the SIG, you will be asking ISO 42001 questions whether or not anyone decided to.

What it does not prove is worth stating as clearly as what it does. It certifies that a governed process for AI risk exists. It does not certify any individual model's accuracy, robustness or safety, it does not establish fitness for a clinical or GxP purpose, and it is not an EU AI Act conformity assessment. Treat it the way you treat ISO 27001: valuable, scoped, and meaningless until you read the scope statement.

Related evidence and next steps

Prescriptive control sets

HITRUST, and where it does and does not carry weight

HITRUST publishes the HITRUST CSF, a harmonised control library, and runs a three-tier assessment portfolio. Its structural difference from both SOC 2 and ISO 27001 is worth understanding: assessments are performed by HITRUST-approved external assessors, but HITRUST itself performs central quality assurance and issues the certification.

The three tiers are straightforward. The e1 is a 43-control foundational assessment valid for one year. The i1 is a 182-control assessment focused on leading security practices, also valid for one year. The r2 is the tailored, risk-factor-driven assessment, valid for two years. Assessments are traversable, so work done for an e1 or i1 can be applied toward a more comprehensive assessment. HITRUST also now offers a purpose-built AI security assessment and certification for deployed AI systems and platforms.

HITRUST matters disproportionately in United States healthcare, payer and provider procurement, where it is frequently the expected instrument and its absence is a real commercial obstacle. It matters considerably less in EU pharmaceutical manufacturing, where Annex 11 and Chapter 7 are the operative expectations and a HITRUST certificate answers a question nobody in the room asked. If your vendor selection spans both, expect to hold both conversations and do not let either side assume its instrument is the universal one.

One practical caution about the marketing. HITRUST's own material describes the CSF as harmonising a large number of authoritative sources, but the figure given differs between pages on its own site, and no version number could be confirmed against a dated release page. Use the control counts, which are consistently published, and avoid quoting a harmonised-source figure or a framework version number in a supplier file where it will be checked.

Related evidence and next steps

The honest summary

What none of them prove

It is easier to compare these instruments by their blind spots than by their strengths, because the blind spots are what your assessment has to cover.

None of them evaluates the product's functional correctness. None of them evaluates model quality, hallucination rate, or behaviour under adversarial input. None of them establishes GxP fitness, and none of them produces documentation you can hold and explain from your own facility. All of them are scoped, and in every case the scope is chosen by the vendor rather than by you.

That is not a criticism of the instruments; they were built for a different purpose and they perform it. It is an argument about sequencing. Read the attestation or the certificate first, because it is fast and it tells you whether a security programme exists. Then ask the questions the instrument does not reach, which for an AI vendor are almost entirely about data flow, model provenance, retention and change control — and which are the subject of the next section.

The corollary for your own organisation is the same. A certificate you hold does not answer a customer's supplier qualification either. If you are on the receiving end of these requests, the work is described under security questionnaire readiness, and the technical evidence that persuades a reviewer fastest is usually a current penetration test with a closed remediation record rather than another badge.

Related evidence and next steps

Ask for three documents before you read anything else

Most vendor assessments start by sending a questionnaire. That is the slowest possible opening move, because it generates work for both sides and produces material you cannot verify. A faster opening is to request three specific documents, all of which already exist if the vendor is credible, and none of which takes more than an hour to read.

Each one is chosen because it resolves an ambiguity that a questionnaire would otherwise leave open for weeks. The scope statement tells you whether a certificate is about the product or about the head office. The results-of-tests section tells you whether the controls actually operated. The model provider clause tells you whose terms govern your data once it leaves the vendor's boundary.

If a vendor cannot produce all three, that is not automatically a reason to walk away — a young company may genuinely have no certificate — but it changes what the contract has to carry. In that situation the contract is doing the work the evidence was supposed to do, and it should be drafted accordingly. The primary texts behind these requests are ISO/IEC 27001:2022, the AICPA trust services criteria, and the model providers' own published data terms.

The scope statement and the Statement of Applicability

A certificate means nothing until you know what it covers and which Annex A controls were excluded, with the justification.

Section 2 and section 4 of the SOC 2 report

The opinion paragraph and the results of tests. Exceptions live in section 4. Section 5 is untested management commentary.

The named model provider and the governing clause

Not a marketing page. The clause, the deployment type, and the sub-processor list you are being asked to approve in advance.

The AI questions a standard questionnaire does not ask well

A SIG or a CAIQ will ask whether data is encrypted at rest, whether access is reviewed, and whether there is an incident response plan. Those are good questions and they are largely settled. The questions below are the ones where AI products differ materially from the SaaS products these instruments were designed for, and where a general answer and the true answer routinely diverge. Each is stated with the evidence a good answer produces, because the aim is not to catch the vendor out but to get to a specific, checkable statement.

Question one

Is our data used to train your models, or the underlying provider's models?

The credible answer is a contractual clause and a named provider, not a paragraph on a marketing page. The clauses exist and are public, which makes this one of the easier claims to verify properly.

Anthropic's Commercial Terms state it flatly: "Anthropic may not train models on Customer Content from Services." OpenAI's platform documentation is equally explicit for the API: "As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)", and its enterprise privacy page states that custom models "are yours alone to use and are not shared with anyone else". Microsoft's Azure AI Foundry documentation states that prompts, completions, embeddings and training data "are NOT used by providers of Models sold by Azure to improve their models or services", and adds that the models are stateless: "no prompts or completions are stored in the model".

So the question to the vendor is not whether such commitments exist. It is which of them applies to the deployment you are being sold, whether the vendor sits on the enterprise tier or the consumer tier of its provider, and what happens if the vendor changes provider. A vendor that reserves the right to substitute model providers at will has an unresolved problem under GDPR Article 28 and, in a GMP context, under Chapter 7 section 7.11.

There is a second-order version of this question that is asked far less often and matters more: is our data used to train the vendor's own models, as distinct from the underlying provider's? A vendor can accurately state that OpenAI does not train on your data while itself fine-tuning a classifier on pooled customer content. If a vendor fine-tunes on pooled customer data, that is a materially different product and it must be disclosed. Azure documents the isolated case as the norm — "fine-tuned models are exclusively available to the customer whose data was used to create the fine-tuned model" — which gives you a reference point for what good looks like.

Related evidence and next steps

Question two

Retention, deletion, and the abuse-monitoring log nobody mentions

This is where "we do not train on your data" quietly stops being the whole answer. Not training is a statement about model weights. It is not a statement about storage, and the storage is where your obligations bite.

OpenAI documents the default plainly: "By default, abuse monitoring logs are generated for all API feature usage and retained for up to 30 days, unless longer retention is required by law." Zero Data Retention and Modified Abuse Monitoring are available but, in the same documentation, "these controls are subject to prior approval by OpenAI and acceptance of additional requirements". Azure operates a parallel application path for modified abuse monitoring and publishes something unusually useful: a verifiable check, where a "ContentLogging" value in the capabilities list appears and reads false when abuse-monitoring logging is off.

The right request is therefore to see the approval and the configuration state, not to receive an assertion. This is one of very few places in AI diligence where a claim can be independently confirmed in a few minutes, and a vendor that has actually obtained modified abuse monitoring will be pleased to show you, because it was not easy to get.

The consequences reach beyond privacy hygiene. A HIPAA business associate agreement must, under 45 CFR section 164.504(e)(2)(ii)(J), require the associate at termination, if feasible, to "return or destroy all protected health information" and "retain no copies of such information". A rolling thirty-day store of prompts is a copy. Equally, the access, amendment and accounting-of-disclosures obligations at (E) to (G) collide with architectures in which content has been transformed into embeddings that cannot be individually located, corrected or enumerated. These two clauses are the first thing a knowledgeable healthcare reviewer asks about, and most AI vendors have not thought about either.

GDPR produces the mirror-image requirement. Article 28(3)(g) requires the processor to delete or return personal data at the end of the service unless retention is legally required. If the vendor's answer to deletion is "the data ages out of the log in thirty days", that is a retention schedule, not a deletion capability, and the difference will matter to your data protection officer.

Ask for the state, not the policy

Configuration evidence that logging is off, or the provider approval letter for modified monitoring.

Map every store

Prompts, completions, embeddings, vector indexes, threads, stored completions, evaluation datasets, support tickets.

Test the deletion path

Can a single subject's content be located and removed, or only a whole tenant?

Write the schedule down

Retention periods per store, in the contract, with the legal basis for any that exceed the service period.

Related evidence and next steps

Question three

Sub-processors, residency, and the deployment type

The model provider is a sub-processor. That single sentence resolves a surprising amount of confusion, because it places the model provider inside a legal framework everyone already understands rather than in a special category.

Under GDPR Article 28(2), a processor may not engage a sub-processor without prior specific or general written authorisation, and under a general authorisation it must inform the controller of intended changes and give the controller the opportunity to object. Article 28(4) requires equivalent obligations to flow down and keeps the original processor fully liable for the sub-processor's performance. A vendor that will not name its model provider, or that reserves the right to swap it silently, has an unresolvable problem under those provisions before you even reach GMP. Chapter 7 section 7.11 then adds the harder requirement in a regulated context: prior evaluation and approval, with the suitability assessment information passed through.

Residency is the question most often answered wrongly, and the reason is that it is now a deployment-type question rather than a company-level one. Microsoft states the mechanics without euphemism: "For any deployment type labeled 'Global,' prompts and responses may be processed in any geography where the relevant model sold by Azure is deployed." A vendor can be an EU company on EU infrastructure and still process prompts globally, and the person filling in your questionnaire may not know it.

Break residency into four separate answers. Where is inference processed? Where does data sit at rest? Where does the abuse-monitoring and human-review store sit? And where are the humans who may read flagged content physically located? Azure answers the last of these for EEA deployments — "the authorized Microsoft employees are located in the European Economic Area" — and documents the access controls in a way that is worth using as a benchmark for other vendors: human reviewers access flagged data "via point wise queries using request IDs, Secure Access Workstations (SAWs), and Just-In-Time (JIT) request approval granted by team managers". That is the level of specificity to hold vendors to. If a vendor cannot describe its own human-review path in comparable terms, it probably has not designed one.

For transfers out of the EEA, the current standard contractual clauses are in Commission Implementing Decision 2021/914 of 4 June 2021, which replaced the three earlier sets adopted under Directive 95/46 and are modular. For United States processing, the EU-US Data Privacy Framework adequacy decision remains the simplest route, and the EU General Court dismissed the Latombe challenge to it on 3 September 2025. Treat the framework as valid but not permanently settled — the court itself stressed that the Commission "is required to monitor continuously the application of the legal framework" on which the decision is based, and an appeal was filed. A vendor relying solely on the framework should still be able to fall back to standard contractual clauses.

Related evidence and next steps

Question four

Tenant isolation, vector stores, and fine-tuning separation

Isolation questions in a standard questionnaire tend to be answered at the level of the application database. For an AI product the interesting boundaries are elsewhere, and they are frequently newer, less tested and less documented than the application itself.

Ask specifically about three things: the abuse-monitoring store, the vector store, and any stateful entity such as threads or stored completions. Azure states that its abuse monitoring data store "is logically separated by customer resource", which is an honest description — logical rather than physical separation is the norm across the industry and should be stated as such rather than dressed up. A vendor claiming physical isolation should be able to explain what that means concretely, because the phrase is often used loosely.

Vector and embedding stores deserve their own line of questioning because they are where the assumptions from the relational world break down. OWASP's Top 10 for LLM Applications names "Vector and Embedding Weaknesses" as LLM08 precisely because embeddings created from one tenant's documents can leak information through a shared index, and because embeddings are difficult to delete selectively, difficult to audit, and difficult to explain to a data protection officer. The same list names "Excessive Agency" as LLM06, which is the risk that matters most as vendors add tool use and autonomous actions to their products.

Fine-tuning separation should be answered in one sentence and backed by architecture. The reference statements exist: Azure documents that fine-tuned models are exclusively available to the customer whose data created them, and OpenAI states that custom models are not shared with anyone else. If the vendor's answer is longer or more conditional than those, the condition is the finding.

Finally, ask what happens on the retrieval side of the product rather than the model side. In our first engagement at a clinical-stage biotech client, connected AI reached employee-information files in Box through existing user permissions — nothing was misconfigured at the vendor and no control failed. The permissions were simply broader than anyone had realised, and the assistant was the first tool efficient enough to notice. A vendor assessment cannot find that, which is why it belongs in a separate piece of work on your own environment.

Related evidence and next steps

Question five

Human oversight, and change control on the model itself

The last two questions are the ones a security questionnaire never asks and a quality reviewer always asks. They are also the two that most often decide whether a tool can be used in a regulated process at all.

Human oversight is not a user-experience question in a GxP context; it is the compliance question. Draft Annex 22 requires that where generative models are used in non-critical applications, "personnel with adequate qualification and training should always be responsible for ensuring that the outputs from such models are suitable for the intended use, i.e. a human-in-the-loop (HITL)". That phrase is only meaningful if your organisation can say who reviews, what evidence they inspect, what qualification they hold, what they are permitted to approve, how exceptions are handled and what record is retained. A vendor can support that design or obstruct it, and the difference is visible in the product: whether outputs carry sources, whether a reviewer can see what the model was given, whether an approval is recorded.

Change control on the model is the question that separates a supplier assessment from a security review most sharply. A silent model version bump is a change to a validated system. Draft Annex 11 section 7.5 addresses it directly by requiring agreement on "the process for release of new system versions and on the regulated user's possibility to test these prior to release". In practice this means asking whether the vendor offers pinned model versions, what notice period applies to a change, whether behaviour-affecting prompt or system-message changes are treated as versioned changes, and what regression evidence the vendor produces. Most AI vendors have a deployment pipeline built for the opposite objective, which is to ship improvements continuously and invisibly.

The EMA reflection paper on artificial intelligence in the medicinal product lifecycle, adopted by CHMP on 9 September 2024, sets the expectation for third-party models where the stakes are high: "if a third-party AI model or service is to be used within the medicinal product lifecycle with high regulatory impact or high patient risk, it is expected that the manufacturer of the system has provided such details through a methodology qualification process … covering the specific context of use". It also notes, in a sentence worth quoting to any vendor that considers these requirements unusual, that the applicable requirements "may in some respects be stricter than what is considered standard practice in the field of data science".

Governance evidence is the cheap part of this and can be requested early: an ISO/IEC 42001 certificate with its scope, alignment with the NIST AI Risk Management Framework, an AI-CAIQ or STAR for AI submission, model cards, and evaluation results. None of these is conclusive and all of them are informative, mostly because of what a vendor chooses not to provide.

A silent model version change is a change to a validated system. That is the question a security questionnaire never asks and a GxP reviewer always asks.

Related evidence and next steps

Draft Annex 22 and generative AI in critical GMP applications

The draft Annex 22 gives additional guidance to Annex 11 for systems with AI models embedded. It covers static, deterministic models only. Dynamic models that "continuously and automatically learn and adapt performance during use" are excluded and "should not be used in critical GMP applications", as are probabilistic ones. The draft adds that it "does not apply to Generative AI and Large Language Models (LLM)" in that setting. It is a draft, and it is not a ban on LLMs in pharma.

Regulatory documentation review for artificial intelligence in a GMP manufacturing quality system

The EU AI Act timeline after the July 2026 Omnibus

Regulation (EU) 2024/1689 entered into force on 1 August 2024 and applied generally from 2 August 2026. The Digital Omnibus on AI, Regulation (EU) 2026/1744, was published in the Official Journal on 24 July 2026 and took effect on 27 July 2026. It defers most Chapter III high-risk duties to 2 December 2027 for stand-alone Annex III systems and 2 August 2028 under Annex I. Article 5, the GPAI rules and the Article 50 transparency duties did not move.

European Union artificial intelligence regulation compliance planning for a life sciences organisation

What Part 11 and MHRA expect a SaaS vendor to be able to produce

21 CFR Part 11 asks for things most SaaS vendors have never been asked for: validation under §11.10(a); "accurate and complete copies of records in both human readable and electronic form suitable for inspection" under §11.10(b); audit trails where "record changes shall not obscure previously recorded information" under §11.10(e); and signatures carrying the printed name, date, time and meaning under §11.50. Section 11.1(e) subjects the system to FDA inspection.

Laboratory data systems producing audit trails and electronic records for regulatory inspection

The instruments you will be offered, and the contracts that carry them

When you ask an AI vendor for evidence, you will normally be offered one of a small number of standard artefacts. Each has a publisher, a licensing model, a current version and a purpose, and each is frequently misdescribed — including by vendor-risk content that has not been updated in years. This section is a short field guide to what you are being handed, followed by the contracts that have to carry whatever the artefacts do not.

Field guide

SIG, CAIQ, AICM, HECVAT — and one that is not current at all

These are the general-purpose instruments. None of them is a life-science instrument, and confusing a good general questionnaire with a supplier qualification is the error this whole page exists to prevent. Within their own scope, though, they are genuinely useful and knowing which one you are looking at saves a great deal of time.

The Standardized Information Gathering questionnaire, or SIG, is published by the Shared Assessments Program as part of its third-party risk management toolkit, on an annual release cycle. SIG Lite is the low-risk screening tier and SIG Core is the comprehensive tier for higher-risk providers; both can be scoped by risk domain and control family. The SIG is licensed rather than free, so access requires membership or a product subscription, and exact question counts are behind that licence — figures circulating on vendor blogs should be attributed to whoever published them rather than stated as fact. The 2026 release added a comprehensive mapping to ISO 42001, and in March 2026 Shared Assessments launched SIG Evolution, moving the questionnaire from a spreadsheet to a browser-based platform while retaining Excel compatibility.

The Consensus Assessments Initiative Questionnaire, or CAIQ, is the Cloud Security Alliance's yes-or-no questionnaire aligned one-to-one with the Cloud Controls Matrix. CCM v4.1, released on 27 January 2026, contains 207 controls across 17 security domains. CAIQ v4 contains 261 questions, down from 310 in v3.1, and CAIQ-Lite contains 124 questions while still covering every CCM domain. A completed CAIQ submitted to the CSA STAR Registry is STAR Level 1, which is a free, public, self-assessment. STAR Level 2 is third-party audited and comes in three variants: STAR Certification, built on ISO/IEC 27001 plus the CCM with three-year validity; STAR Attestation, built on a SOC 2 engagement using AICPA criteria plus the CCM with one-year validity; and C-STAR for Greater China.

For AI products specifically there is now a purpose-built instrument. The CSA AI Controls Matrix, released on 9 July 2025 and updated on 30 October 2025, contains 243 control objectives across 18 security domains for cloud-based AI systems, and ships with a companion AI-CAIQ and a STAR for AI Level 1 submission guide. It is free. For a life-science buyer this is currently the closest thing to a standard AI-specific self-assessment, and asking a vendor whether it has completed one is a fast way to learn how seriously it has thought about AI-specific risk.

The Higher Education Community Vendor Assessment Toolkit, HECVAT, is published by EDUCAUSE with Internet2 and REN-ISAC; the current version is HECVAT 4, revision 4.1.5, and version 4 explicitly added privacy and AI questions. It is free to colleges, universities and their vendors, while third-party risk platforms need a licence to integrate it. It matters to life sciences chiefly when the counterparty is an academic medical centre or a university research office, which for a clinical-stage biotech is a very common partner.

Finally, one negative finding that is worth carrying into any vendor conversation. Google's Vendor Security Assessment Questionnaire, VSAQ, is still listed as a current instrument in a great deal of vendor-risk content. It is not. The repository was archived on 25 November 2022 and is read-only. It was historically influential — it popularised the conditional, branching questionnaire — but it is unmaintained, and any listicle that presents it as current has not been checked recently. That is a useful signal about the rest of the page it appears on.

SIG

Licensed, annual cycle, Lite and Core tiers, ISO 42001 mapping added in 2026, SIG EV platform launched March 2026.

CAIQ

261 questions in v4, 124 in CAIQ-Lite, mapped to CCM v4.1 at 207 controls across 17 domains.

AICM and AI-CAIQ

243 control objectives across 18 domains, free, AI-specific, released July 2025.

HECVAT 4.1.5

Free from EDUCAUSE, privacy and AI questions added in version 4, relevant when the partner is academic.

Related evidence and next steps

The pharma-specific strand

PSCI and Rx-360 are not security instruments, and there is no pharma SIG

There is no single pharmaceutical equivalent of the SIG for information security. There are three separate strands, and treating one as another is a common error that is immediately visible to a pharma quality lead.

The Pharmaceutical Supply Chain Initiative is a member-driven scheme whose Principles for Responsible Supply Chain Management address five areas of responsible business practice: ethics, labour, health and safety, environment, and management systems. Data privacy and security appears as a sub-principle under ethics, not as a security framework. PSCI runs a shared audit programme so that one audit can serve many members. It is a responsible-business and ESG instrument. A page that presents it as a security questionnaire will be wrong in front of the person it is trying to persuade.

Rx-360 is an international pharmaceutical supply chain consortium with an audit operations working group and a shared-audit programme, oriented to GMP and material-integrity supply chain risk rather than to information security. It is relevant to a materials supplier and largely not relevant to a software vendor.

The real pharmaceutical instrument for a software vendor is the GxP supplier qualification questionnaire and audit, driven by Annex 11 sections 3 and the draft 7, EU GMP Chapter 7, ICH Q10 section 2.7, and GAMP 5 Appendix M2 on supplier assessment. These are company-specific documents rather than a published standard instrument, which is precisely why they are unpredictable, why they arrive without warning, and why they are expensive to answer cold.

GAMP 5 Second Edition, published by ISPE in July 2022, is the operational guide most quality organisations will be working from. Its structure is public even though the text is not: main-body Chapter 7 covers supplier activities including sub-supplier assessments at 7.6, section 8.3 covers leveraging supplier input, Appendix M2 covers supplier assessment, Appendix M4 covers software and hardware categories, and Appendix D11 covers artificial intelligence and machine learning. ISPE's own framing of the second edition is that regulated companies should "maximize supplier involvement to leverage knowledge, experience, and documentation where possible" — leverage, not abdicate. The software categories, 1, 3, 4 and 5, drive the depth of both supplier assessment and validation, and a cloud AI product with behaviour that changes between versions is exactly the thing that pushes a system up a category.

Related evidence and next steps

Contracts

The clauses that decide what the evidence was worth

Evidence tells you what is true today. The contract decides what happens when it stops being true. For an AI vendor, where the model, the provider, the deployment type and the feature set can all change without any action by you, the contract is doing more work than in ordinary software procurement.

If protected health information is in scope, a business associate agreement is not optional. Under 45 CFR section 164.502(e)(1)(i) a covered entity may disclose PHI to a business associate only where it "obtains satisfactory assurance that the business associate will appropriately safeguard the information", documented through a written contract. Subcontractors of business associates are themselves business associates under 45 CFR section 160.103, which places the model provider squarely in scope. The required contents at section 164.504(e)(2) include using appropriate safeguards and complying with the Security Rule, reporting unpermitted uses and breaches, ensuring subcontractors agree to the same restrictions, making PHI available for access, amendment and accounting of disclosures, making internal practices and records available to the Secretary, and returning or destroying all PHI at termination where feasible.

There is also a trap in the same regulation that turns diligence into an ongoing duty rather than a one-time act. Section 164.504(e)(1) puts a covered entity out of compliance if it knew of a pattern of activity constituting a material breach by the business associate and failed to take reasonable steps to cure it or, failing that, to terminate. Knowing is enough to create the obligation. That is one of the practical reasons a vendor assessment should be a recurring backlog item rather than a one-off gate.

On the data protection side, GDPR Article 28(1) requires controllers to use only processors "providing sufficient guarantees to implement appropriate technical and organisational measures", and Article 28(3) requires a binding contract covering subject matter, duration, nature and purpose, types of personal data and categories of data subjects, together with obligations (a) to (h). Article 28(3)(h) is the clause that entitles a pharmaceutical customer to audit an AI vendor: it requires the processor to "allow for and contribute to audits, including inspections, conducted by the controller or another auditor mandated by the controller". A data processing agreement that converts that into "you may review our SOC 2 report" is narrowing a statutory right, and a competent reviewer will notice. It is also, not coincidentally, the same audit right that Annex 11 section 3.2 and Chapter 7 already assume you have.

For a GxP context, add the draft Annex 11 section 7.5 items to the negotiation list explicitly: conditions for supplier audits, support during regulatory inspections if requested, communication processes for quality and security issues, an exit strategy by which you retain control of system data, and your ability to test new versions prior to release. These are draft requirements, and we say so, but they are also the clauses that are hardest to add after signature, which makes them cheap to ask for now and expensive to retrofit later.

A useful test of the whole package: ask what happens on the day you terminate. If the answer describes deletion but not return, or return but not in a readable, complete form with audit trails and metadata, then the contract has a hole exactly where the regulator will look. MHRA's data integrity guidance frames the same point for cloud services, requiring that the technical agreement "ensure timely access to data (including metadata and audit trails) to the data owner and national competent authorities upon request", and that arrangements exist "for the restoration of the software/system as per its original validated state".

BAA

Required where PHI is involved. Watch the return-or-destroy clause against the vendor's log retention.

DPA and Article 28(3)(h)

A genuine audit right, not a right to read a report. Sub-processor list with a change-objection mechanism.

Transfers

SCC module selection under Decision 2021/914, with the Data Privacy Framework as an alternative rather than the only route.

GxP clauses

Audit conditions, inspection support, exit strategy with data return format, and pre-release testing of new versions.

Related evidence and next steps

The question, the weak answer, and the evidence that settles it

A working extract from the assessment worksheet. The middle column is not dishonest — it is what a sincere person writes when the question is asked at the wrong level of specificity.

What you askThe answer you usually getThe evidence that settles it
Do you train on our data?No, we never train on customer data.The clause, plus the named model provider and its published terms
How long do you keep prompts?Only as long as necessary.A per-store retention schedule, including abuse-monitoring logs
Can we get zero data retention?Yes, that is available.The provider approval, and the configuration state showing logging off
Where is our data processed?We are hosted in the EU.The deployment type, plus where reviewers of flagged content sit
Who are your sub-processors?A list is available on request.A published list with a change-notification and objection mechanism
Is our tenant isolated?Yes, fully isolated.Whether separation is logical or physical, and for which store
Do you fine-tune on our data?Only to improve your experience.A statement that custom models are exclusive to the customer
How do you handle model updates?We follow a change management policy.Version pinning, notice period, and regression evidence per release
Are you ISO 27001 certified?Yes, here is the certificate.The scope statement and the Statement of Applicability
Can we audit you?We can share our SOC 2 report.A contractual audit right under GDPR Article 28(3)(h)
Will you support an inspection?We have never been asked.A written inspection-support clause in the agreement
What happens if we leave?We delete everything within 30 days.A defined export format including metadata and audit trails

How the assessment runs, and what it is not

The engagement is deliberately small: one named vendor, one intended use, a fixed price and a duration measured in days rather than weeks. That shape is chosen for a specific reason. An open-ended assessment gets deferred until after the purchase decision, at which point it has become documentation rather than diligence. A short one runs before the decision, which is the only time it can change anything.

The work has four parts. We fix the intended use and the risk classification with your quality, IT and legal stakeholders, because that decides the depth of everything else. We review the evidence the vendor has already produced — certificates with their scope statements, the attestation report's opinion and results-of-tests sections, published data terms, sub-processor lists, penetration test summaries. We ask the AI-specific and GxP-specific questions the vendor has probably not been asked before, and record which answers were verified, which were accepted as disclosure, and which were declined. Then we write the contract asks: the specific clauses to negotiate before signature, in language your counsel can use.

The engagement is led by a named security architect on the IntuitionLabs expert bank: an independent consultant with seventeen years in information security, including a decade on the central security team of a major enterprise infrastructure vendor — design and architecture review, security requirements review, source code review, penetration testing and vulnerability response — preceded by client-facing consultancy work leading mobile penetration testing. The specialist is named to you in the engagement documents before work begins, and is a specialist on our expert bank rather than an employee.

What this is not: IntuitionLabs is not a CPA firm, not an accredited certification body and not a notified body. We do not issue SOC 2 reports, ISO certificates or AI Act conformity assessments, and no assessment we write transfers your responsibility as the regulated user. IntuitionLabs itself does not hold SOC 2 or ISO 27001 certification; our own posture is described on our Trust Center and security pages, and we would rather you read that here than discover it later. The boundary is not merely good manners: ISO/IEC 17021-1 clause 5.2.5 bars a certification body from providing management system consultancy, and the AICPA Code of Professional Conduct treats designing, implementing or maintaining internal control as a management responsibility that impairs independence. As one practitioner puts it, the firm constructing your programme cannot issue the report. That prohibition is exactly what creates a legitimate preparation and assessment role, and it works only if the preparer stays on its side of the line.

Days, not weeks

Fixed price, one vendor, one intended use. Short enough to run before the purchase decision rather than after it.

A file, not an opinion

Evidence, gaps, unanswered questions and contract asks, written so your quality unit can adopt it into its own supplier file.

A defensible buy recommendation

Including the honest version: what was not verified, and what the contract has to carry because the evidence did not.

Where vendor assessment sits in the AI security service line

A vendor assessment answers a question about somebody else's organisation. Most of the risk in an AI deployment sits inside your own, which is why these engagements are designed to be run together rather than in isolation. The hub page describes how they sequence.

AI access exposure review

What your assistant can already reach through existing permissions. A vendor assessment cannot find this, because nothing at the vendor is misconfigured.

See the exposure review

Classification and access model

The permission and classification structure a retrieval-enabled assistant inherits. Fixing it is usually cheaper than restricting the assistant.

See the access model

Security questionnaire readiness

The mirror image of this page. Here you assess a vendor; there you prepare to be assessed, with a control narrative and an answer bank that stay consistent.

Prepare to be assessed

Penetration testing and secure code review

Independent technical testing under a signed authorization letter. A current test with a closed remediation record persuades a reviewer faster than a badge.

See testing services

Microsoft Purview for life sciences

Labelling, data loss prevention and audit in a Microsoft estate, so that classification decisions are enforced rather than documented.

See the Purview work

Computer system validation

Where the supplier assessment becomes an input to qualification: risk assessment, requirements, supplier evidence, testing strategy and traceability.

See validation services

Questions about AI vendor assessment

No, and the reason is structural rather than a matter of degree. A SOC 2 report is an attestation under AICPA standards in which a licensed CPA firm gives an opinion on controls the vendor selected, over a window the vendor chose, against the Trust Services Criteria. Those criteria do not address patient safety, product quality, electronic record integrity, or the regulated user's inability to delegate responsibility. The draft revision of EU GMP Annex 11 states in section 7.1 that a regulated user relying on a vendor's qualification "remains fully responsible", and section 7.4 requires that the documentation be "accessible and can be explained from their facility". A report held under NDA that your quality lead has never read and cannot explain does not meet that description. A SOC 2 report is genuinely useful evidence about the vendor's security programme, and it is worth reading properly. It is simply not the instrument that answers the GxP question.
They are mirror images of the same problem. This page is about assessing someone else: you are the life-science company, an AI vendor wants your data, and you have to decide whether to buy and on what terms. Security questionnaire readiness is about being assessed: you are the company receiving a SIG, a CAIQ or a pharma supplier questionnaire from a customer, and you need to answer it accurately, consistently and without overclaiming. The two engagements share a body of knowledge and almost no deliverables. If you both buy AI tools and sell software into life sciences, you will eventually want both, and running the assessment side first tends to make the readiness side much faster, because you have already seen what a serious reviewer looks for.
The current regulatory signal on this is the draft EU GMP Annex 22 on Artificial Intelligence, and it is unusually direct: the draft states that it "does not apply to Generative AI and Large Language Models (LLM), and such models should not be used in critical GMP applications". Two qualifications matter. First, this is a draft; the consultation closed on 7 October 2025 and no final text had been published as of 27 August 2026. Second, the scope is critical GMP applications in the manufacture of medicinal products and active substances, meaning direct impact on patient safety, product quality or data integrity. It is not a ban on LLMs in pharmaceutical companies. For non-critical use the same draft requires a human in the loop: qualified, trained personnel remain responsible for judging whether the output is suitable for its intended use. In practice this means an LLM assistant supporting a scientist is a very different governance question from an LLM making or gating a batch decision, and the two should never be treated as one control problem.
Ask for the contract clause and the named model provider, not the marketing page. The clause exists and is quotable for the major providers: Anthropic's Commercial Terms state that "Anthropic may not train models on Customer Content from Services"; OpenAI's platform data documentation states that "As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)"; Microsoft's Azure AI Foundry data-privacy documentation states that prompts, completions, embeddings and training data "are NOT used by providers of Models sold by Azure to improve their models or services". The question that actually matters is which of these applies to the deployment the vendor is selling you, and whether the vendor can change model provider without telling you. Two vendors with identical marketing copy can sit on completely different sides of that line.
It is the place where "we do not train on your data" quietly stops being the whole answer. OpenAI's documentation states that "By default, abuse monitoring logs are generated for all API feature usage and retained for up to 30 days, unless longer retention is required by law". Zero Data Retention and Modified Abuse Monitoring exist, but the same page states they are "subject to prior approval by OpenAI and acceptance of additional requirements". Azure runs a parallel application process and, helpfully, publishes a verifiable check: a "ContentLogging" value in the capabilities list appears and reads false when abuse-monitoring logging is off. The practical rule is to ask the vendor to show the approval and the configuration state rather than assert it. This matters far beyond privacy hygiene: a HIPAA business associate agreement requires the associate, at termination and where feasible, to "return or destroy all protected health information" and "retain no copies", and a 30-day monitoring store containing prompts is a copy.
Residency is a deployment-type question, not a headquarters question, and this is the single most common way a diligence answer goes wrong. Microsoft states it plainly for its own platform: "For any deployment type labeled 'Global,' prompts and responses may be processed in any geography where the relevant model sold by Azure is deployed." A vendor can be an EU company, on EU infrastructure, and still route inference globally. Break the question into four parts and get four separate answers: where inference is processed; where data is stored at rest; where the abuse-monitoring and human-review store sits; and where the humans who may read flagged content are physically located. Azure documents the last of these for EEA deployments, noting that the authorized reviewers "are located in the European Economic Area". MHRA's data integrity guidance makes the same point from the regulator's side, requiring that "the physical location where the data is held, including the impact of any laws applicable to that geographic location, should be considered".
No. ISO/IEC 42001:2023 is a management system standard, published in December 2023 and developed by ISO/IEC JTC 1/SC 42. It certifies that an organisation has a governed process for AI risk — policies, impact assessments, lifecycle controls, third-party management — using the same Plan-Do-Check-Act structure as ISO 27001. It does not certify any individual model's accuracy, robustness or fitness for a clinical or GxP purpose, and it is not a substitute for an EU AI Act conformity assessment or for your own qualification of the system. It is nonetheless becoming meaningful evidence for a different reason: ISO/IEC 42006:2025, published on 7 July 2025, sets the requirements for bodies that audit and certify AI management systems, which is what turns a self-declared certificate into an accredited one. Read the scope statement and the Statement of Applicability exactly as you would for ISO 27001.
It is a well-structured self-assessment, which is a genuinely useful starting point and a much better artefact than a bespoke Word document. The Consensus Assessments Initiative Questionnaire maps one-to-one to the Cloud Controls Matrix, so answers are comparable across vendors; CCM v4.1, released 27 January 2026, contains 207 controls across 17 security domains, and CAIQ v4 contains 261 questions with a 124-question CAIQ-Lite. A CAIQ published to the CSA STAR Registry is Level 1, which is explicitly a self-assessment. STAR Level 2 is third-party audited. The distinction is the whole point: a completed questionnaire tells you what the vendor says about itself, in a format you can compare. Independent research cited by Vanta reports that only 34% of third-party risk management professionals believe questionnaire responses are accurate, which is why the useful move is to verify a small number of high-consequence answers against primary evidence rather than to send more questions.
Not by itself, and treating it as disqualifying will cost you good vendors. Certifications are demand-driven artefacts; a two-year-old company selling to five customers may reasonably not have one yet. What is disqualifying is the absence of the things that cost little and prove much: a named model provider with its data terms cited by clause; a current penetration test with a remediation record; a written data retention and deletion position; a sub-processor list with a change-notification mechanism; a data processing agreement that preserves the audit right in GDPR Article 28(3)(h) rather than converting it into "you may review our SOC 2 report"; and, for a GxP context, a willingness to accept audit, inspection support, an exit strategy and a pre-release testing window in the contract. A small vendor that can produce those is a better risk than a certified vendor that will not sign them.
No. IntuitionLabs is not a CPA firm, not an accredited certification body and not a notified body. We do not issue SOC 2 reports, ISO certificates or AI Act conformity assessments, and nothing we produce transfers your responsibility as the regulated user — under draft Annex 11 section 7.1 that responsibility cannot be transferred at all. What we produce is a documented, evidence-referenced assessment that your quality, IT, legal and security stakeholders can read, challenge and adopt into your own supplier qualification file, together with the specific contract language and the specific follow-up questions that the evidence did not answer. We should also state our own position plainly, because a page about vendor diligence written by a company concealing its own posture would deserve to be discounted: IntuitionLabs does not hold SOC 2 or ISO 27001 certification. Our current posture is described on our Trust Center and security pages.
It is scoped as a short, fixed-price review of one named vendor and one intended use, sized in days rather than weeks. The fixed price exists for a specific reason: an open-ended assessment gets deferred until after the purchase decision, at which point it is documentation rather than diligence. Because it is small and repeatable, it drops naturally into an AI acceleration backlog as a recurring item — one line per vendor, run before each tool is approved rather than in an annual scramble. The output is written to be reusable: the parts that describe the standards and the questions are stable across vendors, and only the evidence section changes. Where the work belongs inside a wider adoption programme, it is usually sequenced alongside the AI Acceleration Program and the policy work described under AI policy and governance.
As a buyer you are usually a deployer rather than a provider, and the obligations differ accordingly. The dates are the part most published material now gets wrong. Regulation (EU) 2024/1689 entered into force on 1 August 2024 and became generally applicable on 2 August 2026. Then the Digital Omnibus on AI, Regulation (EU) 2026/1744, was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026, deferring the bulk of the Chapter III high-risk obligations to 2 December 2027 for stand-alone Annex III systems and 2 August 2028 for high-risk AI embedded as a safety component of a product regulated under Annex I. What did not move: the Article 5 prohibitions, the general-purpose AI obligations applicable since August 2025, and the Article 50 transparency duties that applied from 2 August 2026. Many widely-cited timeline pages still show the pre-Omnibus schedule, so check the regulation number before you rely on a date. The change is summarised by White & Case and Gibson Dunn, with a consolidated enforcement calendar in the Data Protection Report. Penalties sit in Article 99: up to EUR 35,000,000 or 7% of total worldwide annual turnover for breaches of the Article 5 prohibitions, whichever is higher, with SMEs and start-ups capped at whichever figure is lower.
Bring Us One Vendor

Bring Us One Vendor

Name the vendor, the intended use, and whether a GxP process or patient data is involved. We will tell you what evidence to request, what the contract has to carry, and whether the assessment is worth running at all.

Book a Meeting

© 2026 IntuitionLabs. All rights reserved.