Claude

IntuitionLabs is now a member of the Claude Partner Network – AI training and upskilling with Claude for pharma and biotech. Book a call.

IntuitionLabs
Abstract eTMF evidence network connecting clinical document layers

eTMF Intelligence That Shows Its Evidence

Classify, test completeness, inspect document quality, reconcile clinical systems, and rank remediation while the eTMF remains the governed system of record.

01

Study first

Begin with one study and a bounded document set before portfolio rollout.

05

Connected controls

Classification, completeness, QC, reconciliation, and readiness share one evidence chain.

100%

Cited outputs

Every extracted field and finding points back to its document, page, or source record.

Human

Consequential decisions

Signatures, approvals, and policy decisions stay with authorized reviewers.

Solution overview

A reasoning layer for the trial master file

The eTMF stays the system of record. The agent turns documents, metadata, expected-document logic, and operational data into a reviewable evidence graph.

Synthetic eTMF classification and filing interface with extracted Vault fields and agent activity
Interactive design prototype using entirely synthetic study data. The screen demonstrates the target workflow, not a connection to a production Vault.

An electronic Trial Master File is more than a document repository. It is the evidence base that lets a sponsor, investigator, auditor, or inspector reconstruct how a clinical trial was planned, conducted, supervised, and documented. ICH E6(R3) describes essential records as documents, data, and relevant metadata that support trial management and allow evaluation of the methods, influencing factors, actions, and reliability of the results. That definition reaches beyond the question of whether a PDF exists in a folder.

Conventional eTMF controls are strongest when the expectation is already known. A study template defines an expected document list, filing staff select a type and subtype, lifecycle states control approval, and dashboards count filed artifacts. Those controls remain essential. The gap appears when the evidence needed for a specific study emerges from the content of other records or from activity in connected clinical systems. A delegation log can name people whose CVs should exist. A monitoring plan can define a filing window. A current consent form can conflict with the version recorded for enrolled subjects. Those relationships are difficult to express as a fixed checklist.

The IntuitionLabs eTMF Intelligence Agent adds a traceable reasoning layer across those relationships. It reads the content and metadata the customer has authorized, derives explicit expectations, tests the records against defined quality and consistency rules, and presents the result with the evidence required for review. It does not edit the source document. It does not invent an approval. It does not replace the reviewer who owns a regulated decision. Its job is to narrow the field of attention and make every recommendation inspectable.

The product is designed for sponsors, CROs, and clinical operations teams that already have an eTMF and want stronger continuous oversight. A Veeva Vault eTMF deployment can use Vault APIs, VQL, Direct Data API extracts, native placeholders, EDL items, and workflow tasks. Other eTMF platforms can expose the same logical contract through their own APIs or governed exports. The application layer remains separate so the detection logic, evaluation evidence, and remediation workflow can evolve under the customer’s change-control model.

The screens on this page come from a synthetic design prototype built around a fictional study, CRX-214. Every person, site, vendor, document, score, and finding is fabricated. Each document was generated specifically so the relevant evidence appears on the cited page. That constraint matters: a polished interface is useful only when the underlying result can be challenged against a record. The prototype therefore shows both the operational experience and the standard of evidence expected from a production implementation.

  • Works with the eTMF as system of record
  • Preserves document, page, field, and source attribution
  • Separates model recommendations from governed approvals
  • Supports platform-specific connectors without locking the reasoning model to one vendor

Document intelligence

Classification that shows its work

A reviewer sees the document, the proposed filing location, the extracted Vault fields, the confidence threshold, and the provenance for every value.

Document classification is the first visible use case, but the product treats it as an evidence problem rather than a label-prediction problem. The agent identifies the document type, subtype, and classification, maps that result to the customer’s configured taxonomy, and extracts the metadata required for the record. A Form FDA 1572, for example, can expose the study identifier, site, investigator, country, document date, and version. Each extracted value is paired with the page, section, or field pattern that produced it so the reviewer can verify the value in context.

Confidence is a routing signal, not a claim of truth. A threshold should be set from a labelled sample that represents the customer’s real mix of native PDFs, scans, languages, templates, and document quality. Results above the accepted threshold can follow the agreed low-risk path. Results below it are held for a person. A low-confidence scan is not silently forced into the most likely class, and a conflicting date is not resolved by whichever candidate appeared first in the prompt. The unresolved state is part of the output.

The disposition also depends on the consequence of the action. Creating a draft record with proposed metadata carries a different risk from promoting a document through a lifecycle action that requires a Part 11 electronic signature. A production configuration can therefore apply different thresholds and human gates by document type, lifecycle, and field. High-volume, low-risk administrative filing can be streamlined while regulated approval remains explicit and attributable.

The same intake run can create downstream expectations. If the agent reads a delegation log and identifies eight sub-investigators, it can compare those names with current CV holdings. The initial classification task then becomes the trigger for a completeness rule. This avoids a common fragmentation problem in which one automation classifies the file, a second dashboard counts documents, and neither preserves the semantic relationship between what the document says and what the TMF should contain.

A production implementation begins with a controlled vocabulary and a representative evaluation set. The customer defines the authoritative document types, required fields, matching criteria, and permitted actions. IntuitionLabs builds deterministic checks around the model output, records the model and prompt version, and measures performance per document type. When a model, prompt, extraction schema, or threshold changes, the labelled set provides the regression evidence needed for a documented release decision.

  • Type, subtype, classification, and TMF Reference Model mapping
  • Field-level provenance rather than an opaque completed form
  • Thresholds calibrated by document type and consequence
  • Explicit hold states for low-quality scans and conflicting evidence

Expected-document intelligence

Completeness derived from the study itself

Template expectations stay intact. The agent adds a second layer for evidence implied by people, milestones, plans, approvals, and active-site facts.

Synthetic TMF completeness dashboard with zones, agent-derived expected documents, owners, and due dates
The prototype distinguishes template EDL coverage from nine agent-derived expectations, each linked to the document that created it.

An expected document list gives the study a controlled baseline. The agent does not discard that baseline or substitute a free-form model judgment for it. Instead, it separates two questions. First: are the artifacts in the approved EDL present, correctly matched, and in the expected lifecycle state? Second: did the study generate facts or commitments that imply additional evidence beyond the template? Keeping those questions distinct lets a TMF owner see whether a gap comes from ordinary filing, template design, or study-specific operational reality.

Consider a delegation log that names eight sub-investigators while only five current CVs are filed. The missing three CVs were not guessed from a generic industry checklist. They were derived from a named source document and a countable set of people. The agent can create three proposed expectations with the individual, site, source page, responsible owner, due date, and matching fields needed when the documents arrive. A reviewer can open the delegation log, confirm the names, and accept, amend, or reject the expectations.

The same pattern applies to time-bound commitments. A sponsor oversight plan that requires a quarterly review creates an expected cadence. An IRB approval with an expiration date creates a renewal window. A site activation list paired with a protocol amendment creates site-specific signature expectations. These are not universal rules to hardcode across every sponsor. They are customer-approved reasoning patterns grounded in the study’s own documents and operating model.

Where the eTMF supports native placeholders and EDL items, accepted expectations can be represented inside the governed platform instead of living only in an external dashboard. The matching fields must align with the customer’s configured EDL model so an arriving document can resolve the correct item. When native representation is not appropriate, the product can maintain a parallel expectation register with a link back to the source system. The design decision is explicit because completeness numbers are meaningful only when the denominator is controlled.

Completeness therefore becomes a layered measure. The interface can show coverage by TMF zone, the number of open native EDL items, the number of agent-derived expectations, the items awaiting a person, and the evidence that created each expectation. Teams can filter by study, country, site, owner, due date, source rule, and confidence. Portfolio reporting can aggregate those measures without hiding the distinction between template gaps and derived gaps.

This approach also improves governance of the automation itself. Each reasoning pattern has an owner, a version, a test set, an effective date, and an allowed disposition. If a sponsor changes its CV refresh interval or oversight cadence, the rule changes under control and affected expectations can be recalculated. The system does not pretend that an LLM’s general knowledge is the sponsor’s policy. Policy remains a documented customer input.

  • Native EDL coverage and agent-derived expectations reported separately
  • Every expectation includes source, owner, due date, and matching criteria
  • Customer policy determines which implications become rules
  • Accepted expectations can land as native placeholders when the platform supports them

Document quality control

Presence is not the same as inspection readiness

The agent checks whether the filed artifact is complete, internally consistent, current, legible, and ready for its intended lifecycle state.

Synthetic eTMF document quality control dashboard with running checks and cited findings
The screen separates checks that are running from candidate checks that still require a customer rule and tolerance.

A green completeness tile can still hide weak evidence. The document may be filed under the right classification and still contain an empty signature block, a stale credential, a conflicting date, missing pages, a cropped scan, or an internal identifier that does not match the study. Document QC addresses the quality of the artifact that exists. It reads the content, checks the approved rule set, and returns a finding that states what was observed, where it was observed, why it matters, and what disposition is permitted.

The prototype demonstrates four running checks: signature completeness, date consistency, currency and expiry, and legibility or scan quality. It also shows candidate checks for study-identifier presence, pagination integrity, and duplicate filing. The word candidate is deliberate. A check should not be represented as operational until its source rule, threshold, exceptions, and expected action have been agreed. A model can recognize that a CV is more than two years old, but the sponsor determines whether that age requires replacement for the specific role and process.

Quality findings remain linked to the underlying page. An unsigned monitoring visit report cites its signature block. A protocol signature page with a wet-ink date that conflicts with a footer date preserves both date candidates rather than selecting one without authority. A low-confidence scan records the reason for the hold. This evidence makes the queue reviewable by a TMF specialist and testable by a validation team.

The check engine combines deterministic and model-based methods. Page counts, hash comparisons, lifecycle states, required fields, and date arithmetic are deterministic. Document-type recognition, visual interpretation, semantic contradiction detection, and narrative explanation may use a model. Production design should keep those boundaries visible. A deterministic failure should not be obscured behind a model score, and a model judgment should not be presented as a mathematically certain rule.

Findings can be ranked by consequence and routed accordingly. A missing signature that prevents an effective record from reaching its required state can block progression. An expiring accreditation can create a forward-looking task. A possible duplicate can remain informational until a person confirms the relationship. The product records the original finding, the reviewer decision, the remediation action, and the final state so teams can measure which checks create useful work and which produce noise.

Over time, those dispositions become the operational feedback set. False positives can be traced to a document pattern or rule version. Missed findings can be added to the labelled corpus. Performance can be reported by document type, site, language, scan quality, and check. This is the practical path from an impressive prototype to a controlled system whose limits are known.

  • Signature, date, currency, expiry, legibility, pagination, identifier, and duplicate checks
  • Page-level evidence for each finding
  • Deterministic rules separated from model judgments
  • Block, route, watch, or inform based on consequence

Human oversight

A queue for decisions, not a queue for blind approval

The interface presents the proposed action, the evidence, the consequence, and the identity that will appear in each audit trail before anything consequential occurs.

Human-in-the-loop is often used as a reassuring phrase without defining what the person actually sees or controls. In this solution, a review item carries the original document, the proposed classification or finding, extracted fields, supporting passages, cross-document reasoning, confidence or rule result, intended write, and downstream consequence. The reviewer can confirm, amend, reject, or escalate. A one-click approval without that context would reproduce the trust problem the system is meant to solve.

Identity is equally specific. API writes made with a dedicated service account are recorded against that integration user in the source platform. The product maintains its own reviewer event for the recommendation workflow, but it does not relabel the service account’s Vault action as though the human performed it. When a native lifecycle action requires a regulated electronic signature, the authorized user executes that action in Vault. This preserves the source system’s authentication, signature meaning, and audit attribution.

Permissions follow least privilege. A classification worker does not need the same rights as an administrative connector. Read scope can be limited to the study, document types, renditions, fields, and objects required for the approved use case. Write scope can begin at zero, then expand to draft metadata, placeholders, or workflow tasks only after those actions are tested. Credentials remain in the customer-approved runtime, and every connector call is logged with the study, record, operation, outcome, and correlation identifier.

The queue also exposes uncertainty. A field can have two candidates. A document can be unreadable. A rule can be inapplicable because required context is missing. A connector can be stale. A cross-system discrepancy can reflect legitimate timing rather than an error. Those states are represented directly instead of being collapsed into success or failure. The reviewer is deciding with the limits visible.

Operationally, teams can assign work by study, country, site, document type, finding class, and service-level target. The queue can send a task into the existing clinical operations process rather than demanding that every reviewer adopt a separate application. For Veeva environments, native workflow tasks and placeholders can keep regulated action in Vault while the external layer provides triage, evidence assembly, and portfolio oversight.

Review telemetry is used to improve the system, not to train on client data by default. Acceptance rate, amendment rate, rejection reason, time to decision, and recurrence can show whether a rule is valuable. Any reuse of customer content for model improvement requires a separate, explicit data agreement. The standard production posture is tenant isolation, controlled retention, and no hidden secondary use.

  • Reviewer sees the evidence and intended write before deciding
  • Integration-user actions and human signatures remain correctly attributed
  • Least-privilege scopes introduced in stages
  • Uncertainty and connector freshness are first-class states

Implementation path

Start read-only, earn the write path

A measured pilot proves the use case and the control model before the solution can modify a governed source record.

The first implementation decision is not which model to use. It is which operational question is valuable enough to test and narrow enough to evaluate. A study-level completeness pilot might focus on delegation logs, current CVs, protocol signature pages, and oversight reviews. A document-QC pilot might focus on monitoring visit reports and protocol signatures. A cross-system pilot might test one consent-version reconciliation. A bounded context produces evidence; an open-ended promise to “review the TMF” does not.

The pilot begins with a source and process inventory. The team maps the eTMF platform, Vault or non-Vault configuration, document taxonomies, EDL model, lifecycle states, roles, connected systems, interfaces, retention constraints, and validation boundary. It identifies which records may be processed, which fields are sensitive, where content can be decrypted, and whether the inference endpoint is customer-hosted, private-cloud, or a contracted external service with the required data terms.

Next comes the labelled evaluation set. Subject-matter reviewers select representative documents and record the accepted class, fields, expectations, and findings. The set includes clean native PDFs, realistic scans, common variants, edge cases, negatives, and known contradictions. Precision, recall, field accuracy, false-positive cost, false-negative cost, and abstention behavior are measured by document type and rule. A single aggregate accuracy number is insufficient for a system that takes different actions with different consequences.

The first connected run is read-only. The agent retrieves or receives the approved records, produces findings, and writes only to its own isolated evidence store. Reviewers compare the output with ordinary operations and record dispositions. This phase tests access, document rendering, extraction, rules, latency, queue design, and audit capture without changing the eTMF. It also reveals which proposed use cases merely duplicate native platform capability and which provide additional value.

Write-back is introduced one action at a time. A low-risk first write may be a proposed placeholder or a draft field on a test record. Native workflow routing follows after permissions and attribution are confirmed. Lifecycle changes that carry signatures remain with the authorized user. Each action has an idempotency key, precondition, locked source version, and explicit failure state so a retry cannot duplicate a record or write against stale content.

The pilot closes with a documented acceptance review. The package includes the intended use, architecture, data flow, risk assessment, access design, evaluation protocol and results, model and prompt configuration, deterministic controls, known limitations, SOP impacts, change-control plan, monitoring measures, rollback approach, and prioritized production backlog. The customer can then decide whether the measured outcome supports a production release, a narrower second iteration, or no deployment.

  • Bounded operational question and one-study scope
  • Representative labelled set with per-document-type metrics
  • Read-only shadow mode before any write-back
  • Action-by-action release with idempotency and source-version locks

Product boundary

Built to complement the clinical platform

Native classification, connections, lifecycles, permissions, and audit trails remain valuable. The solution concentrates on study-specific reasoning across documents and systems.

Veeva and other eTMF vendors continue to add automation. A durable product strategy therefore begins with an explicit complement boundary. Native platform services are the right place for platform-governed functions such as record lifecycles, permissions, audit trails, standard document classification, delivered connections, and system-specific workflows. Rebuilding those functions externally would add validation surface without creating a differentiated outcome.

The eTMF Intelligence Agent focuses on the reasoning that crosses an individual document or application boundary. A delegation log changes what the file should expect. An oversight plan creates a cadence. CTMS site status changes the meaning of a missing signature page. EDC consent events can contradict an otherwise correct, Final informed consent document. Those use cases require the agent to assemble evidence from multiple records, preserve each source, and explain the relationship in operational language.

Where Veeva TMF Bot or another native service has already produced a trusted classification, the agent can consume it rather than repeat the task. Where a Veeva-delivered connection already moves study, subject, or document data, the integration design can use the connected data model instead of building a duplicate transfer. Veeva’s published Clinical Operations Connections describe data and document flows across EDC, eCOA, RIM, Safety, Quality, and Clinical Operations. The solution should respect those available paths.

This boundary also keeps the product portable. The reasoning contract can define study, site, person, milestone, expected document, document version, finding, source evidence, and remediation without assuming every customer uses the same platform. A Veeva connector maps that contract to Vault objects and documents. Another eTMF connector maps it to its API or governed export. The core tests operate against the normalized evidence model while source-system actions remain connector-specific.

IntuitionLabs is a Veeva X-Pages Partner in the Veeva Partner Program. The X-Pages designation applies to Vault CRM X-Pages. It is not an eTMF certification and does not mean Veeva endorses this solution. The public prototype is intentionally labelled as a proof of concept using synthetic data. Buyers should judge the production offer by the demonstrated workflow, technical integration plan, evaluation results, validation evidence, and fit with their operating model.

  • Consume trusted native classifications when available
  • Use delivered platform connections before adding duplicate integration
  • Keep cross-document reasoning and evidence portable across eTMF platforms
  • State partnership and product-affiliation boundaries precisely

A focused eTMF intelligence solution family

Start with the core evidence model, then go deeper on the two capabilities that most clearly extend beyond a conventional document checklist.

Cross-System Reconciliation

Compare eTMF evidence with CTMS, EDC, and eConsent facts to surface contradictions a document-only review cannot see.

Explore reconciliation

Inspection Readiness

Turn completeness, quality, timeliness, expiry, and cross-system agreement into an explainable remediation plan.

Explore readiness

Veeva Clinical References

Review public Veeva Vault Clinical interface references for eTMF, CTMS, Study Startup, Site Connect, and shared workflows.

Browse references

eTMF intelligence questions

No. The solution is an intelligence and workflow layer that works with the eTMF as the system of record. In a Veeva deployment, it reads permitted document content and metadata through supported interfaces, reasons across the records, and writes back only approved classifications, metadata, placeholders, or workflow tasks. The same completeness and reconciliation patterns can be adapted to other eTMF platforms. The product is designed to complement the controls, lifecycles, permissions, and audit history already maintained by the source system.
No. The agent can classify a document, draft metadata, propose a disposition, create a placeholder, and route a native workflow task. When a lifecycle action requires an electronic signature, the authorized reviewer completes that action in the governed source system under their own identity. Agent writes are attributed to the dedicated integration user, and reviewer decisions remain attributable to the reviewer. This separation is intentional.
The agent reads operational documents for commitments and entities that imply additional evidence. A delegation log naming eight sub-investigators creates an expectation for eight current CVs. A sponsor oversight plan requiring quarterly reviews creates four expected review records per year. A site activation list creates expectations for site-specific amendment signatures. Each derived expectation preserves the source document, page or section, owner, due date, and matching criteria so the team can challenge the reasoning before it changes the live TMF.
Yes, when the customer authorizes and exposes those sources. Cross-system reconciliation is the differentiator: the agent can compare the current informed consent version filed in the eTMF with subject-level consent events in EDC, compare active-site status in CTMS with filed protocol signature pages, or compare completed remote-consent events with audit certificates in the TMF. Connector design depends on whether those systems share a Vault, use a Veeva-delivered connection, or require a vendor API.
A typical pilot starts with one study, a bounded set of document types, and a labelled evaluation set. The team agrees on the operational question, defines the expected-document rules and quality checks, configures read-only access first, measures precision and recall, and reviews every proposed result. Write-back is introduced only after the evidence, permissions, audit model, and human decision points are accepted. The result is a documented go-forward architecture and measured acceptance baseline, not a generic chatbot demo.
IntuitionLabs is a Veeva X-Pages Partner in the Veeva Partner Program. X-Pages is a Vault CRM designation. It does not certify or endorse this eTMF solution, and this solution is not a Veeva product. Veeva, Vault, and related product names are trademarks of Veeva Systems. We state that boundary directly because buyers should evaluate the product on its architecture, controls, evidence, and fit with their validated environment.
Choose one study and one question worth proving

Choose one study and one question worth proving

We will map the source records, define the evidence standard, and propose a read-only pilot with measurable acceptance criteria.

Book a Working Session

© 2026 IntuitionLabs. All rights reserved.