
eTMF Intelligence That Shows Its Evidence
Classify, test completeness, inspect document quality, reconcile clinical systems, and rank remediation while the eTMF remains the governed system of record.
01
Study first
Begin with one study and a bounded document set before portfolio rollout.
05
Connected controls
Classification, completeness, QC, reconciliation, and readiness share one evidence chain.
100%
Cited outputs
Every extracted field and finding points back to its document, page, or source record.
Human
Consequential decisions
Signatures, approvals, and policy decisions stay with authorized reviewers.
Solution overview
A reasoning layer for the trial master file
The eTMF stays the system of record. The agent turns documents, metadata, expected-document logic, and operational data into a reviewable evidence graph.

An electronic Trial Master File is more than a document repository. It is the evidence base that lets a sponsor, investigator, auditor, or inspector reconstruct how a clinical trial was planned, conducted, supervised, and documented. ICH E6(R3) describes essential records as documents, data, and relevant metadata that support trial management and allow evaluation of the methods, influencing factors, actions, and reliability of the results. That definition reaches beyond the question of whether a PDF exists in a folder.
Conventional eTMF controls are strongest when the expectation is already known. A study template defines an expected document list, filing staff select a type and subtype, lifecycle states control approval, and dashboards count filed artifacts. Those controls remain essential. The gap appears when the evidence needed for a specific study emerges from the content of other records or from activity in connected clinical systems. A delegation log can name people whose CVs should exist. A monitoring plan can define a filing window. A current consent form can conflict with the version recorded for enrolled subjects. Those relationships are difficult to express as a fixed checklist.
The IntuitionLabs eTMF Intelligence Agent adds a traceable reasoning layer across those relationships. It reads the content and metadata the customer has authorized, derives explicit expectations, tests the records against defined quality and consistency rules, and presents the result with the evidence required for review. It does not edit the source document. It does not invent an approval. It does not replace the reviewer who owns a regulated decision. Its job is to narrow the field of attention and make every recommendation inspectable.
The product is designed for sponsors, CROs, and clinical operations teams that already have an eTMF and want stronger continuous oversight. A Veeva Vault eTMF deployment can use Vault APIs, VQL, Direct Data API extracts, native placeholders, EDL items, and workflow tasks. Other eTMF platforms can expose the same logical contract through their own APIs or governed exports. The application layer remains separate so the detection logic, evaluation evidence, and remediation workflow can evolve under the customer’s change-control model.
The screens on this page come from a synthetic design prototype built around a fictional study, CRX-214. Every person, site, vendor, document, score, and finding is fabricated. Each document was generated specifically so the relevant evidence appears on the cited page. That constraint matters: a polished interface is useful only when the underlying result can be challenged against a record. The prototype therefore shows both the operational experience and the standard of evidence expected from a production implementation.
- ■Works with the eTMF as system of record
- ■Preserves document, page, field, and source attribution
- ■Separates model recommendations from governed approvals
- ■Supports platform-specific connectors without locking the reasoning model to one vendor
Primary references
Document intelligence
Classification that shows its work
A reviewer sees the document, the proposed filing location, the extracted Vault fields, the confidence threshold, and the provenance for every value.
Document classification is the first visible use case, but the product treats it as an evidence problem rather than a label-prediction problem. The agent identifies the document type, subtype, and classification, maps that result to the customer’s configured taxonomy, and extracts the metadata required for the record. A Form FDA 1572, for example, can expose the study identifier, site, investigator, country, document date, and version. Each extracted value is paired with the page, section, or field pattern that produced it so the reviewer can verify the value in context.
Confidence is a routing signal, not a claim of truth. A threshold should be set from a labelled sample that represents the customer’s real mix of native PDFs, scans, languages, templates, and document quality. Results above the accepted threshold can follow the agreed low-risk path. Results below it are held for a person. A low-confidence scan is not silently forced into the most likely class, and a conflicting date is not resolved by whichever candidate appeared first in the prompt. The unresolved state is part of the output.
The disposition also depends on the consequence of the action. Creating a draft record with proposed metadata carries a different risk from promoting a document through a lifecycle action that requires a Part 11 electronic signature. A production configuration can therefore apply different thresholds and human gates by document type, lifecycle, and field. High-volume, low-risk administrative filing can be streamlined while regulated approval remains explicit and attributable.
The same intake run can create downstream expectations. If the agent reads a delegation log and identifies eight sub-investigators, it can compare those names with current CV holdings. The initial classification task then becomes the trigger for a completeness rule. This avoids a common fragmentation problem in which one automation classifies the file, a second dashboard counts documents, and neither preserves the semantic relationship between what the document says and what the TMF should contain.
A production implementation begins with a controlled vocabulary and a representative evaluation set. The customer defines the authoritative document types, required fields, matching criteria, and permitted actions. IntuitionLabs builds deterministic checks around the model output, records the model and prompt version, and measures performance per document type. When a model, prompt, extraction schema, or threshold changes, the labelled set provides the regression evidence needed for a documented release decision.
- ■Type, subtype, classification, and TMF Reference Model mapping
- ■Field-level provenance rather than an opaque completed form
- ■Thresholds calibrated by document type and consequence
- ■Explicit hold states for low-quality scans and conflicting evidence
Expected-document intelligence
Completeness derived from the study itself
Template expectations stay intact. The agent adds a second layer for evidence implied by people, milestones, plans, approvals, and active-site facts.

An expected document list gives the study a controlled baseline. The agent does not discard that baseline or substitute a free-form model judgment for it. Instead, it separates two questions. First: are the artifacts in the approved EDL present, correctly matched, and in the expected lifecycle state? Second: did the study generate facts or commitments that imply additional evidence beyond the template? Keeping those questions distinct lets a TMF owner see whether a gap comes from ordinary filing, template design, or study-specific operational reality.
Consider a delegation log that names eight sub-investigators while only five current CVs are filed. The missing three CVs were not guessed from a generic industry checklist. They were derived from a named source document and a countable set of people. The agent can create three proposed expectations with the individual, site, source page, responsible owner, due date, and matching fields needed when the documents arrive. A reviewer can open the delegation log, confirm the names, and accept, amend, or reject the expectations.
The same pattern applies to time-bound commitments. A sponsor oversight plan that requires a quarterly review creates an expected cadence. An IRB approval with an expiration date creates a renewal window. A site activation list paired with a protocol amendment creates site-specific signature expectations. These are not universal rules to hardcode across every sponsor. They are customer-approved reasoning patterns grounded in the study’s own documents and operating model.
Where the eTMF supports native placeholders and EDL items, accepted expectations can be represented inside the governed platform instead of living only in an external dashboard. The matching fields must align with the customer’s configured EDL model so an arriving document can resolve the correct item. When native representation is not appropriate, the product can maintain a parallel expectation register with a link back to the source system. The design decision is explicit because completeness numbers are meaningful only when the denominator is controlled.
Completeness therefore becomes a layered measure. The interface can show coverage by TMF zone, the number of open native EDL items, the number of agent-derived expectations, the items awaiting a person, and the evidence that created each expectation. Teams can filter by study, country, site, owner, due date, source rule, and confidence. Portfolio reporting can aggregate those measures without hiding the distinction between template gaps and derived gaps.
This approach also improves governance of the automation itself. Each reasoning pattern has an owner, a version, a test set, an effective date, and an allowed disposition. If a sponsor changes its CV refresh interval or oversight cadence, the rule changes under control and affected expectations can be recalculated. The system does not pretend that an LLM’s general knowledge is the sponsor’s policy. Policy remains a documented customer input.
- ■Native EDL coverage and agent-derived expectations reported separately
- ■Every expectation includes source, owner, due date, and matching criteria
- ■Customer policy determines which implications become rules
- ■Accepted expectations can land as native placeholders when the platform supports them
Document quality control
Presence is not the same as inspection readiness
The agent checks whether the filed artifact is complete, internally consistent, current, legible, and ready for its intended lifecycle state.

A green completeness tile can still hide weak evidence. The document may be filed under the right classification and still contain an empty signature block, a stale credential, a conflicting date, missing pages, a cropped scan, or an internal identifier that does not match the study. Document QC addresses the quality of the artifact that exists. It reads the content, checks the approved rule set, and returns a finding that states what was observed, where it was observed, why it matters, and what disposition is permitted.
The prototype demonstrates four running checks: signature completeness, date consistency, currency and expiry, and legibility or scan quality. It also shows candidate checks for study-identifier presence, pagination integrity, and duplicate filing. The word candidate is deliberate. A check should not be represented as operational until its source rule, threshold, exceptions, and expected action have been agreed. A model can recognize that a CV is more than two years old, but the sponsor determines whether that age requires replacement for the specific role and process.
Quality findings remain linked to the underlying page. An unsigned monitoring visit report cites its signature block. A protocol signature page with a wet-ink date that conflicts with a footer date preserves both date candidates rather than selecting one without authority. A low-confidence scan records the reason for the hold. This evidence makes the queue reviewable by a TMF specialist and testable by a validation team.
The check engine combines deterministic and model-based methods. Page counts, hash comparisons, lifecycle states, required fields, and date arithmetic are deterministic. Document-type recognition, visual interpretation, semantic contradiction detection, and narrative explanation may use a model. Production design should keep those boundaries visible. A deterministic failure should not be obscured behind a model score, and a model judgment should not be presented as a mathematically certain rule.
Findings can be ranked by consequence and routed accordingly. A missing signature that prevents an effective record from reaching its required state can block progression. An expiring accreditation can create a forward-looking task. A possible duplicate can remain informational until a person confirms the relationship. The product records the original finding, the reviewer decision, the remediation action, and the final state so teams can measure which checks create useful work and which produce noise.
Over time, those dispositions become the operational feedback set. False positives can be traced to a document pattern or rule version. Missed findings can be added to the labelled corpus. Performance can be reported by document type, site, language, scan quality, and check. This is the practical path from an impressive prototype to a controlled system whose limits are known.
- ■Signature, date, currency, expiry, legibility, pagination, identifier, and duplicate checks
- ■Page-level evidence for each finding
- ■Deterministic rules separated from model judgments
- ■Block, route, watch, or inform based on consequence
Primary references
Human oversight
A queue for decisions, not a queue for blind approval
The interface presents the proposed action, the evidence, the consequence, and the identity that will appear in each audit trail before anything consequential occurs.
Human-in-the-loop is often used as a reassuring phrase without defining what the person actually sees or controls. In this solution, a review item carries the original document, the proposed classification or finding, extracted fields, supporting passages, cross-document reasoning, confidence or rule result, intended write, and downstream consequence. The reviewer can confirm, amend, reject, or escalate. A one-click approval without that context would reproduce the trust problem the system is meant to solve.
Identity is equally specific. API writes made with a dedicated service account are recorded against that integration user in the source platform. The product maintains its own reviewer event for the recommendation workflow, but it does not relabel the service account’s Vault action as though the human performed it. When a native lifecycle action requires a regulated electronic signature, the authorized user executes that action in Vault. This preserves the source system’s authentication, signature meaning, and audit attribution.
Permissions follow least privilege. A classification worker does not need the same rights as an administrative connector. Read scope can be limited to the study, document types, renditions, fields, and objects required for the approved use case. Write scope can begin at zero, then expand to draft metadata, placeholders, or workflow tasks only after those actions are tested. Credentials remain in the customer-approved runtime, and every connector call is logged with the study, record, operation, outcome, and correlation identifier.
The queue also exposes uncertainty. A field can have two candidates. A document can be unreadable. A rule can be inapplicable because required context is missing. A connector can be stale. A cross-system discrepancy can reflect legitimate timing rather than an error. Those states are represented directly instead of being collapsed into success or failure. The reviewer is deciding with the limits visible.
Operationally, teams can assign work by study, country, site, document type, finding class, and service-level target. The queue can send a task into the existing clinical operations process rather than demanding that every reviewer adopt a separate application. For Veeva environments, native workflow tasks and placeholders can keep regulated action in Vault while the external layer provides triage, evidence assembly, and portfolio oversight.
Review telemetry is used to improve the system, not to train on client data by default. Acceptance rate, amendment rate, rejection reason, time to decision, and recurrence can show whether a rule is valuable. Any reuse of customer content for model improvement requires a separate, explicit data agreement. The standard production posture is tenant isolation, controlled retention, and no hidden secondary use.
- ■Reviewer sees the evidence and intended write before deciding
- ■Integration-user actions and human signatures remain correctly attributed
- ■Least-privilege scopes introduced in stages
- ■Uncertainty and connector freshness are first-class states
Implementation path
Start read-only, earn the write path
A measured pilot proves the use case and the control model before the solution can modify a governed source record.
The first implementation decision is not which model to use. It is which operational question is valuable enough to test and narrow enough to evaluate. A study-level completeness pilot might focus on delegation logs, current CVs, protocol signature pages, and oversight reviews. A document-QC pilot might focus on monitoring visit reports and protocol signatures. A cross-system pilot might test one consent-version reconciliation. A bounded context produces evidence; an open-ended promise to “review the TMF” does not.
The pilot begins with a source and process inventory. The team maps the eTMF platform, Vault or non-Vault configuration, document taxonomies, EDL model, lifecycle states, roles, connected systems, interfaces, retention constraints, and validation boundary. It identifies which records may be processed, which fields are sensitive, where content can be decrypted, and whether the inference endpoint is customer-hosted, private-cloud, or a contracted external service with the required data terms.
Next comes the labelled evaluation set. Subject-matter reviewers select representative documents and record the accepted class, fields, expectations, and findings. The set includes clean native PDFs, realistic scans, common variants, edge cases, negatives, and known contradictions. Precision, recall, field accuracy, false-positive cost, false-negative cost, and abstention behavior are measured by document type and rule. A single aggregate accuracy number is insufficient for a system that takes different actions with different consequences.
The first connected run is read-only. The agent retrieves or receives the approved records, produces findings, and writes only to its own isolated evidence store. Reviewers compare the output with ordinary operations and record dispositions. This phase tests access, document rendering, extraction, rules, latency, queue design, and audit capture without changing the eTMF. It also reveals which proposed use cases merely duplicate native platform capability and which provide additional value.
Write-back is introduced one action at a time. A low-risk first write may be a proposed placeholder or a draft field on a test record. Native workflow routing follows after permissions and attribution are confirmed. Lifecycle changes that carry signatures remain with the authorized user. Each action has an idempotency key, precondition, locked source version, and explicit failure state so a retry cannot duplicate a record or write against stale content.
The pilot closes with a documented acceptance review. The package includes the intended use, architecture, data flow, risk assessment, access design, evaluation protocol and results, model and prompt configuration, deterministic controls, known limitations, SOP impacts, change-control plan, monitoring measures, rollback approach, and prioritized production backlog. The customer can then decide whether the measured outcome supports a production release, a narrower second iteration, or no deployment.
- ■Bounded operational question and one-study scope
- ■Representative labelled set with per-document-type metrics
- ■Read-only shadow mode before any write-back
- ■Action-by-action release with idempotency and source-version locks
Primary references
Product boundary
Built to complement the clinical platform
Native classification, connections, lifecycles, permissions, and audit trails remain valuable. The solution concentrates on study-specific reasoning across documents and systems.
Veeva and other eTMF vendors continue to add automation. A durable product strategy therefore begins with an explicit complement boundary. Native platform services are the right place for platform-governed functions such as record lifecycles, permissions, audit trails, standard document classification, delivered connections, and system-specific workflows. Rebuilding those functions externally would add validation surface without creating a differentiated outcome.
The eTMF Intelligence Agent focuses on the reasoning that crosses an individual document or application boundary. A delegation log changes what the file should expect. An oversight plan creates a cadence. CTMS site status changes the meaning of a missing signature page. EDC consent events can contradict an otherwise correct, Final informed consent document. Those use cases require the agent to assemble evidence from multiple records, preserve each source, and explain the relationship in operational language.
Where Veeva TMF Bot or another native service has already produced a trusted classification, the agent can consume it rather than repeat the task. Where a Veeva-delivered connection already moves study, subject, or document data, the integration design can use the connected data model instead of building a duplicate transfer. Veeva’s published Clinical Operations Connections describe data and document flows across EDC, eCOA, RIM, Safety, Quality, and Clinical Operations. The solution should respect those available paths.
This boundary also keeps the product portable. The reasoning contract can define study, site, person, milestone, expected document, document version, finding, source evidence, and remediation without assuming every customer uses the same platform. A Veeva connector maps that contract to Vault objects and documents. Another eTMF connector maps it to its API or governed export. The core tests operate against the normalized evidence model while source-system actions remain connector-specific.
IntuitionLabs is a Veeva X-Pages Partner in the Veeva Partner Program. The X-Pages designation applies to Vault CRM X-Pages. It is not an eTMF certification and does not mean Veeva endorses this solution. The public prototype is intentionally labelled as a proof of concept using synthetic data. Buyers should judge the production offer by the demonstrated workflow, technical integration plan, evaluation results, validation evidence, and fit with their operating model.
- ■Consume trusted native classifications when available
- ■Use delivered platform connections before adding duplicate integration
- ■Keep cross-document reasoning and evidence portable across eTMF platforms
- ■State partnership and product-affiliation boundaries precisely
Primary references
A focused eTMF intelligence solution family
Start with the core evidence model, then go deeper on the two capabilities that most clearly extend beyond a conventional document checklist.
Cross-System Reconciliation
Compare eTMF evidence with CTMS, EDC, and eConsent facts to surface contradictions a document-only review cannot see.
Explore reconciliationInspection Readiness
Turn completeness, quality, timeliness, expiry, and cross-system agreement into an explainable remediation plan.
Explore readinessVeeva Clinical References
Review public Veeva Vault Clinical interface references for eTMF, CTMS, Study Startup, Site Connect, and shared workflows.
Browse referenceseTMF intelligence questions
Choose one study and one question worth proving
We will map the source records, define the evidence standard, and propose a read-only pilot with measurable acceptance criteria.
Book a Working Session