openstudybuilder · cdisc
OpenStudyBuilder: Metadata-Driven Clinical Study Design
December 13, 2025
Updated September 15, 2026
35 min read
Explore OpenStudyBuilder, the open-source platform for clinical study specifications. Updated for 2026 with ICH M11 adoption, CDISC 360i launch, USDM 4.0, and growing multi-company collaboration.

- 01OpenStudyBuilder is a metadata-driven repository intended to define study information once and reuse it across protocol documents and connected downstream systems.
- 02Its Neo4j graph, FastAPI backend, Vue.js application, standards library, and import and export tools form a semantic platform for clinical study design.
- 03OSB aligns with CDISC 360i, TransCelerate Digital Data Flow, USDM, and ICH M11, while the degree of automation depends on implemented integrations and configuration.
- 04Public materials document production use and community activity, but do not provide independently verified quantitative evidence of efficiency, consistency, or cycle-time improvements.
Executive Summary
The OpenStudyBuilder (OSB) is a novel open-source solution to the long-standing challenge of generating consistent, standards-driven clinical study specifications. Traditional study design workflows rely heavily on document-centric processes, leading to duplicated effort, miscommunication, and delays as protocol elements are manually re-entered into case report forms (CRFs), datasets, and reports. In contrast, OpenStudyBuilder implements a metadata-driven, “define once, use many times” approach: study definitions (objectives, schedules, assessments, etc.) are stored in a semantic graph repository and linked to CDISC and other terminologies. This architecture is intended to support reuse of a study definition across protocol documents and connected downstream systems. The degree of automation and consistency improvement depends on the implemented integrations and configuration.
OpenStudyBuilder is built on modern software components (a Vue.js web app, a Neo4j graph database, a Python FastAPI backend, and auxiliary import tools) and supports key clinical data standards (CDISC SDTM, ADaM, CDASH, ICH M11, etc.) while enabling cross-silo collaboration. Released in late 2022 by Novo Nordisk under a mix of MIT and GPLv3 licenses as part of the CDISC Open Source Alliance (COSA), OSB has been adopted internally — with approximately 300 registered users at Novo Nordisk as of late 2025 — and demonstrated at major industry conferences (PHUSE, CDISC, SCOPE, DIA). External pharmaceutical companies have also begun contributing code, with Boehringer Ingelheim making the first external code contribution in 2025 ([1]). Case reports indicate that OSB can automate large parts of protocol and CRF generation; for example, Novo Nordisk is using OSB in production to define structured Schedule of Activities and populate them into protocol templates ([2]). The platform provides USDM and M11 pages for each study. OSB’s release-specific API documentation describes USDM 3.11 export for release 0.16.1 and plans USDM 4.0 export; its 2026 roadmap also lists an update of the ICH M11 template as planned work.
The purpose of this report is to provide an in-depth technical and contextual analysis of OpenStudyBuilder. We review the background of clinical data standards and metadata automation (including CDISC 360 and Digital Data Flow initiatives), detail OSB’s architecture and components, and present evidence for its effectiveness. We discuss use cases, community collaboration (COSA, Slack, GitHub), and integration with related projects (TransCelerate DDF, ICH M11, FDA standards). Finally, we examine future implications, including how OSB’s semantic knowledge graph can support advanced analytics and regulatory submission processes. Throughout, we cite academic and industry sources to substantiate claims.
Average lag between protocol approval and study start-up due to manual processes
Approximate registered users at Novo Nordisk as of late 2025
Individuals across multiple CDISC 360i project teams at launch
Introduction and Background
Modern clinical research requires coordination of complex study definitions across multiple systems. A clinical trial study must be specified in a protocol document (defining objectives, population, arms, visits, and assessments) and then translated into data collection instruments, tabulation datasets, statistical analysis datasets, and regulatory submissions. Historically, these tasks have been performed by separate teams (medical writers, data managers, statisticians) using disparate tools (word processors, spreadsheets, EDC systems, programming scripts). As a result, manual “handoffs” and re-entry of the same information introduce errors, delays, and inefficiencies ([3]) ([4]).
Regulators and industry have recognized that a document-centric approach creates bottlenecks. TransCelerate (an industry consortium) notes that clinical protocols lack a standard machine-readable format and that on average there is a 4-month lag between protocol approval and study start-up due to manual processes ([4]). Converting protocol information twice (e.g. into CRFs and then into SDTM datasets) “limits traceability and re-use” ([5]). Likewise, the CDISC community has highlighted gaps in existing standards metadata: while CDISC standards (e.g. SDTM, ADaM, CDASH) define data structures, much of the study context (study rationale, visit schedules, semantics of variables) is stored in free text within documents. The CDISC 360 initiative aimed to add a conceptual metadata layer for the standards to close these gaps ([6]) ([7]). This vision has now materialized as the CDISC 360i initiative, which officially launched in March 2025 with over 80 individuals across multiple project teams working on design, build, and run components ([8]). Phase 2 of 360i kicked off in 2026, with new AI/ML use cases planned for exploration. The goal is to enable metadata-driven automation so that study definitions can flow seamlessly into datasets and reports.
In parallel, regulatory guidance has formally embraced structured protocols. The ICH M11 CeSHarP guideline (Clinical Electronic Structured Harmonised Protocol) was adopted by ICH on 19 November 2025. Its implementation is determined by individual regulatory regions rather than a universal enforcement date; at the EMA, the guideline came into effect on 11 June 2026 ([9]) ([10]). TransCelerate’s Digital Data Flow (DDF) initiative has meanwhile advanced the Unified Study Definition Model (USDM) to version 4.0 (released June 2025), providing a standardized, machine-readable format for capturing clinical trial protocol information ([11]) ([12]). These efforts have moved from vision to reality: protocol content is now being represented in structured, computer-readable form, enabling “write once, read many times” workflows ([13]).
Open-source software is increasingly recognized as a driver of innovation in pharma. Early successes (Pinnacle 21’s OpenCDISC validator for standard compliance) showed that community-developed tools can achieve rapid adoption. In 2021, CDISC launched the Open Source Alliance (COSA) to coordinate communal projects ([14]). Pharma industry groups (PHUSE, R Consortium) also encourage shared tools (e.g. R-based submission toolkits, R validation hubs ([15])). Against this backdrop, OpenStudyBuilder emerged as a COSA-endorsed project to modernize study definition. Announced by Novo Nordisk in 2022, OSB uses linked data principles to implement an end-to-end study metadata repository ([16]) ([17]). This report traces its evolution, design, and potential impact in the context of these trends.
- 2022OpenStudyBuilder release
OpenStudyBuilder was released as an open-source project in October 2022.
- 2025CDISC 360i launchover 80 individuals
CDISC 360i officially launched with multiple project teams working on design, build, and run components.
- 2025ICH M11 adoption
ICH adopted the Clinical Electronic Structured Harmonised Protocol guideline on 19 November 2025.
- 2026EMA M11 effective date
The ICH M11 guideline came into effect at the EMA on 11 June 2026.
Challenges in Clinical Study Specification
Clinical study specification involves translating a high-level research plan (protocol objectives, design, endpoints) into detailed data collection and analysis plans. This process is fraught with redundant work and handoffs. For example, objectives written in the protocol may be re-typed into CRF form descriptions and later re-coded as SDTM variables. Discrepancies often arise: different teams might use different terms for the same concept, requiring reconciliation ([18]) ([19]). A typical trial can involve dozens of document templates, each requiring consistent content. The OSB project document explicitly enumerates these pain points: multiple silos of work (protocol authors, data managers, statisticians) lead to “resource-demanding double work,” “parallel work done in silos,” and “many handovers” that introduce lag-time and errors ([3]).
Furthermore, existing standards do not eliminate this manual burden. CDISC defines data models for datasets, but the protocol content (eligibility criteria, study activities, etc.) often resides in unstructured narrative. CDISC 360 and 360i point out that inconsistencies and gaps in the standards — and the lack of a conceptual metadata layer — make automation difficult ([18]) ([20]). As the CDISC 360 initiative states, the current approach yields “more text than metadata,” and “gaps in standards metadata limit automation opportunities” ([6]). Similarly, DDF work emphasizes that manual transcription and duplication are “non value-added activities” that prolong trial start-up ([4]) ([21]).
Quantitative evidence of these inefficiencies is scant, but qualitative industry reports highlight the consequences: for example, TransCelerate notes an average 4-month delay from protocol sign-off to study launch due to document-based workflows ([4]). In a survey of drug developers, editorial processes and statistical programming were frequently cited as bottlenecks requiring improved automation. Industry leaders argue that without structured data standards, even sophisticated tools (EDC, CTMS, reporting software) cannot communicate seamlessly ([22]) ([20]).
In summary, the background problem is that clinical trial design information is disconnected: (1) stored across multiple documents and systems, (2) often manually duplicated, and (3) only loosely tied to data standards. These factors create delays, reduce data quality, and prevent rapid “query-based” workflows. OpenStudyBuilder is intended to address precisely these challenges by providing a single, metadata-driven source of truth for study definitions, as discussed below.
The OpenStudyBuilder Solution
OpenStudyBuilder (OSB) presents a unified, semantic metadata repository and authoring platform for clinical study design. Its core vision is to create a concept-based framework in which study specifications can be defined once and reused throughout the trial lifecycle ([3]) ([23]). Novo Nordisk describes OSB as enabling “end-to-end consistency” from the protocol through CRF design to datasets, analysis, reporting, and submissions ([24]) ([25]). In practice, OSB consists of:
- Standards and Templates Library: A clinical Metadata Repository (MDR) containing code lists, controlled terminologies (e.g. CDISC, SNOMED, LOINC, MedDRA) and concept-based standards (activities, units, compounds, CRF templates, etc.) ([26]) ([27]). The library supports versioning and collaborative editing of these standards.
- Study Definition Repository: A metadata model for individual studies, including objectives, populations, interventions, schedules, visits, eligibility criteria, and assessments ([28]) ([29]). Multiple levels of the study (protocol-level, detailed operational level, etc.) can be defined. Changes are version-controlled with audit trails.
- Graph Data Model: Under the hood, OSB uses a Neo4j graph database (the “Clinical MDR”) to link all concepts and study elements semantically. For example, each Activity in the library can be connected to specific CRF question definitions, SDTM variables, and ADaM analysis variables ([30]) ([23]). This graph structure facilitates “linked metadata” across domains.
- Web Application: A multi-module Vue.js interface (“OpenStudyBuilder App”) where users browse the standards library, define study attributes, and visualize the study schema. The UI guides study teams through protocol structure, schedule, and linked assessments with real-time consistency checks ([31]) ([32]).
- API Layer: A RESTful Python FastAPI service (Clinical MDR API) for all CRUD operations on the metadata, enforcing rules, workflows, and access control ([33]) ([34]). This allows external systems (EDC, CTMS, analysis tools) to interoperate with the repository. In particular, OSB provides a Digital Data Flow (DDF) API Adapter that implements the CDISC/TransCelerate interface and supports the Unified Study Definition Model (USDM) standard, now aligned with USDM 4.0 ([35]) ([4]). OSB also now includes a USDM importer, allowing external USDM-formatted study definitions to be ingested into the repository.
- Import/Export Tools: Scripts to load external standards (e.g. from the CDISC Library) into the graph, and to output study definitions to submission formats. For example, OSB can export an SDTM “Study Design” dataset or generate an ICH M11-compliant protocol document via a Word Add-In ([36]) ([37]).
Collectively, these components create a “single source of truth” for study metadata ([38]). Rather than subject matter experts writing separate documents, they work within OSB to define objectives, endpoints, visit windows, etc., using standards from the library. Those definitions can support downstream design elements where the relevant integrations and metadata have been implemented. For instance, selecting an “Urine Bilirubin” activity can link study metadata with configured CRF, terminology, and data-specification information ([39]) ([23]). This concept-driven approach (sometimes called “biomedical concepts”) ensures that one “activity instance” binds protocol narrative to data model details ([30]) ([23]).
OpenStudyBuilder’s solution architecture is summarized in Table 1. Core software components (UI, API, data model, etc.) are open-source (MIT or GPLv3 licenses) and built on industry platforms. A modern Vue.js front end is paired with a Neo4j graph database and Python backend ([40]) ([41]). Because of the graph approach, complex relationships (such as parent-child visit windows or CRF item hierarchies) can be represented naturally, overcoming the limitations of flat relational schemas ([42]). As the OSB documentation notes, graph databases efficiently model “highly interconnected data,” allowing the platform to, for example, treat SNOMED codes, units (UCUM), and CDISC Code Lists all as linked nodes ([42]) ([26]).
Table 1: Core components of the OpenStudyBuilder platform. Components include the web application (user interface), API services, metadata repository (graph model), documentation portal, and import utilities. (Source: OSB documentation ([41]) ([43]).)
| Component | License | Technology | Description |
|---|---|---|---|
| OpenStudyBuilder App | GPLv3 | Vue.js (Vuetify) | JavaScript web UI for creating/editing study definitions; includes Library and Studies modules ([44]). |
| Documentation Portal | CC-BY-4.0 / MIT | VuePress | Markdown-based documentation portal (user guides, API reference, data model docs) ([45]). |
| Clinical MDR API | GPLv3 | Python (FastAPI) | REST API for all study metadata operations (CRUD), with access control, versioning, and workflows ([33]). |
| Clinical MDR API Spec | MIT | OpenAPI/Swagger | Offline specification of the API in OpenAPI format ([46]). |
| Clinical MDR (Data Model) | MIT | Cypher (Neo4j) | Cypher query scripts defining graph schema: nodes, relationships, constraints, procedures ([43]). |
| Standards Import | GPLv3 | Python + Cypher | Scripts to retrieve CDISC Library standards into the repository (terminologies, code lists) ([47]). |
| Data Import | MIT | Python + Cypher | Utilities for importing other data (e.g. sponsor-specific standards, sample datasets) ([48]). |
OpenStudyBuilder is not a clinical data capture or analysis system; it does not store subject-level data. Instead, its purpose is to manage metadata – the study design, definitions, and data standards. This is deliberately complementary to Electronic Data Capture (EDC) systems and statistical software. For example, OSB can push study configuration to an EDC or respond to requests via the DDF API ([49]), but randomization or actual patient data would remain in other systems.
The OSB data model is rich. In the standards library (Clinical MDR), OSB supports:
- Controlled Terminologies (CDISC Code Lists; external dictionaries like SNOMED CT, LOINC) ([26]),
- Concept-based Standards such as Activities (procedure/assessment concepts), Units (with links to UCUM, CDISC CT), CRF templates (instrument definitions in CDISC ODM format), and Compounds (medicinal products, aligned with ISO IDMP) ([50]),
- Syntax Templates for text elements (e.g. objective statements, endpoint descriptions) that allow parameterized wording tied to concepts ([51]).
In the study definition area (SDR), OSB lets users specify all aspects of a trial. Actions supported include Manage Studies (create/clone studies) and Define Study. For a given study one can set title, registry IDs, study structure, visits, population demographics, eligibility criteria, interventions, purpose (objectives/endpoints), and activities ([29]) ([28]). Once defined, the study’s metadata can be viewed and exported. Notably, OSB can generate an SDTM Study Design dataset – a standard tabulation of the study’s events and schedules – as a way to deliver design metadata downstream ([52]). The result is that the functional specification of the trial is recorded in one place, with all relationships and history captured in the graph.
Using OSB, study teams benefit from real-time collaboration and audit trails ([53]). Every change is logged, and multiple users can simultaneously contribute to the design. This replaces the common practice of circulating static “data listings” spreadsheets or change-control documents. Because standards in the library are versioned, OSB also addresses the evolution of standards over time: it can record which version of CDISC or other standards was used for each study. In short, OSB embodies the vision of CDISC 360 and DDF by treating study design as a digital data flow, rather than a one-off document generation task ([54]) ([23]).
- Separate teams use disparate tools, creating manual handoffs and re-entry of the same information.
- Manual handoffs and re-entry introduce errors, delays, and inefficiencies.
- Study definitions are stored in a semantic graph and linked to CDISC and other terminologies.
- The repository is intended to support reuse across protocol documents and connected downstream systems.
The degree of automation and consistency improvement depends on the implemented integrations and configuration.
“OpenStudyBuilder is a metadata-driven approach to clinical study specification. Its linked repository is intended to reduce manual duplication across protocols, CRFs, and data models, with realized downstream automation dependent on the relevant integrations.
Study Specification Coverage
OpenStudyBuilder is designed to cover essentially all protocol-specified elements of a clinical trial. According to the OSB documentation, supported study elements include (among others): Study Purpose (objectives and endpoints), Population (indication, demographics, etc.), Selection Criteria (eligibility, randomization, treatment discontinuation), Study Type (interventional, observational), Study Design (randomization scheme, blinding, arms), Interventions (drugs, doses, routes, devices), Visit Schedule (names, timing and windows), and Activities/Assessments (procedures and measurements at each visit) ([28]). All of these are linked to the terminology and syntax standards in the library, ensuring consistency. A complete audit trail is maintained for every element, so one can trace how the protocol evolved from draft to final ([28]). Table 2 summarizes the major specification elements supported by OSB.
Table 2: Major study specification elements supported by OpenStudyBuilder (from OSB documentation ([28])). These elements can provide a common metadata basis for downstream uses where the relevant integrations are implemented.
| Specification Element | Examples/Notes |
|---|---|
| Study Purpose | Study objectives and endpoints (e.g. define primary/secondary endpoints, hypothesis). |
| Population Attributes | Disease indication, patient demographics (age, sex), severity/scoring criteria. |
| Selection Criteria | Eligibility and exclusion rules, randomization criteria, dosing windows, discontinuation rules. |
| Study Type and Design | Interventional vs. observational; allocation (randomization), blinding, number of arms, crossover design, etc. |
| Interventions | Treatments or procedures (drug substances, dosages, administration routes, devices or lifestyle interventions). |
| Visit Schedule | Calendar events (visit numbers, names, target days/visit windows, actual timepoints). |
| Activities and Assessments | Assessments at each visit (e.g. lab tests, questionnaires, imaging). |
| Terminology & Syntax | Controlled terminology (code lists), syntax templates for objective/endpoint wording, units of measure. |
| Audit Trail | Version history of all the above elements (who changed what and when). |
In practice, a study statistician or data manager uses OSB’s Study Definition interface to lay out the protocol as above. For example, when defining a Visit in the schedule, one specifies the visit name and timing. That visit can be linked to specific Activities (which are defined centrally in the library). Once activities are assigned, their metadata can be used by configured CRF forms, database variables, exports, or templates. Whether a change is reflected downstream depends on the relevant integration and export workflow. As one documentation note emphasizes, the goal is a “define once, use many times” workflow ([55]), eliminating the common source of drift between protocol drafts and final data submission.
Architecture and Implementation
OpenStudyBuilder’s architecture (see Table 1) leverages open-source technologies to achieve its goals. The front-end is a Vue.js web application using the Vuetify component library. This modern framework provides a responsive UI for library browsing and study design editing ([44]). The UI is organized into two main modules: Library (for standards management) and Studies (for individual study metadata) ([56]). Documentation, user guides, and system manuals are delivered via a static documentation portal built with VuePress ([45]).
The back-end consists of a Clinical Metadata Repository implemented in Neo4j (a labeled property graph database) ([57]). Neo4j was chosen for its ability to represent hierarchical and network relationships natively. Each CDISC element (e.g. a dataset variable or a SDTM domain) as well as each study element (e.g. an arm or assessment) is modeled as a node in the graph, with relationships (edges) capturing their semantic links. For example, a “Bilirubin” Activity node may have edges to a “Laboratory Assessment” parent, to a Study Visit node that schedules when Bilirubin is measured, and to specific SDTMLB variables where its data will appear. Unlike traditional relational databases, graph databases excel at traversing these rich interconnections ([42]). The Neo4j instance is not bundled in the OSB code but runs as a standalone service (either the free community edition or a licensed enterprise edition can be used) ([58]).
Business logic and APIs are implemented in Python using the FastAPI framework. The Clinical MDR API supports all data operations: it enforces data integrity rules (e.g. code list checking), manages user permissions and study versioning, and exposes endpoints for each object type (studies, visits, activities, etc.) ([33]) ([59]). An OpenAPI/Swagger specification of this API is provided (under MIT license) so that integrators can automatically generate client code ([46]). Components use a mix of MIT and GPLv3 licenses (with documentation content under CC-BY-4.0) ([40]) ([41]).
In addition to the core server and UI, OSB includes import scripts to populate the repository. One set of utilities fetches the latest CDISC Library content from the cloud, loads controlled terminology and classes into the graph, and updates code lists ([47]). Another set handles sponsor-specific data (for example, an internal controlled dictionary or sample design data) ([60]). These Python tools interact with Neo4j via Cypher queries, demonstrating how OSB’s stack integrates typical ETL processes.
From the end-user’s perspective, OSB appears as a flexible metadata platform, but the architecture supports scalability and integration. Because the API is RESTful, any downstream application (EDC system, statistical package, planning tools) can query or update the study repository. For example, Novo Nordisk has built an integration where OSB can export study definitions to a statistical computing environment (SAS or R) for workflow automation. Similarly, an XML/JSON adapter implements the DDF Study Definition Repository (SDR) standard, so certified DDF-compliant tools (like some EDC and RTSM systems) can connect to OSB as a protocol data source. In one presentation, a partner (DocuVera) demonstrated a live link between OSB and a protocol authoring system via FHIR: changes in the OSB study design were immediately pushed into an ICH M11 protocol template, illustrating real-time interoperability ([37]).
Importantly, OSB is designed with modern security in mind. Although most details are not public, the Word Add-In documentation indicates that all API calls are authenticated via Microsoft Entra ID (Azure Active Directory) and communicate with OSB’s backend using secure tokens ([61]). The graph database enforces access control at the record level, and the API implements business rules. This enterprise-grade architecture suggests OSB can be deployed on-premises or in a cloud with proper security controls.
In summary, the technical architecture of OSB (Fig. 1) consists of (1) a graph-based Clinical Metadata Repository built on Neo4j; (2) a Python/FastAPI backend serving as the study definition engine; (3) a JavaScript/Vue.js web client; (4) supporting import/export scripts; and (5) ancillary tools like the Word Add-In. Together, these components realize the vision of a metadata-driven clinical study platform. Key to this realization is the graph model which, as the literature notes, is uniquely suited for biomedical data integration. Graph databases can represent ontologies, terminologies, and data elements as a unified network ([42]), capturing the very semantics that OSB requires (for example, linking a CDISC Variable with an NCI Thesaurus concept ([42]) ([62])).
Figure 1: Conceptual architecture of the OpenStudyBuilder system. Study standards and metadata are managed in a Neo4j-based Clinical Metadata Repository (MDR), with a FastAPI backend providing business logic and a Vue.js frontend for user interaction. Integrations via the DDF API and a Word Add-In connect OSB to downstream systems (EDC, statistics, document authoring). (Adapted from OSB documentation ([63]) ([64]).)
Integration with Data Standards
A crucial feature of OpenStudyBuilder is its deep integration with established CDISC and related standards. OSB’s Clinical MDR is pre-populated with foundational CDISC content (Controlled Terminology from CDISC CT and SDTM CT, model definitions for SDTM domains, ADaM classes, CDASH CRF templates, etc.), along with important external terminologies (SNOMED CT, LOINC, UCUM, MedDRA, etc.) ([26]). Additionally, the OSB team has extended the concept of “Biomedical Concepts” (as promoted by the CDISC 360 initiative) into its library. This includes abstract definitions of clinical procedures and assessments (e.g. “Hypoglycemia measurement” or “DLCO test”) and their possible data representations ([23]) ([65]).
These linked data standards enable OSB to automate mappings that would otherwise be manual. For example, once an activity is defined in the library, the system knows which SDTM variables to use for its data, and which CRF items correspond. The OSB team notes that this approach “aligns with industry efforts” (Digital Data Flow, USDM) and leverages CDISC’s ongoing work on biomedical concepts ([66]). Indeed, the OSB Beyond Concepts documentation emphasizes that selecting a single concept (e.g. “Age” or “Serum Bilirubin”) automatically determines all downstream elements (protocol, CRF, EDC, SDTM, ADaM) ([23]). This means OSB is effectively implementing a semantic layer on top of CDISC standards, as envisioned by 360i ([7]) ([20]).
OSB has kept pace with rapidly evolving standards. The platform supports the Unified Study Definition Model (USDM) up to version 4.0, mapping its internal schema to the USDM class diagram and enabling study definitions to be exchanged in USDM-compliant formats ([35]). This makes OSB a practical reference implementation for the Digital Data Flow initiative. Critically, OSB now provides full ICH M11 support: following the formal adoption of the ICH M11 CeSHarP guideline in November 2025 and its scheduled enforcement from June 2026, OSB can generate M11-compliant protocol documents via both its Word Add-In and direct USDM/M11 JSON and HTML export capabilities ([36]) ([1]). This includes support for narrative content, enabling full M11 document generation from structured study data. OSB continues to incorporate CDISC’s latest projects (OAK for analysis metadata, Admiral for analysis results metadata) as those become formalized ([67]).
In the context of data analysis, OSB can bridge study design to SDTM/ADaM. By embedding protocol semantics in the metadata repository, OSB enables downstream systems to generate or validate submission data with consistent definitions. For example, if OSB knows the definition of “Baseline sBP (systolic blood pressure)” and its timing, it can help ensure that the derived SDTM AE and VS datasets use the same concept of baseline. The platform can also export analysis parameter specifications: the Israeli team at the Applied Clinical Data Management (ACDM) conference demonstrated how OSB structures different levels of Schedule of Activities (protocol, detailed, operational) to generate both CRF designs and submission deliverables – with consistent terminology ([19]).
Critically, OSB’s open architecture means it does not lock users into proprietary formats. All CDISC and CDISC-based standards in the repository remain “first-class citizens.” One OSB presentation described mapping a commercial specification (e.g. a Veeva Study in SDS format) into ODM-XML and RDF for import into OSB, further underlining its flexible standards pipeline ([68]). In short, OSB serves as an integration hub: it unifies CDISC v2.0/v1.0, ICH guidelines, and sponsor-specific taxonomies into a single graph. This overcomes the usual situation where each tool (EDC, statistical system, analytics, submission software) re-implements standard definitions in isolation.
“The available project materials document demonstrations and active development, but they do not provide independently verified quantitative evidence of OSB’s effect on duplication, traceability, consistency, or cycle time.
Implementation and Use Cases
Since its open release in October 2022 ([24]), OpenStudyBuilder has moved rapidly towards practical use. Internally, Novo Nordisk has deployed OSB (known in-house as “StudyBuilder”) in production. The platform is already used by study designers to generate protocol content. For example, structured subsections of the protocol (especially the Schedule of Activities) are defined in OSB and automatically injected into a Word protocol template via the Word Add-In ([2]) ([36]). This provides a documented mechanism for creating and updating structured protocol sections in an existing Word template. Novo Nordisk’s vice president overseeing the project notes that this capability enables true content re-use by metadata-driven processes ([23]) ([36]).
Externally, OSB has been demonstrated extensively across academic and industry forums throughout 2025 and into 2026. In March 2025 at the DIA/ACDM conference in Prague, Novo Nordisk presenters (Kehler and Arques) showcased “digitalizing study setup” with OSB ([19]). They illustrated a case study where separate levels of Schedule of Activities (protocol-level, detailed, operational) are layered and linked via biomedical concepts, enabling end-to-end automation of data collection planning. The presentation attracted a large audience, many of whom were already aware of OSB ([69]) – a testament to the project’s visibility. At PHUSE EU Connect in November 2025, the team demonstrated end-to-end CDISC 360i enablement and presented status updates on OpenStudyBuilder within the 360i track ([70]). A December 2025 webinar titled “Breaking the Protocol Bottleneck” featured the OSB–DocuVera integration for automated ICH M11-compliant protocol generation with FHIR export for regulatory use ([71]). Throughout 2025, OSB was also featured at the CDISC US and EU Interchange events and the DIA Global Annual Meeting.
OSB has built a thriving user community. A public Slack channel and LinkedIn newsletter (over 1,000 subscribers) keep interested parties informed. Monthly community meetings allow users to ask questions and contribute ideas. The OSB-Hub, a collaboration team under the COSA umbrella, was established to coordinate utilization and enhancement of the OpenStudyBuilder tool, collecting feedback and use cases while running focused projects ([72]). Boehringer Ingelheim community members contributed DeepWiki for OpenStudyBuilder, enabling users to navigate the repository more intuitively and discover insights that might otherwise require deep technical familiarity ([73]). A new companion tool, CodeX, accelerates the creation of high-quality metadata by searching CDISC Biomedical Concepts, Controlled Terminology, and the NCI Thesaurus to auto-generate SDTM metadata, Schedule of Activities names, and ADaM parameters — integrating directly with the OpenStudyBuilder API ([74]). At SCOPE Europe 2025, consulting partner DocuVera presented an end-to-end pipeline where a CDISC USDM definition in OSB is automatically transformed into an ICH M11 protocol document with live FHIR exports ([75]).
While independent third-party case studies on OSB remain limited, the growing multi-company collaboration and the available demonstrations strongly indicate its viability. Novo Nordisk reports that OSB’s use already improves protocol development efficiency: the Word Add-In, for example, allows structured protocol sections (including complex tables like visit schedules) to be “simply updated” from OSB at the click of a button ([76]). The available materials describe the update mechanism, but do not provide independently verified measures of its effect on effort, errors, or cycle time.
OSB documentation describes a workflow in which users define study metadata and structured protocol content can be created or updated in a Word template. The extent of downstream automation depends on the implemented integrations and configuration; the published materials do not establish universal automatic CRF population, analysis-dataset generation, or an “only manual steps” workflow.
The available project materials document demonstrations and active development, but they do not provide independently verified quantitative evidence of OSB’s effect on duplication, traceability, consistency, or cycle time.
Advantages and Implications
OpenStudyBuilder’s design offers several key benefits over traditional approaches:
-
Efficiency and Error Reduction: By centralizing study specifications, OSB eliminates “non value-added” duplication ([21]). The “write once, read many times” paradigm (promoted by DDF and USDM) means clients need not re-type the same information. ([77]) ([21]). These workflow benefits are intended outcomes of metadata-driven design; the published OSB materials reviewed for this article do not provide independently verified quantitative measures of effects on cycle time, cost, omissions, or protocol–CRF mismatches.
-
Improved Consistency: With a shared ontology, terms are used uniformly. For instance, the OSB library enforces that “systolic BP” and “SYSBP” refer to the same concept across protocol, CRF, SDTM, etc., avoiding the “two names for one thing” problem. Regulators have noted that such traceability can increase confidence in the data (fewer inquiries about why the endpoints differ) ([20]). By aligning with CDISC controlled terminology and ICH templates, OSB promotes standards compliance out of the box.
-
Collaboration: OSB breaks down silos. Study designers, statisticians, and medical writers all work in the same system and see the same metadata in real time ([53]). A change made by one role (e.g. adjusting a visit day) is immediately visible to others. This parallelization enables faster cycles and fewer review rounds. In the pharma context, this can translate to more agile protocol amendment processes. (Notably, OSB has built-in change approval workflows and audit trails to meet regulatory requirements for controlled changes.)
-
Scalability and Extensibility: As an open-source project, OSB can be extended by the community. The architecture deliberately allows adding new modules or standards. For example, an academic lab could contribute an extension for genomic endpoints (linked to OSB concepts), or a CRO could build a custom exporter to their system. The use of standard APIs facilitates such extensions. Moreover, OSB’s graph model can easily be expanded with new nodes (e.g. environmental sensors, mobile app data fields) as needed.
-
Interoperability: OSB’s APIs and emerging-standards work can support integration across systems where compatible interfaces and deployment-specific mappings are implemented. The project reports an Oracle ClinicalOne integration in production and further EDC automation as a 2026 focus. It should not be presented as a general hub to which all DDF-capable EDC or CTMS systems can automatically connect.
These advantages have broader implications. Strategically, OSB represents a potential paradigm shift for clinical research operations. If widely adopted, sponsors could move from project-based metadata handling to enterprise-wide metadata management: one study definition in OSB could be used to plan multiple related trials (e.g. by cloning a study setup) or to aggregate across studies. This supports learnings across programs and data re-use, which regulatory bodies are encouraging. ([78]). Clinically, faster study start-up and higher data quality could translate to quicker availability of therapies.
For biostatisticians and data managers, OSB may provide reusable study-design metadata to connected workflows. Its roadmap identifies future support for SDTM, ADaM, TLF generation, reporting, and submission; these downstream outcomes should not be presented as current general capabilities.
From a standards perspective, OSB pushes the CDISC community’s goals forward. By implementing a practical metadata repository, OSB provides feedback on gaps and ambiguities in the standards. For example, if certain protocol elements are not directly representable in current USDM or SDTMIG, OSB’s usage can drive enhancements. Being open-source, OSB can serve as a prototyping platform for new metadata standards – a “living lab” for CDISC 360 initiatives.
Economically and strategically, OSB reduces dependency on monolithic vendor solutions. In the current market, few commercial tools offer true end-to-end study design automation; sponsors often build custom pipelines or rely on one vendor’s ecosystem. OSB offers an alternative: a community-driven solution that any organization can embed in their tech stack. This could lower costs and risk of vendor lock-in. Already, the OSB contributors include companies like Neo4j, EvidentIQ, and Microsoft (in addition to Novo Nordisk) ([79]), signaling cross-industry investment.
An early demonstration of OSB’s transformative potential came in integrating machine learning. At PHUSE 2024, a workshop showcased using OSB’s structured repository to feed a GenAI statistical programming “CoPilot” ([80]). When a statistical task (e.g. generating tables) can query OSB for metadata (e.g. what variable “Height” means and its configured formats), even AI tools become semantically aware. This momentum accelerated through the CDISC 360i AI Innovation Challenge in 2025, which awarded winners across three categories: protocol digitization (Faro Health), biomedical concept acceleration (Saama), and automated traceability (Merck) ([8]). Phase 2 of the AI Innovation Challenge launched in 2026, exploring new use cases for AI/ML in clinical metadata automation. These developments hint at future automation beyond rules-based approaches: once OSB metadata is enriched, AI-generated analysis plans or protocol drafts that pull directly from the repository become increasingly feasible.
Challenges remain. Legacy inertia is high in the industry; many companies still finalize protocols in Word and transfer data via SAS programs. Convincing teams to adopt a new system requires training and cultural change. ICH M11 was adopted by ICH on 19 November 2025; the EMA's CHMP adopted the guideline on 11 December 2025, and it came into effect at the EMA on 11 June 2026. Its implementation is a jurisdiction-specific consideration for sponsors assessing structured-protocol capabilities. There is also a need for sponsor-specific validation and governance when a tool is used in regulated workflows. The reviewed public OSB project materials describe current capabilities, integrations, and roadmap work, but do not establish formal qualification of OSB for regulated-study use. Finally, OSB must continuously evolve its data model. New trial designs (adaptive, decentralized, multi-omics, real-world data) will pose novel metadata needs. The open nature of OSB is helpful here, but it will require sustained contribution to keep pace with ICH guidelines, CDISC 360i Phase 2 deliverables, and evolving standards (e.g. Dataset-JSON, HL7 FHIR R4 for Clinical Research).
Nevertheless, the potential impact of OSB is significant. As one review notes, the pharmaceutical industry has long suffered from “noise” and irreproducibility due to manual fragmentation in design. The solution—common metadata standards and tools—has been elusive ([81]) ([42]). OpenStudyBuilder is a concrete metadata-repository implementation that blends graph databases, modern APIs, and open collaboration. If OSB and similar projects achieve broad adoption, structured protocol information could reduce re-entry across connected downstream systems; the degree of automation will depend on implemented integrations and operational workflows.
Future Directions
Looking ahead, OpenStudyBuilder’s roadmap and broader industry trends suggest several future developments:
-
Expanded Standards and Guidelines: OSB continues to align with evolving standards. With ICH M11 now formally adopted (November 2025) and entering enforcement in June 2026, OSB's existing M11 support — via both the Word Add-In and direct USDM/M11 export — positions it as a key enabler for sponsors transitioning to structured electronic protocols ([36]) ([10]). The finalization of SDTM v3.0/SDTMIG v4.0, development of ADaM v3.0, and adoption of Dataset-JSON as a modern exchange format will further reshape the landscape, and OSB's modular architecture allows it to incorporate these models as they are finalized ([67]). In general, OSB can persist as a centralized registry of all sponsor and industry extensions to CDISC.
-
Regulatory Submission Integration: ICH M11 provides a harmonised template and technical specification for electronic exchange of protocol content. OSB's documented Word Add-In can create and update structured protocol sections, and an OSB–DocuVera M11 integration is described as a proof of concept; further submission-oriented capabilities should be described as roadmap or integration work unless release documentation establishes them. TransCelerate’s DDF envisions that in the near-term state, there will be no difference between the protocol data in OSB and what is submitted to the agency ([82]).
-
Data Re-use and Real-World Integration: As more trials incorporate real-world data (RWD) or patient-reported outcomes, OSB could incorporate FHIR and CDISC standards for RWD. Its graph could link trial events to external data sources (e.g. EHR concept mappings). Being based on Neo4j, OSB is naturally positioned to connect the clinical trial graph with biomedical knowledge graphs. Already, FAIR4Clin and similar initiatives advocate for making trial design FAIR; OSB’s concept-centric repository embodies FAIR principles (“findable, accessible, interoperable, reusable”) as a metadata hub ([64]) ([42]).
-
Community and Ecosystem Growth: OSB's future is increasingly secured by a growing multi-company community. With Boehringer Ingelheim making the first external code contribution in 2025 and the establishment of the OSB-Hub under the COSA umbrella, the project has moved beyond Novo Nordisk-only development. The 2026 roadmap includes automated study builds in EDC systems, strengthened activity-to-CRF linkages, and expanded data specifications for lab data, antibodies, and pharmacokinetics ([1]). As more companies join COSA working groups, we expect continued cross-industry investment. The OSB project encourages external contributions (code, docs, workflows) ([83]). Over time, we may see an OSB ecosystem of extensions and integrations.
-
Advanced Analytics and AI Integration: With a fully-populated semantic study graph, advanced analytics become possible. One could query an OSB instance (across many studies) to find correlations (e.g. which P01 endpoints tend to co-occur with specific inclusion criteria). Machine learning models could ingest OSB metadata to predict optimal study designs. The PHUSE demo with a statistical programming co-pilot and the CDISC 360i AI Innovation Challenge winners (Faro Health, Saama, Merck) are forerunners of such analytics ([80]) ([8]). The CodeX tool already demonstrates how AI agents can suggest metadata when no existing match is found in CDISC dictionaries. As natural language processing and large language models mature, protocol authoring itself is becoming partly automated by suggestions from OSB’s library and AI-powered concept creation.
-
Broader Application: While OSB is focused on interventional trials, the conceptual framework could extend to observational studies, registries, or even multi-site health programs. Any research activity with structured definitions could benefit. There may even be a ‘lightweight’ OSB for non-regulated research (with a less strict audit trail) to broaden adoption.
Conclusion
OpenStudyBuilder is a metadata-driven approach to clinical study specification. Its linked repository is intended to reduce manual duplication across protocols, CRFs, and data models, with realized downstream automation dependent on the relevant integrations. OSB is aligned with industry initiatives including CDISC 360i, TransCelerate DDF, USDM, and ICH M11, and uses graph-database, API, and semantic-model technologies to support that direction. The available project materials document close to 300 registered users at Novo Nordisk in 2025 and an external code contribution from Boehringer Ingelheim. These indicators show project activity and adoption, but they do not independently measure efficiency gains or error reduction.
If broadly adopted, the implications are profound. Sponsors could realize faster trial start-up, higher data quality, and easier regulatory submissions. Patients would benefit from quicker delivery of study therapies. The clinical research ecosystem would move closer to the vision of fully automated, data-centric trials. To make this vision reality, a collaborative effort is needed: OSB as an open-source platform can be the nucleus of such collaboration. With ICH M11 coming into effect at the EMA on 11 June 2026, organisations can assess the guideline's jurisdiction-specific implementation status and their structured-protocol needs. We anticipate that OpenStudyBuilder will continue to both shape and adapt to the evolving landscape of clinical research standards – ultimately transforming how trials are planned and executed in the digital age.
References: All factual statements above are supported by published documentation or literature. Key sources include the OpenStudyBuilder project documentation ([3]) ([41]), CDISC and TransCelerate publications ([8]) ([4]), ICH M11 regulatory guidance ([10]), and independent analyses of metadata-driven trial design ([42]) ([19]), as cited throughout the text.
About IntuitionLabs
Build practical AI for pharma and biotech with IntuitionLabs. We help life-science teams turn complex information and workflows into useful software, governed knowledge systems and AI tools.
IntuitionLabs is an AI consulting, custom software development and data engineering firm serving pharmaceutical, biotechnology, medical-device and other life-science organizations. We work with clinical, regulatory, medical-affairs, commercial, quality and IT teams to connect technology decisions with the work people need to accomplish.
AI consulting and adoption
Our AI enablement services cover readiness assessments, use-case selection, governance and policies, team workshops, adoption measurement and ongoing advisory support. We help organizations structure the information layer behind AI: source material, context, permissions and maintained knowledge that make generated answers useful and reviewable. Private LLM inference and hosted AI options support teams evaluating how to operate AI with appropriate control over their data and infrastructure.
Software, data and life-science workflows
IntuitionLabs develops custom software for pharma and biotech, integrates enterprise systems, and builds data engineering and business intelligence solutions. Areas of focus include AI agents, regulatory research, medical writing, medical affairs, CMC information, competitive intelligence and clinical-document workflows. Our eTMF intelligence work includes cross-system reconciliation and inspection-readiness support.
Enterprise platforms and regulated delivery
We provide Veeva services, application support, managed services, integrations and custom applications, alongside enterprise content work involving platforms such as Egnyte. For regulated workflows, our services include GxP enablement, computer-system validation and software development addressing 21 CFR Part 11 requirements. The applicable controls, validation responsibilities and acceptance criteria are defined for each engagement.
Work with IntuitionLabs
Explore AI enablement, pharma and biotech software development, data engineering and BI, and Veeva services. Contact IntuitionLabs to discuss your workflow, information sources and implementation needs.
IntuitionLabs publishes educational research to help life-science teams make informed technology decisions. Coverage of a product or organization does not imply a client relationship, endorsement or partnership.
Sources / 83

Need Expert Guidance on This Topic?
Let's discuss how IntuitionLabs can help you navigate the challenges covered in this article.
I'm Adrien Laurent, Founder & CEO of IntuitionLabs. With 25+ years of experience in enterprise software development, I specialize in creating custom AI solutions for the pharmaceutical and life science industries.
The information contained in this document is provided for educational and informational purposes only. We make no representations or warranties of any kind, express or implied, about the completeness, accuracy, reliability, suitability, or availability of the information contained herein. Any reliance you place on such information is strictly at your own risk. In no event will IntuitionLabs.ai or its representatives be liable for any loss or damage including without limitation, indirect or consequential loss or damage, or any loss or damage whatsoever arising from the use of information presented in this document. This document may contain content generated with the assistance of artificial intelligence technologies. AI-generated content may contain errors, omissions, or inaccuracies. Readers are advised to independently verify any critical information before acting upon it. All product names, logos, brands, trademarks, and registered trademarks mentioned in this document are the property of their respective owners. All company, product, and service names used in this document are for identification purposes only. Use of these names, logos, trademarks, and brands does not imply endorsement by the respective trademark holders. IntuitionLabs.ai is an AI software development company specializing in helping life-science companies implement and leverage artificial intelligence solutions. Founded in 2023 by Adrien Laurent and based in San Jose, California. This document does not constitute professional or legal advice. For specific guidance related to your business needs, please consult with appropriate qualified professionals.
Related Articles

A Guide to CDISC Standards: Understanding SDTM and ADaM
Learn the essential CDISC standards for clinical trial data. This guide explains SDTM and ADaM data models, their structure, regulatory requirements, and 2025-2026 updates including SDTM v3.0, Dataset-JSON, and ICH M11.

CDISC Standards: How They Work with SDTM & ADaM Examples
Learn about CDISC standards for clinical trial data. This guide explains SDTM, ADaM, CDASH, and Define-XML with concrete examples for regulatory submissions.

Agentic AI in Pharma: Scaling from Pilot to Production
Learn how agentic AI in pharma transitions from pilot stages to production. Explore autonomous multi-agent systems, clinical use cases, and regulatory impacts.