Claude

IntuitionLabs is now a member of the Claude Partner Network – AI training and upskilling with Claude for pharma and biotech. Book a call.

IntuitionLabs
Back to Articles
IntuitionLabs

meddra · whodrug

MedDRA & WHODrug: A Guide to Coding in Clinical Trials

November 21, 2025
Updated September 4, 2026
30 min read

Learn how MedDRA (v29.0) and WHODrug (688K+ products) standardize coding for adverse events and medications in clinical trials. Covers structure, AI-assisted coding, regulatory requirements, and 2026 updates

MedDRA & WHODrug: A Guide to Coding in Clinical Trials
Summary
  1. 01MedDRA codes adverse events and WHODrug codes medicines, creating a shared language for clinical trial safety data.
  2. 02Coded data support aggregation, safety signal detection, cross study analysis, and regulatory submissions.
  3. 03Coding quality depends on trained judgment, documented conventions, and careful management of dictionary versions.
  4. 04Automation can assist coding, but ambiguous or complex verbatim descriptions still require human oversight.
01

Executive Summary

In global clinical trials, standardized coding of adverse events (AEs) and concomitant medications is essential for patient safety monitoring, data consistency, and regulatory compliance. The two primary coding dictionaries used in these contexts are MedDRA (Medical Dictionary for Regulatory Activities) for adverse events and WHODrug for medications. MedDRA, an ICH-endorsed hierarchical medical terminology developed in the 1990s to harmonize regulatory reporting globally, is widely used in regulatory safety-reporting systems. Specific terminology requirements depend on the jurisdiction and submission context; for EU clinical trials authorised under the Clinical Trials Regulation, sponsors report SUSARs through EudraVigilance ([1]). It contains tens of thousands of specific medical terms (81,719 current Lowest Level Terms (LLTs), or 91,082 total LLTs including non-current terms, in version 29.0, released March 2026 ([2]) ([3])) organized into a five-level hierarchy (System Organ Class (SOC), High-Level Group Term (HLGT), High-Level Term (HLT), Preferred Term (PT), and Lowest Level Term (LLT) ([4])). WHODrug, maintained by the Uppsala Monitoring Centre (Uppsala WHO Collaborating Centre), is a comprehensive global dictionary of medicinal products (more than 688,000 unique product names from 175 countries as of March 2026 ([5])) that links trade names to active substances via a structured drug code and the ATC classification system ([6]) ([7]).

Employing these dictionaries ensures that free-text entries from case report forms (CRFs) are coded consistently, enabling cross-trial aggregation, safety signal detection, and regulatory submissions. For example, each reported event verbatim is mapped to a MedDRA term ([8]), while each reported medication is mapped to a WHODrug entry ([9]) ([10]). Proper coding facilitates meaningful medical grouping; MedDRA’s hierarchical and multiaxial structure allows aggregation of related events (via PTs, HLTs, etc.), while WHODrug’s classification (by ATC and Standardised Drug Groupings) allows aggregation of related medications ([11]) ([4]). Both dictionaries are regularly updated (MedDRA biannually ([12]); WHODrug frequently to include new products ([13])) and widely adopted—WHODrug is used by thousands of organizations worldwide, including pharmaceutical companies, CROs, and regulatory authorities across nearly 150 countries ([10]) ([14]).

Despite these benefits, coding consistency can be challenging. Studies show inter-coder variability in assigning MedDRA terms: one review found that 12% of codes differed between two coders and 8% deviated from the original description ([15]). MedDRA’s granularity can spread related events across multiple Preferred Terms, which can affect signal detection and should be considered in retrieval strategies ([16]), and complex or ambiguous verbatims can lead coders to select different (often related) terms ([3]) ([17]). For WHODrug, challenges include non-unique trade names across different formulations or countries, and multiple codes for one substance used in different indications ([18]). Technical advances (auto-coding tools, AI like WHODrug Koda ([19])) and standardized query tools (MedDRA SMQs) help mitigate these issues, but expert oversight remains crucial.

This guide explains the structures, coding workflows, governance, and practical use of MedDRA and WHODrug in clinical trials.

12%

Cases in which two experienced coders disagreed on the Preferred Term

8%

Codes judged inaccurate relative to expert judgment

twice yearly

MedDRA update frequency

“

MedDRA and WHODrug are **fundamental instruments** in modern clinical trial data management.

02

Introduction

Background. In clinical trials and pharmacovigilance, adverse event (AE) reporting is fundamental for monitoring patient safety and for regulatory evaluation. Coding these events into standardized terms allows aggregation of data across patients, sites, and studies. Similarly, medications taken by patients (concomitant drugs, study drugs, rescue medications) must be coded to standard vocabularies to analyze drug usage and identify potential drug–drug interactions or medication-related adverse effects. Historically, various coding schemes were used (e.g. COSTART, WHO-ART, ICD) ([20]), leading to variability and difficulty in combining data. The pharmaceutical industry and regulatory bodies recognized the need for global standards. In 1994, a collaborative effort by major regulatory agencies (FDA, EMA, PMDA, etc.) and industry created the Medical Dictionary for Regulatory Activities (MedDRA) to standardize AE coding and electronic submissions globally ([21]) ([4]).

MedDRA is now the regulatory standard for pre- and post-marketing safety data in ICH regions ([21]) ([22]). It is updated twice yearly, and includes a vast range of medical concepts (disease, symptoms, procedures, etc.) ([23]) ([24]). WHODrug evolved from the World Health Organization’s international drug monitoring program, first established post-thalidomide (1968) to facilitate global pharmacovigilance ([25]) ([20]). WHODrug links medications to active ingredients via a structured code and classifies them by the WHO Anatomical Therapeutic Chemical (ATC) system ([26]) ([27]). It is similarly updated regularly and maintained by the Uppsala Monitoring Centre ([28]).

In modern clinical trials, data on AEs and medications are captured in electronic case report forms (eCRFs) and then coded. MedDRA codes adverse events: each reported term is mapped to the most precise (lowest-level) MedDRA term ([29]) ([30]). WHODrug codes medications: each drug verbatim is matched to a WHODrug record that includes the trade name, active moiety, form, strength, country, etc. ([31]) ([10]). This dual coding ensures that safety analyses and regulatory submissions use consistent, internationally-recognized terminology across all sites and countries in a trial ([32]) ([33]).

Purpose of report. This report examines MedDRA and WHODrug in the context of coding adverse events and medications in clinical trials. We cover:

  • Historical context: How and why these dictionaries were developed; organizational stewardship (ICH/MSSO for MedDRA; WHO/UMC for WHODrug) ([21]) ([28]).
  • Structure and content: The hierarchical design of MedDRA (levels from SOC to LLT) and WHODrug (Drug code, linking of trade names to ingredients, ATC classification, Standardised Drug Groupings) ([30]) ([34]).
  • Current usage: Jurisdiction- and submission-specific terminology expectations, including EU clinical-trial SUSAR reporting through EudraVigilance and FDA study-data recommendations; adoption in industry and CROs; and integration into EDC and safety databases ([1]) ([35]).
  • Coding process: How coders use these dictionaries in practice (from CRF verbatim to coded terms), supported by guidelines (ICH points-to-consider, training materials) ([36]) ([3]).
  • Data analysis: How coded data are used for safety analysis—aggregation, signal detection, integrated listings—and how grouping structures (MedDRA’s hierarchy and SMQs; WHODrug’s classifications and SDGs) facilitate this ([37]) ([38]).
  • Challenges and inconsistencies: Empirical evidence on coding variability; potential misclassification or masking of events; difficulties with ambiguous or colloquial terms ([15]) ([39]). Practical issues like version updates (upcoding between dictionary versions) and organizational coding policies ([40]) ([41]).
  • Case studies: Examples from literature or regulatory review illustrating the impact of coding choices (e.g. antidepressant trial mis-coding, use of WHODrug to expand safety signal analysis) ([42]) ([43]).
  • Future directions: Trends such as automation (auto-coding algorithms ([44])), multilingual support (Chinese WHODrug ([45])), and harmonization (mapping between MedDRA, SNOMED, ICD) ([46]) ([4]).
  • Conclusion: Synthesis of findings to emphasize that although MedDRA and WHODrug are indispensable tools for trial safety data management, proper training, consistent practices, and ongoing improvements are needed to ensure data quality.

Throughout the report we provide extensive citations from regulatory guidance, scientific articles, and expert writings to support each claim.

03

MedDRA: Standardized Coding of Adverse Events

History and Stewardship

MedDRA (Medical Dictionary for Regulatory Activities) was created in 1994 by a joint effort of the European, Japanese and US drug regulatory authorities and industry (representatives of the ICH regions) ([21]). Its purpose was to replace disparate late-20th-century coding systems (COSTART, WHO-ART, ICD) and allow standardized electronic submission of safety data across the globe ([21]). The ICH endorsed MedDRA as the official adverse event terminology for harmonized reporting in pre- and post-marketing safety ([21]) ([47]). MedDRA’s maintenance and distribution are managed by the MedDRA Maintenance and Support Services Organization (MSSO, USA) and the Japanese Maintenance Organisation (JMO) ([48]). Both bodies ensure the dictionary remains up-to-date: new terms, modifications, and translations are considered in biannual releases ([48]) ([12]). The MSSO/JMO also issue guidance (e.g. the Term Selection Points-to-Consider) and user tools.

Structure and Content

MedDRA is structured as a five-level hierarchy ([30]) ([4]). The top level consists of 27 primary categories called System Organ Classes (SOCs) (e.g. “Cardiac disorders”, “Nervous system disorders”) ([49]) ([50]). Below SOCs are High-Level Group Terms (HLGTs) and High-Level Terms (HLTs) which form intermediate clinically-relevant clusters. At the core are Preferred Terms (PTs) — unique medical concepts each representing a specific diagnosis, sign, symptom, or process. For each PT there are one or more Lowest Level Terms (LLTs) that capture synonyms, lexical variants, or verbatim expressions ([49]) ([3]). A simplified view: LLTs map up to PT; PTs map to one or more HLTs/HLGTs; and HLTs may occur in more than one SOC through their HLGT links. A PT may be represented in more than one SOC, but each PT has one primary SOC for cumulative SOC-by-SOC outputs. ([51]) ([52]).

Example. The verbatim “upset stomach” might have a corresponding LLT “Gastrointestinal upset” which links to PT “Gastrointestinal disorder” ([53]) ([54]). That PT then appears under HLTs such as “Gastrointestinal disorders” which falls under SOC “Gastrointestinal disorders”. Meanwhile, a different PT “Dyspepsia” exists for a specific clinical entity (“upset stomach” vs “dyspepsia” may require coder judgment) but all would ultimately relate to GI disorders. Because MedDRA is multiaxial, some PTs belong to more than one SOC for retrieval (e.g. “Influenza” is in both “Infections” and “Respiratory disorders” SOCs) ([55]).

Scope. MedDRA covers not only adverse events but also: medical history, therapeutic indications, medical procedures, qualitative results (e.g. “increased”, “absent”), and other related medical concepts ([56]) ([23]). It thereby supports coding of not only AEs but also concomitant conditions and trial-related data (though in practice, most trials only enforce primary use for adverse events and sometimes medical history).

Size. MedDRA has grown substantially. In early versions (~1995) it had only a few thousand terms; by 2012 PT count was ~17,500 ([57]) ([58]). MedDRA is updated twice yearly; users should consult the release documentation for the terminology version used in an analysis. MedDRA is available in multiple languages, and English term names remain the standard for coding.

Regulatory Adoption and Use

MedDRA is widely used by regulators and sponsors, but its required use is jurisdiction- and submission-specific. For EU clinical trials authorised under the Clinical Trials Regulation, sponsors must report SUSARs to EudraVigilance; EMA identifies obtaining a MedDRA licence, where required, as part of preparing to submit those reports ([1]). In the United States, FDA’s Study Data Technical Conformance Guide describes FDA’s current thinking and expressly does not bind FDA or the public; sponsors should confirm the requirements applicable to their submission ([35]). The MSSO/JMO provide free or reduced-cost licences to regulatory agencies and non-profits, and commercial licensing to industry.

The ICH E2B standard for electronic adverse-event reporting integrates MedDRA. In EudraVigilance ICSR reporting, the reaction or event is reported using a MedDRA LLT linked to its PT, together with the MedDRA version; a PT alone is not the minimum ICSR coding requirement. Clinical-study-report tables follow their own applicable submission conventions ([59]). Guidance documents, including the MedDRA Introductory Guide and Points to Consider, address term-selection conventions for simultaneous events, laboratory terms, and signs versus diagnoses.

Table 1. Key features of MedDRA and WHODrug dictionaries summarizes some main differences:

T.01
FeatureMedDRA (Adverse Events)WHODrug (Medicines)
PurposeStandard coding of adverse events, medical history, indications, procedures ([30])Standard coding of medicinal products and active ingredients ([60])
Maintained byMSSO (US), JMO (Japan) under ICH governance ([48]) ([4])Uppsala Monitoring Centre (WHO Collaborating Centre) ([28])
Hierarchy structure5 levels: SOC > HLGT > HLT > PT > LLT ([30]) ([4])Single index with structured Drug Codes linking trade name and ingredient; adjunct ATC classification levels ([34]) ([27])
Terms coveredAll clinical medical concepts (diseases, symptoms, lab findings, etc) ([30])Medicinal products, ingredients, formulations (including conventional, herbal, vaccines, diagnostics) ([61])
Number of entriesSee the release documentation for counts for the MedDRA version used.See the release documentation for counts for the WHODrug version used.
ClassificationSOC/HLGT/HLT/PT hierarchy; multiple SOC assignments allowed (one primary) ([51])Anatomical Therapeutic Chemical (ATC) classification; Standardised Drug Groupings (SDGs) for therapeutic/chemical classes ([62]) ([63])
UpdatesBiannual (every 6 months) releases ([12])Regular (frequent) releases; enhancements annually ([64])
LanguagesAvailable in multiple languages, including English, Japanese, and Chinese ([65])English and Chinese editions ([10])
Regulatory useUsed in regulatory safety-reporting systems; EU clinical-trial SUSARs are reported through EudraVigilance, while FDA study-data guidance is nonbinding ([1]) ([35])Used globally for medicinal-product coding; requirements vary by regulator and submission type.
Cost/licenseProprietary (free to regulators, paid to industry)Proprietary (sold by UMC)
Examples of useCoding AEs from CRFs, signal detection, periodic safety reportsCoding concomitant medications, linking to safety database, defining drug classes in analysis

(Table compiled from multiple sources ([30]) ([28]) ([27]).)

Coding Process and Implementation

In a clinical trial’s workflow, data are first collected as free-text by investigators: patients report symptoms, clinicians record “AEs” or “medical history” verbatims on the case report form. Similarly, concomitant medications are recorded as drug name, sometimes with dose and route. These textual entries must then be assigned codewords from the dictionaries:

  • Adverse Event (MedDRA) Coding: A qualified, trained coder selects the current MedDRA Lowest Level Term (LLT) that most accurately reflects the reported verbatim information. If an exact match is unavailable, the coder applies medical judgment to select an existing LLT that adequately represents the concept; if no MedDRA term adequately reflects the information, the organisation should submit an MSSO change request ([66]). The coder uses the guidance to decide how to code multi-faceted or vague descriptions. EDC (electronic data capture) tools often integrate a MedDRA browser or auto-suggest function to aid this. Consistency within a trial is crucial; sponsors often train coders and document coding conventions. MedDRA users must not alter the dictionary hierarchy or term meanings; change requests should be submitted to MSSO ([66]).

  • Medication (WHODrug) Coding: The verbatim drug name (and possibly strength/form) from CRFs is mapped to an entry in WHODrug. Each WHODrug record (the WHODrug Medicinal Product Data File) includes trade name (brand or generic), active ingredient(s), form, strength, country, marketing authorization holder, etc. Each medication is assigned a unique Drug Record Number plus two sequence parts in the Drug Code (ingredient variant and trade name variant) ([67]). Concomitant drugs often vary by country and brand, so coders use the contextual info (e.g. country or ingredient) to choose the correct WHODrug entry. Many organizations use automated or semi-automated coders (search engines keyed to WHODrug’s trade/generic names) to facilitate this mapping. When coding, a sponsor might list all relevant trade names with a given active substance under one drug code for group analyses.

Once data are coded, adverse events are summarized (e.g. frequency per PT or SOC) and medications are tabulated (e.g. counts by active moiety or drug class). For pooled safety analysis, SDTM uses dictionary-derived variables appropriate to the domain and sponsor implementation. For MedDRA-coded AEs, AEDECOD is the Preferred Term; applicable hierarchy variables include AELLT, AELLTCD, AEPTCD, AESOC, and AESOCCD. For medications, CMDECOD and, where applicable, CMCLAS and CMCLASCD can carry dictionary-derived text and classification values. Dictionary name and version should be documented in submission metadata.

Use of groupings and queries:

  • For MedDRA-coded events, standardized queries (SMQs) may be applied to capture related events (e.g. “Drugs-related hepatic disorder” might include multiple PTs conceptually linked) ([3]). For example, an investigator might query all terms under the SOC “Investigations” with a laboratory abnormality.
  • For WHODrug-coded meds, the ATC classification (e.g. ATC code J01 for “Antibacterials for systemic use”) can be used to group drugs by therapeutic class. WHODrug also provides user-defined Groupings (Standardised Drug Groupings, SDGs) to classify medications by indication or property beyond ATC.

MedDRA and WHODrug versions must be specified in protocols and submissions. Up-versioning (migrating historical data to a new dictionary release) can introduce coding changes. For MedDRA, term changes (added, modified, deprecated) occur semiannually, and sponsors track these with mapping files (e.g. MedDRA Change Analysis Tool). Similarly for WHODrug, new formulations appear frequently and have new codes. Sponsors often freeze a version for a program or re-code key databases to a common version for consistency ([68]) ([41]).

Advantages of Standardized Coding

Using MedDRA and WHODrug yields numerous benefits:

  • Consistency and clarity: Different sites and languages converge on common terminology. An event described as “heart attack”, “MI”, or “myocardial infarction” will all map to a single MedDRA PT (“Myocardial infarction”), ensuring all similar events are aggregated ([69]). Medications like “Tylenol”, “Paracetamol”, and “Acetaminophen” link to the same WHODrug ingredient code.

  • Data integration: Coded data can be pooled and compared across trials/or products. This enables meta-analyses, aggregated safety summaries, and enables regulatory review of multi-site data.

  • Regulatory compliance: Major authorities require specific coding. For instance, the FDA’s Data Standards Catalog specifies MedDRA for adverse event terminology, and the PMDA requires WHODrug in submissions ([70]) ([23]).

  • Signal detection: Post-marketing surveillance relies on coded data. Spontaneous reports in VigiBase and in company safety databases use MedDRA (for events) and WHODrug (for drugs) so that data mining (statistical disproportionality, case evaluation) can be systematic. For example, WHODrug’s ingredient-centric coding allowed UMC to identify a new safety signal of panic attacks related to desogestrel by grouping all reports of desogestrel-containing products ([43]).

  • Multilingual processing: MedDRA has translations in multiple languages (e.g. Japanese, Spanish) and WHODrug is developing non-English versions (Chinese WHODrug) ([45]) ([32]), easing coding in global trials.

Challenges in Coding and Quality Issues

Despite the strengths, the coding process is imperfect and can affect data interpretation:

  • Inter-coder variability: Multiple studies document differences between coders. Toneatti et al. reported that ~12% of adverse events were coded to different PTs by two coders, and 13% were deemed “non-accurate” by adjudicators ([71]). A systematic review found the same 12% figure and noted 8% of codes deviated from source descriptions ([15]). A recent survey of Norwegian PV coders found only 36% of coders chose the “reference” code in ambiguous cases, with common errors being substitution of terms ([72]). Factors contributing include: coder training, interpretation of vague verbatim language, and the granularity of terms (coders sometimes choose a more general LLT/PT or a more specific one inconsistently) ([39]) ([72]). For example, one coder might code “joint pain” as “Arthralgia” (PT), another might split it into “Myalgia” and “Joint swelling” if context differs. Such inconsistencies can slightly alter AE counts in analysis.

  • Granularity and masking of signals: Because MedDRA terms are highly specific, an AE may scatter into multiple categories. The PLOS review noted that “with the introduction of MedDRA, it seems to have become harder to identify adverse events statistically because each code is divided in subgroups” ([73]). In practice, one often has to manually group terms or use SMQs. Conversely, if coders use very broad terms (e.g. coding “rash” vs “severe rash”), subsequent aggregation may under- or over-count certain effects.

  • Loss of nuance: MedDRA coding abstracts away narrative detail. A physician who recorded an event might omit context that could guide coding; coders must infer. For instance, an event like “chest tightness” might be coded under “Angina pectoris” or “Chest discomfort” depending on interpretation, and such choices vary ([69]). In extreme cases, miscoding can skew trial conclusions: the infamous Study 329 trial of paroxetine in adolescents initially reported only mild “emotional lability” in treatment group ([42]). Upon review it was revealed that many cases of suicidal ideation had been coded under less alarming terms (e.g. “emotionally labile”) ([42]).

  • Ambiguity and subjectivity: Verbatim descriptions may be ambiguous or layman’s terms. The Norwegian study highlighted that coders struggle with translating lay descriptions to clinical terms, and with synonyms (one coder’s “dizziness” may map to “Vertigo” PT vs another’s “Dizziness” PT) ([39]) ([72]). Lack of context (e.g. knowing whether a symptom is drug-related or new vs pre-existing) can lead to different PT selections.

  • Training and guidelines: MedDRA includes “Points to Consider” guides, but coding often still relies on the coder’s judgment. Organizations may have internal coding conventions, but differences can exist between companies or CROs. Inconsistent application of term selection rules (e.g. whether to code the lowest level vs a synonym PT) leads to divergence. A common practice is to develop and document coding conventions specific to each trial or sponsor ([41]), but this adds overhead.

  • Versioning issues: When a MedDRA version updates, the hierarchy may change (new PTs, changed SOCs). If a trial spans multiple versions, AE codes may become inconsistent. Sponsors sometimes freeze a dictionary version or re-code earlier data uniformly to one version to avoid this confusion ([68]) ([41]).

For WHODrug, analogous issues include:

  • Non-unique drug names: The same brand name may exist in different countries for different formulations. Without context (strength, country), an identical name could map to multiple WHODrug records. WHODrug includes flags or additional data to distinguish “non-unique trade names” ([34]), but coder vigilance is needed to pick the right one.

  • Multiple identifiers: One active ingredient can have many trade names and salt forms. The WHODrug structure links trade names to ingredients, but a coder must know the ingredient (or at least confirm it) to group drugs properly. E.g., “amoxicillin” vs “amoxicillin + clavulanate” have different codes due to combination ingredients ([74]). If the trial data only lists “Augmentin” (brand) on some and “amoxicillin-clavulanate” on others, misalignment can occur.

  • Indication differences: A single WHODrug code may not cover off-label uses if coded differently by authority requirements. One challenge noted is “different WHODrug identifiers may apply when a single drug is used for different indications” ([18]), because variations (dosage or formulation) can vary by indication.

  • Dictionary updates: WHODrug is updated regularly with new products. In long trials, a medication introduced mid-study may not be in the old dictionary version; testers then either use a placeholder code or recode when updating. UMC provides change analysis tools to recommend merging or splitting codes between versions ([68]). Failure to update may lead to missing coding or catch-all entries.

Data Analysis and Interpretation

The coded data become the basis for safety analysis. Key points:

  • Aggregations by MedDRA: Analysts can tabulate counts of events by SOC or preferred term ([42]). For example, frequency tables of “Headache” (PT) by treatment arm. Because MedDRA is multiaxial, analysis can include or exclude primary vs secondary SOC mapping depending on context.

  • SMQs (Standardised MedDRA Queries): The MSSO defines sets of PTs related to particular conditions (e.g. hepatic events, immune-mediated reactions) ([3]). Analysts use SMQs to capture cases that might span multiple PTs within MedDRA. For instance, an SMQ for “Cardiac Arrhythmias” may include dozens of PTs like “Atrial fibrillation”, “Ventricular tachycardia”, etc. This helps overcome the fragmentation effect. However, SMQs require careful selection and are updated with MedDRA versions ([23]) ([3]).

  • MedDRA in regulatory submissions: In clinical study reports and aggregate safety narratives, reported AEs are typically listed by PT and summarized by SOC. Graphical plots (e.g. volcano plots of AE incidence) often use MedDRA terms. Labeling sections (like the safety section of an investigator brochure or drug label) ultimately derive from MedDRA-coded data.

  • Aggregations by WHODrug: Common analyses include listing the most frequent concomitant medications by ATC class or ingredient. For example, in an oncology trial one might note that 30% of patients took any antiemetic (e.g. ATC A04). In the SDTM CM domain, CMDECOD contains standardized or dictionary-derived medication text, while CMCLAS and CMCLASCD may contain drug-class text and code; analysis-dataset variables should be specified separately. Analysis programs can group by active moiety or ATC level to see class effects (e.g. how many took any NSAID or antihypertensive).

  • Safety signal evaluation: Because WHODrug links related products (same ingredient, different brands) under one code structure, company safety databases can readily pull all ICSRs involving any formulation of a drug. As noted, searching on the active ingredient in WHODrug can reveal signals that might have been missed if searching only one brand name ([43]). This is critical in post-marketing PV.

  • SDGs (Standardised Drug Groupings): To analyze protocol deviations or interactions, sponsors often use WHODrug’s SDGs which cluster drugs by indications or pharmacology (e.g. all CYP3A4 inhibitors, all QT-prolonging antihistamines). This structured grouping can be more intuitive than raw ATC codes ([63]).

Case Studies and Real-World Examples

  • Inter-coder Variation (Toneatti 2005): In an early study by Toneatti et al. ([75]), two experienced coders independently coded 260 AE verbatim reports from a clinical trial using MedDRA. They disagreed on the PT in 12% of cases, often choosing different but related terms. A review committee judged that 8% of the codes were inaccurate relative to expert judgment. This highlights that even trained coders can diverge appreciably, especially on difficult descriptions.

  • Misleading Coding in Published Trial: The paroxetine adolescent depression trial (Study 329) famously misrepresented harms. The coding process labeled suicidality-related events under non-serious terms “emotional lability” or “behavioral problem”, concealing the true safety issue ([42]). This case underscores that coding is not purely mechanical and can be influenced by subjective decisions (or, worst-case, by bias).

  • Signal Detected via WHODrug Query: Lagerlund et al. describe a pharmacovigilance case where a suspected link between the contraceptive desogestrel and panic attacks was first noted by searching WHODrug ([43]). By querying all adverse event reports for any product containing desogestrel (regardless of brand name), analysts identified a cluster of reports coded as “panic attacks and disorders” that putatively implicated desogestrel. This would not have been easily seen if only brand-specific coding were used. The WHODrug structure enabled expanding the search by ingredient.

  • SDG Use in Trials: Baermann and Frischmann (PhUSE 2013) discussed using WHODrug SDGs to compile protocol criteria. For example, an exclusion list might include “CYP2D6 inhibitors” – an SDG could list all medications in that category for consistent capture ([76]). In practice, some sponsors map their eCRF checkboxes for exclusion to an SDG query.

  • Version-upcoding in Long Trials: Datamanagement365 (2021) noted that multi-year studies face the challenge of coding continuity. When a new MedDRA or WHODrug release comes out, some eCRF entries may remain mapped to old codes. If a trial must report in the new version, either historical data must be updated (potentially changing code frequencies, if terms were split/merged) or separate analyses must be handled. Tools like WHODrug’s Change Analysis Tool have been created to quantify these impacts ([68]).

  • Patient vs Regulator Coding (FDA Study 2018): The MedDRA-focused study among patient groups and regulators (not a trial per se) found <3% disagreement in code assignment; patients tended to pick more general PTs than regulators ([77]). This suggests that with training, even novice coders can largely match expert selections, but being aware that patient-reported events might be coded differently is important in pharmacovigilance.

Implications and Future Directions

MedDRA and WHODrug are cornerstones of clinical safety data management, but the landscape is evolving:

  • Interoperability with Electronic Health Records (EHRs): As healthcare systems adopt standards like SNOMED CT for clinical data, mapping between SNOMED and MedDRA becomes important. Richesson et al. (2009) argued that either MedDRA should adapt or robust mappings should be developed to allow re-use of clinical EHR data for AE reporting ([78]) ([79]). Current practice often involves coding AEs to MedDRA after collection, but longer-term, automatic translation from SNOMED-based diagnoses to MedDRA for safety submissions may emerge.

  • Automation and AI: The volume of data motivates use of natural language processing, and significant progress has been made. For MedDRA, companies are now deploying AI-based autocoders that use NLP and deep learning (including multi-task convolutional neural networks) to suggest MedDRA terms from verbatim text, achieving code assignments matching expert human coders 96% of the time and saving approximately 69 hours per 1,000 terms coded ([80]) ([81]). A 2026 study published in Therapeutic Innovation & Regulatory Science further validated AI-assisted MedDRA coding on both English and Chinese adverse event data from clinical trials ([82]). For WHODrug, the AI-powered “WHODrug Koda” coding engine has demonstrated strong real-world performance: compared to a simple direct-match baseline automation level of 61%, Koda increases automation to 89% while maintaining 97% coding accuracy ([83]) ([84]). WHODrug Koda is now integrated into coding platforms such as EvidentIQ Coder and Viedoc, allowing coders to review and apply AI-suggested codes directly in their workflow ([85]). However, these automated solutions still require human oversight due to the subtlety in medical language, particularly for ambiguous or complex verbatim descriptions.

  • Enhanced Dictionaries: Both dictionaries have seen significant structural and functional expansions in 2025–2026:

  • Multilingual expansion:MedDRA documentation is available in multiple languages, and WHODrug Global is available in English and Chinese editions ([10]) ([45]).

  • WHODrug structural changes (2026): The March 2026 release introduced significant structural updates: Preferred Records are now called Substance Records, the old form designation has been replaced with Archived Records, and all WHODrug files (including add-on product files) are now distributed exclusively in standard CSV format rather than TXT files ([86]). Notably, ISO Identification of Medicinal Products (ISO IDMP) data is now linked through WHODrug, facilitating regulatory interoperability ([5]).

  • ICD-10 to MedDRA mapping: MedDRA version 29.0 (March 2026) included an updated ICD-10 to MedDRA mapping release (published January 26, 2026), with new mappings for previously unmapped terms, improving cross-terminology interoperability ([32]).

  • Special queries and groupings: Both dictionaries continue to expand advanced query tools (Standardised MedDRA Queries, MedDRA Clinical Outcomes Assessments (COA) subset, etc.). WHODrug’s SDGs and custom grouping tools improve search strategies for combined medication properties ([63]).

  • Feedback-driven updates: MSSO holds meetings with users (e.g. MedDRA Blue Ribbon Panels) to incorporate emerging industry needs (e.g. new adverse event concepts, better dictionary IDs). Likewise, UMC continuously adds new medicinal products from markets—with over 5 million medicinal product records now in WHODrug from 175 countries.

  • Regulatory and Industry Trends: The globalization of trials continues to drive adoption of these standards. Key recent regulatory developments include:

  • ICH E6(R3) GCP guideline: FDA issued its final E6(R3) guidance in September 2025. The guideline incorporates quality-by-design and risk-proportionate approaches to trial design and conduct, but it does not specifically prescribe MedDRA or WHODrug coding ([87]).

  • FDA study-data standards: For FDA study-data submissions within the scope of the Study Data Technical Conformance Guide, FDA describes MedDRA as used for adverse-event coding and WHODrug Global as typically used for concomitant-medication coding. The guide contains nonbinding recommendations and uses expectation and recommendation language for dictionary versions and dataset variables; sponsors should confirm current submission-specific requirements ([88]).

  • Expanded roles: Point-of-care structured data capture using MedDRA is advancing, with eCRF systems increasingly offering MedDRA-based selectable fields for investigators, minimizing post-hoc coding while requiring user-friendly interfaces.

  • Harmonization with classifications: The ICD-10 to MedDRA mapping update in January 2026 reflects ongoing efforts to improve cross-terminology interoperability. As ICD-11 adoption grows, further mappings between ICD-11 and MedDRA are anticipated. ICH has also initiated new guidelines including E23 on real-world evidence (RWE) and M18 on biosimilar comparative efficacy, which may further expand the role of coded data.

  • Enhanced coding metrics: The research community is developing metrics of coding quality. Routine inter-coder reliability checks are becoming more common in large trial teams, supported by AI-assisted quality assurance tools that flag potential coding discrepancies.

  • Training and Governance: The literature reviewed emphasizes training: targeted programs to reduce the types of errors seen (substitution, omission) ([39]). Professional coding networks and certification (e.g. CDISC or health informatics groups) may play roles in continuing education. Clear documentation of coding decisions is recommended by ICH (compliance and reproducibility).

Overall, the trajectory points toward more integrated and automated coding, while recognizing that tool performance is specific to the evaluated data and workflow and that coding remains a critical process warranting rigor. As new therapies (e.g. gene therapies, cell therapies, complex biologics, and GLP-1 receptor agonists) and novel endpoints arise, the dictionaries continue expanding rapidly: MedDRA version 29.1 (September 2026) is the current MedDRA release, while WHODrug’s March 2026 release introduced ISO IDMP integration. Ensuring consistency in their application remains essential to the reliability of safety evaluations.

“

However, these automated solutions still require human oversight due to the subtlety in medical language, particularly for ambiguous or complex verbatim descriptions.

04

Conclusion

MedDRA and WHODrug are fundamental instruments in modern clinical trial data management. By providing a shared language for reporting adverse events and medications, they enable robust safety monitoring, cross-study comparisons, and streamlined regulatory submissions. Their widespread use by regulators, sponsors, and research organizations underscores their global importance, while specific terminology requirements remain jurisdiction- and submission-dependent.

However, extensive evidence shows that coding is not foolproof: inconsistent term selection by coders can affect data interpretation ([15]) ([39]), and the sheer granularity of these dictionaries can complicate signal detection and data retrieval ([73]) ([3]). Clinical trial organizations must therefore invest in proper coder training, use of coding tools, and quality assurance procedures (e.g. double-coding, audits, clear coding conventions ([41]) ([39])). Regulatory guidance documents (ICH points-to-consider, FDA data standards) and industry consortia (PhUSE, CDISC) continue to provide valuable best practices for term selection, query design, and data integration.

Looking forward, the synergistic use of MedDRA and WHODrug in trial analysis is likely to grow. Their complementarity – one coding what happened to the patient, the other what was given to the patient – means future analytics (such as pharmacoepidemiology within trials) will increasingly rely on their interoperability. Initiatives like embedding MedDRA/WHODrug in EHRs, enhancing cross-terminology mappings, and advancing AI-assisted coding all aim to maximize the value of coded data.

In summary, MedDRA and WHODrug have transformed the handling of safety and medication data in clinical research. They bring order and clarity to otherwise heterogeneous data. Ensuring their continued quality and relevance will require ongoing collaboration among regulators, industry, informaticians, and clinicians. With rigorous application and continuous improvement, these dictionaries will support the ultimate goal of clinical trials: safe and effective patient care.

References:

  • Lagerlund et al. (2020). WHODrug: A Global, Validated and Updated Dictionary for Medicinal Information. Ther Innov Regul Sci. 54(5):1116–1122 ([6]) ([62]).
  • Richesson et al. (2008). Heterogeneous but “Standard” Coding Systems for Adverse Events: Issues in Achieving Interoperability between Apples and Oranges. Contemp Clin Trials. 29(5):635–645 ([89]) ([49]).
  • Bennekou Schroll et al. (2012). Challenges in Coding Adverse Events in Clinical Trials: A Systematic Review. PLOS ONE 7(7):e41174 ([15]) ([30]).
  • Chan et al. (2021). The Utility of Different Data Standards to Document Adverse Drug Event Symptoms and Diagnoses: Mixed Methods Study. J Med Internet Res. 23(12):e27188 ([90]) ([91]).
  • Garmann et al. (2025). Strategies and Challenges in Coding Ambiguous Information Using MedDRA®: An Exploration Among Norwegian Pharmacovigilance Officers. Drug Saf. 48:1253–1269 ([3]) ([72]).
  • Bjerregård Madsen et al. (2022). Automated Drug Coding Using Artificial Intelligence: An Evaluation of WHODrug Koda on Adverse Event Reports. Drug Saf. 45(5):549–560 ([84]).
  • UMC. WHODrug Global—What is WHODrug Global? ([10]).
  • UMC. What's New in WHODrug (March 2026) ([5]).
  • ICH. MedDRA® Data Retrieval and Presentation: Points to Consider, Release 3.26 (based on MedDRA Version 29.0) ([92]).
  • ICH. MedDRA® Term Selection: Points to Consider, Release 4.26 (based on MedDRA Version 29.0) ([93]).
  • FDA. Data Standards Catalog; Study Data Technical Conformance Guide.
  • UMC. WHODrug Global Implementing Guide (various releases) ([94]) ([95]).
  • Additional references cited inline throughout the text (PMC and print journal articles).
The publisher

About IntuitionLabs

Build practical AI for pharma and biotech with IntuitionLabs. We help life-science teams turn complex information and workflows into useful software, governed knowledge systems and AI tools.

IntuitionLabs is an AI consulting, custom software development and data engineering firm serving pharmaceutical, biotechnology, medical-device and other life-science organizations. We work with clinical, regulatory, medical-affairs, commercial, quality and IT teams to connect technology decisions with the work people need to accomplish.

AI consulting and adoption

Our AI enablement services cover readiness assessments, use-case selection, governance and policies, team workshops, adoption measurement and ongoing advisory support. We help organizations structure the information layer behind AI: source material, context, permissions and maintained knowledge that make generated answers useful and reviewable. Private LLM inference and hosted AI options support teams evaluating how to operate AI with appropriate control over their data and infrastructure.

Software, data and life-science workflows

IntuitionLabs develops custom software for pharma and biotech, integrates enterprise systems, and builds data engineering and business intelligence solutions. Areas of focus include AI agents, regulatory research, medical writing, medical affairs, CMC information, competitive intelligence and clinical-document workflows. Our eTMF intelligence work includes cross-system reconciliation and inspection-readiness support.

Enterprise platforms and regulated delivery

We provide Veeva services, application support, managed services, integrations and custom applications, alongside enterprise content work involving platforms such as Egnyte. For regulated workflows, our services include GxP enablement, computer-system validation and software development addressing 21 CFR Part 11 requirements. The applicable controls, validation responsibilities and acceptance criteria are defined for each engagement.

Work with IntuitionLabs

Explore AI enablement, pharma and biotech software development, data engineering and BI, and Veeva services. Contact IntuitionLabs to discuss your workflow, information sources and implementation needs.

IntuitionLabs publishes educational research to help life-science teams make informed technology decisions. Coverage of a product or organization does not imply a client relationship, endorsement or partnership.

Sources / 95
Adrien Laurent

Need Expert Guidance on This Topic?

Let's discuss how IntuitionLabs can help you navigate the challenges covered in this article.

I'm Adrien Laurent, Founder & CEO of IntuitionLabs. With 25+ years of experience in enterprise software development, I specialize in creating custom AI solutions for the pharmaceutical and life science industries.

Disclaimer

The information contained in this document is provided for educational and informational purposes only. We make no representations or warranties of any kind, express or implied, about the completeness, accuracy, reliability, suitability, or availability of the information contained herein. Any reliance you place on such information is strictly at your own risk. In no event will IntuitionLabs.ai or its representatives be liable for any loss or damage including without limitation, indirect or consequential loss or damage, or any loss or damage whatsoever arising from the use of information presented in this document. This document may contain content generated with the assistance of artificial intelligence technologies. AI-generated content may contain errors, omissions, or inaccuracies. Readers are advised to independently verify any critical information before acting upon it. All product names, logos, brands, trademarks, and registered trademarks mentioned in this document are the property of their respective owners. All company, product, and service names used in this document are for identification purposes only. Use of these names, logos, trademarks, and brands does not imply endorsement by the respective trademark holders. IntuitionLabs.ai is an AI software development company specializing in helping life-science companies implement and leverage artificial intelligence solutions. Founded in 2023 by Adrien Laurent and based in San Jose, California. This document does not constitute professional or legal advice. For specific guidance related to your business needs, please consult with appropriate qualified professionals.

Related Articles

Need help with AI?

© 2026 IntuitionLabs. All rights reserved.