Back

AI Framework Standardizes Medication Data for Health Analytics

At a glance

  • A peer-reviewed study on EHR medication data harmonization was published in September 2026.
  • The framework maps multiple drug identifiers to standardized codes for AI use.
  • Retrospective analyses showed high initial mapping rates but required substantial corrections.

Recent academic research has focused on improving the consistency of medication data in electronic health records (EHRs) to support managed care analytics and artificial intelligence applications.

A team from the University at Buffalo’s Department of Pharmacy Practice and Department of Industrial and Systems Engineering published a peer-reviewed article in the Journal of Managed Care & Specialty Pharmacy in September 2026. The authors, Haowen Hsu, Gabriel Gazetta, Chi-Hua Lu, and David M Jacobs, described an informatics framework designed to standardize medication data from diverse sources.

The framework integrates various medication identifiers, such as Multum Drug Synonym IDs, National Drug Codes (NDC), and RxCUI, converting them into standardized ingredient-level (RxCUI [IN]) and pharmacologic-class (ATC) representations. This process aims to facilitate the use of EHR data in AI-driven analytics and managed care scenarios.

In a retrospective analysis, the framework was applied to 214,080 discharge medication records from adults aged 65 or older hospitalized between 2020 and 2024 at Buffalo General Medical Center. The initial mapping using deterministic crosswalks reached 100%, but 30% to 35% of unique records required corrections to their ATC assignments.

What the numbers show

  • 214,080 discharge medication records were analyzed from 2020 to 2024.
  • 100% initial mapping was achieved, but 30–35% of records needed ATC corrections.
  • String-based reconciliation matched 78.5% for RxCUI [IN] and 75.5% for ATC codes, with 57.4% of assignments requiring correction.

String-based reconciliation methods produced initial match rates of 78.5% for RxCUI [IN] and 75.5% for ATC codes. However, 57.4% of these assignments required further corrections due to issues such as branded formulation omissions, lexical misclassifications, and taxonomic ambiguities.

Related work presented at ISPOR 2026 outlined an ontology-informed, layered mapping strategy using RxNorm and ATC to address the challenge of reconciling heterogeneous drug terminologies in real-world EHR data. This approach was tested on 29,123 patient medication records, with 63.9% requiring modifications, including exclusions and code corrections.

The ISPOR analysis detailed the coding distribution among medication records: 62.2% were coded as Multum, 27.3% as NDC, 9.6% as RxNorm, and 0.9% as unknown. Of 5,806 unique Multum-coded records, 3,712 required corrections, which included exclusions, RxNorm corrections, and ATC code adjustments.

These findings highlight the technical steps required to harmonize medication data for AI and managed care analytics, as well as the challenges posed by diverse coding systems and the need for ongoing data reconciliation efforts.

* This article is based on publicly available information at the time of writing.