The Analytics Imperative for Rapidly Evolving Oncology
In this first of a seven-part series, we outline the foundational challenge facing oncology analytics today: the widening gap between how fast the science evolves and how ready commercial analytics infrastructure is to follow. From biomarker-driven eligibility to post-immuno-oncology (IO) treatment sequencing to the simultaneous expansion of antibody-drug conjugates (ADCs) across indications, the science has already outpaced the analytics. The open question is whether commercial infrastructure can close the gap fast enough. This article sets the stage. The six articles that follow drill deeper into each of these challenges and the infrastructure required to solve them.
Oncology drug development is moving faster than commercial analytics can follow. New biomarkers are redefining patient eligibility. IO combinations are reshaping standard of care mid-cycle. ADCs are expanding across indications simultaneously. The science is not slowing down.
The analytics infrastructure most organizations are running assumes a simpler world, one where diagnosis codes defined patients, treatment lines were sequential, and a new approval meant one new market. That world is gone, and the analytics infrastructure built for it cannot describe the one that replaced it.
What the Data Cannot See
In modern solid tumor oncology, eligibility is molecularly defined. A single breast cancer diagnosis now splits into HER2-low, TROP-2-expressing, MSI-H, or KRAS1 G12C mutant subgroups. None of these distinctions appear cleanly in conventional claims data.
HER2-low did not exist as a commercial concept three years ago. TROP-2 has no reimbursed test code. MSI-H determination is buried in laboratory data that rarely reaches analytics pipelines. The patients most likely to benefit from biomarker-driven therapies are systematically invisible to conventional analytics. Market sizes are understated, launch performance looks weaker than it is, and health care provider (HCP) targeting misses the accounts that matter most.
The testing gap compounds this. Not all eligible patients are tested. Community labs and academic centers score the same sample by different conventions, and old tissue often stands in when a current sample would read differently. The distance between the theoretically eligible population and the analytically visible one is one of the most consequential and least discussed problems in oncology commercialization.
The testing gap is not a medical education problem. It is a data infrastructure problem.
What this means analytically:
- Proxy logic is essential. Pathology billing codes, immunohistochemistry (IHC) procedure codes, and treatment sequence signals must be layered to infer what claims alone cannot state.
- Testing penetration must be modeled. Account-level retesting rates matter as much to brand performance as prescribing data.
- Population sizing carries structural uncertainty. That uncertainty must be made explicit and regionalized, with continuous revision as testing practice evolves.
What IO Did to Treatment Sequencing
Line-of-therapy (LOT) logic was built for a sequential world. IO broke it. Consider a patient on pembrolizumab maintenance after induction: is that line one continued or line two begun? IO rechallenge, or restarting immunotherapy after a break, creates ambiguity in prior therapy tracking that no clean rule resolves. Concurrent regimens challenge the idea of a single active treatment line.
Now ADCs are entering first-line settings alongside IO backbones. An agent tracked analytically as third- or fourth-line is being prescribed at diagnosis. LOT logic inherited from that earlier context will misclassify patients and distort eligible population estimates, producing brand performance metrics that do not reflect the market.
Wrong LOT logic does not produce slightly inaccurate reports. It produces a distorted view of the market.
The post-IO LOT challenge in practice:
- IO maintenance versus new line. Distinguishing them requires clinical context, including duration, combination history, and regimen composition rather than claims sequence alone.
- Rechallenge logic. This requires longitudinal patient history across data sources, not only recent claims.
- ADC first-line combinations. Prior therapy logic built for later-line use cannot simply be extended; these combinations require new indexing frameworks.
- Clinical trial gaps. These create invisible discontinuities in commercial claims that naive LOT algorithms misread as treatment cessation.
The ADC Moment
An ADC approval in one indication is rarely the end of the story. The same molecule may be approved in one tumor type, in active launch in a second, and in Phase 3 trials across three or four more, all simultaneously, with different data maturity and different analytical requirements at each stage.
Pre-launch, the challenge is creation: there is no commercial claims history. Patient cohort construction starts from epidemiology data, clinical trial signals, and analog patterns. The business rules and data governance decisions made in this window are hard to reverse once commercial data arrives.
Post-launch, ADC competitive analytics carries complexity that standard share-of-market tools were not built to handle. A switch between two agents sharing a molecular target but differing on payload or drug-to-antibody ratio (DAR) is not equivalent to a class switch. Prior ADC exposure, toxicity history, and biomarker expression level all carry interpretive weight that conventional analytics ignores. Organizations that treat each ADC indication as a standalone project rebuild the same analytics every time. Those that build portfolio-ready infrastructure compound their advantage with every approval.
The Infrastructure Gap
The gap is architectural, not technological. Organizations have invested in platforms and tools, but the underlying data layer still encodes yesterday's oncology. That layer is the ontological framework, the LOT business rules, and the patient identification logic. Patching it incrementally produces analytics that are perpetually one cycle behind the science.
Most organizations are still trying to close the distance to the science. The ones leading now have stopped patching toward it and started building for where the science is going.
What forward-built infrastructure looks like:
- Clinical Ontology Framework. A maintained semantic layer that normalizes drug names, regimen aliases, biomarker terminology, and diagnosis codes across all data sources.
- Mechanism of Action (MOA) Taxonomy. A structured classification of the oncology drug landscape by target, mechanism, payload, and clinical context, updated before commercial data arrives.
- Post-IO LOT Engine. Business rules built for modern solid tumor treatment rather than adapted from chemotherapy-era logic, maintained as standards evolve.
- Patient Identification Framework. A multi-signal, biomarker-aware, eligibility-layered approach built to surface patients that claims data alone cannot find.
- Regulatory-Grade Data Governance. Completeness, conformance, plausibility, and currency; enforced because agents require it, not only because regulators do.
The Operating Model That Works
Technology alone does not solve this. The organizations that lead share a pattern. First, oncology-specialized people embedded permanently, who understand post-IO complexity and biomarker reality. Second, governed process, where business rules are owned and updated continuously rather than found broken at the next reporting cycle. Third, purpose-built technology running on that governed foundation.
None of the three works without the other two. Together, functioning as a data office rather than a vendor relationship, they produce infrastructure that compounds in value as the portfolio grows.
Where the World is Moving
Agentic AI is already here in commercial oncology. It shows up in agents that monitor LOT classifications as claims accumulate, agents that flag testing penetration anomalies before they distort brand reports, and agents that detect competitive patient flow in real time by territory, account, and physician. The constraint is in the data layer they run on, not the agent itself.
An agent on poorly governed LOT logic produces wrong answers at machine speed. An agent on an inconsistently defined testing proxy generates noise. The oncology data infrastructure challenge and the agentic AI readiness challenge are the same challenge. Organizations that understand this have moved from preparing for agents to building the infrastructure that makes them trustworthy, and they are doing it now.
In oncology, the science does not wait. The analytics infrastructure should not either.
In Part 2 of this series, we examine one of the most consequential operational shifts in modern oncology: how immunotherapy broke traditional line-of-therapy logic. We explore the clinical realities that challenged conventional analytics definitions: IO maintenance versus new-line distinction, rechallenge complexity, and the emergence of ADC first-line combinations. We then turn to the analytics framework required to build LOT logic that reflects the oncology market of today, not ten years ago.
1 HER2 (human epidermal growth factor receptor 2), TROP-2 (trophoblast cell-surface antigen 2), MSI-H (microsatellite instability-high), and KRAS (Kirsten rat sarcoma viral oncogene homolog) are biomarkers used to define patient eligibility for targeted oncology therapies.
FAQs
Oncology drug development is evolving faster than most commercial analytics infrastructure can follow, with biomarker-driven eligibility, IO combinations, and ADC expansions across indications outpacing systems built for simpler diagnosis-code-based models. The widening gap between scientific advancement and analytics readiness represents a critical commercial risk for pharma organizations.
The testing gap refers to the distance between the theoretically eligible biomarker-defined patient population and the population that is actually tested and analytically visible, driven by inconsistent community and academic lab practices and outdated tissue use. It is fundamentally a data infrastructure problem that directly undermines accurate launch performance measurement and commercial strategy.
Agentic AI in oncology and pharma commercial analytics is increasingly applied to layer proxy logic across pathology billing codes, IHC procedure codes, and treatment sequence signals to infer biomarker status that claims data alone cannot reveal. This approach helps model testing penetration at the account level rather than relying on assumptions that distort commercial decision-making.
Oncology commercial analytics infrastructure must evolve beyond diagnosis codes to incorporate biomarker proxy logic, account-level testing penetration modeling, and the ability to track complex post-IO treatment sequencing and multi-indication ADC expansion simultaneously. Organizations that fail to modernize this infrastructure risk misreading market size, launch performance, and targeting priorities across their oncology portfolios.
Author details
Dr. Varsha Misra
Director, Data & AI
Axtria
Insights