top of page

Clinical AI and Inherited Statistical Bias: A Case Study in How Old Problems Become New Threats at Scale

  • Jun 7
  • 25 min read

Updated: Jun 13

Subtitle: When the Machine Tells Half the Truth: How Clinical AI Inherited Pharma's Statistical Deception and Why It Matters at the Bedside


Author: Qaisar J. Qayyum, MD Clinical Assistant Professor, Internal Medicine and Geriatric Medicine. Chief Editor, Noor Journal of Complementary and Contemporary Medicine



The same data. Relative risk reduction shows 50%. Absolute risk reduction shows reduction from 2% to 1%. One in 100 patients benefits. Which number should patients see first?

Abstract

Clinical artificial intelligence tools are increasingly embedded in physician workflows as point-of-care decision support resources. A critical and underexamined problem has emerged: these tools routinely present efficacy data using relative risk reduction (RRR) as the primary, and often sole, metric, without accompanying absolute risk reduction (ARR) or natural frequency equivalents. This framing convention, long documented in pharmaceutical marketing literature as a systematic driver of prescribing inflation, now risks institutionalization within AI-generated clinical guidance. Using long-term antiplatelet therapy in chronic atherosclerotic cardiovascular disease (ASCVD) as an anchor case study, and a comparative analysis of UpToDate Expert AI and Open Evidence platforms, this article documents the pattern, quantifies its distorting effect, and offers both a conceptual critique and a practical bedside framework for clinicians navigating AI-assisted decision-making. The patient, not an abstraction, is the one who pays the price.


Keywords: relative risk reduction, absolute risk reduction, natural frequency, clinical AI, antiplatelet therapy, ASCVD, evidence-based medicine, statistical framing, physician decision support, guideline reflexivity


1. Introduction

The promise of clinical artificial intelligence is significant: rapid synthesis of complex evidence, reduction of cognitive load at the point of care, and democratization of specialist-level knowledge across clinical settings. Tools such as UpToDate Expert AI and emerging large language model-based clinical assistants are now routinely consulted by physicians during active patient encounters, often functioning as the final, and unquestioned, reference before a prescribing decision is made.


Yet the quality of a clinical AI response is not determined solely by the accuracy of its source data. It is equally determined by how that data is framed. A response that is mathematically and technically correct but clinically misleading is not a neutral tool, it is a vector for systematic bias in clinical decision-making, operating at the scale of every physician who queries it.


The problem this article addresses is not new to medicine. For decades, pharmaceutical industry communications have preferentially reported efficacy in terms of relative risk reduction, a metric that consistently inflates the perceived benefit of an intervention without technically misrepresenting the data. What is new, and deeply concerning, is that this same framing convention now appears to have been adopted, whether by design or by default, by clinical AI platforms that physicians increasingly trust as objective arbiters of evidence.


Unlike a pharmaceutical sales representative or a journal advertisement, a clinical AI tool carries no visible commercial affiliation. It occupies the epistemic space of a trusted, disinterested consultant. This makes its framing choices more consequential than those of industry, and raises the standard of transparency to which such tools must be held.

2. Background: The Framing Problem in Evidence-Based Medicine


2.1 Relative Versus Absolute Risk — Not the Same Claim

When a therapy reduces a clinical outcome from 10% to 8%, three mathematically valid statements can be made simultaneously:


  • The therapy produces a 20% relative risk reduction (2 ÷ 10 = 20%)

  • The therapy produces a 2% absolute risk reduction (10% − 8% = 2%)

  • Out of every 1,000 patients treated, 20 fewer will experience the outcome — meaning 980 patients take the medication daily, bear its risks and costs, and derive no event-level benefit


Each statement describes identical data. Only the third makes the clinical reality immediately and intuitively visible, to the physician, and to the patient sitting across the desk.


The relative risk figure invites the inference that nearly 1 in 5 patients is meaningfully protected. The absolute and natural frequency figures force the clinician to ask: is treating 1,000 patients to protect 20 worth the cumulative cost, side effect burden, and risk of harm to the remaining 980? This is the essence of individualized, patient-centered evidence-based medicine. Relative risk framing skips this question entirely.


2.2 A Well-Documented Problem — Now Migrating Into AI

The preferential use of RRR in clinical communications has been extensively documented and criticized. Bucher, Weinbacher, and Gyr demonstrated in a seminal 1994 BMJ study that physicians shown RRR alone rated treatments as significantly more effective and expressed higher prescribing intent than those shown absolute numbers for identical underlying data.¹ Nuovo, Melnikow, and Chang, reviewing major journals over a decade, found that absolute risk figures were explicitly reported in only a minority of randomized controlled trial publications despite repeated formal calls for improvement.² Gerd Gigerenzer, German psychologist, director of the

Harding Center for Risk Literacy, and named by Switzerland's Duttweiler Institute as one of the hundred most influential thinkers in the world, spent decades demonstrating that collective statistical illiteracy is widespread not only among patients but among physicians, judges, and policymakers, and that the remedy is not better math education but better data presentation.³ His term for the solution was simple: natural frequencies, telling people that 20 out of 1,000 patients benefit, not that there is a "20% reduction."


What has not been adequately examined is the migration of this framing problem into clinical AI platforms, tools that occupy a position of perceived objectivity and are consulted precisely because clinicians trust them to present evidence free of commercial bias.

3. Case Study: Antiplatelet Therapy in Chronic ASCVD


3.1 The Query

A direct clinical query was submitted to UpToDate Expert AI regarding the rationale for long-term antiplatelet therapy in chronic atherosclerotic cardiovascular disease, in the context of secondary prevention.


3.2 The Initial AI Responses

UpToDate Expert AI responded with this headline claim:

"Antiplatelet therapy reduces subsequent vascular events by about 22% across high-risk patients with established cardiovascular disease."

This figure, drawn from the landmark Antithrombotic Trialists' Collaboration (ATC) meta-analysis⁴, is the relative odds reduction. It was presented without qualification, without the corresponding absolute risk reduction, and without natural frequency equivalents. The control-group event rate, the denominator that gives the 22% figure its entire clinical meaning, was absent.


Open Evidence responded differently but with the same fundamental problem. It opened with guideline recommendations rather than primary evidence, led with the same relative risk figure ("~20% reduction in nonfatal MI and cardiovascular mortality"), and required escalation to produce absolute numbers.


3.3 Escalating the Challenge: How Many Questions to Get Complete Data?

Only after additional direct challenges, first requesting absolute numbers, then explicitly demanding control-group event rates, did either tool produce the underlying data. This is a critical finding in itself: a busy clinician who accepts a first-response AI summary walks away with a relative risk figure and no basis for patient-level benefit-risk calculation.


The complete picture, drawn from the ATC source data both tools cited, is presented in Table 1.


Table 1. Effect of Antiplatelet Therapy on Vascular Events (Antithrombotic Trialists' Collaboration; reproduced in UpToDate, 2026)

Patient Category

Control Event Rate

Antiplatelet Event Rate

Absolute Difference

Patients Helped per 1,000 Treated

Odds Reduction

All trials (high/low risk)

13.2%

10.7%

2.5%

25 per 1,000

22±2%

Prior MI

17.0%

13.5%

3.5%

35 per 1,000

25±4%

Prior stroke / TIA

21.4%

17.8%

3.6%

36 per 1,000

22±4%

Stable angina / CAD

14.1%

9.9%

4.2%

42 per 1,000

33±9%

Unstable angina

13.3%

8.0%

5.3%

53 per 1,000

46±7%

Peripheral arterial disease

7.1%

5.8%

1.3%

13 per 1,000

23±8%

Primary prevention (low risk)

4.9%

4.5%

0.4%

4 per 1,000

10%

Treatment duration 2 years for most subgroups. "Patients helped per 1,000 treated" calculated by the AI from ATC source data.


3.4 What the Full Data Reveals — In Human Terms

Across all trials, the "22% odds reduction" resolves to an absolute benefit of 25 per 1,000 treated over 2 years. In stable angina/CAD, the most common chronic ASCVD phenotype encountered in outpatient practice, the absolute benefit is 42 patients protected per 1,000 treated over two years, which is clinically meaningful but requires careful contextualization against individual bleeding risk. In peripheral arterial disease, the odds reduction of 23% sounds comparable to other subgroups, but the absolute benefit of 13 per 1,000 tells a very different story about the magnitude of benefit.


Most strikingly, in primary prevention, a population in which aspirin prescribing was common for decades, the 10% odds reduction resolves to an absolute benefit of just 4 per 1,000 treated over two years. This is the same dataset that has appropriately driven major guideline revisions away from routine aspirin in primary prevention. The absolute benefit of 4 per 1,000 is not visible in the relative risk figure. It is only visible when the full data is presented.


3.5 A Note on Methodology: This Was a Clinic Query, Not a Literature Search

This interaction with the AI tool was conducted the way a practicing physician asks a clinical question: directly, under time pressure, expecting a complete and clinically actionable answer on the first response. A busy clinician between patients does not submit a multi-part methodologically refined query. They ask a straightforward clinical question and act on what comes back from a trusted source which is specifically developed for physicians.


What came back was a relative risk figure, presented confidently without qualification. It took more than one rounds of escalating follow-up questions, each requiring additional time and statistical sophistication that most clinicians will not invest during an active clinic session, before the complete picture emerged.


The methodology was limited by design, because the patient's physician will be too.


3.6 A Second Data Point: Open Evidence

The identical query was submitted to Open Evidence, a separate AI-powered clinical decision support tool drawing on peer-reviewed literature and major society guidelines.


Open Evidence performed one round better than UpToDate, producing absolute numbers on the second ask rather than the third, and explicitly included harm data that UpToDate required escalation to produce: "aspirin prevents approximately 15 serious vascular events per 1,000 patient-years at the cost of approximately 0.3 additional major bleeds per 1,000 patient-years." On these measures, Open Evidence demonstrated better statistical transparency.


However, both tools failed the test that matters most: the first response test. Both led with relative risk figures without qualification, requiring escalation to produce complete data. Neither tool offered natural frequencies unprompted.


The second failure was unique to Open Evidence: rather than leading with quantitative evidence, the platform led with guideline architecture — Class I recommendations, Level A designations, DAPT duration algorithms, and consensus statements — while absolute risk figures were absent from the opening response entirely.


Guidelines have a legitimate role, but the physician did not ask what guidelines recommend. The physician asked what the evidence shows. These are not the same question.

4. Why This Pattern Exists and Why It Matters


4.1 The Illusion of Objectivity — AI's Most Dangerous Quality

There is a reason pharmaceutical sales representatives no longer have unfettered access to physician offices. Decades of research confirmed what clinicians already suspected: industry-sponsored communications, however accurate in their data, are shaped by commercial intent. Physicians know this. They expect it. They have learned to read a drug company's numbers with appropriate skepticism.


Clinical AI arrives differently. It carries no brand logo, no sales quota, no free lunch. It speaks in the language of citations, meta-analyses, and evidence hierarchies. To the physician consulting it between patients, it feels less like a resource and more like a revelation, the most advanced, most current, most comprehensive way to access scientific truth available at the point of care. That is its appeal. That is also its danger, because that trust is not fully earned.


AI does not write its own rules. It does not choose its own framing. Every default, every presentation hierarchy, every decision about which number appears first and which gets buried in follow-up questions deep, these are human choices, embedded in the system's design and training data. When an AI tool consistently leads with relative risk reduction and omits absolute numbers, that is not a neutral algorithmic output. That is a design choice. Someone decided, or failed to decide otherwise, and the system executes that choice millions of times a day, in clinics around the world, with the full authority of something physicians have been told represents the future of evidence-based medicine.


This is not conspiracy. But it is not innocent either. Bias embedded in training data and default outputs is still bias, it simply operates at a scale and with a credibility that no pharmaceutical sales representative could ever achieve.


A drug company advertisement that cherry-picks statistics is immediately recognizable for what it is. Clinical AI presenting the same cherry-picked statistic is received as scientific truth. The number is identical. The trust is not.

Physicians are not consulting a salesperson. They believe they are consulting science itself. That belief is exactly what is being exploited, whether anyone intended it or not is open for debate.


4.2 Statistical Framing Is Not a Style Choice — It Is a Decision Architecture

Framing is not a presentation preference. It is a decision architecture with predictable, directional consequences. Decades of behavioral economics and medical decision-making research confirm that the same underlying data presented as relative versus absolute figures produces measurably different prescribing intentions in physicians, even experienced ones. The framing effect is not a product of physician naivety, it is a robust cognitive phenomenon that operates independently of clinical expertise or statistical training.


When a clinical AI tool defaults to relative risk reduction, it is not being neutral. It is making a framing choice with a predictable directional effect: toward prescribing, toward intervention, toward the treatment arm. The patient who experiences a preventable drug-related harm because their physician's risk perception was shaped by an incomplete number is the one who ultimately absorbs the cost of that framing choice.


4.3 Institutional Inertia and the Training Data Problem

It is worth acknowledging that relative risk-first framing in clinical summaries predates AI. It appears throughout human-authored clinical reference content, in FDA drug labeling summaries, in cardiology society guideline summaries, and in the abstracts of landmark clinical trials themselves. Clinical AI tools trained on or summarizing this literature will naturally inherit its framing conventions.


This does not excuse the practice. It explains its persistence, and it underscores why explicit corrective standards must be externally imposed rather than assumed to emerge organically from the technology. A system trained on biased inputs does not self-correct. It scales the bias.


4.4 Why Guidelines Are Not the Answer: Conflicts, Commerce, and the Illusion of Universal Truth

When clinical AI defaults to guideline recommendations instead of primary evidence, it is not retreating to safer ground. It is retreating to ground that is compromised in ways most physicians have never been required to examine closely.


Guidelines Are Not Universal — They Were Never Designed to Be

A guideline written in the United States reflects the disease burden, health system economics, payer structures, and patient demographics of the United States. Applied globally without adaptation, it can actively harm patients whose baseline risk profile is fundamentally different.


The aspirin primary prevention story is again instructive. A direct comparison of US guidelines applied to the Japanese population found that the incidence of coronary heart disease in middle-aged Japanese men is nearly four times lower than in American men, while the rate of hemorrhagic stroke is three times higher, meaning the recommended risk threshold for Japanese patients would need to be two to five times higher than the American threshold to achieve the same benefit-to-harm ratio.⁹ The same guideline that protects an American patient could harm a Japanese one.


A global modeling study examining statin prescribing thresholds across 186 countries found that a single CVD risk threshold, the kind embedded in international guidelines, is not appropriate across different populations, with optimal thresholds varying substantially by country, age, sex, and local disease burden.¹⁰


Multiple Guidelines, Same Disease, Different Answers

For virtually every major clinical question in cardiovascular medicine, multiple organizations produce competing recommendations that differ in their thresholds, their drug preferences, and their risk calculations. A direct comparison of statin eligibility for primary prevention under 2021 European Society of Cardiology guidelines versus ACC/AHA and NICE guidelines found substantial differences in which patients qualified for treatment, meaning the same patient might or might not receive a statin depending solely on which guideline their physician consulted.¹¹


The Money Behind the Guidelines

The conflict of interest problem in guideline development is extensively documented and inadequately resolved.


A systematic analysis of European Society of Cardiology guidelines across five major cardiovascular conditions found that up to 64% of studies used to support guideline recommendations were either fully or partially funded by the pharmaceutical industry.¹² The ACC/AHA guidelines have faced similar scrutiny, the 2013 cholesterol guidelines, which dramatically expanded the target population for statin therapy, were criticized at publication for the prevalence of financial relationships between committee members and pharmaceutical companies with direct commercial interests in the recommendations.¹³


A prominent panel of physicians and researchers concluded that medical societies' reliance on industry funds "inevitably creates the perception of and reality of conflicts of interest and jeopardizes public trust", recommending that professional associations aim for zero industry funding for guideline-writing activities and appoint only conflict-free physicians to guideline committees.¹⁴


The Regulator Is Not Independent Either

The deepest layer of this problem reaches the regulatory agency that approves the drugs that guidelines recommend.


The US Food and Drug Administration, the world's most influential drug regulatory body, had pharmaceutical industry user fees account for 66% of the human drugs program budget in fiscal year 2022, meaning two thirds of the budget of the agency responsible for determining which drugs are safe and effective came directly from the companies seeking that determination.¹⁵ By fiscal year 2026, industry user fees account for nearly 51% of the FDA's total program budget.¹⁶


A systematic review found that public speakers at FDA advisory committee hearings with disclosed conflicts of interest were between three and six times more likely to deliver testimony favorable to industry products compared to speakers without conflicts.¹⁷


What This Means for Clinical AI

A clinical AI tool that responds to evidence questions with guideline recommendations is not giving the physician a shortcut to truth. It is presenting the output of a system shaped, at multiple levels, by documented financial relationships, as if it were neutral scientific consensus.


The primary evidence, the actual trial numbers, the absolute event rates, the natural frequencies that show how many patients benefit and how many are harmed — exists independently of these influences. It is available. It is citable. It is what physicians need at the bedside.


Primary evidence is not without its own limitations: publication bias exists, industry-sponsored trials carry their own conflicts, and absolute risk figures are reported in fewer than one in ten published studies. But imperfect primary evidence, read critically, remains a more reliable compass than guidelines shaped by the consensus of committees whose independence is structurally compromised.


This is not an argument against guidelines. In areas where robust trial data is absent, guidelines represent the best available synthesis of expert consensus, and in medico-legal contexts, they remain the defensible standard of care. The argument is narrower: when primary evidence exists and is accessible, it should be what clinical AI offers first. We have learned to be cautious of processed food, stripped of nutrients, optimized for palatability, and shaped by commercial interests rather than health. Processed evidence is no different.

5. How to Combat the AI Algorithm: A Practical Framework for Clinicians

The solution is not to stop using clinical AI. The solution is to use it the way a skilled clinician uses any imperfect tool: with informed skepticism. By the time a physician consults a clinical AI, they have already survived medical school, residency, and years of practice navigating pharmaceutical representatives, sponsored continuing medical education, and guideline committees whose independence is, as documented above, structurally compromised.


This is the same skepticism, applied to a new kind of representative.

Because that is what an AI tool becomes when it reports benefit without harm, leads with relative risk, and answers evidence questions with guideline recommendations: a pharmaceutical representative with no badge, no disclosed conflicts, and the credibility of science itself.


The physician who recognizes this is already most of the way there. The three questions below complete the journey.


Question 1: Freedom From Spoon Feeding — Is This a Relative or Absolute Number?

Any percentage reduction stated without a baseline event rate is mathematically valid but clinically incomplete. The single diagnostic question is:


"Reduced from what to what, and for how many patients in how much time?"


Demand actual numbers from the actual data. In the ATC meta-analysis documented in this article, the headline "33% odds reduction" in stable CAD resolves to 42 patients protected out of every 1,000 treated over two years.⁴ The headline "10% odds reduction" in primary prevention resolves to just 4 patients protected out of every 1,000 treated over two years.⁴ Same confident language. Tenfold difference in absolute benefit. Only the actual numbers reveal the distinction.

If the AI cannot provide this in its first response, the response is incomplete.


We must keep in the back of our minds that under the trusting name of AI, there could be a hidden pharma rep. If a pharmaceutical sales representative had just told you the same thing your AI told you, in the same language, with the same numbers, would you prescribe on that basis alone? If the answer is no, do not prescribe on the AI's basis alone either.


Question 2: Keep It Simple — Who Benefits and Who Is Harmed, Out of How Many?

An honest broker presents both sides of every transaction. Benefit without harm is not evidence-based medicine. It is a sales pitch.


Translate every efficacy claim into two natural frequencies, benefit and harm, in the same currency, for the same 1,000 patients, over the same time horizon:

  • In secondary prevention, aspirin prevents approximately 15 serious vascular events per 1,000 patient-years at the cost of approximately 3 additional major extracranial bleeds per 1,000 patient-years, a favorable ratio for most patients.⁴,¹⁸

  • In primary prevention, aspirin prevents approximately 4 events per 1,000 patient-years while bleeding risk remains unchanged, a ratio that drove guideline reversal after decades of routine prescribing.¹⁹


These two numbers, side by side, are the entire benefit-risk conversation. A clinical AI that gives you only the first number is not giving you half the answer. It is giving you the half that most reliably produces a prescription.


A clinical AI that provides complete data only when interrogated is not a functional decision support tool.


Question 3: Is This Evidence or Is This a Guideline — And Does It Matter Here?

When an AI responds with Class I recommendations before a single absolute number, ask:


"What do the actual trial numbers show, not what the guidelines recommend?"


Guidelines have a legitimate role: where robust trial data is absent, they represent the best available expert consensus, and in medico-legal contexts, they remain the defensible standard of care. The question is not whether guidelines are useful. The question is whether they are being substituted for evidence that already exists and is accessible.


We have learned to be cautious of processed food, stripped of nutrients, optimized for palatability, and shaped by commercial interests rather than health. Processed evidence is no different. The more steps between raw trial data and the AI's response, journal publication, guideline committee, specialty society endorsement, payer modification, the more cautiously that response deserves to be received.


Ask for the compass. Not the map someone else drew from it.


Ask More Questions — It Will Not Cost You As Much Time As You Think

These three questions are not academic extras. They are the minimum your patient's safety requires. Ask them every time. Push back until you get complete answers:


  • Is this a relative or absolute number? I want absolute numbers, real patients, real events, real frequencies.

  • Who benefits and who is harmed — out of how many? I want all of it, the good, the bad, and the ugly. Benefit and harm, side by side, for the same 1,000 patients.

  • Is this evidence or is this a guideline? I want the real evidence, the actual trial numbers I can review myself. Not pre-digested consensus. Not spoon feeding.


If the AI cannot answer all three completely on the first response, it has not finished its job. Keep asking.


And for the physician who says they do not have time, these three questions take sixty seconds. The harm from skipping them can last a lifetime.


Combat the algorithm. Your patient cannot do it for themselves.

6. A Technical Action Plan for Clinical AI Developers

Treat Others As You Like To Be Treated Yourself — Patient Safety Must Take Precedence

You have read the evidence. You understand the problem. Before you read the action plan below, consider one thing: you are also a patient. One day you will sit in an examination room while a physician consults a clinical AI tool, possibly one you helped build. The information that tool surfaces in that moment will shape the decision that follows.


Build the tool you would want your physician to use.


Here is how.


Action 1: Change the Default Output Format

What: Replace relative risk as the default headline metric with natural frequencies.


How:

  • Identify every output template that generates an efficacy statement

  • Insert a mandatory natural frequency conversion after every relative risk figure

  • "22% reduction in vascular events" becomes: "25 fewer patients experience a vascular event out of every 1,000 treated over 2 years" 

  • Relative figures may remain, labeled explicitly as relative, positioned after the natural frequency


Acceptance criterion: Zero first responses that lead with a relative risk figure without an accompanying natural frequency.


Action 2: Make Harm Mandatory — Not Optional

What: Every benefit statement must be paired with a harm statement in identical format. Where validated risk calculators exist, link to them.


How:

  • Build a harm-retrieval trigger into every efficacy output template

  • Benefit and harm must appear in the same response, same format, same patient population, same time horizon

  • For aspirin in secondary prevention, the output must pair both statements: "Aspirin prevents approximately 15 serious vascular events per 1,000 patient-years at the cost of approximately 3 additional major extracranial bleeds per 1,000 patient-years" [ATT Collaboration, Lancet 2009]⁴

  • Where validated risk calculators exist for the specific clinical question, embed or link to them:

    • For antiplatelet-related bleeding risk: HAS-BLED calculator (bleeding risk in anticoagulated patients)

    • For DAPT duration decisions: DAPT Score calculator or PRECISE-DAPT (individualized bleeding risk during dual antiplatelet therapy)

    • For ACS bleeding risk: CRUSADE score (bleeding risk stratification)

  • If harm data is unavailable for the specific query, the response must explicitly state this, not silently omit it


A clinical AI that reports only benefit is functioning as a pharmaceutical representative. That is not the product you are building.


Acceptance criterion: No first response reports efficacy without a corresponding harm statement. Validated risk calculators are linked where they exist for the specific clinical scenario.


Action 3: Tag Every Response — Evidence or Guideline

What: Every clinical response must carry a visible source classification at the point of output.


How:

  • Implement mandatory source tagging in the response header or inline:

    • [Primary trial evidence — specify source and year]

    • [Guideline recommendation, specify organization, year, and evidence class]

    • [Expert consensus — no randomized trial data available]

  • Where guideline-producing organizations have documented industry relationships relevant to the recommendation, add: [Note: relevant financial conflicts of interest exist among guideline authors, see source disclosure]

  • This is a metadata problem. Your pipeline already knows the source. Surface it.


Acceptance criterion: Every response is tagged. No unattributed clinical recommendations.


Action 4: Surface Guideline Contradictions — Do Not Hide Them


What: Where multiple guidelines exist for the same clinical question and differ in their recommendations, present all of them, with their contradictions explicitly disclosed.


How:

  • When a query returns recommendations from more than one guideline-producing organization, present each separately with its source, year, and evidence classification

  • Flag contradictions explicitly at the point of output with their direct clinical implications


Real example 1 — Aspirin in primary prevention: same patient, three different answers

[ACC/AHA 2019: Aspirin may be considered for primary prevention in adults aged 40–70 at higher ASCVD risk and low bleeding risk, Class IIb. Recommended against routinely in adults over 70, Class III.]²


[ESC 2021: Aspirin not recommended for primary prevention due to unfavorable bleeding risk — Class III. Consistent with ACC/AHA on high-risk exception — Class IIb.]³


[USPSTF 2022: Aspirin not recommended for initiation in adults aged 60 and older. Individualized decision for adults aged 40–59 at 10% or greater 10-year CVD risk.]⁴


[Guideline contradiction: ACC/AHA permits aspirin consideration in adults aged 40–70 at high ASCVD risk; USPSTF restricts initiation to adults under 60. A patient aged 62 at high ASCVD risk meets ACC/AHA criteria but not USPSTF criteria for aspirin initiation. The treating physician must be aware of this discrepancy and exercise independent clinical judgment.]


Real example 2 — DAPT duration after PCI: same procedure, different standards

[ACC/AHA 2023: DAPT for 6 months post-PCI for chronic coronary disease, followed by single antiplatelet therapy — Class I, Level A.]⁵


[ESC 2023: In selected low-bleeding-risk patients post-PCI, P2Y12 inhibitor monotherapy after 1–3 months of DAPT may be considered to reduce bleeding — Class IIa, Level A.]⁶


[Guideline contradiction: ACC/AHA recommends 6 months of DAPT as standard; ESC permits discontinuation of aspirin after as little as 1 month in selected patients. The same post-PCI patient may receive meaningfully different therapy depending solely on which guideline their physician consulted.]


On the objection that surfacing multiple guidelines makes responses exhaustively long:

Guideline contradictions are not created by surfacing them, they already exist. A physician unaware of the contradiction between ACC/AHA and USPSTF aspirin recommendations for a 62-year-old patient is not protected by their ignorance, they are exposed by it. Conflict of interest disclosures are already published in guideline documents, surfacing them is a metadata tagging problem, not an original research requirement. The engineering effort is finite. The patient safety consequence of concealment is not.


When response length is cited as a reason to withhold clinically meaningful information, patient safety must take precedence. Always.


Acceptance criterion: No query returns a single guideline recommendation where multiple contradictory guidelines exist, without disclosure of the contradiction and its clinical implications.


Action 5: Audit Training Data for Framing Bias

What: Quantify and correct the framing imbalance inherited from the source literature.


How:

  • Run a systematic audit of training data: what percentage of source documents report ARR and natural frequencies alongside RRR?

  • Build a framing completeness score for each source

  • Up-weight sources that consistently report absolute figures and natural frequencies

  • Down-weight sources that report relative risk only

  • Re-audit after every major training data update


Acceptance criterion: Framing completeness is a documented, tracked metric in the training data quality dashboard, weighted equally alongside impact factor and recency.


Action 6: Implement a Standard QA Stress Test

What: Before any clinical AI output is approved for deployment, stress-test it against a standardized set of clinical efficacy queries.


How:

  • Develop a library of benchmark queries covering major therapeutic areas

  • Submit each query and evaluate the first response against the following pass criteria:

Criterion

Pass

Fail

Natural frequency for benefit stated

First response

Requires follow-up

Natural frequency for harm stated

First response

Absent or requires follow-up

Source tagged as evidence or guideline

First response

Absent

Guideline contradictions disclosed

First response

Single guideline presented without disclosure

Conflicts of interest tagged

First response

Absent

Relative risk labeled as relative

First response

Presented as primary metric

  • Any fail on any criterion = output not cleared for clinical deployment

  • Re-test after every model update


Acceptance criterion: 100% pass rate on benchmark query library before deployment. No exceptions.


Action 7: Adopt One Governing Principle

What: A single ethical standard that supersedes all technical specifications.


The principle:

If a physician acting on your first response alone could make a prescribing decision that harms a patient, a decision they would not have made had they seen the complete data, your first response is inadequate.


Apply this question to every output template, every training data curation decision, and every default setting. When the answer is yes, fix it before deployment, not after.


This is not a technical standard. It is the same standard of honesty that physicians owe their patients, extended, by the people who build these tools, to every physician who trusts them.


The Minimum Acceptable Standard — At a Glance

A clinical AI response is not ready for clinical deployment until it delivers, on the first response, without escalation:


✅ Efficacy in natural frequencies

✅ Harm in natural frequencies, same population, same time horizon

✅ All relevant guidelines surfaced with contradictions disclosed

✅ Conflicts of interest tagged at source level

✅ Source clearly tagged as primary evidence or guideline consensus

✅ Relative risk figures labeled as relative

✅ Treatment duration specified

✅ Validated risk calculators linked where they exist


Every day these defaults remain unchanged, a physician somewhere acts on an incomplete answer. You know what needs to be done. A patient's life could be at risk.

7. Conclusion

This article began with a simple observation: a major clinical AI tool, when asked about the rationale for antiplatelet therapy in chronic ASCVD, led with a relative risk figure and required rounds of escalating questions before producing the absolute numbers that allow individualized clinical decision-making.


That observation expanded into three nested problems:


First: Clinical AI tools systematically present efficacy data using relative risk reduction as the default metric, a statistical framing convention inherited directly from pharmaceutical marketing, a practice long documented to inflate perceived benefit and drive prescribing behavior.


Second: These tools substitute guideline recommendations for primary evidence, presenting the output of committees whose independence is structurally compromised by industry relationships as if it were neutral scientific consensus.


Third: The regulatory agencies, guideline-producing organizations, and AI companies involved in this chain have the power and the resources to fix these problems today.


What This Means for Medicine

A clinical AI tool that presents incomplete, misleading or misrepresenting data is not a neutral tool. It is a vector for systematic bias, operating at the scale of every physician who consults it, in the language of evidence, with the credibility of science itself.


The physician who recognizes this and asks the three questions is practicing medicine. The responsibility for change, however, does not end with the clinician. It lies with the people who built the tool.


What This Means for Patients

A third of adults aged 70 and older, approximately 9 million individuals, continued taking aspirin for primary prevention despite guideline reversal, with a net potential for harm.²⁰


How many patients are currently prescribed medications based on AI-generated responses that lead with relative risk? How many will experience preventable harm because a clinical AI tool presented the half of the truth that most reliably produces a prescription?


We do not yet know. But the mechanism is established, the problem is documented, and the solution is available. Patients deserve access to the complete statistical picture which assist in accurate clinical decision making, not the processed, filtered, commercially optimized derivative that currently passes for evidence-based medicine.


The Path Forward

This article has documented a specific failure in specific tools. But it has identified a systemic problem: the framing conventions of pharmaceutical marketing have been inherited, through training data, through default settings, through institutional inertia, by the AI systems now reshaping clinical decision-making.


The path forward requires simultaneous action at three levels:


Clinicians must demand more. Ask the three questions. Push back on relative risk figures. Refuse to prescribe on the basis of incomplete data.


AI developers must build differently. The seven actions outlined in Section 6 are the minimum acceptable standard. Implement them. Test against them. Deploy only when the standard is met.


Regulators and professional bodies must enforce accountability. Clinical AI tools that present efficacy data without absolute risk figures should not be permitted in clinical settings, not as a best practice, but as a regulatory requirement.

A Final Word

The patient sitting across the desk from a physician consulting a clinical AI tool has no idea what defaults are built into that system.


They do not know whether the recommendation they are about to receive is based on primary evidence or guideline consensus. They do not know whether the benefit number they are hearing is the complete picture or the carefully selected half of one.


Gerd Gigerenzer spent decades teaching the world that humans are not bad at thinking. We are simply given information in the wrong format. When presented with natural frequencies, "25 out of 1,000 patients benefit", people understand risk intuitively.³


The choice of format is not a technical detail. It is a choice about whose interests the information serves.


For too long, that choice has served the interests of people selling medications. It is time for clinical AI to serve the interests of people taking them. The patient is waiting. You know what needs to be done. A patient's life could be at risk.

Acknowledgment:

This article was written with AI assistance. All claims are supported by credible, peer-reviewed references, which were validated for accuracy and authenticity. The AI synthesized information were reviewed by author, ensuring scientific integrity throughout. In the event of any inadvertent errors, the responsibility lies with the AI/authors, and corrections will be made promptly upon identification. I would like to express my sincere gratitude to Dr Tahira Khalid for her thoughtful review and invaluable feedback, and Professor Aftab Ahmad (I.T) for his valuable feedback. Their expertise and guidance have played a pivotal role in refining and enhancing this article.

 

Conflict of Interest Statement:

The author is the developer of a herbal formula and the owner of Dr. Q Formula/Insulinn LLC. However, this affiliation has not influenced the content, analysis, or conclusions of this article

 

Author’s Note on Scope and Intent:

This article does not advocate the replacement of evidence-based conventional care modalities. All complementary interventions are intended to supplement, not supplant, standard clinical practice, and are implemented within a physician-governed, ethically reviewed, and fully documented medical framework.

References

  1. Bucher HC, Weinbacher M, Gyr K. Influence of method of reporting study results on decision of physicians to prescribe drugs to lower cholesterol concentration. BMJ. 1994;309(6957):761–764. Link: https://www.bmj.com/content/309/6957/761

  2. Nuovo J, Melnikow J, Chang D. Reporting number needed to treat and absolute risk reduction in randomized controlled trials. JAMA. 2002;287(21):2813–2814. Link: https://pubmed.ncbi.nlm.nih.gov/12038920/

  3. Gigerenzer G, Gaissmaier W, Kurz-Milcke E, Schwartz LM, Woloshin S. Helping doctors and patients make sense of health statistics. Psychol Sci Public Interest. 2007;8(2):53–96. Link: https://pubmed.ncbi.nlm.nih.gov/26161749/

  4. Antithrombotic Trialists' Collaboration. Collaborative meta-analysis of randomised trials of antiplatelet therapy for prevention of death, myocardial infarction, and stroke in high risk patients. BMJ. 2002;324(7329):71–86. Link: https://www.bmj.com/content/324/7329/71

  5. Office of the Assistant Secretary for Planning and Evaluation. FDA User Fees: Examining Changes in Medical Product Development and Economic Benefits. ASPE Issue Brief. March 2023. Link: https://aspe.hhs.gov/sites/default/files/documents/eafc804ad2c90d9b2a8dd1e04059b378/FDA-User-Fee-Issue-Brief.pdf

  6. Congressional Research Service. FDA Human Medical Product User Fee Programs. Congress.gov. Updated March 2026. Link: https://congress.gov/crs-product/R44750

  7. Antithrombotic Trialists' (ATT) Collaboration. Aspirin in the primary and secondary prevention of vascular disease: collaborative meta-analysis of individual participant data from randomised trials. Lancet. 2009;373(9678):1849–1860. Link: https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(09)60503-1/fulltext

  8. UpToDate Expert AI. Effect of antiplatelet therapy on vascular events [Data Table]. UpToDate, Inc. Accessed June 2026. Link: https://www.uptodate.com/

  9. Mahon N, Takamura M, Aizawa Y, et al. Application of U.S. guidelines in other countries: Aspirin for the primary prevention of cardiovascular events in Japan. Am J Med. 2004;117(7):459–468. Link: https://www.amjmed.com/article/S0002-9343(04)00426-7/fulltext

  10. Yebyo HG, Zappacosta S, Aschmann HE, Haile SR, Puhan MA. Global variation of risk thresholds for initiating statins for primary prevention of cardiovascular disease: a benefit-harm balance modelling study. BMC Cardiovasc Disord. 2020;20(1):418. Link: https://bmccardiovascdisord.biomedcentral.com/articles/10.1186/s12872-020-01697-6

  11. Mortensen MB, Nordestgaard BG, Tybjærg-Hansen A, Afzal S, Saeed S. Statin eligibility for primary prevention of cardiovascular disease according to 2021 European prevention guidelines compared with other international guidelines. JAMA Cardiol. 2022;7(10):1078–1086. Link: https://jamanetwork.com/journals/jamacardiology/fullarticle/2793729

  12. Camm AJ, Califf RM, Dittrich HC, et al. Analysis of conflicts of interest among authors and researchers of European clinical guidelines in cardiovascular medicine. Eur Heart J Open. 2021;1(1):oeab011. Link: https://pmc.ncbi.nlm.nih.gov/articles/PMC8002771/

  13. Ioannidis JPA. Professional societies should abstain from authorship of guidelines and disease definition statements. Circ Cardiovasc Qual Outcomes. 2018;11(1):e004889. Link: https://www.ahajournals.org/doi/10.1161/CIRCOUTCOMES.118.004889

  14. Coyle SL; Ethics and Human Rights Committee, American College of Physicians-American Society of Internal Medicine. Physician-industry relations. Ann Intern Med. 2002;136(5):396–402. Link: https://pubmed.ncbi.nlm.nih.gov/11874314/

  15. [Same as Reference 5 — FDA User Fees document] Link: https://aspe.hhs.gov/sites/default/files/documents/eafc804ad2c90d9b2a8dd1e04059b378/FDA-User-Fee-Issue-Brief.pdf

  16. [Same as Reference 6 — Congressional Research Service] Link: https://congress.gov/crs-product/R44750

  17. Gentilini A, Raymakers AJN, Rand L. Conflicts of interest for FDA advisory committee members and public speakers: systematic review. Value Health. 2026. Primary source Link: https://www.madinamerica.com/2026/05/pharma-cash-creates-conflicts-of-interest-in-fda-testimony-and-clinical-practice-guidelines/

  18. [Same as Reference 7 — ATT Collaboration Lancet 2009] Link: https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(09)60503-1/fulltext

  19. Vandvik PO, Lincoff AM, Gore JM, et al. Primary and secondary prevention of cardiovascular disease: Antithrombotic therapy and prevention of thrombosis, 9th ed: American College of Chest Physicians evidence-based clinical practice guidelines. Chest. 2012;141(2 Suppl):e637S–e668S. Link: https://journal.chestnet.org/article/S0012-3692(12)60134-2/fulltext

  20. Yadav K, Medsker B, Vu A, Aroda VR. Recent trends in aspirin use for cardiovascular disease prevention in the United States, 2015 to 2023. JACC Adv. 2025. Link: https://www.jacc.org/doi/10.1016/j.jacadv.2025.101699


Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating

Chief Editor: Qaisar J Qayyum, MD

drqhealthyliving@gmail.com

Assistant Chief Editor: Tahira Khalid, MD

Publisher: Excellence in Complementary Medicine, LLC, Edmond, OK, USA.

bottom of page