When the Better Hospital Looks Worse

Aurelian’s Casebook

Casebook #1

Good decisions rarely depend on having more numbers. They depend on asking better questions about the numbers already available.

Every day, healthcare leaders rely on performance indicators to guide important decisions.

Hospital mortality, infection rates, readmissions, patient satisfaction scores, and countless other measures help determine where improvement is needed, which programs deserve support, and how an organization evaluates its own performance.

These indicators are valuable.

But before they can support a good decision, they must first answer the right question.

This case begins with what appears to be a straightforward comparison between two hospitals. It ends with one of the most valuable habits an analyst can develop:

Before accepting an important conclusion, pause and ask one more question.

Compared to whom?

Case Presentation

An executive committee is reviewing its monthly hospital performance dashboard.

Among the quality indicators, one measure immediately draws attention: overall hospital mortality.

The dashboard reports the following:

HospitalMoratality Rate
Hospital A8%
Hospital B13%

Based on these figures, the conclusion seems obvious.

Hospital A has the lower mortality rate and therefore appears to provide better care.

At first glance, the evidence seems convincing.

Before accepting that conclusion, however, Aurelian asks a single question.

Compared to whom?

That question changes the entire investigation.

The Case File

Decision Under Review

Determine whether Hospital A is truly outperforming Hospital B based on overall mortality.

Initial Evidence

  • Overall hospital mortality rates
  • Monthly executive performance dashboard

Missing Evidence

• Patient severity

• Case mix

• Distribution of high-risk patients

• Mortality within comparable patient groups

Central Analytical Question

Are these hospitals being compared fairly?

Investigation Status

Evidence Review in Progress

Like any good investigation, we will examine the evidence one piece at a time before reaching a conclusion.

Exhibit A — Overall Mortality

The first exhibit summarizes every hospitalized patient, regardless of diagnosis, severity, or underlying risk.

HospitalMoratality Rate
Hospital A8%
Hospital B13%

The calculation is mathematically correct.

If this were the only evidence available, most decision-makers would reasonably conclude that Hospital A performed better.

But the investigation is not concerned only with whether the calculation is correct. It asks whether the comparison is sufficient for the decision being made

Exhibit B — Patient Case Mix

The next clue appears when we look beyond the overall mortality rate and examine the patients each hospital is treating.

Additional information reveals the following distribution:

Patient SeverityHospital AHospital B
Lower-risk patients 80%50%
Higher-risk patients20%50%

The hospitals are not caring for comparable populations.

Hospital B treats a substantially larger proportion of critically ill patients.

This finding does not invalidate the original mortality calculation.

It introduces another possibility:

Differences in patient severity may influence the overall mortality rate independently of the quality of care.

So the investigation continues.

Exhibit C — Mortality Among Lower-Risk Patients

Rather than combining all patients into a single average, the next exhibit compares mortality among lower-risk patients.

HospitalLower-Risk Mortality Rate
Hospital A3%
Hospital B2%

Within this patient group, Hospital B has the lower mortality rate.

The original conclusion is no longer certain.

One comparison now favors Hospital B.

A second question naturally follows:

Does the same pattern appear among higher-risk patients?

Exhibit D — Mortality Among Higher-Risk Patients

The final exhibit compares patients with greater illness severity.

HospitalHigher-Risk Mortality Rate
Hospital A28%
Hospital B24%

Again, Hospital B has the lower mortality rate.

The evidence now appears contradictory.

Hospital B performs better among lower-risk patients.

Hospital B also performs better among higher-risk patients.

Yet Hospital B reports the higher overall mortality rate.

How can all three statements be true?

Analytical Interpretation

The apparent contradiction disappears once the patient populations are considered.

Hospital B treats a substantially larger proportion of critically ill patients.

Because higher-risk patients experience greater mortality regardless of where they receive care, Hospital B’s overall mortality rate is higher even though its outcomes within comparable severity groups are better.

The original dashboard answered one question accurately:

What happened across all patients combined?

The executive decision required a different question:

How did each hospital perform when treating comparable patients?

These are not equivalent questions.

The overall mortality calculation was mathematically correct.

The comparison was analytically incomplete.

This distinction is subtle, but fundamental.

Statistics rarely mislead because the arithmetic is wrong.

They mislead when we ask them to answer questions they were never designed to answer.

A Name for the Observation

Only after understanding the evidence do we need a statistical name for what occurred.

This type of reversal is commonly described as Simpson’s paradox.

It occurs when a relationship observed within several groups changes after those groups are combined into a single summary measure.

The statistical name matters far less than the habit of thinking carefully about comparisons.

The paradox is not the lesson.

Disciplined comparison is.

Beyond Hospital Mortality

The principle demonstrated in this case extends far beyond mortality statistics.

Healthcare organizations compare surgeons, hospitals, clinical departments, quality indicators, and treatment outcomes.

Researchers compare interventions, populations, and study results.

Public health professionals compare disease rates across communities.

Educational institutions compare schools.

Businesses compare branches, products, and employee performance.

In every setting, a summary indicator is useful only when the underlying comparison is appropriate.

Good analysts resist accepting the first comparison they see.

They first ask whether the groups being compared are genuinely comparable.

Only then do they interpret the numbers.

Key Takeaways

• Overall indicators may answer a different question from the one decision-makers intend to ask.

• Crude mortality rates may be mathematically correct without fairly representing organizational performance.

• Patient case mix can substantially influence summary outcome measures.

• Comparisons within similar patient groups may provide a more meaningful assessment than overall averages.

• Statistical terminology matters less than disciplined reasoning.

Compared to whom?

A Final Thought

This case was never about proving that mortality rates are unreliable.

Nor was it about introducing an interesting statistical paradox.

It was about learning to ask better questions.

Is the calculation correct?

Is the comparison fair?

Does the number answer the question we are actually trying to answer?

Those questions are simple.

The discipline to ask them consistently is not.

Good analysts rarely make better decisions because they have more data.

They make better decisions because they know when a number deserves another question.

Before accepting the conclusion, ask one more:

Compared to whom?

References

• Rothman KJ, Greenland S, Lash TL. Modern Epidemiology.

• Gordis L. Epidemiology.

• Kleinbaum DG, Klein M. Epidemiologic Research: Principles and Quantitative Methods.

• Altman DG. Practical Statistics for Medical Research.

• Pearl J, Mackenzie D. The Book of Why.

Related Reading

Although Aurelian’s Casebook begins with this inaugural publication, readers interested in healthcare indicators and evidence interpretation may also find the following Aurelian’s Surveillance Files useful:

Field Note #1 — Surveillance vs. Research: Same Data, Different Purpose

Field Note #4 — Syndromic Surveillance: Seeing the Pattern Before Knowing the Diagnosis

Field Note #7 — The Surveillance Chain: Every Signal Depends on What Happens Next

As the Academy grows, future Casebooks will examine related topics, including averages, denominators, risk adjustment, dashboard design, and healthcare performance measurement.

Discussion

Have you ever encountered a performance indicator that appeared convincing until additional context changed its interpretation?

What information would you want before making an important decision based on a single summary statistic?

Leave a Comment