When a Good KPI Tells the Wrong Story

Why hospital performance cannot be understood without population, capacity, and context.

Aurelian’s Casebook

Casebook #3

The dashboard was green.

Surgical backlog: 0%.

No patients were waiting beyond the hospital’s defined backlog threshold.

Then the hospital director received a question from regional leadership:

Why is your operating room performing so few surgeries?

It was a reasonable question.

The hospital had an operating room. It had surgical specialists. Available operating-room time was going unused.

A neighboring hospital in the same health system performed far more procedures and maintained much higher operating-room utilization.

At first glance, the comparison seemed obvious.

One hospital was busy.

The other was not.

But before deciding that the smaller hospital was underperforming, there was another question to ask:

How many surgeries should this hospital actually be performing?

That question changes the case.

Because surgical volume is a numerator.

And numerators do not explain themselves.

The Dashboard

Consider two fictional hospitals operating within the same healthcare system.

Hospital A

Hospital A is a regional medical center serving a larger population. Its population is younger, birth volume is higher, and its surgical services include General Surgery, Orthopedics, and OB/GYN.

During the year, Hospital A reports:

Surgical volume: 1,180 procedures

Operating-room utilization: 74%

Surgical backlog: 6%

Hospital B

Hospital B is a smaller community hospital serving a smaller and older population. Birth volume is very low. Orthopedics and OB/GYN are available, but General Surgery is not.

Hospital B reports:

Surgical volume: 310 procedures

Operating-room utilization: 31%

Surgical backlog: 0%

If we look only at activity, Hospital A appears substantially more productive.

Its operating room performs almost four times as many procedures and uses a much greater proportion of available capacity.

Hospital B’s operating room spends much more time unused.

That deserves attention.

But low activity is an observation. Underperformance is a conclusion.

Before moving from one to the other, we need to understand the system producing the numbers.

Low activity is an observation. Underperformance is a conclusion.

Before Judging the Numerator, Understand the Population Generating It

Hospitals do not manufacture surgical demand.

They respond to it.

The number and type of procedures a hospital performs emerge from several interacting factors: population size, age structure, fertility and birth volume, disease burden, clinical indications, available specialties, referral patterns, historical access to care, and the capacity of the institution itself.

That means two hospitals belonging to the same healthcare system should not automatically be expected to generate the same surgical activity.

Look again at our fictional hospitals.

Hospital A serves a larger and younger population. It has more births. It also maintains General Surgery, Orthopedics, and OB/GYN.

Those characteristics create opportunities for a broader surgical case mix.

Hospital B serves a smaller, older population with very low birth volume and no General Surgery service.

Why should those populations generate identical surgical demand?

They should not.

This is where a tool familiar to epidemiologists becomes useful to hospital management: the population pyramid.

A population pyramid is more than a demographic illustration. It helps us understand the population from which healthcare demand emerges.

Populations with different age structures can generate different patterns of healthcare need. Birth volume, for example, directly affects the potential demand facing an obstetric service, while an older population may generate a substantially different mix of health and care needs.

Available specialties also determine which clinically indicated procedures can actually be performed locally.

History can matter as well.

Imagine that Hospital B serves a relatively stable population that had reliable access to General Surgery for several decades before the service disappeared.

Historical access may have influenced the amount of unmet surgical need present when that service changed, while new need continues to emerge according to the characteristics of the population that remains.

None of these observations proves that Hospital B is appropriately configured.

They tell us something more important:

The number 310 cannot tell us by itself whether Hospital B is performing well or poorly.

The denominator matters.

So does the population.

So does the service portfolio.

So does context.

Three KPIs. Three Different Questions.

Part of the problem disappears once we stop asking different indicators to answer the same question.

Surgical Volume

Surgical volume asks:

How many procedures were performed?

Operating-Room Utilization

Operating-room utilization asks:

How much of the available operating-room capacity was used?

Surgical Backlog

Surgical backlog asks:

Among patients already determined to require surgery, are clinically indicated procedures waiting beyond the hospital’s defined backlog threshold?

These measures are related, but they are not interchangeable.

A hospital can therefore report:

0% surgical backlog

low surgical volume

low operating-room utilization

without those findings necessarily contradicting one another.

Suppose every patient currently requiring a procedure within the hospital’s available specialties receives it without unnecessary delay.

The backlog may legitimately be zero.

If relatively few procedures are clinically indicated, total surgical volume may remain low.

And if the institution maintains more operating-room capacity than current demand requires, utilization may also remain low.

The dashboard can therefore be technically correct across all three indicators.

The analytical problem begins when we convert those measurements into conclusions they cannot independently support.

A 0% backlog does not prove that the surgical program is optimally designed.

Low OR utilization does not prove that staff are failing to work.

High surgical volume does not automatically prove superior performance.

Each indicator tells us something.

None tells us everything.

A KPI can measure accurately and still tell the wrong story when interpreted without context.

The Better Management Question

Regional leadership was right to notice Hospital B.

An operating room functioning at 31% utilization deserves investigation.

Perhaps patients are being referred elsewhere unnecessarily.

Perhaps access barriers prevent patients from reaching the service.

Perhaps scheduling is inefficient.

Perhaps staffing shortages suppress activity.

Perhaps the available specialties no longer match the needs of the population.

Perhaps the hospital maintains more surgical capacity than its population currently requires.

Or perhaps low activity is exactly what we should expect from the population and services available.

The KPI cannot distinguish among those explanations.

Analysis can.

That changes the management question from:

Why aren’t you performing more surgeries?

to:

Why is surgical activity low, and is the current level appropriate for the population this hospital serves?

The difference may seem subtle.

It is not.

The first question begins with an expected answer: more surgeries.

The second begins with an investigation.

And that distinction matters because increasing surgical volume should never become an objective independent of clinical need.

A hospital should not perform surgery because a dashboard needs a larger numerator.

It should perform surgery because a patient has an appropriate clinical indication for a procedure that the institution can safely provide.

If demand is lower than available capacity, management may indeed have a problem.

But the problem may be capacity configuration, not clinical productivity.

Perhaps operating-room resources should be redesigned.

Perhaps surgical specialties should be reconsidered.

Perhaps referral relationships should change.

Perhaps capacity should serve a broader population.

Perhaps the existing configuration remains justified for access, resilience, emergency capability, or other reasons not captured by utilization alone.

Those are management decisions.

But they require understanding why the number exists before deciding what the number should become.

Capacity is not demand.

And forcing demand to resemble capacity reverses the purpose of healthcare.

Aurelian’s Casebook

KPIs are indispensable.

Without measurement, organizations cannot reliably identify problems, compare performance, allocate resources, or determine whether interventions are working.

The lesson of this case is therefore not to distrust KPIs.

It is to respect what they can—and cannot—tell us.

Hospital A and Hospital B belong to the same healthcare system.

They share the same mission.

That does not mean they serve the same population, provide the same services, encounter the same demand, or should produce identical numbers.

Population structure matters.

Clinical demand matters.

Available specialties matter.

Referral patterns matter.

Historical context matters.

Capacity matters.

And before interpreting the numerator, we need to understand the population and system generating it.

Low activity is an observation. Underperformance is a conclusion.

Before judging the numerator, understand the population generating it.

Discussion

Consider a KPI used in your own organization.

What question was that indicator originally designed to answer?

And what conclusions are people actually drawing from it?

Are those the same thing?

When two facilities are compared, are differences in population, service portfolio, capacity, case mix, or historical context considered before one is labeled the better performer?

And perhaps the most important question:

If a KPI tells us something is wrong, have we investigated why—or have we simply demanded that the number change?

References

Editorial disclosure: Hospital A, Hospital B, their population characteristics, service configurations, and all numerical data presented in this Casebook are fictional and were created for educational purposes. The analytical principles discussed are evidence-informed; the fictional scenario is not intended to represent any specific institution.

  • World Health Organization. Toolkit for Analysis and Use of Routine Health Facility Data: Integrated Health Services Analysis — District and Facility Level. Geneva: World Health Organization; 2023. ISBN 978-92-4-006060-9.
  • World Health Organization; United Nations Children’s Fund. Primary Health Care Measurement Framework and Indicators: Monitoring Health Systems Through a Primary Health Care Lens. Geneva: World Health Organization; 2022. ISBN 978-92-4-004421-0.
  • World Health Organization. World Report on Ageing and Health. Geneva: World Health Organization; 2015. ISBN 978-92-4-156504-2.
  • Viapiano J, Ward DS. Operating room utilization: the need for data. International Anesthesiology Clinics. 2000;38(4):127–140. doi:10.1097/00004311-200010000-00009.
  • Strum DP, Vargas LG, May JH. Surgical subspecialty block utilization and capacity planning: a minimal cost analysis model. Anesthesiology. 1999;90(4):1176–1185. doi:10.1097/00000542-199904000-00034.
  • Rose J, Weiser TG, Hider P, Wilson L, Gruen RL, Bickler SW. Estimated need for surgery worldwide based on prevalence of diseases: a modelling strategy for the WHO Global Health Estimate. The Lancet Global Health. 2015;3(Suppl 2):S13–S20. doi:10.1016/S2214-109X(15)70087-2.

Related Casebooks

Casebook #1 — When the Better Hospital Looks Worse
How crude mortality comparisons can reverse once differences in patient populations are considered.

Casebook #2 — When Better Numbers Mean Worse Care
Why declining admissions, bed-days, and occupancy do not necessarily mean hospital performance is improving.

Knowledge Applied with Prudence

Leave a Comment