Minimally Important Difference for EQ5D
Defending the indefensible?
Two papers appearing simultaneously in Value in Health stake out opposing positions on a question that matters to anyone interpreting health-related quality of life scores. Should we estimate minimally important differences (MIDs) for preference-weighted measures like the EQ-5D?
Johnson and Al Sayah argue yes – MIDs are useful interpretive aids that simply need better methodological guidelines. My colleagues and I argue no – MIDs for preference-weighted measures are not just unhelpful but conceptually incoherent, and it’s time to stop producing them.
The stakes are real: Al Sayah’s own systematic review identified 840 MID estimates for EQ-5D across 90 studies. These thresholds are used to interpret clinical trials, assess health system performance, and guide treatment decisions. If the concept is flawed, it’s patients that will pay the price.
The clue is in the title: preferences
Johnson and Al Sayah accuse critics of treating preference-weighted measures as having 'mystical status'. This mischaracterises our argument. We're not claiming utilities are mystical – we're pointing out they already incorporate the value judgement that MIDs claim to add separately. That's not mysticism; it's what preference-weighted measures do by definition.
Health state values already embody a measure of importance. That’s what preferences are – statements about relative value. If a health state is valued at 0.8 it is preferred to one valued at 0.7, then any movement from 0.7 to 0.8 is meaningful by definition. The MID literature treats this as an empirical question requiring estimation, but it’s asking the wrong question of the wrong kind of measure.
Johnson and Al Sayah’s central claim is that preference-weighted measures are “just numbers on a scale” requiring interpretation aids like any measurement. But this misses what makes preference-weighted measures distinctive. The preferences already determine relative importance. You can’t then ask “but is this difference important?” without ignoring what the numbers represent.
Consider economic evaluation, where even Johnson and Al Sayah concede MIDs are irrelevant. Why? Because cost-effectiveness thresholds already incorporate the value judgement: a 0.05 QALY gain might be ‘worth it’ at a cost of £1,000 but not at £100,000. Small health gains can be worthwhile if they're cheap to achieve. They even suggest that cost-effectiveness thresholds are analogous to MIDs. But this comparison fails: thresholds incorporate both costs and outcomes. MIDs claim to judge outcome importance without cost context – that's a fundamentally different exercise and one that is doomed to failure.
But their central argument is that EQ-5D is used outside economic evaluation – in clinical trials, PROM programmes, population health monitoring. Does this mean we need MIDs? No – it means we need appropriate analytical methods for those contexts. And MIDs aren’t it.
When Averages Mislead: Importance is an Individual Concept
Imagine 100 patients receive a treatment. Ninety-nine experience no change in health status. One patient has a profound improvement: a utility gain of 0.8, equivalent to moving from severe disability to near-perfect health. The group mean benefit is 0.008 – below almost any published MID threshold.
Would we reject this as ‘not meaningful’? Of course not. It’s a Pareto improvement: one person dramatically better off, nobody worse off. Yet applying an MID to the mean difference would dismiss it as trivial.
This is what happens when you apply scalar thresholds to distributions. MIDs hide who benefits and by how much. They obscure precisely the information needed for decisions: how many patients improved, by what magnitude, and who was left behind. The solution isn’t better MIDs – it’s distributional analysis, responder rates, and transparency about heterogeneity.
The Empirical Crisis
If the conceptual argument leaves you cold, consider the evidence. Al Sayah’s systematic review found MID estimates for EQ-5D-3L ranging from 0.003 to 0.952 – a 300-fold variation (just let that sit for a moment)! For the median estimate of 0.1 to represent ‘minimally important’ we’d have to believe the full range of MIDs they quote including that the highest value of 0.952 are legitimate in the first place. But how can 0.952 possibly be legitimate? It defies all logic. The range isn’t measurement error – it’s evidence that MID isn’t a stable property of the EQ5D instrument but varies with population, context, method, and chance.
Johnson and Al Sayah’s response? We need context-specific MIDs adjusted for baseline score, direction of change, clinical condition, and intervention type. But this concedes the game. If MID varies continuously with context, it’s not an interpretive threshold. The baseline score problem is particularly revealing. At least two studies have found that patients in poor health need larger improvements to reach MID thresholds. Someone with a baseline utility of 0.3 might need a 0.08 gain to be “meaningful,” whilst someone at 0.8 needs only 0.04.
Contrast this result with NICE’s severity modifiers, which explicitly weight smaller QALY gains more heavily when baseline health is poor. NICE says gains matter more for the severely ill. The MID literature says they need to be larger to matter. Both frameworks use the same EQ-5D instrument. They can’t both be right. And since NICE’s approach is grounded in equity considerations and opportunity cost, the MID framework reveals itself as capturing something other than importance – likely statistical artefact.
After a Decade, What Have We Learnt?
In 2014, Coretti and colleagues called for better methodological standards to reduce variability in MID estimates. A decade later, we have 840 estimates and more variability. The response from MID proponents is invariably: better guidelines will help. But this is what was said ten years ago. At what point do we conclude that methodological refinement can’t fix a conceptual error?
Johnson and Al Sayah cite the volume of MID research as evidence of need. I see it as evidence of how easy it is to generate publishable estimates once the conceptual error is made. The barrier to entry is low: collect some EQ-5D data, pick an anchor, run a correlation, publish your MID. The literature has industrialised the production of a fundamentally flawed metric.
Whitehurst and Bryan wrote in 2013 that “just because it is possible to construct a minimally important difference for preference-based measures does not mean that it is useful.” The subsequent decade – with its 840 estimates, 300-fold variation, and continued confusion – has proven them right.
What Should We Do Instead?
The need for interpretive guidance is real. Researchers and clinicians working with EQ-5D scores deserve better than ‘figure it out yourself.’ But MIDs aren’t the answer – they lack precision, are conceptually flawed and will mislead decision-makers.
For clinical trials and effectiveness studies: report effect sizes with confidence intervals, use responder analysis to show distributions, and be transparent about who benefits. For economic evaluation: use explicit cost-effectiveness thresholds that incorporate opportunity cost. For population health monitoring: track distributions over time and across subgroups, not just means against arbitrary thresholds.
And for journals: stop publishing MID estimates for preference-weighted measures. Each new paper adds to a literature that obfuscates rather than illuminates. After 840 estimates and a decade of “better methods,” it’s time to admit the concept doesn’t work.
Johnson and Al Sayah end by calling for consensus-based guidelines to produce more consistent MIDs. I agree we need consensus – consensus that the enterprise has failed and should be abandoned. The data from their own systematic review makes the case. Defending the MID isn’t pragmatism and it does not help decision makers. It’s defending the indefensible.



MIDs here feel like binning a continuous EQ-5D measure for interpretability, like categorizing continuous predictors in regression: simpler story, but information loss and arbitrary cutoffs. Especially since preference weights already encode value judgements, an extra “importance” threshold risks double-counting.
Great work. There is a wider context to what you call out: a disorder Ithat occurs in various degrees of severity from ‘preference aversion’ through ‘preference phobia’ to ‘preference psychopathy’. It is endemic in the medical profession, where it is inculcated in training and reinforced in daily practice - and in the clinical literature - where the word is taboo. Unfortunately it is contagious and infects non-medical healthcare researchers, who should know better but whose careers are dependent on medically-dominated boards of all kinds (funding, awards, ethics) Preferences just make decision making too difficult, especially when one accepts they should be treated as analytically as evidence, rather than given a tokenistic nod (yes, we should ‘take them into account’ when those of patients are mentioned). Deep down the medical doxo ignores the difference between ontology (state 21334) and axiology (u21334). The J and S case of the disorder is particularly interesting insofar as it shows how even those who have been/are heavily involved in the construction of preference measures don’t fully understand what can and can’t de done with them .