Composite Endpoints in Clinical Trials
Shortcut or stumbling block for HTA?
Clinical trials thrive on efficiency. Companies want to demonstrate results quickly, patients want access to effective treatments sooner, and regulators want clarity. Composite endpoints are one of the main tools used to speed things up. By bundling several outcomes into a single measure — like hospitalisation and death — researchers can increase event rates, shorten trials, and deliver headline results that tick regulatory boxes.
From a trial design perspective, this looks like a win. But when the results land on the desk of a health technology assessment (HTA) agency, the picture is rarely so tidy.
Why composites cause headaches in HTA
The problem is that not all outcomes bundled into a composite endpoint are equal. Avoiding hospitalisation is important, but avoiding death is clearly more so. For patients, the stakes are obvious. For HTA bodies, which translate trial data into estimates of cost-effectiveness, the differences are also critical.
A single composite effect estimate masks these differences. If a drug mostly reduces minor outcomes but shows little or no effect on mortality, the composite can still look impressive. When plugged into an economic model, however, the implications diverge: reduced hospitalisations save some money, but preventing deaths delivers far greater health gains. Without disaggregating the data, decision-makers risk over- or under-estimating the true value of the treatment.
But disaggregating also carries risks. A recurring issue is how to handle “non-significant” results of components of the composite. A common temptation is to drop an endpoint entirely if its effect is not statistically significant. But this confuses absence of evidence with evidence of absence. Clinical trials powered for composites are rarely designed to detect effect differences on each component individually. Removing non-significant outcomes hard-codes false certainty into an HTA model.
A real-world example: dapagliflozin for heart failure
This is not a theoretical quibble. In England, NICE was asked to assess dapagliflozin for heart failure with preserved or mildly reduced ejection fraction. The pivotal trial used a composite endpoint of cardiovascular death or hospitalisation for heart failure.
The trial showed an overall benefit, but the two components told a more nuanced story. Hospitalisations fell significantly, while the mortality effect was less clear. When NICE reviewed the evidence, the question was whether to model the mortality effect at all.
NICE’s assessment group suggested dropping it, effectively assuming, with certainty, that dapagliflozin had no impact on death. But this assumption clashed with both the trial data and HTA good practice. Ultimately, NICE judged that mortality benefits could not be ruled out and should be included — with appropriate uncertainty. That shift changed the cost-effectiveness picture and contributed to dapagliflozin being recommended for NHS use.
The lesson here is simple: how composite endpoints are handled can make the difference between rejection and access.
Beyond one trial: when to pool evidence
In our work we didn’t stop with dapagliflozin nor with the specific indication. The same class of drugs — SGLT2 inhibitors — includes empagliflozin, tested in similar heart failure populations with preserved ejection fraction. Both drugs have also been studied in related conditions, such as heart failure with reduced ejection fraction. This raises another methodological challenge: should HTA agencies treat these bodies of evidence as separate silos, or as connected strands of the same story?
Pooling evidence across drugs and indications can reduce uncertainty. Even if individual trials struggle to show significant effects on mortality, combining results may reveal consistent patterns. In the SGLT2 inhibitor case, when results from different trials were pooled, the evidence for mortality benefits strengthened. This matters because HTA is not only about point estimates, it is also about characterising and managing uncertainty. A broader evidence base gives decision-makers more confidence that they are not making costly mistakes.
Of course, pooling requires judgement. Drugs in the same class may differ subtly in effect, and patient populations are not always comparable. But ignoring relevant evidence is rarely the safer path.
Why guidance matters
For regulators, composites make sense. They speed up development, avoid multiple hypothesis testing, and deliver a single neat number. For HTA bodies, which must value treatments in terms of health outcomes and costs, composites create ambiguity. Yet most HTA guidelines provide little direct advice on how to handle them.
That gap leaves too much room for discretion. Companies may naturally favour analyses that show their product in the best light. Assessment groups may apply inconsistent rules. Committees face the unenviable task of weighing decisions that could swing on whether to include or exclude a single uncertain component.
Without clearer principles, decisions risk being inconsistent — and worse, patients risk being denied effective treatments or offered poor value care.
Where next?
In a two-part series of methodological articles, we set out a practical framework for dealing with composite endpoints in HTA. At its core, the message is straightforward:
Use disaggregated data when components clearly differ in both importance and effect.
Avoid treating non-significance as no effect; uncertainty should be propagated through the model.
Consider evidence from related trials, drugs, and indications where it is reasonable to do so.
Above all, prioritise estimation and uncertainty over neat but misleading simplifications based on hypothesis tests that are underpowered.
Composite endpoints will remain central to trial design. The challenge for HTA is to stop them becoming a shortcut that leads to poor decisions. We believe the way forward is not to abandon composites, but to handle them with greater transparency, consistency, and respect for uncertainty.
Closing note
This blog is based on a two-part series of manuscripts where we explore these issues in depth, including the case study of SGLT2 inhibitors in heart failure. For those interested in the detail, the links to the full papers, published in the Journal of Comparative Effectiveness Research, are given below.


