Showing posts with label QALYs. Show all posts
Showing posts with label QALYs. Show all posts

Monday, 29 August 2011

Composite endpoints in RCTs: what are they worth?

Composite endpoints in Phase 3 trials – doncha just love them?  Well I don’t.  As has been pointed out elsewhere, in terms of proving value in an economic evaluation, the first thing I want to do is pick the composite apart because I want to convert each bit into QALYs and savings to understand how that compares against the added cost.

An article has just been published that goes some way to illustrating this point:
‘Weighting components of composite end points in clinical trials: an approach using disability-adjusted life-years’ K.-S. Hong, L. Ali, S. Selco, G. Fonarow and J. Saver  Stroke 2011; 42: 1722-1729.

You need to read the original article but in brief they have focused on vascular endpoints and converted the common components into DALYs left as follows:
7.63 DALYs lost per non-fatal stroke
5.14 DALYs lost per non-fatal MI
11.59 DALYs lost per vascular death
In DALY terms, therefore, if a non-fatal MI = 1 then a non-fatal stroke = 1.48 and vascular death = 2.25.

As a QALY-orientated economist my ideal would have been if they had used QALYs instead of DALYs, but I can understand the DALY disease weightings are more accessible than disutilities in QALY studies.  If you were intending to use these results you also need to understand the different assumptions made – events happen at age 60, US life expectancy data, 3% discount rate, and an assumption good health for older people is worth less than good health for younger people, and so on.  As I said, you have to read the article!

But with those gripes aside I think this is a fantastic illustration of the issue.  I’d like to take it one stage further because Hong and colleagues were thinking as clinicians and trying to produce a measure of health effect alone whereas I would be interested in savings as well.  Supposing I work in a system that is willing to pay £20,000 per QALY, and let’s assume for present purposes DALYs and QALYs are roughly equivalent.  Just to illustrate let’s assume that the lifetime discounted cost of managing events are as follows:
Non-fatal stroke £20,000
Non-fatal MI £4,000 without PCI, £10,000 with PCI
Vascular death £5,000
Then in QALY terms these are worth 1, 0.2 to 0.5, and 0.25 QALYs respectively.

Adding these back in to Hong et al’s figures, we get
Non-fatal stroke = 7.63 (health) + 1 (saving) = 8.63
Non-fatal MI = 5.14 (health) + either 0.2 or 0.5 (saving) = 5.34 to 5.64
Vascular death = 11.59 (health) + 0.25 (saving) = 11.84
Using the higher value of 5.64 for non-fatal MI and setting that to 1, the ratios are 1.53 (non-fatal MI) and 2.1 (vascular death).

I’m a little surprised as my intuition would be that there is a bigger gulf between non-fatal MI on the one hand and non-fatal stroke and vascular death on the other.  I don’t perceive the disability consequence of a non-fatal MI to be any where near that of a stroke.  Vascular death also seems to me a major loss of DALYs or QALYs, losing years of life at 0.6 or 0.7 quality, whereas an MI might be the difference between 70% and 60% over a decade.

But that is to lose sight of the main point of this article which is to put in the public domain something to get this sort of debate started.  Thank you to Hong and colleagues!


Wednesday, 24 August 2011

Utility values for diabetes

The search for consistency (and its desirability)

One of the emerging themes of this blog is the extent to which we can establish standardise aspects of producing HTA evidence and a paper I have just read by Lung et al is an illustration:

Lung TW, Hayes AJ, Hayen A, Farmer A, Clarke PM.
Qual Life Res. 2011 Apr 7. [Epub ahead of print]

This team carried out a literature search so thorough it has me worrying for their psychological stability to identify studies that used one of the QALY-compatible preference measures like EQ-5D or an SF measure, or which used time trade-off or standard gamble in people with diabetes.

They report huge ranges in the values obtained from a humble 14 points for diabetes with no complications (i.e. lowest was 0.74, highest was 0.88) through to 48 points (stroke and end-stage renal disease).  It’s obvious that stroke, for example, would depend on severity and ESRD might depend on whether the person required dialysis and, if they did, whether it was hospital or home based.  However, if an HTA organisation accepts published values from the literature, say because it used its preferred utility elicitation technique, it has handed a substantial element of choice to the people writing the HTA submission.  For ESRD in diabetes will we use a utility value of 0.33 citing source study A or 0.81 citing source study B? 

Using stats techniques I am too dull to understand they then carried out two analyses that I will pick out, a random effects meta-analysis (MA) and a random effects meta regression (MR).

The MA gives a point estimates of the mean utility value across the studies but, as important, it provides a 95% confidence interval (and sample size) ideal for use in sensitivity analysis.  For example, the diabetes with no complications state had values from individual studies from 0.74 to 0.88 but in the MA the mean value was 0.81, 95% confidence interval 0.78 to 0.84.  Fantastic, as a reviewer of HTA submissions, that is so helpful.

In the MR they analyse how much of the variation between the estimates in different studies can be explained by measured features such as the age of the patients, their sex, and the elicitation technique.  Older age and being female led to lower utilities (<<insert joke of your choice>>), and of the elicitation techniques TTO and SG (combined) gave higher values than EQ-5D which, in turn, gave higher values than HUI-3 and SF-6D (combined).    

Tom Lung and team, take a bow, my grateful thanks.  The only other study like this I am aware of is in stroke:

Tengs TO, Lin TH.
Pharmacoeconomics. 2003;21(3):191-200.

Are other people aware of similar studies?  And how should we use them? 

Clearly, the ideal is still that companies measure quality of life in their trials in a way that is compatible with QALYs.  However, in this ‘second best’ situation, I think there is a strong case for making values from these meta-analyses the default settings for an HTA submission.  Of course I would be interested in listening to a company’s arguments for why this should not apply to their particular submission – for example, suppose it could be shown that for a particular treatment the ESRD experienced secondary to diabetes were always of a milder type than a utility value of 0.48 (from Lung et al’s meta-analysis) would imply.

But what we all need to get away from is where an HTA submission can select between two hugely different utility values and cite a supporting reference with equal authority.  This should work for companies as well as it gives them greater certainty when they are estimating the likely cost per QALY at an early stage of a product’s life and it may save them some money on commissioning their own utility surveys. 

Tuesday, 16 August 2011

Mapping to QALYs

Mapping to QALYs: how do I know if that’s good enough?

In HTA, QALYs are the least-bad measure of the value of health outcomes that we have so we have to make them work.  The best way, we think, is by measuring outcomes prospectively in a clinical trial.  When we don’t have that, we reach for a ‘second best’ such as taking values from a previously published study or by using a disease-specific outcome and then estimating a conversion equation to derive quality of life values (utilities) to use as the weights in QALYs.  This is called mapping.
But this poses a challenge to people like me tasked with reviewing the quality of an economic evaluation: what should we be looking out for in mapping studies and what standards should we accept?

Most of the literature I could find on ‘good practice’ focused on the derivation of the relationship between the disease-specific scale and utilities in the first place; less attention has been paid to how these estimates are then used by other researchers.

Useful references were as follows:
A review of studies mapping (or cross walking) non-preference based measures of health to generic preference-based measures.Brazier JE, Yang Y, Tsuchiya A, Rowen DL.  Eur J Health Econ. 2010 Apr;11(2):215-25
and

‘Do estimates of cost-utility based on the EQ-5D differ from those based on the mapping of utility scores?’ Barton GR, Sach TH, Jenkinson C, Avery AJ, Doherty M, Muir KR.  Health Qual Life Outcomes. 2008 Jul 14;6:51.


In terms of the studies estimating the disease-specific to utility relationship, a reviewer should bear in mind the following:

1. Mapping must be based on data, not opinion
There is general agreement that using the opinion of clinicians to predict how a patient would have answered an EQ-5D questionnaire based on their responses to a disease-specific questionnaire is no longer acceptable.

2. The minimum expectation is linear regression analysis. 
Linear regression analysis is commonly used but because the utility scale is bounded the results may be biased and inconsistent.  Especially if a large proportion of subjects are in full health, better options may be Tobit, censored least absolute deviations or restricted maximum likelihood (Brazier).

3. Different models should be tested,
Some evidence suggests simple methods are best with added complexity having little value.  Other studies have found squared terms and interactions, as well as patient characteristics.  For example, Barton et al developed five models to predict utilities in rheumatoid arthritis:
Model A: total WOMAC score only
Model B: pain, stiffness, functioning (sub-scales of WOMAC)
Model C: total WOMAC and total WOMAC2
Model D: pain, functioning, pain*functioning, pain*stiffness, stiffness*functioning, pain2, stiffness2, functioning2
Model E: best of models A through D plus age and sex of patient
The final model, E, was found to have the best fit.

4. A goodness-of-fit test should be reported. 
Brazier et al’s review found R2 was the most commonly used measure – when mapping a generic to a generic (e.g. SF-36 to EQ-5D) a value of 0.5 can be achieved.  Of the 30 studies they identified one of the lowest R2 values for a disease-specific to generic was 0.17 so if your R2 is at this level you have a problem.

5. The key test is the ability to predict.
The main reason for estimating the quantitative relationship between the disease-specific measure and utilities is to be able to predict.  An essential part of the validation of a model is therefore to test it, either by dividing the original sample and estimating the relationship on one half and testing it on the other, by obtaining a second dataset using the same disease-specific outcome, or a similar test.
Brazier et al propose a measure of prediction error such as Mean Absolute Error should be used; an example is Barton et al who defined it as the average value of the difference between actual and predicted values.  Barton et al briefly review other studies and found an MAE of 0.13 was the lowest observed (where low is good) and 0.19 was the highest observed.  Brazier et al found lower MAEs but it was unclear if these were for disease-specific to generic mapping studies.
Plotting errors against EQ-5D-based utilities can also be helpful; there may be a tendency for utilities to be underestimated for those in better health and over-estimated for those in poorer health.

So that gives me some idea of what to look for in the source study, but what about how it was applied in the submission I am reading? 

My first issue is: what is the hierarchy of sources for utility values?  Presumably we place most faith in direct measurement of an instrument such as EQ-5D in a trial but how does mapping compare with – say – time trade-off (TTO) valuation of descriptions of health states?  The balance would seem to lie with mapping because the disease-specific scale was measured in the trial whereas the TTO values are for descriptions which are then applied retrospectively to how the patients MAY have been feeling, which seems to introduce potential biases.

This raises another interesting question: suppose the study in front of me used TTO (or similar) but I know that a secondary outcome measure could have been mapped to estimate QALYs using an existing study – as a reviewer, should I insist this is carried out?  On the one hand I know mapping studies are imperfect but they are based on patients’ self-assessed health.  The answer is that I would probably request a sensitivity analysis using mapping, but I realise that is a way of ducking!

The second issue is how I would know there WAS an existing mapping study in the first place.  The only published review I could find, by Brazier and colleagues, covered 30 studies up to 2007, but several were establishing relationships between generic instruments and others were unpublished studies; none used cancer-specific scales (the topic I was interested in).  So do I have to conduct a literature search each time I see a utility that isn’t based on the gold standard method?  Should the onus be on the pharma company to establish whether a mapping study is available?  A great solution would be to construct a database of mapping studies.  I’ve made a start; please e-mail me if you are interested.

The third issue is: if there are several mapping studies available, which should I use?  For example, I am interested in the EORTC QLQ-C30 measure and even a brief literature search identified three different studies, with a search of the references identifying as many again?  The QLQ-C30 measure is intended to be used across many different types of cancer, but does this mean the mapping algorithm is equally applicable – if the medicine I am reviewing is for colorectal cancer then is a QLQ-C30 model derived from oesophageal cancer patients and validated on breast cancer patients applicable or not?  Should I be influenced by where the sample of patients for the mapping came from – for example, one of the QLQ-C30 studies was in Greek cancer patients, so should I discount it because I work in Scotland?  The most advice I could find was that patients in the original mapping study should have similar QLQ-C30 scores to the ones receiving the treatment I am interested in, but what does that mean in practice?  With six different mapping studies (at least) I can’t even use my normal trick of running a sensitivity analysis as there seems a good chance at least one of the algorithms will give a different answer to the base case.  Ideally I’d like to pick the one with the most robust statistical method, but I don’t think there is currently guidance to help me do that.

Summary
Our approach to mapping seems to be evolving in an ad hoc way.  Some bits of the jigsaw are available but there are a lot of gaps.  I’ve started to piece some of them together but would like to hear from anyone who thinks I’ve got anything wrong or who can fill in the gaps.
Thoughts on a database of mapping studies are very welcome.  (Stop press: I’ve just found another review of published algorithms:
‘Comparing the Incomparable? A Systematic Review of Competing Techniques for Converting Descriptive Measures of Health Status into QALY-Weights’  Duncan Mortimer and Leonie Segal  Med Decis Making 2008; 28; 66.)