{
  "abstract": "Abstract in Full - Workshops & Seminars It is customary to set normal or reference ranges of numerical test results based on two standard deviations of population test results (or ROC curves). The overall performance of a test is then assessed by using the indices of sensitivity and specificity. However, clinicians also learn by experience to interpret individual numerical values. These numerical values usually represent degrees of disease severity, each result mapping to the estimated probability of an outcome. Such personalised probabilities are important for shared decision making and especially if this is done using decision analysis.Normal ranges are sometimes used as simple decision thresholds. For example a ‘high’ albumin excretion rate (AER) of 20mcg/min or above may be used as a criterion for the diagnosis of ‘albuminuria’, which in turn triggers a treatment strategy. However a ‘normal’ AER of 19mcg/min may trigger reassurance, and yet both results are essentially the same. The probability of nephropathy conditional on an abnormal AER will be the average of all probabilities of nephropathy values in a study with patients with AER values from 20 to 200mcg/min (e.g. 30/196 = 0.153) but the estimated probability of nephropathy conditional on an AER of precisely 20mcg/min will be about 0.01 reducing to 0.005 on treatment. To allocate to such a patient a probability of 0.153 of nephropathy, reducing to 0.077 on treatment without taking into account the mildness of the disease would greatly exaggerate the risk of nephropathy and the expected benefit from treatment, leading to over-treatment.Instead of only using a diagnostic or decision threshold based on two standard deviations or an ROC curve, we should also consider creating curves that estimate the probability of an outcome conditional on each numerical value of a test result for those on treatment and placebo (or no treatment). A preliminary step would be to create histograms conditional on narrow ranges of test results. For example in one study only 1 out of 77 = 1.3% of patients with an AER between 20 and 40mc/min develop nephropathy on placebo, whereas on treatment with irbesartan 150mg daily, none of 58 patients developed nephropathy and on irbesartan 300mg daily, 1/68 developed nephropathy. Therefore, in the 77/196 = 39% of patients in the control limb in this range, there is little prospect of developing nephropathy even without treatment so that adding treatment makes little difference. This represents over diagnosis and overtreatment in this AER range between 20 and 40mcg/min.In order to take disease severity into account, there are well established methods for constructing the required curves using logistic regression; the steeper the curve the better the test. The required data are usually recorded during clinical trials but not used in this way currently. A finding that reflects disease severity very well or that can be used to assess the speed of the disease’s progression will reduce over-diagnosis and over-treatment. Knowledge of the disease can be used to design better tests with steeper sigmoid curves. This seminar will explore these issues in detail with worked examples.Objectives The objectives are to reduce the frequency of over-diagnosis by identifying and excluding subgroup that are unlikely to benefit from a diagnosis and treatment. Current RCT results are reported as the average treatment effect (ATE) usually in the form of the average absolute risk reduction (ARR). However there may be hidden heterogeneity of treatment effects (HTE) in subjects of a RCT, this being known as the reference class problem. It can be addressed in a number of ways, including risk modelling using external data independent of an RCT. The approach taken here is different. The object is to prospectively assess the disease severity of subjects in a RCT as a predictor of heterogeneity of treatment effect in order to establish pragmatic diagnostic test thresholds to reduce the frequency of over-diagnosis and over-treatment and also to estimate probabilities of outcomes conditional on each numerical value of a test result.Method The entry criteria for the IRMA2 RCT on Diabetic Albuminuria (DA) included diagnostic criteria to exclude those with the possible differential diagnoses of DA (e.g. UTI and hypertension by controlling the blood pressure in all patients) and excluding those at increased risk of adverse treatment effects (e.g. pregnancy). The albumin excretion rate (AER) was used as an indication of DA severity. The upper 2 standard deviations of the log(AER) in a healthy population is 20mcg/min so it was assumed that if the mean of 3 AERs in someone with Type 2 diabetes mellitus exceeded 20mcg/min, they had DA. Patients were randomised into placebo or irbesartan 150mg daily or irbesartan 300mg daily and those with two AERs >200mcg/min 2 years later were regarded as having ‘diabetic nephropathy’. Histograms and curves were created to display the probability of nephropathy after 2 years conditional on the baseline AER for each trial limb.Results In those people with a pre-randomisation AER range of 20 to 40mg/min, only 1/77 developed nephropathy on placebo within 2 years. On treatment with irbesartan 150mg daily, none of the 58 patients developed nephropathy and on irbesartan 300mg daily, 1/68 developed nephropathy. Therefore, in the 39% of patients on placebo in this range, there is little prospect of developing nephropathy even without treatment so that adding treatment made little difference. The three curves showing the probability of nephropathy for each possible AER in patients on placebo or irbesartan 150mg daily or irbesartan 300mg daily provide a personalised probability estimate of nephropathy for a patient with a particular AER result for use in shared decision making (SDM). The curve showing the ARR at each AER will also aid SDM.Conclusions Basing diagnostic thresholds on the upper or lower 2 standard deviations of a diagnostic test result is flawed. The threshold for a diagnosis should be placed at a point where the probability of some adverse outcome without treatment is significant. The threshold should be regarded as a result beyond which at least some patients would accept treatment in the absence of its adverse effects. Therefore, RCTs should be interpreted in the light of disease severity and other factors that influence the probability of an adverse outcome. This applies particularly to screening tests with numerical results such as the PSA, aortic expansion, tumour size and change, etc. Research should focus on better methods of assessing change in disease severity.",
  "authors": [
    {
      "affiliations": [
        "Aberystwyth University, Aberystwyth, United Kingdom"
      ],
      "name": "Huw Llewelyn"
    }
  ],
  "title": "156 Reducing over-diagnosis and over-treatment by creating continuous functions that estimate the probability of an outcome conditional on each possible test result that represents an individual patient’s disease severity for shared decision making",
  "uid": "1506fc34-16eb-5051-afdb-a4596285e380"
}
