Concept

Why can a positive medical test still mean a low probability of having a disease?

Stephen Davies, Ph.D. Version 2.2.2 Through Discrete Mathematics A Cool Brisk Walk / Chapter 1

"A simple and commonly cited example is that of interpreting medical exam results for the presence of a disease. If your doctor recommends that you undergo a blood test to see if you have some rare condition, you might test positive or negative. But suppose you do indeed test positive. What’s the probability that you actually have the disease? That, of course, is the key point. In symbols, we’re looking for Pr( D | T ), where D is the event that you actually have the disease in question, and T is the event that you test positive for it. But this is hard to approximate with available data. For one thing, most people who undergo this test don’t test positive, so we don’t have a ton of examples of event T occurring whereby we could count the times D also occurred. But worse, it’s hard to tell whether a patient has the disease, at least before advanced symptoms develop — that, after all, is the purpose of our test!\n\nBayes’ Theorem, however, lets us rewrite this as: Pr ( D | T ) = Pr ( T | D ) Pr ( D ) Pr ( T ) . Now we have Pr( D | T ), the hard quantity to compute, in terms of three things we can get data for. To estimate Pr( T | D ), the probability of a person who has the disease testing positive, we can administer the test to unfortunate patients with advanced symptoms and count how many of them test positive. To estimate Pr( D ), the prior probability of having the disease, we can divide the number of known cases by the population as a whole to find how prevalent it is. And getting Pr( T ), the probability of testing positive, is easy since we know the results of the tests we’ve administered. In numbers, suppose our test is 99% accurate — i.e. , if someone actually has the disease, there’s a .99 probability they’ll test positive for it, and if they don’t have it, there’s a .99 probability they’ll test negative. Let’s also assume that this is a very rare disease: only one in a thousand people contracts it. When we interpret those numbers in light of the formula we’re seeking to populate, we realize that Pr( T | D ) = .99, and Pr( D ) = 1 1000 . The other quantity we need is Pr( T ), and we’re all set. But how do we figure out Pr( T ), the probability of testing positive? Answer: use the Law of Total Probability. There are two different “ways” to test positive: (1) to actually have the disease, and (correctly) test positive for it, or (2) to not have the disease, but incorrectly test positive for it anyway because the test was wrong. Let’s compute this: Pr ( T ) = Pr ( T | D ) Pr ( D ) + Pr ( T | D ) Pr ( D ) = . 99 · 1 1000 + . 01 · 999 1000 = . 00099 + . 00999 = . 01098 (4.1)\n\nSee how that works? If I do have the disease (and there’s a 1 in 1,000 chance of that), there’s a .99 probability of me testing positive. On the other hand, if I don’t have the disease (a 999 in 1,000 chance of that), there’s a .01 probability of me testing positive anyway. The sum of those two mutually exclusive probabilities is .01098. Now we can use our Bayes’ Theorem formula to deduce: Pr ( D | T ) = Pr ( T | D ) Pr ( D ) Pr ( T ) = . 99 · 1 1000 . 01098 ≈ . 0902 Wow. We tested positive on a 99% accurate medical exam, yet we only have about a 9% chance of actually having the disease! Great news for the patient, but a head-scratcher for the math student. How can we understand this? Well, the key is to look back at that Total Probability calculation in equation 4.1. Remember that there were two ways to test positive: one where you had the disease, and one where you didn’t. Look at the contribution to the whole that each of those two probabilities produced. The first was .00099, and the second was .00999, over ten times higher. Why? Simply because the disease is so rare. Think about it: the test fails once every hundred times, but a random person only has the disease once every thousand times. If you test positive, it’s far more likely that the test screwed up than that you actually have the disease, which is rarer than blue moons."

Related Ideas

Why can a positive medical test still mean a low probability of having a disease? | Stephen Davies, Ph.D. Version 2.2.2 Through Discrete Mathematics A Cool Brisk Walk | Bifalgorithm | Bifalgorithm