Slow thinking required

What is a false positive?

Slow thinking required

File this in the bucket of items requiring "slow thinking" ala Kahneman.

There is an active outbreak of the Cyclospora parasite in multiple U.S. states, which the formidable epidemiologists at the CDC are investigating. I featured their work in Chapter 2 of Numbers Rule Your World (link). The scientists have narrowed the suspected contaminant to lettuce, which is a common culprit. They then announced a positive test result for Cyclospora in a specific batch of lettuce at Taylor Farms. This is part of the "traceback" investigation. Last week, they rescinded that announcement, calling it a "false positive".

More info here

Taylor Farms is a large producer that supplies Taco Bells and Walmart, amongst other clients. The lettuce in question is grown in Mexico and imported to U.S. As is typical in these situations, these farms voluntarily recall large amounts of produce.

In a short paragraph, the FDA stated:

Due to the complexity in detection of Cyclospora, FDA laboratory experts re-reviewed the sample results and have concluded that the finding does not represent true amplification and should be considered a false positive.

If you just glance at this, you'd miss the problem with this retraction.


I'm referring to the use of the term "false positive".

The definition of a false positive in a diagnostic test is an incorrect positive finding. Said differently, the correct answer should have been negative. I emphasize should have been.

By definition, a false positive is a counterfactual quantity. It is not observable in real life. In any single real-world diagnostic test, we have no right to know whether a positive finding is true or false.

Let's assume I'm an idiot, and in fact, you know that this particular positive test result is erroneous. This implies that you know the answer. If you did know the answer, why would you need to conduct a diagnostic test? If the FDA knew that batch of lettuce from Taylors Farms was not contaminated, they would not have needed to test for it.


What might have happened? Someone could have made a mistake in conducting the test. Repeating the test or reviewing the process of the test might have uncovered an error. But that is not a "false positive"! In that case, one should just correct the record and announce that our lab made a mistake, and in fact, the batch of lettuce tested negative for Cyclospora.


Why is there so much confusion when it comes to test accuracy?

It's fundamental statistics. We always use a test of a few samples to generalize the finding to a bigger population. In this case, if the positive test result would have stood, one would reasonably infer that the lettuce from that farm was a likely source of Cyclospora contamination. It's not about the one test sample but the larger batch.

Moving from the specific to the general always entails uncertainty. Statisticians cannot remove the uncertainty but we try to quantify it. For a diagnostic test, we quantify it by a process of "calibration". This is typically rated by the creator of the test.

You'd start with two groups of samples. One group is known to be contaminated; the other group is known to be uncontaminated. You apply the test to both groups, and you can directly estimate the probability that the test result is correct. You'd get the true positive and true negative rates, which also yields the false negative and false positive rates.

This all happens in the lab of the test designer. This isn't what happens in the field, after the test has gone to market. Any user of the test is someone who does not know the answer. If the test has a 1% false-positive rate, the test vendor is claiming that out of 100 known positive samples, they expect 99 to come back positive, and 1 to be negative. No one knows if that one positive result just observed is that potential mistake or not.

In practice, we are interested in the (Bayesian) inverse. Given that the test result is positive, what is the chance that the finding is incorrect? That's called the "positive predictive value". And it is not 1%. Kaheman-style fast thinking might lead us to think it must be 1%. Slow thinking is required here.