Evaluation & quality

Precision and Recall

Two measures of a system that picks items out of a larger set: precision is the share of the items it picked that were right, and recall is the share of the right items that it picked.

10,000 TRANSACTIONS · 100 FRAUDULENT · illustrativeFlagged: 80Not flagged: 9,920Fraudactually 100TP60caughtFN40missedLegitimateactually 9,900FP20false alarmTN9,880correctly passedPrecision = 60 / 80 = 75%of the 80 flagged, 60 were fraudRecall = 60 / 100 = 60%of the 100 frauds, 60 were flaggedAccuracy = 99.4%flagging nothing would score 99%F1 = 0.67harmonic mean of the twoPrecision shows the false alarms, recall the misses; when cases are rare, accuracy hides both.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/precision-and-recall

In plain terms

A net is cast for tuna. Precision asks how much of the catch is tuna. Recall asks how much of the tuna in the water ended up in the net. A fine-meshed net dragged across the whole bay catches every tuna and a great deal else: high recall, low precision. A single line with the right bait brings up only tuna, and very few of them: high precision, low recall.

Why it matters

Every system that flags, filters, retrieves or classifies makes two kinds of mistake, false alarms and misses, and they rarely cost the same. A missed fraud and a wrongly blocked customer carry different prices; so do a missed tumour and an unnecessary scan. Precision and recall keep the two apart, and the acceptable balance is a business decision that should be made explicitly. Raising one usually lowers the other. A single “accuracy” figure hides the trade-off and can look excellent on a system that misses most of what matters.

Example

A payment company screens 10,000 transactions, of which 100 are fraudulent. The model flags 80. Of these, 60 are fraud and 20 are legitimate; 40 frauds pass unnoticed. Precision is 60 out of 80, or 75%. Recall is 60 out of 100, or 60%. Accuracy is 99.4%, and a model that flagged nothing at all would still reach 99%.

Most often confused with

Precision and Recall vs. Accuracy

Precision and RecallShow false alarms and misses separately
AccuracyOne figure: the share of all decisions that were correct

Accuracy counts every correct decision alike, including the easy ones. When the cases that matter are rare, as with fraud, defects or disease, a system can be highly accurate by ignoring them. Precision and recall look only at the cases that were flagged and the cases that should have been, which is where the cost lies. On balanced data the three figures tell a similar story.

Origin: Both measures come from information retrieval research of the 1950s; the F-measure derives from work published by C. J. van Rijsbergen in 1979.

Under the hood

The four outcomes form the confusion matrix: true positives (TP), false positives (FP), false negatives (FN) and true negatives (TN). Precision = TP / (TP + FP). Recall = TP / (TP + FN), also called sensitivity. F1 is the harmonic mean of the two, 2PR / (P + R), and is high only when both are; in the example, with precision 0.75 and recall 0.60, F1 is 0.67. The F-beta variant gives recall more or less weight. Most classifiers output a score, and the decision threshold sets the balance: a lower threshold raises recall and lowers precision. The precision-recall curve shows every threshold at once. With several classes, the measures are averaged per class (macro) or over all decisions (micro). In retrieval the same measures are applied to the top k results, as precision@k and recall@k. In language-model applications they describe retrieval quality, guardrail filters, extraction tasks and citation checks.

Written by Mehmet Erkek · Last updated: