Skip to main content

πŸ“ Confusion Matrix

Description​

< What is it? >​

  • A confusion matrix counts how a classification model's predictions compare with actual labels. It shows correct predictions and the types of mistakes the model makes.

  • Example: a model predicts whether a user will click a recommended movie. Here, click is the positive class, and the matrix summarizes 100 impressions:

    Actual \ PredictedClickNo click
    Click40 β€” True Positive (TP)10 β€” False Negative (FN)
    No click20 β€” False Positive (FP)30 β€” True Negative (TN)

Key points​

< Reading the four outcomes >​

  • True positive (TP): predicted a click, and the user clicked.
  • False positive (FP): predicted a click, but the user did not click.
  • False negative (FN): predicted no click, but the user clicked.
  • True negative (TN): predicted no click, and the user did not click.
  • Naming convention: β€œtrue/false” describes whether the prediction was correct; β€œpositive/negative” describes the predicted class. Check the axis labels because some displays reverse actual and predicted. Google's confusion matrix guide illustrates the alternative orientation.

< Metrics calculated from the matrix >​

  • Metric calculations: using the counts above:

    MetricCalculationMeaning
    Accuracy(TP+TN)/(TP+TN+FP+FN)=70%(TP + TN) / (TP + TN + FP + FN) = 70\%How often was the prediction correct?
    PrecisionTP/(TP+FP)β‰ˆ66.7%TP / (TP + FP) \approx 66.7\%Of predicted clicks, how many actually happened?
    RecallTP/(TP+FN)=80%TP / (TP + FN) = 80\%Of actual clicks, how many did the model identify?
    F1 score2TP/(2TP+FP+FN)β‰ˆ72.7%2TP / (2TP + FP + FN) \approx 72.7\%Harmonic mean of precision and recall
  • Interpret metrics together: accuracy can hide poor performance on a rare class. Choose metrics based on the cost of missed positives and false alarms. If a formula's denominator is zero, its value is undefined; explicitly document how your evaluation tool handles it. See Google's classification metrics guide.

< Thresholds and multiple classes >​

  • Probability threshold: convert predicted probabilities to class labels before counting outcomes. For example, predict β€œclick” when its probability is at least 0.5. Changing the threshold can change the matrix; select the threshold on validation data and report it with the results.
  • Multiple classes: a single-label classifier with KK classes has a KΓ—KK \times K matrix. With the same class order on both axes, correct predictions lie on the diagonal; off-diagonal cells show which classes are confused.

Reference​