π Confusion Matrix
Descriptionβ
< What is it? >β
-
A confusion matrix counts how a classification model's predictions compare with actual labels. It shows correct predictions and the types of mistakes the model makes.
-
Example: a model predicts whether a user will click a recommended movie. Here, click is the positive class, and the matrix summarizes 100 impressions:
Actual \ Predicted Click No click Click 40 β True Positive (TP) 10 β False Negative (FN) No click 20 β False Positive (FP) 30 β True Negative (TN)
Key pointsβ
< Reading the four outcomes >β
- True positive (TP): predicted a click, and the user clicked.
- False positive (FP): predicted a click, but the user did not click.
- False negative (FN): predicted no click, but the user clicked.
- True negative (TN): predicted no click, and the user did not click.
- Naming convention: βtrue/falseβ describes whether the prediction was correct; βpositive/negativeβ describes the predicted class. Check the axis labels because some displays reverse actual and predicted. Google's confusion matrix guide illustrates the alternative orientation.
< Metrics calculated from the matrix >β
-
Metric calculations: using the counts above:
Metric Calculation Meaning Accuracy How often was the prediction correct? Precision Of predicted clicks, how many actually happened? Recall Of actual clicks, how many did the model identify? F1 score Harmonic mean of precision and recall -
Interpret metrics together: accuracy can hide poor performance on a rare class. Choose metrics based on the cost of missed positives and false alarms. If a formula's denominator is zero, its value is undefined; explicitly document how your evaluation tool handles it. See Google's classification metrics guide.
< Thresholds and multiple classes >β
- Probability threshold: convert predicted probabilities to class labels before counting outcomes. For example, predict βclickβ when its probability is at least
0.5. Changing the threshold can change the matrix; select the threshold on validation data and report it with the results. - Multiple classes: a single-label classifier with classes has a matrix. With the same class order on both axes, correct predictions lie on the diagonal; off-diagonal cells show which classes are confused.
Related ideasβ
- ROCβAUC compares true-positive and false-positive rates across thresholds.
- Confidence Calibration examines whether predicted probabilities match observed frequencies.
- Train, Validation, and Test Sets explains how to separate model selection from final evaluation.