π Normalized Discounted Cumulative Gain (NDCG)
Descriptionβ
< What is it? >β
-
Meaning: Normalized Discounted Cumulative Gain (NDCG) measures how well a ranked list places relevant items near the top. It considers both how relevant each item is and where it appears.
-
Example: suppose three movies have these relevance labels for a user:
Movie Relevance A 3 β highly relevant B 1 β somewhat relevant C 0 β irrelevant Ranking Order Quality Ideal A β B β C Most relevant movie first Worse C β B β A Most relevant movie last NDCG gives the ideal ranking 1.0 and the worse ranking a lower score.
Key pointsβ
< Gain, discount, and normalization >β
-
Gain: reward relevant items. This page uses the common gain formula , where is a nonnegative relevance label.
-
Discounted: reduce the reward for items appearing farther down the list. For the top results:
Here, is the position, starting at 1, and is the label of the item at that position. Sum only available positions when fewer than items are returned.
-
Normalized: divide by the best possible score for the evaluated items' relevance labels:
IDCG is the DCG of the ideal ordering: sort the evaluation set by relevance before taking the top . TensorFlow Ranking documents this normalization and configurable gain and discount functions.
-
Example: using the movie labels above:
Ranking DCG@3 NDCG@3 A β B β C 7.631 1.000 C β B β A 4.131 0.541
< Using NDCG in recommendations >β
- Evaluate each request:
NDCG@10evaluates the ordering of the first 10 recommendations. Calculate it separately for each recommendation request, then average across requests. - Separate labels from predictions: relevance labels come from evaluation data, such as ratings, human judgments, or defined engagement grades. Predicted scores determine the order; they are not the relevance labels used to calculate NDCG.
- Keep the evaluation setup consistent: use the same cutoff, gain formula, candidate-set policy, and tie-breaking rule when comparing models. Evaluating only retrieved candidates measures their ordering; evaluating against the broader catalog can also reflect missed relevant items.
- Handle zero ideal gain: if all relevance labels are zero, the ratio is undefined. The implementation below assigns
0.0; document this convention because evaluation tools may handle these requests differently.
Comparisonβ
< NDCG vs. Mean Reciprocal Rank (MRR) >β
-
What each metric rewards:
Metric What it rewards MRR Placing the first relevant item near the top NDCG Placing multiple relevant items near the top, accounting for relevance grades NDCG can use graded or binary relevance. MRR requires a yes/no relevance decision and ignores relevant items after the first match.
Implementationβ
< Calculate NDCG with Python >β
-
Input: item relevance is
A: 3,B: 1,C: 0. Each score dictionary contains model scores for these same items. This example uses exponential gain and preserves dictionary order to break score ties.from math import log2relevance = {"A": 3, "B": 1, "C": 0}def dcg(labels, k):return sum((2 ** rel - 1) / log2(position + 1)for position, rel in enumerate(labels[:k], start=1))def evaluate(scores, k):order = sorted(scores, key=scores.get, reverse=True)labels = [relevance[item] for item in order]actual = dcg(labels, k)ideal = dcg(sorted(relevance.values(), reverse=True), k)ndcg = actual / ideal if ideal > 0 else 0.0return order, actual, ndcgfor name, scores in [("Ideal", {"A": 0.9, "B": 0.5, "C": 0.1}),("Worse", {"A": 0.1, "B": 0.5, "C": 0.9}),]:order, actual, ndcg = evaluate(scores, k=3)print(f"{name}: {order}, DCG@3={actual:.3f}, NDCG@3={ndcg:.3f}") -
Verified output:
Ideal: ['A', 'B', 'C'], DCG@3=7.631, NDCG@3=1.000Worse: ['C', 'B', 'A'], DCG@3=4.131, NDCG@3=0.541
Related ideasβ
- Mean Reciprocal Rank (MRR) focuses on the first relevant result.
- Recommendation System explains retrieval and ranking.
- Extreme Gradient Boosting (XGBoost) includes an
XGBRankerexample usingrank:ndcg.