๐ Mean Reciprocal Rank (MRR)
Descriptionโ
< What is it? >โ
-
Meaning: Mean Reciprocal Rank (MRR) measures how early the first relevant item appears in ranked results, averaged across requests.
-
Example: each request returns the following results:
Request Ranked results First relevant position Reciprocal rank Alice Irrelevant โ Relevant โ Relevant 2 1/2 = 0.5 Bob Relevant โ Irrelevant โ Relevant 1 1/1 = 1.0 Carol Irrelevant โ Irrelevant โ Relevant 3 1/3 โ 0.333 Average these values:
Key pointsโ
< Reciprocal rank and its mean >โ
-
Reciprocal rank (RR): for each request with a relevant result:
Positions start at 1. If no relevant item is found in the evaluated results, the reciprocal rank is 0. This is the convention described by NIST's TREC reciprocal-rank definition.
-
Mean across requests: for requests:
Higher is better. An MRR of 1.0 means the first result is relevant for every request.
-
Cutoff:
MRR@10only searches the first 10 results. A request with no relevant result there contributes 0. TensorFlow Ranking's MRR metric supports a top-result cutoff.
< Define what counts as relevant >โ
- Binary relevance: MRR treats relevance as a yes/no decision. Define the rule before evaluationโfor example, a movie rated at least 4 stars counts as relevant.
- Only the first match matters: once the first relevant item is found, later relevant items do not affect that request's reciprocal rank.
- Keep all evaluated requests: include requests with no relevant result as zeros when calculating the mean. Predicted scores determine result order; evaluation labels determine whether each result is relevant.
Comparisonโ
< MRR vs. Normalized Discounted Cumulative Gain (NDCG) >โ
- Different priorities: MRR focuses on reaching the first relevant result quickly. NDCG evaluates the placement of multiple relevant results and can account for different relevance grades. See the NDCG comparison table.
Implementationโ
< Calculate MRR with Python >โ
-
Input: lists contain relevance labels already ordered by the model's predicted ranking.
1means relevant and0means irrelevant. This example needs only Python's standard library.def reciprocal_rank(labels, k):for position, relevant in enumerate(labels[:k], start=1):if relevant == 1:return 1.0 / positionreturn 0.0requests = {"Alice": [0, 1, 1],"Bob": [1, 0, 1],"Carol": [0, 0, 1],}values = []for name, labels in requests.items():rr = reciprocal_rank(labels, k=3)values.append(rr)print(f"{name}: RR@3={rr:.3f}")print(f"MRR@3={sum(values) / len(values):.3f}")print(f"No relevant result: RR@3={reciprocal_rank([0, 0, 0], 3):.3f}")print(f"Carol with cutoff 2: RR@2={reciprocal_rank(requests['Carol'], 2):.3f}") -
Verified output:
Alice: RR@3=0.500Bob: RR@3=1.000Carol: RR@3=0.333MRR@3=0.611No relevant result: RR@3=0.000Carol with cutoff 2: RR@2=0.000The mean here covers Alice, Bob, and Carol. The final two lines separately demonstrate the no-match and cutoff cases.
Related ideasโ
- Normalized Discounted Cumulative Gain (NDCG) evaluates multiple relevant results and their grades.
- Recommendation System explains the ranking pipeline.