Skip to main content

๐Ÿ“ Mean Reciprocal Rank (MRR)

Descriptionโ€‹

< What is it? >โ€‹

  • Meaning: Mean Reciprocal Rank (MRR) measures how early the first relevant item appears in ranked results, averaged across requests.

  • Example: each request returns the following results:

    RequestRanked resultsFirst relevant positionReciprocal rank
    AliceIrrelevant โ†’ Relevant โ†’ Relevant21/2 = 0.5
    BobRelevant โ†’ Irrelevant โ†’ Relevant11/1 = 1.0
    CarolIrrelevant โ†’ Irrelevant โ†’ Relevant31/3 โ‰ˆ 0.333

    Average these values:

    MRR=1/2+1+1/33โ‰ˆ0.611MRR=\frac{1/2+1+1/3}{3}\approx0.611

Key pointsโ€‹

< Reciprocal rank and its mean >โ€‹

  • Reciprocal rank (RR): for each request with a relevant result:

    RR=1positionย ofย theย firstย relevantย itemRR=\frac{1}{\text{position of the first relevant item}}

    Positions start at 1. If no relevant item is found in the evaluated results, the reciprocal rank is 0. This is the convention described by NIST's TREC reciprocal-rank definition.

  • Mean across requests: for QQ requests:

    MRR=1Qโˆ‘q=1QRRqMRR=\frac{1}{Q}\sum_{q=1}^{Q}RR_q

    Higher is better. An MRR of 1.0 means the first result is relevant for every request.

  • Cutoff: MRR@10 only searches the first 10 results. A request with no relevant result there contributes 0. TensorFlow Ranking's MRR metric supports a top-result cutoff.

< Define what counts as relevant >โ€‹

  • Binary relevance: MRR treats relevance as a yes/no decision. Define the rule before evaluationโ€”for example, a movie rated at least 4 stars counts as relevant.
  • Only the first match matters: once the first relevant item is found, later relevant items do not affect that request's reciprocal rank.
  • Keep all evaluated requests: include requests with no relevant result as zeros when calculating the mean. Predicted scores determine result order; evaluation labels determine whether each result is relevant.

Comparisonโ€‹

< MRR vs. Normalized Discounted Cumulative Gain (NDCG) >โ€‹

  • Different priorities: MRR focuses on reaching the first relevant result quickly. NDCG evaluates the placement of multiple relevant results and can account for different relevance grades. See the NDCG comparison table.

Implementationโ€‹

< Calculate MRR with Python >โ€‹

  • Input: lists contain relevance labels already ordered by the model's predicted ranking. 1 means relevant and 0 means irrelevant. This example needs only Python's standard library.

    def reciprocal_rank(labels, k):
    for position, relevant in enumerate(labels[:k], start=1):
    if relevant == 1:
    return 1.0 / position
    return 0.0

    requests = {
    "Alice": [0, 1, 1],
    "Bob": [1, 0, 1],
    "Carol": [0, 0, 1],
    }

    values = []
    for name, labels in requests.items():
    rr = reciprocal_rank(labels, k=3)
    values.append(rr)
    print(f"{name}: RR@3={rr:.3f}")

    print(f"MRR@3={sum(values) / len(values):.3f}")
    print(f"No relevant result: RR@3={reciprocal_rank([0, 0, 0], 3):.3f}")
    print(f"Carol with cutoff 2: RR@2={reciprocal_rank(requests['Carol'], 2):.3f}")
  • Verified output:

    Alice: RR@3=0.500
    Bob: RR@3=1.000
    Carol: RR@3=0.333
    MRR@3=0.611
    No relevant result: RR@3=0.000
    Carol with cutoff 2: RR@2=0.000

    The mean here covers Alice, Bob, and Carol. The final two lines separately demonstrate the no-match and cutoff cases.

Referenceโ€‹