Skip to main content

📝 Recommendation System

Description​

< What is it? >​

A recommendation system selects and orders items that are likely to be useful or interesting to a user. Items can be movies, products, articles, or other content. Recommendations can depend on user preferences, previous interactions, item attributes, and the current context.

A common design separates retrieval, which finds a manageable set of candidates, from ranking, which scores those candidates in more detail. This page focuses on two-tower retrieval followed by a ranking model.

Definition​

< Impression >​

  • Meaning: an item being shown to a user — a single instance of a recommendation appearing on screen.
  • Example: a movie poster appears in a Netflix row → 1 impression.

< Conversion Rate (CVR) >​

  • Meaning: conversion rate measures how often users complete a desired action, such as purchasing a recommended product, subscribing, or starting a recommended movie. The product team defines which action counts as a conversion.

  • Calculation: specify the denominator—such as recommendation clicks, impressions, sessions, or users. For a post-click conversion rate:

    Post-click CVR=Clicks that lead to a conversionEligible recommendation clicks×100%\text{Post-click CVR}=\frac{\text{Clicks that lead to a conversion}}{\text{Eligible recommendation clicks}}\times100\%

    This definition counts each click at most once. Define the attribution window too—for example, a purchase within 24 hours after the click. Some platforms count multiple conversions per interaction, so check the counting rule when comparing reports. Google Ads' definition uses conversions per eligible ad interaction.

  • Example: recommended products receive 1,000 impressions and 100 clicks. Of those clicks, 5 lead to a purchase within the chosen window:

    MetricCalculationResult
    Click-through rate (CTR)100 clicks / 1,000 impressions10%
    Post-click conversion rate5 converting clicks / 100 clicks5%
    Impression-to-purchase conversion rate5 converting impressions / 1,000 impressions0.5%

    Here, each click belongs to a distinct impression, and each purchase is attributed to one click. A click shows interest; the conversion records the chosen downstream action. Always state which conversion rate you report.

Key points​

< Retrieval and ranking pipeline >​

Retrieval must search a large catalog efficiently. Ranking operates on the smaller candidate set, allowing a more expensive model to examine each user–item pair. The numbers below are illustrative:

1 million movies
|
v
Two-tower retrieval
|
v
500 candidate movies
|
v
Ranking model scores each user–movie pair
|
v
Top 10 movies
  • Example: a movie service retrieves 500 movies using embedding similarity, then sorts them by predicted user rating to select 10 recommendations. See TensorFlow's ranking tutorial for a rating-prediction model.

< Collaborative filtering with matrix factorization >​

  • Learn embeddings from ratings: matrix factorization learns user and movie vectors from patterns in observed ratings. In the classic Netflix Prize setting, the objective was to predict how many stars a user would give an unrated movie. This historical approach is described by Koren, Bell, and Volinsky (2009).

  • Example: start with a user–movie rating matrix. Suppose the observed ratings are:

    UserMovie AMovie BMovie CMovie D
    Alice54?1
    Bob4?21
    Carol125?

    ? means unknown, not a zero-star rating. Similar rating patterns provide evidence about shared preferences.

  • Create two trainable embedding tables: choose an embedding dimension dd, such as 32 for a real dataset, and initialize the vectors with small random values.

    TableShapeEach row represents
    UUNumber of users × ddOne user's embedding
    VVNumber of movies × ddOne movie's embedding

    In the basic model without biases, learn values whose products approximate the known ratings:

    R≈UV⊤R \approx UV^\top

    This represents a large rating matrix using two smaller matrices when dd is much smaller than the user and movie counts. The coordinates are latent factors learned from behavior; they need not correspond to named genres. See Google's matrix factorization guide.

  • Predict a rating: a common extension adds user and movie biases:

    r^ui=μ+bu+bi+Uu⊤Vi\hat r_{ui}=\mu+b_u+b_i+U_u^\top V_i
    TermMeaning
    μ\muOverall average training rating
    bub_uThe user's tendency to rate generously or harshly
    bib_iThe movie's tendency to receive higher or lower ratings
    Uu⊤ViU_u^\top V_iThe personalized user–movie interaction
  • Example: calculate the training error. If Alice gave Movie A 5 stars and the model predicts 3.2, the squared error is:

    (5−3.2)2=3.24(5-3.2)^2=3.24

    Training minimizes errors on observed ratings, plus regularization to discourage overfitting. Gradient descent updates both embedding tables and the biases. Repeated updates connect users and movies through their shared rating history.

  • Predict missing entries: after training, retrieve Alice's vector and Movie C's vector, calculate the predicted rating, and use it to rank candidate movies. The system can score selected pairs without materializing the entire predicted matrix.

  • Connect this to profile-based towers: basic matrix factorization uses user ID → embedding lookup and movie ID → embedding lookup. Age, genre, cast, and descriptions are optional extensions. The JSON profile example instead computes embeddings from those fields through feature-based towers.

  • Use the trained scoring function: this rating model uses a dot product plus biases. Replacing it with cosine similarity changes the scoring function because cosine removes vector magnitude.

< User and item towers >​

  • A tower is a model, typically a neural network that converts user or item features into a user or item embedding. A profile describes the input data or a learned representation. See Embeddings for the input flow and the distinction between profiles, towers, and embeddings.

In dot-product retrieval, both embeddings have the same dimension and are trained to be compatible:

retrievalScore⁡(u,i)=fθ(xu)⊤gϕ(xi)\operatorname{retrievalScore}(u,i) = f_\theta(x_u)^\top g_\phi(x_i)

Here, xux_u and xix_i are user and item features, and fθf_\theta and gϕg_\phi are their towers. Item embeddings can be computed ahead of time and indexed for fast retrieval. At request time, the user tower produces a vector used to search that index. TensorFlow's retrieval tutorial demonstrates this design.

  • Learn compatibility from interaction data: millions of labeled impressions can train the user and item MLP towers together. For click prediction, a clicked impression supplies label 1, and an unclicked impression supplies label 0.

    For each user–item pair:

    1. The user MLP and item MLP produce embeddings.
    2. Their similarity produces a compatibility score.
    3. The loss compares that score with the interaction label.
    4. Backpropagation computes gradients, and an optimizer updates both MLPs and any trainable embedding tables.

    With a suitable objective, training encourages clicked pairs to receive higher similarity scores and negative pairs to receive lower scores. For L2-normalized embeddings, higher cosine similarity also means smaller L2 distance. The towers therefore learn representations that make user–item pairs likely to receive clicks more compatible. Google's Machine Learning Crash Course explains how task objectives shape learned embeddings.

    Labels can be individual click/no-click events; an aggregated click-through rate is not required. Unclicked impressions provide noisy negative evidence, and an item that was never shown is not automatically a negative example. The resulting similarity score is not automatically a calibrated click probability.

< What the ranking model predicts >​

A ranking model typically produces one score per candidate. The score can represent predicted click probability, purchase probability, a rating, or a learned relevance value. The training labels and loss determine its meaning; an arbitrary relevance score is not automatically a probability.

The ranker can jointly process user features, item features, and context, learning interactions beyond an embedding dot product. It can reuse retrieval embeddings or learn its own embeddings for the ranking objective.

Comparison​

< Models for ranking >​

ModelHow it ranksTypical use
Multilayer Perceptron (MLP)Combines user/item embeddings and other features through neural-network layersA straightforward neural recommendation ranker
Gradient-boosted trees, such as XGBoost with LambdaMARTLearns ranking scores from features such as retrieval similarity, popularity, and user–item interaction statisticsRanking with structured or tabular features
Deep & Cross Network (DCN)Combines deep layers with explicit feature interactionsRecommendations where combinations of user, item, and context features matter
Transformer cross-encoderProcesses a query and candidate text jointly to predict relevanceSearch and document reranking

An MLP is a useful starting point for understanding neural ranking. XGBoost offers a tree-based learning-to-rank implementation, while DCN explicitly models feature crosses. A text cross-encoder is particularly relevant when the task involves matching a query to candidate documents; applying it to personalized recommendations requires suitable user context and training data.

Concrete examples are available in TensorFlow's MLP tutorial, XGBoost's learning-to-rank guide, TensorFlow's DCN tutorial, and Sentence Transformers' reranking guide.

Implementation​

< Train a simple two-tower model from JSON profiles >​

  • What this example trains: two small MLPs, a shared genre embedding table, and an actor embedding table. User inputs contain genre preference and age; item inputs contain cast and genre. Both towers output three learned coordinates. The three output dimensions are a teaching choice, not three named profile fields.

  • Install and run: install PyTorch, then save the Python block as recommendation_demo.py and run python recommendation_demo.py. It runs on CPU and needs no external dataset or pretrained model.

    python -m pip install torch
  • Example: encode profiles by key, train on four explicitly supplied interaction labels, and retrieve the highest-scoring item for each user. The vocabulary dictionaries are fixed; the vectors returned by nn.Embedding are trainable.

    import json
    import torch
    from torch import nn
    from torch.nn import functional as F

    torch.manual_seed(42)

    users = json.loads('''[
    {"preferred_genre": "horror", "age": 25},
    {"preferred_genre": "history", "age": 40}
    ]''')
    items = json.loads('''[
    {"cast": ["Jason", "David"], "genre": "horror"},
    {"genre": "history", "cast": ["Michael"]}
    ]''')
    # Illustrative observed labels, not inferred from matching genre names.
    # Rows: users; columns: items. All four pairs are labeled in this toy data.
    clicked = torch.tensor([[1., 0.], [0., 1.]])
    genres = {"horror": 1, "history": 2} # 0 = unknown
    actors = {"Jason": 1, "David": 2, "Michael": 3}

    class TwoTower(nn.Module):
    def __init__(self):
    super().__init__()
    self.genre = nn.Embedding(len(genres) + 1, 4)
    self.actor = nn.Embedding(len(actors) + 1, 4)
    # nn.Sequential is a PyTorch container that chains layers in order.
    # Calling self.user_mlp(x) passes x through the first layer, then feeds
    # each layer's output into the next and returns the final output.
    # Here: Linear(5, 8) → ReLU → Linear(8, 3).
    self.user_mlp = nn.Sequential(
    nn.Linear(5, 8), # 4 genre components + 1 scaled age → 8 learned values.
    nn.ReLU(), # Replace negative values with 0; adds nonlinearity.
    nn.Linear(8, 3), # 8 hidden values → a 3-dimensional user embedding.
    ) # encode_users() L2-normalizes this embedding afterward.
    self.item_mlp = nn.Sequential(nn.Linear(8, 8), nn.ReLU(), nn.Linear(8, 3))

    def encode_users(self, profiles):
    ids = torch.tensor([genres.get(p["preferred_genre"], 0) for p in profiles])
    age = torch.tensor([[p["age"] / 100] for p in profiles], dtype=torch.float32)
    # 4 genre components + 1 scaled age = 5 input features.
    features = torch.cat([self.genre(ids), age], dim=1)
    return F.normalize(self.user_mlp(features), dim=1)

    def encode_items(self, profiles):
    ids = torch.tensor([genres.get(p["genre"], 0) for p in profiles])
    casts = []
    for p in profiles:
    cast_ids = torch.tensor([actors.get(a, 0) for a in p["cast"]] or [0])
    casts.append(self.actor(cast_ids).mean(dim=0))
    # Cast comes first here: input positions need not match the user tower.
    features = torch.cat([torch.stack(casts), self.genre(ids)], dim=1)
    return F.normalize(self.item_mlp(features), dim=1)

    model = TwoTower() # CPU; no pretrained models or downloads required.
    optimizer = torch.optim.Adam(model.parameters(), lr=0.03)

    for step in range(200):
    model.train()
    optimizer.zero_grad()
    u = model.encode_users(users)
    v = model.encode_items(items)
    logits = (u @ v.T) / 0.2 # Scaled cosine scores for all four pairs.
    loss = F.binary_cross_entropy_with_logits(logits, clicked)
    if step == 0:
    initial_loss = loss.item()
    loss.backward() # Gradients reach both MLPs and both embedding tables.
    optimizer.step()

    model.eval()
    with torch.no_grad():
    u = model.encode_users(users)
    v = model.encode_items(items)
    cosine = u @ v.T
    distances = torch.cdist(u, v, p=2)
    final_loss = F.binary_cross_entropy_with_logits(cosine / 0.2, clicked).item()
    best_items = cosine.argmax(dim=1)

    print("User vectors:", tuple(u.shape))
    print("Item vectors:", tuple(v.shape))
    print("Loss decreased:", final_loss < initial_loss)
    print("Top item for each user:", [items[i]["genre"] for i in best_items.tolist()])
    print("L2 gives the same choices:", torch.equal(best_items, distances.argmin(dim=1)))

    Expected output:

    User vectors: (2, 3)
    Item vectors: (2, 3)
    Loss decreased: True
    Top item for each user: ['horror', 'history']
    L2 gives the same choices: True
  • Follow the transformations: the user tower receives five numbers (four genre components plus scaled age); the item tower receives eight (four cast components plus four genre components). Their separate MLPs transform these different inputs into a shared three-dimensional output space. Reading by key makes the reordered item JSON fields harmless.

  • Training labels drive alignment: clicked supplies the target for each user–item pair. Binary cross-entropy compares these labels with scaled cosine scores; backpropagation updates both MLPs and both embedding tables together. The fixed dictionaries, age-scaling rule, and labels are not trained. Google's embedding guide describes learning embeddings through a task objective.

  • Scope of the demonstration: these synthetic examples show the training mechanics and memorization of four labels, not recommendation quality on unseen users or items. Use real positive and negative interaction examples and held-out evaluation for a real system. Missing interactions are not automatically negative labels. Save vocabulary mappings and preprocessing alongside model weights; category IDs must remain consistent at inference. Cosine scores are compatibility scores, not calibrated click probabilities.

< A basic neural ranker >​

  • A neural network can itself be a learning-to-rank model. “Neural network” describes the model architecture; “learning to rank” describes the task and training approach.

    Retrieve candidate items
    ↓
    Build features for each user–item pair
    ↓
    Ranking model produces one score per candidate
    ↓
    Sort candidates by score → return top items
  • Combine features: concatenate the user embedding, item embedding, and any additional features, then feed them into an MLP with one output neuron.

    User embedding: 32 dimensions
    Item embedding: 32 dimensions
    Additional features: 2 dimensions (e.g., popularity, freshness)
    ──
    MLP input: 66 dimensions
    MLP output: 1 score
  • Example — pointwise training: this PyTorch snippet shows the core training and ranking operations. It assumes features contains encoded user–item pairs with shape [batch_size, 66], clicked contains their 0/1 labels with shape [batch_size], and candidate_features contains candidates for one request with shape [num_candidates, 66]. These snippets require your prepared data.

    import torch
    from torch import nn

    ranker = nn.Sequential(
    nn.Linear(66, 32),
    nn.ReLU(),
    nn.Linear(32, 1), # One score per user–item pair
    )

    optimizer = torch.optim.Adam(ranker.parameters(), lr=0.001)

    # Training: repeat over batches of historical impressions.
    scores = ranker(features).squeeze(-1) # [batch_size]
    loss = nn.BCEWithLogitsLoss()(scores, clicked.float())

    optimizer.zero_grad()
    loss.backward()
    optimizer.step()

    # Inference: candidate_features contains items for ONE user/request.
    ranker.eval()
    with torch.inference_mode():
    scores = ranker(candidate_features).squeeze(-1)
    ranked_indices = scores.argsort(descending=True)
  • Interpret the score: the score is a click logit. Applying sigmoid converts it into a predicted click probability, but sorting logits gives the same order.

    P^(click⁡=1∣u,i,c)=σ ⁣(MLP⁡([eu;ei;c]))\widehat P(\operatorname{click}=1\mid u,i,c) = \sigma\!\left(\operatorname{MLP}([e_u;e_i;c])\right)

    Here, eue_u and eie_i are embeddings, cc contains encoded context features, and semicolons denote concatenation. This is pointwise ranking: the model learns from each pair’s label individually. Rating prediction with a regression loss is another pointwise approach, demonstrated in TensorFlow's ranking examples.

  • Example context: context could include the device type and time of day. The model could learn that the same user is more likely to click a short video on a phone during a commute than a full-length movie.

< Train specifically for relative ranking >​

  • Choose a training objective: you can train the same MLP using a ranking objective.

    ApproachTraining labelsWhat training encourages
    PointwiseClick/no-click or rating for each itemPredict each item’s target
    PairwiseItem A should rank above item B for the same requestGive A a higher score than B
    ListwiseRelevance labels for a candidate listProduce a better ordering of the list
  • Example — RankNet: use the difference between two scores. Here, features_a and features_b are equally sized batches of 66-dimensional feature vectors; each paired row belongs to the same user/query context.

    # Alternative training objective for the MLP above.
    ranker.train()
    score_a = ranker(features_a).squeeze(-1)
    score_b = ranker(features_b).squeeze(-1)

    # Label 1 means A should rank above B.
    loss = nn.BCEWithLogitsLoss()(
    score_a - score_b,
    torch.ones_like(score_a),
    )

    optimizer.zero_grad()
    loss.backward()
    optimizer.step()

    Backpropagation trains the model to make score_a > score_b. At inference, it still scores each candidate and sorts the scores. Neural ranking frameworks support pointwise, pairwise, and listwise objectives. TensorFlow Ranking paper

< Use a tree-based learning-to-rank model >​

  • XGBoost's XGBRanker uses boosted decision trees rather than an MLP. Prepare one row per candidate, with a group ID identifying the recommendation request:

    Request IDUserItemFeaturesRelevance label
    101User AMovie 1User/item/context features3
    101User AMovie 2User/item/context features0
    101User AMovie 3User/item/context features1
    102User BMovie 1User/item/context features0
    102User BMovie 2User/item/context features2

    Larger labels mean greater relevance. Candidates are compared within the same request.

  • Example: assume X is a NumPy array of numeric feature rows, relevance and request_ids are aligned one-dimensional NumPy arrays, and candidate_features uses the same feature schema for one new request.

    import numpy as np
    from xgboost import XGBRanker

    # X: numeric feature rows
    # relevance: one relevance label per row
    # request_ids: identifies which candidates belong together
    order = np.argsort(request_ids)

    ranker = XGBRanker(
    objective="rank:ndcg",
    n_estimators=100,
    max_depth=4,
    )

    ranker.fit(
    X[order],
    relevance[order],
    qid=request_ids[order],
    )

    # Candidates for ONE new recommendation request.
    scores = ranker.predict(candidate_features)
    ranked_indices = np.argsort(-scores)
  • Interpret the objective: rank:ndcg uses LambdaMART with NDCG-based weighting to favor useful ranking changes, especially near the top of the list. Its scores are relative ranking values, not click probabilities. XGBoost learning-to-rank documentation

  • Connect the ranker to two-tower retrieval: the resulting pipeline could be:

    User tower + item embeddings
    ↓
    Cosine similarity retrieves 100 candidates
    ↓
    MLP ranker OR XGBoost ranker scores those candidates
    ↓
    Return the top 10

Q & A​

< What is an embedding coordinate? >​

  • A coordinate is one number in a vector. In [0.2, -0.7, 0.4], coordinate 1 is 0.2, accessed as embedding[0] in Python. This vector has three coordinates, or dimensions.
  • Output coordinates are learned combinations of inputs. The user tower's first output can depend on both genre preference and age. It has no required meaning such as “age” or “interest in horror.” The item tower's outputs are aligned with the user's outputs through joint training.

< What is feature encoding, and how is it implemented? >​

  • Feature encoding converts profile fields into numeric model inputs. The application defines which keys to read and how to transform their values:

    FieldEncoding operationResult
    age: 25Divide by 100 in the simple exampleOne scaled number: 0.25
    preferred_genre: "horror"Map the category to an ID, then use nn.EmbeddingA trainable genre vector
    genre: "horror"Use the shared genre vocabulary and lookup tableA trainable genre vector
    cast: ["Jason", "David"]Look up each actor and average the vectorsOne cast vector
  • IDs select rows; they do not express magnitude. Assigning "horror" ID 1 makes it select row 1 of the genre table. The table starts with random weights and is updated during training. The category-to-ID mapping stays fixed.

  • Concatenation creates the tower input. Join the encoded field vectors and numeric values in a consistent order for each tower. The runnable two-tower example implements vocabulary lookup, age scaling, cast pooling, concatenation, and training in one script.

< What model is trained, and how does it learn user–item compatibility? >​

  • Train the whole two-tower model jointly. Its trainable parameters include the category embedding tables and the weights and biases of the user and item MLPs. In the simple implementation above, every one of these components is trained from scratch. In the richer JSON profile example, the pretrained description encoder remains frozen.
  • Interaction labels supply the learning signal. A record contains a user profile, an item profile, and an observed target such as clicked = 1. The model encodes both profiles, predicts a compatibility score, calculates a loss, and backpropagates gradients. An optimizer then updates the trainable parameters.
  • Feature mappings and compatibility play different roles. The application explicitly reads preferred_genre on the user side and genre on the item side. Positive and negative interaction examples teach the model how these features relate to preference. Equal vector dimensions enable comparison; the shared training objective makes that comparison useful.

< Do user and item profiles need the same JSON schema for cosine similarity or L2 distance? >​

  • Different schemas are supported. The user and item towers can accept different keys, field types, and input dimensions. Cosine similarity compares their output vectors; it does not compare JSON keys. Each profile must follow the input schema expected by its own tower.

    RequirementMust match?Explanation
    JSON keys and fieldsNoUsers can have preference and age; items can have genre, cast, and length_minutes
    Encoded input dimensionNoEach tower has its own input layers
    Tower architecture and weightsNoThe towers can process their inputs differently
    Output embedding dimensionYesBoth vectors must contain the same number of components
    Learned embedding spaceYes, for useful scoresBoth outputs must be aligned so their similarity reflects user–item compatibility
  • JSON field order and embedding coordinates are different things. The preprocessing code reads fields by key and puts each feature into the input slot expected by its own tower. The tower transforms those inputs into a new vector. Its output coordinate positions are learned combinations of features, not the positions of keys in the JSON object.

  • Example: different genre keys and positions. These two item objects produce the same encoded input when fields are extracted by key:

    {"cast": ["Jason"], "genre": "horror"}
    {"genre": "horror", "cast": ["Jason"]}

    For a user profile such as {"preferred_genre": "horror", "age": 25}, each side follows its own explicit mapping. The following is pseudocode:

    user_features = concatenate(
    encode_user_genre(user["preferred_genre"]),
    scale_age(user["age"]),
    )
    item_features = concatenate(
    encode_cast(item["cast"]),
    encode_item_genre(item["genre"]),
    )

    user_vector = user_tower(user_features) # d learned coordinates
    item_vector = item_tower(item_features) # d learned coordinates

    Genre can occupy different positions in the two input vectors because the towers have different input weights. Each tower must keep its own feature mapping consistent between training and inference. A genre lookup table may be shared when both fields use the same genre vocabulary, but separate genre encoders can also learn alignment through the joint objective.

  • Example: the JSON profile implementation. In the Embeddings page's PyTorch example, the user features form a 19-dimensional input, while the item features form a 417-dimensional input, including the 384-dimensional description vector:

    User JSON: age, gender, preference, income_usd, watch_hours_weekly
    → user feature encoding → 19 numbers → user tower → 32 numbers

    Item JSON: genre, length_minutes, cast, director, company, short_description
    → item feature encoding → 417 numbers → item tower → 32 numbers

    User embedding + item embedding → cosine similarity → retrieval score

    The output coordinates are learned components. Coordinate 1 does not simply mean age on the user side or genre on the item side. The towers combine their input features into representations suitable for comparison.

  • Training aligns the vectors. Jointly train the towers using user–item interactions and a loss based on their compatibility scores. For example, positive interactions can encourage a user interested in horror to score highly with relevant horror movies, while negative examples provide contrasting signals. The relation between preference and genre comes from feature encoding and training, not matching key names. TensorFlow's retrieval tutorial demonstrates jointly trained user and movie models; Google's embedding guide explains task-based embedding learning.

  • L2 distance compares corresponding output coordinates. Once the towers produce vectors in a shared space, Euclidean distance is:

    ∥eu−ei∥2=∑k=1d(eu,k−ei,k)2\lVert e_u-e_i\rVert_2 =\sqrt{\sum_{k=1}^{d}(e_{u,k}-e_{i,k})^2}

    The subtraction pairs learned coordinate kk with learned coordinate kk. It does not subtract age from cast or compare the first JSON values. For smaller distances to indicate better recommendations, train with a compatible objective, such as a contrastive or triplet loss, or use normalized embeddings trained for cosine similarity.

  • Example: user 1 is closer to item 1. Suppose trained towers produce these illustrative unit-length vectors:

    ProfileOutput embeddingL2 distance from user 1
    User 1[1.0, 0.0]—
    Item 1[0.8, 0.6](1−0.8)2+(0−0.6)2≈0.632\sqrt{(1-0.8)^2+(0-0.6)^2}\approx0.632
    Item 2[0.0, 1.0](1−0)2+(0−1)2≈1.414\sqrt{(1-0)^2+(0-1)^2}\approx1.414

    Item 1 is closer because its learned vector has the smaller distance. Training data and the objective determine whether that closeness predicts user preference; these illustrative coordinates cannot be inferred from the JSON schemas alone.

  • Cosine similarity needs nonzero vectors of equal dimension. For user vector eue_u and item vector eie_i:

    cosine⁡(eu,ei)=eu⊤ei∥eu∥2∥ei∥2\operatorname{cosine}(e_u,e_i) = \frac{e_u^\top e_i}{\lVert e_u\rVert_2\lVert e_i\rVert_2}

    L2-normalizing both vectors makes their dot product equal cosine similarity. Use a scoring and normalization scheme consistent with training. Independently trained or randomly initialized vectors can yield a numerical cosine score without useful recommendation meaning; the JSON example's towers need interaction training before their scores become useful.

    For unit-length vectors, squared L2 distance and cosine similarity satisfy:

    ∥eu−ei∥22=2−2cosine⁡(eu,ei)\lVert e_u-e_i\rVert_2^2=2-2\operatorname{cosine}(e_u,e_i)

    Thus, minimizing L2 distance gives the same ordering as maximizing cosine similarity when every compared vector is normalized. For unnormalized vectors, these rankings can differ.

  • Embedding JSON as text is a separate option. One pretrained text encoder can embed differently structured profile descriptions, but its text similarity measures semantic relatedness rather than learned personal preference. Validate this baseline on recommendation data. Two unrelated text encoders also need alignment before their vectors can be compared meaningfully, even if their output dimensions match.

Reference​