Skip to main content

πŸ“ Extreme Gradient Boosting (XGBoost)

Description​

< What is it? >​

Extreme Gradient Boosting (XGBoost) is a machine-learning library best known for its efficient implementation of gradient-boosted decision trees. It supports classification, regression, and ranking, and is commonly used with structured, tabular data.

The tree booster builds an ensemble in stages. Each new tree adds a correction to the current prediction, guided by the training loss. Regularization and efficient tree-building algorithms help control complexity and scale training. The library also supports other boosters; this page focuses on boosted trees.

Key points​

< How boosting builds a prediction >​

Let Ftβˆ’1(x)F_{t-1}(x) be the ensemble's current raw score and ft(x)f_t(x) the new tree's output. A boosting step updates the score as follows:

Ft(x)=Ftβˆ’1(x)+Ξ·ft(x)F_t(x)=F_{t-1}(x)+\eta f_t(x)

Here, Ξ·\eta is the learning rate, which scales each tree's contribution. Each tree receives the original feature vector xx and returns a leaf score. Training is sequential because the next tree depends on the current ensemble's loss derivatives.

Input features x
| | |
Tree 1 Tree 2 Tree 3 ...
| | |
score score score
+---------+---------+
|
Sum with base score
|
Task-specific output

The diagram shows prediction, with learning-rate scaling included in the stored tree contributions. For regression with squared-error loss, the raw score is the prediction. For binary classification with a logistic objective, a sigmoid converts it into a probability:

P^(y=1∣x)=Οƒ(FT(x))=11+eβˆ’FT(x)\widehat P(y=1\mid x)=\sigma(F_T(x)) =\frac{1}{1+e^{-F_T(x)}}
  • Example: in squared-error regression, suppose the current prediction is 60 and the target is 80. If the next tree predicts a correction of 15 and Ξ·=0.1\eta=0.1, the updated prediction is 60+0.1Γ—15=61.560+0.1\times15=61.5. Trees learn corrections across many training rows, so a single update need not eliminate an individual error.

Google's gradient boosting lesson explains how residual fitting for squared error generalizes to other losses.

< Loss derivatives and regularization >​

XGBoost uses first- and second-order loss derivatives to evaluate tree splits and leaf scores. For squared error, these updates relate directly to residuals; for classification and ranking, they follow the chosen objective. Training also penalizes tree complexity and leaf weights. See the boosted-tree tutorial for the derivation.

< Important parameters >​

ParameterWhat it controls
n_estimatorsMaximum number of boosting rounds in the scikit-learn interface
learning_rateContribution of each new tree; smaller values often need more rounds
max_depthMaximum tree depth; deeper trees can learn more complex interactions
min_child_weightMinimum sum of instance Hessians needed in a child; larger values discourage splits
subsampleFraction of training rows sampled for each boosting round
colsample_bytreeFraction of features sampled for each tree
reg_alpha, reg_lambdaL1 and L2 penalties on leaf weights
gammaMinimum loss reduction required to make a split
tree_method="hist"Histogram-based split search
early_stopping_roundsStops training after the monitored validation metric fails to improve for this many rounds

Use validation data to select model complexity and training duration. The parameter reference documents the individual controls.

< Ranking in a recommendation system >​

After retrieval, XGBoost can score candidate items using features such as embedding similarity, item popularity, user–category interaction counts, and request context. Each row represents a candidate within a recommendation request.

XGBRanker supports LambdaMART-style learning to rank. With objective="rank:ndcg", training uses pairwise updates weighted by their effect on Normalized Discounted Cumulative Gain (NDCG), which rewards placing relevant items near the top. Rows must be grouped by query or recommendation request using query IDs (qid). The resulting scores are used to order candidates within a request.

  • Example: two-tower retrieval returns 500 movies. XGBoost receives 500 feature rows for that request and produces a score for each movie. Sorting those scores gives the ranked candidate list.

A ranker produces relevance scores; a classifier trained on click labels can instead estimate click probabilities. These objectives serve different training goals. See XGBoost's learning-to-rank guide.

Comparison​

Implementation​

< Binary classification with early stopping >​

Install the Python dependencies:

python -m pip install xgboost scikit-learn

This synthetic example uses separate training, validation, and test splits. Validation controls early stopping; the test set is reserved for the final measurement.

from sklearn.datasets import make_classification
from sklearn.metrics import log_loss
from sklearn.model_selection import train_test_split
from xgboost import XGBClassifier

X, y = make_classification(
n_samples=2000, n_features=20, n_informative=10, random_state=42
)
X_dev, X_test, y_dev, y_test = train_test_split(
X, y, test_size=0.2, stratify=y, random_state=42
)
X_train, X_valid, y_train, y_valid = train_test_split(
X_dev, y_dev, test_size=0.25, stratify=y_dev, random_state=42
)

model = XGBClassifier(
objective="binary:logistic",
tree_method="hist",
n_estimators=500,
max_depth=4,
learning_rate=0.05,
subsample=0.8,
colsample_bytree=0.8,
eval_metric="logloss",
early_stopping_rounds=20,
random_state=42,
n_jobs=2,
)
model.fit(X_train, y_train, eval_set=[(X_valid, y_valid)], verbose=False)

probabilities = model.predict_proba(X_test)[:, 1]
print("Test log loss:", log_loss(y_test, probabilities))
print("Best boosting round (zero-based):", model.best_iteration)

The scikit-learn interface automatically uses the best iteration for prediction after early stopping. The example's parameter values are illustrative, not universally optimal. See the estimator guide.

< Training a ranking model >​

  • What is a request? A request is one occasion when the application asks for recommendationsβ€”for example, a user opens the home page. The application assigns the request ID. During training, qid tells XGBoost which candidates belong to the same ranking task. One user can make many requests.

    Request 101: Alice opens the home page β†’ rank 3 candidate movies
    Request 102: Bob opens the home page β†’ rank 3 candidate movies
    Request 103: Alice refreshes later β†’ rank another candidate list
  • Training input: prepare feature rows and relevance labels for each request. Use nonnegative integer relevance grades with rank:ndcg, keep each request in one data split, and sort rows by qid before fitting. This small, synthetic example uses three features in a fixed order: [embedding_similarity, genre_match, popularity]. genre_match is 1 when the movie matches the user's preferred genre and 0 otherwise; popularity is scaled to the range 0–1.

  • Who defines embedding_similarity? You choose the similarity measure as a feature, and your application calculates it. In a two-tower system, it can be the cosine similarity between the trained user and item embeddings. The ranking example below manually supplies synthetic values such as 0.95 and 0.35; in production, the feature pipeline would calculate them from embeddings.

    ComponentWhat it does
    You, the developerChoose cosine similarity as a ranking feature
    Trained user/item towersProduce embeddings from profiles
    Feature-processing codeCalculate their cosine similarity
    XGBoost rankerCombine that similarity with genre match, popularity, and other features to produce a final ranking score
    User profile β†’ User tower β†’ User embedding ─┐
    β”œβ†’ Cosine similarity ─┐
    Item profile β†’ Item tower β†’ Item embedding β”€β”˜ β”‚
    β”œβ†’ XGBoost β†’ Ranking score
    Genre match + popularity β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
  • Example β€” calculate the similarity feature: these concrete sample vectors illustrate the calculation; a trained pair of towers would supply the vectors in production.

    import numpy as np

    user_embedding = np.array([1.0, 0.0])
    item_embedding = np.array([0.8, 0.6])

    embedding_similarity = (
    np.dot(user_embedding, item_embedding)
    / (np.linalg.norm(user_embedding) * np.linalg.norm(item_embedding))
    )

    print(embedding_similarity)

    Output:

    0.8
  • Inference input: for one new recommendation request, pass a numeric matrix with one row per candidate and the same feature definitions and column order used during training. Here, X_candidates has shape (3, 3). Movie IDs are kept separately so the scores can be mapped back to movies.

    Movie IDEmbedding similarityGenre matchPopularity
    movie_A0.3500.90
    movie_B0.9510.70
    movie_C0.6510.50

    predict() needs neither relevance labels nor qid for this NumPy input. The application keeps track of which candidates belong to each request and sorts each request separately.

  • Complete example: install numpy, xgboost, and scikit-learn, then run the following code. The training rows below represent two historical requests with three candidates each; relevance grades range from 0 (irrelevant) to 3 (highly relevant).

    import numpy as np
    from xgboost import XGBRanker

    # Columns: [embedding_similarity, genre_match, popularity]
    X_train = np.array([
    [0.95, 1, 0.70], # Request 101: highly relevant
    [0.65, 1, 0.50], # Request 101: somewhat relevant
    [0.35, 0, 0.90], # Request 101: irrelevant
    [0.90, 1, 0.40], # Request 102: highly relevant
    [0.60, 1, 0.80], # Request 102: somewhat relevant
    [0.20, 0, 0.60], # Request 102: irrelevant
    ], dtype=np.float32)
    relevance_train = np.array([3, 1, 0, 3, 1, 0])
    request_ids_train = np.array([101, 101, 101, 102, 102, 102])

    ranker = XGBRanker(
    objective="rank:ndcg",
    tree_method="hist",
    n_estimators=50,
    max_depth=2,
    learning_rate=0.1,
    min_child_weight=0, # Allows splits in this tiny teaching dataset
    random_state=42,
    n_jobs=1,
    )
    ranker.fit(X_train, relevance_train, qid=request_ids_train)

    # Inference: three candidates for ONE new recommendation request.
    movie_ids = np.array(["movie_A", "movie_B", "movie_C"])
    X_candidates = np.array([
    [0.35, 0, 0.90], # movie_A
    [0.95, 1, 0.70], # movie_B
    [0.65, 1, 0.50], # movie_C
    ], dtype=np.float32)

    scores = ranker.predict(X_candidates) # Shape: (3,), in input row order
    top_indices = np.argsort(-scores, kind="stable")[:3]

    print("Input shape:", X_candidates.shape)
    print("Output shape:", scores.shape)
    for movie_id, score in zip(movie_ids, scores):
    print(f"{movie_id}: {score:.3f}")
    print("Ranked movie IDs:", movie_ids[top_indices].tolist())
  • Inference output: predict() returns one relevance score per input row, not a sorted list or a click probability. Scores may be negative or greater than 1. The application sorts them in descending order and uses the same indices to reorder movie IDs. See XGBoost's learning-to-rank guide.

    For a concrete illustration, if the model returns the scores below, the resulting ordering is:

    Input shape: (3, 3)
    Output shape: (3,)
    movie_A: -0.800
    movie_B: 1.200
    movie_C: 0.300
    Ranked movie IDs: ['movie_B', 'movie_C', 'movie_A']

    These score values illustrate the input-to-output mapping; they are not captured output from the training code above. Running the code prints the actual learned scores, which depend on the fitted model and library version.

  • Evaluate separately: this tiny example demonstrates the mechanics and reuses feature patterns from training; it does not measure generalization. For interaction logs, choose a time-based split when evaluating predictions of future behavior, and build features only from information available at each request. Evaluate ranking quality per request rather than treating the scores as class probabilities.

Video Tutorial​

Reference​