Skip to main content

📝 Term Frequency–Inverse Document Frequency (TF-IDF)

Description​

< What is it? >​

Term Frequency–Inverse Document Frequency (TF-IDF) weights terms in a document using two signals: how often a term appears in that document and how rare it is across the document collection, or corpus. A term gets a larger weight when it appears frequently in a document but in relatively few documents overall. See Introduction to Information Retrieval.

ComponentWhat it measuresIntuition
Term frequency (TF)Frequency within one documentIs this term prominent here?
Inverse document frequency (IDF)Rarity across the corpusDoes this term help distinguish documents?
TF-IDFThe product of TF and IDFHow strongly should this term characterize this document?
  • Example: “neuron” can receive a high weight in a neural-network article if it appears often there but in few other articles.

TF-IDF supports lexical retrieval and can also provide input features for text classifiers, as illustrated in Google's text-classification guide.

Key points​

< Calculating TF and IDF >​

TF has several conventions. This page's worked example uses the fraction of a document's tokens that match the term:

TF⁡(t,d)=count⁡(t,d)number of tokens in d\operatorname{TF}(t,d)=\frac{\operatorname{count}(t,d)}{\text{number of tokens in }d}

For a term present in the corpus, a basic IDF definition is:

IDF⁡(t)=ln⁡(Ndf⁡(t))\operatorname{IDF}(t)=\ln\left(\frac{N}{\operatorname{df}(t)}\right)

Here, NN is the total number of documents and df⁡(t)\operatorname{df}(t) is the number of documents containing the term. Document frequency counts documents, not total occurrences. The information-retrieval textbook's IDF explanation describes why corpus-wide rarity matters.

"Inverse" means that, for a fixed corpus size, the more documents contain a term, the lower its IDF. Document frequency is in the denominator: increasing df⁡(t)\operatorname{df}(t) decreases N/df⁡(t)N/\operatorname{df}(t) and therefore its logarithm.

  • Example: in a collection of 100 documents, the basic formula above gives:

    Documents containing the termIDF (natural logarithm)Meaning
    14.61Rare; helps distinguish documents
    102.30Moderately common
    1000Appears everywhere; cannot distinguish documents by its presence

Multiply the two components:

TF-IDF⁡(t,d)=TF⁡(t,d)×IDF⁡(t)\operatorname{TF\text{-}IDF}(t,d)=\operatorname{TF}(t,d)\times\operatorname{IDF}(t)
  • Example: take three documents: cat sat cat, dog sat, and bird sat. In the first document, cat has TF 2/32/3 and IDF ln⁡(3/1)\ln(3/1), giving TF-IDF approximately 0.7324.

< Turning documents into vectors >​

Assign one vector component to each vocabulary term and fill it with that term's TF-IDF weight. Terms absent from a document receive zero, so these vectors are usually sparse. A vocabulary of 10,000 terms produces a 10,000-dimensional vector regardless of document length. This representation follows the TF-IDF vector model.

For retrieval, transform the query using the same vocabulary and corpus IDF values, then compare it with document vectors using cosine similarity. Do not fit a separate vocabulary or IDF calculation for each query.

Optional word stemming merges variants such as connected and connecting into one vocabulary entry, connect, before computing TF-IDF. Apply the same preprocessing to documents and queries. The scikit-learn example below uses its default tokenizer without stemming.

< Formula variants in libraries >​

Library defaults may differ from the worked example. In scikit-learn, TfidfVectorizer defaults to raw term counts, smoothed IDF, and L2 normalization of each resulting vector. Its smoothed IDF is:

IDF⁡smooth(t)=ln⁡(1+N1+df⁡(t))+1\operatorname{IDF}_{\text{smooth}}(t)=\ln\left(\frac{1+N}{1+\operatorname{df}(t)}\right)+1

A term present in every document therefore has IDF 1, rather than 0, with these defaults. See scikit-learn's TF-IDF documentation.

Comparison​

< TF-IDF vs. learned embeddings >​

AspectTF-IDFLearned dense embeddings
ComponentsWeights for explicit vocabulary termsLearned numerical features
Vector sizeDetermined by vocabulary sizeDetermined by model architecture
MatchingTerm overlapLearned semantic relationships
FittingComputes vocabulary and corpus statistics; no relevance labels requiredUses a model learned from a training objective, often available pretrained

Learned embeddings can place related concepts near one another even when they have different words. Google's Machine Learning Crash Course explains this distinction between sparse representations and learned semantic representations.

Implementation​

< Calculating a term weight in Python >​

This example uses only Python's standard library and the unsmoothed, length-normalized formula above:

from math import log

documents = ["cat sat cat", "dog sat", "bird sat"]
tokens = [document.split() for document in documents]
term = "cat"

tf = tokens[0].count(term) / len(tokens[0])
df = sum(term in document for document in tokens)
idf = log(len(documents) / df)

print(f"TF: {tf:.4f}")
print(f"IDF: {idf:.4f}")
print(f"TF-IDF: {tf * idf:.4f}")

Expected output:

TF: 0.6667
IDF: 1.0986
TF-IDF: 0.7324

< Retrieving documents with scikit-learn >​

TF-IDF search represents documents and the query as vectors, then ranks documents by cosine similarity to the query:

  1. Fit TF-IDF on the document collection to build its vocabulary and IDF weights.
  2. Transform the query using that same fitted vectorizer.
  3. Calculate similarity and return the highest-scoring matches.
  • Example: with scikit-learn installed, search three documents for cat dog:

    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.metrics.pairwise import cosine_similarity

    documents = [
    "cat dog",
    "cat bird",
    "fish bird",
    ]

    # Indexing: compute document vectors once.
    vectorizer = TfidfVectorizer()
    document_vectors = vectorizer.fit_transform(documents)

    # Searching: reuse the vocabulary and IDF weights.
    query = "cat dog"
    query_vector = vectorizer.transform([query])

    scores = cosine_similarity(query_vector, document_vectors)[0]
    ranked_indices = scores.argsort()[::-1]

    for i in ranked_indices:
    if scores[i] > 0:
    print(f"{scores[i]:.3f} | {documents[i]}")

    Expected output:

    1.000 | cat dog
    0.428 | cat bird
    DocumentWhy it ranks here
    cat dogMatches both query terms; its vector is identical to the query vector
    cat birdShares only cat, so similarity is lower
    fish birdShares no query terms; its score is zero and the code excludes it

The score is a similarity measure, not a probability. TF-IDF matches vocabulary terms; it does not automatically know that words such as car and automobile have similar meanings.

Troubleshoot​

SymptomWhat to check
Every query score is zeroQuery terms may be absent from the vocabulary or removed during preprocessing; return no lexical match instead of an arbitrary document
Hand calculations differ from library resultsCheck TF scaling, IDF smoothing, and final vector normalization
Similar meanings receive low similarityTF-IDF does not inherently connect synonyms; consider semantic or hybrid retrieval

Reference​