📝 Term Frequency–Inverse Document Frequency (TF-IDF)
Description
< What is it? >
Term Frequency–Inverse Document Frequency (TF-IDF) weights terms in a document using two signals: how often a term appears in that document and how rare it is across the document collection, or corpus. A term gets a larger weight when it appears frequently in a document but in relatively few documents overall. See Introduction to Information Retrieval.
| Component | What it measures | Intuition |
|---|---|---|
| Term frequency (TF) | Frequency within one document | Is this term prominent here? |
| Inverse document frequency (IDF) | Rarity across the corpus | Does this term help distinguish documents? |
| TF-IDF | The product of TF and IDF | How strongly should this term characterize this document? |
- Example: “neuron” can receive a high weight in a neural-network article if it appears often there but in few other articles.
TF-IDF supports lexical retrieval and can also provide input features for text classifiers, as illustrated in Google's text-classification guide.
Key points
< Calculating TF and IDF >
TF has several conventions. This page's worked example uses the fraction of a document's tokens that match the term:
For a term present in the corpus, a basic IDF definition is:
Here, is the total number of documents and is the number of documents containing the term. Document frequency counts documents, not total occurrences. The information-retrieval textbook's IDF explanation describes why corpus-wide rarity matters.
"Inverse" means that, for a fixed corpus size, the more documents contain a term, the lower its IDF. Document frequency is in the denominator: increasing decreases and therefore its logarithm.
-
Example: in a collection of 100 documents, the basic formula above gives:
Documents containing the term IDF (natural logarithm) Meaning 1 4.61 Rare; helps distinguish documents 10 2.30 Moderately common 100 0 Appears everywhere; cannot distinguish documents by its presence
Multiply the two components:
- Example: take three documents:
cat sat cat,dog sat, andbird sat. In the first document,cathas TF and IDF , giving TF-IDF approximately 0.7324.
< Turning documents into vectors >
Assign one vector component to each vocabulary term and fill it with that term's TF-IDF weight. Terms absent from a document receive zero, so these vectors are usually sparse. A vocabulary of 10,000 terms produces a 10,000-dimensional vector regardless of document length. This representation follows the TF-IDF vector model.
For retrieval, transform the query using the same vocabulary and corpus IDF values, then compare it with document vectors using cosine similarity. Do not fit a separate vocabulary or IDF calculation for each query.
Optional word stemming merges variants such as connected and connecting into one vocabulary entry, connect, before computing TF-IDF. Apply the same preprocessing to documents and queries. The scikit-learn example below uses its default tokenizer without stemming.
< Formula variants in libraries >
Library defaults may differ from the worked example. In scikit-learn, TfidfVectorizer defaults to raw term counts, smoothed IDF, and L2 normalization of each resulting vector. Its smoothed IDF is:
A term present in every document therefore has IDF 1, rather than 0, with these defaults. See scikit-learn's TF-IDF documentation.
Comparison
< TF-IDF vs. learned embeddings >
| Aspect | TF-IDF | Learned dense embeddings |
|---|---|---|
| Components | Weights for explicit vocabulary terms | Learned numerical features |
| Vector size | Determined by vocabulary size | Determined by model architecture |
| Matching | Term overlap | Learned semantic relationships |
| Fitting | Computes vocabulary and corpus statistics; no relevance labels required | Uses a model learned from a training objective, often available pretrained |
Learned embeddings can place related concepts near one another even when they have different words. Google's Machine Learning Crash Course explains this distinction between sparse representations and learned semantic representations.
Implementation
< Calculating a term weight in Python >
This example uses only Python's standard library and the unsmoothed, length-normalized formula above:
from math import log
documents = ["cat sat cat", "dog sat", "bird sat"]
tokens = [document.split() for document in documents]
term = "cat"
tf = tokens[0].count(term) / len(tokens[0])
df = sum(term in document for document in tokens)
idf = log(len(documents) / df)
print(f"TF: {tf:.4f}")
print(f"IDF: {idf:.4f}")
print(f"TF-IDF: {tf * idf:.4f}")
Expected output:
TF: 0.6667
IDF: 1.0986
TF-IDF: 0.7324
< Retrieving documents with scikit-learn >
TF-IDF search represents documents and the query as vectors, then ranks documents by cosine similarity to the query:
- Fit TF-IDF on the document collection to build its vocabulary and IDF weights.
- Transform the query using that same fitted vectorizer.
- Calculate similarity and return the highest-scoring matches.
-
Example: with scikit-learn installed, search three documents for
cat dog:from sklearn.feature_extraction.text import TfidfVectorizerfrom sklearn.metrics.pairwise import cosine_similaritydocuments = ["cat dog","cat bird","fish bird",]# Indexing: compute document vectors once.vectorizer = TfidfVectorizer()document_vectors = vectorizer.fit_transform(documents)# Searching: reuse the vocabulary and IDF weights.query = "cat dog"query_vector = vectorizer.transform([query])scores = cosine_similarity(query_vector, document_vectors)[0]ranked_indices = scores.argsort()[::-1]for i in ranked_indices:if scores[i] > 0:print(f"{scores[i]:.3f} | {documents[i]}")Expected output:
1.000 | cat dog0.428 | cat birdDocument Why it ranks here cat dogMatches both query terms; its vector is identical to the query vector cat birdShares only cat, so similarity is lowerfish birdShares no query terms; its score is zero and the code excludes it
The score is a similarity measure, not a probability. TF-IDF matches vocabulary terms; it does not automatically know that words such as car and automobile have similar meanings.
Troubleshoot
| Symptom | What to check |
|---|---|
| Every query score is zero | Query terms may be absent from the vocabulary or removed during preprocessing; return no lexical match instead of an arbitrary document |
| Hand calculations differ from library results | Check TF scaling, IDF smoothing, and final vector normalization |
| Similar meanings receive low similarity | TF-IDF does not inherently connect synonyms; consider semantic or hybrid retrieval |
Related ideas
- Word Stemming explains how word normalization changes the vocabulary used by TF-IDF.
- Search & Retrieval places TF-IDF within lexical, semantic, and hybrid retrieval.
- Embeddings explains learned vector representations.
- Scikit-learn provides text vectorizers and machine-learning models.
- Retrieval Augmented Generation (RAG) can use lexical retrieval to select supporting passages.