π Word Stemming
Descriptionβ
< What is it? >β
Word stemming reduces words to a shared root-like form, called a stem, usually by removing or rewriting suffixes. The stem serves as a matching key and does not have to be a real dictionary word. Elastic's stemming guide explains how this helps word variants match during search.
-
Example: the English Porter stemming algorithm produces these stems:
Original word Stem connect connect connected connect connecting connect connections connect studies studi
The output studi is useful for matching even though it is not an English word. Different stemming algorithms can produce different results.
Key pointsβ
< Stemming in search and TF-IDF >β
Stemming is a preprocessing step before building a lexical vocabulary or calculating TF-IDF weights. Word forms that share a stem use the same vocabulary entry, so their occurrences contribute to the same term statistics.
Document: "connected" β stem: "connect"
Query: "connecting" β stem: "connect" β lexical match
Use the same stemming process for indexed documents and incoming queries. Elastic recommends consistent stemmer filters during indexing and search analysis.
< Benefits and limitations >β
Stemming can retrieve relevant documents that use a different word form, but it can also merge words whose meanings differ. Its effect on search quality depends on the queries and collection. The information-retrieval textbook discusses this tradeoff between finding more matches and retaining precise matches.
Stemming rules are language-specific. Stemming also does not inherently match synonyms such as car and automobile; synonym expansion or semantic retrieval addresses that separate problem.
Comparisonβ
< Stemming vs. lemmatization >β
| Aspect | Stemming | Lemmatization |
|---|---|---|
| Method | Applies word-form rules, such as suffix removal | Uses vocabulary and morphological analysis, often with part-of-speech information |
| Output | A stem that may not be a dictionary word | A dictionary base form, called a lemma |
| Example | studies β studi with Porter stemming | studies β study |
| Purpose | Group word variants for matching | Recover a word's linguistic base form |
See Introduction to Information Retrieval for the distinction and its implications for retrieval.
Implementationβ
< Porter stemming in Python >β
Stemming is the general technique; Porter is a specific English stemming algorithm that applies suffix-reduction rules. The Natural Language Toolkit (NLTK) provides PorterStemmer; its default mode includes NLTK's extensions to the original algorithm. See the NLTK PorterStemmer documentation and the Porter algorithm description.
Install NLTK once:
python -m pip install nltk
-
Example: stem individual words with
stem(). This example needs no downloaded corpora or trained model.from nltk.stem import PorterStemmerstemmer = PorterStemmer()words = ["connect", "connected", "connecting", "connections", "studies"]for word in words:print(f"{word} β {stemmer.stem(word)}")Expected output:
connect β connectconnected β connectconnecting β connectconnections β connectstudies β studi
The connection-related forms share the matching key connect. The result studi illustrates that a stem need not be a dictionary word. For full sentences, tokenize the text first, then stem each token.
Related ideasβ
- Search & Retrieval places text preprocessing within a search pipeline.
- Term FrequencyβInverse Document Frequency (TF-IDF) assigns weights to the terms produced by preprocessing.
- Embeddings provides learned representations for semantic matching.