Skip to main content

πŸ“ Word Stemming

Description​

< What is it? >​

Word stemming reduces words to a shared root-like form, called a stem, usually by removing or rewriting suffixes. The stem serves as a matching key and does not have to be a real dictionary word. Elastic's stemming guide explains how this helps word variants match during search.

  • Example: the English Porter stemming algorithm produces these stems:

    Original wordStem
    connectconnect
    connectedconnect
    connectingconnect
    connectionsconnect
    studiesstudi

The output studi is useful for matching even though it is not an English word. Different stemming algorithms can produce different results.

Key points​

< Stemming in search and TF-IDF >​

Stemming is a preprocessing step before building a lexical vocabulary or calculating TF-IDF weights. Word forms that share a stem use the same vocabulary entry, so their occurrences contribute to the same term statistics.

Document: "connected" β†’ stem: "connect"
Query: "connecting" β†’ stem: "connect" β†’ lexical match

Use the same stemming process for indexed documents and incoming queries. Elastic recommends consistent stemmer filters during indexing and search analysis.

< Benefits and limitations >​

Stemming can retrieve relevant documents that use a different word form, but it can also merge words whose meanings differ. Its effect on search quality depends on the queries and collection. The information-retrieval textbook discusses this tradeoff between finding more matches and retaining precise matches.

Stemming rules are language-specific. Stemming also does not inherently match synonyms such as car and automobile; synonym expansion or semantic retrieval addresses that separate problem.

Comparison​

< Stemming vs. lemmatization >​

AspectStemmingLemmatization
MethodApplies word-form rules, such as suffix removalUses vocabulary and morphological analysis, often with part-of-speech information
OutputA stem that may not be a dictionary wordA dictionary base form, called a lemma
Examplestudies β†’ studi with Porter stemmingstudies β†’ study
PurposeGroup word variants for matchingRecover a word's linguistic base form

See Introduction to Information Retrieval for the distinction and its implications for retrieval.

Implementation​

< Porter stemming in Python >​

Stemming is the general technique; Porter is a specific English stemming algorithm that applies suffix-reduction rules. The Natural Language Toolkit (NLTK) provides PorterStemmer; its default mode includes NLTK's extensions to the original algorithm. See the NLTK PorterStemmer documentation and the Porter algorithm description.

Install NLTK once:

python -m pip install nltk
  • Example: stem individual words with stem(). This example needs no downloaded corpora or trained model.

    from nltk.stem import PorterStemmer

    stemmer = PorterStemmer()
    words = ["connect", "connected", "connecting", "connections", "studies"]

    for word in words:
    print(f"{word} β†’ {stemmer.stem(word)}")

    Expected output:

    connect β†’ connect
    connected β†’ connect
    connecting β†’ connect
    connections β†’ connect
    studies β†’ studi

The connection-related forms share the matching key connect. The result studi illustrates that a stem need not be a dictionary word. For full sentences, tokenize the text first, then stem each token.

Reference​