๐ Named Entity Recognition (NER)
Descriptionโ
< What is it? >โ
Named Entity Recognition (NER) finds spans of text that name real-world entities and assigns each one a category, such as a person, organization, location, date, or product. It is commonly treated as a sequence-labeling or token-classification task.
-
Example: in โAva Chen joined Acme Labs in Berlin,โ an NER system may extract:
Text span Entity type Ava Chen person Acme Labs organization Berlin location
NER is a task, not one specific model. Rules, conditional random fields, recurrent networks, and Transformers can all perform it.
Key pointsโ
< Entity labels and the BIO scheme >โ
NER datasets often label tokens with the BIO scheme:
| Tag | Meaning |
|---|---|
B-PER | Beginning of a person entity |
I-PER | Inside the same person entity |
O | Outside every labeled entity |
- Example: โAva Chen joined Acme Labsโ can be labeled as
B-PER I-PER O B-ORG I-ORG. The token labels are then combined back into the spans โAva Chenโ and โAcme Labs.โ
< How a modern NER model works >โ
A Transformer encoder creates a contextual representation for each token . A small classification head predicts that token's entity label:
The surrounding words matter. For example, โJordan scoredโ is likely a person, while โJordan borders Iraqโ is likely a location.
< Example: extract entities with a pretrained model >โ
This example uses Hugging Face's token-classification pipeline. The aggregation strategy merges BIO-tagged token pieces into complete entity spans:
Input text: "Apple opened an office in London."
# pip install transformers torch
from transformers import pipeline
ner = pipeline(
task="token-classification",
model="dslim/bert-base-NER",
aggregation_strategy="simple",
)
text = "Apple opened an office in London."
for entity in ner(text):
print(
f"{entity['word']}: {entity['entity_group']} "
f"(score={entity['score']:.2f}, "
f"span={entity['start']}:{entity['end']})"
)
An illustrative result is:
word | entity_group | score | start:end |
|---|---|---|---|
Apple | ORG | approximately 0.99 | 0:5 |
London | LOC | approximately 0.99 | 26:32 |
The score is the model's confidence for the predicted entity type. The character offsets identify where the entity appears in the original text; exact scores can vary by model and version.
Note: entity_group is not a list of words. It is a category label selected from the model's finite label set. The text span can be many different wordsโfor example, Apple, Microsoft, or OpenAIโbut each may receive the same ORG (organization) group. For this checkpoint, the merged entity groups are typically PER, ORG, LOC, and MISC; another NER model may define a different set, such as DATE or PRODUCT.
With aggregation_strategy="simple", the pipeline combines BIO token labels such as B-ORG and I-ORG into one complete span, then reports the shared type as entity_group: "ORG". Tokens labeled O mean โnot an entityโ and are usually omitted from the output.
< Training data and tokenization >โ
Training examples need both the text and annotated entity spans. A tokenizer may split one word into several subword tokens, so labels must be aligned carefully; special tokens are ignored, and a common convention labels only the first subword of each original word.
< Evaluation and uses >โ
NER is usually evaluated with entity-level precision, recall, and F1. A predicted entity normally counts as correct only when both its span and entity type match the reference. Token accuracy alone can hide boundary mistakes.
Common uses include search, document extraction, content moderation, medical or legal document processing, and redacting personally identifiable information.
Comparisonโ
< NER vs text classification >โ
| Named Entity Recognition | Text classification | |
|---|---|---|
| Prediction unit | Tokens or text spans | A whole document, sentence, or input |
| Output | B-PER, I-PER, B-ORG, O, and similar labels | One or more document-level labels |
| Example | Find โAva Chenโ as a person | Classify a review as positive or negative |
| Typical challenge | Correct entity boundaries and token-label alignment | Choosing the correct class for the complete input |
Related ideasโ
- Classification introduces supervised classification more generally.
- Embeddings explains token representations.
- Transformer explains the encoder architecture commonly used for modern NER.
- Softmax explains how the model converts label scores into probabilities.
Referenceโ
- Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition
- Hugging Face: Token classification
- spaCy for NLP (geeksforgeeks.org)
- spaCy projects (spacy.io/usage/projects)
- explosion.ai/ner (explosion.ai)
- NER github repo (Tanish-Sarkar)
- Token classification (Huggingface.co)
- NER drugs (explosion.ai)
- token classification (Huggingface.co)
- NER with spaCy (NLP with spaCy | DataCamp)
- EntityRuler for NER (NLP with spaCy | DataCamp)
- NER in code (NLP in Python | DataCamp)