Skip to main content

πŸ“ Large Language Models (LLMs)

Description​

< What is it? >​

A large language model (LLM) is a neural network trained on a very large collection of text to model language. Most modern generative LLMs use a decoder-only Transformer and generate one token at a time.

β€œLarge” has no fixed parameter-count threshold. It usually means the model was trained at a scale large enough to acquire broad capabilities, such as writing, summarizing, translating, answering questions, and producing code. These capabilities are learned statistical patterns, not a database of guaranteed facts.

Key points​

< The basic task: predict the next token >​

Given earlier tokens, an LLM predicts a probability distribution for the next token:

pΞΈ(xt∣x1,…,xtβˆ’1)p_\theta(x_t\mid x_1,\ldots,x_{t-1})

Generation repeats that one operation:

prompt β†’ tokenize β†’ token IDs β†’ embeddings β†’ Transformer blocks
β†’ logits β†’ softmax probabilities β†’ choose a token β†’ append it β†’ repeat
  • Example: for Paris is the capital of, a model may assign high probability to the token France. It appends the selected token, then predicts the next one from the longer sequence.

Tokens are not necessarily whole words: a token can be a word piece, punctuation mark, number, or whitespace-prefixed text. See Embeddings for how token IDs become vectors and Softmax for how logits become probabilities.

< Pretraining learns from the text itself >​

During pretraining, a document provides both the input and its target. For example, the prefix The sky is can train the model to predict blue. This is called self-supervised learning because the next-token targets come from the text rather than from manually written labels.

The usual objective is next-token cross-entropy loss:

L=βˆ’1Tβˆ‘t=1Tlog⁑pΞΈ(xt∣x<t)\mathcal{L} = -\frac{1}{T} \sum_{t=1}^{T} \log p_\theta(x_t\mid x_{<t})

Backpropagation changes the model’s weights so the observed next token becomes more likely. Repeating this over vast and varied text teaches grammar, facts, writing styles, and useful patterns of reasoningβ€”but it does not make every generated statement correct or current.

< From a base model to an assistant >​

Pretraining creates a base model that is good at continuing text. Later training changes its behavior into a usable assistant or specialist:

Model or systemMain additional stepTypical result
Base modelNext-token pretrainingContinues text, but may not reliably follow an instruction
Instruct modelPost-training, such as supervised fine-tuning and preference optimizationAnswers questions, follows formats, and learns desired refusal behavior
Specialist modelFine-tuning on a task or domainAdapts to a particular writing style, workflow, or narrow task
RAG applicationRetrieves external documents at inference timeCan use current or private information with sources

Fine-tuning changes model weights; retrieval-augmented generation (RAG) supplies relevant information in the prompt without changing them.

< Choosing the next token >​

The model’s final layer produces one logit for every vocabulary token. Softmax converts those logits into probabilities, then a decoding strategy selects the next token:

StrategyChoiceTypical effect
Greedy decodingAlways choose the most probable tokenRepeatable, but can be bland or get stuck
Temperature samplingRescale the distribution before samplingLower temperature is more predictable; higher temperature is more varied
Top-pp samplingSample only from the smallest set whose cumulative probability reaches ppAvoids sampling extremely unlikely tokens

Sampling explains why the same prompt can produce different answers. It also means a fluent answer is not, by itself, evidence that the answer is true.

< Context, knowledge, and limitations >​

An LLM works from the tokens in its context windowβ€”the prompt, conversation history, retrieved documents, and any generated tokens so far. It does not search its full training corpus at answer time.

  • Hallucination: the model can generate plausible but unsupported or false content.
  • Knowledge freshness: training knowledge can be old, incomplete, or missing; RAG and tools can provide newer evidence.
  • Context limits: long prompts cost more computation and can still omit or dilute important information.
  • No automatic verification: citations, calculations, code execution, and tool results need independent checking.

For an application needing reliable current facts, use retrieval, tools, and evaluation around the LLM instead of treating its generated text as the sole source of truth.

Comparison​

< Common Transformer language-model architectures >​

ArchitectureAttention patternTypical objectiveCommon use
Decoder-only (GPT-like)Causal: each token sees earlier tokens onlyNext-token predictionOpen-ended generation, chat, code completion
Encoder-only (BERT-like)Bidirectional: tokens can use context on both sidesMasked-token predictionSearch, classification, token representations
Encoder–decoder (T5-like)Encoder reads the input; decoder attends to encoded input and earlier output tokensConditional next-token predictionTranslation, summarization, structured text-to-text tasks

In everyday usage, β€œLLM” often refers to a decoder-only generative model, but the term can also include large encoder-only and encoder–decoder language models.

Reference​