π Large Language Models (LLMs)
Descriptionβ
< What is it? >β
A large language model (LLM) is a neural network trained on a very large collection of text to model language. Most modern generative LLMs use a decoder-only Transformer and generate one token at a time.
βLargeβ has no fixed parameter-count threshold. It usually means the model was trained at a scale large enough to acquire broad capabilities, such as writing, summarizing, translating, answering questions, and producing code. These capabilities are learned statistical patterns, not a database of guaranteed facts.
Key pointsβ
< The basic task: predict the next token >β
Given earlier tokens, an LLM predicts a probability distribution for the next token:
Generation repeats that one operation:
prompt β tokenize β token IDs β embeddings β Transformer blocks
β logits β softmax probabilities β choose a token β append it β repeat
- Example: for
Paris is the capital of, a model may assign high probability to the tokenFrance. It appends the selected token, then predicts the next one from the longer sequence.
Tokens are not necessarily whole words: a token can be a word piece, punctuation mark, number, or whitespace-prefixed text. See Embeddings for how token IDs become vectors and Softmax for how logits become probabilities.
< Pretraining learns from the text itself >β
During pretraining, a document provides both the input and its target. For example, the prefix The sky is can train the model to predict blue. This is called self-supervised learning because the next-token targets come from the text rather than from manually written labels.
The usual objective is next-token cross-entropy loss:
Backpropagation changes the modelβs weights so the observed next token becomes more likely. Repeating this over vast and varied text teaches grammar, facts, writing styles, and useful patterns of reasoningβbut it does not make every generated statement correct or current.
< From a base model to an assistant >β
Pretraining creates a base model that is good at continuing text. Later training changes its behavior into a usable assistant or specialist:
| Model or system | Main additional step | Typical result |
|---|---|---|
| Base model | Next-token pretraining | Continues text, but may not reliably follow an instruction |
| Instruct model | Post-training, such as supervised fine-tuning and preference optimization | Answers questions, follows formats, and learns desired refusal behavior |
| Specialist model | Fine-tuning on a task or domain | Adapts to a particular writing style, workflow, or narrow task |
| RAG application | Retrieves external documents at inference time | Can use current or private information with sources |
Fine-tuning changes model weights; retrieval-augmented generation (RAG) supplies relevant information in the prompt without changing them.
< Choosing the next token >β
The modelβs final layer produces one logit for every vocabulary token. Softmax converts those logits into probabilities, then a decoding strategy selects the next token:
| Strategy | Choice | Typical effect |
|---|---|---|
| Greedy decoding | Always choose the most probable token | Repeatable, but can be bland or get stuck |
| Temperature sampling | Rescale the distribution before sampling | Lower temperature is more predictable; higher temperature is more varied |
| Top- sampling | Sample only from the smallest set whose cumulative probability reaches | Avoids sampling extremely unlikely tokens |
Sampling explains why the same prompt can produce different answers. It also means a fluent answer is not, by itself, evidence that the answer is true.
< Context, knowledge, and limitations >β
An LLM works from the tokens in its context windowβthe prompt, conversation history, retrieved documents, and any generated tokens so far. It does not search its full training corpus at answer time.
- Hallucination: the model can generate plausible but unsupported or false content.
- Knowledge freshness: training knowledge can be old, incomplete, or missing; RAG and tools can provide newer evidence.
- Context limits: long prompts cost more computation and can still omit or dilute important information.
- No automatic verification: citations, calculations, code execution, and tool results need independent checking.
For an application needing reliable current facts, use retrieval, tools, and evaluation around the LLM instead of treating its generated text as the sole source of truth.
Comparisonβ
< Common Transformer language-model architectures >β
| Architecture | Attention pattern | Typical objective | Common use |
|---|---|---|---|
| Decoder-only (GPT-like) | Causal: each token sees earlier tokens only | Next-token prediction | Open-ended generation, chat, code completion |
| Encoder-only (BERT-like) | Bidirectional: tokens can use context on both sides | Masked-token prediction | Search, classification, token representations |
| Encoderβdecoder (T5-like) | Encoder reads the input; decoder attends to encoded input and earlier output tokens | Conditional next-token prediction | Translation, summarization, structured text-to-text tasks |
In everyday usage, βLLMβ often refers to a decoder-only generative model, but the term can also include large encoder-only and encoderβdecoder language models.
Related ideasβ
- Transformer explains the architecture underlying most modern generative LLMs.
- Embeddings explains how text tokens become model inputs.
- Post-Training explains how a base model becomes an instruct model.
- Fine-Tuning adapts a model to a specific task or domain.
- Prompt Engineering shapes behavior without changing model weights.
- Retrieval Augmented Generation (RAG) grounds an LLM with external information.
Referenceβ
- Attention Is All You Need (Vaswani et al., 2017)
- Language Models are Few-Shot Learners (Brown et al., 2020)
- Training language models to follow instructions with human feedback (Ouyang et al., 2022)
- LLM Visualization (Ben Bycroft)
- Foundations of Large Language Models - crash courses (Numeryst | Youtube playlist)