Skip to content

Large Language Models Explained: Architecture, Training, and Limits

Large language models (LLMs): how transformer self-attention and pretraining work, what context windows limit, and why LLMs produce hallucinations despite scale.

Google Gemini large language model platform
Google Gemini · Credit: Google

A large language model is a neural network that predicts the next token. Large language models, or LLMs, power chat assistants, coding tools, and search, and they handle translation, summarization, code generation, and multi-step reasoning. LLMs achieve this breadth through scale rather than task-specific programming: the same trained weights handle a contract clause and a Python function without separate rule sets for either. LLMs now sit behind chat assistants, coding tools, and search, yet the mechanics that make them fluent are also the source of their sharpest limits. The boundaries that matter most are structural, set by the transformer architecture, the size of the context window, and the way pretraining fixes knowledge in place. Understanding those three constraints is the difference between using an LLM well and trusting it where it should not be trusted.

What an LLM Actually Is

Hugging Face Open LLM Leaderboard showing model size variants and benchmark scores
Open LLM Leaderboard · Credit: Hugging Face

A large language model is a neural network trained on enough text that it learns the statistical relationships between tokens well enough to generate fluent, contextually appropriate language. The training signal is deceptively narrow. The model reads a sequence of tokens and predicts the most likely next token, then repeats that step to produce a sentence, a paragraph, or a working code block. This is next-token prediction, and it is the only thing an LLM is directly optimized to do. Fluency, factual recall, apparent reasoning, and few-shot learning are downstream effects of doing that one task across trillions of tokens. The transformer architecture is what makes the approach scale, and the size of the training data is what turns a next-token predictor into something that can hold a conversation. For the wider family of systems this sits inside, see how modern generative AI systems work.

  • Token: the unit an LLM reads and writes, usually a word fragment rather than a full word.
  • Next-token prediction: the training objective, predicting the most probable next token given everything before it.
  • Parameters: the learned weights that encode statistical patterns; modern large language models hold billions of them.
  • Transformer architecture: the network design that lets the model relate every token to every other token in the input.

How the Transformer Architecture Works

Video thumbnail shows How Large Language Models Work
How Large Language Models Work. Video: IBM Technology via YouTube.

The transformer is the architectural backbone of virtually every production LLM, built around a mechanism called self-attention that lets the model weigh how relevant each token in the input is to every other token before producing an output. Earlier sequence models read text left to right and lost track of distant words. Self-attention removes that bottleneck by computing, for every token, a weighted view of the entire input at once. When a model resolves what "it" refers to three sentences back, self-attention is the operation carrying that link. Stacking many attention layers lets the transformer architecture build progressively abstract representations, from surface grammar in early layers to meaning and intent in later ones. Production systems also refine the base design: Google's Gemma 4 model card describes a hybrid attention mechanism that interleaves local sliding-window attention with full global attention, keeping the final layer always global so the model retains a view of the whole sequence (Google AI for Developers, Gemma 4 model card).

NVIDIA H100 GPU data center hardware used for training large language models at scale
H100 Tensor Core GPU · Credit: NVIDIA
  1. Tokenize: the input text is split into tokens and mapped to numeric vectors.
  2. Attend: self-attention scores how strongly each token should influence every other token.
  3. Transform: stacked attention and feed-forward layers refine those representations into richer meaning.
  4. Predict: the final layer outputs a probability distribution over the vocabulary for the next token.
  5. Repeat: the chosen token is appended to the input and the cycle runs again until the response is complete.

Pretraining and Fine-Tuning: The Two-Stage Pipeline

LLM training runs in two distinct phases: pretraining on an enormous unlabeled corpus followed by fine-tuning on a smaller supervised dataset, a pipeline OpenAI used in its early GPT work to show that the same core model can adapt to very different tasks with minimal modification. OpenAI described the approach as training a transformer on a very large amount of data in an unsupervised manner using language modeling as the signal, then fine-tuning that model on much smaller supervised datasets to solve specific tasks (OpenAI, Improving Language Understanding by Generative Pre-Training). Pretraining is where the model absorbs grammar, facts, and reasoning patterns through next-token prediction across trillions of tokens. Fine-tuning is far cheaper and reshapes that general capability toward a narrower goal, whether following instructions, writing code, or matching a support tone. The economic consequence is direct: the expensive pretraining run happens once, and many specialized models branch off it. OpenAI reported that the same core model could be fine-tuned for very different tasks with minimal adaptation, which is why a single base model now underpins assistants, classifiers, and coding tools alike. Scale also unlocked few-shot learning, where a model handles a new task from a handful of examples placed in the prompt rather than a fresh training run, a behavior GPT-3 first demonstrated at scale.

  1. Pretraining: unsupervised learning on a massive unlabeled corpus, optimizing next-token prediction to absorb broad language patterns.
  2. Fine-tuning: supervised training on a smaller curated dataset that specializes the pretrained model for a target task.
  3. Alignment: further tuning, often with human feedback, that shapes how the model responds before it ships.

Context Windows and the Limits of LLM Memory

A context window defines the maximum number of tokens an LLM can process in one inference call, a boundary that acts like short-term memory for the model, per Google's Gemini API documentation. Google frames the context window directly as an analogy for short-term memory, and notes that large language models were historically limited by how much text could be passed in at once (Google AI for Developers, Long context). Anything outside the window does not exist for the model on that call. Early generative models could process only about 8,000 tokens at a time, per the same Google documentation, which forced developers to truncate or chunk long inputs. The window has since expanded sharply: Google's Gemini API documentation describes reaching 1 million tokens, enough to hold an entire codebase or a long document in working memory at once. A larger window reduces but does not erase the limit, and retrieval-augmented generation (RAG) remains the standard way to feed an LLM relevant external text on demand rather than packing everything into the prompt. The mechanics of why bigger windows help and where they still break are covered in how context window size shapes LLM performance and where it breaks.

  • Context window: the maximum tokens an LLM reads in a single inference call, covering both prompt and response.
  • Token limit: the hard ceiling; content beyond it is invisible to the model and silently dropped.
  • Retrieval-augmented generation (RAG): fetching relevant external text at query time so the model works from current, specific sources.

Architecture Variants: Dense, MoE, and Beyond

Not all LLMs share the same internal structure: dense models activate all parameters for every token, while Mixture of Experts (MoE) models route each token to a subset of specialized sub-networks. The trade-off is compute against capacity. A dense model is simpler and predictable but pays the full parameter cost on every token. An MoE model can hold far more total parameters while activating only a fraction per token, which raises capacity without a proportional rise in compute. Both designs are now shipping in the same product family: Google's Gemma 4 model card states that Gemma 4 features both Dense and Mixture of Experts architectures, suiting it to text generation, coding, and reasoning (Google AI for Developers, Gemma 4 model card). Scaling laws explain why these choices matter. Larger models trained on more data tend to perform better in measurable, repeatable ways, and architecture choices like MoE are partly a response to the cost pressure that scaling laws create.

AttributeDense modelMixture of Experts (MoE)
Parameters used per tokenAll parametersA routed subset of experts
Compute per tokenHigher, scales with sizeLower relative to total capacity
Total capacity ceilingLimited by compute budgetMuch higher for the same compute
Routing complexityNoneAdds a routing layer to manage

Where LLMs Fail: Hallucinations, Opacity, and Adaptation Costs

Despite rapid capability growth, large language models carry well-documented failure modes that stem directly from their architecture. The model predicts statistically likely text, not verified facts, so it can produce a fluent, confident answer that is simply wrong. This hallucination is a property of next-token prediction, not a bug awaiting a patch, and the same property that makes the output sound authoritative is what makes the error hard to catch. Opacity compounds the problem. OpenAI has stated that language models have become more capable and more broadly deployed, but that understanding of how they work internally is still very limited (OpenAI, Language models can explain neurons in language models). Interpretability research is active rather than settled: Anthropic's work on tracing the internal computation of a model shows how much effort it takes to read even partial reasoning out of the weights (Anthropic, Tracing the thoughts of a language model). The mitigations are practical rather than total. Why models fabricate, and how to reduce it, is covered in why AI models produce hallucinations.

  • Hallucination: confident, plausible output that is factually wrong, rooted in next-token prediction rather than fact retrieval.
  • Opacity: limited insight into why a model produced a given answer, which makes errors hard to predict or audit.
  • Adaptation cost: keeping a model current means fine-tuning or retrieval, since pretraining freezes knowledge at a cutoff.
  • Mitigation: retrieval-augmented generation and human review reduce risk but do not eliminate the underlying failure modes.

How Today's Production LLMs Compare

Major large language models in active deployment illustrate how the architectural trade-offs above play out across context window size, modality support, and task positioning. The clearest recent shift is context length. OpenAI's developer changelog reports that GPT-5.5 supports a 1 million token context window along with image input and tool integrations such as function calling and web search (OpenAI Developer Platform, API changelog). Google's Gemma 4 model card documents a different profile: context windows of 128K tokens on small models and 256K on medium models, multimodal text and image input with audio on the small tier, and a mix of Dense and MoE architectures (Google AI for Developers, Gemma 4 model card). The table below maps these vendor-documented figures side by side. Whether to reach for an open-weight model or a hosted one depends on those trade-offs, covered in open-weight versus closed AI model trade-offs for builders, and the modality differences are explored in how multimodal AI models combine text, image, and audio.

ModelContext windowModalityArchitecture note
GPT-5.5 (OpenAI)1M tokensText and image inputTool integrations including function calling and web search
Gemma 4 small (Google)128K tokensText, image, audio inputDense and MoE; hybrid sliding-window and global attention
Gemma 4 medium (Google)256K tokensText and image inputDense and MoE architectures

References

Frequently Asked Questions

What is a context window in a large language model?

A context window is the maximum number of tokens an LLM can read at once during a single inference call. Models with small windows lose access to earlier parts of a conversation; models with large windows (1 million tokens or more, per Google's Gemini API documentation) can hold entire codebases or long documents in working memory at once.

What is the difference between pretraining and fine-tuning an LLM?

Pretraining is the initial phase where the model learns language patterns from a very large unlabeled dataset using unsupervised objectives such as next-token prediction. Fine-tuning is a subsequent, lower-cost phase where the pretrained model is further trained on a smaller supervised dataset to specialize it for a particular task, as OpenAI described for its early GPT work.

Why can LLMs still produce incorrect or fabricated outputs?

LLMs generate text by predicting statistically likely continuations of input tokens rather than retrieving verified facts. OpenAI has stated that understanding of how models work internally remains very limited, which means the model can produce plausible-sounding text that is factually wrong, a pattern commonly called hallucination and a fundamental architectural limit, not a fixable bug.

What is the Mixture-of-Experts architecture used in some LLMs?

Mixture of Experts (MoE) activates only a subset of a model's specialized sub-networks for each input token rather than running all parameters every time. Google's Gemma 4 model card confirms Gemma 4 uses both Dense and MoE architectures, with MoE enabling larger total parameter counts without a proportional increase in compute per token.

Share this guide

Julian Beaumont

Julian Beaumont covers artificial intelligence and large language models for techshooked, following the path from research paper to deployed feature. His standard is anti-hype: ask what a model actually does, what trained it, how it fails, and whether a benchmark measures what the announcement claims.