A diffusion model is a generative model that learns to reconstruct a coherent image by iteratively reversing a noise-injection process applied during training. The approach inverts a deceptively simple intuition from the score-based SDE paper by Song et al.: creating noise from data is easy, while creating data from noise is the generative task (Song et al., Score-Based Generative Modeling through Stochastic Differential Equations). That single inversion sits behind Stable Diffusion, DALL-E, Imagen, and the text-to-image systems that have reset public expectations for image synthesis. The internal machinery is consistent across vendors: a forward process that turns clean data into Gaussian noise, a reverse process that walks that path back, and a neural network trained to predict the noise at each step. Understanding the forward and reverse process, the score function, the role of the U-Net, and the latent-space compression that powers Stable Diffusion is the difference between treating a diffusion model as a black box and reasoning about why a sampler choice changes the output.
What a Diffusion Model Actually Does
The forward process of a diffusion model is a fixed Markov chain that progressively adds small amounts of Gaussian noise to a real training image across a predefined number of timesteps until the original image is completely obscured. Each step is independent of the model's weights, which is what makes the design tractable: the trajectory from clean image to pure noise is defined entirely by a variance schedule chosen up front. Ho et al. ran their DDPM with 1,000 forward steps on CIFAR-10, then trained a neural network to invert that exact trajectory, reporting an Inception score of 9.46 and a then-leading FID of 3.17 on the unconditional CIFAR-10 dataset (Ho et al., Denoising Diffusion Probabilistic Models). One useful property follows from the Markov-chain structure: any intermediate noisy version of an image can be sampled directly from the original in a single closed-form step, which means training does not require simulating the full chain.
- Start from a real image: a clean training sample drawn from the dataset acts as the chain's origin.
- Apply a noise schedule: a predefined variance schedule sets how much Gaussian noise is added at each timestep.
- Advance the Markov chain: each timestep adds a small noise increment, conditioning only on the previous state.
- Reach a noise terminus: after the full schedule, the image is statistically indistinguishable from pure Gaussian noise.
- Sample arbitrary timesteps: training draws a random timestep and computes its noisy image in closed form for efficiency.
The Reverse Process: How the Model Learns to Denoise
The reverse process is where a diffusion model does its real work: a neural network, typically a U-Net, learns to predict and subtract the noise added at each timestep so that a pure noise input gradually resolves into a recognizable image. The training objective for DDPM is straightforward in practice. Pick a real image, pick a random timestep, add the prescribed noise in closed form, then ask the U-Net to recover the noise component from the noisy image. The loss is the squared error between the predicted noise and the actual noise added. Once trained, the network knows how to walk any noisy image one step closer to a clean one, and sampling chains that prediction across the schedule. Ho et al. reported that this same training recipe extended beyond CIFAR-10, producing sample quality on 256x256 LSUN comparable to ProgressiveGAN (Ho et al., Denoising Diffusion Probabilistic Models). The U-Net is well suited to the task because its encoder-decoder shape with skip connections preserves spatial detail while still aggregating wide context, which a denoising step needs at every resolution.
- Sample pure noise: inference begins from a tensor drawn from a standard Gaussian, matching the forward terminus.
- Predict the noise: the U-Net receives the noisy image and the current timestep and outputs an estimate of the noise component.
- Subtract a denoising step: the predicted noise is removed according to the sampler's update rule, advancing the chain one step toward a clean image.
- Iterate across timesteps: the procedure repeats from the last step to the first, refining the image at every pass.
- Read the final sample: the result of the last denoising step is the generated image, ready for any further post-processing.
Score-Based Generative Modeling and the SDE Framework
A diffusion model can be reframed as a score-based generative model, where the network learns the gradient of the log-probability of the noisy data distribution rather than predicting noise directly. Song et al. formalized this view with a stochastic differential equation that smoothly transforms a complex data distribution to a known prior distribution by slowly injecting noise, paired with a reverse-time SDE that transforms the prior back into the data distribution by slowly removing the noise (Song et al., Score-Based Generative Modeling through Stochastic Differential Equations). The reverse-time SDE depends only on the time-dependent gradient field, the score, of the perturbed data distribution, which is the quantity the network is trained to estimate. Song et al. showed that this framework encapsulates previous approaches in score-based generative modeling and diffusion probabilistic modeling and reported record-breaking performance on CIFAR-10 with an Inception score of 9.89 and FID of 2.20, alongside high-fidelity 1024x1024 image generation for the first time from a score-based generative model. For an adjacent view of how neural networks absorb statistical patterns from data at scale, see how large language models learn statistical patterns from data.
- Score function: the gradient of the log-probability of a noisy data distribution, which the network learns to estimate.
- Forward SDE: a continuous-time stochastic differential equation that injects noise to push data toward a known prior.
- Reverse-time SDE: the matching equation that, given the score, transforms samples from the prior back into the data distribution.
- Predictor-corrector sampler: a hybrid sampling scheme introduced by Song et al. that corrects discretization error in the reverse-time SDE.
- Neural ODE counterpart: a deterministic ODE derived from the SDE that supports exact likelihood computation and more efficient sampling.
Latent Diffusion Models and Stable Diffusion
A latent diffusion model runs the entire forward and reverse diffusion process inside a compressed latent space rather than on raw pixels, which is the design that makes Stable Diffusion practical for consumer hardware. Pixel-space diffusion is expensive because the U-Net must operate at full image resolution at every denoising step. A latent diffusion model instead trains a separate autoencoder, typically a VAE, that maps images into a much smaller latent grid, runs the noise process there, then decodes the final clean latent back to pixels at the end. Cao et al.'s survey of diffusion models places this efficiency push at the center of the field's recent direction, categorizing diffusion-model research into efficient sampling, improved likelihood estimation, and handling data with special structures (Cao et al., A Survey on Generative Diffusion Models). Stable Diffusion is the best-known latent diffusion system, but the design generalizes: any image-domain pipeline that needs to keep memory and compute manageable benefits from running the chain in a lower-dimensional latent space.
| Property | Pixel-space diffusion model | Latent diffusion model |
|---|---|---|
| Denoising domain | Raw RGB pixels at full resolution | Compressed latent grid from a VAE encoder |
| U-Net compute per step | Scales with full image resolution | Scales with the latent grid, far smaller |
| Memory footprint | High; constrains training and inference | Reduced; fits consumer GPUs at higher resolutions |
| Example system | DDPM, Imagen pixel-stage decoder | Stable Diffusion and related latent text-to-image models |
| Decoder role | None; output is already pixels | VAE decoder maps the final clean latent back to an image |
Text-to-Image: Conditioning a Diffusion Model on Language

A text-to-image diffusion model steers the reverse process using a text embedding so that the denoised output reflects the content described in a prompt. The conditioning signal usually comes from a frozen text encoder such as CLIP, which maps a prompt into a vector the U-Net consumes alongside the noisy latent at every denoising step. Classifier-free guidance, the standard recipe across modern text-to-image systems, runs the network twice per step, once with the prompt and once without, then blends the two outputs at a guidance scale that trades prompt adherence against image diversity. The technique is broad enough that production vendors describe it as the steering primitive behind their text-to-image systems; OpenAI announced its 4o image generation as a model that aims for instruction-following and accurate text rendering inside image generation (OpenAI, Introducing 4o Image Generation). Cao et al. group these guidance methods under the broader theme of conditional generation, a category their survey covers alongside image synthesis, video generation, and molecule design (Cao et al., A Survey on Generative Diffusion Models).
- Text encoder: a model such as CLIP that turns a prompt into an embedding the diffusion model can read.
- Cross-attention conditioning: the mechanism that injects the text embedding into the U-Net at each denoising step.
- Classifier-free guidance: a two-pass sampling trick that blends conditional and unconditional predictions to amplify prompt adherence.
- Guidance scale: a scalar that controls how strongly the model follows the prompt versus producing diverse outputs.
- Negative prompts: an auxiliary unconditional input used at inference to push the reverse process away from unwanted features.
Sampling Methods and the Speed-Quality Tradeoff
Generating an image with a diffusion model requires running the reverse process sequentially from pure noise, and the number of sampling steps directly controls the tradeoff between generation speed and output quality. The original DDPM sampler used the same number of reverse steps as the forward chain, 1,000 in the Ho et al. setup, which produced high-quality samples but made inference slow. Subsequent samplers exploited the fact that the learned score does not require so many steps to produce sharp images. Song et al. derived a neural ODE counterpart to the SDE that samples from the same distribution while enabling exact likelihood computation and improved sampling efficiency (Song et al., Score-Based Generative Modeling through Stochastic Differential Equations). The cost of sampling is the dominant reason diffusion inference is heavy: each step is a full neural-network evaluation, and modern systems still typically run dozens to hundreds of them per image. That cost has become a research target in its own right; Cao et al.'s survey lists efficient sampling as one of three primary research directions for the field, alongside improved likelihood estimation and special-structure data (Cao et al., A Survey on Generative Diffusion Models). For a side-by-side look at how three of the best-known systems differ in practice, see how Midjourney, DALL-E, and Stable Diffusion compare as image generators.
| Sampler family | Typical steps | Underlying mechanism | Practical tradeoff |
|---|---|---|---|
| DDPM ancestral sampler | Hundreds to a thousand | Discretized reverse Markov chain | High sample quality, slowest inference |
| SDE solvers (Song et al.) | Hundreds | Numerical integration of the reverse-time SDE | Strong fidelity, predictor-corrector reduces error |
| Probability-flow ODE solvers | Tens to low hundreds | Deterministic ODE counterpart of the SDE | Faster, deterministic, supports likelihood computation |
| Few-step distilled samplers | Single digits to tens | Student model trained to mimic many-step output | Cheapest inference, mild quality cost on hard prompts |
References
- Ho et al., Denoising Diffusion Probabilistic Models (arXiv 2006.11239)
- Song et al., Score-Based Generative Modeling through Stochastic Differential Equations (arXiv 2011.13456)
- Cao et al., A Survey on Generative Diffusion Models (arXiv 2209.00796)
- OpenAI, Introducing 4o Image Generation
Further reading
Frequently Asked Questions
What is the difference between the forward and reverse process in a diffusion model?
The forward process gradually adds Gaussian noise to a training image over many steps until only noise remains. The reverse process trains a neural network to undo that noise step by step, reconstructing a coherent image from random noise at inference time.
How does a latent diffusion model differ from a standard diffusion model?
A latent diffusion model runs the denoising process in a compressed latent space rather than on raw pixel values. This cuts memory and compute significantly, which is why Stable Diffusion generates high-resolution images far faster than pixel-space models of comparable quality.
Why does generating an image with a diffusion model require many steps?
Each denoising step removes only a small amount of noise so that the reverse process stays stable and controllable. This multi-step design lets the model learn a precise reverse Markov chain, but it also means the neural network must run dozens to hundreds of times per image.
What is classifier-free guidance in text-to-image diffusion models?
Classifier-free guidance adjusts the reverse process to steer image generation toward a text prompt without requiring a separate classifier network. The model runs twice per step, once conditioned on the prompt and once unconditioned, and the outputs are blended at a guidance scale to balance prompt adherence against image diversity.
What is DDPM?
DDPM stands for Denoising Diffusion Probabilistic Model, the formulation introduced in the 2020 Ho et al. paper that defined the standard training objective for image diffusion models. The paper reported an Inception score of 9.46 and a FID of 3.17 on CIFAR-10, the benchmark that subsequent diffusion model research built on.









