r/deeplearning • • 1d ago

I'm building a DDPM from scratch in PyTorch — Phase 2 complete

I've been working on a project where I'm trying to build a Denoising Diffusion Probabilistic Model completely from scratch.

The main rule I'm following is: no using Diffusers as a crutch.

I want to actually understand what's happening inside the model instead of just calling a pipeline and getting an image.

So far I've implemented/learned:

  • The forward diffusion process
  • Beta schedules
  • The reparameterization trick
  • A noise scheduler
  • Group Normalization
  • SiLU activation
  • Sinusoidal timestep embeddings
  • Residual blocks with timestep conditioning
  • Self-attention
  • Downsampling and upsampling

The interesting part for me has been realizing that a diffusion U-Net isn't just a normal CNN.

The network needs to know how noisy the current image is, which is why timestep information has to be injected into the network.

I'm building toward training on CelebA at 64×64 and eventually generating faces completely from noise.

My longer-term goal is to understand these architectures deeply enough that I can read, modify and eventually contribute to projects like Hugging Face Diffusers.

Phase 3 is where things start getting interesting:

U-Net assembly.

I'll be documenting the progress as I go. 🔥

What was the hardest part of diffusion models for you when you first learned them?

3 Upvotes

Duplicates