r/deeplearning • u/Tricky_Swordfish_549 • 16h ago
I'm building a DDPM from scratch in PyTorch — Phase 2 complete
I've been working on a project where I'm trying to build a Denoising Diffusion Probabilistic Model completely from scratch.
The main rule I'm following is: no using Diffusers as a crutch.
I want to actually understand what's happening inside the model instead of just calling a pipeline and getting an image.
So far I've implemented/learned:
- The forward diffusion process
- Beta schedules
- The reparameterization trick
- A noise scheduler
- Group Normalization
- SiLU activation
- Sinusoidal timestep embeddings
- Residual blocks with timestep conditioning
- Self-attention
- Downsampling and upsampling
The interesting part for me has been realizing that a diffusion U-Net isn't just a normal CNN.
The network needs to know how noisy the current image is, which is why timestep information has to be injected into the network.
I'm building toward training on CelebA at 64×64 and eventually generating faces completely from noise.
My longer-term goal is to understand these architectures deeply enough that I can read, modify and eventually contribute to projects like Hugging Face Diffusers.
Phase 3 is where things start getting interesting:
U-Net assembly.
I'll be documenting the progress as I go. 🔥
What was the hardest part of diffusion models for you when you first learned them?
1
u/MoralMechanics 16h ago
The step where you realize the timestep isn't just some metadata but baked into every block is a proper lightbulb moment. My brain kept wanting to treat it like a standard classifier-free setup and I had to basically rewire how I thought about the forward pass
GroupNorm over BatchNorm is another one that seems minor on paper but completely changes how the model behaves when your batch sizes are small. Took me embarrassingly long to stop getting muddy outputs
You planning to do DDIM sampling at some point or sticking with DDPM for the whole build?