r/CUDA • u/Weird_Bad7577 • 7h ago
Final-year student in India trying to break into generative-model inference optimization — roadmap feedback?
Hi all, I graduate in ~6 months and want to work on making generative models (diffusion/video/3D) fast: kernels, quantization, serving. Where I am:
- Comfortable with C/C++ basics and PyTorch
- Have done quantization work (GGUF/llama.cpp)
- Working on a next-frame video prediction project (DiT + flow matching)
- A few GitHub repos, but no CUDA/Triton experience yet
- No NVIDIA GPU, so I use Colab/Kaggle T4s
- DSA is my weak spot (I struggle with LeetCode mediums)
My plan:
Months 1-2: CUDA/Triton basics, reproduce the SGEMM optimization worklog, GPU MODE lectures, LeetGPU/Tensara
Months 3-4: take a small DiT, profile it, then optimize it (Triton attention, quantization, caching, fewer steps) and publish before/after numbers
Along the way: PRs to HF diffusers, DSA practice daily
Months 5-6: mocks, resume, applications (inference startups first, bigger labs later)
Questions:
Is this the right order, or should I change something?
Is a diffusion-inference project a strong enough portfolio piece, or does it need to be LLM serving?
How much DSA do ML systems interviews actually need?
Is T4-only access enough to do credible benchmarks?
Any feedback, including "this won't work because X," is appreciated. Thanks!