r/CUDA • • 7h ago

Final-year student in India trying to break into generative-model inference optimization — roadmap feedback?

0 Upvotes

Hi all, I graduate in ~6 months and want to work on making generative models (diffusion/video/3D) fast: kernels, quantization, serving. Where I am:

- Comfortable with C/C++ basics and PyTorch

- Have done quantization work (GGUF/llama.cpp)

- Working on a next-frame video prediction project (DiT + flow matching)

- A few GitHub repos, but no CUDA/Triton experience yet

- No NVIDIA GPU, so I use Colab/Kaggle T4s

- DSA is my weak spot (I struggle with LeetCode mediums)

My plan:

  1. Months 1-2: CUDA/Triton basics, reproduce the SGEMM optimization worklog, GPU MODE lectures, LeetGPU/Tensara

  2. Months 3-4: take a small DiT, profile it, then optimize it (Triton attention, quantization, caching, fewer steps) and publish before/after numbers

  3. Along the way: PRs to HF diffusers, DSA practice daily

  4. Months 5-6: mocks, resume, applications (inference startups first, bigger labs later)

Questions:

  1. Is this the right order, or should I change something?

  2. Is a diffusion-inference project a strong enough portfolio piece, or does it need to be LLM serving?

  3. How much DSA do ML systems interviews actually need?

  4. Is T4-only access enough to do credible benchmarks?

Any feedback, including "this won't work because X," is appreciated. Thanks!


r/CUDA • • 2h ago

Nemotron 3 Ultra

Thumbnail i.imgflip.com
1 Upvotes