CUDA / GPU Memory Hierarchy Deep Dive

Question

NVIDIA's deep-dive round goes below the algorithm into systems fundamentals. Expect open-ended GPU questions with multiple follow-ups.

What this round covers

  • Memory coalescing and why it matters for throughput
  • Shared memory vs global vs registers; bank conflicts
  • How you profile a kernel to find bottlenecks (nsight, occupancy)
  • Synchronization (__syncthreads), reductions, warp divergence

Tips

Build understanding from first principles — these companies prefer engineers who can derive behavior, not recite it. Bring real low-level debugging stories.

added …
LeaderboardSalaryAccount