CUDA / GPU Memory Hierarchy Deep Dive
Question
NVIDIA's deep-dive round goes below the algorithm into systems fundamentals. Expect open-ended GPU questions with multiple follow-ups.
What this round covers
- Memory coalescing and why it matters for throughput
- Shared memory vs global vs registers; bank conflicts
- How you profile a kernel to find bottlenecks (nsight, occupancy)
- Synchronization (__syncthreads), reductions, warp divergence
Tips
Build understanding from first principles — these companies prefer engineers who can derive behavior, not recite it. Bring real low-level debugging stories.
added …