Persistent Threads and Megakernels
Reading about recent megakernels for LLM inference brought me back to persistent threads and megakernels I used in my own GPU work. This post traces both ideas from their origins in ray tracing to modern attention kernels and full-model LLM inference