Persistent Threads and Megakernels

Reading about recent megakernels for LLM inference brought me back to persistent threads and megakernels I used in my own GPU work. This post traces both ideas from their origins in ray tracing to modern attention kernels and full-model LLM inference

Read More

Adaptive Software Rasterization with CUDA

A CUDA software rasterizer that runs the graphics pipeline entirely on the GPU. The framework manages rasterization and adaptive sampling in software, generating additional work when needed while achieving interactive frame rates.

Read More

Uncertain Transport in Unsteady Flows

We model uncertainty in time-dependent flows with stochastic differential equations and identify surfaces that resist or enhance diffusive transport. The method avoids expensive Monte Carlo simulation while also showing the absolute scale of uncertainty, making the resulting flow structures easier to interpret.

Read More

Stochastic Volume Rendering of Multi‐Phase SPH Data

We render large, unstructured SPH simulations directly, without first converting the particle data into a volume. Particle sampling guided by the view and local data complexity makes ray marching faster, allowing the method to scale from interactive previews to more accurate renderings with multi-phase and single-scattering effects.

Read More