29 posts in total
2026
Memory Models and Forward-Progress Guarantees
Introduction to OpenMP
What Is a GPU Warp? SIMT, Thread IDs, Divergence, and Synchronization
vLLM Platform System
Local CUDA vLLM Setup for Python-Only Development Using a Precompiled Wheel
Compile NEFF Executables from NKI Kernels
What is S3?
vLLM Internals — PagedAttention and Custom Accelerator Compilation
Exporting Compute Graphs, LLM Shape Dynamics, and Serving Runtimes
Schedules in Machine Learning Computation: What They Are and Who Needs to Know About Them