16 posts in total
2026
What Is a GPU Warp? SIMT, Thread IDs, Divergence, and Synchronization
Kernel Regression and the Pairwise Nature of Attention
Local CUDA vLLM Setup for Python-Only Development Using a Precompiled Wheel
Compile NEFF Executables from NKI Kernels
vLLM Internals — PagedAttention and Custom Accelerator Compilation
Hugging Face Models: Repositories, Serving, and the Inference Engine Landscape
Exporting Compute Graphs, LLM Shape Dynamics, and Serving Runtimes
Schedules in Machine Learning Computation: What They Are and Who Needs to Know About Them
Learning MLIR and HLO by Building a Tiny StableHLO-to-LLVM IR Compiler
Jie Liu's B-Exam: Abstractions and Optimizations for Sparse Tensor Computation on Modern Hardware