34 posts in total
2026
vLLM Internals — PagedAttention and Custom Accelerator Compilation
Hugging Face Models: Repositories, Serving, and the Inference Engine Landscape
Exporting Compute Graphs, LLM Shape Dynamics, and Serving Runtimes
Schedules in Machine Learning Computation: What They Are and Who Needs to Know About Them
Important Locations on Jailbroken iOS
Learning MLIR and HLO by Building a Tiny StableHLO-to-LLVM IR Compiler
Using MLIR as a C++ Library with a Relocatable Install
Docker and Podman Containers as Lightweight VMs for Interactive Work
Kleene Algebra, NetKAT, StacKAT, GKAT, CF-GKAT
PyTorch + CUDA vs. XLA + TPU: Two Execution Models for ML Systems