14 posts in total
2026
llama.cpp Workflow: Download, Convert, Quantize, and Serve Models
Compile NEFF Executables from PyTorch Code
vLLM Platform System
How the KV Cache Works in HuggingFace Transformers
Local CUDA vLLM Setup for Python-Only Development Using a Precompiled Wheel
vLLM Internals — PagedAttention and Custom Accelerator Compilation
Hugging Face Models: Repositories, Serving, and the Inference Engine Landscape
Exporting Compute Graphs, LLM Shape Dynamics, and Serving Runtimes
Schedules in Machine Learning Computation: What They Are and Who Needs to Know About Them
Main Takeaways from a Group Discussion on AI Coding