
LOCAL LLM FIELD GUIDES
Choose local LLMs with evidence
Choose hardware, models, quantization, KV cache, serving frameworks, and performance paths for local LLMs.
START HERE
Start from the question you need answered
Quantization: Where VRAM is Reduced and What Quality Trades AwayIn-depth guide to 4-bit and 8-bit quantization for weights and KV cache, summarizing QAT, MXFP4, AWQ, GPTQ, and GGUF with VRAM calculations, quality trade-offs, and hallucination risks.Why LLM tok/s Hits Memory Bandwidth BottlenecksDeep-dive into the Roofline Model for Local LLMs: why Decode is memory bandwidth bound, why Prefill is compute (FLOPS) bound, distinguishing TTFT, TPOT, ITL, and reading calculator results like a pro.Choose Local LLM Hardware by Workload, Not GPU NamesComprehensive guide to choosing hardware for Local LLMs. Summarizing pros and cons across Apple Silicon Mac, NVIDIA CUDA, and AMD ROCm with VRAM, memory bandwidth, specs, and Thailand market planning prices.