
LOCAL LLM FIELD GUIDES
Choose local LLMs with evidence
Choose hardware, models, quantization, KV cache, serving frameworks, and performance paths for local LLMs.
START HERE
Start from the question you need answered
Quantization: what is reduced, and how to chooseChoose local LLM weight and KV-cache quantization from the runtime, memory budget, and quality risk.Why LLM tok/s often hits memory bandwidthWhy LLM decode tok/s is often memory-bandwidth-limited, and how to read calculator estimates responsibly.Choose Local LLM hardware from the job, not a GPU nameChoose local LLM hardware across NVIDIA, AMD, and Mac plus consumer, workstation, and datacenter GPUs by workload and memory budget.