LOCAL LLM FIELD GUIDE
ทำไม tok/s ของ LLM มักติด Memory Bandwidth
ระหว่าง decode ระบบต้องอ่าน weights และ state ที่จำเป็นเพื่อสร้าง token ถัดไปทีละตัว จึงมักติดการเคลื่อนข้อมูลมากกว่า FLOPS. ตัวเลขจริงยังถูกปรับด้วย framework, kernel, attention และ speculative decoding.

ROOFLINE PERFORMANCE MODEL
สูตรนี้เป็น model สำหรับวางแผน ไม่ใช่ benchmark
ให้คิดว่า tok/s เริ่มจากข้อมูลที่ hardware เคลื่อนย้ายได้ แล้วถูกปรับด้วยเส้นทาง software และจำนวน bytes ที่อ่าน.


ใช้กับงานจริง
อ่าน tok/s ให้ตรงกับสิ่งที่วัด
Decode สร้าง output ทีละ token จึงมักอ่าน weights และ KV cache ซ้ำ. Prefill รับ prompt หลาย token พร้อมกันและมีคอขวดต่างกัน.
ตัวเลขจาก calculator ใช้สำหรับจัดลำดับทางเลือก. ใช้ benchmark เดียวกันยืนยันก่อนตัดสินใจเรื่อง latency, capacity หรือค่าใช้จ่าย.
| ตัวชี้วัด | ใช้ตอบคำถาม | ไม่แทนค่า |
|---|---|---|
| TTFT | ผู้ใช้รอก่อน token แรกนานเท่าไร | ความเร็วระหว่าง generate |
| ITL | ช่วงห่างระหว่าง output token | เวลา prefill |
| output tok/s | ความเร็ว decode ของหนึ่ง request | aggregate throughput |
| peak memory | headroom ก่อน OOM | คุณภาพคำตอบ |
เงื่อนไขที่ต้องคงเดิมเมื่อเปรียบเทียบ
- model, quant และ runtime เดียวกัน
- working context และ batch เดียวกัน
- prompt, sampling และ power mode ใกล้กัน
- วัด TTFT, ITL, tok/s และ peak memory แยกกัน
readPerStepGB = (activeWeightsBytes + N * kvReadPerStreamBytes) / 1e9
tok/s_total = bandwidthGBps * frameworkEff * quantKernelEff
* attentionEff * specMult(N) * N / readPerStepGBDecode ไม่เหมือน prefill
Prefill ประมวลผล prompt หลาย token พร้อมกันและอาจใช้ compute ได้มากกว่า. Decode สร้างทีละ token จึงอ่าน weights ซ้ำและมักอ่อนไหวต่อ bandwidth. อย่าตี tok/s เดียวว่าแทน latency ทุกช่วง.
Bytes per token คือแกนของโมเดล
เมื่อ precision ของ weights ลดลง bytes ที่ต้องอ่านต่อ token อาจลดลง แต่ผลจริงต้องมี kernel และ memory access ที่เหมาะสม. Context ยาวยังเพิ่ม attention/KV work จน throughput ลดได้.
อ่านผล estimate
Calculator ใช้ hardware bandwidth เป็นจุดตั้งต้นแล้วปรับตาม framework, quant-kernel, attention และ spec decode. นี่คือ planning estimate ไม่ใช่คำสัญญา benchmark. request mix และ clock/power มีผล.
ยืนยันด้วย workload
วัด TTFT, inter-token latency, output tok/s, concurrency และ memory peak แยกกัน. ใช้ prompt, context, batching และ sampling ใกล้ production เพื่อรู้ว่า bottleneck เปลี่ยนหรือไม่.
แหล่งอ้างอิง
ลองกับ configuration ของคุณ
นำ model, quantization, context, framework และ hardware ที่คิดไว้ไปลองใน calculator จากนั้นยืนยันด้วย benchmark ที่ใกล้ production.
เปิดเครื่องคำนวณ