LLM Cal

LOCAL LLM FIELD GUIDE

ทำไม tok/s ของ LLM มักติด Memory Bandwidth

ระหว่าง decode ระบบต้องอ่าน weights และ state ที่จำเป็นเพื่อสร้าง token ถัดไปทีละตัว จึงมักติดการเคลื่อนข้อมูลมากกว่า FLOPS. ตัวเลขจริงยังถูกปรับด้วย framework, kernel, attention และ speculative decoding.

มาสคอตแมวดำ STH ประกอบคู่มือ ทำไม tok/s ของ LLM มักติด Memory Bandwidth

ROOFLINE PERFORMANCE MODEL

สูตรนี้เป็น model สำหรับวางแผน ไม่ใช่ benchmark

ให้คิดว่า tok/s เริ่มจากข้อมูลที่ hardware เคลื่อนย้ายได้ แล้วถูกปรับด้วยเส้นทาง software และจำนวน bytes ที่อ่าน.

tok/sBandwidth × Framework × Quant-Kernel × Attention × Spec-Decode÷Bytes Read per Token
แมวดำ STH กำลังนำกระแสข้อมูลผ่านหน่วยความจำ
แผนภาพสรุปสรุปกลไกสำหรับใช้เทียบกับค่าใน calculator
แผนภาพอธิบาย ทำไม tok/s ของ LLM มักติด Memory Bandwidth

ใช้กับงานจริง

อ่าน tok/s ให้ตรงกับสิ่งที่วัด

Decode สร้าง output ทีละ token จึงมักอ่าน weights และ KV cache ซ้ำ. Prefill รับ prompt หลาย token พร้อมกันและมีคอขวดต่างกัน.

ตัวเลขจาก calculator ใช้สำหรับจัดลำดับทางเลือก. ใช้ benchmark เดียวกันยืนยันก่อนตัดสินใจเรื่อง latency, capacity หรือค่าใช้จ่าย.

ตัวชี้วัดใช้ตอบคำถามไม่แทนค่า
TTFTผู้ใช้รอก่อน token แรกนานเท่าไรความเร็วระหว่าง generate
ITLช่วงห่างระหว่าง output tokenเวลา prefill
output tok/sความเร็ว decode ของหนึ่ง requestaggregate throughput
peak memoryheadroom ก่อน OOMคุณภาพคำตอบ

เงื่อนไขที่ต้องคงเดิมเมื่อเปรียบเทียบ

  • model, quant และ runtime เดียวกัน
  • working context และ batch เดียวกัน
  • prompt, sampling และ power mode ใกล้กัน
  • วัด TTFT, ITL, tok/s และ peak memory แยกกัน
รูปแบบ roofline ที่ calculator ใช้
readPerStepGB = (activeWeightsBytes + N * kvReadPerStreamBytes) / 1e9
tok/s_total = bandwidthGBps * frameworkEff * quantKernelEff
              * attentionEff * specMult(N) * N / readPerStepGB

Decode ไม่เหมือน prefill

Prefill ประมวลผล prompt หลาย token พร้อมกันและอาจใช้ compute ได้มากกว่า. Decode สร้างทีละ token จึงอ่าน weights ซ้ำและมักอ่อนไหวต่อ bandwidth. อย่าตี tok/s เดียวว่าแทน latency ทุกช่วง.

Bytes per token คือแกนของโมเดล

เมื่อ precision ของ weights ลดลง bytes ที่ต้องอ่านต่อ token อาจลดลง แต่ผลจริงต้องมี kernel และ memory access ที่เหมาะสม. Context ยาวยังเพิ่ม attention/KV work จน throughput ลดได้.

อ่านผล estimate

Calculator ใช้ hardware bandwidth เป็นจุดตั้งต้นแล้วปรับตาม framework, quant-kernel, attention และ spec decode. นี่คือ planning estimate ไม่ใช่คำสัญญา benchmark. request mix และ clock/power มีผล.

ยืนยันด้วย workload

วัด TTFT, inter-token latency, output tok/s, concurrency และ memory peak แยกกัน. ใช้ prompt, context, batching และ sampling ใกล้ production เพื่อรู้ว่า bottleneck เปลี่ยนหรือไม่.

แหล่งอ้างอิง

ลองกับ configuration ของคุณ

นำ model, quantization, context, framework และ hardware ที่คิดไว้ไปลองใน calculator จากนั้นยืนยันด้วย benchmark ที่ใกล้ production.

เปิดเครื่องคำนวณ