LLM Cal

LOCAL LLM FIELD GUIDE

Speculative Decoding และ Attention Backend: เร็วขึ้นเมื่อเงื่อนไขตรง

Speculative decoding ใช้ draft model หรือวิธี draft token ล่วงหน้าแล้วให้ target model ตรวจพร้อมกัน. จะช่วยเมื่อ acceptance rate ดีและ overhead คุ้ม. attention backend ช่วยลดต้นทุนของ attention แต่ขึ้นกับ GPU, runtime, context และ kernel support.

มาสคอตแมวดำ STH ประกอบคู่มือ Speculative Decoding และ Attention Backend: เร็วขึ้นเมื่อเงื่อนไขตรง

VERIFY BEFORE YOU TRUST

draft เร็วขึ้นได้เมื่อ target รับ token ที่ draft มา

  1. วัด baseline: TTFT, ITL, output tok/s
  2. เปิด spec และดู acceptance rate/length
  3. เทียบ p95 และ cost บน traffic เดียวกัน
แมวดำ STH ตรวจ token ที่ร่างล่วงหน้าต่อ attention weave
แผนภาพสรุปสรุปกลไกสำหรับใช้เทียบกับค่าใน calculator
แผนภาพอธิบาย Speculative Decoding และ Attention Backend: เร็วขึ้นเมื่อเงื่อนไขตรง

ใช้กับงานจริง

เปิด feature แล้ววัดผลเป็นลำดับ

Speculative decoding ให้ draft เสนอ token แล้ว target ตรวจ. ผลดีเกิดเมื่อ token ที่ draft เสนอถูกยอมรับมากพอที่จะคุ้มกับงานเพิ่ม.

Attention backend เป็นเส้นทาง kernel และ memory access. ความเข้ากันได้ขึ้นกับ model, GPU, runtime release, context และ KV dtype.

สัญญาณค่าที่ดีสัญญาณให้หยุด
acceptance rate หรือ lengthdraft ที่ยอมรับต่อเนื่องยอมรับน้อยจน overhead เพิ่ม
TTFTไม่แย่ลงจนผู้ใช้รู้สึกได้first token ช้ากว่า baseline
ITL และ output tok/sdecode ดีขึ้นใน workload เดิมเร็วเฉพาะ short prompt
p95 และ peak memoryยังอยู่ใน SLO และ headroomtail แย่ลงหรือใกล้ OOM

ลำดับ benchmark ที่ควรทำ

  • วัด baseline โดยไม่เปิด spec decode
  • เปิด method ที่ runtime และ model รองรับ
  • ดู acceptance, TTFT, ITL, tok/s, p95 และ peak memory
  • เทียบ prompt, context และ concurrency ชุดเดิม
  • เก็บ feature เฉพาะเมื่อผลดีคงอยู่
บันทึกผลให้เปรียบเทียบได้
method = MTP
working_context = 32768
concurrency = 4
measure = TTFT, ITL, output_tok_s, p95, peak_memory

กลไก speculative decoding

draft model, MTP หรือ n-gram เสนอ token ต่อหน้า แล้ว target model ตรวจและยอมรับเป็นช่วง. ถ้ายอมรับน้อย งานตรวจและการจัดการ draft อาจไม่คุ้ม. จึงต้องวัด acceptance length/rate ควบคู่กับ latency.

อย่าดูความเร็วเฉลี่ยอย่างเดียว

speculation อาจช่วย output throughput แต่ TTFT, tail latency, sampling, tool calls และ prompt distribution เปลี่ยนผลได้. เปิดเฉพาะ model/runtime ที่มี support ชัดเจน และเทียบ baseline เดิมบน traffic เดียวกัน.

Attention backend

FlashAttention, FlashInfer และ backend รุ่นใหม่เป็นวิธีจัด memory access/kernel สำหรับ attention. ความเข้ากันได้ขึ้นกับ CUDA/ROCm/Metal, GPU architecture, context, KV dtype และรุ่น runtime.ไม่ใช่ toggle สากล.

วิธีเลือก

เริ่มจาก backend ที่ runtime แนะนำและรองรับ KV quant ที่คุณใช้. จากนั้น benchmark prefill, decode, long-context และ concurrency. ปิด feature เมื่อเพิ่ม complexity แต่ไม่ช่วย p95 หรือ cost ของ workload จริง.

แหล่งอ้างอิง

ลองกับ configuration ของคุณ

นำ model, quantization, context, framework และ hardware ที่คิดไว้ไปลองใน calculator จากนั้นยืนยันด้วย benchmark ที่ใกล้ production.

เปิดเครื่องคำนวณ