บทที่ 27 · Part 7 — Production Engineering

Measure, Profile and Optimize

ใช้ benchmark, benchstat, pprof, trace, PGO และ GC evidence ก่อนแก้ performance

การเปลี่ยน struct ทั้งระบบเป็น pointer, เพิ่ม sync.Pool หรือย้ายจาก Chi ไป framework เร็วกว่าโดยยังไม่รู้ ว่า latency อยู่ที่ database เป็น optimization ที่เพิ่ม complexity โดยไม่แตะ bottleneck Performance work ที่ดีเริ่มจาก user-visible objective, วัด baseline, หาเหตุ แล้วเปลี่ยนทีละอย่าง

จบบทนี้คุณจะ

  • เขียน benchmark ด้วย b.Loop และเปรียบเทียบหลาย run ด้วย benchstat
  • เลือก CPU/heap/block/mutex profile หรือ execution trace ตามคำถาม
  • ใช้ PGO และ GC/runtime tuning หลังพิสูจน์ workload แล้ว

Performance Loop ที่ย้อนกลับได้

objective ต้องระบุ metric และ workload เช่น reduce P99 latency ของ submit endpoint ภายใต้ concurrency ที่กำหนดโดยไม่เพิ่ม error rate/memoryเกิน budget ไม่ใช้คำว่า “ทำให้เร็วขึ้น” เพราะ throughput, latency, CPU และ memory อาจแลกกัน

เริ่มจาก trace/metrics แยกเวลาที่รอ MySQL, Redis, network และ queue ออกจาก on-CPU ถ้า 90% อยู่ที่ query การลด allocationใน JSON 20% แทบไม่เปลี่ยน end-to-end latency แต่ยังเพิ่ม maintenance cost

Benchmark ให้สิ่งที่ตั้งใจวัด

Go 1.24+ ควรใช้ b.Loop() สำหรับ benchmark ใหม่ มันจัด timing loop และลดความเสี่ยง dead-code elimination:

func BenchmarkNormalizeBatchIDs(b *testing.B) {
    input := loadBatchIDs(1_000)
    b.ReportAllocs()

    for b.Loop() {
        NormalizeBatchIDs(input)
    }
}

ใช้ input sizes/distributions ที่แทน production และแยก setup ออกจาก measured body รันหลายครั้งบน environment ควบคุมแล้วใช้ benchstat ไม่สรุปจากตัวเลขครั้งเดียว:

go test -run=^$ -bench=BenchmarkNormalize -benchmem -count=10 ./internal/settlement \
    > before.txt
go test -run=^$ -bench=BenchmarkNormalize -benchmem -count=10 ./internal/settlement \
    > after.txt
go tool benchstat before.txt after.txt

Pin benchstat เป็น tool dependency ใน repo จริงและวัด variants แบบ serial เพราะ benchmark ที่รันพร้อมกัน แย่ง CPU/thermal budget ผลที่ไม่มี statistical significance ไม่ควรถูกเรียกว่า improvement

เลือก Diagnostic ให้ตรงคำถาม

คำถามเครื่องมือ
CPU อยู่ที่ function ใดCPU profile + go tool pprof
allocation churn มาจากไหนheap alloc_objects/alloc_space
memory ที่ยังอยู่มาจากไหนheap inuse_objects/inuse_space
goroutine รอ lock/channel/network เมื่อไรexecution trace, block/mutex/goroutine profile
compiler ทำให้ value escape/inlining ไหมgo build -gcflags=-m=2 หลัง profile ชี้จุด

pprof เป็น sampling/aggregate view ว่า ที่ไหน ใช้ resource ส่วน trace บอก timeline ว่า เมื่อไรและ ทำไม goroutine block/schedule/GC อย่าเปิด block/mutex profiling rate สูงตลอดโดยไม่วัด overhead และ ต้องป้องกัน profile endpoint ตามบทก่อน

Optimization Patterns หลังมีหลักฐาน

allocation profile อาจนำไปสู่ preallocation ที่มี bound, buffer reuse, streaming หรือเลิก conversion ซ้ำ CPU profile อาจชี้ algorithm, parsing หรือ reflection Lock profileอาจชี้ critical section ใหญ่เกิน แต่ไม่ควรใช้ sync.Pool, unsafe, custom JSON หรือ cache ก่อน profileระบุ cost และ benchmarkพิสูจน์ ผลภายใต้ correctness tests

Optimization comment ควรบันทึกเหตุผลและ benchmark/profile reference เพื่อไม่ให้คนถัดไป “ทำให้อ่านง่าย” แล้วลบ behavior สำคัญ หรือในทางกลับกันทำให้ code ซับซ้อนค้างไว้หลัง workload เปลี่ยน

PGO และ Go Runtime

Profile-Guided Optimization ใช้ representative CPU profile ช่วย compiler optimize hot paths วาง default.pgo ใน main package หรือกำหนด profileตาม Go tooling แล้ว build/test artifactจริง Profile ต้องมาจาก workload ที่แทน productionและทบทวนเมื่อ code/traffic เปลี่ยน PGO ไม่ใช่ guarantee ว่าทุก โปรแกรมเร็วขึ้น จึงเปรียบเทียบ latency/CPU/binary behavior ก่อนและหลัง

GC tuning เริ่มจากลด allocationและตั้ง memory limitให้สัมพันธ์กับ container/headroom ผ่าน GOMEMLIMIT หรือ runtime API จากนั้นดู GC CPU, heap goal, RSS/OOM และ latency อย่าคัดลอก percentage เดียวทุก service เพราะ non-Go memory, mmap, cgo และ sidecar ใช้ budgetด้วย GOGC ต่ำลด heapแต่เพิ่ม CPU; สูงทำตรงข้าม ต้องวัดภายใต้ memory limitจริง

Production Toolbox

Default: testing.B + b.Loop, benchstat, pprof และ trace ใช้ PGO เมื่อมี representative profile และ release comparison ใช้ continuous profilerเมื่อ cost/retentionเหมาะ Optimization library หรือ runtime knob ทุกตัวต้องมี baseline, hypothesis, comparison และ rollback

Checklist ก่อนเรียกว่างาน Performance

  • objective ระบุ workload, metric และ budgetหลายมิติ
  • external wait ถูกแยกจาก on-CPU ก่อน micro-optimize
  • benchmark ป้องกัน setup/noise และรันหลาย samples
  • profile type ตรงกับคำถามและมาจาก representative load
  • เปลี่ยนทีละ hypothesis พร้อม correctness/race tests
  • PGO profile มี provenance และ refresh policy
  • GC/memory tuningคิดรวม container และ non-Go memory
  • หลักฐานถูกเก็บใน PR/decision ไม่ใช่ตัวเลขจากเครื่องเดียวแบบไร้ context

อ่านเพิ่ม: Profiling Go Programs, Go Diagnostics, Profile-guided optimization และ testing.B