บทที่ 27 · Part 7 — Production Engineering
Measure, Profile and Optimize
ใช้ benchmark, benchstat, pprof, trace, PGO และ GC evidence ก่อนแก้ performance
การเปลี่ยน struct ทั้งระบบเป็น pointer, เพิ่ม sync.Pool หรือย้ายจาก Chi ไป framework เร็วกว่าโดยยังไม่รู้
ว่า latency อยู่ที่ database เป็น optimization ที่เพิ่ม complexity โดยไม่แตะ bottleneck Performance work
ที่ดีเริ่มจาก user-visible objective, วัด baseline, หาเหตุ แล้วเปลี่ยนทีละอย่าง
จบบทนี้คุณจะ
- เขียน benchmark ด้วย
b.Loopและเปรียบเทียบหลาย run ด้วยbenchstat - เลือก CPU/heap/block/mutex profile หรือ execution trace ตามคำถาม
- ใช้ PGO และ GC/runtime tuning หลังพิสูจน์ workload แล้ว
Performance Loop ที่ย้อนกลับได้
objective ต้องระบุ metric และ workload เช่น reduce P99 latency ของ submit endpoint ภายใต้ concurrency ที่กำหนดโดยไม่เพิ่ม error rate/memoryเกิน budget ไม่ใช้คำว่า “ทำให้เร็วขึ้น” เพราะ throughput, latency, CPU และ memory อาจแลกกัน
เริ่มจาก trace/metrics แยกเวลาที่รอ MySQL, Redis, network และ queue ออกจาก on-CPU ถ้า 90% อยู่ที่ query การลด allocationใน JSON 20% แทบไม่เปลี่ยน end-to-end latency แต่ยังเพิ่ม maintenance cost
Benchmark ให้สิ่งที่ตั้งใจวัด
Go 1.24+ ควรใช้ b.Loop() สำหรับ benchmark ใหม่ มันจัด timing loop และลดความเสี่ยง dead-code
elimination:
func BenchmarkNormalizeBatchIDs(b *testing.B) {
input := loadBatchIDs(1_000)
b.ReportAllocs()
for b.Loop() {
NormalizeBatchIDs(input)
}
}
ใช้ input sizes/distributions ที่แทน production และแยก setup ออกจาก measured body รันหลายครั้งบน
environment ควบคุมแล้วใช้ benchstat ไม่สรุปจากตัวเลขครั้งเดียว:
go test -run=^$ -bench=BenchmarkNormalize -benchmem -count=10 ./internal/settlement \
> before.txt
go test -run=^$ -bench=BenchmarkNormalize -benchmem -count=10 ./internal/settlement \
> after.txt
go tool benchstat before.txt after.txt
Pin benchstat เป็น tool dependency ใน repo จริงและวัด variants แบบ serial เพราะ benchmark ที่รันพร้อมกัน
แย่ง CPU/thermal budget ผลที่ไม่มี statistical significance ไม่ควรถูกเรียกว่า improvement
เลือก Diagnostic ให้ตรงคำถาม
| คำถาม | เครื่องมือ |
|---|---|
| CPU อยู่ที่ function ใด | CPU profile + go tool pprof |
| allocation churn มาจากไหน | heap alloc_objects/alloc_space |
| memory ที่ยังอยู่มาจากไหน | heap inuse_objects/inuse_space |
| goroutine รอ lock/channel/network เมื่อไร | execution trace, block/mutex/goroutine profile |
| compiler ทำให้ value escape/inlining ไหม | go build -gcflags=-m=2 หลัง profile ชี้จุด |
pprof เป็น sampling/aggregate view ว่า ที่ไหน ใช้ resource ส่วน trace บอก timeline ว่า เมื่อไรและ ทำไม goroutine block/schedule/GC อย่าเปิด block/mutex profiling rate สูงตลอดโดยไม่วัด overhead และ ต้องป้องกัน profile endpoint ตามบทก่อน
Optimization Patterns หลังมีหลักฐาน
allocation profile อาจนำไปสู่ preallocation ที่มี bound, buffer reuse, streaming หรือเลิก conversion
ซ้ำ CPU profile อาจชี้ algorithm, parsing หรือ reflection Lock profileอาจชี้ critical section ใหญ่เกิน
แต่ไม่ควรใช้ sync.Pool, unsafe, custom JSON หรือ cache ก่อน profileระบุ cost และ benchmarkพิสูจน์
ผลภายใต้ correctness tests
Optimization comment ควรบันทึกเหตุผลและ benchmark/profile reference เพื่อไม่ให้คนถัดไป “ทำให้อ่านง่าย” แล้วลบ behavior สำคัญ หรือในทางกลับกันทำให้ code ซับซ้อนค้างไว้หลัง workload เปลี่ยน
PGO และ Go Runtime
Profile-Guided Optimization ใช้ representative CPU profile ช่วย compiler optimize hot paths วาง
default.pgo ใน main package หรือกำหนด profileตาม Go tooling แล้ว build/test artifactจริง Profile
ต้องมาจาก workload ที่แทน productionและทบทวนเมื่อ code/traffic เปลี่ยน PGO ไม่ใช่ guarantee ว่าทุก
โปรแกรมเร็วขึ้น จึงเปรียบเทียบ latency/CPU/binary behavior ก่อนและหลัง
GC tuning เริ่มจากลด allocationและตั้ง memory limitให้สัมพันธ์กับ container/headroom ผ่าน GOMEMLIMIT
หรือ runtime API จากนั้นดู GC CPU, heap goal, RSS/OOM และ latency อย่าคัดลอก percentage เดียวทุก service
เพราะ non-Go memory, mmap, cgo และ sidecar ใช้ budgetด้วย GOGC ต่ำลด heapแต่เพิ่ม CPU; สูงทำตรงข้าม
ต้องวัดภายใต้ memory limitจริง
Production Toolbox
Default: testing.B + b.Loop, benchstat, pprof และ trace ใช้ PGO เมื่อมี representative profile
และ release comparison ใช้ continuous profilerเมื่อ cost/retentionเหมาะ Optimization library หรือ
runtime knob ทุกตัวต้องมี baseline, hypothesis, comparison และ rollback
Checklist ก่อนเรียกว่างาน Performance
- objective ระบุ workload, metric และ budgetหลายมิติ
- external wait ถูกแยกจาก on-CPU ก่อน micro-optimize
- benchmark ป้องกัน setup/noise และรันหลาย samples
- profile type ตรงกับคำถามและมาจาก representative load
- เปลี่ยนทีละ hypothesis พร้อม correctness/race tests
- PGO profile มี provenance และ refresh policy
- GC/memory tuningคิดรวม container และ non-Go memory
- หลักฐานถูกเก็บใน PR/decision ไม่ใช่ตัวเลขจากเครื่องเดียวแบบไร้ context
อ่านเพิ่ม: Profiling Go Programs,
Go Diagnostics,
Profile-guided optimization และ
testing.B