บทที่ 26 · Part 7 — Production Engineering

Service Lifecycle and Observability

ประกอบ application, readiness, graceful shutdown, slog, metrics และ traces ให้มี ownership เดียวกัน

service ที่ตอบ /health ว่า 200 แต่หยุดรับ request ไม่ได้, log error ซ้ำ 5 ครั้ง และไม่มี metric ว่า queue เต็มหรือ DB pool รออยู่ยังไม่พร้อม production Lifecycle กับ observability ต้องออกแบบพร้อม component เพราะเจ้าของที่สร้าง resource เป็นผู้เดียวที่รู้ว่าจะเริ่ม หยุด และอธิบายสภาพมันอย่างไร

จบบทนี้คุณจะ

  • ประกอบ startup/readiness/shutdown ตาม dependency order
  • เลือก logs, metrics, traces และ profiles ตามคำถาม ไม่เก็บทุกอย่างแบบไร้ขอบเขต
  • ใช้ slog และ samber/slog-multi ในตำแหน่งที่เหมาะกับ Go 1.26

Composition Root เป็นเจ้าของ Lifecycle

main หรือ application package ควร load/validate config, สร้าง logger/telemetry, เปิด clients, สร้าง services/adapters แล้วเริ่ม servers/workers ไม่มี package ย่อยควรเรียก log.Fatal หรือ os.Exit เพราะ มันข้าม deferred cleanup และแย่ง policy จาก owner

shutdown ใช้ signal-derived context เช่น signal.NotifyContext และมี grace deadline ลำดับสำคัญ: ประกาศ not-ready/หยุด traffic ใหม่, http.Server.Shutdown, หยุดรับ worker jobs, รอ in-flight ตาม policy, flush telemetry แล้วปิด DB/Redis หากปิด pool ก่อน handler/activity จบ จะสร้าง failure ระหว่าง drain

liveness ตอบว่า process ควรถูก restart ไหม ไม่ควรล้มเพียงเพราะ dependency กระพริบชั่วคราว Readiness ตอบว่าควรรับงานใหม่ไหมและอาจพิจารณา dependency ที่จำเป็น Startup check ช่วย fail fast เมื่อ config/DSN ผิด แต่ health check ต้องมี timeout, caching หรือ rate ที่ไม่กลายเป็น load source

สัญญาณแต่ละชนิดตอบคนละคำถาม

Signalคำถามตัวอย่าง Settlement
Structured logsเหตุการณ์ใดเกิดพร้อม context อะไรbatch rejected, shutdown timed out
Metricsปริมาณ/อัตรา/latency/saturation เป็นอย่างไรrequest rate, DB wait, queue depth
Tracesเวลาเดินทางไปอยู่ span ใดHTTP → MySQL → Temporal client
ProfilesCPU/memory/lock ใช้ที่ code ใดallocation hot path, mutex contention

ใช้ log/slog เป็น default และ log ที่ handling boundary ครั้งเดียวด้วย message template คงที่ + structured attributes อย่าใส่ batch/user ID เป็น Prometheus label เพราะ cardinality ไม่จำกัด แต่ใส่ใน sampled trace หรือ log ที่มี retention/access policyได้

func logAccepted(ctx context.Context, logger *slog.Logger, batch Batch) {
    if !logger.Enabled(ctx, slog.LevelInfo) {
        return
    }
    logger.InfoContext(
        ctx,
        "settlement batch accepted",
        slog.String("batch_id", batch.ID),
        slog.String("status", string(batch.Status)),
    )
}

ห้าม log payload, credential, token หรือข้อมูลส่วนบุคคลโดยไม่มี classification/redaction policy และ อย่าสับสน operational log กับ immutable audit trail Audit event ต้องมี schema, integrity, access, retention และเจ้าของต่างหาก

slog Multi-Handler ในปี 2026

Go 1.26 มี slog.NewMultiHandler สำหรับ fan-out ง่ายไปหลาย handlers จึงควรใช้ stdlib ก่อนเมื่อแค่ส่ง JSON stdout พร้อม audit handler ถ้าต้อง routing ตาม level/attribute, failover, fanout pipeline ซับซ้อน หรือ bridge ecosystem คอร์สแนะนำ samber/slog-multi และ packages ใน samber slog ecosystem

เนื่องจาก lab core ของคอร์สยังประกาศ go 1.25.0 ห้ามใช้ slog.NewMultiHandler ใน package ที่ต้อง compile ด้วย Go 1.25 ให้ใช้ composite handler เล็กของทีม หรือ samber/slog-multi จนกว่าจะยก minimum เป็น 1.26 นี่คือตัวอย่างของการแยก current-toolchain recommendation ออกจาก minimum-version contract

library เหล่านี้แทน handler composition boilerplate ไม่ได้กำหนดว่า event ใดควร log หรือ redaction ใด ถูกต้อง Handler ที่ส่ง network ต้องมี buffering/backpressure/failure policy; logging ห้ามทำ request path ค้างไม่มีกำหนด

Metrics และ Tracing ที่ใช้งานได้

ใช้ Prometheus counters สำหรับ totals, gauges สำหรับ current state และ histograms สำหรับ latency/size ที่ต้อง aggregate ข้าม instances ออกแบบ label จาก bounded vocabulary เช่น method, route pattern, status class, operation, outcome ไม่ใช้ raw path, error text, user/batch ID

OpenTelemetry เป็น default สำหรับ distributed tracing ให้ instrument entry/exit boundaries ก่อน: HTTP middleware, application operation, database/external calls และ Temporal client Propagate context ผ่าน header/SDK และกำหนด sampling/cost budget ไม่สร้าง span ทุก helper เล็กจน trace อ่านไม่ได้

pprof/continuous profiling ช่วยหา CPU/heap/goroutine/mutex/block แต่ endpoint มีข้อมูลภายในและอาจแพง ต้องแยก listenerหรือป้องกัน auth/network policy ไม่เปิดสู่ public internet

Alert จาก User Impact และ Saturation

alert ที่ดีเริ่มจาก SLO/user impact เช่น error ratio/latency burn แล้วตามด้วย saturation ที่ actionable เช่น DB pool wait, worker slots, queue age และ Redis timeout ไม่ alert ทุก log errorหรือ goroutine count ที่เปลี่ยนเล็กน้อย Dashboard/alert ต้องมี runbook, owner และทดสอบว่า query ใช้หน่วย/label ถูก

Production Toolbox

Default: log/slog, Prometheus client, OpenTelemetry และ secured net/http/pprof ใช้ slog.NewMultiHandler สำหรับ fan-out ง่ายเมื่อ minimum เป็น Go 1.26; ใช้ samber/slog-multi เมื่อ ยังรองรับ Go 1.25 หรือ routing/composition ซับซ้อน Signals ทุกตัวต้องมี cardinality, sampling, retention และ failure budget

Checklist ก่อน Ready

  • startup validate config และ dependency ด้วย timeout
  • readiness/liveness ตอบคำถามต่างกันและไม่ยิง dependencyถี่เกิน
  • shutdown order หยุด ingress ก่อน drain worker แล้วค่อยปิด clients/exporters
  • error log ครั้งเดียวที่ boundary พร้อม trace correlation
  • metric labels bounded และมี unit/name ชัด
  • spans อยู่ที่ meaningful boundaries และมี sampling policy
  • profile/debug endpoint ถูกป้องกัน
  • dashboard/alert มี SLO, owner และ runbook

อ่านเพิ่ม: log/slog, OpenTelemetry Go, Prometheus Go client และ Go Diagnostics