บทที่ 17 · Part 4 — Reliability and Operations

CI and Suite Operations

วาง pipeline, browser dependencies, retries, sharding และ merged reports พร้อม policy ดูแล suite ระยะยาว

CI and Suite Operations

Suite ที่ดีบน laptop ยังไม่สร้าง feedback ถ้า CI ใช้ browser image คนละรุ่น, artifacts หายเมื่อ job fail หรือ shards แย่งข้อมูลเดียวกัน Operations ของ test suite จึงต้อง version, observe และมี ownership เหมือน production tooling

จบบทนี้คุณจะ

สร้าง CI baseline จาก lockfile + browser dependencies ตั้ง fail/flaky policy ใช้ sharding กับ blob report อย่างถูก และวาง upgrade/metrics/quarantine governance ที่ทีมดูแลต่อได้

CI Baseline

ลำดับ official baseline คือ checkout, install dependencies จาก lockfile, install Playwright browsers + OS dependencies และ run tests ตัวอย่าง GitHub Actions แบบ 1 browser:

name: playwright
on:
  pull_request:
  push:
    branches: [main]

jobs:
  test:
    runs-on: ubuntu-latest
    timeout-minutes: 30
    steps:
      - uses: actions/checkout@v5
      - uses: actions/setup-node@v5
        with:
          node-version: 24
          cache: npm
      - run: npm ci
      - run: npx playwright install --with-deps chromium
      - run: npx playwright test --project=chromium
      - uses: actions/upload-artifact@v4
        if: ${{ !cancelled() }}
        with:
          name: playwright-report
          path: playwright-report/

Action versions เป็น infrastructure dependency ต้อง review/update ตาม policy ของ repository ตัวอย่างใช้ Node 24 ซึ่งอยู่ใน system requirements ของ Playwright 1.62 แต่ต้องตรวจ installation page ใหม่ทุกครั้งที่อัปเกรด

Install เฉพาะ browsers ที่ job ใช้ ถ้า matrix รัน WebKit/Firefox ให้ติดตั้งใน job นั้น Package version, lockfile, browser binaries และ Docker image tag ต้องตรงกัน Official docs ไม่แนะนำ cache browser binaries เป็น default เพราะ restore cost ใกล้ download และ Linux dependencies cache ไม่ได้; ถ้าทีมวัดแล้วคุ้มให้ key ด้วย Playwright version

Config สำหรับ CI Signal

export default defineConfig({
  forbidOnly: Boolean(process.env.CI),
  retries: process.env.CI ? 2 : 0,
  failOnFlakyTests: Boolean(process.env.CI),
  retryStrategy: process.env.CI ? 'isolated' : 'immediate',
  globalTimeout: process.env.CI ? 25 * 60_000 : undefined,
  reporter: process.env.CI
    ? [['line'], ['blob'], ['html', { open: 'never' }]]
    : [['list'], ['html', { open: 'never' }]],
})

ค่า retry/global timeout เป็น illustrative ต้องวัด duration และ environment budget จริง forbidOnly กัน committed focus test failOnFlakyTests บังคับให้ผ่าน retry ยังต้อง triage Global timeout หยุด systemic outage ไม่ใช่เร่ง test เดี่ยว

Sharding เมื่อ Isolation พร้อม

npx playwright test --shard=1/4
npx playwright test --shard=2/4
npx playwright test --shard=3/4
npx playwright test --shard=4/4

เมื่อ fullyParallel: true Playwright balance shards ระดับ individual tests ได้ดีกว่า file-level แต่ sharding ไม่แก้ suite ที่ใช้ shared mutable data ก่อนเปิด shard ให้รัน repeated high-worker tests, ตรวจ account pool/environment capacity และ exact cleanup

แต่ละ shard ใช้ blob reporter แล้ว merge:

npx playwright merge-reports --reporter html ./all-blob-reports

Pipeline ต้อง upload blob-report จากทุก shard, download มา directory เดียวใน merge job แล้ว publish HTML ครั้งเดียว Blob รวม results และ attachments เช่น trace/screenshot diff อย่ารัน HTML reporter แยกแล้วพยายามต่อไฟล์เอง

Browser Matrix ด้วย Risk

ตัวอย่าง policy:

  • PR smoke: Chromium + critical API tests
  • main/nightly: Chromium, Firefox, WebKit ตาม supported matrix
  • focused mobile/a11y/visual projects เฉพาะ tags ที่เกี่ยว
  • real-device smoke ในระบบแยกเมื่อ hardware/browser-shell risk สำคัญ

นี่เป็น course heuristic ทีมต้องใช้ product support policy, usage evidence และ defect history ตัดสิน อย่าเพิ่มทุก browser × role × locale × theme ให้ทุก test เพราะ duration เพิ่มแบบคูณโดย claim ไม่เพิ่ม

Artifact Upload และ Security

ใช้ if: !cancelled() เพื่อยัง upload report เมื่อ tests fail แต่ไม่เสียเวลาหลัง run ถูก cancel Artifact access และ retention ต้องตรง data sensitivity ของ trace/HAR/screenshot อย่า hardcode จำนวนวันจากตัวอย่าง ให้ security/data owner กำหนดพร้อมชื่อและวันที่ ใช้ synthetic users และ redact network attachments ก่อน upload

Suite Health Metrics

วัดอย่างน้อย:

  • first-attempt pass rate และ flaky rate แยก project/test
  • p50/p95 duration และ top slow tests/fixtures
  • quarantine count, age และ critical coverage gaps
  • failure category: product, test, environment, dependency
  • time-to-triage จาก first failure ถึง owner/root cause
  • browser/project-specific defect yield

จำนวน tests ไม่ใช่ health metric หลัก Dashboard ต้องช่วยตัดสินใจ เช่น top flaky tests มี owner และ CI budget overrun ชี้ fixture/network bottleneck

Upgrade Playwright อย่างควบคุม

  1. ตรวจ release notes และ breaking/deprecation notes
  2. bump @playwright/test + lockfile
  3. install browser binaries ใหม่
  4. pin Docker image tag ให้ตรงถ้าใช้ container
  5. รัน typecheck/lint + all projects + visual review
  6. ตรวจ auth state, component testing และ custom fixtures ที่แตะ API เปลี่ยน
  7. commit snapshot changes แยกพร้อมเหตุผล ไม่ regenerate เงียบ

สำหรับ baseline นี้ 1.62.1 แก้ regressions ของ tsconfig, branded primitive และ ARIA snapshots จึงใช้ patch นี้แทน 1.62.0 Future release ต้อง re-research ไม่อ้างบทนี้ว่าล่าสุดตลอดไป

Definition of Operational Ready

  • Lockfile, browsers และ image versions สอดคล้องกัน
  • PR/main/nightly matrices มี risk rationale
  • only, flaky และ global timeout policies enforce บน CI
  • Shards ใช้ isolated data และ merge blob reports สำเร็จ
  • Failure artifacts upload แม้ test fail และควบคุม access/retention
  • Flaky/quarantine/slow metrics มี owner review
  • Upgrade runbook ครอบ release notes, binaries, types และ baselines

สรุปบทนี้

CI ที่เชื่อถือได้ทำ version และ evidence reproducible Sharding เพิ่ม throughput หลัง isolation พร้อมเท่านั้น ส่วน flaky policy, artifact security และ upgrade cadence ทำให้ suite รักษาความน่าเชื่อถือหลังวันเปิดตัว

อ่านเพิ่มเติม