บทที่ 15 · Part 4 — Reliability and Operations
Time, Parallelism and Flakiness
ควบคุมนาฬิกา ออกแบบ isolation และวิเคราะห์ retry เพื่อกำจัด race, shared state และ arbitrary waits
Time, Parallelism and Flakiness
test countdown ผ่านบน laptop แต่ fail ตอนเที่ยงคืนใน timezone ของ CI ส่วน reservation tests 2 รายการแย่ง account เดียวกันและผ่านเฉพาะเมื่อ retry ปัญหาเหล่านี้ไม่ใช่ “Playwright ช้า” แต่เป็น time/state contract ที่ suite ไม่ได้ควบคุม
จบบทนี้คุณจะ
ควบคุม browser clock แยกจาก backend time ออกแบบ suite ให้ parallel-safe ใช้ retries/timeouts เป็น diagnostic signal และจำแนก root cause ของ flaky test แทนการเพิ่ม sleep
Clock 2 แบบ
page.clock.setFixedTime() fix ค่าที่ Date.now()/new Date() เห็น แต่ timers ยังเดิน เหมาะกับ label “วันนี้”:
test('TIME-001 - labels a workshop starting today', async ({ page }) => {
await page.clock.setFixedTime(new Date('2030-06-15T08:00:00.000Z'))
await page.goto('/workshops')
await expect(page.getByText('Today at 09:00')).toBeVisible()
})
สำหรับ timer/countdown ใช้ install() ก่อน page เรียก time APIs แล้วขยับเวลา:
test('TIME-002 - expires an unconfirmed hold after ten minutes', async ({ page }) => {
await page.clock.install({ time: new Date('2030-06-15T08:00:00.000Z') })
await page.goto('/reservation-hold')
await expect(page.getByRole('status')).toContainText('10:00 remaining')
await page.clock.fastForward('10:00')
await expect(page.getByRole('alert')).toContainText('Reservation hold expired')
})
Clock API ควบคุมเวลาใน browser context ไม่ได้เปลี่ยน database/backend clock ถ้า expiry decision อยู่ server ต้องมี injectable clock/test seam ที่ service หรือสร้าง server state ผ่าน API อย่าใช้ browser clock แล้วสรุปว่า backend expiry ถูกต้อง
Parallelism Model
Playwright รัน test files ขนานโดย default Tests ในไฟล์เดียวกันรันตามลำดับ เว้นเปิด fullyParallel
แต่ละ test ได้ BrowserContext ใหม่จึงแยก cookies/local storage ส่วน server-side state ยังต้องออกแบบเอง
เพื่อให้ test รันเดี่ยว, reorder, repeat และ shard ได้:
- unique resource id ต่อ test
- account ต่อ worker เมื่อ account state ถูก mutate
- ไม่มี global mutable variable ที่เก็บ resource จาก test ก่อนหน้า
- teardown exact ids ไม่ truncate shared store
- mock/clock อยู่ใน test context ไม่เปลี่ยน shared service
ทดสอบสมมติฐานด้วย:
npx playwright test tests/ui/reservations --workers=1 --repeat-each=20
npx playwright test tests/ui/reservations --workers=8 --repeat-each=20
ถ้า fail เฉพาะหลาย workers ให้ตรวจ data collision, account quota, environment capacity และ dependency rate limit ก่อนลด worker ถ้าระบบทดสอบมี limit จริงให้บันทึก measurement/owner และตั้ง concurrency ตาม budget ไม่เดา
Retry เป็น Classification
Playwright แบ่งผลเป็น passed, flaky (fail ก่อนแล้ว pass retry) และ failed Default ไม่มี retry คอร์สนี้แนะนำ local
zero retries; CI อาจใช้ 1–2 เพื่อเก็บ artifacts และเปิด failOnFlakyTests ไม่ให้ flaky กลายเป็น success ปกติ:
export default defineConfig({
retries: process.env.CI ? 2 : 0,
failOnFlakyTests: Boolean(process.env.CI),
retryStrategy: process.env.CI ? 'isolated' : 'immediate',
})
Playwright 1.62 เพิ่ม retryStrategy: 'isolated' ซึ่งรวบ retries ไปท้าย run และรันทีละ test ใน worker เดียว
ช่วยทดสอบสมมติฐานว่า contention ทำให้ fail แต่ถ้า pass เมื่อ isolated ก็ยังเป็น flaky ที่ต้องแก้ root cause
immediate เป็น default และเหมาะกับ feedback เร็วตามปกติ
อย่า retry helper เฉพาะ non-idempotent action แบบอัตโนมัติ เพราะ request แรกอาจสำเร็จแล้ว response หาย Playwright test retry เริ่ม test ใหม่ทั้งรายการ จึงต้องให้ data/idempotency/cleanup รองรับด้วย
Timeout Hierarchy
Default test timeout และ expect timeout มีคนละ scope ตาม docs ให้เพิ่มตรง operation ที่มี latency budget ต่างจริง:
await expect(page.getByRole('status')).toContainText('Export ready', {
timeout: 30_000,
})
ไม่เพิ่ม global test timeout เป็น 120 วินาทีเพราะ export test เดียวช้า Fixture ราคาแพงตั้ง timeout ของ fixture ได้
และ CI ควรมี globalTimeout กัน environment outage แขวน run ทั้งหมด
Playwright 1.62 รองรับ AbortSignal สำหรับ operations/assertions จำนวนมาก ใช้ cancellation เมื่อ orchestration ภายนอก
ต้องหยุดงาน แต่ signal ไม่ปิด default timeout และไม่ใช่ตัวแทนของ state-based wait
Flaky Failure Taxonomy
| Symptom | หลักฐานที่ควรดู | Root cause ที่พบบ่อย |
|---|---|---|
| element not visible | trace DOM/actionability | locator ผิด, overlay, state ไม่ถึง |
| pass เมื่อ sleep ยาว | network/timeline | wait ผูกเวลาแทน outcome |
| fail เฉพาะ parallel | ids/accounts/server logs | shared mutable state, capacity |
| fail เฉพาะ browser | trace + support matrix | browser behavior/unsupported feature |
| screenshot diff | actual/diff + environment | font, browser/OS, dynamic data, real UI change |
| API status สลับ | exchange + request id | environment dependency, duplicate effect |
| fail ช่วง timezone | clock/config/payload | implicit local time, DST, date boundary |
เริ่มด้วย reproduce test เดี่ยว + repeat จากนั้นเปรียบเทียบ serial/parallel และ project อย่าเปลี่ยน locator, timeout, retry และ data พร้อมกันเพราะจะไม่รู้ causal fix
Quarantine ที่รับผิดชอบ
หาก flaky test block release และยังแก้ไม่ทัน อาจแยก quarantine job แต่ต้องมี issue, owner, first-seen,
reproduction evidence และ expiry/review date ห้าม skip เงียบหรือปล่อย quarantine สีเขียวถาวร Critical journey
ที่ถูก quarantine หมายถึง coverage gap ต้องมี manual/runtime control ชั่วคราวตาม risk owner
รายการตรวจสอบ
- Browser clock ติดตั้งก่อน application ใช้ time API
- Backend time มี seam ของตน ไม่สมมติว่าตาม browser
- Test ผ่านแบบเดี่ยว, repeat, parallel และ shard ด้วย data ownership เดิม
- CI แยก flaky จาก passed และเก็บ first-failure trace
- Timeout เพิ่มเฉพาะ operation ที่มี latency contract
- Quarantine มี owner, issue และวันทบทวน
สรุปบทนี้
Flakiness คือข้อมูลว่า test หรือระบบมี input ที่ควบคุมไม่ได้ Clock, isolation และ artifacts ช่วยทำ input นั้นให้เห็น Retry มีไว้จำแนกและเก็บหลักฐาน ไม่ใช่ลบ failure ออกจากความรับผิดชอบ