บทที่ 2 · Part 1 — Security Foundations

Threat Modeling & Risk Assessment

เปลี่ยน architecture diagram ให้เป็น DFD, หา threat ด้วย STRIDE, จัดลำดับด้วย risk และเลือก control ที่ตรวจสอบผลได้

ก่อนเปิดฟีเจอร์โอนเงินไปบัญชีธนาคาร ทีมพัฒนาได้รับ checklist ยาว 3 หน้า: เปิด WAF, scan dependency, ใช้ TLS และเปิด audit log ทุกข้อดูมีประโยชน์ แต่ไม่มีใครตอบได้ว่า ถ้า bank callback ถูกส่งซ้ำ หรือ operations agent เปลี่ยนบัญชีปลายทางหลังอนุมัติแล้ว control ข้อไหนจะหยุดความเสียหาย

Threat modeling ทำให้ทีมเลิกถามว่า “ใส่ Security ครบหรือยัง” แล้วเปลี่ยนเป็น “ระบบนี้พังอย่างไรได้บ้าง เพราะ assumption ไหน และมีวิธีลดหรือตรวจจับผลนั้นอย่างไร”

Learning Outcomes

  • กำหนด scope และวาด Data Flow Diagram ที่เห็น trust boundary ได้
  • ใช้ STRIDE สร้าง threat ที่ผูกกับ component และ data flow จริง
  • เขียน threat statement และ attack tree ที่ทีมลงมือแก้ต่อได้
  • ประเมิน likelihood, impact และ residual risk โดยเปิดเผย assumption
  • เลือก preventive, detective และ corrective controls ที่เสริมกัน
  • จัด threat modeling workshop และรักษา model ให้ทันระบบ

Threat Modeling Is a Design Activity

Threat modeling คือการสร้างแบบจำลองระบบจากมุมของผู้ก่อความเสียหาย แล้วใช้แบบจำลองนั้น ตัดสินใจเรื่อง design ก่อนที่ต้นทุนการแก้จะสูง ไม่ใช่เอกสารที่ Security ทำหลังเขียน code เสร็จ

ผลลัพธ์ที่ดีไม่ใช่จำนวน threat มากที่สุด แต่คือ:

  1. ทีมมีภาพระบบและ assumption ตรงกัน
  2. พบ design flaw ที่ test หรือ scanner มองไม่เห็น
  3. เรื่องสำคัญมี owner และ control ที่พิสูจน์ผลได้
  4. residual risk ถูกยอมรับโดยคนที่มีอำนาจ ไม่ใช่ซ่อนอยู่ใน backlog

Threat model ไม่แทน code review, penetration test หรือ incident learning แต่ช่วยบอกว่า กิจกรรมเหล่านั้นควรให้ความสำคัญกับอะไร

Choose the Right Scope

scope ที่กว้างเกินไปจะจบด้วย diagram อ่านไม่ออก ส่วน scope ที่แคบเกินไปจะมองไม่เห็น เส้นทางโจมตีข้ามระบบ ให้เริ่มจาก 1 business capability หรือ 1 การเปลี่ยนแปลงสำคัญ

ตัวอย่าง scope ที่ดี:

  • “ลูกค้าผูกบัญชีธนาคารใหม่และถอนเงินครั้งแรก”
  • “operations agent ทำ manual refund เกิน 500000 สตางค์”
  • “mobile app อัปโหลดเอกสาร KYC แล้ว reviewer เปิดดู”
  • “CI/CD deploy image ไป production account”

กำหนดขอบเขตด้วยข้อความสั้น ๆ:

In scope:
- mobile withdrawal request
- wallet API, payout service, ledger, bank adapter
- bank callback and operations review

Out of scope for this session:
- bank internal implementation
- customer device OS internals
- card payment flow

Assumptions to verify:
- bank callback has a unique event ID
- ledger entries cannot be updated after posting
- operations access requires managed device

Out of scope ไม่ได้แปลว่าปลอดภัย

สิ่งที่อยู่นอก scope ต้องมี owner หรือ threat model อื่นรองรับ ถ้า dependency ภายนอก รับเงินไปแล้วแต่ไม่มีใครวิเคราะห์ callback contract การตีกรอบเพียงทำให้ blind spot มีชื่อเท่านั้น

Model the System with a DFD

Data Flow Diagram หรือ DFD สำหรับ threat modeling ต้องแสดงสิ่งที่มีผลต่อ trust ไม่จำเป็นต้องสวยหรือใช้ notation ซับซ้อน แต่ควรมีอย่างน้อย 4 ชนิด:

Elementคำถาม Security
External entityใครอยู่นอกขอบเขตควบคุมของระบบ และพิสูจน์ identity อย่างไร
Processcode นี้รันด้วย privilege อะไร และเชื่อ input ใด
Data storeเก็บ asset อะไร ใครอ่าน/เขียน/ลบได้
Data flowข้อมูลอะไรข้ามไป ใช้ protocol ใด และป้องกัน replay อย่างไร

วาด trust boundary ทุกครั้งที่ ownership, network, identity หรือ privilege เปลี่ยน

เลขบนลูกศรมีไว้ให้ threat อ้างถึง flow ได้ตรงกัน เช่น “T-07: replay flow 7” diagram ควรบอก data สำคัญด้วย ไม่ใช่เขียนเพียง request ทุกเส้น

Record Assumptions Beside the Diagram

architecture diagram มักซ่อน assumption เช่น:

  • API gateway ตรวจ token แล้ว service ภายในจึงไม่ตรวจ authorization ซ้ำ
  • traffic ใน VPC เชื่อถือได้
  • callback ที่มาจาก IP ของธนาคารเป็นของแท้
  • operations user ไม่มีทางถูก phishing
  • queue ส่ง message ครั้งเดียว

เขียน assumption เหล่านี้ออกมา เพราะ threat จำนวนมากเกิดจาก assumption ที่ไม่จริง บทบาทของ threat modeling ไม่ใช่พิสูจน์ว่าทุกอย่างอันตราย แต่คือทำให้ทีมรู้ว่า กำลังฝากความถูกต้องไว้กับอะไร

Generate Threats with STRIDE

STRIDE เป็น prompt 6 แบบสำหรับตรวจทุก element และ data flow ไม่ใช่ checklist ที่ เลือก threat ทั่วไปมาวางโดยไม่อ่าน architecture

CategorySecurity propertyคำถามตัวอย่าง withdrawal
SpoofingAuthenticationใครปลอมเป็นใครได้ผู้โจมตีใช้ callback credential ที่รั่วปลอมเป็นธนาคาร
TamperingIntegrityใครแก้ข้อมูลหรือคำสั่งได้เปลี่ยนบัญชีปลายทางหลัง customer ยืนยันแล้ว
RepudiationAccountabilityใครปฏิเสธการกระทำได้operator ทำ refund แต่ audit ไม่มี before/after
Information DisclosureConfidentialityข้อมูลหลุดตรงไหนerror ส่งเลขบัญชีเต็มเข้า centralized log
Denial of ServiceAvailabilityอะไรทำให้ resource หมดcallback ปลอมจำนวนมากใช้ worker และ DB connection จนเต็ม
Elevation of PrivilegeAuthorizationใครได้สิทธิ์สูงกว่าที่ควรsupport role เรียก internal payout endpoint ได้

วิธีใช้ที่ได้ผลคือเดินทีละ element:

  1. เลือก flow 7: bank callback
  2. ถาม STRIDE ทั้ง 6 แบบ
  3. ตัด category ที่อธิบายไม่ได้พร้อมเหตุผล
  4. เขียน threat ที่มี actor, action, weakness และ impact
  5. เชื่อม threat กลับไปยัง requirement, control, test และ owner

STRIDE เป็นตัวช่วยตั้งคำถาม ไม่ใช่ขอบเขตสูงสุด

business logic abuse เช่น ลูกค้าที่ผ่าน authentication ใช้ promotion ซ้ำ 1,000 ครั้ง อาจไม่ลงกล่องเดียวอย่างสวยงาม แต่ยังเป็น threat จริง ให้เพิ่ม abuse case จากคนที่เข้าใจธุรกิจเสมอ

Write Actionable Threat Statements

ข้อความ “มี SQL injection” บอกชื่อช่องโหว่แต่ไม่บอกเส้นทางหรือผลกระทบ ใช้รูปแบบนี้แทน:

Threat actor CAN perform action
AGAINST asset or component
BECAUSE assumption or weakness
LEADING TO business or security impact.

ตัวอย่าง:

ผู้ที่ได้ bank callback credential สามารถส่ง event เดิมซ้ำ ไปยัง payout endpoint เพราะระบบตรวจ signature แต่ไม่บันทึก event ID ที่ประมวลผลแล้ว ส่งผลให้ ledger post payout ซ้ำและเงินออกมากกว่า 1 ครั้ง

ข้อความนี้นำไปออกแบบได้ทันที:

  • ป้องกัน: unique constraint บน (provider, event_id) และ idempotent state transition
  • ตรวจจับ: metric callback duplicate และ alert เมื่อ payout reference ซ้ำ
  • กู้คืน: reconciliation หา duplicate แล้วเปิด compensation case โดยไม่แก้ ledger ทับ

Keep Threats Separate from Controls

อย่าเขียน “ใช้ WAF ป้องกัน injection” เป็น threat เพราะทีมจะยึดกับ solution ก่อนเข้าใจปัญหา Threat ควรคงอยู่แม้เปลี่ยน technology ส่วน control เป็นทางเลือกที่ประเมิน trade-off ได้

Explore Paths with Attack Trees

STRIDE เดินตาม component ได้ดี แต่ attack tree เริ่มจากผลที่ผู้โจมตีต้องการแล้วแตกเส้นทางย้อนกลับ

ประโยชน์ของ tree คือเห็นว่าการแก้เส้นเดียวไม่ปิด goal ทั้งหมด การบังคับ MFA ลูกค้าไม่ช่วย ถ้า support recovery หรือ object authorization ยังอ่อน และบางเส้นอาจใช้ control ร่วมกันได้ เช่น independent notification ช่วยตรวจทั้ง token theft กับ operator abuse

Add Abuse Cases from the Business

Misuse ที่สำคัญไม่จำเป็นต้องใช้ exploit ทางเทคนิค:

  • ผู้ใช้เปิดหลายบัญชีเพื่อรับ promotion ซ้ำ
  • merchant กับลูกค้าสมคบกันสร้าง refund หลังถอนเงินแล้ว
  • operator แบ่ง refund ก้อนใหญ่เป็นรายการเล็กเพื่อหลบ approval threshold
  • vendor ใช้ production data นอกวัตถุประสงค์
  • ลูกค้าส่ง request พร้อมกันเพื่อใช้ balance เดียว 2 ครั้ง

ให้ Product, Operations, Fraud, Finance และ Support เข้าร่วม session เพราะคนเหล่านี้รู้ exception path ที่ diagram ของ engineer มักไม่มี

Assess Risk with Evidence

หลังสร้าง threat แล้วต้องจัดลำดับ ไม่เช่นนั้นรายการ 100 ข้อจะทำให้ทุกอย่างดูสำคัญเท่ากัน

Estimate Likelihood

พิจารณาหลักฐานหลายด้าน:

  • Reachability — เปิด Internet, partner network หรือ internal only
  • Preconditions — ต้องมี account, privileged credential หรือ physical device หรือไม่
  • Exploitability — ทำซ้ำอัตโนมัติได้ไหม ต้องชนะ race window แคบแค่ไหน
  • Existing controls — control ทำงานจริงและมีผลทดสอบหรือเพียงอยู่ในเอกสาร
  • Detectability for the attacker — ทดลองแล้วรู้ผลทันทีหรือไม่
  • Threat activity — เคยเกิดในระบบที่ประเมิน, vendor หรืออุตสาหกรรมหรือไม่

Estimate Impact

อย่ารวมทุกอย่างเป็นคำว่า “High” โดยไม่บอกเหตุผล:

Impact dimensionคำถาม
Financialเงินออกสูงสุดเท่าไรต่อรายการ ต่อวัน และก่อนตรวจพบ
Customerกระทบกี่คน เข้าถึงเงินไม่ได้หรือข้อมูลใดรั่ว
Operationalต้องหยุดระบบหรือทำ manual recovery นานเท่าไร
Legal and regulatoryมี obligation ใดและใครเป็นผู้ประเมิน
Reputationลูกค้าและ partner จะสูญเสียความเชื่อมั่นแบบใด
Safety and systemicกระทบบริการสำคัญหรือระบบภายนอกต่อเนื่องหรือไม่

จำนวนเงินตัวอย่างต้องระบุว่าเป็น scenario assumption ไม่ใช่ตัวเลขจริง เช่น “บัญชีหนึ่งถอนได้ไม่เกิน 2000000 สตางค์ต่อวัน แต่ credential เดียวเข้าถึง 50,000 บัญชี” ทำให้เห็นว่า per-user limit ไม่ได้จำกัด blast radius ของ service credential

Use a Transparent Risk Scale

ตัวอย่าง scale แบบ 4 ระดับ:

RatingLikelihoodImpact
1 — Lowต้องมีเงื่อนไขหลายชั้นและยังไม่มีเส้นทางที่พิสูจน์ได้จำกัดใน test data หรือกู้คืนง่าย
2 — Moderateมีเส้นทางแต่ต้องมี access หรือ timing เฉพาะกระทบลูกค้าจำนวนจำกัดและหยุดได้เร็ว
3 — Highทำซ้ำได้ด้วย credential ทั่วไปหรือ endpoint publicเงิน/ข้อมูล/บริการกระทบอย่างมีนัยสำคัญ
4 — Criticalทำอัตโนมัติได้กว้างหรือกำลังถูกใช้จริงsystemic loss, mass compromise หรือกู้คืนยาก

อย่าให้ multiplication กลบ uncertainty: Likelihood 2 × Impact 4 ควรเก็บเหตุผล, confidence และสิ่งที่จะทำให้ rating เปลี่ยนด้วย

Maintain a Useful Risk Register

risk register ที่ทีมใช้ทำงานควรเชื่อม threat กับการตัดสินใจ:

Fieldตัวอย่าง
IDPAY-T07
Threatreplay bank callback ทำ payout ซ้ำ
Asset/flowcustomer funds / DFD flow 7
Existing controlssignature validation
Likelihood/impactHigh / Critical พร้อมเหตุผล
Treatmentidempotent state transition + unique event ID
Verificationconcurrency test และ duplicate metric
Owner/due datePayout team / ก่อน public launch
Residual riskduplicate จาก provider ที่เปลี่ยน event ID ยังต้อง reconcile
Decision ownerHead of Payments Risk

owner ของ action กับ owner ที่ยอมรับ risk อาจเป็นคนละคน engineer แก้ code ได้ แต่ไม่ควร ตัดสินใจแทนธุรกิจว่า loss scenario แบบใดยอมรับได้

Design Controls as a System

มี 4 แนวทางหลักต่อ risk:

  • Avoid — ตัด capability หรือ data ที่ไม่จำเป็น เช่นไม่เก็บเลขเอกสารเต็มใน analytics
  • Reduce — ลด likelihood/impact ด้วย technical หรือ process controls
  • Transfer/Share — กระจายผลทางสัญญาหรือประกัน แต่ไม่ได้ลบ operational responsibility
  • Accept — ยอมรับอย่างมี owner, เหตุผล, expiry และเงื่อนไขทบทวน

control ควรครอบคลุมหลายจังหวะ:

Control typeหน้าที่ตัวอย่าง callback replay
Preventiveหยุดเหตุsignature, timestamp window, idempotent transition
Detectiveเห็นสิ่งที่หลุดduplicate event metric, ledger anomaly alert
Correctiveจำกัดและแก้ผลdisable provider credential, reconcile และ compensate
Recoveryกลับสู่สภาพควบคุมrestore service, rotate key, verify all pending payouts

Control ที่ไม่มี verification เป็นเพียงความหวัง

“เปิด audit log” ยังไม่พอ ต้องทดสอบว่า event สำคัญถูกบันทึกโดยไม่รั่ว secret, ส่งถึงปลายทางแม้ component ล้ม, แก้ย้อนหลังไม่ได้ และมีคนเห็น alert ทันเวลา

Prefer Controls Near the Invariant

ถ้า invariant คือ “provider event 1 รายการ post ledger ได้ไม่เกิน 1 ครั้ง” control ที่ database transaction และ unique constraint อยู่ใกล้ invariant กว่า in-memory cache หรือ WAF rule การวาง control ที่ boundary ยังมีประโยชน์ แต่ไม่ควรเป็นชั้นเดียวที่ค้ำความถูกต้องของเงิน

Run a Threat Modeling Workshop

session 60–90 นาทีที่มี focus ดีกว่าประชุมครึ่งวันโดยไม่มี artifact:

Before the Session

  • owner เตรียม scope, diagram, data classification และ known constraints
  • เชิญ engineering, product/operations และ Security ตาม risk
  • ส่ง pre-read และระบุ decision ที่ต้องการ

During the Session

  1. ยืนยัน scope และ success/failure ของ business flow
  2. เดิน DFD ทีละเส้นและแก้ diagram ให้ตรงของจริง
  3. ทำเครื่องหมาย trust boundary กับ assumption
  4. ใช้ STRIDE และ abuse cases สร้าง threat
  5. จัดกลุ่ม duplicate และเขียน threat statement
  6. ประเมิน risk แบบหยาบเพื่อหาเรื่องที่ต้อง deep dive
  7. มอบ owner ให้ action และ decision

After the Session

  • ส่ง diagram, threat/risk register และ decision record
  • สร้าง requirement/test ไม่ใช่เก็บทุกอย่างไว้ใน presentation
  • ระบุสิ่งที่ยังไม่รู้และวันที่ทบทวน
  • re-run เมื่อ architecture, data, trust boundary หรือ threat landscape เปลี่ยน

Treat the Model as Living Design

trigger ที่ควรเปิด threat model เดิมกลับมา:

  • เพิ่ม endpoint, privileged action หรือ external integration
  • เปลี่ยน authentication/authorization model
  • เก็บข้อมูลชนิดใหม่หรือย้าย data residency
  • เปลี่ยน network/account boundary
  • มี incident, penetration-test finding หรือ control failure
  • dependency สำคัญเปลี่ยน contract

เพิ่ม threat-model link ใน design document และ pull request ที่เปลี่ยน boundary ไม่จำเป็นต้องประชุมเต็มรูปแบบทุกครั้ง การแก้ diagram และเพิ่ม threat 2 ข้อใน review ก็รักษา model ได้ดีกว่ารอ annual exercise

Review Checklist

  • scope ผูกกับ business capability และระบุสิ่งที่อยู่นอก scope
  • DFD มี external entity, process, data store, data flow และ trust boundary
  • flow สำคัญบอกชนิดข้อมูลและ protocol ไม่ใช่เขียนเพียง request
  • assumption ถูกเขียนและมีคนรับผิดชอบตรวจสอบ
  • ใช้ STRIDE ทุก element/flow และเพิ่ม business abuse case
  • threat statement มี actor, action, weakness และ impact
  • risk rating มีเหตุผล, evidence, confidence และ impact หลายมิติ
  • control ครอบคลุม prevention, detection, correction และ recovery ตามความเสี่ยง
  • ทุก action มี owner; residual risk มี decision owner และ review date
  • model ถูกเชื่อมกับ requirement, test, telemetry และ incident learning

สรุป

Threat modeling คือสะพานจากหลักคิดในบทที่ 1ไปสู่ engineering decision: วาดของจริง, เปิดเผย assumption, สร้าง threat อย่างเป็นระบบ, จัดลำดับด้วย impact และเลือก control ที่ทดสอบได้ เป้าหมายไม่ใช่ทำนายทุกการโจมตี แต่ทำให้ทีมเห็น failure สำคัญก่อนผู้โจมตีและก่อน production

Further Reading