บทที่ 25 · Part 5 — Cloud Security on AWS

AWS Logging, Detection & Incident Response

CloudTrail, Config, GuardDuty, Macie, Security Hub, centralized evidence, detection engineering, containment และ cloud incident readiness

cloud API ทำให้สร้าง เปลี่ยน และลบ infrastructure ได้ภายในไม่กี่วินาที ความสามารถเดียวกันทำให้ attacker สร้าง persistence หรือปิดหลักฐานได้เร็ว Logging ที่เปิดไว้แต่ไม่ครอบคลุม data events, Region หรือ account จึงอาจให้ความมั่นใจที่ไม่ตรงกับสิ่งที่ตรวจสอบได้จริง

Learning Outcomes

  • แยกหน้าที่และ coverage ของ CloudTrail, Config, GuardDuty, Macie และ Security Hub ได้
  • ออกแบบ centralized evidence ที่ workload account แก้ย้อนหลังไม่ได้ง่ายได้
  • สร้าง detection use case จาก threat, telemetry, owner และ runbook ได้
  • เตรียม investigation/containment สำหรับ identity, network, data และ workload incidents ได้
  • วาง tabletop, evidence preservation และ recovery ที่คำนึงถึง availability ได้

Telemetry Is an Architecture

logging ไม่ใช่ checkbox ของแต่ละ service แต่เป็น data pipeline ที่มี:

  • Source: event, configuration, flow, DNS, application หรือ finding
  • Collection: service integration, trail, subscription, agent หรือ API
  • Transport: delivery path, retry, buffer และ failure signal
  • Storage: account, Region, retention, encryption, integrity และ access
  • Analysis: query, rule, correlation และ enrichment
  • Response: owner, ticket/page, runbook, containment และ evidence

หากไม่มี monitoring ของ log pipeline การหยุดส่ง log อาจเงียบกว่าการโจมตีที่ต้องการตรวจ

Detection Pipeline

automation ควร enrich และทำ reversible action ที่มี confidence/authority ชัด ส่วน action ที่ทำลาย evidence, ปิดระบบสำคัญ หรือกระทบลูกค้าจำนวนมากต้องมี human decision ตาม policy

CloudTrail

AWS CloudTrail บันทึก activity จาก console, CLI, SDK และ AWS services ตาม event coverage ที่เลือก

Event Types

  • Management events: control-plane operations เช่นสร้าง role, เปลี่ยน Security Group
  • Data events: high-volume resource operations เช่น S3 object หรือ Lambda invocation บางชนิด
  • Insights events: unusual API call/error-rate activity เมื่อเปิดและมี baseline
  • Network activity events: visibility สำหรับ supported VPC endpoint scenarios ตาม service capability

การเปิด management events ไม่ได้ทำให้เห็น object read/write ทุกประเภท ต้องเลือก data event selectors ตาม asset/risk พร้อมประเมิน volume และ cost

Event History Is Not a Durable Audit Design

CloudTrail Event History ที่เปิดอัตโนมัติแสดง 90 วันของ management events ต่อ Region และไม่มี data events สำหรับ durable/centralized history ต้องสร้าง trail หรือ CloudTrail Lake event data store

ข้อพลาดที่พบบ่อย:

  • ตรวจเฉพาะ Region หลักแต่มี activity ใน opt-in/unused Region
  • ไม่มี organization trail ทำให้ account ใหม่หลุด coverage
  • S3 destination policy/KMS key ทำให้ delivery fail
  • selector ไม่ครอบคลุม sensitive data resource
  • retention สั้นกว่า investigation/regulatory need
  • analyst เชื่อว่า absence of event เท่ากับ absence of action

Organization Trail

organization trail เก็บ event จาก member accounts และควรส่งสำเนาไป Log Archive account ที่ workload administrator แก้/ลบไม่ได้ง่าย

design decisions:

  • multi-Region และ opt-in Region coverage
  • management read/write selection
  • data event selectors ต่อ critical resource
  • S3/Lake retention และ query path
  • KMS/key policy dependency
  • log bucket write/read/delete separation
  • delivery health alert
  • account lifecycle enrollment

trail console default บางอย่างช่วยตั้งต้น แต่ architecture ต้องตรวจ effective configuration จริง

Log File Integrity Validation

CloudTrail log-file integrity validation สร้าง digest chain ที่ใช้ SHA-256 และ RSA signature เพื่อช่วยตรวจ ว่าไฟล์ถูกเปลี่ยน/ลบ/เพิ่มหลัง delivery หรือไม่

สิ่งที่กลไกนี้ไม่ได้ทำ:

  • ไม่ validate file ให้อัตโนมัติ ต้องใช้ CLI/tooling ตรวจ
  • ไม่ป้องกัน attacker ที่มีสิทธิ์ลบ log/digest
  • ไม่แก้ missing event selector หรือ delivery outage
  • chain ขาดได้เมื่อย้าย/ลบไฟล์ที่จำเป็น

จึงต้องใช้ storage access, retention/Object Lock ตาม requirement และ periodic validation exercise ร่วมกัน

AWS Config

AWS Config บันทึก resource configuration/change และประเมิน resource กับ Config rules/conformance packs เหมาะกับคำถามเช่น:

  • Security Group เปิด public เมื่อใด
  • bucket encryption/public setting เปลี่ยนอย่างไร
  • resource ใดไม่ตรง baseline
  • relationship ระหว่าง supported resources เป็นอย่างไร

แต่ Config:

  • ไม่ใช่ runtime threat detector
  • ไม่ block change ด้วยตัวเอง
  • coverage ขึ้นกับ supported resource, recorder, Region และ delivery
  • compliant rule ไม่พิสูจน์ว่า application/data ปลอดภัย
  • remediation automation ต้องมี permission, failure และ rollback control

Config กับ CloudTrail เสริมกัน: Config แสดง state/timeline ส่วน CloudTrail ช่วยตอบว่า principal/API ใดสร้าง change

Amazon GuardDuty

GuardDuty วิเคราะห์ threat จาก foundational data sources แบบ independent service streams ได้แก่ CloudTrail management events, EC2 VPC flow data และ Route 53 Resolver DNS logs พร้อม protection plans เพิ่มสำหรับ EKS, S3, RDS, Lambda, runtime และ feature อื่นตาม configuration

Important Nuances

  • ไม่ต้องเปิด customer VPC Flow Logs เพื่อให้ GuardDuty ใช้ foundational flow source
  • แต่ควรเปิด Flow Logs เองตาม investigation/retention need เพราะ GuardDuty finding ไม่ใช่ raw history
  • เป็น regional service ต้อง enable/delegate ทุก Region ที่ใช้หรือต้อง monitor
  • custom/external DNS resolver query ไม่อยู่ใน Route 53 Resolver DNS source
  • protection plan ต้องเปิดและมี coverage/resource prerequisites ของตน
  • finding คือ signal ที่ต้อง triage ไม่ใช่ proof สมบูรณ์หรือ automatic breach declaration

central delegated administrator ช่วย aggregate แต่ต้อง test account/Region enrollment และ finding delivery

Amazon Macie

Macie ช่วย discovery/classification sensitive data และสร้าง policy findings สำหรับ S3 ตาม supported storage, file formats และ permissions

automated discovery ใช้ sampling/representative objects และโดยทั่วไปพิจารณา latest versions ดังนั้นผลว่า ไม่พบ sensitive data ไม่ได้พิสูจน์ว่าทุก object/version ไม่มีข้อมูลดังกล่าว

ข้อควรออกแบบ:

  • scope/bucket inventory และ unsupported object/format
  • sampling กับ targeted sensitive-data discovery job
  • custom data identifier และ false positive review
  • result repository ที่ encrypt และ retain นานพอ
  • access control ต่อ sample/result ที่ sensitive
  • finding owner และ remediation path

Macie เก็บ findings/discovery results ใน service 90 วัน ควรส่ง discovery results ไป encrypted S3 repository เมื่อจำเป็นต้องเก็บนานกว่า

Security Hub and Security Hub CSPM

ชื่อ service ปัจจุบันต้องแยกหน้าที่:

  • Security Hub CSPM: ประเมิน standards/controls และรวบรวม posture findings โดย controls ส่วนใหญ่พึ่ง AWS Config
  • Security Hub: correlate security signals เป็น exposure/attack-path context และ unused-access context ตาม capability ที่เปิดใช้

ข้อควรระวัง:

  • findings/controls มี regional behavior
  • finding ก่อน enable ไม่ถูก ingest ย้อนหลังโดยอัตโนมัติ
  • disabled/suppressed finding ต้องมี owner/reason/expiry
  • failed control อาจเป็น real risk, false positive, unsupported architecture หรือ evidence gap
  • score/dashboard ไม่แทน risk acceptance และ remediation verification

อย่าเรียกทุก Security Hub capability ว่า CSPM เพราะทำให้ architecture และ expectation คลาดเคลื่อน

Native Service Comparison

Serviceคำถามหลักไม่ได้แทน
CloudTrailใครเรียก AWS API อะไร เมื่อไร จาก context ใดapplication transaction log ทุกชนิด
Configresource configuration เปลี่ยน/ตรง rule หรือไม่runtime attack detection
GuardDutythreat/anomaly signal จาก supported sources คืออะไรraw-log archive/IR decision
MacieS3 มี sensitive-data/policy exposure signal ใดdata catalog ทุก service
Security Hub CSPMposture controls/standards มีสถานะอย่างไรpreventive enforcement
Security Hubsignals เชื่อมเป็น exposure/path ใดowner/runbook/containment

Application Security Telemetry

AWS native logs ไม่รู้ business intent จึงต้องมี application events เช่น:

  • authentication success/failure/recovery/MFA change
  • authorization denial และ privileged action
  • beneficiary/payment/instruction creation/change/approval
  • transaction amount/currency/source/destination binding
  • idempotency/replay/rate/concurrency control result
  • admin override และ maker-checker decision
  • data export/search/bulk read
  • webhook/signature validation

ไม่ log credential, raw token, secret, full payment instrument หรือ sensitive payload โดยไม่จำเป็น ใช้ stable correlation ID และ actor/subject/object/action/result/reason/context ที่ data policy อนุญาต

Centralized Evidence

Log Archive design ควรให้ source account เขียนผ่าน controlled service path แต่แก้/ลบย้อนหลังไม่ได้

ควรพิจารณา:

  • organization-level delivery
  • separate write/read/administration roles
  • Object Lock/retention เมื่อ requirement ต้องการ immutability
  • KMS key availability และ key-admin separation
  • least-privilege analyst query path
  • masking/partitioning สำหรับ sensitive logs
  • cross-account/Region replication ตาม recovery need
  • deletion/legal hold process
  • cost/query performance และ schema/version

immutable log ที่ไม่มี parser/query/runbook ยังใช้ตอบ incident ได้ช้า ต้องซ้อม retrieval ด้วย

Detection Engineering

ทุก detection use case ควรระบุ:

  1. threat/abuse story และ protected asset
  2. prerequisite และ expected event sequence
  3. required telemetry กับ blind spots
  4. query/rule/correlation window
  5. severity/confidence/context enrichment
  6. owner, SLA และ escalation
  7. containment authority/rollback
  8. simulation/test และ expected evidence
  9. false-positive/false-negative feedback

ตัวอย่าง cloud identity use case:

  • privileged role assume จาก source/attribute ใหม่
  • ตามด้วย GetSecretValue หรือ snapshot export
  • ตามด้วย trust policy/logging/security-control change

แต่ละ event เดี่ยวอาจถูกต้อง การ correlate ตาม identity, session, time และ asset sensitivity ช่วยเพิ่ม signal

Alert Quality

metric ที่มีประโยชน์:

  • telemetry coverage และ delivery health
  • mean time to acknowledge/triage/contain
  • true/false positive แยก use case
  • stale alert และ alert without owner
  • finding age ตาม severity/asset
  • repeat finding หลัง remediation
  • detection test pass rate
  • accounts/Regions/resources ที่หลุด enrollment

จำนวน finding ที่ปิดไม่เท่ากับ risk ลด หากปิดแบบ suppress โดยไม่แก้ root cause

Incident Readiness

เตรียมก่อนเกิดเหตุ:

  • incident roles, on-call และ executive/legal/compliance contacts
  • severity/decision authority
  • out-of-band communication
  • security tooling/evidence access ที่ไม่พึ่ง compromised IdP path เดียว
  • approved investigation tools และ isolated forensic account
  • service quota/support/provider contacts
  • containment runbooks ต่อ identity, data, network และ workload
  • backup/key/restore procedure
  • customer/partner/regulator communication process ที่คนมี authority ตัดสิน

AWS Support หรือ managed incident service ช่วยบางส่วน แต่ไม่แทน ownership, detection coverage หรือ legal advice

Investigation Workflow

  1. Validate: finding/event จริงหรือ expected change
  2. Scope: accounts, Regions, identities, sessions, resources, data และ time window
  3. Preserve: CloudTrail, Config snapshot, application/IdP/network/runtime evidence
  4. Contain: ลด access/path ด้วย action ที่เหมาะสมและ reversible เมื่อทำได้
  5. Eradicate: ลบ persistence และแก้ root cause
  6. Recover: restore known-good state, monitor และทยอยคืน service
  7. Learn: timeline, control gap, owner, due date และ detection regression test

ต้องบันทึก actor, time, reason, command/change และ evidence reference ของ responder เองด้วย

Containment Patterns

Identity

  • disable/delete access key
  • revoke role sessions ตาม supported mechanism
  • add scoped explicit deny หรือแก้ trust path
  • disable federated identity/session
  • rotate downstream secret ที่ถูกอ่านได้

Network

  • quarantine workload ด้วย restrictive SG/NACL/firewall/route
  • block egress destination แบบ time-bound
  • remove public route/endpoint/peering
  • preserve flow/DNS/firewall evidence

Workload

  • isolate instance/task/pod/function
  • capture memory/disk/runtime evidence ตาม forensic plan
  • stop deployment pipeline และ pin known-good artifact
  • replace immutable workload จาก trusted image/config

Data

  • deny object/database/export path
  • revoke pre-signed URL/session/replication trust
  • preserve versions/snapshot/logs
  • assess accessed data ไม่ใช่เพียง resource ที่ถูกแตะ

Containment อาจทำลาย Evidence หรือ Availability

terminate instance, rotate key, disable KMS key หรือเปลี่ยน route อาจทำให้หลักฐานหายและบริการหยุด responder ต้องใช้ pre-approved authority, preserve volatile evidence ตามความเหมาะสม และบันทึก decision การเปลี่ยน SG อาจไม่ตัด tracked connection ที่มีอยู่ทันที ต้อง test behavior จริง

Automation Safety

automation เหมาะกับ enrichment, tagging, snapshot, ticket และ reversible quarantine ที่ confidence สูง

ก่อน auto-remediation ต้องกำหนด:

  • exact trigger และ trust ของ source
  • idempotency/concurrency
  • permission boundary ของ automation role
  • dry-run/test account
  • blast radius/quota/rate limit
  • rollback และ kill switch
  • human notification
  • evidence preservation
  • behavior เมื่อ dependency ล่ม

automation ที่ถูก attacker trigger ได้อาจกลายเป็น denial-of-service tool

Tabletop and Game Days

scenario ที่ควรซ้อม:

  • privileged federated session ถูกขโมย
  • public S3/data export exposure
  • compromised workload role และ secret access
  • malicious image/deployment pipeline
  • logging delivery ถูกปิด
  • KMS key disable/delete request
  • DDoS/partner outage พร้อม degraded mode
  • ransomware/operator deletion และ restore

exercise ต้องเก็บเวลาจริง, missing permission/log/contact, decision ambiguity และ corrective actions ไม่ใช่เพียง ประชุมอ่าน runbook

Review Checklist

  • organization trail ครอบคลุม accounts/Regions พร้อม data-event selectors ตาม risk
  • Event History ไม่ถูกใช้แทน durable centralized audit trail
  • log delivery failure, selector/config และ integrity validation ถูก monitor/test
  • Config recorder/rules/remediation มี coverage, owner และ rollback
  • GuardDuty enable/delegate ทุก Region และ protection plan ที่ตั้งใจ
  • customer Flow Logs เก็บตาม investigation need แม้ GuardDuty ไม่ต้องพึ่ง
  • Macie sampling/format/version limitations ถูกบันทึกและ targeted jobs เติมช่องว่าง
  • Security Hub กับ Security Hub CSPM ใช้ตามหน้าที่ปัจจุบัน
  • native cloud findings เชื่อม application security events
  • evidence store แยก write/read/admin พร้อม retention/key/retrieval test
  • detection มี threat, telemetry, owner, runbook, test และ feedback
  • incident access, contacts, containment, evidence และ recovery ถูก exercise
  • automation reversible, scoped, idempotent และมี kill switch

สรุป

CloudTrail, Config, GuardDuty, Macie และ Security Hub ตอบคำถามคนละแบบ ไม่มี service เดียวให้ทั้ง prevention, evidence, detection และ response ระบบที่พร้อมรับ incident ต้องรวม cloud กับ application telemetry, ปกป้อง pipeline เอง, ทดสอบ detection และให้คนที่มี authority ตัดสิน containment/recovery ที่มีผลต่อธุรกิจ

Further Reading