บทที่ 25 · Part 5 — Cloud Security on AWS
AWS Logging, Detection & Incident Response
CloudTrail, Config, GuardDuty, Macie, Security Hub, centralized evidence, detection engineering, containment และ cloud incident readiness
cloud API ทำให้สร้าง เปลี่ยน และลบ infrastructure ได้ภายในไม่กี่วินาที ความสามารถเดียวกันทำให้ attacker สร้าง persistence หรือปิดหลักฐานได้เร็ว Logging ที่เปิดไว้แต่ไม่ครอบคลุม data events, Region หรือ account จึงอาจให้ความมั่นใจที่ไม่ตรงกับสิ่งที่ตรวจสอบได้จริง
Learning Outcomes
- แยกหน้าที่และ coverage ของ CloudTrail, Config, GuardDuty, Macie และ Security Hub ได้
- ออกแบบ centralized evidence ที่ workload account แก้ย้อนหลังไม่ได้ง่ายได้
- สร้าง detection use case จาก threat, telemetry, owner และ runbook ได้
- เตรียม investigation/containment สำหรับ identity, network, data และ workload incidents ได้
- วาง tabletop, evidence preservation และ recovery ที่คำนึงถึง availability ได้
Telemetry Is an Architecture
logging ไม่ใช่ checkbox ของแต่ละ service แต่เป็น data pipeline ที่มี:
- Source: event, configuration, flow, DNS, application หรือ finding
- Collection: service integration, trail, subscription, agent หรือ API
- Transport: delivery path, retry, buffer และ failure signal
- Storage: account, Region, retention, encryption, integrity และ access
- Analysis: query, rule, correlation และ enrichment
- Response: owner, ticket/page, runbook, containment และ evidence
หากไม่มี monitoring ของ log pipeline การหยุดส่ง log อาจเงียบกว่าการโจมตีที่ต้องการตรวจ
Detection Pipeline
automation ควร enrich และทำ reversible action ที่มี confidence/authority ชัด ส่วน action ที่ทำลาย evidence, ปิดระบบสำคัญ หรือกระทบลูกค้าจำนวนมากต้องมี human decision ตาม policy
CloudTrail
AWS CloudTrail บันทึก activity จาก console, CLI, SDK และ AWS services ตาม event coverage ที่เลือก
Event Types
- Management events: control-plane operations เช่นสร้าง role, เปลี่ยน Security Group
- Data events: high-volume resource operations เช่น S3 object หรือ Lambda invocation บางชนิด
- Insights events: unusual API call/error-rate activity เมื่อเปิดและมี baseline
- Network activity events: visibility สำหรับ supported VPC endpoint scenarios ตาม service capability
การเปิด management events ไม่ได้ทำให้เห็น object read/write ทุกประเภท ต้องเลือก data event selectors ตาม asset/risk พร้อมประเมิน volume และ cost
Event History Is Not a Durable Audit Design
CloudTrail Event History ที่เปิดอัตโนมัติแสดง 90 วันของ management events ต่อ Region และไม่มี data events สำหรับ durable/centralized history ต้องสร้าง trail หรือ CloudTrail Lake event data store
ข้อพลาดที่พบบ่อย:
- ตรวจเฉพาะ Region หลักแต่มี activity ใน opt-in/unused Region
- ไม่มี organization trail ทำให้ account ใหม่หลุด coverage
- S3 destination policy/KMS key ทำให้ delivery fail
- selector ไม่ครอบคลุม sensitive data resource
- retention สั้นกว่า investigation/regulatory need
- analyst เชื่อว่า absence of event เท่ากับ absence of action
Organization Trail
organization trail เก็บ event จาก member accounts และควรส่งสำเนาไป Log Archive account ที่ workload administrator แก้/ลบไม่ได้ง่าย
design decisions:
- multi-Region และ opt-in Region coverage
- management read/write selection
- data event selectors ต่อ critical resource
- S3/Lake retention และ query path
- KMS/key policy dependency
- log bucket write/read/delete separation
- delivery health alert
- account lifecycle enrollment
trail console default บางอย่างช่วยตั้งต้น แต่ architecture ต้องตรวจ effective configuration จริง
Log File Integrity Validation
CloudTrail log-file integrity validation สร้าง digest chain ที่ใช้ SHA-256 และ RSA signature เพื่อช่วยตรวจ ว่าไฟล์ถูกเปลี่ยน/ลบ/เพิ่มหลัง delivery หรือไม่
สิ่งที่กลไกนี้ไม่ได้ทำ:
- ไม่ validate file ให้อัตโนมัติ ต้องใช้ CLI/tooling ตรวจ
- ไม่ป้องกัน attacker ที่มีสิทธิ์ลบ log/digest
- ไม่แก้ missing event selector หรือ delivery outage
- chain ขาดได้เมื่อย้าย/ลบไฟล์ที่จำเป็น
จึงต้องใช้ storage access, retention/Object Lock ตาม requirement และ periodic validation exercise ร่วมกัน
AWS Config
AWS Config บันทึก resource configuration/change และประเมิน resource กับ Config rules/conformance packs เหมาะกับคำถามเช่น:
- Security Group เปิด public เมื่อใด
- bucket encryption/public setting เปลี่ยนอย่างไร
- resource ใดไม่ตรง baseline
- relationship ระหว่าง supported resources เป็นอย่างไร
แต่ Config:
- ไม่ใช่ runtime threat detector
- ไม่ block change ด้วยตัวเอง
- coverage ขึ้นกับ supported resource, recorder, Region และ delivery
- compliant rule ไม่พิสูจน์ว่า application/data ปลอดภัย
- remediation automation ต้องมี permission, failure และ rollback control
Config กับ CloudTrail เสริมกัน: Config แสดง state/timeline ส่วน CloudTrail ช่วยตอบว่า principal/API ใดสร้าง change
Amazon GuardDuty
GuardDuty วิเคราะห์ threat จาก foundational data sources แบบ independent service streams ได้แก่ CloudTrail management events, EC2 VPC flow data และ Route 53 Resolver DNS logs พร้อม protection plans เพิ่มสำหรับ EKS, S3, RDS, Lambda, runtime และ feature อื่นตาม configuration
Important Nuances
- ไม่ต้องเปิด customer VPC Flow Logs เพื่อให้ GuardDuty ใช้ foundational flow source
- แต่ควรเปิด Flow Logs เองตาม investigation/retention need เพราะ GuardDuty finding ไม่ใช่ raw history
- เป็น regional service ต้อง enable/delegate ทุก Region ที่ใช้หรือต้อง monitor
- custom/external DNS resolver query ไม่อยู่ใน Route 53 Resolver DNS source
- protection plan ต้องเปิดและมี coverage/resource prerequisites ของตน
- finding คือ signal ที่ต้อง triage ไม่ใช่ proof สมบูรณ์หรือ automatic breach declaration
central delegated administrator ช่วย aggregate แต่ต้อง test account/Region enrollment และ finding delivery
Amazon Macie
Macie ช่วย discovery/classification sensitive data และสร้าง policy findings สำหรับ S3 ตาม supported storage, file formats และ permissions
automated discovery ใช้ sampling/representative objects และโดยทั่วไปพิจารณา latest versions ดังนั้นผลว่า ไม่พบ sensitive data ไม่ได้พิสูจน์ว่าทุก object/version ไม่มีข้อมูลดังกล่าว
ข้อควรออกแบบ:
- scope/bucket inventory และ unsupported object/format
- sampling กับ targeted sensitive-data discovery job
- custom data identifier และ false positive review
- result repository ที่ encrypt และ retain นานพอ
- access control ต่อ sample/result ที่ sensitive
- finding owner และ remediation path
Macie เก็บ findings/discovery results ใน service 90 วัน ควรส่ง discovery results ไป encrypted S3 repository เมื่อจำเป็นต้องเก็บนานกว่า
Security Hub and Security Hub CSPM
ชื่อ service ปัจจุบันต้องแยกหน้าที่:
- Security Hub CSPM: ประเมิน standards/controls และรวบรวม posture findings โดย controls ส่วนใหญ่พึ่ง AWS Config
- Security Hub: correlate security signals เป็น exposure/attack-path context และ unused-access context ตาม capability ที่เปิดใช้
ข้อควรระวัง:
- findings/controls มี regional behavior
- finding ก่อน enable ไม่ถูก ingest ย้อนหลังโดยอัตโนมัติ
- disabled/suppressed finding ต้องมี owner/reason/expiry
- failed control อาจเป็น real risk, false positive, unsupported architecture หรือ evidence gap
- score/dashboard ไม่แทน risk acceptance และ remediation verification
อย่าเรียกทุก Security Hub capability ว่า CSPM เพราะทำให้ architecture และ expectation คลาดเคลื่อน
Native Service Comparison
| Service | คำถามหลัก | ไม่ได้แทน |
|---|---|---|
| CloudTrail | ใครเรียก AWS API อะไร เมื่อไร จาก context ใด | application transaction log ทุกชนิด |
| Config | resource configuration เปลี่ยน/ตรง rule หรือไม่ | runtime attack detection |
| GuardDuty | threat/anomaly signal จาก supported sources คืออะไร | raw-log archive/IR decision |
| Macie | S3 มี sensitive-data/policy exposure signal ใด | data catalog ทุก service |
| Security Hub CSPM | posture controls/standards มีสถานะอย่างไร | preventive enforcement |
| Security Hub | signals เชื่อมเป็น exposure/path ใด | owner/runbook/containment |
Application Security Telemetry
AWS native logs ไม่รู้ business intent จึงต้องมี application events เช่น:
- authentication success/failure/recovery/MFA change
- authorization denial และ privileged action
- beneficiary/payment/instruction creation/change/approval
- transaction amount/currency/source/destination binding
- idempotency/replay/rate/concurrency control result
- admin override และ maker-checker decision
- data export/search/bulk read
- webhook/signature validation
ไม่ log credential, raw token, secret, full payment instrument หรือ sensitive payload โดยไม่จำเป็น ใช้ stable correlation ID และ actor/subject/object/action/result/reason/context ที่ data policy อนุญาต
Centralized Evidence
Log Archive design ควรให้ source account เขียนผ่าน controlled service path แต่แก้/ลบย้อนหลังไม่ได้
ควรพิจารณา:
- organization-level delivery
- separate write/read/administration roles
- Object Lock/retention เมื่อ requirement ต้องการ immutability
- KMS key availability และ key-admin separation
- least-privilege analyst query path
- masking/partitioning สำหรับ sensitive logs
- cross-account/Region replication ตาม recovery need
- deletion/legal hold process
- cost/query performance และ schema/version
immutable log ที่ไม่มี parser/query/runbook ยังใช้ตอบ incident ได้ช้า ต้องซ้อม retrieval ด้วย
Detection Engineering
ทุก detection use case ควรระบุ:
- threat/abuse story และ protected asset
- prerequisite และ expected event sequence
- required telemetry กับ blind spots
- query/rule/correlation window
- severity/confidence/context enrichment
- owner, SLA และ escalation
- containment authority/rollback
- simulation/test และ expected evidence
- false-positive/false-negative feedback
ตัวอย่าง cloud identity use case:
- privileged role assume จาก source/attribute ใหม่
- ตามด้วย
GetSecretValueหรือ snapshot export - ตามด้วย trust policy/logging/security-control change
แต่ละ event เดี่ยวอาจถูกต้อง การ correlate ตาม identity, session, time และ asset sensitivity ช่วยเพิ่ม signal
Alert Quality
metric ที่มีประโยชน์:
- telemetry coverage และ delivery health
- mean time to acknowledge/triage/contain
- true/false positive แยก use case
- stale alert และ alert without owner
- finding age ตาม severity/asset
- repeat finding หลัง remediation
- detection test pass rate
- accounts/Regions/resources ที่หลุด enrollment
จำนวน finding ที่ปิดไม่เท่ากับ risk ลด หากปิดแบบ suppress โดยไม่แก้ root cause
Incident Readiness
เตรียมก่อนเกิดเหตุ:
- incident roles, on-call และ executive/legal/compliance contacts
- severity/decision authority
- out-of-band communication
- security tooling/evidence access ที่ไม่พึ่ง compromised IdP path เดียว
- approved investigation tools และ isolated forensic account
- service quota/support/provider contacts
- containment runbooks ต่อ identity, data, network และ workload
- backup/key/restore procedure
- customer/partner/regulator communication process ที่คนมี authority ตัดสิน
AWS Support หรือ managed incident service ช่วยบางส่วน แต่ไม่แทน ownership, detection coverage หรือ legal advice
Investigation Workflow
- Validate: finding/event จริงหรือ expected change
- Scope: accounts, Regions, identities, sessions, resources, data และ time window
- Preserve: CloudTrail, Config snapshot, application/IdP/network/runtime evidence
- Contain: ลด access/path ด้วย action ที่เหมาะสมและ reversible เมื่อทำได้
- Eradicate: ลบ persistence และแก้ root cause
- Recover: restore known-good state, monitor และทยอยคืน service
- Learn: timeline, control gap, owner, due date และ detection regression test
ต้องบันทึก actor, time, reason, command/change และ evidence reference ของ responder เองด้วย
Containment Patterns
Identity
- disable/delete access key
- revoke role sessions ตาม supported mechanism
- add scoped explicit deny หรือแก้ trust path
- disable federated identity/session
- rotate downstream secret ที่ถูกอ่านได้
Network
- quarantine workload ด้วย restrictive SG/NACL/firewall/route
- block egress destination แบบ time-bound
- remove public route/endpoint/peering
- preserve flow/DNS/firewall evidence
Workload
- isolate instance/task/pod/function
- capture memory/disk/runtime evidence ตาม forensic plan
- stop deployment pipeline และ pin known-good artifact
- replace immutable workload จาก trusted image/config
Data
- deny object/database/export path
- revoke pre-signed URL/session/replication trust
- preserve versions/snapshot/logs
- assess accessed data ไม่ใช่เพียง resource ที่ถูกแตะ
Containment อาจทำลาย Evidence หรือ Availability
terminate instance, rotate key, disable KMS key หรือเปลี่ยน route อาจทำให้หลักฐานหายและบริการหยุด responder ต้องใช้ pre-approved authority, preserve volatile evidence ตามความเหมาะสม และบันทึก decision การเปลี่ยน SG อาจไม่ตัด tracked connection ที่มีอยู่ทันที ต้อง test behavior จริง
Automation Safety
automation เหมาะกับ enrichment, tagging, snapshot, ticket และ reversible quarantine ที่ confidence สูง
ก่อน auto-remediation ต้องกำหนด:
- exact trigger และ trust ของ source
- idempotency/concurrency
- permission boundary ของ automation role
- dry-run/test account
- blast radius/quota/rate limit
- rollback และ kill switch
- human notification
- evidence preservation
- behavior เมื่อ dependency ล่ม
automation ที่ถูก attacker trigger ได้อาจกลายเป็น denial-of-service tool
Tabletop and Game Days
scenario ที่ควรซ้อม:
- privileged federated session ถูกขโมย
- public S3/data export exposure
- compromised workload role และ secret access
- malicious image/deployment pipeline
- logging delivery ถูกปิด
- KMS key disable/delete request
- DDoS/partner outage พร้อม degraded mode
- ransomware/operator deletion และ restore
exercise ต้องเก็บเวลาจริง, missing permission/log/contact, decision ambiguity และ corrective actions ไม่ใช่เพียง ประชุมอ่าน runbook
Review Checklist
- organization trail ครอบคลุม accounts/Regions พร้อม data-event selectors ตาม risk
- Event History ไม่ถูกใช้แทน durable centralized audit trail
- log delivery failure, selector/config และ integrity validation ถูก monitor/test
- Config recorder/rules/remediation มี coverage, owner และ rollback
- GuardDuty enable/delegate ทุก Region และ protection plan ที่ตั้งใจ
- customer Flow Logs เก็บตาม investigation need แม้ GuardDuty ไม่ต้องพึ่ง
- Macie sampling/format/version limitations ถูกบันทึกและ targeted jobs เติมช่องว่าง
- Security Hub กับ Security Hub CSPM ใช้ตามหน้าที่ปัจจุบัน
- native cloud findings เชื่อม application security events
- evidence store แยก write/read/admin พร้อม retention/key/retrieval test
- detection มี threat, telemetry, owner, runbook, test และ feedback
- incident access, contacts, containment, evidence และ recovery ถูก exercise
- automation reversible, scoped, idempotent และมี kill switch
สรุป
CloudTrail, Config, GuardDuty, Macie และ Security Hub ตอบคำถามคนละแบบ ไม่มี service เดียวให้ทั้ง prevention, evidence, detection และ response ระบบที่พร้อมรับ incident ต้องรวม cloud กับ application telemetry, ปกป้อง pipeline เอง, ทดสอบ detection และให้คนที่มี authority ตัดสิน containment/recovery ที่มีผลต่อธุรกิจ