บทที่ 21 · Part 5 — Cloud Security on AWS

AWS Security Foundations & Multi-Account Design

Shared responsibility, account/OU boundaries, Control Tower, SCP guardrails, root protection, break-glass และ secure landing zone

การย้าย workload ขึ้น AWS เปลี่ยนผู้ดูแลบางชั้นของระบบ แต่ไม่ได้โอนความรับผิดชอบด้าน Security ทั้งหมดให้ cloud provider ความผิดพลาดที่สร้างผลกระทบรุนแรงจำนวนมากยังเกิดจาก identity, public exposure, logging gap และ account architecture ที่ลูกค้าต้องออกแบบเอง

Learning Outcomes

  • แยก Security of the Cloud กับ Security in the Cloud ตาม service model ได้
  • ใช้ AWS account และ OU เป็น security boundary โดยไม่ผูกกับ organization chart ได้
  • อธิบาย landing zone, Control Tower และบทบาทของ security accounts ได้
  • ใช้ SCP เป็น permission guardrail โดยไม่เข้าใจผิดว่า SCP grant permission ได้
  • ออกแบบ root protection, delegated administration และ break-glass ที่ตรวจสอบย้อนหลังได้

Shared Responsibility Model

AWS รับผิดชอบ Security of the Cloud เช่น physical facility, hardware, virtualization layer และ managed-service infrastructure ส่วนลูกค้ารับผิดชอบ Security in the Cloud เช่น data, identity, configuration, network policy และ application

ขอบเขตของลูกค้าเปลี่ยนตาม service:

Service modelAWS ดูแลเพิ่มลูกค้ายังต้องดูแล
EC2facility, hardware, hypervisorguest OS, patch, host firewall, application, IAM, data
ECS on Fargatehost OS และ container runtime infrastructureimage, task role, network, secret, application, data
RDSdatabase host/engine operations ตาม serviceaccount, database user, schema/query, network, backup policy, data
S3storage infrastructurebucket/object policy, public access, lifecycle, classification, data
Lambdaruntime infrastructurecode/dependency, execution role, trigger, secret, concurrency, data

Managed ไม่ได้แปลว่า Secure by Default ทุกมิติ

managed service ลด operational burden บางส่วน แต่ authorization, data classification, tenant isolation, secure code, logging coverage และ business controls ยังเป็นความรับผิดชอบของผู้ใช้บริการ

สร้าง Responsibility Matrix

แต่ละ workload ควรมีตารางที่ระบุ:

  • component/service และ data classification
  • control objective เช่น patching, encryption, backup, access review
  • AWS capability ที่เกี่ยวข้อง
  • owner ของ configuration และ owner ของ evidence
  • shared dependency เช่น IdP, KMS, DNS, CI/CD
  • failure/incident path และ recovery test

คำว่า “AWS ดูแลให้” ต้องแปลงเป็นขอบเขตที่พิสูจน์ได้เสมอ

Why Multi-Account

AWS account เป็นทั้ง permission, quota, billing และ isolation boundary การรวม production, development, security logs และ sandbox ไว้ account เดียวทำให้ policy ซับซ้อนและขยาย blast radius

เหตุผลหลักของ multi-account:

  • แยก production จาก non-production
  • จำกัดผลจาก credential หรือ workload compromise
  • แยก sensitive data/workload ตาม regulatory หรือ risk boundary
  • รักษา log/evidence นอก account ที่ถูกโจมตี
  • กระจาย quota และลด resource-name/policy collision
  • ทำ account-level guardrail และ cost attribution

account ไม่ใช่ boundary สมบูรณ์ หาก cross-account trust, organization policy, network และ central services เปิดกว้างเกินไป

Reference Organization

นี่เป็น reference ไม่ใช่สูตรตายตัว จำนวน account ขึ้นกับ ownership, data classification, deployment independence และ risk การแยกมากเกินโดยไม่มี automation จะเพิ่ม policy drift และ operational cost

OU Design

Organizational Unit หรือ OU ใช้จัดกลุ่ม account เพื่อ apply governance policy ควรออกแบบตาม function และ control profile ไม่ใช่ mirror organization chart ซึ่งเปลี่ยนตามการปรับโครงสร้างบุคลากร

ตัวอย่าง control profile:

  • Security: log archive และ security operations
  • Infrastructure: network, DNS, identity หรือ shared CI/CD
  • Workloads: production/non-production แยกตาม sensitivity
  • Sandbox: experimentation พร้อม spend/data/service restrictions
  • Suspended/Quarantine: account ที่ปิดใช้งานหรืออยู่ระหว่าง incident

ควรหลีกเลี่ยง OU tree ลึกโดยไม่มีเหตุผล เพราะ policy inheritance และ exception จะตรวจยาก

Management Account

management account มีอำนาจสูงและไม่ถูกจำกัดด้วย SCP จึงควร:

  • ไม่มี application workload
  • จำกัดผู้เข้าถึงและใช้เฉพาะ task ที่ต้องทำจาก account นี้
  • ใช้ federated, short-lived administrative access
  • delegate supported service administration ไป security/platform account
  • เปิด centralized logging และ alert สำหรับการใช้งาน
  • แยก billing task จาก security administration เมื่อทำได้
  • ทบทวน contact, MFA และ recovery path

การสร้าง IAM user/long-lived access key เพื่อ routine administration ใน management account เพิ่มความเสี่ยง โดยไม่จำเป็น

Security Accounts

Log Archive Account

ใช้รับ immutable หรือ tightly controlled copies ของ organization logs เช่น CloudTrail, Config และ selected service/application logs หลักสำคัญคือ workload account ไม่ควรแก้หรือลบหลักฐานย้อนหลังได้

Security Tooling Account

ใช้ delegated administration, detection aggregation, investigation tooling และ controlled automation การแยกจาก Log Archive ลดโอกาสที่ analyst/tooling permission จะลบ evidence ต้นฉบับ

Network and Shared Services

centralized network หรือ identity ช่วยลดความซ้ำซ้อน แต่กลายเป็น high-blast-radius dependency ต้องมี change control, redundant path, quota monitoring, scoped delegation และ recovery plan

Landing Zone and Control Tower

landing zone คือ environment baseline ที่พร้อมให้ workload เข้าใช้งาน เช่น:

  • organization/account/OU structure
  • identity federation และ access model
  • centralized logging/detection
  • network/DNS/shared services
  • guardrails และ baseline configuration
  • account vending/lifecycle automation
  • tagging, cost และ ownership metadata

AWS Control Tower ช่วยสร้างและ govern landing zone ด้วย account factory, controls และ dashboard แต่ไม่แทน threat model, application security หรือการตรวจว่า control ครอบคลุม requirement จริง

Baseline ต้อง Versioned

policy, account template, infrastructure และ exceptions ควรอยู่ใน version control ผ่าน review/test manual console change ที่ไม่มี reconciliation ทำให้ account ใหม่กับ account เก่ามี security posture ต่างกัน

Service Control Policies

SCP กำหนด maximum available permissions ให้ account/OU แต่ไม่ grant permission การเรียก action ต้องมี allow จาก IAM/resource policy ที่เกี่ยวข้องด้วย

ตัวอย่าง guardrail objectives:

  • deny การปิด centralized logging/configuration
  • จำกัด Region ที่อนุญาตโดยมี exception สำหรับ global services
  • deny public access configuration บางประเภท
  • ป้องกันการออกจาก Organization
  • จำกัดการเปลี่ยน security service delegate
  • require selected resource conditions/tags เมื่อ service รองรับ

สิ่งที่ SCP ไม่ทำ

  • ไม่ grant permission
  • ไม่จำกัด principal ใน management account
  • ไม่กระทบ service-linked role
  • ไม่แทน resource policy, permissions boundary หรือ application authorization
  • ไม่รับประกันว่า resource configuration ปลอดภัยทั้งหมด

Safe Rollout

  1. อ่าน service last accessed data และ inventory dependency
  2. เขียน policy ที่แคบและอธิบาย objective/exception
  3. test ใน account หรือ OU ขนาดเล็ก
  4. ตรวจ denied events และ workload health
  5. ขยาย scope แบบ staged
  6. มี rollback, owner และ expiry สำหรับ exception

การ attach deny กว้างที่ organization root โดยไม่ทดสอบอาจทำให้ incident tooling, backup หรือ deployment หยุดพร้อมกัน

Root User Protection

root user ผูกกับ account และทำ task บางชนิดที่ IAM principal ทำไม่ได้ แนวทางสำคัญ:

  • ห้ามใช้ใน daily operations
  • ไม่มี root access key
  • ใช้ phishing-resistant MFA เมื่อ account capability รองรับ
  • ปกป้อง email, phone และ recovery channel แยกจาก routine administrator
  • เก็บ credential/recovery material ด้วย dual control ตาม risk
  • alert ทุก root sign-in และ root action
  • ทดสอบ recovery procedure โดยไม่เปิดเผย secret

Root Recovery เป็นทั้ง Security และ Availability

การปิดกั้นจนไม่มี recovery path อาจทำให้ account สูญเสียการควบคุม แต่ recovery ที่คนเดียวทำได้ง่าย ก็เป็น account-takeover path ต้องกำหนด decision authority, evidence และ out-of-band verification

Break-Glass Access

break-glass ใช้เมื่อ federation/control plane ปกติล่มหรือเกิด incident ไม่ใช่ shortcut สำหรับงานเร่ง

ควรมี:

  • dedicated path ที่ไม่ใช้ dependency เดียวกับ normal access
  • strong authentication และ material ที่แบ่ง custody ตาม risk
  • minimal but sufficient permissions หรือ elevation process
  • explicit trigger/approver และ time limit
  • alert แบบ real time
  • session/action logging ไป account ที่ผู้ใช้แก้ไม่ได้
  • post-use rotation, review และ incident record
  • periodic exercise ที่ไม่รบกวน production

break-glass ที่ไม่เคยซ้อมอาจใช้ไม่ได้เมื่อเกิดเหตุ ส่วน break-glass ที่ใช้ประจำคือ standing privilege

Account Lifecycle

Provision

  • assign owner, business purpose, environment และ data classification
  • enroll baseline logs, detection, backup และ identity
  • apply OU/SCP/tag/budget/quota baseline
  • verify no default/public resource exposure
  • register contacts และ incident route

Operate

  • inventory accounts และ inactive resources
  • reconcile baseline drift
  • review cross-account trust และ exceptions
  • rotate contacts/owners เมื่อคนหรือทีมเปลี่ยน
  • test backup, recovery และ security telemetry

Close or Quarantine

  • preserve required records/evidence
  • revoke sessions, roles, keys และ external trust
  • remove network/CI/CD/partner paths
  • handle data retention/deletion ตาม policy
  • monitor closure period และ update inventory

การย้าย account ไป Suspended OU อย่างเดียวไม่ลบ credential, data หรือ trust relationship

Governance without Bottlenecks

central platform ควรกำหนด paved road ที่ปลอดภัยและใช้ง่าย:

  • reusable account/workload templates
  • policy-as-code tests ก่อน deploy
  • self-service พร้อม scoped role
  • exception process ที่มี owner, reason, compensating control และ expiry
  • metrics เช่น baseline coverage, exception age, root usage และ log delivery health

guardrail ที่ข้ามได้ง่ายหรือทำให้ delivery หยุดโดยไม่มี alternative มักนำไปสู่ shadow infrastructure

Review Checklist

  • responsibility matrix ระบุ AWS/customer/third-party owner และ evidence
  • account แยก production, non-production, security logs และ high-risk workloads ตาม need
  • OU สะท้อน control profile ไม่ mirror organization chart
  • management account ไม่มี workload และจำกัด routine access
  • Log Archive และ Security Tooling แยก duties/permissions
  • landing zone baseline versioned, tested และตรวจ drift
  • SCP ใช้เป็น maximum-permission guardrail และ staged rollout
  • root ไม่มี access key/daily use พร้อม MFA, alert และ recovery control
  • break-glass independent, time-bound, audited, rotated และ exercised
  • account provisioning, operation, quarantine และ closure มี owner/runbook

สรุป

AWS Security เริ่มจากการแบ่ง responsibility และ blast radius ให้ชัด multi-account architecture, landing zone และ SCP สร้าง governance layer ส่วน identity, workload configuration, application และ data ยังต้องมี control เฉพาะชั้น Account boundary ที่ดีช่วยจำกัดเหตุ แต่ต้องรักษา logging, recovery และ automation ให้ทำงานได้จริงตลอด lifecycle

Further Reading