ABOUT ME

-

Today
-
Yesterday
-
Total
-
  • [AWS RDS][운영] promoted replica handoff 뒤 manual snapshot leftovers와 retained automated backup cleanup을 어떤 순서표로 따로 남기나
    기타개발지식/풀스택개발 2026. 8. 26. 20:15

    IT 리서치 노트

    [AWS RDS][운영] promoted replica handoff 뒤 manual snapshot leftovers와 retained automated backup cleanup을 어떤 순서표로 따로 남기나

    Amazon RDS에서 source DB 삭제 뒤 promoted replica handoff가 끝나면 많은 팀이 retained automated backup과 snapshot 남은 것을 한 묶음으로 지운다. 하지만 2026년 8월 26일 기준 AWS 공식 문서를 다시 보면 manual snapshot leftovers와 retained automated backup cleanup은 생성 이유도, 삭제 경로도, 남겨야 할 evidence도 다르다. 이 글은 promoted replica handoff 뒤 surviving instance 운영, manual snapshot leftovers, retained automated backup cleanup을 어떤 순서표로 따로 관리해야 하는지 정리한다.

    1. 개요

    결론부터 말하면 promoted replica handoff 뒤에는 manual snapshot leftovers와 retained automated backup cleanup을 같은 큐로 두지 않는 편이 좋다. manual snapshot은 intentional restore asset일 수 있고, retained automated backup은 retention과 deletion workflow를 따르는 비용 cleanup 대상일 수 있다.

    source 삭제와 handoff가 끝났다고 restore 자산이 하나의 목록으로 정리되는 것은 아니다. 이미 LatestRestorableTime 종료 글이 PITR 창 closure를 다뤘다면 이번 글은 snapshot leftovers와 retained backup cleanup 자체를 더 세밀하게 나누는 단계다.

    2. 어디서 실제로 막히는가

    가장 흔한 실패는 promoted replica handoff 완료 직후 Snapshots와 Retained 탭을 같이 훑고 모두 '남은 backup'으로만 적는 것이다. 이렇게 되면 restore asset으로 의도적으로 남긴 manual snapshot과 빨리 정리해야 하는 retained automated backup이 같은 우선순위로 보인다.

    두 번째 실패는 automated backup을 장기 보존하려고 manual snapshot으로 복사한 뒤, incident note에는 그 생성 근거를 안 남기는 것이다. 그러면 한 달 뒤 snapshot leftovers가 왜 남아 있는지와 삭제해도 되는지 판단하는 시간이 길어진다.

    세 번째 실패는 surviving instance handoff와 restore asset cleanup을 같은 owner가 끝냈다고 가정하는 것이다. promoted replica는 운영팀의 health check와 failover memo가 먼저고, retained backup cleanup은 비용과 복구팀의 검토가 먼저일 수 있다.

    • 증상: handoff는 끝났는데 snapshot leftovers와 retained backup이 한 목록으로만 남는다.
    • 실패: manual snapshot을 intentional keep인지 leftover인지 구분하지 않는다.
    • 막힘: retained automated backup 삭제 경로와 snapshot 삭제 경로를 같은 작업으로 적는다.
    • 누락: restore asset 유지 근거와 다음 review date를 남기지 않는다.
    헷갈리는 상태 실제 의미 먼저 볼 곳
    snapshot이 아직 남아 있다 manual restore asset일 수도 있다 Snapshots와 keep reason 메모
    retained backup이 아직 남아 있다 retention 만료 전 cleanup 대상일 수 있다 Retained 탭과 DbiResourceId 기록
    handoff는 끝났다 surviving instance 운영은 정리됐다는 뜻일 뿐 restore asset과 비용 cleanup은 별도 확인

    3. 실무에서 적용하는 순서

    가장 실용적인 순서는 여섯 단계다. 먼저 surviving instance handoff를 운영 관점에서 닫는다. 다음으로 manual snapshot leftovers를 inventory 하고 keep reason을 적는다. 세 번째로 retained automated backup 목록을 분리한다. 네 번째로 automated backup을 장기 보존하려면 manual snapshot 복사 여부를 기록한다. 다섯 번째로 retained backup cleanup을 실행한다. 마지막으로 manual snapshot leftovers를 유지할지 삭제할지 다음 review date와 함께 남긴다.

    1. surviving promoted replica handoff를 먼저 닫는다.
    2. manual snapshot leftovers를 inventory 하고 keep reason을 적는다.
    3. retained automated backup 목록을 별도 추출한다.
    4. 장기 보존이면 automated backup copy를 manual snapshot으로 전환했는지 확인한다.
    5. retained backup cleanup을 Retained 탭 또는 CLI로 실행한다.
    6. manual snapshot leftovers의 next review date를 남긴다.

    콘솔에서는 먼저 DB instances 화면에서 surviving instance 상태를 조회하고 handoff owner 메모를 저장한다. 그다음 Snapshots 메뉴를 클릭해 manual snapshot 이름과 keep reason을 확인한다. 이어서 Automated backups의 Retained 탭을 클릭해 DbiResourceId와 retention 값을 조회하고, 삭제 대상이면 콘솔이나 CLI 명령을 실행해 delete time을 기록한다. 마지막으로 cleanup 파일에 next review date와 delete owner를 입력한다.

    실무 메모는 최소한 manual_snapshot_reason, manual_snapshot_delete_owner, retained_backup_deleted_at, dbi_resource_id를 포함하는 편이 좋다. snapshot이 남는 이유와 retained backup 삭제 여부가 분리돼야 월말 비용과 restore 준비 상태를 동시에 설명할 수 있다.

    post-handoff memo 예시
    surviving_instance=prod-db-replica-1
    handoff_completed_at=2026-08-26T10:08+09:00
    
    manual_snapshot_name=prod-db-clean-exit
    manual_snapshot_reason=retain_restore_asset
    manual_snapshot_delete_owner=db-ops
    manual_snapshot_next_review=2026-09-30
    
    retained_backup_present=true
    dbi_resource_id=db-ABCDEFG123456789
    retained_backup_deleted_at=2026-08-26T10:19+09:00
    
    incident_close=pending_snapshot_review

    이 구조를 쓰면 surviving instance 운영 안정화와 restore asset 검토, 비용 cleanup을 같은 티켓 안에서도 서로 다른 큐로 움직일 수 있다. handoff가 끝났다고 restore asset 정리까지 자동으로 끝나는 것은 아니다.

    4. 공식 문서와 예시 화면으로 확인하기

    첫 공식 화면은 source DB 삭제 후에도 manual snapshot은 자동으로 사라지지 않는다는 문장이다. promoted replica handoff가 끝나도 restore 자산이 따로 남을 수 있다는 뜻이다.

    delete-db-instance 문서는 manual DB snapshots가 source 삭제와 함께 자동 삭제되지 않는다고 설명한다.
    delete-db-instance 문서는 manual DB snapshots가 source 삭제와 함께 자동 삭제되지 않는다고 설명한다.

    즉 deleting 종료 뒤 수동 snapshot leftovers는 cleanup 실패일 수도 있지만 intentional restore asset일 수도 있다. retained automated backup과 같은 버킷으로만 보면 판단이 불명확해진다.

    두 번째 자료는 final snapshot과 retained automated backup의 독립성이다. cleanup 메모는 snapshot leftovers와 retained backup을 별도 열로 가져가야 한다.

    AWS는 final snapshot과 retained automated backup을 서로 독립적인 객체로 설명한다.
    AWS는 final snapshot과 retained automated backup을 서로 독립적인 객체로 설명한다.

    manual snapshot leftovers는 보존 근거를, retained backup은 만료와 삭제 계획을 따로 적는 편이 맞다. 하나는 사람이 지울 때까지 남고 다른 하나는 retention과 deletion workflow를 탄다.

    세 번째 화면은 automated backup을 장기 보존 manual snapshot으로 바꾸는 경계다. promoted replica handoff 뒤 남아 있는 snapshot이 단순 leftover인지 intentional copy인지 여기서 갈린다.

    AWS는 장기 보존이 필요하면 automated backup을 manual snapshot으로 복사해 두라고 안내한다.
    AWS는 장기 보존이 필요하면 automated backup을 manual snapshot으로 복사해 두라고 안내한다.

    따라서 incident close 때는 manual snapshot이 생성 근거와 삭제 owner를 가져야 한다. retained automated backup 삭제와 같은 칼럼에 넣으면 why-kept evidence가 사라진다.

    네 번째 자료는 retained automated backup 삭제 경로다. manual snapshot leftovers와 달리 retained backup은 Retained 탭과 DbiResourceId 중심으로 지우는 절차를 탄다.

    retained automated backup cleanup은 Retained 탭과 delete-db-instance-automated-backup 경로를 따른다.
    retained automated backup cleanup은 Retained 탭과 delete-db-instance-automated-backup 경로를 따른다.

    이 경로가 snapshot 삭제 경로와 다르다는 점을 메모에 남겨야 한다. cleanup owner가 snapshot과 retained backup을 같은 큐로 처리한다고 가정하면 삭제 누락이 생기기 쉽다.

    다섯 번째 공식 화면은 manual snapshot 삭제 경로다. retained automated backup cleanup과 삭제 버튼 위치부터 다르다.

    manual snapshot leftovers는 Snapshots 경로에서 delete snapshot으로 별도 정리한다.
    manual snapshot leftovers는 Snapshots 경로에서 delete snapshot으로 별도 정리한다.

    즉 promoted replica handoff 뒤 cleanup 순서는 surviving instance, retained backup, manual snapshot 세 줄로 나눠야 한다. restore asset을 비용 cleanup 대상과 섞으면 남겨야 할 snapshot까지 지우거나 반대로 불필요한 backup을 계속 남길 수 있다.

    실무에서는 handoff 뒤 순서표를 운영 자산, restore 자산, 비용 cleanup 세 칸으로 나눠 두는 편이 가장 빠르다.

    promoted replica handoff 뒤 manual snapshot leftovers와 retained automated backup cleanup을 다른 큐로 나눈 순서표다.
    promoted replica handoff 뒤 manual snapshot leftovers와 retained automated backup cleanup을 다른 큐로 나눈 순서표다.

    이미 LatestRestorableTime 종료 글이 복구 창 closure를 다뤘다면, 이번 표는 restore asset inventory 자체를 더 좁혀 보는 후속편이다. deleting 이후 cleanup 글과도 직접 이어진다.

    5. 주의사항과 리스크

    첫 번째 리스크는 manual snapshot leftovers를 단순 leftover로 보고 삭제하는 것이다. 두 번째는 retained automated backup cleanup을 snapshot review가 끝날 때까지 같이 미루는 것이다. 세 번째는 automated backup을 manual snapshot으로 복사해 놓고 그 이유를 메모에 안 남기는 것이다.

    운영 문서에는 surviving instance id, manual snapshot keep reason, retained backup delete time, next review date를 같이 남기는 편이 좋다. 이 필드가 있어야 restore readiness와 비용 cleanup을 동시에 추적할 수 있다.

    특히 Snapshots 메뉴와 Retained 탭에서 각각 무엇을 클릭했고 어떤 값을 확인했는지, 어떤 삭제 명령을 실행했는지, 어떤 파일에 다음 review 일자를 저장했는지 남겨 두면 후속 조회가 빨라진다.

    • manual snapshot leftovers는 restore asset일 수 있다.
    • retained automated backup cleanup은 별도 비용 큐다.
    • handoff 완료와 restore asset 정리는 같은 종료 신호가 아니다.

    6. 결론

    promoted replica handoff 뒤 snapshot이 남아 있다는 이유만으로 모두 같은 cleanup 대상으로 보면 restore 자산과 비용 자산이 섞인다. manual snapshot leftovers와 retained automated backup cleanup을 다른 순서표로 나누면 surviving instance 운영, restore asset 유지, 비용 cleanup을 훨씬 짧게 설명할 수 있다.

    여기서 더 좁혀 manual snapshot keep reason은 이미 남겼는데 final snapshot 시각만 비어 있는 상황을 다루고 싶다면 후속편인 manual snapshot keep reason과 final snapshot closure 메모 줄 글을 이어서 보면 좋다. 이 글이 restore asset inventory와 retained cleanup을 나눴다면, 후속편은 final snapshot 시간축 한 줄을 추가하는 단계다.

    • surviving instance handoff를 먼저 닫는다.
    • manual snapshot leftovers는 keep reason과 함께 적는다.
    • retained automated backup cleanup은 별도 큐로 처리한다.

    7. 참고 링크

    1. https://docs.aws.amazon.com/cli/latest/reference/rds/delete-db-instance.html
    2. https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_WorkingWithAutomatedBackups.Retaining.html
    3. https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_CopySnapshot.html
    4. https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_WorkingWithAutomatedBackups-Deleting.html
    5. https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_DeleteSnapshot.html
Designed by Tistory.