ABOUT ME

-

Today
-
Yesterday
-
Total
-
  • [Claude][운영] Claude Sonnet 5 fallback second turn은 read hit인데 incomplete가 날 때 stop reason과 max_tokens를 어떤 순서로 나누나
    기타개발지식/AI 2026. 9. 1. 20:14

    IT 리서치 노트

    [Claude][운영] Claude Sonnet 5 fallback second turn은 read hit인데 incomplete가 날 때 stop reason과 max_tokens를 어떤 순서로 나누나

    Claude Sonnet 5 fallback second turn이 read hit인데도 incomplete가 날 때 많은 팀이 cache 자체가 흔들렸다고 적는다. 하지만 2026년 9월 1일 KST 기준 Anthropic 공식 문서를 다시 보면 cached prompt prefixes는 context window를 계속 차지하고, Sonnet 5는 adaptive thinking이 기본이며, thinking과 응답은 같은 max_tokens 한도를 공유할 수 있다. 이 글은 fallback second turn이 read hit인데 incomplete가 날 때 stop reason과 max_tokens를 어떤 순서로 나눠야 조치가 빨라지는지 정리한다.

    1. 개요

    결론부터 말하면 read hit 뒤 incomplete가 나면 먼저 stop reason을 나누고 확인한 뒤, 그다음 max_tokens와 context window 압박을 따로 봐야 한다. cache read hit는 비용 신호일 뿐이고, incomplete 원인을 설명해 주는 값은 아니다.

    Claude Sonnet 5 fallback 운영에서는 cache_read_hit, stop_reason, max_tokens, effort, output 길이를 같은 trace에 두고 저장하고 비교해야 한다. 이미 read hit와 effort triage 글이 느림의 상위 분기를 다뤘다면, 이번 글은 incomplete가 섞였을 때의 더 좁은 분기다.

    2. 어디서 실제로 막히는가

    현장에서 먼저 꼬이는 지점은 세 가지다. 첫째, read hit가 보이면 prompt cache는 정상이라고만 적고 stop reason을 조회하지 않는다. 둘째, output이 잘렸는데도 max_tokens 부족인지 context window 포화인지 구분하지 않고 같은 메모에 적는다. 셋째, Sonnet 5의 adaptive thinking이 기본이라는 사실을 빼먹고 예전 budget 감각으로 fallback trace를 해석하고 비교한다.

    Anthropic 문서는 cached prompt prefixes가 context window를 계속 차지한다고 설명하고, stop reasons 문서는 max_tokens와 model_context_window_exceeded를 구분한다. 또 Sonnet 5 prompting 문서는 thinking이 큰 비중을 차지하면 응답이 거의 thinking으로 채워진 뒤 잘릴 수 있다고 적는다. 즉 read hit 뒤 incomplete는 cache 안정성보다 budget 배분과 window 압박을 먼저 확인하고 기록해야 한다.

    문제는 incident note가 흔히 cache hit인데도 느리고 답이 짧다 한 줄로 끝난다는 점이다. 이렇게 적으면 max_tokens를 올려야 하는지, effort를 낮춰야 하는지, 오래된 tool 결과를 줄여야 하는지 갈리지 않는다. stop reason을 먼저 저장하고 자르지 않으면 조치가 계속 빗나간다.

    • 증상: read hit인데 답이 중간에서 끊기거나 incomplete처럼 보인다.
    • 실패: stop reason을 남기지 않고 cache 상태만 기록하고 끝낸다.
    • 막힘: max_tokens 부족과 context window 포화를 같은 사건처럼 적고 비교한다.
    • 누락: Sonnet 5 adaptive thinking과 effort 값을 trace에 저장하지 않는다.
    겉으로 보이는 현상 먼저 확인할 값 판단 기준
    답이 짧고 중간에서 끊긴다 stop_reason, max_tokens max_tokens 상한이 먼저 닿았는지 확인한다
    cache hit인데도 다음 turn이 안 이어진다 context window, cached prefix 길이 cached prefix가 window를 이미 채우는지 확인한다
    비용과 응답 길이가 같이 흔들린다 effort, usage, output 길이 thinking 비중과 budget headroom을 함께 확인하고 비교한다

    3. 실무에서 적용하는 순서

    가장 실용적인 triage 순서는 다섯 단계다. 먼저 read hit 여부와 stop reason을 한 줄에 기록한다. 다음으로 stop reason이 max_tokens인지 model_context_window_exceeded인지 분리해 저장한다. 세 번째로 max_tokens, effort, observed output 길이를 함께 조회하고 비교한다. 네 번째로 cached prefix와 tool 결과가 context window를 얼마나 차지했는지 확인한다. 마지막으로 이 값을 다음 turn trace와 비교하고 다음 조치를 선택한다.

    1. cache read hit와 stop reason을 같은 trace에 기록하고 저장한다.
    2. max_tokens와 model_context_window_exceeded를 먼저 갈라 적고 확인한다.
    3. effort와 observed output 길이를 함께 조회하고 비교한다.
    4. cached prefix와 tool tail 길이를 재검토하고 필요한 부분을 줄인다.
    5. 다음 turn 조치를 headroom 확대와 tail 축소 중 하나로 고르고 실행한다.

    이렇게 보면 조치가 단순해진다. max_tokens면 output budget 문제이므로 headroom을 늘리거나 effort를 낮추고 다시 실행하면 된다. model_context_window_exceeded면 prompt cache hit 여부와 별개로 긴 tool 결과나 오래된 turn을 덜어야 한다. read hit는 여전히 유용한 신호지만, incomplete의 원인 라벨로 쓰면 안 된다.

    cache_read_hit=true
    stop_reason=model_context_window_exceeded|max_tokens
    max_tokens=2400
    effort=high
    cached_prefix_kept=true
    next_action=raise_budget|trim_tail

    4. 공식 문서와 예시 화면으로 확인하기

    첫 자료는 context windows 문서다. cached prefix는 돈을 줄여 줄 수 있어도 context window에서 사라지지는 않는다고 문서가 직접 적고 있다.

    cache read hit가 나도 cached prompt prefix는 context window를 계속 차지한다.
    cache read hit가 나도 cached prompt prefix는 context window를 계속 차지한다.

    즉 second turn이 read hit였다는 사실만으로 incomplete 원인이 사라지지 않는다. cache 비용과 context budget은 다른 축이므로, read hit 뒤 incomplete가 나면 stop reason을 먼저 확인하고 window 압박을 따로 비교해야 한다.

    두 번째 자료는 stop reasons 문서다. Claude는 context window 한도에 닿았을 때와 max tokens 상한에 닿았을 때를 다른 stop reason으로 반환한다.

    read hit 뒤 incomplete가 나면 stop reason이 max_tokens인지 model_context_window_exceeded인지 먼저 갈라야 한다.
    read hit 뒤 incomplete가 나면 stop reason이 max_tokens인지 model_context_window_exceeded인지 먼저 갈라야 한다.

    이 분기가 없으면 output budget 부족을 cache instability처럼 적게 된다. 같은 prompt shape에서도 조치가 달라지므로 stop reason을 먼저 기록하고 확인하는 절차가 필요하다.

    세 번째 자료는 Sonnet 5 prompting 문서다. thinking이 켜진 상태에서 headroom이 부족하면 생각 블록이 먼저 예산을 먹고 잘린 답변과 max_tokens stop reason이 나올 수 있다고 설명한다.

    Sonnet 5에서는 thinking과 응답이 같은 max_tokens 한도를 공유하므로 headroom 부족이 incomplete를 만들 수 있다.
    Sonnet 5에서는 thinking과 응답이 같은 max_tokens 한도를 공유하므로 headroom 부족이 incomplete를 만들 수 있다.

    특히 fallback second turn은 prompt cache가 살아 있어도 adaptive thinking과 output budget이 그대로 남아 있기 때문에, read hit와 max_tokens는 서로 독립적으로 기록하고 비교해야 한다. read hit인데도 incomplete가 나는 이유가 여기서 갈린다.

    실무에서는 read hit, output budget, stop reason이 한 줄 incident note에 뭉쳐 있다. 아래 표는 무엇부터 나눌지 정리한 예시다.

    read hit 뒤 incomplete가 날 때는 stop reason과 max_tokens를 먼저 분리해야 한다.
    read hit 뒤 incomplete가 날 때는 stop reason과 max_tokens를 먼저 분리해야 한다.

    이미 usage와 tail diff 글이 숫자 비교를 다뤘다면 이번 글은 incomplete 분기 순서다. 또 read hit와 effort triage 글과 함께 보면 fallback second-turn branch가 자연스럽게 이어진다.

    마지막 자료는 trace log 예시다. 핵심은 read hit 플래그와 stop reason과 max_tokens를 같은 record에 두는 것이다.

    fallback trace는 cache read hit와 stop reason과 max_tokens를 같은 record에 남겨야 한다.
    fallback trace는 cache read hit와 stop reason과 max_tokens를 같은 record에 남겨야 한다.

    이 정도 구조면 post 430의 fallback checklist, post 443의 tools mismatch, post 449의 effort triage가 하나의 incomplete 운영표로 연결된다.

    5. 주의사항과 리스크

    첫 번째 리스크는 read hit가 나왔다는 이유만으로 cache branch에서만 사고를 복기하는 것이다. 두 번째 리스크는 adaptive thinking이 기본이라는 점을 잊고 예전 max_tokens 상한을 그대로 쓰고 비교하는 것이다. 세 번째 리스크는 context window 포화를 output budget 부족처럼 처리해 잘못된 조정을 반복하는 것이다.

    운영 전에는 최소한 cache_read_hit, stop_reason, max_tokens, effort, observed output length, tail size를 같은 incident note에 두고 저장하고 비교하는 편이 좋다. 그래야 next-turn cost spike나 regenerated cache branch와도 자연스럽게 연결된다.

    • read hit는 비용 신호이지 incomplete 원인 라벨이 아니다.
    • max_tokens와 context window 포화를 먼저 갈라 적는다.
    • adaptive thinking headroom을 trace 없이 추정하지 않는다.

    6. 결론

    Claude Sonnet 5 fallback second turn이 read hit인데도 incomplete가 날 수 있다. 이때는 cache 자체보다 stop reason과 max_tokens와 context window 압박을 먼저 나눠야 한다. 그 분기만 잡아도 조치가 훨씬 빨라진다.

    • stop reason을 먼저 자른다.
    • max_tokens와 context window 포화를 다른 조치로 본다.
    • read hit와 effort와 output 길이를 같은 trace에 남긴다.

    7. 참고 링크

    1. https://platform.claude.com/docs/en/build-with-claude/prompt-caching
    2. https://platform.claude.com/docs/en/build-with-claude/handling-stop-reasons
    3. https://platform.claude.com/docs/en/build-with-claude/context-windows
    4. https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5
    5. https://platform.claude.com/docs/en/models/sonnet-5/whats-new-sonnet-5
Designed by Tistory.