ABOUT ME

-

Today
-
Yesterday
-
Total
-
  • [Claude][운영] tool use 루프에서 prompt caching usage가 예상과 다를 때 thinking blocks와 ephemeral_5m 토큰을 어떻게 같이 읽나
    기타개발지식/풀스택개발 2026. 7. 29. 20:19

    IT 리서치 노트

    [Claude][운영] tool use 루프에서 prompt caching usage가 예상과 다를 때 thinking blocks와 ephemeral_5m 토큰을 어떻게 같이 읽나

    Claude에서 tool use 루프를 돌릴 때 prompt caching usage가 예상과 다르게 보이는 순간이 있다. 새 사용자 입력은 많지 않은데 cache write가 크게 잡히거나, 분명 캐시를 읽은 것 같은데 context window가 줄지 않는 식이다. 2026년 7월 29일 기준 Anthropic 공식 문서를 다시 보면 tool results를 포함한 follow-up 요청에서 thinking blocks도 함께 캐시될 수 있고, input과 cache read와 cache creation 토큰은 모두 context window 계산에 들어가며, 조직·배치 단위 집계에는 ephemeral_5m_input_tokens 같은 TTL별 필드도 따로 노출된다. 이 글은 tool use 루프에서 prompt caching usage가 예상과 다를 때 어떤 순서로 읽어야 덜 헤매는지 정리한 것이다.

    1. 개요

    결론부터 말하면 tool use 루프에서 prompt caching usage는 사용자 메시지 길이만으로 읽으면 안 된다. follow-up 요청이 tool results를 포함하는지, thinking blocks가 함께 캐시되는지, 그 결과 write와 read와 uncached input이 어떻게 갈라지는지를 같이 봐야 한다.

    또 per-request usage와 조직 또는 배치 단위 집계는 보는 층위가 다르다. 응답 하나를 볼 때는 cache_creation_input_tokens와 cache_read_input_tokens를 보고, 기간 집계에서는 ephemeral_5m_input_tokens 같은 TTL별 필드를 같이 봐야 write 폭증 원인을 설명하기 쉬워진다.

    2. 어디서 실제로 막히는가

    현장에서 가장 흔한 오독은 세 가지다. 첫째, 새 사용자 입력이 짧으니 cache write가 작아야 한다고 생각한다. 둘째, cache read가 생기면 context window도 자동으로 줄 거라고 생각한다. 셋째, tool_choice나 thinking 설정이 바뀌어 cache block이 무효화됐는데도 같은 대화니까 hit가 유지될 것이라고 본다.

    Anthropic 문서는 thinking blocks가 tool results와 함께 캐시될 수 있다고 설명하고, context windows 문서는 input, cache read, cache creation 토큰이 모두 창 계산에 들어간다고 적는다. 또 tool use 문서는 tool_choice 변화가 cached message blocks를 무효화할 수 있다고 짚는다. 따라서 tool-use loop에서 usage가 이상해 보이는 것은 버그보다 구조 문제인 경우가 더 많다.

    • 증상: follow-up turn에서 cache_creation_input_tokens가 예상보다 크게 뛴다.
    • 실패: 사용자 질문 길이만 보고 write 크기를 해석한다.
    • 막힘: context window 증가를 캐시 실패로만 오해한다.
    • 누락: tool_choice, thinking, tool results 포함 여부를 메모하지 않는다.
    이상 신호 먼저 볼 것 해석 포인트
    follow-up에서 write가 급증 thinking blocks + tool results 포함 여부 직전 사고 흔적과 도구 결과가 한 번에 캐시됐는지 본다
    cache read가 있는데도 창이 빡빡함 context window 규칙 read와 creation 토큰도 창 계산에 들어간다
    비슷한 루프인데 hit가 흔들림 tool_choice, thinking 설정 메시지 내용 외 설정 변화가 breakpoint를 깨는지 본다

    3. 실무에서 적용하는 순서

    가장 짧은 운영 순서는 여섯 단계다. 1단계에서 per-request usage의 input_tokens, cache_creation_input_tokens, cache_read_input_tokens를 그대로 저장한다. 2단계에서 그 turn이 tool result follow-up인지 표시한다. 3단계에서 thinking 설정과 tool_choice를 같이 기록한다. 4단계에서 context window 해석은 세 필드를 합쳐 본다. 5단계에서 배치나 조직 집계에서는 ephemeral_5m_input_tokens와 같은 TTL별 필드로 다시 합산한다. 6단계에서 이상 신호가 나면 사용자 입력보다 thinking/tool result 포함 여부를 먼저 의심한다.

    1. 응답 usage의 세 필드를 turn별로 저장한다.
    2. 해당 turn이 tool result follow-up인지 표시한다.
    3. thinking과 tool_choice를 함께 메모한다.
    4. context window는 read와 creation까지 포함해 계산한다.
    5. 집계 분석에서는 ephemeral_5m 또는 1h 필드를 같이 본다.
    6. 이상 신호가 나면 텍스트 길이보다 loop 구조를 먼저 본다.

    특히 tool-use loop는 '답변 1회'가 아니라 '사고 + 도구 결과 + follow-up'의 묶음으로 봐야 한다. thinking blocks는 tool results를 포함한 후속 요청에서 캐시될 수 있으므로, write가 큰 이유가 직전 assistant 내부 맥락까지 묶여 들어갔기 때문일 수 있다. 이때 per-request usage와 org usage를 분리해 적어 두면 월말 비용 설명도 쉬워진다.

    turn = "tool_result_follow_up"
    tool_choice = "auto"
    thinking = "unchanged"
    save = [
      input_tokens,
      cache_creation_input_tokens,
      cache_read_input_tokens,
      ephemeral_5m_input_tokens
    ]
    
    if tool_choice changed:
      expect cached message blocks to miss
    if follow_up includes tool results:
      expect prior thinking blocks to join cache write

    관련 흐름으로는 Claude TTL 분기 글이 write-read 손익을 보는 기준을 다루고, 이번 글은 tool loop 안에서 왜 그 숫자가 튀는지를 읽는 쪽에 가깝다.

    4. 공식 문서와 예시 화면으로 확인하기

    첫 자료는 Anthropic prompt caching 문서의 usage 필드 설명이다. 여기서는 cache_creation_input_tokens, cache_read_input_tokens, input_tokens를 어떻게 해석해야 하는지 직접 설명한다.

    Anthropic 문서는 cache_creation_input_tokens, cache_read_input_tokens, input_tokens를 함께 보라고 안내한다.
    Anthropic 문서는 cache_creation_input_tokens, cache_read_input_tokens, input_tokens를 함께 보라고 안내한다.

    즉 per-request usage만 봐도 write와 read와 uncached input을 분리해서 볼 수 있다. tool use 루프에서 예상 밖 write가 잡혔다면 먼저 이 셋이 어느 구간에서 늘었는지 나눠 보는 편이 맞다.

    두 번째 자료는 extended thinking models 문서다. Anthropic은 tool-use loop에서 follow-up 요청이 tool results를 포함할 때 이전 대화와 thinking blocks가 함께 캐시될 수 있다고 설명한다.

    Anthropic 문서는 tool results를 포함한 follow-up 요청에서 thinking blocks도 함께 캐시될 수 있다고 적는다.
    Anthropic 문서는 tool results를 포함한 follow-up 요청에서 thinking blocks도 함께 캐시될 수 있다고 적는다.

    이 문장 하나 때문에 usage 해석이 달라진다. write가 늘어난 원인이 새 사용자 질문이 아니라 직전 thinking과 tool results를 포함한 follow-up 구조일 수 있기 때문이다.

    세 번째 자료는 context windows 문서다. 여기서는 prompt caching을 쓰면 input_tokens, cache_read_input_tokens, cache_creation_input_tokens 세 필드가 모두 context window 계산에 들어간다고 설명한다.

    Prompt caching 사용 시 input, cache read, cache creation 토큰이 모두 context window에 반영된다.
    Prompt caching 사용 시 input, cache read, cache creation 토큰이 모두 context window에 반영된다.

    따라서 '캐시니까 공짜'처럼 보면 안 된다. read와 write를 나눠서 비용을 봐야 할 뿐 아니라, context window와 max output 계획에도 함께 반영해야 한다.

    네 번째 자료는 tool use 구현 문서의 주의사항이다. Anthropic은 prompt caching을 쓸 때 tool_choice 변화가 cached message blocks를 무효화할 수 있다고 적고 있다.

    tool_choice 변화는 cached message blocks를 무효화할 수 있다.
    tool_choice 변화는 cached message blocks를 무효화할 수 있다.

    실무에서 같은 대화처럼 보이는데 hit가 흔들리는 이유가 여기서 자주 나온다. 메시지 내용이 같아도 tool_choice를 auto, any, tool로 바꾸면 message-level breakpoint가 miss로 돌아설 수 있다.

    다섯 번째 자료는 batch results 문서의 usage 필드 예시다. 여기서는 조직·배치 단위 집계에서 ephemeral_5m_input_tokens와 관련 필드가 따로 노출된다는 점을 확인할 수 있다.

    Batch results 예시는 ephemeral_5m_input_tokens 같은 TTL별 집계 필드가 따로 존재함을 보여 준다.
    Batch results 예시는 ephemeral_5m_input_tokens 같은 TTL별 집계 필드가 따로 존재함을 보여 준다.

    즉 per-request usage와 조직·배치 집계 usage는 보는 층위가 다르다. 개별 응답에서 cache_creation과 cache_read를 보고, 월간 또는 배치 추적에서는 ephemeral_5m 또는 1h 집계를 같이 보면 원인이 덜 엉킨다.

    마지막 자료는 tool-use follow-up 응답에서 usage를 읽는 최소 예시다. thinking과 tool results가 이어진 뒤 어떤 필드를 먼저 적어 두면 다음 분석이 쉬운지 한 줄로 정리했다.

    Tool-use follow-up 응답에서 cache write, cache read, uncached input을 같이 남기는 예시다.
    Tool-use follow-up 응답에서 cache write, cache read, uncached input을 같이 남기는 예시다.

    핵심은 usage 총합만 보지 않는 것이다. 어떤 follow-up에서 write가 생겼는지, 어떤 turn에서 read가 늘었는지, 어떤 turn이 context window를 키웠는지를 분리해서 남겨야 다음 조정이 된다.

    5. 주의사항과 리스크

    첫 번째 리스크는 per-request usage만 보고 조직 집계를 무시하는 것이다. 두 번째 리스크는 cache read가 있으니 context window도 자동으로 가벼워진다고 착각하는 것이다. 세 번째 리스크는 tool_choice나 thinking 설정 변화로 인한 miss를 캐시 기능 자체의 불안정성으로 오해하는 것이다.

    운영 로그에는 최소한 turn 종류, tool results 포함 여부, thinking 설정, tool_choice, 세 usage 필드, TTL별 집계 필드를 같이 남겨 두는 편이 좋다. 그래야 비용 급증이 실제 사용자 트래픽 때문인지, 한두 번의 tool-use 루프 구조 때문인지 짧게 가를 수 있다.

    • Tool-use follow-up은 사용자 입력보다 더 큰 cache write를 만들 수 있다.
    • Cache read와 cache creation 토큰도 context window에 포함된다.
    • 설정 변화는 텍스트 변화가 없어도 cached blocks를 무효화할 수 있다.

    6. 결론

    Tool use 루프에서 prompt caching usage가 예상과 다를 때는 메시지 길이보다 loop 구조를 먼저 봐야 한다. thinking blocks, tool results, tool_choice, context window 규칙이 함께 작동하기 때문이다.

    정리하면 응답 usage의 세 필드와 조직 집계의 ephemeral_5m 계열 필드를 같이 보고, follow-up turn과 설정 변화를 함께 적어 두면 왜 write가 커졌는지와 왜 hit가 흔들렸는지를 훨씬 짧게 설명할 수 있다.

    이 글이 tool-use follow-up과 usage 해석을 다뤘다면, 바로 다음 단계인 thinking config와 tool_choice가 함께 바뀔 때 cache miss 원인을 로그 순서로 나누는 글은 같은 Claude 가지 안에서 settings diff를 어떻게 분리할지까지 이어 준다.

    7. 참고 링크

    1. https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching
    2. https://docs.anthropic.com/en/docs/about-claude/models/extended-thinking-models
    3. https://docs.anthropic.com/en/docs/build-with-claude/context-windows
    4. https://docs.anthropic.com/en/docs/agents-and-tools/tool-use/implement-tool-use
    5. https://docs.anthropic.com/en/api/retrieving-message-batch-results
Designed by Tistory.