본문 바로가기
Agentic AI/AI_AGENT

AI Agent(2/3): 다단계 도구 호출(Multi-step Tool Calling) 및 자율적 루프(Autonomous Re-act Loop)

by 아톨 2026. 10. 3.

첫 번째 코드와 비교했을 때, 이번 코드는 질문 하나를 해결하기 위해 여러 도구를 연쇄적으로 판단하여 호출하는 에이전트 구조로 진화했습니다. 다단계 도구 호출(Multi-step Tool Calling) 및 자율적 루프(Autonomous Re-act Loop)를 구현한 대단히 완성도 높은 AI 에이전트 코드입니다.

 

1. 이 코드가 작동하는 전체 시나리오 (Chaining Process)

사용자가 "현재 서울의 평균 기온을 계산해 줘."라는 질문을 보냈을 때 내부 동작 순서입니다.

                  [사용자 질문] "서울의 평균 기온 계산해 줘"
                                 │
                                 ▼
                     ┌──────────────────────┐
                     │ 1차 LLM 판단 (while)  │
                     └──────────────────────┘
                                 │
         ┌───────────────────────┴───────────────────────┐
         ▼                                               ▼
[도구 호출 1] get_weather("서울")               [시스템 제한] "암산 금지"
         │                                               │
         ▼                                               │
[조회 결과] min_temp: 12, max_temp: 22                   │
         │                                               │
         └───────────────────────┬───────────────────────┘
                                 ▼
                     ┌──────────────────────┐
                     │ 2차 LLM 판단 (while)  │ ──> "조회한 12와 22로 평균 코드를 작성하자"
                     └──────────────────────┘
                                 │
                                 ▼
[도구 호출 2] code_interpreter(code="print((12 + 22) / 2)")
                                 │
                                 ▼
[실행 결과] "17.0"
                                 │
                                 ▼
                     ┌──────────────────────┐
                     │ 3차 LLM 판단 (while)  │
                     └──────────────────────┘
                                 │
                                 ▼ (더 이상 필요한 도구 호출 없음 -> tool_calls가 빈 리스트)
                  [최종 답변] "서울의 평균 기온은 17°C입니다."

2. 코드의 핵심 요소 상세 설명

① 자율 루프 제어 (while True: 구문)

이 코드의 가장 강력한 특징입니다.

  • 첫 번째 코드는 1번만 도구를 실행하고 마쳤지만, 이 코드는 while True:를 통해 LLM이 스스로 "더 도구를 써야 할까? 아니면 답변할 수 있을까?"를 반복해서 평가합니다.
  • tool_calls가 존재하면: 도구를 실행하고 그 결과를 messages 배열에 누적(Append)한 뒤 while 루프의 다음 바퀴로 돌아갑니다.
  • tool_calls가 없으면: 충분한 정보를 얻었다고 판단하여 최종 답변을 출력하고 break로 루프를 탈출합니다.

② 샌드박스형 코드 실행기 (code_interpreter)

파이썬의 동적 실행 능력을 활용하여 안전하게 코드를 수행하도록 설계되었습니다.

  • io.StringIO() & redirect_stdout: 파이썬의 print() 출력값을 화면으로 보내지 않고, 메모리 버퍼로 빼돌려(캡처) 문자열로 읽어옵니다.
  • exec(code, {}): 문자열 파이썬 코드를 실행합니다. 두 번째 인자로 {}를 전달함으로써 외부 전역 변수에 접근하거나 오염시키는 것을 차단합니다.
  • 보안 주석 # noqa:S102: Bandit 등 정적 분석 도구가 exec() 사용에 따른 보안 경고를 남기는 것을 방지합니다.
  • 에러 수집: 실행 중 산술 에러(예: ZeroDivisionError)나 문법 에러가 터지더라도 프로그램 전체가 뻗지 않고 [실행오류] 메시지를 만들어 LLM에게 반환합니다. 그럼 LLM은 이 에러를 읽고 코드를 수정해서 재시도하게 됩니다.

③ 엄격한 시스템 프롬프트 (Strict Guardrail)

"너는 절대로 수학적 연산을 직접 수행(암산/추론)해서는 안 된다."
"평균/합/차/나눗셈 등 모든 계산은 반드시 code_interpreter 도구로만 수행해야 한다."
  • LLM은 언어 모델 특성상 간단한 산술 연산에서도 환각(Hallucination)이 발생할 위험이 있습니다.
  • 시스템 프롬프트로 "직접 계산 금지" 제약을 강하게 걸고, 반드시 Python 코드 실행기를 거치도록 강제하여 100% 정확한 수치 계산을 보장합니다.

④ 동적 도구 라우팅 (avaiable_functions 및 분기문)

if function_name == "code_interpreter":
    code = function_args.get("code", "")
    function_response = function_to_call(code)
else:
    function_response = function_to_call(location=function_args.get("location"))
  • LLM이 요청한 함수명(function_name)에 따라 인자(code vs location)를 구분하여 전달하도록 깔끔한 파이프라인 처리가 되어 있습니다.

3. 코드 설계의 우수한 점

  1. 에이전트 연쇄 능력 (Tool Chaining): 단순한 단발성 Q&A가 아니라 "날씨 데이터 수집 -> 수치 가공 -> 수학 계산 -> 최종 응답"으로 이어지는 복합적 문제를 하나의 루프 구조 안에서 스스로 추론하고 해결합니다.
  2. 안전장치 설계 (Exception Handling): API 응답 오류나 Python 실행 에러가 발생해도 전체 스크립트가 멈추지 않고, 에러 문자열을 에이전트의 컨텍스트(Context)로 다시 되돌려주도록 구성되어 있습니다.
  3. 결과 가시성: 중간 판단 과정(tool_calls), 도구 인자, 도구 결과가 터미널에 명확히 출력되므로 에이전트의 사고 과정을 시각적으로 추적(Tracing)할 수 있습니다.

💡 소소한 팁 & 개선 가이드

  1. 변수 오타 확인:
    • 코드 중반에 avaiable_functions (l 오타) 로 작성된 변수명이 있습니다. available_functions로 맞춰주시면 더 직관적입니다.
  2. code_interpreter 내 모듈 import 지원:
    • 현재 exec(code, {}) 방식은 기본 내장 함수만 사용 가능합니다.
    • 만약 LLM이 math, datetime, numpy 등의 모듈을 쓰려고 코드를 생성하면 NameError: name 'math' is not defined가 발생할 수 있습니다.
    • 필요하다면 exec(code, {"__builtins__": __builtins__, "math": math})처럼 유용한 라이브러리를 전역 공간에 미리 넘겨주면 활용도가 훨씬 높아집니다.

한 줄 요약: OpenAI 최신 Responses API의 Multi-step Tool Calling과 Code Interpreter 패턴을 파이썬 코드로 완성도 높게 잘 구현해 낸 훌륭한 예제입니다!

 

 
    import os
    from dotenv import load_dotenv
    from openai import OpenAI
    import json
    import io
    import contextlib

    load_dotenv(override=True)

    client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
    model = os.getenv("OPENAI_MODEL", "gpt-5.4-mini")

    MOCK_WEATHER_DATA = {
        "서울":{
            "condition":"맑음",
            "temperature_c":22,
            "min_temp_c":12,
            "max_temp_c":22,
            "humidity":45,
            "summary":"서울은 맑고 쾌청한 가을 하늘을 보입니다.",
        },
        "대전":{
            "condition":"비",
            "temperature_c":21,
            "min_temp_c":17,
            "max_temp_c":24,
            "humidity":70,
            "summary":"부산은 비가 내리고 습도가 높습니다.",        
        },
    }

    def get_weather(location):
        weather = MOCK_WEATHER_DATA.get(location)
        if weather is None:
            return json.dumps({"error":"날씨 정보를 찾을 수 없습니다."}, ensure_ascii=False)    
       
        return json.dumps(
            {
                "location":location,
                "condition":weather["condition"],
                "min_temp":weather["min_temp_c"],
                "max_temp":weather["max_temp_c"],
            },
            ensure_ascii=False,
        )
    def code_interpreter(code:str)->str:
        stdout_capture = io.StringIO()
        try:
            with contextlib.redirect_stdout(stdout_capture):
                exec(code, {}) #noqa:S102  
            output=stdout_capture.getvalue()
            return output.strip() if output.strip() else "(출력없음)"
        except Exception as e:
            return f"[실행오류] {type(e).__name__}:{e}"
       
    avaiable_functions = {
        "get_weather":get_weather,
        "code_interpreter":code_interpreter,
    }

    tools = [
        {
            "type":"function",
            "name":"get_weather",
            "description":"특정 지역의 현재 최고 및 최저 기온을 조회한다.",
            "parameters":{
                "type":"object",
                "properties":{
                    "location":{"type":"string","description":"도시 이름(예: 서울, Newyork...)"},
                },
                "required":["location"]
            },        
        },
        {
            "type":"function",
            "name":"code_interpreter",
            "description":(
                "Python 코드를 실행하여 계산 결과를 반환한다. "
                "수학 계산, 통계, 데이터 분석 등 모든 연산에 사용한다. "
                "결과는 반드시 print()로 출력해야 한다."
            ),
            "parameters":{
                "type":"object",
                "properties":{
                    "code":{"type":"string","description":"실행할 Python 코드. print()로 결과를 출력해야 한다."},
                },
                "required":["code"]
            },        
        },
    ]

    def call_llm(user_prompt):
        messages = [
            {
                "role":"system",
                "content":(
                    "너는 절대로 수학적 연산을 직접 수행(암산/추론)해서는 안 된다. "
                    "평균/합/차/나눗셈 등 모든 계산은 반드시 code_interpreter 도구로만 수행해야 한다. "
                    "만약 code_interpreter 도구를 사용할 수 없는 상황이라면, "
                    "계산 결과를 절대 추측하거나 직접 계산하지 말고 "
                    "'code_interpreter 도구가 없어 계산할 수 없습니다'라고 명시적으로 답해야 한다."            
                ),
            },
            {
                "role":"user",
                "content":user_prompt,
            }
        ]

        try:
            while True:
                print("[시스템] 요청을 보냅니다...")
                response = client.responses.create(
                    model=model,
                    input=messages,
                    tools=tools,
                )
                tool_calls = [item for item in response.output if item.type == "function_call"]
                if not tool_calls:
                    print("="*100)
                    print(f"[최종 답변] {response.output_text}")
                    print("="*100)
                    break
               
                print("="*100)
                print("response: ",response)
                print("="*100)    
                print("tool_calls: ",tool_calls)
                print("="*100)  
               
                if tool_calls:
                    messages.append(
                        {
                            "role":"assistant",
                            "content":response.output_text or "도구 호출 요청",
                        }
                    )
                    for tool_call in tool_calls:
                        function_name = tool_call.name
                        function_to_call = avaiable_functions[function_name]
                        function_args = json.loads(tool_call.arguments)
                       
                        if function_name == "code_interpreter":
                            code = function_args.get("code","")
                            print(f"[도구실행] {function_name}({function_args})호출중...")
                            function_response = function_to_call(code) if function_to_call else json.dumps(
                                {"error":f"지원하지 않는 도구입니다: {function_name}"}, ensure_ascii=False,
                            )
                        else:
                            print(f"[도구실행] {function_name}({function_args})호출중...")    
                            if function_to_call is None:
                                function_response = json.dumps(
                                    {"error":f"지원하지 않는 도구입니다: {function_name}"},
                                    ensure_ascii=False,
                                )
                            else:
                                function_response = function_to_call(
                                    location=function_args.get("location")
                                )
                        print(f"[실행 결과] {function_response}")
   
                        messages.append(
                            {
                                "role":"user",
                                "content":f"[Tool output for {function_name}]: {function_response}"
                            }
                        )
        except Exception as e:
            print(f"[에러발생]: {e}")
           
       
    def main():
        user_ask = "현재 서울의 평균 기온을 계산해 줘."
        print("="*100)
        print("질문: ", user_ask)
        print("="*100)    
        answer_5 = call_llm(user_ask)

    if __name__ =="__main__":
        main()
반응형