Agentic AI/AI_AGENT
AI Agent(2/3): 다단계 도구 호출(Multi-step Tool Calling) 및 자율적 루프(Autonomous Re-act Loop)
아톨
2026. 10. 3. 14:50
첫 번째 코드와 비교했을 때, 이번 코드는 질문 하나를 해결하기 위해 여러 도구를 연쇄적으로 판단하여 호출하는 에이전트 구조로 진화했습니다. 다단계 도구 호출(Multi-step Tool Calling) 및 자율적 루프(Autonomous Re-act Loop)를 구현한 대단히 완성도 높은 AI 에이전트 코드입니다.
1. 이 코드가 작동하는 전체 시나리오 (Chaining Process)
사용자가 "현재 서울의 평균 기온을 계산해 줘."라는 질문을 보냈을 때 내부 동작 순서입니다.
[사용자 질문] "서울의 평균 기온 계산해 줘"
│
▼
┌──────────────────────┐
│ 1차 LLM 판단 (while) │
└──────────────────────┘
│
┌───────────────────────┴───────────────────────┐
▼ ▼
[도구 호출 1] get_weather("서울") [시스템 제한] "암산 금지"
│ │
▼ │
[조회 결과] min_temp: 12, max_temp: 22 │
│ │
└───────────────────────┬───────────────────────┘
▼
┌──────────────────────┐
│ 2차 LLM 판단 (while) │ ──> "조회한 12와 22로 평균 코드를 작성하자"
└──────────────────────┘
│
▼
[도구 호출 2] code_interpreter(code="print((12 + 22) / 2)")
│
▼
[실행 결과] "17.0"
│
▼
┌──────────────────────┐
│ 3차 LLM 판단 (while) │
└──────────────────────┘
│
▼ (더 이상 필요한 도구 호출 없음 -> tool_calls가 빈 리스트)
[최종 답변] "서울의 평균 기온은 17°C입니다."
2. 코드의 핵심 요소 상세 설명
① 자율 루프 제어 (while True: 구문)
이 코드의 가장 강력한 특징입니다.
- 첫 번째 코드는 1번만 도구를 실행하고 마쳤지만, 이 코드는 while True:를 통해 LLM이 스스로 "더 도구를 써야 할까? 아니면 답변할 수 있을까?"를 반복해서 평가합니다.
- tool_calls가 존재하면: 도구를 실행하고 그 결과를 messages 배열에 누적(Append)한 뒤 while 루프의 다음 바퀴로 돌아갑니다.
- tool_calls가 없으면: 충분한 정보를 얻었다고 판단하여 최종 답변을 출력하고 break로 루프를 탈출합니다.
② 샌드박스형 코드 실행기 (code_interpreter)
파이썬의 동적 실행 능력을 활용하여 안전하게 코드를 수행하도록 설계되었습니다.
- io.StringIO() & redirect_stdout: 파이썬의 print() 출력값을 화면으로 보내지 않고, 메모리 버퍼로 빼돌려(캡처) 문자열로 읽어옵니다.
- exec(code, {}): 문자열 파이썬 코드를 실행합니다. 두 번째 인자로 {}를 전달함으로써 외부 전역 변수에 접근하거나 오염시키는 것을 차단합니다.
- 보안 주석 # noqa:S102: Bandit 등 정적 분석 도구가 exec() 사용에 따른 보안 경고를 남기는 것을 방지합니다.
- 에러 수집: 실행 중 산술 에러(예: ZeroDivisionError)나 문법 에러가 터지더라도 프로그램 전체가 뻗지 않고 [실행오류] 메시지를 만들어 LLM에게 반환합니다. 그럼 LLM은 이 에러를 읽고 코드를 수정해서 재시도하게 됩니다.
③ 엄격한 시스템 프롬프트 (Strict Guardrail)
"너는 절대로 수학적 연산을 직접 수행(암산/추론)해서는 안 된다."
"평균/합/차/나눗셈 등 모든 계산은 반드시 code_interpreter 도구로만 수행해야 한다."
- LLM은 언어 모델 특성상 간단한 산술 연산에서도 환각(Hallucination)이 발생할 위험이 있습니다.
- 시스템 프롬프트로 "직접 계산 금지" 제약을 강하게 걸고, 반드시 Python 코드 실행기를 거치도록 강제하여 100% 정확한 수치 계산을 보장합니다.
④ 동적 도구 라우팅 (avaiable_functions 및 분기문)
if function_name == "code_interpreter":
code = function_args.get("code", "")
function_response = function_to_call(code)
else:
function_response = function_to_call(location=function_args.get("location"))
- LLM이 요청한 함수명(function_name)에 따라 인자(code vs location)를 구분하여 전달하도록 깔끔한 파이프라인 처리가 되어 있습니다.

3. 코드 설계의 우수한 점
- 에이전트 연쇄 능력 (Tool Chaining): 단순한 단발성 Q&A가 아니라 "날씨 데이터 수집 -> 수치 가공 -> 수학 계산 -> 최종 응답"으로 이어지는 복합적 문제를 하나의 루프 구조 안에서 스스로 추론하고 해결합니다.
- 안전장치 설계 (Exception Handling): API 응답 오류나 Python 실행 에러가 발생해도 전체 스크립트가 멈추지 않고, 에러 문자열을 에이전트의 컨텍스트(Context)로 다시 되돌려주도록 구성되어 있습니다.
- 결과 가시성: 중간 판단 과정(tool_calls), 도구 인자, 도구 결과가 터미널에 명확히 출력되므로 에이전트의 사고 과정을 시각적으로 추적(Tracing)할 수 있습니다.
💡 소소한 팁 & 개선 가이드
- 변수 오타 확인:
- 코드 중반에 avaiable_functions (l 오타) 로 작성된 변수명이 있습니다. available_functions로 맞춰주시면 더 직관적입니다.
- code_interpreter 내 모듈 import 지원:
- 현재 exec(code, {}) 방식은 기본 내장 함수만 사용 가능합니다.
- 만약 LLM이 math, datetime, numpy 등의 모듈을 쓰려고 코드를 생성하면 NameError: name 'math' is not defined가 발생할 수 있습니다.
- 필요하다면 exec(code, {"__builtins__": __builtins__, "math": math})처럼 유용한 라이브러리를 전역 공간에 미리 넘겨주면 활용도가 훨씬 높아집니다.
한 줄 요약: OpenAI 최신 Responses API의 Multi-step Tool Calling과 Code Interpreter 패턴을 파이썬 코드로 완성도 높게 잘 구현해 낸 훌륭한 예제입니다!
import os
from dotenv import load_dotenv
from openai import OpenAI
import json
import io
import contextlib
load_dotenv(override=True)
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
model = os.getenv("OPENAI_MODEL", "gpt-5.4-mini")
MOCK_WEATHER_DATA = {
"서울":{
"condition":"맑음",
"temperature_c":22,
"min_temp_c":12,
"max_temp_c":22,
"humidity":45,
"summary":"서울은 맑고 쾌청한 가을 하늘을 보입니다.",
},
"대전":{
"condition":"비",
"temperature_c":21,
"min_temp_c":17,
"max_temp_c":24,
"humidity":70,
"summary":"부산은 비가 내리고 습도가 높습니다.",
},
}
def get_weather(location):
weather = MOCK_WEATHER_DATA.get(location)
if weather is None:
return json.dumps({"error":"날씨 정보를 찾을 수 없습니다."}, ensure_ascii=False)
return json.dumps(
{
"location":location,
"condition":weather["condition"],
"min_temp":weather["min_temp_c"],
"max_temp":weather["max_temp_c"],
},
ensure_ascii=False,
)
def code_interpreter(code:str)->str:
stdout_capture = io.StringIO()
try:
with contextlib.redirect_stdout(stdout_capture):
exec(code, {}) #noqa:S102
output=stdout_capture.getvalue()
return output.strip() if output.strip() else "(출력없음)"
except Exception as e:
return f"[실행오류] {type(e).__name__}:{e}"
avaiable_functions = {
"get_weather":get_weather,
"code_interpreter":code_interpreter,
}
tools = [
{
"type":"function",
"name":"get_weather",
"description":"특정 지역의 현재 최고 및 최저 기온을 조회한다.",
"parameters":{
"type":"object",
"properties":{
"location":{"type":"string","description":"도시 이름(예: 서울, Newyork...)"},
},
"required":["location"]
},
},
{
"type":"function",
"name":"code_interpreter",
"description":(
"Python 코드를 실행하여 계산 결과를 반환한다. "
"수학 계산, 통계, 데이터 분석 등 모든 연산에 사용한다. "
"결과는 반드시 print()로 출력해야 한다."
),
"parameters":{
"type":"object",
"properties":{
"code":{"type":"string","description":"실행할 Python 코드. print()로 결과를 출력해야 한다."},
},
"required":["code"]
},
},
]
def call_llm(user_prompt):
messages = [
{
"role":"system",
"content":(
"너는 절대로 수학적 연산을 직접 수행(암산/추론)해서는 안 된다. "
"평균/합/차/나눗셈 등 모든 계산은 반드시 code_interpreter 도구로만 수행해야 한다. "
"만약 code_interpreter 도구를 사용할 수 없는 상황이라면, "
"계산 결과를 절대 추측하거나 직접 계산하지 말고 "
"'code_interpreter 도구가 없어 계산할 수 없습니다'라고 명시적으로 답해야 한다."
),
},
{
"role":"user",
"content":user_prompt,
}
]
try:
while True:
print("[시스템] 요청을 보냅니다...")
response = client.responses.create(
model=model,
input=messages,
tools=tools,
)
tool_calls = [item for item in response.output if item.type == "function_call"]
if not tool_calls:
print("="*100)
print(f"[최종 답변] {response.output_text}")
print("="*100)
break
print("="*100)
print("response: ",response)
print("="*100)
print("tool_calls: ",tool_calls)
print("="*100)
if tool_calls:
messages.append(
{
"role":"assistant",
"content":response.output_text or "도구 호출 요청",
}
)
for tool_call in tool_calls:
function_name = tool_call.name
function_to_call = avaiable_functions[function_name]
function_args = json.loads(tool_call.arguments)
if function_name == "code_interpreter":
code = function_args.get("code","")
print(f"[도구실행] {function_name}({function_args})호출중...")
function_response = function_to_call(code) if function_to_call else json.dumps(
{"error":f"지원하지 않는 도구입니다: {function_name}"}, ensure_ascii=False,
)
else:
print(f"[도구실행] {function_name}({function_args})호출중...")
if function_to_call is None:
function_response = json.dumps(
{"error":f"지원하지 않는 도구입니다: {function_name}"},
ensure_ascii=False,
)
else:
function_response = function_to_call(
location=function_args.get("location")
)
print(f"[실행 결과] {function_response}")
messages.append(
{
"role":"user",
"content":f"[Tool output for {function_name}]: {function_response}"
}
)
except Exception as e:
print(f"[에러발생]: {e}")
def main():
user_ask = "현재 서울의 평균 기온을 계산해 줘."
print("="*100)
print("질문: ", user_ask)
print("="*100)
answer_5 = call_llm(user_ask)
if __name__ =="__main__":
main()
반응형