Several tool calls at once
Parallel tool calls are several tool requests the model makes in one message, which the agent runs together before it calls the model again.
Last updated: 27 Sep, 2026 · LangChain 1.4
ShopModel, a small stand-in model written below, not Groq. It always puts every tool call in one AI message, which is what this lesson is about. Groq's gpt-oss-120b usually asks for tools one at a time instead, so it would not show parallel calls on cue.In the streaming-steps lesson, a question about A17 and C40 made the real model look the orders up one at a time, in separate AI messages. A model can also ask for several tools in one AI message, and the agent then runs them all before it calls the model again. Each call carries its own id, and each result comes back tagged with it, so the model can tell which answer belongs to which request. The lesson ends with return_direct, which lets a tool's result be the answer and skips the last model call.
The return_direct and parallel_tool_calls options
@tool(return_direct=True) # end the run as soon as this tool returns
def my_tool(...): ...
# turn parallel calls off when binding (OpenAI, Anthropic)
model.bind_tools([t], parallel_tool_calls=False)The stand-in model
ShopModel is built like the one in the BaseChatModel lesson, with three additions: a tools field that bind_tools fills in, replies that are tool calls, and a decide method that holds the reply logic. This is the last version of ShopModel. Save the class as shop_model.py: the desk lessons import it from there, and the middleware lessons repeat it at the top of their file so each one runs on its own. Start with the imports and the type name.
import re
from langchain.chat_models import BaseChatModel
from langchain.messages import AIMessage, ToolMessage
from langchain_core.outputs import ChatGeneration, ChatResult
class ShopModel(BaseChatModel):
tools: list = []
@property
def _llm_type(self):
return "shop"bind_tools returns a copy holding the tools. _generate hands the conversation to decide and wraps the message it returns. With the logic in one method, a later stand-in such as StuckModel in the call-limits lesson changes the replies by overriding decide alone.
def bind_tools(self, tools, **kwargs):
return self.model_copy(update={"tools": tools}) # a copy holding the tools
def _generate(self, messages, stop=None, run_manager=None, **kwargs):
message = self.decide(messages) # the reply comes from decide
return ChatResult(generations=[ChatGeneration(message=message)])decide first checks whether tool results have come back. It collects the tool messages at the end of the conversation and, if there are any, answers by joining their text.
def decide(self, messages):
results = [] # the tool results at the end
for m in reversed(messages):
if not isinstance(m, ToolMessage):
break
results.insert(0, m.text)
if results: # results are back: answer with them
return AIMessage(" ".join(results))Otherwise it looks for order ids in the last message. When the message mentions a refund it asks for refund_order, and otherwise for lookup_order: one call per order id, all in one AI message. If that tool is not bound, it says it has no way to look the order up, as the BaseChatModel lesson's version did.
text = messages[-1].text
orders = re.findall(r"\b[A-Z]\d+\b", text)
tool = "refund_order" if "refund" in text.lower() else "lookup_order"
if orders and tool in [t.name for t in self.tools]: # one call per order id
calls = [{"name": tool, "args": {"order_id": o}, "id": f"call_{o}"}
for o in orders]
return AIMessage("", tool_calls=calls)
if orders: # that tool is not bound
return AIMessage(f"I have no way to look up {orders[0]} yet.")
return AIMessage("Hello. Which order is this about?")Building the agent
This lesson's agent answers order questions with lookup_order, the tool built in Tools: a function the model can call. Start the file with it.
from langchain.tools import tool
ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}
@tool
def lookup_order(order_id: str) -> str:
"""Look up an order's shipping status by its id, such as A17."""
status = ORDERS.get(order_id)
return f"{order_id} {status}." if status else f"{order_id} is not an order we have."Build the agent with the stand-in.
from shop_model import ShopModel
from langchain.agents import create_agent
agent = create_agent(ShopModel(), tools=[lookup_order])Three calls in one run
Ask about three orders in one message; ShopModel puts all three requests in one AI message. Then walk the messages between the question and the final answer, printing each tool call's id and each result tagged with the id it answers. Three calls go out together and three results come back, all before the model is called a second time.
result = agent.invoke({"messages": [{"role": "user", "content": "Where are A17, B22 and C40?"}]})
for message in result["messages"][1:-1]:
if message.type == "ai":
for call in message.tool_calls:
print("asked ", call["id"])
else:
print("result", message.tool_call_id, "->", message.text)asked call_A17 asked call_B22 asked call_C40 result call_A17 -> A17 shipped on 3 March. result call_B22 -> B22 is not an order we have. result call_C40 -> C40 waiting for stock.
Skipping the last model call
Without return_direct there is one more model call after the tools: the loop above skipped that last AI message with [1:-1], and ShopModel's answer there only repeats what the tools said. A tool marked return_direct=True ends the run the moment it returns, and its result becomes the answer.
from langchain.tools import tool
ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}
@tool(return_direct=True)
def lookup_order(order_id: str) -> str:
"""Look up an order's shipping status by its id, such as A17."""
status = ORDERS.get(order_id)
return f"{order_id} {status}." if status else f"{order_id} is not an order we have."Use this version in place of the plain one and run the same question and print every message with its type. The run ends on the tool results, with no AI message after them.
from langchain.agents import create_agent
agent = create_agent(ShopModel(), tools=[lookup_order])
result = agent.invoke({"messages": [{"role": "user", "content": "Where are A17 and C40?"}]})
for message in result["messages"]:
print(f"{message.type:<5} {message.text or len(message.tool_calls)}")human Where are A17 and C40? ai 2 tool A17 shipped on 3 March. tool C40 waiting for stock.
What the ids matched
- Three calls, three results. All three ran before the model was called a second time, and each result is tagged with the id of the call it answers.
- To switch parallel calls off, OpenAI and Anthropic accept
parallel_tool_calls=Falsewhen you bind tools yourself, as in bind_tools: asking for a tool. - return_direct ends the run on the tool result: one model call instead of two, and its text is the answer.
A normal tool vs return_direct
| Normal tool | return_direct | |
|---|---|---|
| Result goes | Back to the model | Straight out as the answer |
| Model calls | Two: ask, then answer | One: ask only |
| Run ends | After the model's reply | As soon as the tool returns |
When to run tools in parallel
- Looking up several items in one request instead of a round trip to the model per item.
- A tool whose result is the finished answer, such as a lookup that needs no rewording.
return_direct ends the run only if every one of them has it. Mix in one plain tool and all the results go back to the model as usual.Related
- Previous: Validation errors, and the retry
- Next: Short-term memory with a checkpointer
- Reference: create_agent and tools
- Ask about one order with
return_directon and count the messages. - Stream the
return_directagent and check which step comes last. - Print each tool message's
tool_call_idand compare it with theidof each call in the AI message.
Slow is fine. Stopping is the only problem.