LangChain (YT style)LangChain 1.4 · Python 3.12+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

Several tool calls at once

Parallel tool calls are several tool requests the model makes in one message, which the agent runs together before it calls the model again.

Last updated: 27 Sep, 2026 · LangChain 1.4

A written stand-in model here
The runs in this lesson use ShopModel, a small stand-in model written below, not Groq. It always puts every tool call in one AI message, which is what this lesson is about. Groq's gpt-oss-120b usually asks for tools one at a time instead, so it would not show parallel calls on cue.

In the streaming-steps lesson, a question about A17 and C40 made the real model look the orders up one at a time, in separate AI messages. A model can also ask for several tools in one AI message, and the agent then runs them all before it calls the model again. Each call carries its own id, and each result comes back tagged with it, so the model can tell which answer belongs to which request. The lesson ends with return_direct, which lets a tool's result be the answer and skips the last model call.

The return_direct and parallel_tool_calls options

python
@tool(return_direct=True)          # end the run as soon as this tool returns
def my_tool(...): ...

# turn parallel calls off when binding (OpenAI, Anthropic)
model.bind_tools([t], parallel_tool_calls=False)

The stand-in model

ShopModel is built like the one in the BaseChatModel lesson, with three additions: a tools field that bind_tools fills in, replies that are tool calls, and a decide method that holds the reply logic. This is the last version of ShopModel. Save the class as shop_model.py: the desk lessons import it from there, and the middleware lessons repeat it at the top of their file so each one runs on its own. Start with the imports and the type name.

python
import re

from langchain.chat_models import BaseChatModel
from langchain.messages import AIMessage, ToolMessage
from langchain_core.outputs import ChatGeneration, ChatResult


class ShopModel(BaseChatModel):
    tools: list = []

    @property
    def _llm_type(self):
        return "shop"

bind_tools returns a copy holding the tools. _generate hands the conversation to decide and wraps the message it returns. With the logic in one method, a later stand-in such as StuckModel in the call-limits lesson changes the replies by overriding decide alone.

python
    def bind_tools(self, tools, **kwargs):
        return self.model_copy(update={"tools": tools})   # a copy holding the tools

    def _generate(self, messages, stop=None, run_manager=None, **kwargs):
        message = self.decide(messages)                   # the reply comes from decide
        return ChatResult(generations=[ChatGeneration(message=message)])

decide first checks whether tool results have come back. It collects the tool messages at the end of the conversation and, if there are any, answers by joining their text.

python
    def decide(self, messages):
        results = []                                # the tool results at the end
        for m in reversed(messages):
            if not isinstance(m, ToolMessage):
                break
            results.insert(0, m.text)
        if results:                                 # results are back: answer with them
            return AIMessage(" ".join(results))

Otherwise it looks for order ids in the last message. When the message mentions a refund it asks for refund_order, and otherwise for lookup_order: one call per order id, all in one AI message. If that tool is not bound, it says it has no way to look the order up, as the BaseChatModel lesson's version did.

python
        text = messages[-1].text
        orders = re.findall(r"\b[A-Z]\d+\b", text)
        tool = "refund_order" if "refund" in text.lower() else "lookup_order"
        if orders and tool in [t.name for t in self.tools]:   # one call per order id
            calls = [{"name": tool, "args": {"order_id": o}, "id": f"call_{o}"}
                     for o in orders]
            return AIMessage("", tool_calls=calls)
        if orders:                                  # that tool is not bound
            return AIMessage(f"I have no way to look up {orders[0]} yet.")
        return AIMessage("Hello. Which order is this about?")

Building the agent

This lesson's agent answers order questions with lookup_order, the tool built in Tools: a function the model can call. Start the file with it.

python
from langchain.tools import tool

ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}


@tool
def lookup_order(order_id: str) -> str:
    """Look up an order's shipping status by its id, such as A17."""
    status = ORDERS.get(order_id)
    return f"{order_id} {status}." if status else f"{order_id} is not an order we have."

Build the agent with the stand-in.

python
from shop_model import ShopModel
from langchain.agents import create_agent

agent = create_agent(ShopModel(), tools=[lookup_order])

Three calls in one run

Ask about three orders in one message; ShopModel puts all three requests in one AI message. Then walk the messages between the question and the final answer, printing each tool call's id and each result tagged with the id it answers. Three calls go out together and three results come back, all before the model is called a second time.

Example
result = agent.invoke({"messages": [{"role": "user", "content": "Where are A17, B22 and C40?"}]})

for message in result["messages"][1:-1]:
    if message.type == "ai":
        for call in message.tool_calls:
            print("asked ", call["id"])
    else:
        print("result", message.tool_call_id, "->", message.text)

Skipping the last model call

Without return_direct there is one more model call after the tools: the loop above skipped that last AI message with [1:-1], and ShopModel's answer there only repeats what the tools said. A tool marked return_direct=True ends the run the moment it returns, and its result becomes the answer.

python
from langchain.tools import tool

ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}


@tool(return_direct=True)
def lookup_order(order_id: str) -> str:
    """Look up an order's shipping status by its id, such as A17."""
    status = ORDERS.get(order_id)
    return f"{order_id} {status}." if status else f"{order_id} is not an order we have."

Use this version in place of the plain one and run the same question and print every message with its type. The run ends on the tool results, with no AI message after them.

Example
from langchain.agents import create_agent

agent = create_agent(ShopModel(), tools=[lookup_order])
result = agent.invoke({"messages": [{"role": "user", "content": "Where are A17 and C40?"}]})

for message in result["messages"]:
    print(f"{message.type:<5} {message.text or len(message.tool_calls)}")

What the ids matched

  • Three calls, three results. All three ran before the model was called a second time, and each result is tagged with the id of the call it answers.
  • To switch parallel calls off, OpenAI and Anthropic accept parallel_tool_calls=False when you bind tools yourself, as in bind_tools: asking for a tool.
  • return_direct ends the run on the tool result: one model call instead of two, and its text is the answer.

A normal tool vs return_direct

Normal toolreturn_direct
Result goesBack to the modelStraight out as the answer
Model callsTwo: ask, then answerOne: ask only
Run endsAfter the model's replyAs soon as the tool returns

When to run tools in parallel

  • Looking up several items in one request instead of a round trip to the model per item.
  • A tool whose result is the finished answer, such as a lookup that needs no rewording.
Watch out. When the model calls several tools in one step, return_direct ends the run only if every one of them has it. Mix in one plain tool and all the results go back to the model as usual.
Try it yourself
  • Ask about one order with return_direct on and count the messages.
  • Stream the return_direct agent and check which step comes last.
  • Print each tool message's tool_call_id and compare it with the id of each call in the AI message.

Slow is fine. Stopping is the only problem.