AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

LLM gateways

An LLM gateway is a proxy layer between an application and its model providers that every model call passes through, adding routing, retries, timeouts, fallbacks, load balancing, caching, rate limits and logging in one place.

Last updated: 09 Oct, 2026 · OpenAI SDK 3.3

On the whiteboard in AI guardrails, the video ties each property of a production app to the thing that provides it: security to guardrails and fault tolerance to gateways, the subject of the class before it. Installing Python for AI security already showed the failure a gateway is for: a model id that worked in June returned a 404 by October. This lesson covers what a gateway does and builds its central idea, the fallback, by hand.

What a gateway adds to a model call

Core capabilities of an LLM gateway · from the What Are LLM Gateways video · 7:47 to 11:44

This part of the What Are LLM Gateways With Detailed Implementation video starts at 0:07:47. It lists the core capabilities one after another: a unified API, so one function call reaches every provider; automatic fallbacks to a backup model; routing of different requests to different models; load balancing across keys and providers; caching of repeated questions; and observability, with every prompt, reply, token and cost logged. Gateways can also host guardrails and evaluation hooks, which is why the two layers are often confused.

The repository of the gateways class explains the idea with a restaurant. The application is the kitchen and the model provider is a supplier. Without a gateway the chef phones one supplier directly, and when that supplier is out of stock the kitchen stops. A gateway is the purchasing manager in between: it sends each order to the right supplier, tries a backup when the first one fails, keeps a record of every order, remembers frequent orders so they are not paid for twice, and cancels an order that takes too long.

Problem in productionWithout a gatewayThe gateway feature
The provider returns 429, rate limit reachedThe request fails and the user sees an errorRetries with a growing wait
The provider or one model is down or retiredEvery request fails until someone changes the codeFallback to the next target
A response stallsA worker hangs and requests pile up behind itA request timeout
The same question arrives a thousand timesA thousand paid model callsA cache
One key or model carries all the trafficIt reaches its limit firstLoad balancing across targets
A new provider is addedEach call site is rewrittenRouting by configuration, one API
Nobody knows what was sent or what it costNo record to debug or bill fromLogs and metadata per request

Gateway vs guardrails

A user's message first meets guardrails, which ask whether the request should happen at all and reject a blocked one with a message. A passed request goes to the LLM gateway, which asks how the request should be sent: routing, retries, timeouts, fallbacks, load balancing, caching, rate limits and logs. The gateway sends it to provider or model A and falls back to provider or model B when A fails.
GuardrailsGateway
AsksShould this request happen at all?How should this request be sent?
Looks atThe content of the message and the replyStatus codes, timings, targets and cost
StopsJailbreaks, off-topic requests, personal dataNothing by content in its core job: it routes, retries and limits
OWASP risk it addressesLLM01, LLM02, LLM05, LLM07LLM10 Unbounded Consumption
In the video's materialNeMo GuardrailsPortkey

The two work together: guardrails at the gate, the gateway for every request that passes.

Gateway features in the Portkey demos

The repository of the gateways class uses Portkey, a hosted gateway with a Python SDK. Shown, not run here: the blocks need a Portkey account, a Portkey API key in PORTKEY_API_KEY and the portkey-ai package. The repository's targets name llama-3.3-70b-versatile and llama-3.1-8b-instant, since retired on Groq; the blocks below name the two current models. The repository's first demos pass the provider as virtual_key; the blocks use the @provider-slug/model string of the current Portkey documentation, with a placeholder for the slug in your Portkey account.

Routing a call through the gateway

python
import os

from portkey_ai import Portkey

portkey = Portkey(api_key=os.environ["PORTKEY_API_KEY"])

response = portkey.chat.completions.create(
    model="@your-provider-slug/openai/gpt-oss-120b",
    messages=[{"role": "user", "content": "What is Kubernetes?"}],
)

The call has the same shape as a direct one. Only the client and the model string change, and from then on every call appears in the gateway's log with the model used, its tokens and its cost.

Retries and a timeout

python
portkey = Portkey(api_key=os.environ["PORTKEY_API_KEY"], config={
    "request_timeout": 10000,                                   # milliseconds
    "retry": {"attempts": 3, "on_status_codes": [429, 500, 502, 503, 504]},
})

On one of the listed status codes the gateway waits and sends the request again, up to three more times with a longer wait each time, and the application receives only the final result. request_timeout ends a call that has not answered after 10 seconds with a 408 error.

A fallback target

python
portkey = Portkey(api_key=os.environ["PORTKEY_API_KEY"], config={
    "strategy": {"mode": "fallback"},
    "targets": [
        {"override_params": {"model": "@your-provider-slug/openai/gpt-oss-120b"}},   # primary
        {"override_params": {"model": "@your-provider-slug/openai/gpt-oss-20b"}},    # fallback
    ],
})

Targets are tried in order. When the primary fails, the same request goes to the next one.

Load balancing and caching

python
loadbalance = {"strategy": {"mode": "loadbalance"}, "targets": [
    {"override_params": {"model": "@your-provider-slug/openai/gpt-oss-120b"}, "weight": 0.7},
    {"override_params": {"model": "@your-provider-slug/openai/gpt-oss-20b"}, "weight": 0.3},
]}
cache = {"cache": {"mode": "simple"}}

With loadbalance, each request is sent to one target at random in proportion to the weights, here about 70 and 30 in a hundred. With the simple cache, a request identical to an earlier one is answered from the stored reply without a model call.

Building a fallback and retry wrapper

The central behaviour of a gateway fits in a short function over the OpenAI SDK. Portkey applies these rules on its servers from a configuration; the wrapper here applies them in your process, on the course key. Save the pieces as fallback.py.

The client and the errors worth retrying

python
import os
import time

import openai
from openai import OpenAI

client = OpenAI(base_url="https://api.groq.com/openai/v1",
                api_key=os.environ["GROQ_API_KEY"], max_retries=0)
RETRYABLE = (openai.RateLimitError, openai.APITimeoutError, openai.InternalServerError)

The OpenAI SDK retries some failed calls by itself. max_retries=0 turns that off so that every attempt is made, and printed, by the wrapper. RETRYABLE holds the temporary failures: a rate limit (429), a timeout and a server error (5xx). Trying again after a short wait can succeed for these.

The wrapper

python
def chat_with_fallback(messages, models, attempts=2, timeout=20):
    for model in models:
        for attempt in range(1, attempts + 1):
            start = time.perf_counter()
            try:
                reply = client.chat.completions.create(
                    model=model, messages=messages, temperature=0, max_tokens=300, timeout=timeout)
                print(f"{model}: ok after {time.perf_counter() - start:.2f} s, attempt {attempt}")
                return reply.choices[0].message.content
            except RETRYABLE as error:               # temporary: wait, try the same model again
                print(f"{model}: {type(error).__name__}, attempt {attempt}, waiting {attempt} s")
                time.sleep(attempt)
            except openai.APIStatusError as error:   # permanent: go to the next model
                print(f"{model}: {error.status_code} {error.code} after {time.perf_counter() - start:.2f} s")
                break
    raise RuntimeError("every model in the list failed")
A request goes to model 1, llama-3.3-70b-versatile, which returns 404 model_not_found; the wrapper falls back to model 2, openai/gpt-oss-20b, which answers. Below, the rule for each outcome: a reply is returned; a 429, a timeout or a 5xx error means wait and try the same model again; any other error, such as 404 or 401, means move to the next model in the list.
  • The outer loop is the fallback: it walks the list of models in order.
  • The inner loop is the retry: up to attempts tries per model, with a wait of 1 second after the first failure and 2 seconds after the second.
  • timeout=timeout ends a call that has not answered in 20 seconds; the SDK raises APITimeoutError, which is in RETRYABLE.
  • Any other status error, such as a 404 for a model that no longer exists, will fail the same way every time, so the wrapper leaves the retry loop with break and moves to the next model.

Running the wrapper with a retired first model

The first model in the list is llama-3.3-70b-versatile, the primary model of the gateway repository, which Groq has retired. The second is openai/gpt-oss-20b.

ExampleAPI keyRun on Groq
import os
import time

import openai
from openai import OpenAI

client = OpenAI(base_url="https://api.groq.com/openai/v1",
                api_key=os.environ["GROQ_API_KEY"], max_retries=0)
RETRYABLE = (openai.RateLimitError, openai.APITimeoutError, openai.InternalServerError)


def chat_with_fallback(messages, models, attempts=2, timeout=20):
    for model in models:
        for attempt in range(1, attempts + 1):
            start = time.perf_counter()
            try:
                reply = client.chat.completions.create(
                    model=model, messages=messages, temperature=0, max_tokens=300, timeout=timeout)
                print(f"{model}: ok after {time.perf_counter() - start:.2f} s, attempt {attempt}")
                return reply.choices[0].message.content
            except RETRYABLE as error:               # temporary: wait, try the same model again
                print(f"{model}: {type(error).__name__}, attempt {attempt}, waiting {attempt} s")
                time.sleep(attempt)
            except openai.APIStatusError as error:   # permanent: go to the next model
                print(f"{model}: {error.status_code} {error.code} after {time.perf_counter() - start:.2f} s")
                break
    raise RuntimeError("every model in the list failed")


answer = chat_with_fallback(
    [{"role": "user", "content": "In one sentence, what is an LLM gateway?"}],
    models=["llama-3.3-70b-versatile", "openai/gpt-oss-20b"],
)
print(answer)

What the fallback did

  • The first model failed in 0.12 seconds with 404 model_not_found. A 404 is not in RETRYABLE, so there was no second attempt and no wait: asking again for a model that does not exist cannot help.
  • The second model answered in 0.60 seconds on its first attempt, and its sentence is what the caller received.
  • The caller saw one answer and no error. The retired model cost about a tenth of a second. Without the wrapper the same request ends in the traceback shown in Installing Python for AI security.
  • The timings are from this run and change with the network and the load on the provider; the order of events does not.

A hand-written wrapper vs a gateway

The wrapper aboveA gateway such as Portkey
RunsInside each application processAs one proxy all applications call
Rules change byEditing and redeploying codeEditing a configuration
Cache and rate limitsPer process, lost on restartShared across every instance
LogsWhatever you printA dashboard of requests, tokens and cost
Good forOne service, learning the ideaSeveral services or providers

Where you use an LLM gateway

  • An app with more than one provider or model, so that an outage or a retired model becomes a slower answer and not an error.
  • Several teams sharing provider keys, where per-team budgets and rate limits have to be enforced in one place.
  • Cost control: cached answers for repeated questions and a record of which feature spends the tokens.
Watch out. A fallback changes which model answers. The backup may be smaller, follow instructions differently and fail guardrail or evaluation checks the primary passes. Test the whole app on every model in the list, and log which model served each request.
Try it yourself
  • Swap the two models in the list. The working model answers first and the retired one is never called.
  • Pass only ["llama-3.3-70b-versatile"]. After the 404 the list is used up and the wrapper raises RuntimeError: every model in the list failed.
  • Call it with models=["openai/gpt-oss-20b"] and timeout=0.01. No reply can arrive in a hundredth of a second, so you see APITimeoutError twice, with waits of 1 and 2 seconds, before the wrapper gives up.

Little by little, you're building something great.