LLM gateways
An LLM gateway is a proxy layer between an application and its model providers that every model call passes through, adding routing, retries, timeouts, fallbacks, load balancing, caching, rate limits and logging in one place.
Last updated: 09 Oct, 2026 · OpenAI SDK 3.3
On the whiteboard in AI guardrails, the video ties each property of a production app to the thing that provides it: security to guardrails and fault tolerance to gateways, the subject of the class before it. Installing Python for AI security already showed the failure a gateway is for: a model id that worked in June returned a 404 by October. This lesson covers what a gateway does and builds its central idea, the fallback, by hand.
What a gateway adds to a model call
This part of the What Are LLM Gateways With Detailed Implementation video starts at 0:07:47. It lists the core capabilities one after another: a unified API, so one function call reaches every provider; automatic fallbacks to a backup model; routing of different requests to different models; load balancing across keys and providers; caching of repeated questions; and observability, with every prompt, reply, token and cost logged. Gateways can also host guardrails and evaluation hooks, which is why the two layers are often confused.
The repository of the gateways class explains the idea with a restaurant. The application is the kitchen and the model provider is a supplier. Without a gateway the chef phones one supplier directly, and when that supplier is out of stock the kitchen stops. A gateway is the purchasing manager in between: it sends each order to the right supplier, tries a backup when the first one fails, keeps a record of every order, remembers frequent orders so they are not paid for twice, and cancels an order that takes too long.
| Problem in production | Without a gateway | The gateway feature |
|---|---|---|
| The provider returns 429, rate limit reached | The request fails and the user sees an error | Retries with a growing wait |
| The provider or one model is down or retired | Every request fails until someone changes the code | Fallback to the next target |
| A response stalls | A worker hangs and requests pile up behind it | A request timeout |
| The same question arrives a thousand times | A thousand paid model calls | A cache |
| One key or model carries all the traffic | It reaches its limit first | Load balancing across targets |
| A new provider is added | Each call site is rewritten | Routing by configuration, one API |
| Nobody knows what was sent or what it cost | No record to debug or bill from | Logs and metadata per request |
Gateway vs guardrails
| Guardrails | Gateway | |
|---|---|---|
| Asks | Should this request happen at all? | How should this request be sent? |
| Looks at | The content of the message and the reply | Status codes, timings, targets and cost |
| Stops | Jailbreaks, off-topic requests, personal data | Nothing by content in its core job: it routes, retries and limits |
| OWASP risk it addresses | LLM01, LLM02, LLM05, LLM07 | LLM10 Unbounded Consumption |
| In the video's material | NeMo Guardrails | Portkey |
The two work together: guardrails at the gate, the gateway for every request that passes.
Gateway features in the Portkey demos
The repository of the gateways class uses Portkey, a hosted gateway with a Python SDK. Shown, not run here: the blocks need a Portkey account, a Portkey API key in PORTKEY_API_KEY and the portkey-ai package. The repository's targets name llama-3.3-70b-versatile and llama-3.1-8b-instant, since retired on Groq; the blocks below name the two current models. The repository's first demos pass the provider as virtual_key; the blocks use the @provider-slug/model string of the current Portkey documentation, with a placeholder for the slug in your Portkey account.
Routing a call through the gateway
import os
from portkey_ai import Portkey
portkey = Portkey(api_key=os.environ["PORTKEY_API_KEY"])
response = portkey.chat.completions.create(
model="@your-provider-slug/openai/gpt-oss-120b",
messages=[{"role": "user", "content": "What is Kubernetes?"}],
)The call has the same shape as a direct one. Only the client and the model string change, and from then on every call appears in the gateway's log with the model used, its tokens and its cost.
Retries and a timeout
portkey = Portkey(api_key=os.environ["PORTKEY_API_KEY"], config={
"request_timeout": 10000, # milliseconds
"retry": {"attempts": 3, "on_status_codes": [429, 500, 502, 503, 504]},
})On one of the listed status codes the gateway waits and sends the request again, up to three more times with a longer wait each time, and the application receives only the final result. request_timeout ends a call that has not answered after 10 seconds with a 408 error.
A fallback target
portkey = Portkey(api_key=os.environ["PORTKEY_API_KEY"], config={
"strategy": {"mode": "fallback"},
"targets": [
{"override_params": {"model": "@your-provider-slug/openai/gpt-oss-120b"}}, # primary
{"override_params": {"model": "@your-provider-slug/openai/gpt-oss-20b"}}, # fallback
],
})Targets are tried in order. When the primary fails, the same request goes to the next one.
Load balancing and caching
loadbalance = {"strategy": {"mode": "loadbalance"}, "targets": [
{"override_params": {"model": "@your-provider-slug/openai/gpt-oss-120b"}, "weight": 0.7},
{"override_params": {"model": "@your-provider-slug/openai/gpt-oss-20b"}, "weight": 0.3},
]}
cache = {"cache": {"mode": "simple"}}With loadbalance, each request is sent to one target at random in proportion to the weights, here about 70 and 30 in a hundred. With the simple cache, a request identical to an earlier one is answered from the stored reply without a model call.
Building a fallback and retry wrapper
The central behaviour of a gateway fits in a short function over the OpenAI SDK. Portkey applies these rules on its servers from a configuration; the wrapper here applies them in your process, on the course key. Save the pieces as fallback.py.
The client and the errors worth retrying
import os
import time
import openai
from openai import OpenAI
client = OpenAI(base_url="https://api.groq.com/openai/v1",
api_key=os.environ["GROQ_API_KEY"], max_retries=0)
RETRYABLE = (openai.RateLimitError, openai.APITimeoutError, openai.InternalServerError)The OpenAI SDK retries some failed calls by itself. max_retries=0 turns that off so that every attempt is made, and printed, by the wrapper. RETRYABLE holds the temporary failures: a rate limit (429), a timeout and a server error (5xx). Trying again after a short wait can succeed for these.
The wrapper
def chat_with_fallback(messages, models, attempts=2, timeout=20):
for model in models:
for attempt in range(1, attempts + 1):
start = time.perf_counter()
try:
reply = client.chat.completions.create(
model=model, messages=messages, temperature=0, max_tokens=300, timeout=timeout)
print(f"{model}: ok after {time.perf_counter() - start:.2f} s, attempt {attempt}")
return reply.choices[0].message.content
except RETRYABLE as error: # temporary: wait, try the same model again
print(f"{model}: {type(error).__name__}, attempt {attempt}, waiting {attempt} s")
time.sleep(attempt)
except openai.APIStatusError as error: # permanent: go to the next model
print(f"{model}: {error.status_code} {error.code} after {time.perf_counter() - start:.2f} s")
break
raise RuntimeError("every model in the list failed")- The outer loop is the fallback: it walks the list of models in order.
- The inner loop is the retry: up to
attemptstries per model, with a wait of 1 second after the first failure and 2 seconds after the second. timeout=timeoutends a call that has not answered in 20 seconds; the SDK raisesAPITimeoutError, which is inRETRYABLE.- Any other status error, such as a 404 for a model that no longer exists, will fail the same way every time, so the wrapper leaves the retry loop with
breakand moves to the next model.
Running the wrapper with a retired first model
The first model in the list is llama-3.3-70b-versatile, the primary model of the gateway repository, which Groq has retired. The second is openai/gpt-oss-20b.
import os
import time
import openai
from openai import OpenAI
client = OpenAI(base_url="https://api.groq.com/openai/v1",
api_key=os.environ["GROQ_API_KEY"], max_retries=0)
RETRYABLE = (openai.RateLimitError, openai.APITimeoutError, openai.InternalServerError)
def chat_with_fallback(messages, models, attempts=2, timeout=20):
for model in models:
for attempt in range(1, attempts + 1):
start = time.perf_counter()
try:
reply = client.chat.completions.create(
model=model, messages=messages, temperature=0, max_tokens=300, timeout=timeout)
print(f"{model}: ok after {time.perf_counter() - start:.2f} s, attempt {attempt}")
return reply.choices[0].message.content
except RETRYABLE as error: # temporary: wait, try the same model again
print(f"{model}: {type(error).__name__}, attempt {attempt}, waiting {attempt} s")
time.sleep(attempt)
except openai.APIStatusError as error: # permanent: go to the next model
print(f"{model}: {error.status_code} {error.code} after {time.perf_counter() - start:.2f} s")
break
raise RuntimeError("every model in the list failed")
answer = chat_with_fallback(
[{"role": "user", "content": "In one sentence, what is an LLM gateway?"}],
models=["llama-3.3-70b-versatile", "openai/gpt-oss-20b"],
)
print(answer)llama-3.3-70b-versatile: 404 model_not_found after 0.12 s openai/gpt-oss-20b: ok after 0.60 s, attempt 1 An LLM gateway is a middleware layer that authenticates, routes, and manages client requests to large language model services, often adding features such as rate limiting, caching, and policy enforcement.
What the fallback did
- The first model failed in 0.12 seconds with
404 model_not_found. A 404 is not inRETRYABLE, so there was no second attempt and no wait: asking again for a model that does not exist cannot help. - The second model answered in 0.60 seconds on its first attempt, and its sentence is what the caller received.
- The caller saw one answer and no error. The retired model cost about a tenth of a second. Without the wrapper the same request ends in the traceback shown in Installing Python for AI security.
- The timings are from this run and change with the network and the load on the provider; the order of events does not.
A hand-written wrapper vs a gateway
| The wrapper above | A gateway such as Portkey | |
|---|---|---|
| Runs | Inside each application process | As one proxy all applications call |
| Rules change by | Editing and redeploying code | Editing a configuration |
| Cache and rate limits | Per process, lost on restart | Shared across every instance |
| Logs | Whatever you print | A dashboard of requests, tokens and cost |
| Good for | One service, learning the idea | Several services or providers |
Where you use an LLM gateway
- An app with more than one provider or model, so that an outage or a retired model becomes a slower answer and not an error.
- Several teams sharing provider keys, where per-team budgets and rate limits have to be enforced in one place.
- Cost control: cached answers for repeated questions and a record of which feature spends the tokens.
Related
- Previous: Guardrail frameworks
- Next: NeMo Guardrails
- See also: Redis caching for RAG
- Reference: the gateways class repository and Portkey documentation
- Swap the two models in the list. The working model answers first and the retired one is never called.
- Pass only
["llama-3.3-70b-versatile"]. After the 404 the list is used up and the wrapper raisesRuntimeError: every model in the list failed. - Call it with
models=["openai/gpt-oss-20b"]andtimeout=0.01. No reply can arrive in a hundredth of a second, so you seeAPITimeoutErrortwice, with waits of 1 and 2 seconds, before the wrapper gives up.
Little by little, you're building something great.