Rate limits and retries
A rate limit is a provider's cap on how many tokens or requests one key may send per minute, and DeepEval's AsyncConfig and ErrorConfig are the settings that keep an evaluation under that cap and keep it running when a call still fails.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
The evaluate() run scored one test case at a time, with AsyncConfig(max_concurrent=1). This lesson removes that line, shows the error Groq's free plan returns, and goes through each setting that prevents it.
One judge call per metric, per golden
Someone in the video asks why the judge does not score everything in one call. Because a judge, like a person, works best with little on its plate: LLMs perform best when only part of their context window is full, and a judge with too much to read can hallucinate, like a biased judge in court. So every metric gets its own call for every golden. With five goldens and five metrics, that is 25 calls, and 25 calls sent to Groq at once on the free tier hit the rate limit.
The video's app avoids it with cooldowns: a sample cooldown of 25 seconds and an experiment cooldown of 35 seconds, plain sleep calls that let the API limits go back to normal. The video's app scores with RAGAS and writes those pauses itself. DeepEval has settings for the same job, passed to evaluate().
- written in Custom judge model
- written in TechNest RAG app
- written in TechNest RAG app
- written in Datasets and goldens
- written in evaluate()
View the code here
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
import json
import os
import re
from openai import OpenAI
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
CHAT_MODEL = "qwen/qwen3.8-27b"
SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""
with open("catalog.json", encoding="utf-8") as f:
CATALOG = json.load(f)
SKIP = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i", "in", "is", "it",
"long", "much", "my", "of", "on", "s", "technest", "the", "to", "what", "with", "you", "your"}
def words(text):
"""The words in a text that carry meaning, in lower case."""
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in SKIP}
def retrieve(question, top_k=3):
asked = words(question)
ranked = sorted(CATALOG, key=lambda item: len(asked & words(item["title"] + " " + item["content"])), reverse=True)
return [item["content"] for item in ranked[:top_k]]
def generate(question, contexts):
context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
]
response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
return response.choices[0].message.content.strip()
def answer(question, top_k=3):
contexts = retrieve(question, top_k)
return generate(question, contexts), contexts
[
{
"id": "prod_001",
"category": "product",
"title": "ProBook X1 Laptop",
"content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
},
{
"id": "prod_002",
"category": "product",
"title": "PixelPhone 15",
"content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
},
{
"id": "prod_003",
"category": "product",
"title": "SoundPods Pro",
"content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
},
{
"id": "prod_004",
"category": "product",
"title": "UltraTab S2 Tablet",
"content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
},
{
"id": "prod_005",
"category": "product",
"title": "SmartWatch X",
"content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
},
{
"id": "prod_006",
"category": "product",
"title": "ProCam 4K Action Camera",
"content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
},
{
"id": "prod_007",
"category": "product",
"title": "BassBuds Max Headphones",
"content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
},
{
"id": "prod_008",
"category": "product",
"title": "SoundBar 360",
"content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
},
{
"id": "policy_001",
"category": "policy",
"title": "Return Policy",
"content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
},
{
"id": "policy_002",
"category": "policy",
"title": "Shipping Policy",
"content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
},
{
"id": "policy_003",
"category": "policy",
"title": "Warranty Policy",
"content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
},
{
"id": "policy_004",
"category": "policy",
"title": "Payment Policy",
"content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
},
{
"id": "faq_001",
"category": "faq",
"title": "Order Tracking",
"content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
},
{
"id": "faq_002",
"category": "faq",
"title": "International Shipping",
"content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
},
{
"id": "faq_003",
"category": "faq",
"title": "Bulk and Business Orders",
"content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
}
]
[
{
"name": "g001",
"input": "What is TechNest's return policy?",
"expected_output": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."
},
{
"name": "g002",
"input": "What are the RAM and storage specs of the ProBook X1?",
"expected_output": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."
},
{
"name": "g003",
"input": "How long is the battery life on the SoundPods Pro?",
"expected_output": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours."
},
{
"name": "g004",
"input": "What are TechNest's shipping options and how long do returns take to process?",
"expected_output": "TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received."
},
{
"name": "g005",
"input": "What is the price of the PixelPhone 15?",
"expected_output": "The TechNest PixelPhone 15 is priced at $899."
}
]
from deepeval.dataset import EvaluationDataset
from deepeval.metrics import AnswerRelevancyMetric, GEval
from deepeval.test_case import LLMTestCase, SingleTurnParams
from judge import judge
from technest import answer
dataset = EvaluationDataset()
dataset.add_goldens_from_json_file(file_path="goldens.json")
def build_test_cases(goldens):
"""Ask the bot every golden's question and wrap each reply in a test case."""
test_cases = []
for golden in goldens:
response, contexts = answer(golden.input)
test_cases.append(LLMTestCase(
name=golden.name,
input=golden.input,
actual_output=response,
expected_output=golden.expected_output,
retrieval_context=contexts,
))
return test_cases
relevancy = AnswerRelevancyMetric(model=judge)
correctness = GEval(
name="Correctness",
evaluation_steps=[
"Check whether the facts in 'actual output' contradict any facts in 'expected output'.",
"Penalize facts from 'expected output' that 'actual output' leaves out.",
"Do not penalize extra details or different wording.",
],
evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
model=judge,
)
Running evaluate() without a limit
All five goldens and the two metrics from technest_eval.py, built in evaluate(), with the default settings. The except block prints the status code and the token numbers from Groq's message; the full message also names your Groq organisation.
import re
from deepeval import evaluate
from openai import RateLimitError
from technest_eval import build_test_cases, correctness, dataset, relevancy
test_cases = build_test_cases(dataset.goldens) # all five goldens
try:
evaluate(test_cases=test_cases, metrics=[relevancy, correctness])
except RateLimitError as err:
print("status:", err.status_code)
found = re.search(r"tokens per (minute|day).*?Requested \d+", err.message)
print(found.group() if found else "no token count in the message")✨ You're running DeepEval's latest Answer Relevancy Metric! (using
openai/gpt-oss-120b, strict=False, async_mode=True)...
✨ You're running DeepEval's latest Correctness [GEval] Metric! (using
openai/gpt-oss-120b, strict=False, async_mode=True)...
Evaluating 5 test case(s) in parallel ━━━━━╸ 20% 0:00:43
🎯 Evaluating test case #0 0% 0:00:43
🎯 Evaluating test case #1 ━━━━━━━━━━━━━━╸ 50% 0:00:43
🎯 Evaluating test case #2 ━━━━━━━━━━━━━━╸ 50% 0:00:43
🎯 Evaluating test case #3 ━━━━━━━━━━━━━━╸ 50% 0:00:43
🎯 Evaluating test case #4 ━━━━━━━━━━━━━━╸ 50% 0:00:43
status: 429
tokens per minute (TPM): Limit 8000, Used 7177, Requested 1007The progress bar shows Evaluating 5 test case(s) in parallel, so all five test cases and their ten metrics started together. At 20% the judge asked for 1,007 more tokens in a minute that had already used 7,177 of its 8,000, the client's retries did not get it through, and the RateLimitError stopped evaluate(). No report and no scores. Your run can fail at a different point, or not at all, depending on what else used the key that minute.
The AsyncConfig and ErrorConfig API
from deepeval.evaluate import AsyncConfig, ErrorConfig
AsyncConfig(
run_async=True, # False: one metric at a time, nothing in parallel
max_concurrent=1, # how many test cases are scored at the same time
throttle_value=10, # seconds to wait after starting each test case
)
ErrorConfig(
ignore_errors=True, # a failed metric is recorded, the run goes on
skip_on_missing_params=True, # skip a metric when its test case lacks a field
)Fewer test cases at a time
max_concurrent caps how many test cases are in flight. The metrics of one test case still run together, so one test case at a time means two metrics in flight at once here, not the ten of the run above. run_async=False goes further and runs every metric one after the other; it is the gentlest of the no-pause options, and it can still finish sooner than a run with long throttle_value pauses.
A pause between test cases
throttle_value is the number of seconds evaluate() waits after it starts each test case, before it starts the next. It only applies when run_async is on. The docs call a combination of throttle_value and max_concurrent the best way to handle rate limits.
Retries inside the judge
When a call still gets a 429, the judge's own client retries it. judge.py from Custom judge model builds its clients with max_retries=8, so each call waits and tries again up to eight times before the error reaches DeepEval. DeepEval's own retry settings, the DEEPEVAL_RETRY_* variables, drive the retries of its built-in model classes such as OpenAIModel; a custom judge retries however its client does.
Keeping the run alive
ErrorConfig(ignore_errors=True) catches an error from any metric, records it on that test case and moves on, so one call that fails after all its retries costs one score instead of the whole run.
Running three goldens inside the limit
The same metrics with both settings, on three of the goldens to keep the run short.
from deepeval import evaluate
from deepeval.evaluate import AsyncConfig, ErrorConfig
from technest_eval import build_test_cases, correctness, dataset, relevancy
test_cases = build_test_cases(dataset.goldens[:3]) # g001, g002, g003
evaluate(
test_cases=test_cases,
metrics=[relevancy, correctness],
async_config=AsyncConfig(max_concurrent=1, throttle_value=10),
error_config=ErrorConfig(ignore_errors=True),
)✨ You're running DeepEval's latest Answer Relevancy Metric! (using openai/gpt-oss-120b, strict=False, async_mode=True)... ✨ You're running DeepEval's latest Correctness [GEval] Metric! (using openai/gpt-oss-120b, strict=False, async_mode=True)... ╭──────────────────────────────────────────────────────────────────────────────╮ │ 🚀 DeepEval Evaluation Results │ ╰──────────────────────────────────────────────────────────────────────────────╯ ╭──────────────────────────────────────────────────────────────────────────────╮ │ ✅ g001 (Passed 2 metrics) │ ╰──────────────────────────────────────────────────────────────────────────────╯ ╭──────────────────────────────────────────────────────────────────────────────╮ │ ✅ g002 (Passed 2 metrics) │ ╰──────────────────────────────────────────────────────────────────────────────╯ ╭──────────────────────────────────────────────────────────────────────────────╮ │ │ │ ❌ g003 │ │ ├── Input: How long is the battery life on the SoundPods │ │ │ Pro? │ │ │ Actual Output: The TechNest SoundPods Pro offer 8 hours of │ │ │ playback on a single charge. When used with │ │ │ the charging case, the total battery life │ │ │ extends to 24 hours. │ │ │ Expected Output: The SoundPods Pro offer 8 hours of playback │ │ │ per charge and an additional 24 hours from the │ │ │ charging case, giving a total of 32 hours. │ │ └── Metrics │ │ Status ┃ Metric ┃ Score ┃ Threshold ┃ Reason │ │ ━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━╇━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━ │ │ PASS │ Answer Relevancy │ 1.00 │ 0.50 │ The score is 1.00 │ │ │ │ │ │ because the answer │ │ │ │ │ │ directly a... │ │ FAIL │ Correctness │ 0.00 │ 0.50 │ The actual output │ │ │ [GEval] │ │ │ states the total │ │ │ │ │ │ battery life with │ │ │ │ │ │ the case is 24 │ │ │ │ │ │ hours, │ │ │ │ │ │ contradicting the │ │ │ │ │ │ expected total of │ │ │ │ │ │ 32 hours, and it │ │ │ │ │ │ omits the expected │ │ │ │ │ │ 32‑hour total, so │ │ │ │ │ │ it fails the │ │ │ │ │ │ evaluation │ │ │ │ │ │ criteria. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────╯ ╭──────────────────────────────────────────────────────────────────────────────╮ │ Aggregate Metrics │ │ │ │ Metric ┃ Average Score ┃ Pass Rate ┃ Total │ │ ━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━ │ │ Answer Relevancy │ 1.00 │ 100.00% | passed=3 | │ 3 │ │ │ │ failed=0 │ │ │ Correctness [GEval] │ 0.67 │ 66.67% | passed=2 | failed=1 │ 3 │ ╰──────────────────────────────────────────────────────────────────────────────╯ ⚠ WARNING: No hyperparameters logged. » Log hyperparameters to attribute prompts and models to your test runs. ================================================================================ ✓ Evaluation completed 🎉! (time taken: 50.29s | token cost: None) » Test Results (3 total tests): » Pass Rate: 66.67% | Passed: 2 | Failed: 1 =============================================================================== = » Want to share evals with your team, or a place for your test cases to live? ❤️ 🏡 » Run 'deepeval view' to analyze and save testing results on Confident AI.
What changed in the run
- No rate-limit error stopped the run. All three test cases were scored and the report printed in full.
- g001 and g002 passed both metrics. g003 failed correctness for the same reason as in the evaluate() lesson: the bot says 24 hours in total where the catalog adds up to 32.
- The time taken was 50.29 seconds.
throttle_value=10spaces out when each test case starts. The pause overlaps with scoring, so it adds time only when scoring a case is faster than the gap. - No metric errored, so
ignore_errors=Truehad nothing to catch in this run. Its job is the run where a call fails after all its retries.
A 429 that names tokens per day (TPD) is different: the model's daily budget is spent, and no retry or pause helps. Wait for it to reset, or switch the judge with export JUDGE_MODEL=openai/gpt-oss-20b, which has its own daily budget.
The time budget per test case
Retries take time, and DeepEval gives each test case's metrics a time budget that includes them. When the budget runs out, the metric fails with Timed out/cancelled while evaluating metric. The budget is the setting DEEPEVAL_PER_TASK_TIMEOUT_SECONDS; set DEEPEVAL_PER_TASK_TIMEOUT_SECONDS_OVERRIDE in the environment to change it. Read the values your install uses:
from deepeval.config.settings import get_settings
settings = get_settings()
print("per task:", settings.DEEPEVAL_PER_TASK_TIMEOUT_SECONDS, "seconds")
print("per attempt:", settings.DEEPEVAL_PER_ATTEMPT_TIMEOUT_SECONDS, "seconds")per task: 180 seconds per attempt: 88.5 seconds
In this install a test case's metric gets 180 seconds, and each call of one of DeepEval's built-in models gets 88.5 of them. Both are computed from the retry settings unless you override them, and they can change between versions, so read them as above instead of assuming a number. If a slow judge on a busy key times out, raise the per-task override, for example export DEEPEVAL_PER_TASK_TIMEOUT_SECONDS_OVERRIDE=300.
Cooldowns vs DeepEval's settings
| Sleep calls (the video's app) | AsyncConfig | Client retries (max_retries) | ErrorConfig(ignore_errors=True) | |
|---|---|---|---|---|
| What it does | Waits a fixed time between samples | Caps parallel test cases and spaces them out | Waits and repeats a call that got a 429 | Records a failed metric and carries on |
| Prevents 429s | Yes, if the pause is long enough | Makes them much rarer | No, it recovers from them | No, it limits the damage |
| Cost | Slower runs, even when there is room | Slower runs | Slower calls during a burst | A missing score |
| Where it lives | Your loop | evaluate() | judge.py | evaluate() |
When to change these settings
- On a free or low tier: start with
max_concurrent=1, and addthrottle_valueif 429s still show up. - On a paid tier with high limits: raise
max_concurrentto finish a large dataset faster. - For a long nightly run: add
ignore_errors=Trueso one bad call does not throw away an hour of scores, and check the report for errored metrics.
ignore_errors=True hides failures as well as rate limits. A metric that errors on every test case, because of a wrong model name or a missing field, still finishes the run, with no score. Read the report, or the result object in EvaluationResult, before you trust a pass rate.Related
- Previous: evaluate()
- Next: EvaluationResult
- Reference: Flags and configs
- Run the settings example after
export DEEPEVAL_PER_TASK_TIMEOUT_SECONDS_OVERRIDE=300: the per-task budget prints 300.0 and the per-attempt time grows with it. - Run it again after
export DEEPEVAL_RETRY_MAX_ATTEMPTS=4(unset the first variable): the same 180 seconds are now split across four attempts, so each attempt gets less time. - In the run inside the limit, replace the
AsyncConfigwithAsyncConfig(run_async=False): the metrics run one after another and the run still finishes with the same three test cases.
You understood something today that you didn't yesterday.