Rate limits and cooldowns
A rate limit is a cap a model provider puts on how many requests or tokens you may send per minute and per day; a cooldown is a pause between evaluation calls that keeps a run under that cap.
Last updated: 29 Sep, 2026 · RAGAS 0.4.3
An evaluation is many judge calls in a row. On a free tier, sending them all at once gets you 429 Too Many Requests errors instead of scores. The video's app handles this with two cooldowns, and this lesson writes them.
Counting the judge calls
The video counts. With five goldens and five metrics, each metric scores each golden: 25 judge calls. Hit Groq's free tier with all 25 at once and it answers with rate-limit errors.
One metric per call, with a pause
Someone in the class asks why not put everything in one call. The video's answer: one call per metric, because a judge grades better when it has one job and a mostly empty context window, the way a person handles one or two projects better than ten. So the calls stay separate, and the app spaces them out. After each sample it waits 25 seconds, the sample cooldown; after each metric's full pass it waits 35 seconds, the experiment cooldown. It is a plain sleep that lets the provider's per-minute window recover.
One more detail the count leaves out: most RAGAS metrics make more than one judge call per sample. Faithfulness asks once for the claims and once for the verdicts; answer relevancy asks once per generated question. The real number of calls is higher than 25, which makes the pauses more important, not less.
The cooldown and retry API
await asyncio.sleep(SAMPLE_COOLDOWN) # pause between samples
try:
... # one judge call
except Exception as error: # a 429 names the rate limit in its message
await asyncio.sleep(RETRY_WAIT) # wait for the window to clear, then retry onceScoring one sample with a retry
This is the video's _score_one, shortened. A rate-limit error waits and retries once; any other error is raised, so it is never hidden.
async def score_one(metric, inputs):
try:
return (await metric.ascore(**inputs)).value
except Exception as error:
if "429" not in str(error) and "rate" not in str(error).lower():
raise
print(f" rate limit hit, waiting {RETRY_WAIT}s then retrying once")
await asyncio.sleep(RETRY_WAIT)
return (await metric.ascore(**inputs)).valueScoring the samples one at a time
async def score_experiment(metric, rows):
scores = []
for i, inputs in enumerate(rows):
scores.append(await score_one(metric, inputs))
print(f"{i + 1}/{len(rows)} scored: {scores[-1]}")
if i < len(rows) - 1:
await asyncio.sleep(SAMPLE_COOLDOWN)
return scoresOne sample at a time, with the cooldown between them and none after the last. The video's score_experiment does the same.
- written in TechNest RAG app
- written in TechNest RAG app
- written in LLM as a judge
- written in Goldens
View the code here
import json
import os
import numpy as np
from google import genai
from openai import OpenAI
gemini = genai.Client() # reads GOOGLE_API_KEY
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
EMBED_MODEL = "gemini-embedding-2"
CHAT_MODEL = "qwen/qwen3.8-27b"
SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""
with open("catalog.json", encoding="utf-8") as f:
CATALOG = json.load(f)
def embed(texts):
# gemini-embedding-2 turns everything in one call into one embedding, so send one text per call
vectors = np.array([gemini.models.embed_content(model=EMBED_MODEL, contents=t).embeddings[0].values for t in texts])
return vectors / np.linalg.norm(vectors, axis=1, keepdims=True)
DOC_VECTORS = embed([f"{item['title']}. {item['content']}" for item in CATALOG])
def retrieve(question, top_k=3):
scores = DOC_VECTORS @ embed([question])[0]
best = np.argsort(scores)[::-1][:top_k]
return [CATALOG[i]["content"] for i in best]
def generate(question, contexts):
context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
]
response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
return response.choices[0].message.content.strip()
def answer(question, top_k=3):
contexts = retrieve(question, top_k)
return generate(question, contexts), contexts
[
{
"id": "prod_001",
"category": "product",
"title": "ProBook X1 Laptop",
"content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
},
{
"id": "prod_002",
"category": "product",
"title": "PixelPhone 15",
"content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
},
{
"id": "prod_003",
"category": "product",
"title": "SoundPods Pro",
"content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
},
{
"id": "prod_004",
"category": "product",
"title": "UltraTab S2 Tablet",
"content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
},
{
"id": "prod_005",
"category": "product",
"title": "SmartWatch X",
"content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
},
{
"id": "prod_006",
"category": "product",
"title": "ProCam 4K Action Camera",
"content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
},
{
"id": "prod_007",
"category": "product",
"title": "BassBuds Max Headphones",
"content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
},
{
"id": "prod_008",
"category": "product",
"title": "SoundBar 360",
"content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
},
{
"id": "policy_001",
"category": "policy",
"title": "Return Policy",
"content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
},
{
"id": "policy_002",
"category": "policy",
"title": "Shipping Policy",
"content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
},
{
"id": "policy_003",
"category": "policy",
"title": "Warranty Policy",
"content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
},
{
"id": "policy_004",
"category": "policy",
"title": "Payment Policy",
"content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
},
{
"id": "faq_001",
"category": "faq",
"title": "Order Tracking",
"content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
},
{
"id": "faq_002",
"category": "faq",
"title": "International Shipping",
"content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
},
{
"id": "faq_003",
"category": "faq",
"title": "Bulk and Business Orders",
"content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
}
]
import os
from google import genai
from openai import AsyncOpenAI
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory
groq = AsyncOpenAI(
api_key=os.environ.get("JUDGE_GROQ", os.environ["GROQ_API_KEY"]),
base_url="https://api.groq.com/openai/v1",
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=groq)
class OneTextPerCall(GoogleEmbeddings):
"""gemini-embedding-2 turns a list into one embedding, so embed each text on its own."""
def embed_texts(self, texts, **kwargs):
return [self.embed_text(text) for text in texts]
async def aembed_texts(self, texts, **kwargs):
return [await self.aembed_text(text) for text in texts]
embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")
[
{
"id": "g001",
"metric_focus": "faithfulness",
"user_input": "What is TechNest's return policy?",
"reference": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."
},
{
"id": "g002",
"metric_focus": "answer_relevancy",
"user_input": "What are the RAM and storage specs of the ProBook X1?",
"reference": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."
},
{
"id": "g003",
"metric_focus": "context_precision",
"user_input": "How long is the battery life on the SoundPods Pro?",
"reference": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours."
},
{
"id": "g004",
"metric_focus": "context_recall",
"user_input": "What are TechNest's shipping options and how long do returns take to process?",
"reference": "TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received."
},
{
"id": "g005",
"metric_focus": "answer_correctness",
"user_input": "What is the price of the PixelPhone 15?",
"reference": "The TechNest PixelPhone 15 is priced at $899."
}
]
Faithfulness over three goldens with cooldowns
The app answers each golden (phase one), then faithfulness scores the answers one by one (phase two of one metric). Save it as cooldown.py next to the other files. It scores the first three goldens, enough to see the pacing; it takes about a minute, most of it the two pauses.
import asyncio
import json
from judge import judge
from ragas.metrics.collections import Faithfulness
from technest import answer
SAMPLE_COOLDOWN = 25 # seconds between samples, as in the video
RETRY_WAIT = 65 # seconds to wait after a 429 before one retry
async def score_one(metric, inputs):
try:
return (await metric.ascore(**inputs)).value
except Exception as error:
if "429" not in str(error) and "rate" not in str(error).lower():
raise
print(f" rate limit hit, waiting {RETRY_WAIT}s then retrying once")
await asyncio.sleep(RETRY_WAIT)
return (await metric.ascore(**inputs)).value
async def score_experiment(metric, rows):
scores = []
for i, inputs in enumerate(rows):
scores.append(await score_one(metric, inputs))
print(f"{i + 1}/{len(rows)} scored: {scores[-1]}")
if i < len(rows) - 1:
await asyncio.sleep(SAMPLE_COOLDOWN)
return scores
goldens = json.load(open("goldens.json", encoding="utf-8"))[:3] # three goldens keep the demo short
rows = []
for golden in goldens:
response, contexts = answer(golden["user_input"])
rows.append({"user_input": golden["user_input"], "response": response, "retrieved_contexts": contexts})
scores = asyncio.run(score_experiment(Faithfulness(llm=judge), rows))
for golden, score in zip(goldens, scores):
print(golden["id"], "faithfulness", round(score, 2))1/3 scored: 1.0 2/3 scored: 1.0 3/3 scored: 1.0 g001 faithfulness 1.0 g002 faithfulness 1.0 g003 faithfulness 1.0
Reading the paced run
- The progress lines arrive about 25 seconds apart, the sample cooldown between them.
- The scores are faithfulness for each golden's real answer, so they change when the app's answers change.
- No rate-limit line appears if the pauses were long enough; if one does, the retry waited and scored that sample instead of losing it.
All at once vs paced
| All calls at once | One sample at a time with cooldowns | |
|---|---|---|
| Speed | Fast until the first 429 | Slow and steady |
| On a free tier | Rate-limit errors, lost scores | Stays under the per-minute cap |
| What the video's app uses | Yes: 25 s per sample, 35 s per metric |
When you need cooldowns
- On any free or low tier, where the per-minute token cap is small next to an evaluation's calls.
- When one key serves both the app and the judge: a separate judge key,
JUDGE_GROQ, keeps the app working while evaluations run.
Related
- Previous: RAGAS metrics
- Next: Faithfulness
- Reference: Groq rate limits
- Set
SAMPLE_COOLDOWN = 0, run it again, and watch whether a rate-limit line appears. - Print how long the whole run took with
time.perf_counter()before and afterasyncio.run.
Every expert started right here.