Release gate: an eval suite that blocks bad releases
An evaluation gate that runs on every change and turns the build red when answer quality drops below a line.
The problem
Someone edits a prompt or bumps a model version, it looks fine in a quick manual check, and it ships. A week later answers are quietly worse on the cases nobody re-tested. There was no gate, so a regression walked straight through.
You want an evaluation suite that runs on every change, scores the answers against known-good cases, and fails the build when quality drops below a line. The point is not a dashboard; it is a red build that blocks a bad change.
Architecture
A gate in the pipeline. A change triggers an eval suite over a fixed dataset, the answers are scored, and a threshold turns the scores into pass or fail that CI enforces. The dataset and the threshold are the real work; the model that grades is just a tool.
The gate is only as honest as its dataset. Build the set of cases from real failures and answers you trust, and keep it under version control with the code, so the gate tightens as you learn what breaks.
What it draws on
Everything here comes from this stage of the roadmap; the project is where those courses meet.
- Topic: golden datasets
- DeepEval or Ragas: scoring answers with a judge
- Promptfoo: regression suites in CI
- OpenAI Agents SDK: tracing and streaming
What done looks like
| Requirement | Done when |
|---|---|
| Golden set | Built from real questions, including ones that must be refused |
| Scoring | Runs without a person, with the scoring method written down |
| Gate | A deliberately broken prompt fails the check |
| Report | Shows the score over time and the worst failures |
Where to start
Write ten cases you would be embarrassed to get wrong, with the answer you expect, before you touch a scorer. Wire those into a suite that returns a number, set a threshold, and make a failing case turn the build red. Grow the dataset every time something slips through — that is the whole discipline.
Every expert started right here.