1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you evaluate code generation? Explain pass@k.
30-second answerSay your answer out loud first, then reveal.
Metrics
| Metric | Meaning |
|---|---|
| pass@1 | Chance a single sample is correct (most relevant for users) |
| pass@k | Chance any of k samples is correct (relevant when you can test and pick) |
| Compile / lint rate | Syntactic and style validity |
| Test pass rate | Fraction of tests passed (partial credit) |
| Security findings | Static analysis (e.g. injection, hard-coded secrets) |
| Edit acceptance | Real-world: suggestions accepted and retained |
Worked example. n = 10 samples, c = 3 correct.
pass@1 = 1 − C(7,1)/C(10,1) = 1 − 7/10 = 0.3
pass@5 = 1 − C(7,5)/C(10,5) = 1 − 21/252 ≈ 0.92Practical considerations
- Sandboxing: generated code is untrusted. Isolate execution with no network, resource limits and timeouts.
- Test quality: weak tests inflate scores; use hidden tests and property-based tests.
- Contamination: public benchmarks may be in training data, so build internal tasks from your codebase.
Related
Every expert started right here.