The benchmark runner behind most published model scores. Built by EleutherAI.
Others in evals and testing
Common questions
Is lm-evaluation-harness free to use?
lm-evaluation-harness is open source and free to run yourself. You still pay whichever model provider you point it at, and this tutorial is free with no signup.
Do I need to know Python to use lm-evaluation-harness?
Basic Python is enough. Functions, dictionaries and imports cover most of what lm-evaluation-harness asks of you.
When should I not use lm-evaluation-harness?
When something else in evals and testing fits the job better. The others in that group are listed below.
Lessons for lm-evaluation-harness are being written. Meanwhile the APIs for AI tutorial covers the same ground: state, tools, loops, memory and human approval. Most of it carries straight over.
Start the APIs for AI tutorial →