← All frameworks

HELM

Stanford's holistic evaluation: many scenarios and metrics, not one leaderboard number.

Evals and testing

Stanford's holistic evaluation: many scenarios and metrics, not one leaderboard number. Built by Stanford CRFM.

Others in evals and testing

Common questions

Is HELM free to use?

HELM is open source and free to run yourself. You still pay whichever model provider you point it at, and this tutorial is free with no signup.

Do I need to know Python to use HELM?

Basic Python is enough. Functions, dictionaries and imports cover most of what HELM asks of you.

When should I not use HELM?

When something else in evals and testing fits the job better. The others in that group are listed below.

Lessons for HELM are being written. Meanwhile the APIs for AI tutorial covers the same ground: state, tools, loops, memory and human approval. Most of it carries straight over.

Start the APIs for AI tutorial →