Research · Benchmarks

Benchmarks that keep progress honest

We publish evaluations so anyone can see what a model actually understands, and so the next step toward useful, human-level intelligence is measured fairly. Browse current programmes below.

How we approach benchmarks

Advancing HLI means knowing what works. Our evaluations favour clarity over hype: controlled protocols, real-world sectors, and results that help builders choose and improve models people can actually run.

Measure real understanding

Fluency is not enough. We evaluate whether models handle the vocabulary, reasoning, and tasks that show up in real sectors of life.

Fair and reproducible

Same items, same prompts, same scoring, so rankings stay trustworthy as architectures and training recipes change.

Guide efficient progress

Benchmarks that help compact, accessible models improve, aligned with intelligence that runs where compute is modest.

Expand with impact

Start where signal is clear, then grow coverage across domains that matter for education, health, livelihoods, and beyond.

Current benchmarks

Open a programme for methods, leaderboards, and resources.