Research · Data

Data that advances intelligence for everyone

We build datasets that make machine intelligence measurable, trainable, and useful, especially where data and compute are scarce. Current releases sit below; the vision is broader than any single programme.

How we approach data

Our research goals call for intelligence that works with minimal data, remains accessible everywhere, and serves people in their own languages and contexts. That is the standard we hold every release to.

Small data, strong signal

Toward computationally efficient intelligence: learn and evaluate with less, so progress does not depend on unlimited corpora or compute.

Local and useful

Data grounded in languages, cultures, and settings people live in, so models serve real work, not only abstract multilingual averages.

Shared yardsticks

Clear sourcing, review, and scoring rules so researchers can compare fairly as we push toward human-level capabilities.

Open for impact

Public previews, reports, and access paths where they help others reproduce results and build systems that benefit more people.

Current datasets

Open a release for methods, coverage, and access links.