Skip to content
Trust & transparency A
THE OPEN INTELLIGENCE HUBCOMMUNITY PREVIEW
EVIDENCE, NOT A UNIVERSAL SCORE

Choose models with proof.

Open model benchmarks are only useful when their methods, revisions and limits travel with the result.

ALTOPENAI EVALUATION STANDARD

Make every comparison reproducible.

We surface benchmark sources and treat an upstream score as evidence, not a verdict. Compare like for like: task, model revision, precision, prompt settings, data version and hardware all matter.

Result evidence recordModel + data revision · harness · configuration · precision · hardware · logs

Benchmark registry 6

Source-linked benchmark records to investigate before you promote a model.

How evaluation results work
ReasoningSource-linked

GSM8K

Grade-school multi-step mathematics questions for reasoning evaluations.

Use for
Math reasoning candidates.
Compare only when
The same split, prompting method, answer extraction and model revision are used.
Open source & methodology
KnowledgeSource-linked

MMLU-Pro

A challenging multi-domain knowledge benchmark with multiple-choice questions.

Use for
Broad academic and professional knowledge candidates.
Compare only when
The same task subset, few-shot setting, prompt and dataset revision are used.
Open source & methodology
CodeSource-linked

SWE-bench Verified

Human-validated GitHub issue-to-pull-request tasks scored through repository test resolution.

Use for
Software-engineering agents and code-repair systems.
Compare only when
The benchmark revision, sandbox, agent tools, model settings and verifier version are held constant.
Open source & methodology
AudioSource-linked

Open ASR Leaderboard

Speech-recognition evaluation across short, long and multilingual audio conditions.

Use for
Automatic-speech-recognition models.
Compare only when
Language, audio corpus, decoding configuration, speed measurement and error metric are the same.
Open source & methodology
Document AISource-linked

olmOCR-bench

Document OCR to Markdown tasks assessed with a large set of unit tests.

Use for
Document understanding and OCR systems.
Compare only when
The same source documents, rendering pipeline, parser version and test suite are used.
Open source & methodology
VisionSource-linked

Video-MME v2

Video-understanding evaluation for multimodal models.

Use for
Video and multimodal reasoning candidates.
Compare only when
The same clips, frames, visual preprocessing, prompts and model revision are used.
Open source & methodology
READ RESULTS WITH CONTEXT

What AltOpenAI will label clearly.

01

Method attached

The record includes artifacts and configuration to inspect. Independent execution and verification remain separate requirements.

02

Publisher-reported

The publisher reports a result; inspect their model card and methodology.

03

Community-submitted

Useful signal, but independently reproduce before a production decision.

04

Not comparable

Different model type, revision, precision or task means scores must not be ranked together.