THE OPEN INTELLIGENCE HUB COMMUNITY PREVIEW ↗
EVIDENCE, NOT A UNIVERSAL SCORE Choose models with proof. Open model benchmarks are only useful when their methods, revisions and limits travel with the result.
ALTOPENAI EVALUATION STANDARD Make every comparison reproducible. We surface benchmark sources and treat an upstream score as evidence, not a verdict. Compare like for like: task, model revision, precision, prompt settings, data version and hardware all matter.
Result evidence record Model + data revision · harness · configuration · precision · hardware · logs All Reasoning Knowledge Code Audio Document AI Vision
Reasoning Source-linked
GSM8K Grade-school multi-step mathematics questions for reasoning evaluations.
Use for Math reasoning candidates.
Compare only when The same split, prompting method, answer extraction and model revision are used. Open source & methodology Knowledge Source-linked
MMLU-Pro A challenging multi-domain knowledge benchmark with multiple-choice questions.
Use for Broad academic and professional knowledge candidates.
Compare only when The same task subset, few-shot setting, prompt and dataset revision are used. Open source & methodology Code Source-linked
SWE-bench Verified Human-validated GitHub issue-to-pull-request tasks scored through repository test resolution.
Use for Software-engineering agents and code-repair systems.
Compare only when The benchmark revision, sandbox, agent tools, model settings and verifier version are held constant. Open source & methodology Audio Source-linked
Open ASR Leaderboard Speech-recognition evaluation across short, long and multilingual audio conditions.
Use for Automatic-speech-recognition models.
Compare only when Language, audio corpus, decoding configuration, speed measurement and error metric are the same. Open source & methodology Document AI Source-linked
olmOCR-bench Document OCR to Markdown tasks assessed with a large set of unit tests.
Use for Document understanding and OCR systems.
Compare only when The same source documents, rendering pipeline, parser version and test suite are used. Open source & methodology Vision Source-linked
Video-MME v2 Video-understanding evaluation for multimodal models.
Use for Video and multimodal reasoning candidates.
Compare only when The same clips, frames, visual preprocessing, prompts and model revision are used. Open source & methodology READ RESULTS WITH CONTEXT
What AltOpenAI will label clearly. 01 Method attached The record includes artifacts and configuration to inspect. Independent execution and verification remain separate requirements.
02 Publisher-reported The publisher reports a result; inspect their model card and methodology.
03 Community-submitted Useful signal, but independently reproduce before a production decision.
04 Not comparable Different model type, revision, precision or task means scores must not be ranked together.
Find models to evaluate