Skip to content
Trust & transparency A
THE OPEN INTELLIGENCE HUBCOMMUNITY PREVIEW
FRONTIER RESEARCH

Explore open AI research.

Read selected papers on evaluation, agents, routing and provenance. Bring useful ideas into a plan you can test on your own workload.

11 selected papersPrimary arXiv sourcesSources checked Curated library · not a live feed

11 papers

Agents·v1

From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents

Laksh Advani

The study compares agents' completion claims with environment state. It finds that confident language can conceal failed tasks and that language-model judges struggle to detect this reliably.

Where the evidence stops

Findings are benchmark-specific; the suggested detectors need domain calibration and do not replace state verification.

Put it into practice

Require an observable outcome and attach its evidence before a run can be marked successful.

Verified task successFalse completion rateEvidence coverage
Source & review details
arXiv
2606.09863v1
First submitted
Version date
Source reports
FAGEN at ICML 2026, as reported on arXiv
Reviewed
Limitation note
Scope and application caution
Open benchmark project
Agents·v3

Towards a Science of AI Agent Reliability

Stephan Rabanser et al.

A multi-dimensional evaluation separates task accuracy from repeatability, robustness, failure prediction and harmful outcomes. The revised study evaluates 15 models across two benchmarks.

Where the evidence stops

The evaluated tasks cover only two environments; results depend on scaffold and metric version.

Put it into practice

Record repeated trials, prompt perturbations, resource variance and policy violations alongside task success.

Repeated-trial successPrompt robustnessPolicy violations
Source & review details
arXiv
2602.16666v3
First submitted
Version date
Source reports
ICML 2026, as reported on arXiv
Reviewed
Limitation note
Scope and author revision note
Open benchmark project
Provenance·v1

Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity

James Jewitt et al.

An audit of linked datasets, models and applications finds frequent gaps between permissive license labels and the accompanying license or attribution documents.

Where the evidence stops

Automated document evidence does not determine legal permission for every artifact or jurisdiction.

Put it into practice

Separate a publisher's license label from observed license files, notice files and upstream lineage; surface unknowns for review.

License file presenceNotice presenceUpstream lineage coverage
Source & review details
arXiv
2602.08816v1
First submitted
Version date
Source reports
arXiv preprint
Reviewed
Limitation note
Application caution
Agents·v1

τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Victor Barres et al.

A simulated telecom environment allows both the agent and user to act. Its tasks expose failures in communication and coordination that single-actor tests can miss.

Where the evidence stops

Simulated users and bounded domains cannot establish reliability for every real customer workflow.

Put it into practice

Add workflow evaluation plans that capture tool schemas, policy rules, user simulator, final state and repeated runs.

Task successPolicy adherenceTool errors
Source & review details
arXiv
2506.07982v1
First submitted
Version date
Source reports
arXiv research paper
Reviewed
Limitation note
Scope inference
Open benchmark project
Provenance·v1

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

Nikhil Kandpal et al.

The authors release openly licensed and public-domain pretraining text, source-processing code, mixtures and checkpoints, demonstrating a practical path to documented training data.

Where the evidence stops

The experiments concern particular 7B models and data mixtures; they do not certify every downstream reuse.

Put it into practice

Link datasets, source terms and training-mixture evidence to model records instead of reducing openness to weight availability.

Source documentation coverageLicense evidence coverageDataset revision
Source & review details
arXiv
2506.05209v1
First submitted
Version date
Source reports
arXiv research paper
Reviewed
Limitation note
Scope inference
Open research artifact
Evaluation·v2

The Leaderboard Illusion

Shivalika Singh et al.

An analysis of Chatbot Arena argues that selective disclosure and unequal evaluation access can distort apparent model rankings.

Where the evidence stops

The analysis addresses a particular evaluation system and historical period; it does not invalidate every preference comparison.

Put it into practice

Display score provenance, sample counts, submission history and evaluation conditions; keep incompatible score families separate.

Sample countConfidence intervalProtocol completeness
Source & review details
arXiv
2504.20879v2
First submitted
Version date
Source reports
arXiv research paper
Reviewed
Limitation note
Scope inference
Provenance·v2

Atlas: A Framework for ML Lifecycle Provenance & Transparency

Marcin Spoczynski, Marcela S. Melara and Sebastian Szyller

Atlas prototypes verifiable lineage across the model lifecycle using artifact measurements, attestations and transparency infrastructure.

Where the evidence stops

The authors identify hardware trust boundaries and unresolved trade-offs between disclosure and confidentiality.

Put it into practice

Start with revision-pinned provenance manifests; add signature verification and independently verifiable attestations as infrastructure matures.

Pinned artifact identityLineage completenessAttestation verification
Source & review details
arXiv
2502.19567v2
First submitted
Version date
Source reports
arXiv research paper
Reviewed
Limitation note
Author-stated limitation
Evaluation·v2

LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Colin White et al.

LiveBench refreshes tasks from recent sources and uses objective scoring to reduce exposure to stale test sets and subjective judging.

Where the evidence stops

Refreshes limit contamination risk rather than proving its absence; comparisons require the same benchmark release.

Put it into practice

Require benchmark release, dataset revision and evaluation date on every imported score; flag cross-release comparisons.

Release-matched scoreBenchmark release dateObjective scorer version
Source & review details
arXiv
2406.19314v2
First submitted
Version date
Source reports
ICLR 2025 Spotlight, as reported on arXiv
Reviewed
Limitation note
Scope inference consistent with revised title
Open benchmark project
Routing·v4

RouteLLM: Learning to Route LLMs with Preference Data

Isaac Ong et al.

Learned routing selects between stronger and cheaper models using preference data, demonstrating that model choice can be optimized for a quality-cost trade-off.

Where the evidence stops

The paper studies two-model routing; distribution shifts and latency requirements still need workload-specific validation.

Put it into practice

Build a routing experiment that compares each model and the proposed router on the same workload, budget and quality floor.

Cost per successful taskQuality floorP95 latency
Source & review details
arXiv
2406.18665v4
First submitted
Version date
Source reports
arXiv research paper
Reviewed
Limitation note
Author-stated limitation
Open research artifact
Evaluation·v2

Holistic Evaluation of Language Models

Percy Liang et al.

HELM evaluates models across shared scenarios and multiple metrics, exposing trade-offs that a single accuracy leaderboard can hide.

Where the evidence stops

Its scenario coverage is deliberately bounded, and historical scores do not describe today's deployments.

Put it into practice

Create a task-specific scorecard with quality, robustness, safety and efficiency metrics plus explicit missing evidence.

Task accuracyRobustnessSafety evaluation
Source & review details
arXiv
2211.09110v2
First submitted
Version date
Source reports
Foundational research paper
Reviewed
Limitation note
Scope inference
Open benchmark project
Provenance·v2

Model Cards for Model Reporting

Margaret Mitchell et al.

Model cards organize intended use, evaluation conditions and known limitations into documentation that supports responsible model selection.

Where the evidence stops

A completed document is self-reported evidence, not independent verification of the model or its suitability.

Put it into practice

Add a model passport that records source, revision, intended use, limitations, evaluations and evidence ownership.

Intended-use coverageEvaluation provenanceLimitation disclosure
Source & review details
arXiv
1810.03993v2
First submitted
Version date
Source reports
FAT* 2019, as reported on arXiv
Reviewed
Limitation note
Application caution

This library includes preprints and published research. Summaries and product implications are AltOpenAI interpretations. Checking a source does not independently reproduce its results or certify a model for your use case.