The study compares agents' completion claims with environment state. It finds that confident language can conceal failed tasks and that language-model judges struggle to detect this reliably.
Where the evidence stops
Findings are benchmark-specific; the suggested detectors need domain calibration and do not replace state verification.
Put it into practice
Require an observable outcome and attach its evidence before a run can be marked successful.
A multi-dimensional evaluation separates task accuracy from repeatability, robustness, failure prediction and harmful outcomes. The revised study evaluates 15 models across two benchmarks.
Where the evidence stops
The evaluated tasks cover only two environments; results depend on scaffold and metric version.
Put it into practice
Record repeated trials, prompt perturbations, resource variance and policy violations alongside task success.
An audit of linked datasets, models and applications finds frequent gaps between permissive license labels and the accompanying license or attribution documents.
Where the evidence stops
Automated document evidence does not determine legal permission for every artifact or jurisdiction.
Put it into practice
Separate a publisher's license label from observed license files, notice files and upstream lineage; surface unknowns for review.
A simulated telecom environment allows both the agent and user to act. Its tasks expose failures in communication and coordination that single-actor tests can miss.
Where the evidence stops
Simulated users and bounded domains cannot establish reliability for every real customer workflow.
Put it into practice
Add workflow evaluation plans that capture tool schemas, policy rules, user simulator, final state and repeated runs.
The authors release openly licensed and public-domain pretraining text, source-processing code, mixtures and checkpoints, demonstrating a practical path to documented training data.
Where the evidence stops
The experiments concern particular 7B models and data mixtures; they do not certify every downstream reuse.
Put it into practice
Link datasets, source terms and training-mixture evidence to model records instead of reducing openness to weight availability.
Learned routing selects between stronger and cheaper models using preference data, demonstrating that model choice can be optimized for a quality-cost trade-off.
Where the evidence stops
The paper studies two-model routing; distribution shifts and latency requirements still need workload-specific validation.
Put it into practice
Build a routing experiment that compares each model and the proposed router on the same workload, budget and quality floor.
This library includes preprints and published research. Summaries and product implications are AltOpenAI interpretations. Checking a source does not independently reproduce its results or certify a model for your use case.