SIMJECTURE BENCHMARK RESULTS
Model leaderboard
Measured scientific coding tasks. Compare completion, time and API-equivalent cost.
Task leaderboard
Finish time ranking
Median verified finish time · shorter bars are better
Cost ranking
API-equivalent cost per attempt · shorter bars are better
Red marks unfinished runs. They have no verified finish-time rank; cost bars still show recorded spending, including labelled lower bounds. Configurations without inference are omitted.
Cost–time Pareto plot
Upper left is better. Dashed lines connect the same model's reasoning efforts.
Published API tariffs
Your runs & community contributions
Test any supported model with your installed agent or API connection. Local and imported runs appear in Community & local. Official results are maintained by Simjecture.