Showing All tasks — 467 tasks, 44,362 trials. Back to the default 20-subtask chain view.
APIFlow-Bench 1.0 Leaderboard
Deterministic API engineering benchmark spanning 7 capability axes, over REST APIs only (GraphQL and MCP are planned). Showing all tasks: 44,362 trials across 19 models and 467 tasks covering all 7 axes — authentication, discovery, error recovery, multi-step workflow, pagination, schema, statefulness. Switch task set above; click any card to drill into a slice with the full interactive detail view; click any model name to see its per-task results.
☑ 44,362 trials
▤ 19 models
⛁ 467 tasks
⌖ Overall pass rate: 85.2%
Overall Ranking
Columns are grouped into metric families.
Pass rate = passed ÷ all trials;
API recovery = passed ÷ trials that hit a 4xx/5xx. Difficulty bucket is judged per task from its content and reference solve transcripts (Easy / Long-horizon / Hard / Hard + Long-horizon) — a separate taxonomy from the creation-time “Difficulty (as created)” filter, so the two don't map one-to-one.
Per-trial details (verifier groups, checkpoints, step counts) live on the model → task → trial pages.
| Rank | Model | Scoring | Difficulty bucket (pass rate) | Horizon (pass rate) | Cost / efficiency | Behavioral |
| Pass rate? | API recovery? | Easy? | Long-horizon? | Hard? | Hard + Long-horizon? | Short? | Medium? | Long? | Tool calls / trial? | Time / trial? | Tokens / trial? | Cost / trial? | Abstain %? | Refusal %? |
| 1 | claude-sonnet-4-6 | 89.0%2077/2335 | 621/728 | 100%505/505 | 100%335/335 | 89%357/400 | 81%879/1085 | 97%675/695 | 95%462/485 | 82%939/1145 | 18.9 | 1.3m | 920.7K | $0.31 | 0.2%4/2335 | 0.0%0/2335 |
| 2 | claude-opus-4-8 | 88.6%2068/2335 | 728/835 | 100%505/505 | 99%330/335 | 90%362/400 | 80%871/1085 | 99%685/695 | 94%457/485 | 81%926/1145 | 12.1 | 48.2s | 70.9K | $0.085 | 2.4%55/2335 | 0.0%0/2335 |
| 3 | deepseek-v4-flash | 88.4%2065/2335 | 805/921 | 99%499/505 | 97%326/335 | 91%365/400 | 80%872/1085 | 97%674/695 | 95%460/485 | 81%928/1145 | 19.6 | 53.7s | 153.2K | $0.013 | 0.2%4/2335 | 0.0%0/2335 |
| 4 | kimi-k2p6 | 87.9%2052/2335 | 776/874 | 98%495/505 | 97%324/335 | 92%366/400 | 80%865/1085 | 97%672/695 | 94%458/485 | 80%920/1145 | 17.0 | 1.4m | 205.7K | $0.071 | 0.3%8/2335 | 0.0%0/2335 |
| 5 | gpt-5.5 | 87.8%2051/2335 | 646/759 | 100%504/505 | 98%327/335 | 88%353/400 | 80%867/1085 | 98%678/695 | 93%451/485 | 81%922/1145 | 14.8 | 36.6s | 52.1K | $0.11 | 0.6%15/2335 | 0.0%0/2335 |
| 6 | claude-sonnet-5 | 87.8%2050/2335 | 717/842 | 100%503/505 | 99%330/335 | 92%369/400 | 78%847/1085 | 99%689/695 | 94%458/485 | 79%902/1145 | 15.3 | 53.8s | 119.6K | $0.054 | 1.0%24/2335 | 0.0%0/2335 |
| 7 | kimi-k2p7-code | 87.6%2046/2335 | 713/827 | 99%502/505 | 97%326/335 | 90%360/400 | 79%856/1085 | 97%677/695 | 94%456/485 | 80%911/1145 | 13.9 | 2.3m | 81.5K | $0.031 | 0.7%16/2335 | 0.0%0/2335 |
| 8 | qwen3p6-plus | 87.3%2038/2334 | 744/878 | 99%498/505 | 98%329/335 | 90%360/400 | 79%851/1084 | 98%678/695 | 93%453/485 | 79%907/1144 | 16.3 | 1.5m | 143.9K | $0.057 | 0.1%2/2334 | 0.0%0/2334 |
| 9 | glm-5p2 | 87.1%2033/2335 | 686/816 | 99%499/505 | 97%326/335 | 87%347/400 | 79%861/1085 | 97%675/695 | 91%441/485 | 80%917/1145 | 16.3 | 3.9m | 140.8K | $0.10 | 0.3%8/2335 | 0.0%0/2335 |
| 10 | deepseek-v4-pro | 86.7%2024/2335 | 894/1058 | 99%502/505 | 98%327/335 | 88%350/400 | 78%844/1085 | 98%678/695 | 92%446/485 | 79%899/1145 | 20.4 | 3.6m | 114.3K | $0.12 | 0.1%3/2335 | 0.0%0/2335 |
| 11 | claude-haiku-4-5-20251001 | 86.0%2007/2335 | 773/904 | 97%492/505 | 97%325/335 | 86%343/400 | 78%847/1085 | 96%664/695 | 91%440/485 | 79%903/1145 | 16.5 | 35.5s | 110.1K | $0.028 | 1.0%24/2335 | 0.0%0/2335 |
| 12 | kimi-k2p5 | 85.9%2005/2335 | 723/849 | 99%499/505 | 97%326/335 | 88%350/400 | 76%828/1085 | 96%670/695 | 93%450/485 | 77%883/1145 | 12.3 | 31.9s | 69.6K | $0.029 | 0.6%15/2335 | 0.0%0/2335 |
| 13 | qwen3p7-plus | 85.5%1997/2335 | 645/776 | 98%497/505 | 95%318/333 | 89%351/395 | 77%831/1073 | 97%670/694 | 92%444/480 | 78%883/1132 | 12.1 | 4.0m | 120.7K | $0.035 | 1.2%28/2335 | 0.0%0/2335 |
| 14 | minimax-m2p7 | 85.5%1995/2334 | 829/981 | 98%493/505 | 96%322/335 | 86%344/400 | 77%834/1084 | 95%662/695 | 91%443/485 | 78%888/1144 | 12.9 | 26.8s | 58.9K | $0.014 | 0.4%10/2334 | 0.0%0/2334 |
| 15 | gpt-5.6-terra | 85.4%1993/2335 | 615/755 | 99%500/505 | 98%327/335 | 82%327/400 | 77%839/1085 | 95%663/695 | 90%436/485 | 78%894/1145 | 13.0 | 1.2m | 37.9K | $0.038 | 1.3%31/2335 | 0.0%0/2335 |
| 16 | gpt-oss-120b | 82.0%1914/2335 | 715/901 | 94%474/505 | 94%315/335 | 84%334/400 | 73%791/1085 | 92%642/695 | 89%430/485 | 74%842/1145 | 13.5 | 25.5s | 62.4K | $0.006 | 0.3%7/2335 | 0.0%0/2335 |
| 17 | gpt-oss-20b | 80.9%1889/2334 | 768/965 | 91%462/505 | 91%304/334 | 84%335/400 | 73%788/1085 | 91%630/695 | 87%421/485 | 73%838/1144 | 14.9 | 1.6m | 72.3K | $0.003 | 1.9%45/2334 | 0.0%0/2334 |
| 18 | gpt-5-mini | 80.6%1881/2335 | 682/852 | 95%479/505 | 94%314/335 | 81%324/400 | 70%763/1085 | 93%646/695 | 86%418/485 | 71%816/1145 | 11.5 | 34.3s | 51.0K | $0.006 | 5.5%128/2335 | 0.0%0/2335 |
| 19 | gpt-5.4-mini | 68.5%1599/2335 | 607/964 | 82%416/505 | 79%263/335 | 69%276/400 | 59%644/1085 | 80%553/695 | 75%366/485 | 59%680/1145 | 11.0 | 14.9s | 22.4K | $0.008 | 3.1%72/2335 | 0.0%0/2335 |
Slices
▥Behavior Mix
Top 10 of 19 models
passed?failed?abstained (clarify / report_blocked)?no-score?
| Rank? | Model? | Mix? | Pass rate? |
| 1 | claude-sonnet-4-6 |
| 89.0% |
| 2 | claude-opus-4-8 |
| 88.6% |
| 3 | deepseek-v4-flash |
| 88.4% |
| 4 | kimi-k2p6 |
| 87.9% |
| 5 | gpt-5.5 |
| 87.8% |
| 6 | claude-sonnet-5 |
| 87.8% |
| 7 | kimi-k2p7-code |
| 87.6% |
| 8 | qwen3p6-plus |
| 87.3% |
| 9 | glm-5p2 |
| 87.1% |
| 10 | deepseek-v4-pro |
| 86.7% |
🔐Authentication
6,079 trials
| Rank? | Model? | Pass rate? |
| 1 | gpt-5.5 | 87.5% |
| 2 | claude-sonnet-4-6 | 86.6% |
| 3 | deepseek-v4-flash | 85.3% |
| 4 | deepseek-v4-pro | 85.0% |
| 4 | glm-5p2 | 85.0% |
| 6 | kimi-k2p6 | 84.4% |
| 6 | qwen3p6-plus | 84.4% |
| 8 | kimi-k2p7-code | 83.8% |
| 9 | claude-sonnet-5 | 82.8% |
| 9 | qwen3p7-plus | 82.8% |
+9 more models below — click for the full ranking of all 19.
| Rank? | Model? | Pass rate? |
| 1 | claude-sonnet-4-6 | 82.5% |
| 1 | claude-opus-4-8 | 82.5% |
| 3 | glm-5p2 | 79.5% |
| 4 | gpt-5.5 | 78.9% |
| 5 | deepseek-v4-flash | 78.6% |
| 6 | kimi-k2p5 | 77.8% |
| 6 | qwen3p6-plus | 77.8% |
| 8 | kimi-k2p6 | 77.5% |
| 8 | deepseek-v4-pro | 77.5% |
| 10 | claude-haiku-4-5-20251001 | 77.3% |
+9 more models below — click for the full ranking of all 19.
↻Error Recovery
7,220 trials
| Rank? | Model? | Pass rate? |
| 1 | claude-opus-4-8 | 89.5% |
| 2 | claude-sonnet-5 | 88.9% |
| 3 | deepseek-v4-flash | 87.9% |
| 4 | claude-sonnet-4-6 | 87.6% |
| 5 | kimi-k2p7-code | 87.4% |
| 6 | kimi-k2p6 | 87.1% |
| 6 | qwen3p7-plus | 87.1% |
| 8 | qwen3p6-plus | 86.8% |
| 9 | minimax-m2p7 | 85.8% |
| 10 | gpt-5.5 | 85.5% |
+9 more models below — click for the full ranking of all 19.
⛓Multi-step Workflow
6,460 trials
| Rank? | Model? | Pass rate? |
| 1 | deepseek-v4-pro | 95.0% |
| 2 | glm-5p2 | 94.4% |
| 3 | claude-opus-4-8 | 94.1% |
| 3 | deepseek-v4-flash | 94.1% |
| 3 | claude-sonnet-5 | 94.1% |
| 3 | kimi-k2p7-code | 94.1% |
| 7 | kimi-k2p6 | 93.8% |
| 7 | qwen3p6-plus | 93.8% |
| 9 | gpt-5.5 | 93.2% |
| 10 | claude-sonnet-4-6 | 92.9% |
+9 more models below — click for the full ranking of all 19.
| Rank? | Model? | Pass rate? |
| 1 | gpt-5.5 | 92.2% |
| 2 | deepseek-v4-flash | 91.4% |
| 2 | kimi-k2p7-code | 91.4% |
| 4 | claude-haiku-4-5-20251001 | 91.0% |
| 4 | claude-opus-4-8 | 91.0% |
| 4 | claude-sonnet-5 | 91.0% |
| 4 | claude-sonnet-4-6 | 91.0% |
| 8 | gpt-5.6-terra | 90.2% |
| 8 | kimi-k2p6 | 90.2% |
| 10 | kimi-k2p5 | 89.4% |
+9 more models below — click for the full ranking of all 19.
| Rank? | Model? | Pass rate? |
| 1 | kimi-k2p6 | 93.2% |
| 2 | kimi-k2p5 | 92.5% |
| 2 | deepseek-v4-flash | 92.5% |
| 4 | claude-haiku-4-5-20251001 | 91.2% |
| 4 | qwen3p6-plus | 91.2% |
| 6 | kimi-k2p7-code | 90.8% |
| 7 | claude-sonnet-4-6 | 90.5% |
| 7 | claude-sonnet-5 | 90.5% |
| 9 | gpt-5.5 | 90.2% |
| 10 | deepseek-v4-pro | 89.5% |
+9 more models below — click for the full ranking of all 19.
≣Statefulness
7,218 trials
| Rank? | Model? | Pass rate? |
| 1 | claude-sonnet-4-6 | 92.4% |
| 2 | claude-sonnet-5 | 92.1% |
| 3 | claude-opus-4-8 | 91.8% |
| 4 | qwen3p6-plus | 91.3% |
| 5 | deepseek-v4-flash | 90.8% |
| 6 | kimi-k2p6 | 90.5% |
| 7 | minimax-m2p7 | 90.3% |
| 7 | kimi-k2p7-code | 90.3% |
| 9 | claude-haiku-4-5-20251001 | 90.0% |
| 10 | kimi-k2p5 | 89.7% |
+9 more models below — click for the full ranking of all 19.
Per-Model Results
Drill into any model to see per-task pass rates across its attempted tasks. This build contains aggregate results only (no per-trial pages).