Showing Solo subtasks created as hard — 82 tasks, 7,790 trials. Back to the default 20-subtask chain view.
APIFlow-Bench 1.0 Leaderboard
Deterministic API engineering benchmark spanning 7 capability axes, over REST APIs only (GraphQL and MCP are planned). Showing the solo subtasks task set: 7,790 trials across 19 models and 82 tasks covering all 7 axes — authentication, discovery, error recovery, multi-step workflow, pagination, schema, statefulness. Switch task set above; click any card to drill into a slice with the full interactive detail view; click any model name to see its per-task results.
☑ 7,790 trials
▤ 19 models
⛁ 82 tasks
⌖ Overall pass rate: 93.1%
Overall Ranking
Columns are grouped into metric families.
Pass rate = passed ÷ all trials;
API recovery = passed ÷ trials that hit a 4xx/5xx. Difficulty bucket is judged per task from its content and reference solve transcripts (Easy / Long-horizon / Hard / Hard + Long-horizon) — a separate taxonomy from the creation-time “Difficulty (as created)” filter, so the two don't map one-to-one.
Per-trial details (verifier groups, checkpoints, step counts) live on the model → task → trial pages.
| Rank | Model | Scoring | Difficulty bucket (pass rate) | Horizon (pass rate) | Cost / efficiency | Behavioral |
| Pass rate? | API recovery? | Easy? | Long-horizon? | Hard? | Short? | Medium? | Long? | Tool calls / trial? | Time / trial? | Tokens / trial? | Cost / trial? | Abstain %? | Refusal %? |
| 1 | claude-sonnet-5 | 97.6%400/410 | 128/136 | 99%138/140 | 100%90/90 | 98%171/175 | 99%233/235 | 98%156/160 | 100%10/10 | 14.0 | 55.4s | 138.6K | $0.060 | 1.2%5/410 | 0.0%0/410 |
| 2 | claude-sonnet-4-6 | 96.8%397/410 | 98/106 | 100%140/140 | 100%90/90 | 95%166/175 | 98%230/235 | 98%156/160 | 100%10/10 | 12.0 | 47.7s | 118.9K | $0.062 | 0.0%0/410 | 0.0%0/410 |
| 3 | kimi-k2p7-code | 96.6%396/410 | 128/134 | 99%138/140 | 96%86/90 | 97%170/175 | 97%229/235 | 97%155/160 | 100%10/10 | 13.5 | 1.9m | 96.6K | $0.036 | 1.0%4/410 | 0.0%0/410 |
| 4 | gpt-5.5 | 96.3%395/410 | 104/114 | 100%140/140 | 97%87/90 | 96%168/175 | 99%233/235 | 95%152/160 | 100%10/10 | 14.1 | 31.6s | 54.0K | $0.10 | 0.0%0/410 | 0.0%0/410 |
| 5 | claude-opus-4-8 | 96.1%394/410 | 121/134 | 100%140/140 | 100%90/90 | 94%164/175 | 100%235/235 | 93%149/160 | 100%10/10 | 11.8 | 42.6s | 68.6K | $0.084 | 3.9%16/410 | 0.0%0/410 |
| 6 | glm-5p2 | 95.6%392/410 | 116/126 | 97%136/140 | 97%87/90 | 97%169/175 | 98%230/235 | 95%152/160 | 100%10/10 | 14.3 | 3.2m | 85.7K | $0.077 | 0.0%0/410 | 0.0%0/410 |
| 7 | kimi-k2p6 | 95.4%391/410 | 134/140 | 96%135/140 | 94%85/90 | 97%169/175 | 97%227/235 | 95%152/160 | 100%10/10 | 15.6 | 1.2m | 236.5K | $0.071 | 0.5%2/410 | 0.0%0/410 |
| 7 | qwen3p6-plus | 95.4%391/410 | 117/129 | 97%136/140 | 98%88/90 | 95%167/175 | 98%231/235 | 94%150/160 | 100%10/10 | 15.1 | 1.4m | 136.5K | $0.053 | 0.0%0/410 | 0.0%0/410 |
| 7 | deepseek-v4-pro | 95.4%391/410 | 141/152 | 99%139/140 | 99%89/90 | 93%162/175 | 98%230/235 | 94%150/160 | 100%10/10 | 18.9 | 3.5m | 130.3K | $0.11 | 0.2%1/410 | 0.0%0/410 |
| 10 | deepseek-v4-flash | 95.1%390/410 | 139/146 | 96%135/140 | 96%86/90 | 95%166/175 | 96%225/235 | 95%152/160 | 100%10/10 | 17.9 | 54.6s | 183.9K | $0.015 | 0.0%0/410 | 0.0%0/410 |
| 11 | kimi-k2p5 | 93.9%385/410 | 122/130 | 97%136/140 | 96%86/90 | 92%161/175 | 95%224/235 | 93%149/160 | 100%10/10 | 12.1 | 28.6s | 75.1K | $0.031 | 0.7%3/410 | 0.0%0/410 |
| 11 | gpt-5.6-terra | 93.9%385/410 | 99/112 | 98%137/140 | 99%89/90 | 91%159/175 | 95%224/235 | 94%151/160 | 100%10/10 | 12.3 | 1.0m | 36.6K | $0.036 | 2.0%8/410 | 0.0%0/410 |
| 13 | minimax-m2p7 | 93.4%383/410 | 144/159 | 98%137/140 | 94%85/90 | 91%159/175 | 96%225/235 | 92%147/160 | 90%9/10 | 13.0 | 25.6s | 66.2K | $0.016 | 0.2%1/410 | 0.0%0/410 |
| 14 | qwen3p7-plus | 93.2%382/410 | 102/113 | 97%136/140 | 97%87/90 | 92%159/172 | 97%227/234 | 92%146/158 | 90%9/10 | 11.9 | 3.5m | 146.6K | $0.042 | 1.7%7/410 | 0.0%0/410 |
| 15 | claude-haiku-4-5-20251001 | 92.4%379/410 | 133/147 | 95%133/140 | 94%85/90 | 92%161/175 | 95%223/235 | 91%146/160 | 100%10/10 | 14.8 | 29.5s | 101.6K | $0.026 | 1.7%7/410 | 0.0%0/410 |
| 16 | gpt-oss-120b | 91.0%373/410 | 130/145 | 91%127/140 | 98%88/90 | 90%158/175 | 91%215/235 | 92%148/160 | 100%10/10 | 12.4 | 20.3s | 56.6K | $0.006 | 0.0%0/410 | 0.0%0/410 |
| 17 | gpt-5-mini | 89.5%367/410 | 115/129 | 95%133/140 | 90%81/90 | 87%152/175 | 93%218/235 | 86%138/160 | 100%10/10 | 11.1 | 29.7s | 50.4K | $0.006 | 3.7%15/410 | 0.0%0/410 |
| 18 | gpt-oss-20b | 87.1%357/410 | 129/150 | 86%121/140 | 89%80/90 | 89%156/175 | 89%208/235 | 88%140/160 | 90%9/10 | 13.5 | 1.5m | 65.4K | $0.003 | 2.0%8/410 | 0.0%0/410 |
| 19 | gpt-5.4-mini | 74.4%305/410 | 103/149 | 81%114/140 | 63%57/90 | 77%134/175 | 79%185/235 | 72%115/160 | 50%5/10 | 10.3 | 14.2s | 20.9K | $0.008 | 2.7%11/410 | 0.0%0/410 |
Slices
▥Behavior Mix
Top 10 of 19 models
passed?failed?abstained (clarify / report_blocked)?no-score?
| Rank? | Model? | Mix? | Pass rate? |
| 1 | claude-sonnet-5 |
| 97.6% |
| 2 | claude-sonnet-4-6 |
| 96.8% |
| 3 | kimi-k2p7-code |
| 96.6% |
| 4 | gpt-5.5 |
| 96.3% |
| 5 | claude-opus-4-8 |
| 96.1% |
| 6 | glm-5p2 |
| 95.6% |
| 7 | kimi-k2p6 |
| 95.4% |
| 7 | qwen3p6-plus |
| 95.4% |
| 7 | deepseek-v4-pro |
| 95.4% |
| 10 | deepseek-v4-flash |
| 95.1% |
🔐Authentication
1,425 trials
| Rank? | Model? | Pass rate? |
| 1 | claude-sonnet-5 | 100.0% |
| 1 | claude-opus-4-8 | 100.0% |
| 1 | qwen3p6-plus | 100.0% |
| 4 | glm-5p2 | 98.7% |
| 5 | gpt-5.5 | 97.3% |
| 6 | kimi-k2p6 | 96.0% |
| 7 | kimi-k2p7-code | 94.7% |
| 7 | deepseek-v4-pro | 94.7% |
| 7 | qwen3p7-plus | 94.7% |
| 10 | gpt-oss-120b | 93.3% |
+9 more models below — click for the full ranking of all 19.
| Rank? | Model? | Pass rate? |
| 1 | 8 models tied at 100% | 100.0% |
| 9 | kimi-k2p5 | 98.7% |
| 9 | deepseek-v4-flash | 98.7% |
| 11 | qwen3p6-plus | 97.3% |
| 11 | deepseek-v4-pro | 97.3% |
| 13 | claude-haiku-4-5-20251001 | 96.0% |
| 14 | claude-sonnet-5 | 94.7% |
| 15 | qwen3p7-plus | 93.3% |
| 16 | gpt-oss-120b | 85.3% |
| 16 | gpt-5-mini | 85.3% |
| 18 | gpt-5.4-mini | 78.7% |
+1 more model below — click for the full ranking of all 19.
↻Error Recovery
1,045 trials
| Rank? | Model? | Pass rate? |
| 1 | claude-sonnet-4-6 | 100.0% |
| 1 | claude-opus-4-8 | 100.0% |
| 1 | claude-sonnet-5 | 100.0% |
| 4 | gpt-5.6-terra | 98.2% |
| 4 | deepseek-v4-pro | 98.2% |
| 6 | minimax-m2p7 | 96.4% |
| 6 | qwen3p6-plus | 96.4% |
| 6 | qwen3p7-plus | 96.4% |
| 9 | gpt-oss-120b | 94.5% |
| 9 | gpt-5.5 | 94.5% |
+9 more models below — click for the full ranking of all 19.
⛓Multi-step Workflow
1,520 trials
| Rank? | Model? | Pass rate? |
| 1 | deepseek-v4-flash | 97.5% |
| 2 | kimi-k2p6 | 96.2% |
| 3 | kimi-k2p5 | 95.0% |
| 3 | kimi-k2p7-code | 95.0% |
| 3 | deepseek-v4-pro | 95.0% |
| 6 | gpt-oss-120b | 93.8% |
| 6 | gpt-5.5 | 93.8% |
| 6 | claude-sonnet-4-6 | 93.8% |
| 6 | claude-sonnet-5 | 93.8% |
| 6 | glm-5p2 | 93.8% |
+9 more models below — click for the full ranking of all 19.
| Rank? | Model? | Pass rate? |
| 1 | 11 models tied at 100% | 100.0% |
| 12 | claude-opus-4-8 | 96.0% |
| 12 | deepseek-v4-pro | 96.0% |
| 14 | gpt-5.4-mini | 92.0% |
| 14 | gpt-oss-120b | 92.0% |
| 14 | qwen3p6-plus | 92.0% |
| 14 | glm-5p2 | 92.0% |
| 18 | claude-haiku-4-5-20251001 | 88.0% |
| 18 | gpt-oss-20b | 88.0% |
| Rank? | Model? | Pass rate? |
| 1 | 9 models tied at 100% | 100.0% |
| 10 | deepseek-v4-flash | 97.8% |
| 10 | gpt-5.5 | 97.8% |
| 10 | gpt-oss-20b | 97.8% |
| 10 | deepseek-v4-pro | 97.8% |
| 14 | gpt-oss-120b | 95.6% |
| 14 | gpt-5-mini | 95.6% |
| 14 | gpt-5.6-terra | 95.6% |
| 14 | claude-opus-4-8 | 95.6% |
| 18 | minimax-m2p7 | 93.3% |
| 19 | gpt-5.4-mini | 82.2% |
≣Statefulness
1,045 trials
| Rank? | Model? | Pass rate? |
| 1 | claude-sonnet-5 | 98.2% |
| 2 | claude-sonnet-4-6 | 94.5% |
| 2 | kimi-k2p7-code | 94.5% |
| 4 | gpt-5.5 | 92.7% |
| 5 | claude-opus-4-8 | 90.9% |
| 6 | deepseek-v4-pro | 89.1% |
| 7 | deepseek-v4-flash | 87.3% |
| 7 | qwen3p6-plus | 87.3% |
| 7 | glm-5p2 | 87.3% |
| 10 | gpt-5.6-terra | 85.5% |
+9 more models below — click for the full ranking of all 19.
Per-Model Results
Drill into any model to see per-task pass rates across its attempted tasks. This build contains aggregate results only (no per-trial pages).