Showing All tasks created as hard — 308 tasks, 29,258 trials. Back to the default 20-subtask chain view.
APIFlow-Bench 1.0 Leaderboard
Deterministic API engineering benchmark spanning 7 capability axes, over REST APIs only (GraphQL and MCP are planned). Showing all tasks: 29,258 trials across 19 models and 308 tasks covering all 7 axes — authentication, discovery, error recovery, multi-step workflow, pagination, schema, statefulness. Switch task set above; click any card to drill into a slice with the full interactive detail view; click any model name to see its per-task results.
☑ 29,258 trials
▤ 19 models
⛁ 308 tasks
⌖ Overall pass rate: 81.3%
Overall Ranking
Columns are grouped into metric families.
Pass rate = passed ÷ all trials;
API recovery = passed ÷ trials that hit a 4xx/5xx. Difficulty bucket is judged per task from its content and reference solve transcripts (Easy / Long-horizon / Hard / Hard + Long-horizon) — a separate taxonomy from the creation-time “Difficulty (as created)” filter, so the two don't map one-to-one.
Per-trial details (verifier groups, checkpoints, step counts) live on the model → task → trial pages.
| Rank | Model | Scoring | Difficulty bucket (pass rate) | Horizon (pass rate) | Cost / efficiency | Behavioral |
| Pass rate? | API recovery? | Easy? | Long-horizon? | Hard? | Hard + Long-horizon? | Short? | Medium? | Long? | Tool calls / trial? | Time / trial? | Tokens / trial? | Cost / trial? | Abstain %? | Refusal %? |
| 1 | claude-sonnet-4-6 | 85.8%1321/1540 | 469/556 | 100%140/140 | 100%130/130 | 95%171/180 | 81%879/1085 | 98%230/235 | 98%156/160 | 82%934/1140 | 21.8 | 1.5m | 1.3M | $0.44 | 0.2%3/1540 | 0.0%0/1540 |
| 2 | claude-opus-4-8 | 84.7%1305/1540 | 537/625 | 100%140/140 | 96%125/130 | 94%169/180 | 80%871/1085 | 100%235/235 | 93%149/160 | 81%921/1140 | 12.9 | 48.5s | 76.7K | $0.092 | 2.8%43/1540 | 0.0%0/1540 |
| 3 | deepseek-v4-flash | 84.6%1303/1540 | 580/674 | 96%135/140 | 94%122/130 | 95%171/180 | 80%872/1085 | 96%225/235 | 95%152/160 | 81%923/1140 | 20.4 | 53.5s | 137.7K | $0.012 | 0.1%2/1540 | 0.0%0/1540 |
| 4 | gpt-5.5 | 84.5%1302/1540 | 491/581 | 100%140/140 | 94%122/130 | 96%173/180 | 80%867/1085 | 99%233/235 | 95%152/160 | 80%917/1140 | 15.9 | 37.6s | 57.5K | $0.12 | 0.4%6/1540 | 0.0%0/1540 |
| 5 | kimi-k2p6 | 84.2%1296/1540 | 560/638 | 96%135/140 | 92%120/130 | 97%174/180 | 80%865/1085 | 97%227/235 | 95%152/160 | 80%915/1140 | 18.2 | 1.5m | 227.6K | $0.078 | 0.5%7/1540 | 0.0%0/1540 |
| 6 | glm-5p2 | 84.0%1294/1540 | 521/613 | 97%136/140 | 95%123/130 | 97%174/180 | 79%861/1085 | 98%230/235 | 95%152/160 | 80%912/1140 | 17.2 | 4.1m | 159.4K | $0.12 | 0.5%8/1540 | 0.0%0/1540 |
| 7 | kimi-k2p7-code | 83.9%1292/1540 | 528/618 | 99%138/140 | 93%121/130 | 97%175/180 | 79%856/1085 | 97%229/235 | 97%155/160 | 79%906/1140 | 14.5 | 2.4m | 78.3K | $0.031 | 0.5%7/1540 | 0.0%0/1540 |
| 8 | claude-sonnet-5 | 83.6%1287/1540 | 516/618 | 99%138/140 | 96%125/130 | 98%176/180 | 78%847/1085 | 99%233/235 | 98%156/160 | 79%897/1140 | 16.4 | 57.8s | 127.7K | $0.059 | 1.0%16/1540 | 0.0%0/1540 |
| 9 | qwen3p6-plus | 83.4%1283/1539 | 544/650 | 97%136/140 | 95%124/130 | 96%172/180 | 79%851/1084 | 98%231/235 | 94%150/160 | 79%902/1139 | 17.4 | 1.6m | 161.2K | $0.063 | 0.0%0/1539 | 0.0%0/1539 |
| 10 | deepseek-v4-pro | 82.8%1275/1540 | 670/800 | 99%139/140 | 95%124/130 | 93%167/180 | 78%844/1085 | 98%230/235 | 94%150/160 | 78%894/1140 | 22.1 | 3.8m | 123.9K | $0.12 | 0.1%2/1540 | 0.0%0/1540 |
| 11 | claude-haiku-4-5-20251001 | 82.3%1267/1540 | 524/624 | 95%133/140 | 93%121/130 | 92%166/180 | 78%847/1085 | 95%223/235 | 91%146/160 | 79%898/1140 | 17.5 | 35.8s | 117.6K | $0.029 | 1.0%15/1540 | 0.0%0/1540 |
| 12 | gpt-5.6-terra | 82.1%1264/1540 | 470/570 | 98%137/140 | 95%124/130 | 91%164/180 | 77%839/1085 | 95%224/235 | 94%151/160 | 78%889/1140 | 14.0 | 1.3m | 42.4K | $0.042 | 1.4%22/1540 | 0.0%0/1540 |
| 13 | minimax-m2p7 | 81.7%1257/1539 | 564/679 | 98%137/140 | 92%120/130 | 91%164/180 | 77%834/1084 | 96%225/235 | 92%147/160 | 78%883/1139 | 13.7 | 26.2s | 61.6K | $0.015 | 0.5%8/1539 | 0.0%0/1539 |
| 14 | kimi-k2p5 | 81.4%1253/1540 | 509/606 | 97%136/140 | 93%121/130 | 92%166/180 | 76%828/1085 | 95%224/235 | 93%149/160 | 77%878/1140 | 13.0 | 32.1s | 69.3K | $0.029 | 0.5%8/1540 | 0.0%0/1540 |
| 15 | qwen3p7-plus | 81.3%1252/1540 | 472/578 | 97%136/140 | 94%121/129 | 93%164/177 | 77%831/1073 | 97%227/234 | 92%146/158 | 78%879/1127 | 12.9 | 4.1m | 123.8K | $0.037 | 1.1%17/1540 | 0.0%0/1540 |
| 16 | gpt-oss-120b | 77.9%1200/1540 | 518/666 | 91%127/140 | 92%120/130 | 90%162/180 | 73%791/1085 | 91%215/235 | 92%148/160 | 73%837/1140 | 14.3 | 26.8s | 67.8K | $0.007 | 0.3%5/1540 | 0.0%0/1540 |
| 17 | gpt-oss-20b | 76.9%1184/1540 | 557/712 | 86%121/140 | 88%114/130 | 89%161/180 | 73%788/1085 | 89%208/235 | 88%140/160 | 73%836/1140 | 15.8 | 1.7m | 78.0K | $0.004 | 1.8%28/1540 | 0.0%0/1540 |
| 18 | gpt-5-mini | 75.8%1168/1540 | 489/621 | 95%133/140 | 88%115/130 | 87%156/180 | 70%763/1085 | 93%218/235 | 86%138/160 | 71%811/1140 | 12.2 | 36.0s | 54.8K | $0.006 | 5.8%90/1540 | 0.0%0/1540 |
| 19 | gpt-5.4-mini | 63.6%980/1540 | 432/692 | 81%114/140 | 65%84/130 | 77%138/180 | 59%644/1085 | 79%185/235 | 72%115/160 | 60%680/1140 | 11.8 | 15.4s | 23.9K | $0.009 | 3.2%49/1540 | 0.0%0/1540 |
Slices
▥Behavior Mix
Top 10 of 19 models
passed?failed?abstained (clarify / report_blocked)?no-score?
| Rank? | Model? | Mix? | Pass rate? |
| 1 | claude-sonnet-4-6 |
| 85.8% |
| 2 | claude-opus-4-8 |
| 84.7% |
| 3 | deepseek-v4-flash |
| 84.6% |
| 4 | gpt-5.5 |
| 84.5% |
| 5 | kimi-k2p6 |
| 84.2% |
| 6 | glm-5p2 |
| 84.0% |
| 7 | kimi-k2p7-code |
| 83.9% |
| 8 | claude-sonnet-5 |
| 83.6% |
| 9 | qwen3p6-plus |
| 83.4% |
| 10 | deepseek-v4-pro |
| 82.8% |
🔐Authentication
4,369 trials
| Rank? | Model? | Pass rate? |
| 1 | gpt-5.5 | 86.5% |
| 2 | claude-sonnet-4-6 | 83.9% |
| 3 | qwen3p6-plus | 83.5% |
| 4 | glm-5p2 | 83.0% |
| 5 | kimi-k2p6 | 82.6% |
| 6 | deepseek-v4-flash | 82.2% |
| 6 | deepseek-v4-pro | 82.2% |
| 8 | kimi-k2p7-code | 81.7% |
| 9 | qwen3p7-plus | 80.9% |
| 10 | claude-opus-4-8 | 79.6% |
+9 more models below — click for the full ranking of all 19.
| Rank? | Model? | Pass rate? |
| 1 | claude-sonnet-4-6 | 76.4% |
| 2 | claude-opus-4-8 | 74.4% |
| 3 | gpt-5.5 | 70.8% |
| 3 | glm-5p2 | 70.8% |
| 5 | deepseek-v4-flash | 68.8% |
| 6 | kimi-k2p5 | 68.4% |
| 7 | claude-haiku-4-5-20251001 | 68.0% |
| 8 | qwen3p6-plus | 67.6% |
| 8 | gpt-5.6-terra | 67.6% |
| 8 | kimi-k2p6 | 67.6% |
+9 more models below — click for the full ranking of all 19.
↻Error Recovery
4,560 trials
| Rank? | Model? | Pass rate? |
| 1 | claude-opus-4-8 | 87.1% |
| 2 | claude-sonnet-4-6 | 85.8% |
| 3 | deepseek-v4-flash | 85.4% |
| 4 | claude-sonnet-5 | 85.0% |
| 5 | qwen3p6-plus | 84.6% |
| 6 | kimi-k2p7-code | 84.2% |
| 6 | qwen3p7-plus | 84.2% |
| 8 | glm-5p2 | 83.3% |
| 9 | kimi-k2p6 | 82.5% |
| 10 | minimax-m2p7 | 82.1% |
+9 more models below — click for the full ranking of all 19.
⛓Multi-step Workflow
4,655 trials
| Rank? | Model? | Pass rate? |
| 1 | kimi-k2p6 | 95.1% |
| 1 | deepseek-v4-pro | 95.1% |
| 1 | glm-5p2 | 95.1% |
| 4 | deepseek-v4-flash | 93.9% |
| 4 | claude-sonnet-5 | 93.9% |
| 4 | kimi-k2p7-code | 93.9% |
| 7 | gpt-5.5 | 93.1% |
| 7 | claude-opus-4-8 | 93.1% |
| 7 | qwen3p6-plus | 93.1% |
| 10 | claude-sonnet-4-6 | 91.8% |
+9 more models below — click for the full ranking of all 19.
| Rank? | Model? | Pass rate? |
| 1 | gpt-5.5 | 89.7% |
| 2 | deepseek-v4-flash | 89.0% |
| 3 | claude-haiku-4-5-20251001 | 88.3% |
| 3 | gpt-5.6-terra | 88.3% |
| 3 | kimi-k2p7-code | 88.3% |
| 6 | claude-opus-4-8 | 87.6% |
| 6 | claude-sonnet-5 | 87.6% |
| 6 | claude-sonnet-4-6 | 87.6% |
| 9 | kimi-k2p6 | 86.2% |
| 10 | minimax-m2p7 | 85.5% |
+9 more models below — click for the full ranking of all 19.
| Rank? | Model? | Pass rate? |
| 1 | kimi-k2p6 | 89.5% |
| 2 | kimi-k2p5 | 88.4% |
| 2 | deepseek-v4-flash | 88.4% |
| 4 | claude-haiku-4-5-20251001 | 87.4% |
| 5 | kimi-k2p7-code | 86.3% |
| 5 | qwen3p6-plus | 86.3% |
| 7 | claude-sonnet-4-6 | 85.3% |
| 7 | claude-sonnet-5 | 85.3% |
| 9 | gpt-5.5 | 84.7% |
| 10 | qwen3p7-plus | 83.7% |
+9 more models below — click for the full ranking of all 19.
≣Statefulness
4,559 trials
| Rank? | Model? | Pass rate? |
| 1 | claude-sonnet-5 | 91.2% |
| 2 | claude-sonnet-4-6 | 90.4% |
| 3 | claude-opus-4-8 | 89.6% |
| 4 | claude-haiku-4-5-20251001 | 88.3% |
| 4 | kimi-k2p7-code | 88.3% |
| 4 | glm-5p2 | 88.3% |
| 7 | minimax-m2p7 | 87.9% |
| 7 | kimi-k2p6 | 87.9% |
| 9 | qwen3p6-plus | 87.9% |
| 10 | gpt-5.5 | 87.5% |
+9 more models below — click for the full ranking of all 19.
Per-Model Results
Drill into any model to see per-task pass rates across its attempted tasks. This build contains aggregate results only (no per-trial pages).