Showing Solo subtasks created as easy — 69 tasks, 6,555 trials. Back to the default 20-subtask chain view.
APIFlow-Bench 1.0 Leaderboard
Deterministic API engineering benchmark spanning 7 capability axes, over REST APIs only (GraphQL and MCP are planned). Showing the solo subtasks task set: 6,555 trials across 19 models and 69 tasks covering all 7 axes — authentication, discovery, error recovery, multi-step workflow, pagination, schema, statefulness. Switch task set above; click any card to drill into a slice with the full interactive detail view; click any model name to see its per-task results.
☑ 6,555 trials
▤ 19 models
⛁ 69 tasks
⌖ Overall pass rate: 93.5%
Overall Ranking
Columns are grouped into metric families.
Pass rate = passed ÷ all trials;
API recovery = passed ÷ trials that hit a 4xx/5xx. Difficulty bucket is judged per task from its content and reference solve transcripts (Easy / Long-horizon / Hard / Hard + Long-horizon) — a separate taxonomy from the creation-time “Difficulty (as created)” filter, so the two don't map one-to-one.
Per-trial details (verifier groups, checkpoints, step counts) live on the model → task → trial pages.
| Rank | Model | Scoring | Difficulty bucket (pass rate) | Horizon (pass rate) | Cost / efficiency | Behavioral |
| Pass rate? | API recovery? | Easy? | Long-horizon? | Hard? | Short? | Medium? | Tool calls / trial? | Time / trial? | Tokens / trial? | Cost / trial? | Abstain %? | Refusal %? |
| 1 | claude-sonnet-5 | 97.4%336/345 | 90/91 | 100%195/195 | 100%75/75 | 88%66/75 | 98%231/235 | 95%105/110 | 11.4 | 37.2s | 68.1K | $0.033 | 0.0%0/345 | 0.0%0/345 |
| 2 | claude-opus-4-8 | 95.9%331/345 | 78/81 | 100%195/195 | 100%75/75 | 81%61/75 | 96%225/235 | 96%106/110 | 9.4 | 48.5s | 47.7K | $0.058 | 0.6%2/345 | 0.0%0/345 |
| 3 | qwen3p7-plus | 95.7%330/345 | 76/80 | 100%195/195 | 97%73/75 | 83%62/75 | 96%226/235 | 95%104/110 | 8.7 | 3.4m | 72.4K | $0.025 | 0.3%1/345 | 0.0%0/345 |
| 4 | kimi-k2p5 | 95.4%329/345 | 96/102 | 99%194/195 | 100%75/75 | 80%60/75 | 95%224/235 | 95%105/110 | 8.7 | 29.5s | 30.4K | $0.014 | 0.0%0/345 | 0.0%0/345 |
| 4 | deepseek-v4-flash | 95.4%329/345 | 90/97 | 99%194/195 | 100%75/75 | 80%60/75 | 95%224/235 | 95%105/110 | 14.7 | 46.5s | 87.7K | $0.008 | 0.0%0/345 | 0.0%0/345 |
| 4 | kimi-k2p7-code | 95.4%329/345 | 84/87 | 99%194/195 | 100%75/75 | 80%60/75 | 95%224/235 | 95%105/110 | 11.2 | 1.9m | 46.8K | $0.020 | 0.3%1/345 | 0.0%0/345 |
| 7 | claude-sonnet-4-6 | 94.8%327/345 | 63/66 | 100%195/195 | 100%75/75 | 76%57/75 | 96%225/235 | 93%102/110 | 10.5 | 49.6s | 82.7K | $0.043 | 0.0%0/345 | 0.0%0/345 |
| 7 | gpt-5.5 | 94.8%327/345 | 62/64 | 99%194/195 | 100%75/75 | 77%58/75 | 95%224/235 | 94%103/110 | 11.0 | 32.3s | 33.5K | $0.075 | 0.3%1/345 | 0.0%0/345 |
| 7 | kimi-k2p6 | 94.8%327/345 | 95/96 | 97%190/195 | 99%74/75 | 84%63/75 | 94%220/235 | 97%107/110 | 12.9 | 55.5s | 101.4K | $0.041 | 0.0%0/345 | 0.0%0/345 |
| 7 | deepseek-v4-pro | 94.8%327/345 | 90/97 | 100%195/195 | 100%75/75 | 76%57/75 | 96%225/235 | 93%102/110 | 14.3 | 2.7m | 72.6K | $0.079 | 0.0%0/345 | 0.0%0/345 |
| 11 | minimax-m2p7 | 94.5%326/345 | 120/131 | 98%191/195 | 100%75/75 | 80%60/75 | 94%220/235 | 96%106/110 | 10.2 | 27.1s | 42.2K | $0.011 | 0.0%0/345 | 0.0%0/345 |
| 12 | qwen3p6-plus | 94.2%325/345 | 82/93 | 99%193/195 | 100%75/75 | 76%57/75 | 95%223/235 | 93%102/110 | 11.1 | 55.3s | 65.1K | $0.029 | 0.0%0/345 | 0.0%0/345 |
| 13 | claude-haiku-4-5-20251001 | 93.6%323/345 | 107/115 | 98%191/195 | 99%74/75 | 77%58/75 | 94%221/235 | 93%102/110 | 12.5 | 33.8s | 73.1K | $0.022 | 0.6%2/345 | 0.0%0/345 |
| 14 | gpt-5.6-terra | 92.8%320/345 | 62/71 | 99%194/195 | 100%75/75 | 68%51/75 | 94%220/235 | 91%100/110 | 9.9 | 52.3s | 25.0K | $0.026 | 0.3%1/345 | 0.0%0/345 |
| 14 | glm-5p2 | 92.8%320/345 | 67/78 | 99%194/195 | 97%73/75 | 71%53/75 | 95%223/235 | 88%97/110 | 13.6 | 2.9m | 83.0K | $0.065 | 0.0%0/345 | 0.0%0/345 |
| 16 | gpt-5-mini | 91.9%317/345 | 90/101 | 95%186/195 | 99%74/75 | 76%57/75 | 92%217/235 | 91%100/110 | 9.2 | 29.4s | 37.0K | $0.005 | 3.8%13/345 | 0.0%0/345 |
| 17 | gpt-oss-120b | 91.3%315/345 | 94/103 | 95%185/195 | 95%71/75 | 79%59/75 | 91%215/235 | 91%100/110 | 11.2 | 23.8s | 47.2K | $0.005 | 0.6%2/345 | 0.0%0/345 |
| 18 | gpt-oss-20b | 90.7%313/345 | 95/101 | 93%182/195 | 93%70/75 | 81%61/75 | 91%213/235 | 91%100/110 | 12.2 | 1.3m | 53.9K | $0.003 | 1.2%4/345 | 0.0%0/345 |
| 19 | gpt-5.4-mini | 81.4%281/345 | 70/108 | 85%165/195 | 95%71/75 | 60%45/75 | 82%192/235 | 81%89/110 | 8.5 | 13.5s | 18.0K | $0.007 | 3.5%12/345 | 0.0%0/345 |
Slices
▥Behavior Mix
Top 10 of 19 models
passed?failed?abstained (clarify / report_blocked)?no-score?
| Rank? | Model? | Mix? | Pass rate? |
| 1 | claude-sonnet-5 |
| 97.4% |
| 2 | claude-opus-4-8 |
| 95.9% |
| 3 | qwen3p7-plus |
| 95.7% |
| 4 | kimi-k2p5 |
| 95.4% |
| 4 | deepseek-v4-flash |
| 95.4% |
| 4 | kimi-k2p7-code |
| 95.4% |
| 7 | claude-sonnet-4-6 |
| 94.8% |
| 7 | gpt-5.5 |
| 94.8% |
| 7 | kimi-k2p6 |
| 94.8% |
| 7 | deepseek-v4-pro |
| 94.8% |
🔐Authentication
380 trials
| Rank? | Model? | Pass rate? |
| 1 | claude-sonnet-5 | 80.0% |
| 1 | claude-opus-4-8 | 80.0% |
| 1 | qwen3p7-plus | 80.0% |
| 1 | glm-5p2 | 80.0% |
| 5 | minimax-m2p7 | 75.0% |
| 5 | kimi-k2p5 | 75.0% |
| 5 | kimi-k2p6 | 75.0% |
| 5 | deepseek-v4-flash | 75.0% |
| 5 | gpt-5.6-terra | 75.0% |
| 5 | gpt-5.5 | 75.0% |
+9 more models below — click for the full ranking of all 19.
| Rank? | Model? | Pass rate? |
| 1 | 12 models tied at 100% | 100.0% |
| 13 | gpt-5-mini | 98.3% |
| 13 | claude-haiku-4-5-20251001 | 98.3% |
| 13 | kimi-k2p6 | 98.3% |
| 13 | glm-5p2 | 98.3% |
| 17 | gpt-5.4-mini | 95.0% |
| 18 | gpt-oss-120b | 91.7% |
| 19 | gpt-oss-20b | 86.7% |
↻Error Recovery
1,330 trials
| Rank? | Model? | Pass rate? |
| 1 | claude-sonnet-5 | 92.9% |
| 2 | kimi-k2p6 | 90.0% |
| 3 | claude-opus-4-8 | 87.1% |
| 4 | kimi-k2p5 | 85.7% |
| 4 | gpt-oss-20b | 85.7% |
| 4 | minimax-m2p7 | 85.7% |
| 4 | deepseek-v4-flash | 85.7% |
| 4 | qwen3p7-plus | 85.7% |
| 4 | kimi-k2p7-code | 85.7% |
| 10 | gpt-oss-120b | 84.3% |
+9 more models below — click for the full ranking of all 19.
⛓Multi-step Workflow
855 trials
| Rank? | Model? | Pass rate? |
| 1 | 7 models tied at 100% | 100.0% |
| 8 | kimi-k2p5 | 97.8% |
| 8 | qwen3p6-plus | 97.8% |
| 8 | deepseek-v4-flash | 97.8% |
| 8 | gpt-5.5 | 97.8% |
| 8 | qwen3p7-plus | 97.8% |
| 13 | minimax-m2p7 | 95.6% |
| 13 | claude-haiku-4-5-20251001 | 95.6% |
| 13 | glm-5p2 | 95.6% |
| 16 | gpt-oss-120b | 91.1% |
| 17 | kimi-k2p6 | 88.9% |
+2 more models below — click for the full ranking of all 19.
| Rank? | Model? | Pass rate? |
| 1 | 14 models tied at 100% | 100.0% |
| 15 | qwen3p6-plus | 98.0% |
| 15 | gpt-oss-20b | 98.0% |
| 17 | minimax-m2p7 | 96.0% |
| 18 | gpt-5-mini | 94.0% |
| 19 | gpt-5.4-mini | 90.0% |
| Rank? | Model? | Pass rate? |
| 1 | 13 models tied at 100% | 100.0% |
| 14 | kimi-k2p7-code | 98.2% |
| 14 | glm-5p2 | 98.2% |
| 16 | gpt-5.6-terra | 96.4% |
| 17 | gpt-oss-120b | 92.7% |
| 18 | gpt-5-mini | 90.9% |
| 19 | gpt-5.4-mini | 81.8% |
| Rank? | Model? | Pass rate? |
| 1 | 14 models tied at 100% | 100.0% |
| 15 | claude-haiku-4-5-20251001 | 97.8% |
| 15 | gpt-5-mini | 97.8% |
| 15 | claude-opus-4-8 | 97.8% |
| 15 | gpt-oss-20b | 97.8% |
| 19 | gpt-5.4-mini | 80.0% |
Per-Model Results
Drill into any model to see per-task pass rates across its attempted tasks. This build contains aggregate results only (no per-trial pages).