Showing 20-subtask chain — 11 tasks, 1,045 trials.
APIFlow-Bench 1.0 Leaderboard
Deterministic API engineering benchmark spanning 7 capability axes, over REST APIs only (GraphQL and MCP are planned). Showing the 20-subtask chain task set: 1,045 trials across 19 models and 11 tasks covering 5 of the 7 axes — authentication, error recovery, multi-step workflow, pagination, statefulness (all 7 under All tasks). Switch task set above; click any card to drill into a slice with the full interactive detail view; click any model name to see its per-task results.
☑ 1,045 trials
▤ 19 models
⛁ 11 tasks
⌖ Overall pass rate: 60.9%
Overall Ranking
Columns are grouped into metric families.
Pass rate = passed ÷ all trials;
API recovery = passed ÷ trials that hit a 4xx/5xx. Difficulty/horizon breakdowns are omitted: every task in this slice is judged Hard + Long-horizon and Long-horizon.
Per-trial details (verifier groups, checkpoints, step counts) live on the model → task → trial pages.
| Rank | Model | Scoring | Cost / efficiency | Behavioral |
| Pass rate? | API recovery? | Tool calls / trial? | Time / trial? | Tokens / trial? | Cost / trial? | Abstain %? | Refusal %? |
| 1 | gpt-5.5 | 72.7%40/55 | 20/25 | 21.7 | 49.6s | 91.5K | $0.17 | 0.0%0/55 | 0.0%0/55 |
| 2 | claude-opus-4-8 | 69.1%38/55 | 18/23 | 13.4 | 44.3s | 69.7K | $0.086 | 0.0%0/55 | 0.0%0/55 |
| 2 | deepseek-v4-flash | 69.1%38/55 | 21/29 | 21.0 | 44.7s | 80.6K | $0.008 | 0.0%0/55 | 0.0%0/55 |
| 4 | glm-5p2 | 65.5%36/55 | 17/23 | 17.3 | 3.2m | 92.8K | $0.079 | 0.0%0/55 | 0.0%0/55 |
| 4 | qwen3p7-plus | 65.5%36/55 | 16/24 | 12.0 | 5.3m | 125.3K | $0.039 | 0.0%0/55 | 0.0%0/55 |
| 6 | claude-haiku-4-5-20251001 | 63.6%35/55 | 15/20 | 14.9 | 26.6s | 66.2K | $0.022 | 0.0%0/55 | 0.0%0/55 |
| 6 | gpt-5.6-terra | 63.6%35/55 | 16/22 | 14.5 | 1.3m | 41.3K | $0.043 | 0.0%0/55 | 0.0%0/55 |
| 6 | qwen3p6-plus | 63.6%35/55 | 18/28 | 14.5 | 1.1m | 72.5K | $0.035 | 0.0%0/55 | 0.0%0/55 |
| 6 | deepseek-v4-pro | 63.6%35/55 | 22/29 | 23.1 | 3.5m | 82.8K | $0.11 | 0.0%0/55 | 0.0%0/55 |
| 10 | claude-sonnet-4-6 | 61.8%34/55 | 19/25 | 14.1 | 50.4s | 70.3K | $0.048 | 0.0%0/55 | 0.0%0/55 |
| 10 | kimi-k2p7-code | 61.8%34/55 | 15/21 | 13.9 | 2.3m | 50.4K | $0.022 | 0.0%0/55 | 0.0%0/55 |
| 12 | claude-sonnet-5 | 60.0%33/55 | 16/25 | 19.3 | 1.2m | 140.1K | $0.070 | 0.0%0/55 | 0.0%0/55 |
| 13 | minimax-m2p7 | 58.2%32/55 | 18/25 | 13.6 | 19.5s | 43.7K | $0.010 | 0.0%0/55 | 0.0%0/55 |
| 13 | kimi-k2p6 | 58.2%32/55 | 16/21 | 21.8 | 1.9m | 255.9K | $0.087 | 0.0%0/55 | 0.0%0/55 |
| 15 | kimi-k2p5 | 56.4%31/55 | 15/20 | 11.6 | 22.8s | 34.6K | $0.017 | 0.0%0/55 | 0.0%0/55 |
| 15 | gpt-oss-20b | 56.4%31/55 | 18/25 | 16.7 | 2.1m | 80.3K | $0.004 | 0.0%0/55 | 0.0%0/55 |
| 17 | gpt-oss-120b | 52.7%29/55 | 15/24 | 14.8 | 25.8s | 68.8K | $0.007 | 0.0%0/55 | 0.0%0/55 |
| 18 | gpt-5-mini | 50.9%28/55 | 16/24 | 12.7 | 31.7s | 54.4K | $0.006 | 7.3%4/55 | 0.0%0/55 |
| 19 | gpt-5.4-mini | 43.6%24/55 | 14/32 | 13.8 | 16.0s | 26.1K | $0.010 | 5.5%3/55 | 0.0%0/55 |
Slices
▥Behavior Mix
Top 10 of 19 models
passed?failed?abstained (clarify / report_blocked)?no-score?
| Rank? | Model? | Mix? | Pass rate? |
| 1 | gpt-5.5 |
| 72.7% |
| 2 | claude-opus-4-8 |
| 69.1% |
| 2 | deepseek-v4-flash |
| 69.1% |
| 4 | glm-5p2 |
| 65.5% |
| 4 | qwen3p7-plus |
| 65.5% |
| 6 | claude-haiku-4-5-20251001 |
| 63.6% |
| 6 | gpt-5.6-terra |
| 63.6% |
| 6 | qwen3p6-plus |
| 63.6% |
| 6 | deepseek-v4-pro |
| 63.6% |
| 10 | claude-sonnet-4-6 |
| 61.8% |
🔐Authentication
285 trials
| Rank? | Model? | Pass rate? |
| 1 | gpt-5.5 | 33.3% |
| 2 | claude-sonnet-4-6 | 26.7% |
| 3 | claude-opus-4-8 | 20.0% |
| 4 | claude-sonnet-5 | 6.7% |
| 4 | gpt-oss-20b | 6.7% |
| 4 | deepseek-v4-flash | 6.7% |
| 4 | qwen3p7-plus | 6.7% |
| 4 | glm-5p2 | 6.7% |
| 9 | gpt-5.4-mini | 0.0% |
| 9 | minimax-m2p7 | 0.0% |
+9 more models below — click for the full ranking of all 19.
↻Error Recovery
190 trials
| Rank? | Model? | Pass rate? |
| 1 | deepseek-v4-flash | 70.0% |
| 2 | qwen3p6-plus | 60.0% |
| 3 | minimax-m2p7 | 50.0% |
| 3 | claude-haiku-4-5-20251001 | 50.0% |
| 3 | kimi-k2p5 | 50.0% |
| 3 | gpt-5-mini | 50.0% |
| 3 | claude-sonnet-4-6 | 50.0% |
| 3 | claude-opus-4-8 | 50.0% |
| 3 | gpt-oss-20b | 50.0% |
| 3 | gpt-5.6-terra | 50.0% |
+9 more models below — click for the full ranking of all 19.
⛓Multi-step Workflow
380 trials
| Rank? | Model? | Pass rate? |
| 1 | 8 models tied at 100% | 100.0% |
| 9 | qwen3p6-plus | 95.0% |
| 9 | kimi-k2p7-code | 95.0% |
| 11 | minimax-m2p7 | 85.0% |
| 11 | claude-sonnet-5 | 85.0% |
| 11 | kimi-k2p6 | 85.0% |
| 14 | gpt-5.4-mini | 80.0% |
| 14 | kimi-k2p5 | 80.0% |
| 14 | gpt-5-mini | 80.0% |
| 17 | gpt-oss-120b | 75.0% |
| 17 | claude-sonnet-4-6 | 75.0% |
+1 more model below — click for the full ranking of all 19.
| Rank? | Model? | Pass rate? |
| 1 | 17 models tied at 100% | 100.0% |
| 18 | gpt-5-mini | 40.0% |
| 19 | gpt-5.4-mini | 20.0% |
| Rank? | Model? | Pass rate? |
| 1 | All 19 models tied at 100% | 100.0% |
Per-Model Results
Drill into any model to see per-task pass rates across its attempted tasks. This build contains aggregate results only (no per-trial pages).