Showing Solo subtasks created as medium — 90 tasks, 8,549 trials. Back to the default 20-subtask chain view.
APIFlow-Bench 1.0 Leaderboard
Deterministic API engineering benchmark spanning 7 capability axes, over REST APIs only (GraphQL and MCP are planned). Showing the solo subtasks task set: 8,549 trials across 19 models and 90 tasks covering all 7 axes — authentication, discovery, error recovery, multi-step workflow, pagination, schema, statefulness. Switch task set above; click any card to drill into a slice with the full interactive detail view; click any model name to see its per-task results.
☑ 8,549 trials
▤ 19 models
⛁ 90 tasks
⌖ Overall pass rate: 92.0%
Overall Ranking
Columns are grouped into metric families.
Pass rate = passed ÷ all trials;
API recovery = passed ÷ trials that hit a 4xx/5xx. Difficulty bucket is judged per task from its content and reference solve transcripts (Easy / Long-horizon / Hard / Hard + Long-horizon) — a separate taxonomy from the creation-time “Difficulty (as created)” filter, so the two don't map one-to-one.
Per-trial details (verifier groups, checkpoints, step counts) live on the model → task → trial pages.
| Rank | Model | Scoring | Difficulty bucket (pass rate) | Horizon (pass rate) | Cost / efficiency | Behavioral |
| Pass rate? | API recovery? | Easy? | Long-horizon? | Hard? | Short? | Medium? | Long? | Tool calls / trial? | Time / trial? | Tokens / trial? | Cost / trial? | Abstain %? | Refusal %? |
| 1 | deepseek-v4-flash | 96.2%433/450 | 135/150 | 100%170/170 | 99%129/130 | 92%134/145 | 100%225/225 | 94%203/215 | 100%5/5 | 20.5 | 1.0m | 256.9K | $0.020 | 0.4%2/450 | 0.0%0/450 |
| 2 | claude-opus-4-8 | 96.0%432/450 | 113/129 | 100%170/170 | 100%130/130 | 91%132/145 | 100%225/225 | 94%202/215 | 100%5/5 | 11.7 | 47.2s | 69.1K | $0.084 | 2.2%10/450 | 0.0%0/450 |
| 3 | qwen3p6-plus | 95.6%430/450 | 118/135 | 99%169/170 | 100%130/130 | 90%131/145 | 100%224/225 | 93%201/215 | 100%5/5 | 16.3 | 1.4m | 144.8K | $0.056 | 0.4%2/450 | 0.0%0/450 |
| 4 | claude-sonnet-4-6 | 95.3%429/450 | 89/106 | 100%170/170 | 100%130/130 | 89%129/145 | 98%220/225 | 95%204/215 | 100%5/5 | 15.2 | 52.7s | 226.4K | $0.096 | 0.2%1/450 | 0.0%0/450 |
| 4 | kimi-k2p6 | 95.3%429/450 | 121/140 | 100%170/170 | 100%130/130 | 89%129/145 | 100%225/225 | 93%199/215 | 100%5/5 | 15.7 | 1.2m | 210.6K | $0.069 | 0.2%1/450 | 0.0%0/450 |
| 6 | claude-sonnet-5 | 94.9%427/450 | 111/133 | 100%170/170 | 100%130/130 | 88%127/145 | 100%225/225 | 92%197/215 | 100%5/5 | 14.6 | 53.1s | 131.5K | $0.056 | 1.8%8/450 | 0.0%0/450 |
| 7 | kimi-k2p7-code | 94.4%425/450 | 101/122 | 100%170/170 | 100%130/130 | 86%125/145 | 100%224/225 | 91%196/215 | 100%5/5 | 14.0 | 2.5m | 119.0K | $0.039 | 1.8%8/450 | 0.0%0/450 |
| 8 | kimi-k2p5 | 94.0%423/450 | 118/141 | 99%169/170 | 100%130/130 | 86%124/145 | 99%222/225 | 91%196/215 | 100%5/5 | 12.5 | 32.9s | 100.9K | $0.038 | 1.6%7/450 | 0.0%0/450 |
| 9 | gpt-5.5 | 93.8%422/450 | 93/114 | 100%170/170 | 100%130/130 | 84%122/145 | 98%221/225 | 91%196/215 | 100%5/5 | 13.8 | 36.3s | 47.9K | $0.100 | 1.8%8/450 | 0.0%0/450 |
| 9 | deepseek-v4-pro | 93.8%422/450 | 134/161 | 99%168/170 | 98%128/130 | 87%126/145 | 99%223/225 | 90%194/215 | 100%5/5 | 19.0 | 3.3m | 113.6K | $0.11 | 0.2%1/450 | 0.0%0/450 |
| 11 | glm-5p2 | 93.1%419/450 | 98/125 | 99%169/170 | 100%130/130 | 83%120/145 | 99%222/225 | 89%192/215 | 100%5/5 | 15.1 | 3.8m | 121.5K | $0.089 | 0.0%0/450 | 0.0%0/450 |
| 12 | claude-haiku-4-5-20251001 | 92.7%417/450 | 142/165 | 99%168/170 | 100%130/130 | 82%119/145 | 98%220/225 | 89%192/215 | 100%5/5 | 16.2 | 35.5s | 112.8K | $0.028 | 1.6%7/450 | 0.0%0/450 |
| 13 | qwen3p7-plus | 92.2%415/450 | 97/118 | 98%166/170 | 96%124/129 | 87%125/143 | 96%217/225 | 92%194/212 | 80%4/5 | 12.1 | 4.0m | 147.1K | $0.038 | 2.2%10/450 | 0.0%0/450 |
| 14 | minimax-m2p7 | 91.6%412/450 | 145/171 | 97%165/170 | 98%127/130 | 83%120/145 | 96%217/225 | 88%190/215 | 100%5/5 | 12.4 | 28.5s | 62.4K | $0.015 | 0.4%2/450 | 0.0%0/450 |
| 15 | gpt-5.6-terra | 90.9%409/450 | 83/114 | 99%169/170 | 98%128/130 | 77%112/145 | 97%219/225 | 86%185/215 | 100%5/5 | 11.9 | 1.1m | 32.6K | $0.034 | 1.8%8/450 | 0.0%0/450 |
| 16 | gpt-oss-120b | 88.7%399/450 | 103/132 | 95%162/170 | 95%124/130 | 78%113/145 | 94%212/225 | 85%182/215 | 100%5/5 | 12.5 | 22.5s | 55.6K | $0.006 | 0.0%0/450 | 0.0%0/450 |
| 17 | gpt-5-mini | 88.0%396/450 | 103/130 | 94%160/170 | 96%125/130 | 77%111/145 | 94%211/225 | 84%180/215 | 100%5/5 | 11.1 | 32.3s | 48.6K | $0.006 | 5.6%25/450 | 0.0%0/450 |
| 18 | gpt-oss-20b | 87.3%392/449 | 116/152 | 94%159/170 | 93%120/129 | 78%113/145 | 93%209/225 | 84%181/215 | 50%2/4 | 14.0 | 1.6m | 66.9K | $0.003 | 2.9%13/449 | 0.0%0/449 |
| 19 | gpt-5.4-mini | 75.1%338/450 | 105/164 | 81%137/170 | 83%108/130 | 64%93/145 | 78%176/225 | 75%162/215 | 0%0/5 | 10.3 | 14.2s | 20.9K | $0.008 | 2.4%11/450 | 0.0%0/450 |
Slices
▥Behavior Mix
Top 10 of 19 models
passed?failed?abstained (clarify / report_blocked)?no-score?
| Rank? | Model? | Mix? | Pass rate? |
| 1 | deepseek-v4-flash |
| 96.2% |
| 2 | claude-opus-4-8 |
| 96.0% |
| 3 | qwen3p6-plus |
| 95.6% |
| 4 | claude-sonnet-4-6 |
| 95.3% |
| 4 | kimi-k2p6 |
| 95.3% |
| 6 | claude-sonnet-5 |
| 94.9% |
| 7 | kimi-k2p7-code |
| 94.4% |
| 8 | kimi-k2p5 |
| 94.0% |
| 9 | gpt-5.5 |
| 93.8% |
| 9 | deepseek-v4-pro |
| 93.8% |
🔐Authentication
1,330 trials
| Rank? | Model? | Pass rate? |
| 1 | claude-sonnet-4-6 | 98.6% |
| 1 | deepseek-v4-flash | 98.6% |
| 3 | deepseek-v4-pro | 97.1% |
| 4 | claude-sonnet-5 | 95.7% |
| 5 | gpt-5.5 | 94.3% |
| 6 | claude-opus-4-8 | 92.9% |
| 6 | kimi-k2p6 | 92.9% |
| 6 | kimi-k2p7-code | 92.9% |
| 6 | glm-5p2 | 92.9% |
| 10 | claude-haiku-4-5-20251001 | 90.0% |
+9 more models below — click for the full ranking of all 19.
| Rank? | Model? | Pass rate? |
| 1 | 6 models tied at 100% | 100.0% |
| 7 | kimi-k2p7-code | 98.2% |
| 7 | glm-5p2 | 98.2% |
| 9 | kimi-k2p5 | 96.4% |
| 9 | claude-haiku-4-5-20251001 | 96.4% |
| 11 | minimax-m2p7 | 94.5% |
| 11 | qwen3p7-plus | 94.5% |
| 13 | gpt-5.5 | 92.7% |
| 14 | claude-sonnet-4-6 | 90.9% |
| 14 | gpt-5.6-terra | 90.9% |
| 16 | gpt-oss-120b | 89.1% |
+3 more models below — click for the full ranking of all 19.
↻Error Recovery
1,330 trials
| Rank? | Model? | Pass rate? |
| 1 | 9 models tied at 100% | 100.0% |
| 10 | minimax-m2p7 | 98.6% |
| 10 | claude-sonnet-5 | 98.6% |
| 10 | deepseek-v4-flash | 98.6% |
| 10 | qwen3p7-plus | 98.6% |
| 14 | claude-haiku-4-5-20251001 | 97.1% |
| 14 | deepseek-v4-pro | 97.1% |
| 16 | gpt-oss-120b | 95.7% |
| 17 | gpt-5-mini | 91.4% |
| 17 | gpt-oss-20b | 91.4% |
| 19 | gpt-5.4-mini | 81.4% |
⛓Multi-step Workflow
950 trials
| Rank? | Model? | Pass rate? |
| 1 | claude-opus-4-8 | 94.0% |
| 1 | qwen3p6-plus | 94.0% |
| 1 | qwen3p7-plus | 94.0% |
| 4 | claude-sonnet-4-6 | 92.0% |
| 4 | deepseek-v4-flash | 92.0% |
| 4 | kimi-k2p6 | 92.0% |
| 7 | claude-haiku-4-5-20251001 | 90.0% |
| 7 | kimi-k2p5 | 90.0% |
| 7 | claude-sonnet-5 | 90.0% |
| 7 | minimax-m2p7 | 90.0% |
+9 more models below — click for the full ranking of all 19.
| Rank? | Model? | Pass rate? |
| 1 | claude-opus-4-8 | 91.7% |
| 1 | gpt-5.5 | 91.7% |
| 1 | kimi-k2p5 | 91.7% |
| 1 | claude-sonnet-4-6 | 91.7% |
| 1 | claude-sonnet-5 | 91.7% |
| 1 | kimi-k2p7-code | 91.7% |
| 1 | kimi-k2p6 | 91.7% |
| 8 | gpt-5-mini | 90.0% |
| 8 | claude-haiku-4-5-20251001 | 90.0% |
| 8 | deepseek-v4-flash | 90.0% |
+9 more models below — click for the full ranking of all 19.
| Rank? | Model? | Pass rate? |
| 1 | 12 models tied at 100% | 100.0% |
| 13 | gpt-oss-120b | 98.0% |
| 13 | minimax-m2p7 | 98.0% |
| 13 | gpt-5-mini | 98.0% |
| 16 | claude-haiku-4-5-20251001 | 96.0% |
| 16 | gpt-oss-20b | 96.0% |
| 16 | qwen3p7-plus | 96.0% |
| 19 | gpt-5.4-mini | 88.0% |
≣Statefulness
1,804 trials
| Rank? | Model? | Pass rate? |
| 1 | qwen3p6-plus | 95.8% |
| 2 | claude-opus-4-8 | 94.7% |
| 2 | deepseek-v4-flash | 94.7% |
| 4 | claude-sonnet-4-6 | 93.7% |
| 5 | kimi-k2p5 | 92.6% |
| 5 | kimi-k2p6 | 92.6% |
| 7 | gpt-oss-120b | 91.6% |
| 7 | minimax-m2p7 | 91.6% |
| 7 | qwen3p7-plus | 91.6% |
| 10 | claude-haiku-4-5-20251001 | 90.5% |
+9 more models below — click for the full ranking of all 19.
Per-Model Results
Drill into any model to see per-task pass rates across its attempted tasks. This build contains aggregate results only (no per-trial pages).