Postman APIFlow-Bench 1.0 Leaderboard
Task set ?20-subtask chain (11)15-subtask chain (11)10-subtask chain (12)5-subtask chain (13)Solo subtasks (241)All tasks (467)

Showing 20-subtask chain — 11 tasks, 1,045 trials.

APIFlow-Bench 1.0 Leaderboard

Deterministic API engineering benchmark spanning 7 capability axes, over REST APIs only (GraphQL and MCP are planned). Showing the 20-subtask chain task set: 1,045 trials across 19 models and 11 tasks covering 5 of the 7 axes — authentication, error recovery, multi-step workflow, pagination, statefulness (all 7 under All tasks). Switch task set above; click any card to drill into a slice with the full interactive detail view; click any model name to see its per-task results.

☑ 1,045 trials ▤ 19 models ⛁ 11 tasks ⌖ Overall pass rate: 60.9%

Overall Ranking

Columns are grouped into metric families. Pass rate = passed ÷ all trials; API recovery = passed ÷ trials that hit a 4xx/5xx. Difficulty/horizon breakdowns are omitted: every task in this slice is judged Hard + Long-horizon and Long-horizon. Per-trial details (verifier groups, checkpoints, step counts) live on the model → task → trial pages.

RankModelScoringCost / efficiencyBehavioral
Pass rate?API recovery?Tool calls / trial?Time / trial?Tokens / trial?Cost / trial?Abstain %?Refusal %?
1gpt-5.572.7%40/5520/2521.749.6s91.5K$0.170.0%0/550.0%0/55
2claude-opus-4-869.1%38/5518/2313.444.3s69.7K$0.0860.0%0/550.0%0/55
2deepseek-v4-flash69.1%38/5521/2921.044.7s80.6K$0.0080.0%0/550.0%0/55
4glm-5p265.5%36/5517/2317.33.2m92.8K$0.0790.0%0/550.0%0/55
4qwen3p7-plus65.5%36/5516/2412.05.3m125.3K$0.0390.0%0/550.0%0/55
6claude-haiku-4-5-2025100163.6%35/5515/2014.926.6s66.2K$0.0220.0%0/550.0%0/55
6gpt-5.6-terra63.6%35/5516/2214.51.3m41.3K$0.0430.0%0/550.0%0/55
6qwen3p6-plus63.6%35/5518/2814.51.1m72.5K$0.0350.0%0/550.0%0/55
6deepseek-v4-pro63.6%35/5522/2923.13.5m82.8K$0.110.0%0/550.0%0/55
10claude-sonnet-4-661.8%34/5519/2514.150.4s70.3K$0.0480.0%0/550.0%0/55
10kimi-k2p7-code61.8%34/5515/2113.92.3m50.4K$0.0220.0%0/550.0%0/55
12claude-sonnet-560.0%33/5516/2519.31.2m140.1K$0.0700.0%0/550.0%0/55
13minimax-m2p758.2%32/5518/2513.619.5s43.7K$0.0100.0%0/550.0%0/55
13kimi-k2p658.2%32/5516/2121.81.9m255.9K$0.0870.0%0/550.0%0/55
15kimi-k2p556.4%31/5515/2011.622.8s34.6K$0.0170.0%0/550.0%0/55
15gpt-oss-20b56.4%31/5518/2516.72.1m80.3K$0.0040.0%0/550.0%0/55
17gpt-oss-120b52.7%29/5515/2414.825.8s68.8K$0.0070.0%0/550.0%0/55
18gpt-5-mini50.9%28/5516/2412.731.7s54.4K$0.0067.3%4/550.0%0/55
19gpt-5.4-mini43.6%24/5514/3213.816.0s26.1K$0.0105.5%3/550.0%0/55

Slices

Behavior Mix
Top 10 of 19 models
passed?failed?abstained (clarify / report_blocked)?no-score?
Rank?Model?Mix?Pass rate?
1gpt-5.5
72.7%
2claude-opus-4-8
69.1%
2deepseek-v4-flash
69.1%
4glm-5p2
65.5%
4qwen3p7-plus
65.5%
6claude-haiku-4-5-20251001
63.6%
6gpt-5.6-terra
63.6%
6qwen3p6-plus
63.6%
6deepseek-v4-pro
63.6%
10claude-sonnet-4-6
61.8%
🔐Authentication
285 trials
Rank?Model?Pass rate?
1gpt-5.533.3%
2claude-sonnet-4-626.7%
3claude-opus-4-820.0%
4claude-sonnet-56.7%
4gpt-oss-20b6.7%
4deepseek-v4-flash6.7%
4qwen3p7-plus6.7%
4glm-5p26.7%
9gpt-5.4-mini0.0%
9minimax-m2p70.0%
+9 more models below — click for the full ranking of all 19.
Error Recovery
190 trials
Rank?Model?Pass rate?
1deepseek-v4-flash70.0%
2qwen3p6-plus60.0%
3minimax-m2p750.0%
3claude-haiku-4-5-2025100150.0%
3kimi-k2p550.0%
3gpt-5-mini50.0%
3claude-sonnet-4-650.0%
3claude-opus-4-850.0%
3gpt-oss-20b50.0%
3gpt-5.6-terra50.0%
+9 more models below — click for the full ranking of all 19.
Multi-step Workflow
380 trials
Rank?Model?Pass rate?
18 models tied at 100%100.0%
9qwen3p6-plus95.0%
9kimi-k2p7-code95.0%
11minimax-m2p785.0%
11claude-sonnet-585.0%
11kimi-k2p685.0%
14gpt-5.4-mini80.0%
14kimi-k2p580.0%
14gpt-5-mini80.0%
17gpt-oss-120b75.0%
17claude-sonnet-4-675.0%
+1 more model below — click for the full ranking of all 19.
Pagination
95 trials
Rank?Model?Pass rate?
117 models tied at 100%100.0%
18gpt-5-mini40.0%
19gpt-5.4-mini20.0%
Statefulness
95 trials
Rank?Model?Pass rate?
1All 19 models tied at 100%100.0%

Per-Model Results

Drill into any model to see per-task pass rates across its attempted tasks. This build contains aggregate results only (no per-trial pages).

claude-haiku-4-5-20251001
55 trials· 63.6%
claude-opus-4-8
55 trials· 69.1%
claude-sonnet-4-6
55 trials· 61.8%
claude-sonnet-5
55 trials· 60.0%
deepseek-v4-flash
55 trials· 69.1%
deepseek-v4-pro
55 trials· 63.6%
glm-5p2
55 trials· 65.5%
gpt-5-mini
55 trials· 50.9%
gpt-5.4-mini
55 trials· 43.6%
gpt-5.5
55 trials· 72.7%
gpt-5.6-terra
55 trials· 63.6%
gpt-oss-120b
55 trials· 52.7%
gpt-oss-20b
55 trials· 56.4%
kimi-k2p5
55 trials· 56.4%
kimi-k2p6
55 trials· 58.2%
kimi-k2p7-code
55 trials· 61.8%
minimax-m2p7
55 trials· 58.2%
qwen3p6-plus
55 trials· 63.6%
qwen3p7-plus
55 trials· 65.5%