Postman APIFlow-Bench 1.0 Leaderboard
Task set ?20-subtask chain (11)15-subtask chain (11)10-subtask chain (12)5-subtask chain (13)Solo subtasks (241)All tasks (467)

Showing 10-subtask chain — 12 tasks, 1,140 trials. Back to the default 20-subtask chain view.

APIFlow-Bench 1.0 Leaderboard

Deterministic API engineering benchmark spanning 7 capability axes, over REST APIs only (GraphQL and MCP are planned). Showing the 10-subtask chain task set: 1,140 trials across 19 models and 12 tasks covering 5 of the 7 axes — authentication, discovery, error recovery, schema, statefulness (all 7 under All tasks). Switch task set above; click any card to drill into a slice with the full interactive detail view; click any model name to see its per-task results.

☑ 1,140 trials ▤ 19 models ⛁ 12 tasks ⌖ Overall pass rate: 70.6%

Overall Ranking

Columns are grouped into metric families. Pass rate = passed ÷ all trials; API recovery = passed ÷ trials that hit a 4xx/5xx. Difficulty/horizon breakdowns are omitted: every task in this slice is judged Hard + Long-horizon and Long-horizon. Per-trial details (verifier groups, checkpoints, step counts) live on the model → task → trial pages.

RankModelScoringCost / efficiencyBehavioral
Pass rate?API recovery?Tool calls / trial?Time / trial?Tokens / trial?Cost / trial?Abstain %?Refusal %?
1minimax-m2p775.0%45/6024/2812.419.6s41.5K$0.0090.0%0/600.0%0/60
1claude-haiku-4-5-2025100175.0%45/6020/2215.129.7s67.2K$0.0220.0%0/600.0%0/60
1claude-sonnet-4-675.0%45/6024/2410.033.0s37.3K$0.0270.0%0/600.0%0/60
1gpt-5.575.0%45/6021/2112.328.5s28.8K$0.0730.0%0/600.0%0/60
1qwen3p6-plus75.0%45/6025/2912.938.8s52.3K$0.0220.0%0/600.0%0/60
1deepseek-v4-flash75.0%45/6024/2516.941.4s56.5K$0.0060.0%0/600.0%0/60
1kimi-k2p675.0%45/6022/2513.345.2s55.8K$0.0280.0%0/600.0%0/60
1glm-5p275.0%45/6023/2613.02.0m44.5K$0.0390.0%0/600.0%0/60
1deepseek-v4-pro75.0%45/6029/3818.12.8m66.7K$0.0870.0%0/600.0%0/60
10claude-sonnet-573.3%44/6025/2814.335.6s70.7K$0.0320.0%0/600.0%0/60
10gpt-oss-20b73.3%44/6022/2514.61.4m67.3K$0.0030.0%0/600.0%0/60
10kimi-k2p7-code73.3%44/6024/2611.81.9m32.1K$0.0150.0%0/600.0%0/60
10qwen3p7-plus73.3%44/6023/249.93.0m66.3K$0.0180.0%0/600.0%0/60
14gpt-5.6-terra71.7%43/6025/2513.51.2m35.3K$0.0350.0%0/600.0%0/60
15kimi-k2p568.3%41/6024/2510.528.8s31.9K$0.0150.0%0/600.0%0/60
15claude-opus-4-868.3%41/6025/2610.537.5s53.3K$0.0630.0%0/600.0%0/60
17gpt-oss-120b61.7%37/6023/2813.124.8s56.1K$0.0060.0%0/600.0%0/60
18gpt-5-mini58.3%35/6016/1910.633.0s44.2K$0.00511.7%7/600.0%0/60
19gpt-5.4-mini45.0%27/6015/2911.014.7s22.0K$0.0086.7%4/600.0%0/60

Slices

Behavior Mix
Top 10 of 19 models
passed?failed?abstained (clarify / report_blocked)?no-score?
Rank?Model?Mix?Pass rate?
1minimax-m2p7
75.0%
1claude-haiku-4-5-20251001
75.0%
1claude-sonnet-4-6
75.0%
1gpt-5.5
75.0%
1qwen3p6-plus
75.0%
1deepseek-v4-flash
75.0%
1kimi-k2p6
75.0%
1glm-5p2
75.0%
1deepseek-v4-pro
75.0%
10claude-sonnet-5
73.3%
🔐Authentication
190 trials
Rank?Model?Pass rate?
19 models tied at 100%100.0%
10claude-sonnet-590.0%
10deepseek-v4-flash90.0%
10kimi-k2p7-code90.0%
13gpt-5.6-terra80.0%
13gpt-oss-20b80.0%
15gpt-5-mini70.0%
16kimi-k2p560.0%
16claude-opus-4-860.0%
18gpt-5.4-mini30.0%
19gpt-oss-120b20.0%
🔎Discovery
475 trials
Rank?Model?Pass rate?
1deepseek-v4-flash64.0%
1gpt-oss-20b64.0%
3minimax-m2p760.0%
3claude-sonnet-560.0%
3claude-sonnet-4-660.0%
3gpt-oss-120b60.0%
3gpt-5.560.0%
3claude-opus-4-860.0%
3kimi-k2p560.0%
3claude-haiku-4-5-2025100160.0%
+9 more models below — click for the full ranking of all 19.
Error Recovery
190 trials
Rank?Model?Pass rate?
1minimax-m2p750.0%
1kimi-k2p550.0%
1gpt-oss-120b50.0%
1gpt-5.550.0%
1claude-haiku-4-5-2025100150.0%
1deepseek-v4-flash50.0%
1claude-sonnet-4-650.0%
1qwen3p6-plus50.0%
1claude-sonnet-550.0%
1claude-opus-4-850.0%
+9 more models below — click for the full ranking of all 19.
Schema
95 trials
Rank?Model?Pass rate?
116 models tied at 100%100.0%
17gpt-5.4-mini80.0%
17gpt-5-mini80.0%
17qwen3p7-plus80.0%
Statefulness
190 trials
Rank?Model?Pass rate?
118 models tied at 100%100.0%
19gpt-5.4-mini50.0%

Per-Model Results

Drill into any model to see per-task pass rates across its attempted tasks. This build contains aggregate results only (no per-trial pages).

claude-haiku-4-5-20251001
60 trials· 75.0%
claude-opus-4-8
60 trials· 68.3%
claude-sonnet-4-6
60 trials· 75.0%
claude-sonnet-5
60 trials· 73.3%
deepseek-v4-flash
60 trials· 75.0%
deepseek-v4-pro
60 trials· 75.0%
glm-5p2
60 trials· 75.0%
gpt-5-mini
60 trials· 58.3%
gpt-5.4-mini
60 trials· 45.0%
gpt-5.5
60 trials· 75.0%
gpt-5.6-terra
60 trials· 71.7%
gpt-oss-120b
60 trials· 61.7%
gpt-oss-20b
60 trials· 73.3%
kimi-k2p5
60 trials· 68.3%
kimi-k2p6
60 trials· 75.0%
kimi-k2p7-code
60 trials· 73.3%
minimax-m2p7
60 trials· 75.0%
qwen3p6-plus
60 trials· 75.0%
qwen3p7-plus
60 trials· 73.3%