Postman APIFlow-Bench 1.0 Leaderboard
Task set ?20-subtask chain (11)15-subtask chain (11)10-subtask chain (12)5-subtask chain (13)Solo subtasks (241)All tasks (467)

Showing 5-subtask chain — 13 tasks, 1,235 trials. Back to the default 20-subtask chain view.

APIFlow-Bench 1.0 Leaderboard

Deterministic API engineering benchmark spanning 7 capability axes, over REST APIs only (GraphQL and MCP are planned). Showing the 5-subtask chain task set: 1,235 trials across 19 models and 13 tasks covering all 7 axes — authentication, discovery, error recovery, multi-step workflow, pagination, schema, statefulness. Switch task set above; click any card to drill into a slice with the full interactive detail view; click any model name to see its per-task results.

☑ 1,235 trials ▤ 19 models ⛁ 13 tasks ⌖ Overall pass rate: 89.8%

Overall Ranking

Columns are grouped into metric families. Pass rate = passed ÷ all trials; API recovery = passed ÷ trials that hit a 4xx/5xx. Difficulty bucket is judged per task from its content and reference solve transcripts (Easy / Long-horizon / Hard / Hard + Long-horizon) — a separate taxonomy from the creation-time “Difficulty (as created)” filter, so the two don't map one-to-one. Difficulty/horizon breakdowns are omitted: every task in this slice is Long-horizon. Per-trial details (verifier groups, checkpoints, step counts) live on the model → task → trial pages.

RankModelScoringDifficulty bucket (pass rate)Cost / efficiencyBehavioral
Pass rate?API recovery?Long-horizon?Hard + Long-horizon?Tool calls / trial?Time / trial?Tokens / trial?Cost / trial?Abstain %?Refusal %?
1kimi-k2p6100.0%65/6535/35100%5/5100%60/6022.11.8m193.8K$0.0880.0%0/650.0%0/65
2deepseek-v4-flash98.5%64/6536/37100%5/598%59/6021.957.6s101.0K$0.0090.0%0/650.0%0/65
3claude-haiku-4-5-2025100196.9%63/6531/33100%5/597%58/6032.055.5s333.1K$0.0580.0%0/650.0%0/65
4kimi-k2p7-code95.4%62/6532/35100%5/595%57/6016.93.0m73.4K$0.0320.0%0/650.0%0/65
5kimi-k2p592.3%60/6530/35100%5/592%55/6012.831.5s44.6K$0.0210.0%0/650.0%0/65
5claude-opus-4-892.3%60/6526/31100%5/592%55/6013.745.4s74.4K$0.0880.0%0/650.0%0/65
5minimax-m2p792.3%60/6533/38100%5/592%55/6018.330.6s103.8K$0.0231.5%1/650.0%0/65
5gpt-5.592.3%60/6528/33100%5/592%55/6018.645.0s68.0K$0.140.0%0/650.0%0/65
5deepseek-v4-pro92.3%60/6540/45100%5/592%55/6025.53.9m111.9K$0.130.0%0/650.0%0/65
5glm-5p292.3%60/6530/35100%5/592%55/6020.65.8m243.4K$0.150.0%0/650.0%0/65
11gpt-5.6-terra90.8%59/6525/30100%5/590%54/6016.21.6m50.3K$0.0500.0%0/650.0%0/65
11qwen3p7-plus90.8%59/6527/32100%5/590%54/6013.04.7m153.1K$0.0420.0%0/650.0%0/65
13gpt-oss-120b86.2%56/6530/39100%5/585%51/6016.527.3s80.5K$0.0080.0%0/650.0%0/65
13gpt-oss-20b86.2%56/6534/42100%5/585%51/6017.61.5m85.8K$0.0046.2%4/650.0%0/65
13qwen3p6-plus86.2%56/6528/35100%5/585%51/6026.22.1m224.0K$0.0890.0%0/650.0%0/65
13claude-sonnet-4-686.2%56/6521/30100%5/585%51/6077.74.8m9.4M$2.950.0%0/650.0%0/65
17claude-sonnet-584.6%55/6523/33100%5/583%50/6019.81.0m118.8K$0.0590.0%0/650.0%0/65
18gpt-5-mini83.1%54/6531/34100%5/582%49/6013.334.5s59.4K$0.00713.8%9/650.0%0/65
19gpt-5.4-mini67.7%44/6521/36100%5/565%39/6013.617.3s28.3K$0.0104.6%3/650.0%0/65

Slices

Behavior Mix
Top 10 of 19 models
passed?failed?abstained (clarify / report_blocked)?no-score?
Rank?Model?Mix?Pass rate?
1kimi-k2p6
100.0%
2deepseek-v4-flash
98.5%
3claude-haiku-4-5-20251001
96.9%
4kimi-k2p7-code
95.4%
5kimi-k2p5
92.3%
5claude-opus-4-8
92.3%
5minimax-m2p7
92.3%
5gpt-5.5
92.3%
5deepseek-v4-pro
92.3%
5glm-5p2
92.3%
🔐Authentication
95 trials
Rank?Model?Pass rate?
115 models tied at 100%100.0%
16claude-opus-4-880.0%
17gpt-oss-20b60.0%
18gpt-5.4-mini20.0%
19claude-sonnet-50.0%
🔎Discovery
190 trials
Rank?Model?Pass rate?
115 models tied at 100%100.0%
16gpt-oss-120b90.0%
16gpt-oss-20b90.0%
18qwen3p6-plus80.0%
19gpt-5.4-mini50.0%
Error Recovery
475 trials
Rank?Model?Pass rate?
115 models tied at 100%100.0%
16gpt-oss-120b96.0%
16gpt-5.6-terra96.0%
16gpt-oss-20b96.0%
19gpt-5-mini92.0%
Multi-step Workflow
95 trials
Rank?Model?Pass rate?
117 models tied at 100%100.0%
18gpt-5.4-mini80.0%
19gpt-5-mini0.0%
Pagination
95 trials
Rank?Model?Pass rate?
113 models tied at 100%100.0%
14minimax-m2p780.0%
14gpt-oss-20b80.0%
16gpt-5-mini60.0%
16qwen3p6-plus60.0%
18gpt-5.4-mini40.0%
19claude-sonnet-4-60.0%
Schema
95 trials
Rank?Model?Pass rate?
1kimi-k2p6100.0%
2gpt-5-mini80.0%
2gpt-oss-20b80.0%
2deepseek-v4-flash80.0%
5claude-haiku-4-5-2025100160.0%
6kimi-k2p7-code40.0%
7claude-opus-4-820.0%
7minimax-m2p720.0%
7claude-sonnet-4-620.0%
10gpt-5.4-mini0.0%
+9 more models below — click for the full ranking of all 19.
Statefulness
190 trials
Rank?Model?Pass rate?
114 models tied at 100%100.0%
15gpt-5-mini90.0%
15qwen3p7-plus90.0%
17gpt-oss-120b80.0%
18gpt-5.4-mini70.0%
18gpt-oss-20b70.0%

Per-Model Results

Drill into any model to see per-task pass rates across its attempted tasks. This build contains aggregate results only (no per-trial pages).

claude-haiku-4-5-20251001
65 trials· 96.9%
claude-opus-4-8
65 trials· 92.3%
claude-sonnet-4-6
65 trials· 86.2%
claude-sonnet-5
65 trials· 84.6%
deepseek-v4-flash
65 trials· 98.5%
deepseek-v4-pro
65 trials· 92.3%
glm-5p2
65 trials· 92.3%
gpt-5-mini
65 trials· 83.1%
gpt-5.4-mini
65 trials· 67.7%
gpt-5.5
65 trials· 92.3%
gpt-5.6-terra
65 trials· 90.8%
gpt-oss-120b
65 trials· 86.2%
gpt-oss-20b
65 trials· 86.2%
kimi-k2p5
65 trials· 92.3%
kimi-k2p6
65 trials· 100.0%
kimi-k2p7-code
65 trials· 95.4%
minimax-m2p7
65 trials· 92.3%
qwen3p6-plus
65 trials· 86.2%
qwen3p7-plus
65 trials· 90.8%