Postman

APIFlow-Bench 1.0 Leaderboard

5-subtask chainHard

Deterministic API engineering benchmark spanning 7 capability axes, over REST APIs only (GraphQL and MCP are planned). Showing the 5-subtask chain task set: 1,560 trials across 24 models and 13 tasks covering all 7 axes — authentication, discovery, error recovery, multi-step workflow, pagination, schema, statefulness. Switch task set above; click any card to drill into a slice with the full interactive detail view; click any model name to see its per-task results.

Overall pass rate83.8%Trials1,560Models24Tasks13

Overall Ranking

Columns are grouped into metric families. Pass rate = passed ÷ all trials; API recovery = passed ÷ trials that hit a 4xx/5xx. Difficulty bucket is judged per task from its content and reference solve transcripts (Easy / Long-horizon / Hard / Hard + Long-horizon) — a separate taxonomy from the creation-time “Difficulty (as created)” filter, so the two don't map one-to-one. Difficulty/horizon breakdowns are omitted: every task in this slice is Long-horizon. Per-trial details (verifier groups, checkpoints, step counts) live on the model → task → trial pages.

Rank
Model
Scoring
Difficulty bucket (pass rate)
Cost / efficiency
Behavioral
Pass rate
API recovery
Long-horizon
Hard + Long-horizon
Tool calls / trial
Time / trial
Tokens / trial
Cost / trial
Abstain %
Refusal %
2deepseek-v4-flash98.5%64/6536/37100%5/598%59/6021.957.6s101.0K$0.0090.0%0/650.0%0/65
3claude-haiku-4-5-2025100196.9%63/6531/33100%5/597%58/6032.055.5s333.1K$0.0580.0%0/650.0%0/65
4kimi-k2p7-code95.4%62/6532/35100%5/595%57/6016.93.0m73.4K$0.0320.0%0/650.0%0/65
5kimi-k393.8%61/6533/37100%5/593%56/6020.213.1m164.6K$0.200.0%0/650.0%0/65
6kimi-k2p592.3%60/6530/35100%5/592%55/6012.831.5s44.6K$0.0210.0%0/650.0%0/65
6claude-opus-4-892.3%60/6526/31100%5/592%55/6013.745.4s74.4K$0.0880.0%0/650.0%0/65
6minimax-m2p792.3%60/6533/38100%5/592%55/6018.330.6s103.8K$0.0231.5%1/650.0%0/65
6gpt-5.592.3%60/6528/33100%5/592%55/6018.645.0s68.0K$0.140.0%0/650.0%0/65
6gpt-5.6-sol92.3%60/6525/30100%5/592%55/6016.33.0m48.5K$0.110.0%0/650.0%0/65
6deepseek-v4-pro92.3%60/6540/45100%5/592%55/6025.53.9m111.9K$0.130.0%0/650.0%0/65
6glm-5p292.3%60/6530/35100%5/592%55/6020.65.8m243.4K$0.150.0%0/650.0%0/65
13gpt-5.6-terra90.8%59/6525/30100%5/590%54/6016.21.6m50.3K$0.0500.0%0/650.0%0/65
13qwen3p7-plus90.8%59/6527/32100%5/590%54/6013.04.7m153.1K$0.0420.0%0/650.0%0/65
15gpt-5.6-luna87.7%57/6522/30100%5/587%52/6015.01.8m40.5K$0.0160.0%0/650.0%0/65
16gpt-oss-120b86.2%56/6530/39100%5/585%51/6016.527.3s80.5K$0.0080.0%0/650.0%0/65
16gpt-oss-20b86.2%56/6534/42100%5/585%51/6017.61.5m85.8K$0.0046.2%4/650.0%0/65
16qwen3p6-plus86.2%56/6528/35100%5/585%51/6026.22.1m224.0K$0.0890.0%0/650.0%0/65
16claude-sonnet-4-686.2%56/6521/30100%5/585%51/6077.74.8m9.4M$2.950.0%0/650.0%0/65
20claude-sonnet-584.6%55/6523/33100%5/583%50/6019.81.0m118.8K$0.0590.0%0/650.0%0/65
21gpt-5-mini83.1%54/6531/34100%5/582%49/6013.334.5s59.4K$0.00713.8%9/650.0%0/65
22gpt-5.4-mini67.7%44/6521/36100%5/565%39/6013.617.3s28.3K$0.0104.6%3/650.0%0/65
23minimax-m321.5%14/6514/6560%3/518%11/60185.031.4m5.3M$0.3653.8%35/650.0%0/65
24claude-fable-510.8%7/650/4100%5/53%2/603.41.1m16.0K$0.0310.0%0/6589.2%58/65

Slices

Per-Model Results

Drill into any model to see per-task pass rates across its attempted tasks. This build contains aggregate results only (no per-trial pages).