Postman

APIFlow-Bench 1.0 Leaderboard

20-subtask chainHard

Deterministic API engineering benchmark spanning 7 capability axes, over REST APIs only (GraphQL and MCP are planned). Showing the 20-subtask chain task set: 1,320 trials across 24 models and 11 tasks covering 5 of the 7 axes — authentication, error recovery, multi-step workflow, pagination, statefulness (all 7 under All tasks). Switch task set above; click any card to drill into a slice with the full interactive detail view; click any model name to see its per-task results.

Overall pass rate57.5%Trials1,320Models24Tasks11

Overall Ranking

Columns are grouped into metric families. Pass rate = passed ÷ all trials; API recovery = passed ÷ trials that hit a 4xx/5xx. Difficulty/horizon breakdowns are omitted: every task in this slice is judged Hard + Long-horizon and Long-horizon. Per-trial details (verifier groups, checkpoints, step counts) live on the model → task → trial pages.

Rank
Model
Scoring
Cost / efficiency
Behavioral
Pass rate
API recovery
Tool calls / trial
Time / trial
Tokens / trial
Cost / trial
Abstain %
Refusal %
1kimi-k372.7%40/5521/2620.712.8m171.4K$0.200.0%0/550.0%0/55
3claude-opus-4-869.1%38/5518/2313.444.3s69.7K$0.0860.0%0/550.0%0/55
3deepseek-v4-flash69.1%38/5521/2921.044.7s80.6K$0.0080.0%0/550.0%0/55
5glm-5p265.5%36/5517/2317.33.2m92.8K$0.0790.0%0/550.0%0/55
5qwen3p7-plus65.5%36/5516/2412.05.3m125.3K$0.0390.0%0/550.0%0/55
7claude-haiku-4-5-2025100163.6%35/5515/2014.926.6s66.2K$0.0220.0%0/550.0%0/55
7gpt-5.6-luna63.6%35/5515/2314.01.6m31.8K$0.0140.0%0/550.0%0/55
7gpt-5.6-terra63.6%35/5516/2214.51.3m41.3K$0.0430.0%0/550.0%0/55
7qwen3p6-plus63.6%35/5518/2814.51.1m72.5K$0.0350.0%0/550.0%0/55
7deepseek-v4-pro63.6%35/5522/2923.13.5m82.8K$0.110.0%0/550.0%0/55
7gpt-5.6-sol63.6%35/5516/2117.23.4m67.2K$0.150.0%0/550.0%0/55
13claude-sonnet-4-661.8%34/5519/2514.150.4s70.3K$0.0480.0%0/550.0%0/55
13kimi-k2p7-code61.8%34/5515/2113.92.3m50.4K$0.0220.0%0/550.0%0/55
15claude-sonnet-560.0%33/5516/2519.31.2m140.1K$0.0700.0%0/550.0%0/55
16minimax-m2p758.2%32/5518/2513.619.5s43.7K$0.0100.0%0/550.0%0/55
16kimi-k2p658.2%32/5516/2121.81.9m255.9K$0.0870.0%0/550.0%0/55
18kimi-k2p556.4%31/5515/2011.622.8s34.6K$0.0170.0%0/550.0%0/55
18gpt-oss-20b56.4%31/5518/2516.72.1m80.3K$0.0040.0%0/550.0%0/55
20gpt-oss-120b52.7%29/5515/2414.825.8s68.8K$0.0070.0%0/550.0%0/55
21gpt-5-mini50.9%28/5516/2412.731.7s54.4K$0.0067.3%4/550.0%0/55
22gpt-5.4-mini43.6%24/5514/3213.816.0s26.1K$0.0105.5%3/550.0%0/55
23claude-fable-516.4%9/554.01.2m18.8K$0.0420.0%0/5576.4%42/55
24minimax-m37.3%4/554/55210.336.2m6.3M$0.4356.4%31/550.0%0/55

Slices

Per-Model Results

Drill into any model to see per-task pass rates across its attempted tasks. This build contains aggregate results only (no per-trial pages).