Postman

APIFlow-Bench 1.0 Leaderboard

15-subtask chainHard

Deterministic API engineering benchmark spanning 7 capability axes, over REST APIs only (GraphQL and MCP are planned). Showing the 15-subtask chain task set: 1,320 trials across 24 models and 11 tasks covering 6 of the 7 axes — authentication, discovery, error recovery, multi-step workflow, schema, statefulness (all 7 under All tasks). Switch task set above; click any card to drill into a slice with the full interactive detail view; click any model name to see its per-task results.

Overall pass rate62.7%Trials1,320Models24Tasks11

Overall Ranking

Columns are grouped into metric families. Pass rate = passed ÷ all trials; API recovery = passed ÷ trials that hit a 4xx/5xx. Difficulty/horizon breakdowns are omitted: every task in this slice is judged Hard + Long-horizon and Long-horizon. Per-trial details (verifier groups, checkpoints, step counts) live on the model → task → trial pages.

Rank
Model
Scoring
Cost / efficiency
Behavioral
Pass rate
API recovery
Tool calls / trial
Time / trial
Tokens / trial
Cost / trial
Abstain %
Refusal %
2gpt-5.572.7%40/5516/2111.924.8s26.4K$0.0690.0%0/550.0%0/55
2claude-sonnet-572.7%40/5515/2014.535.3s68.4K$0.0310.0%0/550.0%0/55
2claude-opus-4-872.7%40/5516/2110.635.9s51.2K$0.0620.0%0/550.0%0/55
2gpt-5.6-sol72.7%40/5515/2012.42.0m27.2K$0.0630.0%0/550.0%0/55
2kimi-k2p7-code72.7%40/5515/2010.81.9m28.2K$0.0130.0%0/550.0%0/55
2kimi-k372.7%40/5516/2112.16.6m44.5K$0.0750.0%0/550.0%0/55
8deepseek-v4-flash70.9%39/5516/2217.938.8s52.3K$0.0050.0%0/550.0%0/55
8kimi-k2p670.9%39/5514/1912.334.9s41.5K$0.0200.0%0/550.0%0/55
10qwen3p6-plus69.1%38/5520/2511.229.3s42.3K$0.0180.0%0/550.0%0/55
10claude-sonnet-4-669.1%38/5513/1810.332.2s37.2K$0.0270.0%0/550.0%0/55
10gpt-oss-120b69.1%38/5513/1813.741.6s60.2K$0.0060.0%0/550.0%0/55
10gpt-5.6-luna69.1%38/5514/1711.71.4m25.8K$0.0100.0%0/550.0%0/55
14qwen3p7-plus67.3%37/5515/209.23.5m77.5K$0.0310.0%0/550.0%0/55
15gpt-5.6-terra65.5%36/5511/1612.41.2m30.6K$0.0320.0%0/550.0%0/55
15deepseek-v4-pro65.5%36/5519/3119.92.9m65.1K$0.0850.0%0/550.0%0/55
17minimax-m2p761.8%34/5512/2012.021.7s39.7K$0.0090.0%0/550.0%0/55
17kimi-k2p561.8%34/5510/1610.527.0s30.0K$0.0140.0%0/550.0%0/55
17gpt-oss-20b61.8%34/5515/2314.81.4m65.2K$0.0033.6%2/550.0%0/55
20claude-haiku-4-5-2025100158.2%32/5511/1713.625.5s57.8K$0.0211.8%1/550.0%0/55
21gpt-5-mini56.4%31/5510/1811.530.5s46.4K$0.00510.9%6/550.0%0/55
22gpt-5.4-mini47.3%26/558/1611.814.7s20.8K$0.0083.6%2/550.0%0/55
23minimax-m316.4%9/559/55190.532.7m5.5M$0.3852.7%29/550.0%0/55
24claude-fable-512.7%7/550/42.751.9s13.3K$0.0270.0%0/5587.3%48/55

Slices

Per-Model Results

Drill into any model to see per-task pass rates across its attempted tasks. This build contains aggregate results only (no per-trial pages).