Postman

APIFlow-Bench 1.0 Leaderboard

10-subtask chainHard

Deterministic API engineering benchmark spanning 7 capability axes, over REST APIs only (GraphQL and MCP are planned). Showing the 10-subtask chain task set: 1,440 trials across 24 models and 12 tasks covering 5 of the 7 axes — authentication, discovery, error recovery, schema, statefulness (all 7 under All tasks). Switch task set above; click any card to drill into a slice with the full interactive detail view; click any model name to see its per-task results.

Overall pass rate66.2%Trials1,440Models24Tasks12

Overall Ranking

Columns are grouped into metric families. Pass rate = passed ÷ all trials; API recovery = passed ÷ trials that hit a 4xx/5xx. Difficulty/horizon breakdowns are omitted: every task in this slice is judged Hard + Long-horizon and Long-horizon. Per-trial details (verifier groups, checkpoints, step counts) live on the model → task → trial pages.

Rank
Model
Scoring
Cost / efficiency
Behavioral
Pass rate
API recovery
Tool calls / trial
Time / trial
Tokens / trial
Cost / trial
Abstain %
Refusal %
1claude-haiku-4-5-2025100175.0%45/6020/2215.129.7s67.2K$0.0220.0%0/600.0%0/60
1claude-sonnet-4-675.0%45/6024/2410.033.0s37.3K$0.0270.0%0/600.0%0/60
1gpt-5.575.0%45/6021/2112.328.5s28.8K$0.0730.0%0/600.0%0/60
1qwen3p6-plus75.0%45/6025/2912.938.8s52.3K$0.0220.0%0/600.0%0/60
1deepseek-v4-flash75.0%45/6024/2516.941.4s56.5K$0.0060.0%0/600.0%0/60
1kimi-k2p675.0%45/6022/2513.345.2s55.8K$0.0280.0%0/600.0%0/60
1gpt-5.6-sol75.0%45/6023/2313.02.1m29.8K$0.0700.0%0/600.0%0/60
1glm-5p275.0%45/6023/2613.02.0m44.5K$0.0390.0%0/600.0%0/60
1deepseek-v4-pro75.0%45/6029/3818.12.8m66.7K$0.0870.0%0/600.0%0/60
11claude-sonnet-573.3%44/6025/2814.335.6s70.7K$0.0320.0%0/600.0%0/60
11gpt-oss-20b73.3%44/6022/2514.61.4m67.3K$0.0030.0%0/600.0%0/60
11kimi-k2p7-code73.3%44/6024/2611.81.9m32.1K$0.0150.0%0/600.0%0/60
11qwen3p7-plus73.3%44/6023/249.93.0m66.3K$0.0180.0%0/600.0%0/60
15gpt-5.6-terra71.7%43/6025/2513.51.2m35.3K$0.0350.0%0/600.0%0/60
15kimi-k371.7%43/6024/3014.36.9m55.6K$0.0941.7%1/600.0%0/60
17kimi-k2p568.3%41/6024/2510.528.8s31.9K$0.0150.0%0/600.0%0/60
17claude-opus-4-868.3%41/6025/2610.537.5s53.3K$0.0630.0%0/600.0%0/60
17gpt-5.6-luna68.3%41/6023/2513.51.5m29.4K$0.0120.0%0/600.0%0/60
20gpt-oss-120b61.7%37/6023/2813.124.8s56.1K$0.0060.0%0/600.0%0/60
21gpt-5-mini58.3%35/6016/1910.633.0s44.2K$0.00511.7%7/600.0%0/60
22gpt-5.4-mini45.0%27/6015/2911.014.7s22.0K$0.0086.7%4/600.0%0/60
23minimax-m325.0%15/6015/60122.220.5m2.6M$0.1935.0%21/600.0%0/60
24claude-fable-58.3%5/600/13.053.9s13.6K$0.0280.0%0/6091.7%55/60

Slices

Per-Model Results

Drill into any model to see per-task pass rates across its attempted tasks. This build contains aggregate results only (no per-trial pages).