Berkeley Function Calling Leaderboard v4

A function-calling benchmark for tool selection, schema adherence, and argument correctness.

  • In index
  • Agentic
ModelsModels with a result.22
Top resultBest result on this test.88.5%BTL-3
Top-3 spreadPoints from first to third.13.5 pts
YearYear of release.N/A

Result and price

020406080100$0.03$0.1$0.3$1$3$10RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

22 results
#ModelResultThe score from the source.
01BTL-388.5%
04BTL-473.5%
18ZAYA1-8B39.2%