Shishir G. Patil et al.
Berkeley Function Calling Leaderboard v3
A function-calling benchmark for tool selection, schema adherence, and argument correctness, covering single-turn, parallel, irrelevance and multi-turn subsets.
ModelsModels with a result.1
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.2025
Results
1 result#ModelResultThe score from the source.