Finance Agent v1.1

Evaluating agents on core financial analyst tasks

ModelsModels with a result.50
Top resultBest result on this test.64.4%Claude Opus 4.7
Top-3 spreadPoints from first to third.3.8 pts
UpdatedDate the source changed the results.4 Jun 2026

Result and cost

020406080100$0.01$0.03$0.1$0.3$1$3$10RESULTUS DOLLARS PER TASK · LOG SCALE

Results

50 results
#ModelResultThe score from the source.CostUS dollars to run one task.TimeTime to run one task.
01Claude Opus 4.764.4%$0.8025 min
02Claude Sonnet 4.663.3%$1.446 min
03Muse Spark60.6%$0.0637 min
04DeepSeek V4 Pro 042360.4%$0.58510 min
06GPT-5.560.0%$1.3314 min
07Gemini 3.1 Pro Preview59.7%$0.8734 min
08Claude Opus 4.5 Thinking58.8%$1.53 min
09GPT-5.258.5%$0.97810 min
10GLM-5.157.7%$0.2928 min
11GPT-5.457.2%$1.4111 min
12GPT-5.155.3%$0.47410 min
13Gemini 3 Pro55.2%$0.5593 min
14Qwen3.6 Plus54.6%$0.1525 min
16Qwen3.5 Plus Thinking54.5%$0.2366 min
17Grok 4.353.8%$0.44213 min
18Grok 453.5%$1.075 min
19GPT-5.4 mini53.4%$0.49215 min
20GLM-5 Thinking53.2%$0.5019 min
21Qwen 3.6 Max (preview)52.8%$0.58831 min
22Grok 4.1 Fast (Reasoning)52.4%$0.0582 min
23Grok 4.20 (Reasoning)52.3%$0.3622 min
24GPT-552.2%$0.58615 min
25GPT-5 mini51.9%$0.14211 min
26Gemma 4 31B50.8%$089 min
27Kimi K2.5 Thinking50.6%$0.1794 min
28MiniMax M2.748.4%$0.1622 min
29GPT-5.4 nano47.8%$0.0135 min
30Gemini 3 Flash Preview47.6%$0.3715 min
31Claude Haiku 4.5 Thinking46.9%$0.3772 min
33Mistral Medium 3.546.1%$1.159 min
34Grok 4 Fast (Reasoning)46.1%$0.04759 s
35GLM-4.746.0%$0.24410 min
36Qwen3.5 Flash45.6%$0.0554 min
37Grok 4.1 Fast Non Reasoning44.4%$0.0892 min
38Qwen3 Max44.3%$0.8025 min
39Gemini 2.5 Pro41.6%$0.34131 min
40MiniMax M2.538.6%$0.16523 min
41Kimi K2 Thinking36.6%$0.2511 min
42GLM-4.636.5%$0.18712 min
43MiniMax M2.133.4%$0.11422 min
44GPT-OSS 120B21.5%$0.0653 min
45Mistral Large18.0%$0.0862 min
46GPT-4o8.1%$0.15831 s
47Command A4.2%$2.3410 min
48DeepSeek V3.2 (Thinking)2.3%$0.20929 min
49Jamba Large 1.70.4%$5.0439 min
50DeepSeek V3.20.0%$0.042 min