General AI Assistants
GAIA evaluates AI models on real-world tasks that are conceptually simple for humans but require multi-step reasoning, web browsing, tool use, and multimodal understanding for AI. Tasks span three difficulty levels and test practical assistant capabilities rather than academic knowledge.
ModelsModels with a result.1
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.2024
Results
1 result#ModelResultThe score from the source.