BIG-Bench Hard

A suite of 23 challenging tasks from the BIG-Bench collaborative benchmark where prior language models failed to exceed average human performance, even with chain-of-thought prompting.

ModelsModels with a result.3
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.2022

Results

3 results
#ModelResultThe score from the source.
Gemma 4 12B53.0%