MLS-Bench Lite

A 30-task subset of MLS-Bench that evaluates whether AI systems can invent generalizable and scalable machine-learning methods.

ModelsModels with a result.4
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.2026

Results

4 results
#ModelResultThe score from the source.
Kimi K348.3%
Qwen3.8 Max41.0%