LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

A continuously updated coding benchmark built from newly collected LeetCode, AtCoder, and Codeforces problems. Fresh problem windows reduce one contamination path, but results still need a release and setup check.

ModelsModels with a result.7
Top resultBest result on this test.91.6%Qwen3.7 Max
Top-3 spreadPoints from first to third.3.8 pts
YearYear of release.2024

Result and price

020406080100$0.3$1$3$10RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

7 results
#ModelResultThe score from the source.
04GLM-4.784.9%