Live Benchmarks

E2B Benchmark

As AI writes more of the code now, we measure how well it understands yours - benchmarking real-world usage of your library across models, tasks, and evolving releases.

View on GitHubTotal tasks: 10Last run: 6/16/2026

Model Performance

ModelPassedAvg DurationSuccess Rate
#1
claude-4-6-sonnetNEW
7222.5s
70%
#2
gemini-3.1-pro
6202.0s
60%
#3
glm-4.7
4336.6s
40%
#4
gpt-5.2-codex
4148.9s
40%