Live Benchmarks

E2B Benchmark

As AI writes more of the code now, we measure how well it understands yours - benchmarking real-world usage of your library across models, tasks, and evolving releases.

View on GitHubTotal tasks: 10Last run: 3/31/2026

Model Performance

ModelPassedAvg DurationSuccess Rate
#1
glm-4.7-with-skillsNEW
8260.3s
80%
#2
claude-4-6-sonnet
7222.5s
70%
#3
gemini-3.1-pro
6202.0s
60%
#4
gpt-5.2-codex
4148.9s
40%
#5
glm-4.7
4336.6s
40%