Live Benchmarks

Gel Benchmark

Performance results of AI coding models on gel tasks, measuring success rate and execution time with high precision.

View on GitHubTotal tasks: 13Last run: 8/7/2026

Model Performance

ModelPassedAvg DurationSuccess Rate
#1
claude-5-sonnetNEW
11570.6s
85%
#2
gemini-3.5-flash
8686.0s
62%
#3
deepseek-4-pro
4996.0s
31%
#4
glm-5.2
21048.2s
15%