As AI writes more of the code now, we measure how well it understands yours - benchmarking real-world usage of your library across models, tasks, and evolving releases.
| Model | Passed | Avg Duration | Success Rate |
|---|---|---|---|
| #1 claude-4-6-sonnetNEW | 7 | 222.5s | 70% |
| #2 gemini-3.1-pro | 6 | 202.0s | 60% |
| #3 glm-4.7 | 4 | 336.6s | 40% |
| #4 gpt-5.2-codex | 4 | 148.9s | 40% |