Performance results of AI coding models on TanStack tasks, measuring success rate and execution time with high precision.
| Model | Passed | Avg Duration | Success Rate |
|---|---|---|---|
| #1 gemini-3.5-flashNEW | 11 | 564.8s | 79% |
| #2 gemini-3.1-pro | 8 | 407.8s | 57% |
| #3 claude-4-6-sonnet | 3 | 503.5s | 21% |
| #4 glm-5.1 | 2 | 599.0s | 14% |
| #5 deepseek-4-pro | 1 | 756.0s | 7% |