Live Benchmarks

TanStack Benchmark

Performance results of AI coding models on TanStack tasks, measuring success rate and execution time with high precision.

View on GitHubTotal tasks: 14Last run: 7/23/2026

Model Performance

ModelPassedAvg DurationSuccess Rate
#1
gemini-3.5-flashNEW
11564.8s
79%
#2
gemini-3.1-pro
8407.8s
57%
#3
claude-4-6-sonnet
3503.5s
21%
#4
glm-5.1
2599.0s
14%
#5
deepseek-4-pro
1756.0s
7%