Live Benchmarks

Mastra Benchmark

Performance results of AI coding models on Mastra.ai tasks, measuring success rate and execution time with high precision.

View on GitHubTotal tasks: 13Last run: 4/24/2026

Model Performance

ModelPassedAvg DurationSuccess Rate
#1
claude-4-6-sonnetNEW
13235.3s
100%
#2
gpt-5.2-codex
9248.5s
69%
#3
glm-4.7
9444.5s
69%
#4
gemini-3-flash
8354.5s
62%