Can AI agents reliably complete real-world tasks? τ-bench measures how well agents converse with users, call tools, retrieve knowledge, and follow policy across enterprise domains — in text and voice.
| # | Model | Pass^1 |
|---|---|---|
| 🥇 | Qwen 3.8 MaxQwen | 55.2% |
| 🥈 | Claude Opus 5Anthropic | 48.7% |
| 🥉 | Grok 4.5xAI | 47.9% |
| # | Model | Pass^1 |
|---|---|---|
| 🥇 | Pine Voice PreviewPine AI | 75.4% |
| 🥈 | grok-voice-think-fast-1.0xAI | 67.3% |
| 🥉 | grok-voice-think-fast-2.0xai | 62.5% |
| # | Model | Pass^1 |
|---|---|---|
| 🥇 | Qwen3.5-397B-A17BAlibaba Cloud | 87.9% |
| 🥈 | Gemini 3.0 ProGoogle | 85.4% |
| 🥉 | Claude Opus 4.5Anthropic | 85.3% |
Audited and fixed 50+ tasks across the airline and retail domains — correcting expected actions, ambiguous instructions, and impossible constraints.