NEWτ³-bench is here: τ-knowledge evaluates agents on knowledge-intensive tasks, and τ-voice benchmarks real-time voice agents.

τ³-Banking Leaderboard

Text agents resolving banking customer-service tasks over a ~700-document knowledge base. Published as τ-knowledge.

RankModelReleasedRetrievalReasoningUser Sim
#1
GPT-5.5
Apr 22, 2026🔍 alltoolsxhighgpt-5.2
46.4%
#2
GPT-5.4
Mar 5, 2026🔍 alltoolsxhighgpt-5.2
39.4%
#3
GPT-5.2
Dec 11, 2025🔍 alltoolshighgpt-5.2
32.2%
#4
Claude Opus 4.7
Apr 16, 2026🔍 alltoolsmaxgpt-5.2
30.1%
#5
GLM-5.2
Jun 16, 2026🔍 alltoolsmaxGLM-5.2
29.6%
#6
Gemini 3 Flash
Dec 17, 2025🔍 Terminalhighgpt-5.2
27.3%
#7
Claude Opus 4.6
Feb 5, 2026🔍 alltoolsmaxgpt-5.2
27.3%
#8
Gemini 3.1 Pro Preview
Feb 19, 2026🔍 alltoolshighgpt-5.2
26.0%
#9
Claude Sonnet 4.5
Sep 29, 2025🔍 Terminalenabledgpt-5.2
25.3%
#10
Claude Opus 4.5
Nov 24, 2025🔍 alltoolshighgpt-5.2
24.7%
#11
Gemini 3 Pro
Nov 18, 2025🔍 Terminalhighgpt-5.2
18.0%
#12
Grok 4.2
Mar 10, 2026🔍 alltoolshighgpt-5.2
18.0%
#13
Grok 4 fast
Sep 19, 2025🔍 alltoolshighgpt-5.2
15.7%
#14
Gemini 2.5 Pro
Jun 17, 2025🔍 alltoolshighgpt-5.2
13.7%
#15
Grok 4.1 fast
Nov 19, 2025🔍 alltoolshighgpt-5.2
13.1%
#16
GPT-5.2
Dec 11, 2025🔍 qwen_embeddingsnonegpt-5.2
12.6%
#17
GLM-5
Feb 11, 2026🔍 text-emb-3-largeenabledgpt-5.2
9.8%
#18
Qwen3.5-397B-A17B
Feb 16, 2026🔍 text-emb-3-largeenabledgpt-5.2
9.8%
τ³-Banking Banking pass@1 by release date
Pass@1 on τ³-Banking · banking · plotted at each model's public release date. 17 models.
0%10%20%30%40%50%60%70%80%90%100%BANKING PASS@1Jun 2025Aug 2025Oct 2025Dec 2025Feb 2026Apr 2026Jun 2026Gemini 2.5 Pro (Google) · 2025-06-17 · 13.7% bankingGrok 4 fast (xAI) · 2025-09-19 · 15.7% bankingClaude Sonnet 4.5 (Anthropic) · 2025-09-29 · 25.3% bankingGemini 3 Pro (Google) · 2025-11-18 · 18.0% bankingGrok 4.1 fast (xAI) · 2025-11-19 · 13.1% bankingClaude Opus 4.5 (Anthropic) · 2025-11-24 · 24.7% bankingGPT-5.2 (OpenAI) · 2025-12-11 · 32.2% bankingGemini 3 Flash (Google) · 2025-12-17 · 27.3% bankingClaude Opus 4.6 (Anthropic) · 2026-02-05 · 27.3% banking🧠GLM-5 (Zhipu AI) · 2026-02-11 · 9.8% bankingQwen3.5-397B-A17B (Alibaba Cloud) · 2026-02-16 · 9.8% bankingGemini 3.1 Pro Preview (Google) · 2026-02-19 · 26.0% bankingGPT-5.4 (OpenAI) · 2026-03-05 · 39.4% bankingGrok 4.2 (xAI) · 2026-03-10 · 18.0% bankingClaude Opus 4.7 (Anthropic) · 2026-04-16 · 30.1% bankingGPT-5.5 (OpenAI) · 2026-04-22 · 46.4% bankingZGLM-5.2 (Z.ai) · 2026-06-16 · 29.6% bankingGemini 2.5 Pro13.7%Grok 4 fast15.7%Claude Sonnet 4.525.3%GPT-5.232.2%GPT-5.439.4%GPT-5.546.4%
GooglexAIAnthropicOpenAIZhipu AIAlibaba CloudZ.aiFrontier (best so far)
Markers show each model's banking pass@1 plotted at its public release date. Hover any marker for details. The dashed line tracks the running best. Submissions without a published model_release.release_date fall back to the evaluation date.

Submit Your Results

Have new results to share? Submit your model evaluation results by creating a pull request to add your JSON submission file. See our submission guidelines for the required format and process.