τ-Rec extends τ-bench to conversational recommendation. It evaluates whether an agent can uncover a user's preferences over multiple turns, use catalog tools, follow policy, and make a recommendation that is checked programmatically—not graded by another language model. Across four repeated trials, the strongest evaluated agent falls from roughly 57% pass^1 to 35% pass^4.
Recommender agents must do more than retrieve a plausible item. They have to ask the right questions, recognize what the user has not said, inspect a changing catalog, respect operational policy, and commit to one recommendation. A polished conversation can still end with the wrong item.
τ-Rec brings this setting to the τ-bench family through 60 multi-turn movie-recommendation tasks. Each task pairs a simulated user persona with structured constraints, a 153-item catalog, search and metadata tools, and deterministic checks for both recommendation quality and policy compliance. The benchmark and its evaluation artifacts were published in the ACM RecSys 2026 Reproducibility and Resource Track.
Three ideas τ-Rec brings together
τ-Rec combines three pieces needed to evaluate conversational recommender agents as complete systems—not just as rankers.
Verifiable by construction
Structured catalog predicates and programmatic policy checks replace subjective LLM-as-a-judge scoring.
Reliability across trials
Repeated runs separate one-shot capability from consistent success. To our knowledge, pass^k is novel to recommendation evaluation.
Reveal-tagged elicitation
Constraints tagged volunteer, on_ask, or hidden force agents to uncover needs through dialogue.
How a τ-Rec task works
A task starts with a persona and a checklist of requirements. Some constraints are volunteered immediately, some appear only when the agent asks, and some remain hidden. The agent must use conversation and tools to narrow the catalog before making one final recommendation.
Simulated user
A persona responds naturally while revealing constraints at controlled points.
Conversation and tools
Deterministic scorer
The selected item is checked against the catalog and every applicable policy rule.
- constraints satisfied
- available on service
- policy checks pass
Condensed from the public trace archive; repeated catalog lookups are omitted.
Each task is run four times. τ-Rec reports pass^1, pass^2, and pass^4, so an agent receives full reliability credit only when it succeeds repeatedly rather than getting one favorable trajectory.
Paper results
This snapshot shows all nine configurations evaluated in the paper. Switch metrics to see how rankings change as the reliability requirement becomes stricter.
| # | Model | Mode | pass^1 |
|---|---|---|---|
| 1 | DeepSeek V4 Flash | Max thinking | 57.1% |
| 2 | DeepSeek V4 Flash | High thinking | 56.0% |
| 3 | GPT-5.4 | Medium thinking | 55.1% |
| 4 | DeepSeek V4 Flash | No thinking | 54.6% |
| 5 | Sonnet 4.6 | Default | 53.7% |
| 6 | GPT-5.4 | No thinking | 47.1% |
| 7 | GPT-5 mini | Default | 41.7% |
| 8 | Gemini 2.5 Flash | Default | 27.5% |
| 9 | Qwen3-32B | Default | 27.1% |
Key findings
The central result is a reliability gap. Models that look capable on one attempt become much less dependable when the same task must succeed across repeated trials.
The strongest evaluated agent completes about 57% of tasks on one try, but only about 35% on all four tries. The gap shows why one-shot accuracy can overstate whether an agent is ready to behave reliably in a user-facing setting.
Preference elicitation is another major failure mode. For the highlighted configuration, one-try success falls from 85% when every need is volunteered to 20% when one constraint is never stated. An agent that recommends too quickly can sound helpful while missing the requirement that matters most.
The best agents verify before they commit. Strong configurations use catalog and availability tools to check their choice. Common failures include declining to recommend anything or recommending an item without verifying that it is actually available.
Paper, code, and current results
Read the RecSys paper for the full methodology, run the benchmark from GitHub, or see the latest verified submissions.