Community Spotlight

τ-Rec: A Verifiable Benchmark for Agentic Recommender Systems

September 2026 · ACM RecSys 2026

TL;DR

τ-Rec extends τ-bench to conversational recommendation. It evaluates whether an agent can uncover a user's preferences over multiple turns, use catalog tools, follow policy, and make a recommendation that is checked programmatically—not graded by another language model. Across four repeated trials, the strongest evaluated agent falls from roughly 57% pass^1 to 35% pass^4.

Recommender agents must do more than retrieve a plausible item. They have to ask the right questions, recognize what the user has not said, inspect a changing catalog, respect operational policy, and commit to one recommendation. A polished conversation can still end with the wrong item.

τ-Rec brings this setting to the τ-bench family through 60 multi-turn movie-recommendation tasks. Each task pairs a simulated user persona with structured constraints, a 153-item catalog, search and metadata tools, and deterministic checks for both recommendation quality and policy compliance. The benchmark and its evaluation artifacts were published in the ACM RecSys 2026 Reproducibility and Resource Track.

Three ideas τ-Rec brings together

τ-Rec combines three pieces needed to evaluate conversational recommender agents as complete systems—not just as rankers.

Verifiable

Verifiable by construction

Structured catalog predicates and programmatic policy checks replace subjective LLM-as-a-judge scoring.

Reliable

Reliability across trials

Repeated runs separate one-shot capability from consistent success. To our knowledge, pass^k is novel to recommendation evaluation.

Elicitation

Reveal-tagged elicitation

Constraints tagged volunteer, on_ask, or hidden force agents to uncover needs through dialogue.

How a τ-Rec task works

A task starts with a persona and a checklist of requirements. Some constraints are volunteered immediately, some appear only when the agent asks, and some remain hidden. The agent must use conversation and tools to narrow the catalog before making one final recommendation.

1

Simulated user

A persona responds naturally while revealing constraints at controlled points.

genre: Action
2

Conversation and tools

“I want an action movie for Friday night.”
search_catalog(...) → candidates
“Do you have a runtime or streaming preference?”
check_availability(...) → true
3

Deterministic scorer

The selected item is checked against the catalog and every applicable policy rule.

  • constraints satisfied
  • available on service
  • policy checks pass
Sample conversation trace
Task 040 · DeepSeek V4 Flash (max thinking) · passing trial

Condensed from the public trace archive; repeated catalog lookups are omitted.

User Volunteer
“Genres: Horror (this is a must). Minimum rating: 5.5 (no lower). I’m a cinephile who follows indie film festivals and prefers auteur-driven, original films with a cult following.”
Agent
search_catalog("horror")
get_metadata(candidate_id) …
“Before I commit, could you tell me which streaming services you have? That’ll help me narrow things down to what’s actually available to you.”
User On ask
“I have Hulu.”
Agent
check_availability(candidate_id, ["Hulu"])
get_metadata("tmdb_1285965")
check_content_preference("R")
“Dangerous Animals really stands out for someone with your taste. Before I recommend it though—it’s rated R for mature content. Are you comfortable with that?”
User On ask
“Yes—R is fine, and that sounds like a perfect pick for my taste.”
Agent
recommend(item_id="tmdb_1285965")
“Recommended: Dangerous Animals (2025). You’re all set—enjoy the horror!”
✓ Deterministic result: pass 5/5 constraints satisfied · no policy violations

Each task is run four times. τ-Rec reports pass^1, pass^2, and pass^4, so an agent receives full reliability credit only when it succeeds repeatedly rather than getting one favorable trajectory.

Paper results

This snapshot shows all nine configurations evaluated in the paper. Switch metrics to see how rankings change as the reliability requirement becomes stricter.

Benchmark snapshot
60 tasks × 4 trials · GPT-5 mini user simulator
Metric:
τ-Rec paper results, ranked by the selected pass metric
# Model Mode pass^1
1 DeepSeek V4 Flash Max thinking 57.1%
2 DeepSeek V4 Flash High thinking 56.0%
3 GPT-5.4 Medium thinking 55.1%
4 DeepSeek V4 Flash No thinking 54.6%
5 Sonnet 4.6 Default 53.7%
6 GPT-5.4 No thinking 47.1%
7 GPT-5 mini Default 41.7%
8 Gemini 2.5 Flash Default 27.5%
9 Qwen3-32B Default 27.1%

Key findings

The central result is a reliability gap. Models that look capable on one attempt become much less dependable when the same task must succeed across repeated trials.

Paper result: reliability falls across repeated trials, from 57 percent pass one to 35 percent pass four for the strongest evaluated agent
Paper results across nine evaluated configurations. The highlighted configuration drops by 22 percentage points when success is required on all four trials.

The strongest evaluated agent completes about 57% of tasks on one try, but only about 35% on all four tries. The gap shows why one-shot accuracy can overstate whether an agent is ready to behave reliably in a user-facing setting.

Paper result: hidden preferences make tasks harder, with one-try success falling from 85 percent for volunteered constraints to 20 percent when a constraint remains hidden
One-try success for the highlighted configuration by reveal difficulty; whiskers show the range across all nine configurations.

Preference elicitation is another major failure mode. For the highlighted configuration, one-try success falls from 85% when every need is volunteered to 20% when one constraint is never stated. An agent that recommends too quickly can sound helpful while missing the requirement that matters most.

The best agents verify before they commit. Strong configurations use catalog and availability tools to check their choice. Common failures include declining to recommend anything or recommending an item without verifying that it is actually available.

Explore τ-Rec

Paper, code, and current results

Read the RecSys paper for the full methodology, run the benchmark from GitHub, or see the latest verified submissions.