Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
Pythonsierra-research/tau2-bench

tau2-bench

τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

81.5/100
2.0KForks: 502
View on GitHubHomepage →
Loading report...

Similar Projects

AI_Diplomacy

49

Frontier Models playing the board game Diplomacy.

Python699

meta-agents-research-environments

68

Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic, realistic scenarios. Unlike static benchmarks, this platform introduces evolving environments where agents must adapt their strategies as new information becomes available, mirroring real-world challenges.

Python551

hermes-agent

90

The agent that grows with you

Python243.1K

AutoGPT

96

AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.

Python187.2K
Back to List