MultiVerse: A Trustworthy Tool-Use Benchmark for Real-World Task Delegation under Harnesses

Abstract

As agents improve, users delegate increasingly complex tasks involving sequences of tool calls across multiple services. These user requests are often incomplete or constrained, requiring agents to ask for clarification or follow instructions that change how they ordinarily use tools and communicate. Moreover, users often make requests through end-user harnesses, such as Hermes and OpenClaw, which change how models act. Therefore, a tool-use benchmark for such tasks must use realistic service interfaces, support multi-turn user interaction, assess instruction-based behavioral steering and compare models across harnesses. Trustworthy evaluation also requires distinguishing agent failures from errors introduced by the infrastructure, harness integrations, user simulators, or verifiers. We present MULTIVERSE, a benchmark for cross-service delegation that simulates 35 real-world services with 809 tools in a self-contained sandbox. Its tasks include behavioral instructions and multi-turn settings that test whether agents seek missing information from a reticent user simulator. The benchmark compares twelve models on the same tasks under one bare harness (a simple loop) and four end-user harnesses. The strongest model–harness pair achieves a mean success rate of 69.8%. Instruction-following weaknesses differ across constraints: call budgets and final-response requirements have the highest average violation rates, while the most compliant model differs by instruction category. Agents struggle to obtain user-held information: in the multi-turn setting, 29.5% of attempts terminate without the agent consulting the user for information needed to solve the task. Among the models and settings evaluated, no harness is consistently best: a harness that helps one model can hurt another. Averaged across the evaluated models and settings, end-user harnesses perform nearly identically to the bare harness, despite large gains and losses for particular model–harness–setting combinations. We plan to release the benchmark and evaluation code and provide a leaderboard for the hidden test set.

Publication
To appear in arXiv, 2026

Related