Menus Limit Available Actions
A tool menu is a short, ordered subset of available actions that an agent receives before carrying out a task. The agent can call only the tools in this menu.
This means the menu’s contents and order affect whether the agent can complete the task. A practical criterion is to check not only whether the required action is included, but also whether the tools that prepare data for it are available.
Why Relevance Does Not Guarantee Completion
A multi-step task requires a final action and predecessor tools. These tools must create the necessary inputs in the right order.
Ranking by relevance to the query may find the final action but miss or defer less obvious tools that produce the required data. A selection criterion is to evaluate the menu based on the completeness of the executable chain, not just how closely the tools match the query.
| Approach | What it considers | Risk |
|---|---|---|
| Relevance ranking | How closely a tool relates to the query | Tools that produce the required data may be missing or listed too late |
| State-path selection | Whether transitions can be executed and the order of dependencies | The menu must account for state and the relationships between inputs and outputs |
How Menus Are Formed Using State Paths
A state path is a route from the observable state of a query to the desired result. The State-Path Tool Menu proposed by the authors assembles a menu based on whether this route can be executed.
The mechanism involves several steps:
- The encoder considers the tools available from the current state.
- It links tool outputs to the inputs of subsequent actions.
- It takes into account action sequences that recur in training paths.
- Search selects an executable starting tool, tools that produce missing data, and the final action.
- Reranking places producers before consumers.

A practical criterion is to check that the selected menu includes the starting action, the necessary producer tools, and the final action in an executable order.
What the ToolBench Results Show
The authors report that on ToolBench, the proposed menu increased online execution success from 0.737 to 0.898. It also outperformed baseline search, reranking, generation, and routing methods without modifying the agent.
The State-Path menu covered more complete chains with 32 tools than the official list did with 128 tools. The success gain persisted across executor families with models of varying capabilities.
These results reflect the authors’ evaluation on ToolBench. A criterion for interpreting them is to compare methods based on online execution success and complete-chain coverage, without attributing the result solely to menu size.
How to Choose an Approach to Menu Formation
For tasks with dependencies between actions, searching for relevant tools alone is not enough. The menu should include tools that produce missing inputs and place them before the actions that consume those inputs.
When comparing approaches, check three properties:
- Executability: Is the starting action available from the current state?
- Completeness: Are the producer tools and final action included?
- Order: Do producers come before consumers?
Use online execution success and complete-chain coverage to evaluate the result. This criterion reflects not only how well the tools match the query, but also whether the agent can follow the path to the result.
Source
The paper “The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents,” arXiv:2609.09395. The authors reported that the paper was accepted to EMNLP 2026 Main; the stated submission date is September 8, 2026.



