Design WebMCP Tools Around Tasks, Not Buttons

October 7, 2026

A web app can expose tools to an AI agent, but that does not mean every button deserves a tool. A table with sorting, filters, row menus, and edit dialogs might have dozens of controls. Mirroring them all gives the agent a longer menu without telling it which sequence completes a user's request.

The more interesting design question is: what work should the agent be able to finish?

A to-do app makes the problem concrete. Instead of publishing separate tools for every small action, it can expose a small set of flexible operations, such as searching and editing tasks. The point is familiar to anyone who has designed an API: the shape of the interface should reflect useful capabilities, not the accidental layout of a screen.

Start with the user's operation

Imagine a user asks, “Find my overdue tasks for this project and move the urgent ones to tomorrow.” A tool named click_next_page is too close to the UI. A tool named manage_everything is too vague. A better interface might offer search_tasks and update_task with explicit, validated inputs.

// An illustrative tool contract, not a copy of any application's API.
type SearchTasksInput = {
  projectId: string
  dueBefore?: string
  status?: 'open' | 'done'
}

type UpdateTaskInput = {
  taskId: string
  changes: {
    dueDate?: string
    priority?: 'normal' | 'urgent'
  }
}

This split gives the agent a way to inspect state before changing it. It also makes the write operation narrow enough for the app to validate. If the agent has to discover which rows are overdue, search_tasks can return stable IDs and relevant fields. If the user asks for a change, update_task can require an ID and reject invalid dates or unauthorized edits.

The tool's output matters as much as its input. A search should return enough information to tell similar tasks apart, including the stable ID the update needs. An update should say which record changed and what state it now has. If the request fails, the error should distinguish “task not found,” “you cannot edit this project,” and “date is invalid.” An agent can recover from a specific failure; a generic success flag or silent no-op makes it guess.

The right granularity depends on the product. A design editor may need a tool for changing a selection's properties; a checkout flow may need a carefully bounded submit_order operation. The test is whether the tool helps complete a recognizable task without hiding consequential work behind an ambiguous name.

Tool design is also context design

An agent chooses from the tools it can see. Every extra tool adds another description to interpret and another chance to pick the wrong operation. Combining related actions can make that choice easier, but only if the combined contract remains understandable. One giant tool with an unstructured action string merely moves the confusion into its arguments.

I would review a proposed tool using four questions:

  1. Can a developer explain when an agent should call it in one sentence?
  2. Do its name and schema distinguish a read from a write?
  3. Can the app validate the input and report a useful failure?
  4. Can the user see or confirm the effect when the operation matters?

Those questions are especially useful when a team starts with an existing UI. It is tempting to export the component tree as a tool list because the controls already exist. But a visual control is often only one step in a larger operation. Conversely, a single screen may contain several meaningful tasks with different permissions and risks.

Keep the browser's authority in view

WebMCP's proposed imperative API lets a page register a tool with a name, description, input schema, and execute function. The important architectural detail is that the page implements the operation. The tool should use the application's existing state, authorization checks, and validation. Exposing a tool does not grant an agent permission the user does not have.

For a read operation, the description should say what data is returned. For a write operation, it should say what changes and whether the effect is reversible. The WebMCP guidance on secure tools discusses the security considerations for tools exposed by a page, while the imperative API documentation describes the registration shape and tool annotations. Those annotations can help describe intent, but the app still needs to enforce the operation's rules.

There is a practical maintenance benefit here too. If an agent must click through a menu, modal, and form, every UI change can break that sequence. A task-level tool gives the app one place to maintain the operation's contract. That does not eliminate maintenance: schemas, permissions, and error messages still evolve. It makes the boundary explicit enough to test.

A review method for an existing app

Take one common user request and trace it through the product. Write down the data the agent needs to read, the state it needs to change, and any moment that deserves user confirmation. Then propose the smallest tool set that supports that workflow. Try a second request that is similar but not identical. If the tools cannot express it without adding a new one for every button, the boundary is probably too narrow. If the tool can perform unrelated changes from the same vague payload, it is probably too broad.

WebMCP is still a developing web platform proposal, so I would treat a published tool contract as an experiment to review with real tasks and failure cases. The durable lesson is older than the proposal: good interfaces name meaningful operations, constrain their inputs, and make their effects legible.

For more on this topic, watch Talks with Ido Evergreen: Accessible Agent Web? WebMCP with Alex Nahas.