An AI coding agent can finish a request with an impressive summary: it changed six files, fixed the bug, and ran a command that exited successfully. That account is useful, but it is not proof that the user can complete the task. The code may compile while the page fails at runtime. The new route may load while the button does nothing. A screenshot may look correct while a form silently discards its data.
The gap between an edit and a working product is where I think agent workflows need the most attention. Consider an agent using Goose for a user interface migration and connecting Chrome DevTools through MCP to inspect the application it has changed. The example makes a simple point concrete: an agent needs a way to observe the result of its own work.
Start with a task the agent can actually verify
“Improve the app” is too broad to evaluate. “Move this page to the new framework while preserving the navigation and form flow” gives the agent a boundary. I would make that boundary more explicit with acceptance criteria:
- The page opens at the expected URL without a console error.
- The primary navigation reaches the same destinations.
- Submitting the form produces the expected success or validation state.
- The changed files contain no unrelated refactor.
These checks are not a magic prompt. They are a definition of what evidence the agent should bring back. The agent can then inspect the repository, make a focused change, run the application, and test the described flow. If it cannot access the required environment, it should report that limitation plainly.
A React-to-Next.js UI migration illustrates why a build alone is insufficient. A framework migration affects routing, data loading, assets, and the way a page behaves in a browser. The meaningful question is not whether the agent produced Next.js files. It is whether the behavior people rely on survived the move.
Give the agent a feedback loop
The strongest part of a browser-aware workflow is its feedback loop. The agent makes a change, opens the page, observes the result, and adjusts. Through a tool such as Chrome DevTools MCP, it can inspect browser state rather than relying only on a terminal log or its own prediction of what the code should do.
The observation must match the task. For a layout issue, a screenshot and computed styles may help. For a failed request, inspect the network response and the browser console. For a broken interaction, exercise the control and compare the visible state before and after. More tool access is valuable only when it closes the gap between the claimed result and the actual behavior.
I would ask an agent to report evidence in a short, reviewable format: which route it opened, which action it took, what it observed, and what remains unverified. “The app works” tells me little. “I opened /settings, changed the display name, saved, reloaded, and saw the new name persist” gives me a specific claim I can repeat.
This also changes how I read failures. A failed browser check is not merely an obstacle to hide from the user. It may reveal that the task was underspecified, the environment is missing a dependency, or the proposed implementation misunderstood the product. That is exactly the information a good development loop should surface.
Treat MCP access as a capability boundary
MCP can connect an agent to a browser, repository, design file, or other tool. Each connection changes what the agent can read or do. I would review a server's source and permissions before giving it access to a sensitive project, and I would keep access scoped to the task. Convenience does not remove the need to understand what a third-party tool can reach.
There is a second boundary inside the content the tool returns. A web page, issue description, or terminal output can contain instructions written by someone other than the developer. The agent should use that material as task data, not as authority to change its goals. This matters particularly when a browser tool opens untrusted pages while the agent has write access elsewhere.
Keep the human review specific
A developer should refine generated code instead of merging it as soon as the agent declares success. The agent can do the repetitive inspection and first implementation pass; the developer still owns the product decision and the merge.
My review would focus on four questions. Did the agent change only what the task required? Is the code understandable to the next maintainer? Does the runtime evidence cover the risky user path? Did the agent flag the parts it could not verify? A green test suite is part of that review, but it is not a substitute for it.
The agent workflow I want is therefore a loop, not a handoff: define the behavior, make the edit, observe the running system, compare the result with the acceptance criteria, and review the diff. If any step is missing, I know which claim remains uncertain. That is far more useful than a polished completion message.
For more on this topic, watch Talks with Ido Evergreen: Vibe Coding, MCP and AI Agents with Rizel Scarlett.