Test the Contract, Not the Implementation

December 16, 2025

The easiest test to write is often the one closest to the code in front of you. A component calls a helper, so we spy on the helper. A function maps an array, so we assert that map was called. The test passes, and the coverage number rises. Then a harmless refactor breaks it.

The difficult question is not whether we can test a detail. It is whether a failure of that test would tell us something the user or another developer actually needs to know.

Tests should protect an intention: what the software promises to do. Take a passkey login as an example. The useful requirement is that a user can authenticate with a passkey. The location of the button, the name of an internal helper, and the sequence of private function calls are implementation choices. They may change while the promise remains intact.

Translate a feature into a contract

Before choosing a test runner, write the behavior in a sentence a product teammate could understand. For example:

A signed-in user can save an edited task and see the new value after reloading.

This sentence exposes several conditions that a test must cover: authentication, persistence, and the value shown after a fresh read. It does not prescribe a React hook, a REST endpoint, or a database query. Those are ways to implement the contract.

The test can then choose the narrowest boundary that proves the behavior. A pure formatting rule may need a small unit test. A task editor that coordinates form validation, a request, and UI state may need a component or integration test. Persistence after reload may require an end-to-end test against a representative environment. “Unit,” “integration,” and “end to end” describe the radius of the test; they do not rank its value.

The responsibility of a test changes with the boundary it covers. Testing a function in isolation is different from testing the whole application from the user's perspective. The useful question is which boundary can reveal the failure you care about.

Design the failure before the assertion

Design tests around failure. We write a test because we want an alarm when an expectation stops holding. An alarm that rings for every internal rearrangement is noise. An alarm that never rings when the user journey breaks is false comfort.

Consider two tests for a search form. One checks that clicking Search calls setQuery with a particular string. The other enters a term, submits the form, and checks that matching results appear. The first may be useful if setQuery is a public contract of a library. In an application, it mostly records how today's component is wired. The second checks the user-visible effect.

The distinction matters during maintenance. If the form switches from local state to URL state, the first test fails although the search still works. If the request fails silently and the UI remains empty, the first test can pass while the user is stuck. Playwright's testing guidance recommends verifying user-visible behavior and avoiding implementation details for this reason.

An assertion should also make a failure easy to diagnose. “The saved task appears after reload” is a stronger signal than a long chain of assertions checking every intermediate variable. Intermediate checks are useful when they distinguish different failure modes. They become a burden when they duplicate the same contract or pin down incidental steps.

Coverage is a map, not a shipping decision

Coverage tells you which lines or branches a suite exercised. It does not tell you whether the tests protect the right promises. A suite can visit every line of a checkout component without proving that a customer can complete an order. A single end-to-end test of that flow can provide more confidence about the customer journey while leaving many implementation lines uncovered.

This does not make coverage useless. A large untested area can be a useful prompt to investigate. The mistake is treating the percentage as the objective. Overtesting often comes from the urge to cover every internal detail because the developer knows those details well. A better question is what the software is intended to do and whether that behavior is covered.

For a new feature, I would make a short contract inventory:

  • What should work when the input is valid?
  • What should the user see when the input is invalid?
  • What happens when a dependency is slow or fails?
  • Which result must survive a reload or a new session?
  • Which public API or type contract is consumed by other code?

That inventory usually identifies more valuable tests than a target coverage percentage. It also tells us where a lower-level test is enough and where a real browser or service boundary is necessary.

Let the product shape the test suite

A library may need type tests and compatibility checks across runtime and TypeScript versions. A product dashboard has a different risk profile. Its highest-value checks might be authentication, data editing, permissions, and the few journeys that generate revenue or support requests. The same testing rule applies, but the test mix changes.

Suppose a team maintains five login methods. Organizing tests by provider can make failures easier to locate. A passkey flow may need a browser-level test because the browser's credential behavior matters. A function that normalizes a callback URL can be tested directly. A permission rule should be checked at the server boundary even if the UI hides a button. One test type cannot substitute for all the others.

I also would not make every test an end-to-end test. Browser tests provide fidelity, but they cost more to set up and can fail for reasons outside the feature under test. Lower-level tests give faster feedback for rules with a clear input and output. The trick is to avoid mocking away the very behavior you wanted to prove.

Keep the suite readable under pressure

A test suite earns its keep when a change breaks it at an inconvenient moment. At that point, the person reading the failure may not be the person who wrote the test. Names, setup, and assertions become an interface for that future reader.

Short names that state the expectation make tests easier to review. “Supports passkey authentication” carries more information than “calls auth handler,” because it points to the contract. Test code also needs restraint: a large test with mutated setup state can be difficult to debug. Reuse can be helpful, but shared hooks, loops, and conditional branches can make the cause of a failure harder to see than a little duplication would.

A practical review question is: if this test fails six months from now, will the failure tell us which promise broke? If the answer is no, revise the test before adding another one like it.

The goal is a suite that preserves the freedoms we want in implementation while sounding an alarm when the software stops doing its job. That is a more useful target than simply writing more tests.

For more on this topic, watch Talks with Ido Evergreen: Software Testing 101 with Artem Zakharchenko.