Skip to content

Engineering Hygiene

Coverage Is Not Confidence

A test suite's job is not to prove your code works. It is to give you permission to change it. Most suites are optimized for the first goal and quietly make the second one harder.

By Anshul Kapoor·October 2026·7 min read
TestingCode QualityEngineeringFrontend
LinkedIn

TL;DR

  • Coverage measures which lines ran during a test. It says nothing about whether anything meaningful was asserted, which is why a suite can hit 90 percent and still catch nothing.
  • A test that fails when you refactor working code is not protection. It is a second copy of the implementation that you now have to maintain.
  • The useful question is not "is this tested" but "if this broke, what would tell me, and how long would it take."
  • Test at the boundary a user or caller actually depends on. Everything below that line is implementation, and pinning it down is what makes suites brittle.

The number went up and the bugs did not go down

A team I worked with spent most of a quarter on test coverage. It had been flagged in an audit, a target was set, and the number moved from the low fifties to the high eighties. The dashboard turned green. Leadership was pleased.

Production incidents over the following quarter were unchanged.

That result is not as strange as it sounds, and it is worth sitting with rather than explaining away. The tests written during that push were real tests. They ran real code. They just happened to be written the way tests get written when the goal is a number: aimed at whatever was cheapest to cover. Getters, mappers, formatting helpers, a lot of expect(result).toBeDefined(). The hard parts of the system, the ones with branching logic and external calls and genuine ambiguity about correct behavior, were exactly as expensive to test as they had been before, so they stayed uncovered.

Coverage tells you which lines executed while the suite ran. It does not tell you whether anything was asserted about them, whether the assertion was meaningful, or whether the scenario resembled anything a user does. You can reach a very high number with tests that would pass if you deleted half the logic in the functions under test.

“Coverage measures which lines ran. It has no opinion about whether anything was checked when they did.”

What a test suite is actually for

The framing that fixed this for me is that a test suite is not primarily evidence. It is permission.

When you have a good suite and you need to restructure a module, you make the change, run the tests, and find out in ninety seconds whether you broke something a caller depends on. That speed is the entire value. It converts "we should clean that up eventually" into something a person can do on a Tuesday afternoon without booking a risk review.

When the suite is bad, that same change turns red in forty places, none of which represent a real defect, and you spend the afternoon updating tests to match the code you just wrote. Do that twice and you learn the lesson the codebase is teaching: changing this area is expensive. So you stop. You add a parameter instead of restructuring. You write the new thing beside the old thing. The system gets worse, not because anyone chose that, but because the tests made the better path costly.

That is the failure mode I care about most, because it is invisible in every metric. Nobody files a ticket saying the test suite prevented a refactor. The refactor simply does not happen, and the reason never gets written down.

“A bad suite does not announce itself with failures. It shows up as the refactors your team quietly stopped attempting.”

Tests that fail for the wrong reason

There is a clean way to sort tests: when this fails, have I learned that something is broken, or only that something is different?

Tests that answer "different" tend to share a cause. They reach past the interface and assert on machinery.

// Fails whenever the implementation changes, passes whenever behavior breaks
it('fetches the user', async () => {
  const getJson = jest.fn().mockResolvedValue({ id: '1', name: 'Ada' });
  const cache = { get: jest.fn(), set: jest.fn() };
  const service = new UserService(getJson, cache);
 
  await service.load('1');
 
  expect(getJson).toHaveBeenCalledTimes(1);
  expect(getJson).toHaveBeenCalledWith('/api/users/1');
  expect(cache.set).toHaveBeenCalledWith('user:1', expect.anything());
});

Read what that actually asserts. Not that loading a user returns the user. Not that a second load avoids a redundant fetch. It asserts that the code makes a particular call, with a particular URL string, in a particular order, to collaborators the caller does not know exist. Switch to a different HTTP client, change the cache key format, add a retry, and it goes red while the behavior is completely intact.

The version that earns its place asserts the thing a caller depends on.

it('returns the user and serves the second read from cache', async () => {
  const server = fakeUserApi({ '1': { id: '1', name: 'Ada' } });
  const service = new UserService(server.fetch, new MemoryCache());
 
  expect(await service.load('1')).toEqual({ id: '1', name: 'Ada' });
 
  await service.load('1');
  expect(server.requestCount).toBe(1);
});

Same coverage. Completely different relationship with change. The second test still catches a broken cache, which is the behavior anyone cares about, but it does not care how the cache is keyed or which client makes the request. You can rewrite the internals entirely and it keeps holding you to the promise that matters.

Snapshot tests deserve a specific mention, because they are the purest form of this problem. A large snapshot asserts that the output is identical to what it was, which means every intentional change produces a diff no reviewer reads carefully, and the ritual response is to press u. After the third or fourth time, the snapshot has stopped being an assertion and become a transcript.

The pyramid nobody actually has

Most teams can describe the testing pyramid. Many unit tests, fewer integration tests, a handful of end-to-end tests.

The suites I have actually opened rarely look like that. They look like a wide base of unit tests covering the code that was easy to unit test, a thin and unloved middle, and a few end-to-end tests that are quarantined because they are flaky. The shape is real, but it was produced by what was convenient to write rather than by where the risk lives.

The distortion has an obvious cause. Pure functions are pleasant to test, so they get thorough tests they mostly do not need. Anything that touches a network boundary, a database, or a state machine spanning several components is annoying to test, so it gets a mock and a shallow assertion. The result is a suite that is extremely confident about a date formatter and has nothing to say about whether checkout works.

Worth being blunt about the arithmetic: a unit test proves a unit behaves correctly in isolation. Systems mostly fail at the seams, in the wiring between units that each work fine. A suite composed entirely of unit tests is testing the parts of the system least likely to be the thing that broke.

A useful diagnostic: take your three worst incidents from the last year and ask, for each one, which test would have caught it if it had existed. If the honest answer is repeatedly "an integration test we do not have," the shape of the suite is wrong regardless of what the coverage number says.

Testing when you cannot test everything

You cannot test everything, and pretending otherwise is how teams end up with a suite that is both enormous and unhelpful. The question is where finite effort goes.

The most useful sorting I know weighs two things: how likely is this to break, and how bad is it if it does. That sounds obvious written down, and almost nobody applies it deliberately, because coverage tooling reports every file as equally worth covering.

Under that lens, the payment path and the auth flow deserve tests at several levels, including the slow expensive kind, because the cost of a failure is unbounded. A settings page that toggles a boolean deserves considerably less than it usually gets. Code that has broken before deserves a test now, since past breakage is the best available predictor of future breakage. Code nobody has touched in two years and that has never failed can wait.

There is one more category worth naming, because it is where the highest-value tests usually live: the places where correct behavior is genuinely ambiguous. Rounding rules. Timezone handling. Retry and backoff semantics. What happens when a request succeeds after the user navigated away. Those tests are valuable in a way the coverage number cannot see, because they are documentation of a decision. When someone changes that behavior in a year, the test failing is the only thing standing between them and silently reversing a choice that was made carefully.

The question that replaces the number

I am not arguing against measuring coverage. It is a fine smoke detector. A module at zero percent is worth a look, and a sudden drop usually means something got merged without tests. As a signal it is cheap and occasionally informative.

The trouble is exclusively what happens when it becomes a target. Numbers that become targets get satisfied in the cheapest available way, and for coverage the cheapest way is to write tests that execute code without meaningfully checking it. The metric goes up. The property it was supposed to proxy for does not move.

The question worth asking in review is harder to automate and much more useful: if this behavior broke tomorrow, what would tell us, and how quickly? Sometimes the honest answer is a test, and you write it. Sometimes it is a type, and the test would have been redundant. Sometimes it is a monitor or an alert, because the failure is operational and no test was ever going to catch it. Sometimes the answer is "a customer would tell us, and that is acceptable for this feature," which is a legitimate engineering decision as long as it is made on purpose.

A suite built by asking that question looks strange next to a coverage dashboard. It is uneven. Some files are covered exhaustively and others not at all, and the ratio does not tell a tidy story. It is also the kind of suite that people trust enough to lean on, and that is the only outcome that ever mattered.

The goal was never to have tested the code. It was to be able to change it tomorrow without being afraid.

“Coverage answers whether a line ran. The question worth asking is whether anyone would find out if it stopped working.”

Found this useful?

Share it with someone who might be wrestling with the same problem.

ShareLinkedIn

Thanks for reading.

If this resonated and you're hiring for Senior/Staff Frontend roles, I'd love to chat.