Ask ten engineers what “good test coverage” means and you’ll get ten different numbers, usually pulled from a percentage someone saw in a code-quality dashboard rather than from any principled reasoning about what actually needs testing. Unit testing strategy suffers from a metric problem: coverage percentage is easy to measure and easy to put in a dashboard, which makes it the thing teams optimize for, even when it’s a weak proxy for whether the tests actually catch the bugs that matter.
The uncomfortable truth is that a codebase can hit 100% line coverage and still ship broken releases, because coverage measures which lines executed during a test run, not whether the test actually verified the behavior that matters. Here’s how to think about unit testing as a strategy rather than a number to hit.
What actually deserves a unit test?
Not everything, and treating “test everything” as the goal produces a test suite that’s expensive to maintain and doesn’t concentrate effort where bugs actually hide. Prioritize:
Business logic with real branching complexity — pricing calculations, permission checks, state-machine transitions, anything with more than one or two conditional paths where a wrong branch produces a wrong answer silently.
Code that has broken before. A bug that shipped once is a strong signal that the surrounding logic is more subtle than it looks. A regression test for the specific case that broke is one of the highest-value tests you can write, because it directly prevents a known failure mode from recurring.
Edge cases at data boundaries — empty inputs, maximum values, off-by-one conditions, null/undefined handling — the places where developers’ mental model of “the happy path” tends to diverge from what actually happens.
Deliberately skip or minimize testing on thin wrapper code (a function that does nothing but call a well-tested library function), simple data transfer objects with no logic, and framework-generated boilerplate. Tests on code with no branching logic mostly verify that the language still works, at real maintenance cost every time the wrapper’s shape changes.
Why doesn’t code coverage percentage tell you what you actually want to know?
Coverage answers “did this line execute during a test,” not “did a test verify this line’s behavior is correct.” A test that calls a function and asserts nothing about its output will show that function as covered, while catching zero bugs. NIST’s often-cited 2002 study on the economic costs of inadequate software testing infrastructure (“The Economic Impacts of Inadequate Infrastructure for Software Testing”) found that a large share of the cost of software bugs comes from defects that make it well past unit-level testing into integration and production — a reminder that coverage percentage was never intended as the whole story, and treating it as a completeness proxy misreads what it measures (NIST Planning Report 02-3, 2002).
A more useful question than “what’s our coverage percentage” is “if this specific function had a bug, would any of our tests fail?” That’s a mutation-testing framing rather than a coverage-percentage framing, and it’s a genuinely different (and more expensive) thing to measure — most teams don’t need full mutation testing, but keeping the question in mind while writing tests produces meaningfully better ones than chasing a coverage number.
How should mocking fit into a unit testing strategy?
Mocking — replacing a real dependency (a database call, an external API, a file-system operation) with a fake stand-in that returns controlled values — is necessary for genuine unit isolation, but it’s also the single easiest way to write tests that pass without verifying anything real.
Mock at the boundary of your system, not inside your own logic. Mocking an external payment API so you can test your own retry logic without making real network calls is a legitimate use — you don’t control that dependency’s behavior, and you shouldn’t need a live connection to test your handling of it. Mocking an internal function so heavily that the test only verifies “this function was called with these arguments” — without any of the real logic executing — produces a test that will still pass after the underlying logic breaks, because it never actually exercised that logic.
A rough heuristic: if a test would still pass after you deleted the body of the function it’s testing and replaced it with return None, the test isn’t testing what you think it’s testing.
How do integration and end-to-end tests fit alongside unit tests?
Unit tests verify a single function or class in isolation and should be fast — hundreds or thousands running in seconds. Integration tests verify that multiple components work together correctly (your code plus a real database, for instance) and are inherently slower and more brittle. End-to-end tests verify the whole system behaves correctly from a user’s perspective and are the slowest and most expensive to maintain of the three.
The common failure mode is either testing everything at the end-to-end layer (slow test suites that take an hour and fail for unrelated reasons) or testing everything at the unit layer with heavy mocking (fast suites that pass while the system is actually broken, because the mocks hid a real integration problem). A workable default: most of your test volume in fast, well-targeted unit tests; a smaller layer of integration tests for the specific seams where components actually need to cooperate correctly (database queries, API contracts); a thin layer of end-to-end tests for the critical user paths that would be catastrophic to break.
Setting this strategy up early is closely tied to how a team already thinks about its backend framework choice and its git branching strategy — a fast test suite that runs on every branch push is what makes trunk-based or short-lived-branch workflows safe in the first place. Teams that skip investing in test speed early often find their branching model breaking down later, not because the model was wrong, but because nobody trusts the test suite enough to merge quickly.
How do you retrofit a testing strategy onto a codebase that has almost none?
This is a more common situation than teams like to admit, and the instinct to “just start writing tests” without a plan usually produces a scattered handful of tests on whatever code someone happened to be touching, with no coherent coverage of the areas that actually matter most.
A more effective approach: start with the highest-risk, highest-change-frequency code first — the modules where bugs are most expensive and most likely, not the easiest code to test. Write characterization tests that capture current behavior before making any other changes to that code, even if the current behavior isn’t ideal; this gives you a safety net for future changes without requiring you to first decide what the “correct” behavior should have been. Resist the temptation to chase a coverage percentage across the whole codebase during this phase — it rewards testing easy, low-value code just to move the number, instead of testing the code that’s actually risky.
Set a policy that new and meaningfully changed code requires tests going forward, even while legacy code remains untested. This “coverage ratchet” approach — the tested proportion of the codebase only grows, never shrinks, because new work always ships with tests even if old work doesn’t yet have them — is a far more sustainable path to a well-tested codebase than a big-bang effort to test everything retroactively, which rarely survives contact with the next urgent feature deadline.
What does a healthy test suite look like at the team-process level, not just the code level?
Beyond the tests themselves, a few process signals distinguish a testing strategy that’s actually working from one that exists on paper. Tests run automatically on every change, not manually before a release — a test suite that requires someone to remember to run it gets skipped under deadline pressure, which is exactly when tests matter most. Failing tests block merges rather than producing a warning someone can ignore; a test suite whose failures are treated as optional stops being trusted within a few sprints. And test failures get investigated the same day, not left to accumulate — a backlog of “known failing” tests is functionally the same as having no tests for that code, except with the added cost of maintaining tests nobody’s acting on.
Frequently Asked Questions
What’s a reasonable code coverage target?
There isn’t a universal number worth chasing. A more useful practice is tracking coverage on new and changed code specifically — requiring meaningful tests on what’s actually being written today — rather than treating an aggregate percentage across the whole codebase as a quality goal in itself.
Should every pull request include new tests?
For any change to logic with branching complexity, yes. For pure refactors that don’t change behavior, existing tests passing unchanged is itself the signal of correctness — adding new tests there is optional unless the refactor also expands what the code needs to handle.
Is TDD (test-driven development) necessary for a good testing strategy?
No. Writing tests before implementation is one valid workflow, not a prerequisite for a healthy test suite. What matters is that tests exist, are meaningful, and get maintained — whether they’re written before, during, or immediately after the corresponding code.
How do you fix a test suite that’s become slow and unreliable?
Start by separating true unit tests (no I/O, no network, no real database) from everything else, since misclassified integration tests running in the “unit” suite are the most common cause of slowness. Then audit for flaky tests — ones that fail intermittently for reasons unrelated to actual bugs — and either fix or quarantine them, since a team that ignores intermittent failures quickly stops trusting the suite at all.
