Every codebase eventually accumulates the same three copies of the same date-formatting function, written by three different people who each didn’t know the other two existed. Duplicated logic isn’t a sign of a bad team — it’s what happens by default in any codebase that grows faster than anyone has time to go back and organize it. The question isn’t how to prevent duplication from ever appearing; it’s how to build a habit of finding it and extracting it before it compounds into three subtly different implementations that all handle edge cases differently.
Organizing reusable code well is a different skill from writing correct code, and it’s the one that determines whether a codebase gets easier or harder to work in as it grows.
When is duplicated code actually a problem?
Not immediately, and this is worth saying explicitly because “don’t repeat yourself” gets treated as an absolute rule when it’s really a judgment call. Two pieces of code that look similar today but represent genuinely different concepts — a formatDate for user-facing display and a formatDate for a log timestamp — will diverge in requirements over time even if they started identical. Merging them prematurely into one shared function creates a different problem: a function with an ever-growing list of configuration flags trying to serve unrelated callers, which is often worse than the duplication it replaced.
The useful signal isn’t “this code looks similar,” it’s “this code represents the same concept and needs to change together.” Two functions that happen to look alike but serve genuinely different business rules aren’t duplication in the sense that matters — they’re coincidental similarity, and forcing them into one shared abstraction usually creates coupling between things that should be allowed to evolve independently.
What’s the difference between a utility function and something that deserves its own internal package?
A utility function lives close to where it’s used — in a shared utils module within a single service or application — and its blast radius if it changes is limited to that one codebase. An internal package is versioned and consumed by multiple, independently deployed services or applications, which means a breaking change to it has to be coordinated across every consumer.
The jump from “utility function” to “internal package” should be deliberate, not automatic, because it comes with real overhead: a package needs its own versioning scheme, its own test suite, its own release process, and — critically — a plan for how consumers pick up updates. Extracting a function into a shared package the moment two files use it is usually premature; the overhead of maintaining a package is only worth paying once the reuse is genuinely cross-team or cross-service, not just cross-file.
How do you know when to extract shared code versus leaving it duplicated a little longer?
A practical heuristic: wait for the third occurrence, not the second. The first implementation is just code. The second occurrence might be coincidence, or it might be genuine duplication — you often can’t tell yet which requirements are shared and which are incidental. By the third occurrence, a real pattern is usually visible, and you have enough examples to design an abstraction that actually fits all three use cases rather than guessing from two data points and getting the interface wrong.
This isn’t a rigid rule so much as a discipline against the more common failure: extracting an abstraction too early, from a sample size of two, and then bending every subsequent caller to fit an interface that was designed around too little information. An abstraction extracted too early tends to accumulate special-case parameters as new callers reveal requirements the original two didn’t have — which is a worse outcome than the duplication it was meant to prevent.
Where should shared code physically live?
This depends heavily on your repository topology and, more directly, on whether you’re operating a single repository or several. In a monorepo, shared code typically lives in a top-level packages/ or libs/ directory, imported directly by consumers with no publishing step required — the same commit that changes the shared code and updates its callers can merge atomically, which is one of the concrete advantages of that structure. In a polyrepo setup, shared code has to be extracted into its own repository, versioned, and published to a package registry (internal or public), with consumers pinning to a specific version and upgrading deliberately.
Neither location is universally correct; it follows from whichever repository structure the team has already chosen, and retrofitting shared-code organization onto the wrong repository topology is usually more painful than picking the topology with reuse patterns in mind from the start. Getting the underlying build tooling right matters here too — a monorepo without incremental, dependency-aware builds turns every change to shared code into a full rebuild of everything downstream, which quickly makes teams avoid touching shared code at all, defeating the point of consolidating it.
What are the warning signs that reusable code has been organized badly?
A shared function with more than four or five boolean flags. Each flag usually represents a caller-specific requirement that got bolted onto a general-purpose function instead of being handled by a caller-specific wrapper or a genuinely separate function.
Nobody knows who owns a shared package. Shared code without a clear owning team tends to accumulate inconsistent conventions as different consumers patch it for their own needs, with no one accountable for keeping the whole thing coherent.
A “utils” or “common” module that’s become a dumping ground. If a module’s name doesn’t describe what it actually contains, it’s probably become a catch-all for “code that didn’t fit anywhere else” rather than a coherent, purposeful abstraction — and that’s a sign it’s due for a split into more specific, better-named modules.
Consumers copy-pasting from a shared package instead of importing it, because importing the actual dependency is harder than duplicating three lines. That’s a signal the extraction itself — its packaging, its documentation, its ease of adoption — needs attention, not just the code inside it.
How do you name and document a shared package so people actually find and use it?
A shared package’s discoverability matters as much as its correctness — a well-designed internal library that nobody knows exists gets reinvented independently by the next team that hits the same problem. Give shared packages names that describe what they do, not who built them or when (“date-formatting” rather than “platform-utils-v2”), since a descriptive name is itself a form of documentation that survives long after the original context is forgotten.
Document the package’s intended scope explicitly, including what it’s not for — a README that says “this package handles currency formatting for display; it does not handle currency arithmetic or rounding for financial calculations” prevents a common failure mode where a consumer stretches a package to cover a use case it was never designed for, discovers an edge case it doesn’t handle, and either patches around it locally (defeating the purpose of sharing) or files a change request that pulls the package’s scope in a direction its original design didn’t anticipate.
Does organizing reusable code differently change as a team grows?
Yes, substantially. A five-person team can rely on shared context and direct conversation to know what shared utilities exist and who’s responsible for them — an informal utils folder with no separate ownership works fine at that scale because the tacit-knowledge cost of not documenting it is low. A fifty-person organization spanning multiple teams can’t rely on that same shared context; the same informal approach produces duplicate implementations because no single person has visibility into what already exists across every team.
The organizational answer scales alongside the technical one: as headcount grows, invest more deliberately in discoverability (a searchable internal package registry or catalog, consistent naming conventions enforced by convention or tooling) and in explicit ownership (a team or individual accountable for a shared package’s quality and roadmap, rather than an ownerless artifact that accumulates inconsistent patches from whoever touched it last). Retrofitting this structure onto an already-large, already-informally-organized codebase is possible but harder than building the habit in from a smaller scale, which is the strongest argument for taking reuse organization seriously well before it feels urgent.
Frequently Asked Questions
Is the “rule of three” for extracting shared code a strict rule?
No — it’s a heuristic against extracting abstractions too early from too little information, not a mandate to wait exactly until the third occurrence. If a second occurrence is clearly the same concept with an obvious shared interface, extracting then is fine. The point is resisting the urge to abstract from a sample size of one or two when the shape of the real requirement isn’t yet clear.
How do you handle a shared package that needs a breaking change?
Version it explicitly and let consumers upgrade on their own schedule rather than forcing an instant, coordinated migration — the same pattern any external dependency uses. In a monorepo, this is less of a concern since changes can often be made atomically across all consumers in one commit; in a polyrepo, a deprecation period with both old and new versions available briefly is usually the safest path.
What’s the risk of over-abstracting shared code?
An abstraction that tries to serve too many different callers’ needs accumulates configuration complexity until it’s harder to understand and modify than the duplicated code it replaced. Over-abstraction is a real and common failure mode, not just a theoretical one — it’s usually caused by extracting too early, before the actual shape of the reusable concept was clear.
Should every team have a dedicated platform or shared-libraries team?
Only once the organization is large enough that shared code genuinely needs a dedicated owner to keep it coherent — usually multiple product teams actively depending on the same internal libraries. Smaller organizations are often better served by rotating ownership or a lightweight review process for changes to shared code, rather than standing up a dedicated team before there’s enough shared surface area to justify it.
