Paradox of "Build" vs "Buy"

INTRODUCTION - "The Question Nobody Answers Honestly"
Picture a meeting room. 12 people around a table. A slide deck projected at the front. On screen: a project proposal for a new core insurance platform, built in-house, estimated at $20M and 12 months till Go-Live.
Someone raises their hand and asks whether the company has evaluated commercial alternatives. The room shifts. The head of engineering explains, patiently, that they have looked at vendors. None of them understand the nuances of the business. None of them can handle the specific workflow. None of them offer the level of control the company needs. The team has the talent. The budget is approved.
18 months later, the project is at $35M and 9 months from delivery. The slide estimates were wrong. They always are.
I have been in that room on both sides of the table. And in all that time, across every industry and every technology generation, I have never once heard a logically compelling argument for building software that a purpose-built vendor already offers at scale. What I have heard, repeatedly, is a variation of the same words: “We need or want to do it ourselves.”
That is not a strategy. That is a preference. And preferences, when they carry 9-figure consequences, deserve a great deal more scrutiny than they typically receive.
The desire to build is not a strategy. It is a bias. And it is one of the most expensive biases in modern business.
This article is my attempt to end the argument. Not by declaring build always wrong and buy always right - that would be too simple, and I will address the exceptions honestly later. But by laying out what the data, the behavioral science, the organizational dynamics, and three decades of documented failures actually show. The picture is not ambiguous. And it is time we stopped pretending it is.
This is not a new debate. But it has become an urgent one. The emergence of generative AI, agentic systems, and AI-powered coding tools has added a dangerous new dimension: the illusion that because software is now faster to prototype, it is cheaper to own. It is not. The planning fallacy is alive and well in the age of LLMs. And the psychological traps that drove the wrong decision in 2005 are driving it again in 2025, this time with faster prototypes and bigger budgets.
CHAPTER # 1 - "The Seduction of Building"
There is something deeply human about wanting to build things. Not just in software, but in everything. A child prefers the drawing they made over a printed copy of a better drawing. A chef who cooks a meal at home will rate it more highly than restaurant food of comparable quality. A homeowner who renovates their kitchen will, on average, overestimate its value by tens of thousands of dollars. We are wired to overvalue what we create. It is biology.
In 2012, researchers Michael Norton, Daniel Mochon, and Dan Ariely at Harvard published a study that gave this phenomenon its name. They gave two groups of participants identical IKEA storage boxes. One group assembled them from instructions. The other group received them pre-built. When asked how much they would pay for the boxes, the assemblers valued theirs at 63% more than the pre-built group valued identical boxes. The researchers called it the IKEA Effect: labor alone is sufficient to make us overvalue what we create, regardless of quality.
The effect becomes even more striking when participants build something genuinely mediocre. Amateur origami creators, in a follow-up experiment, valued their own folded creations nearly as highly as expert origami and expected strangers to share their inflated opinions. They did not. The bias is entirely internal. The creator sees something others simply do not.
Now apply this to software. A development team spends 18 months building an internal platform. They have made hundreds of micro-decisions along the way: the database architecture, the API structure, the authentication system, the UI framework, the deployment pipeline. They have debugged it at midnight. They have defended it in steering committees. They have bonded over it in late-night pizza sessions. By the time it ships, they believe it is excellent. Not because it objectively is but because they built it. The IKEA Effect does not distinguish between a flat-pack cabinet and a million-dollar software system.
The IKEA Effect has one critical boundary condition: it dissolves when the builder fails to complete the work. But in software, teams routinely declare completion at launch, before the years of maintenance, security patching, integration upkeep, and feature parity work begin. The project is “done.” The bias has permanently fused to the product. And when a vendor later offers equivalent functionality at a fraction of the cost, the team will reject it not on rational grounds, but because the internal system has become “theirs.”
The Endowment Effect: Ownership Changes Everything
Closely related to the IKEA Effect is the endowment effect, first named by Nobel laureate Richard Thaler in 1980. The endowment effect describes how people assign more value to objects they own than to identical objects they do not own simply because ownership creates psychological attachment. In Thaler’s classic experiments, people who were randomly assigned a coffee mug demanded roughly twice as much to sell it as others were willing to pay to buy an identical mug. Ownership alone, with no change in the object, changed its perceived value.
Wharton researcher Michael Platt took this further, studying the biological mechanism behind the endowment effect. His findings showed that owners and non-owners literally process information about an object differently. Owners focus attention on desirable features; buyers focus on undesirable ones. The same object, seen through different eyes, looks like a different object. When a company owns its internal software, its team sees all the features it has and forgets about the features it lacks. When evaluating a vendor’s product, the same team scrutinizes every gap and limitation. The comparison is structurally unfair before a single number is written down.
Owen Jones at Vanderbilt Law School has argued that the endowment effect has evolutionary roots: in ancestral environments, holding on to what you possessed was often more rational than trading it. The brain developed a default bias toward retention. What served human survival in a resource-scarce world now costs organizations billions in unnecessarily retained legacy software.
The Planning Fallacy: Why the Timeline Is Always Wrong
In 1979, Daniel Kahneman and Amos Tversky identified what they called the planning fallacy: a systematic tendency to underestimate the time, cost, and risk of future actions while simultaneously overestimating the benefits. The mechanism is what Kahneman later described as the “inside view” problem: when planning a project, people naturally focus on the specific features of their project and construct a best-case scenario, rather than asking how similar projects have actually performed historically.
The planning fallacy is so pervasive that researchers studying it have produced their own grim irony: students asked to estimate completion dates with 99% confidence, meaning they were almost certain, actually finished by that date only 45% of the time. People are dramatically overconfident in their predictions, and the overconfidence persists even when they are explicitly told to account for it.
In software specifically, researcher Magne Jørgensen at the University of Oslo found that software professionals’ 90% confidence intervals contained the actual effort only 60 to 70% of the time. Projects consume 30 to 40% more effort than estimated on average. This finding has not improved in 3 decades of research. It appears to be a structural feature of how human beings think about complex future work, not a correctable project management problem.
Tom Cargill of Bell Labs captured the consequence in what became known as the Ninety-Ninety Rule: “The first 90% of the code accounts for the first 90% of the development time. The remaining 10% of the code accounts for the other 90% of the development time.” Every engineer who has shipped software recognizes this immediately. Every project plan that preceded it pretended otherwise. The planning fallacy does not disappear with experience. If anything, experience with complex projects can make the bias worse, because experienced teams have more elaborate mental models with which to construct internally consistent, plausible-sounding, and entirely wrong estimates.
Optimism Bias and the Iron Law
Beneath the planning fallacy lies a more fundamental force: optimism bias. Human beings are constitutionally inclined to believe that their own projects will succeed, that their own forecasts are reliable, and that risks will not materialize for them the way they do for others. This is not irrational in all contexts - optimism is a functional emotional state that sustains effort in the face of uncertainty. But in the specific context of large project planning, it is catastrophically expensive.
Bent Flyvbjerg, professor at Oxford University and the world’s leading researcher on megaproject failure, has spent his career documenting what he calls the Iron Law of Projects. His database covers more than 16,000 projects across 136 countries and nearly every category of large initiative. The findings are consistent to the point of being almost mechanical: 91.5% of projects go over budget, over schedule, or both. Less than 0.5% deliver on budget, on time, and with all promised benefits.
For IT projects specifically, Flyvbjerg found that 18% have cost overruns above 50%, and among those projects, the average overrun is 447%. Not 47%. Four hundred and forty-seven. This is not a story about bad project managers. It is a story about a systematic human failure to forecast accurately when personal involvement, optimism, and complexity collide.
Flyvbjerg identifies 2 causes operating simultaneously. The first is unconscious optimism bias: teams genuinely believe their project will be the exception. The second is what he calls the Machiavelli Factor: deliberate underestimation in order to get projects approved. Both forces push in the same direction. The result, Flyvbjerg concludes, is inverted Darwinism: not the best projects get built, but the projects that look best on paper.
- 91.5% of IT projects go over budget, over schedule, or both (Flyvbjerg, Oxford)
- 45% average cost overrun for large IT projects (McKinsey / University of Oxford)
- 56% less value delivered than projected, on average (McKinsey)
- 447% average cost overrun for IT projects with overruns above 50% (Flyvbjerg)
- 31% of software projects succeed on time, on budget, with full scope (Standish CHAOS 2020)
CHAPTER # 2 - "The Real Cost Model"
When a company decides to build software, the conversation almost always centers on the initial development budget. This is almost always the wrong number to examine. The right number is the Total Cost of Ownership over 5 to 10 years. And that number, when calculated honestly, is almost always a shock.
The reason the shock comes is not that companies are careless with money. It is that the costs of building software are structured in a way that systematically hides the largest expenses until it is too late to reverse the decision. The visible costs appear upfront. The larger costs accumulate silently, year after year, in maintenance budgets, talent costs, opportunity costs, and the compounding weight of technical debt.
What People Count: Direct Costs
Direct costs are the numbers that appear in the business case. They include initial development, testing, infrastructure setup, integration work, security reviews, and deployment. These are the numbers most likely to be understated, because they are estimated at the beginning of the project, before scope creep, technical complexity, and the planning fallacy have done their work.
The Standish Group’s CHAOS report, tracking software projects since 1994, found that the average cost overrun is 189% of the original estimate. Not 89% over. The final cost is nearly 2x the budgeted amount. For large company projects, the failure rate is even more severe: only 9% of projects at companies with revenues above $500M are delivered successfully. Every project plan in that boardroom was wrong. Most of them were wrong by a lot.
A Harvard Business Review analysis found that 1 in 6 large IT projects carries a cost overrun above 200%. Some exceed 700%. These are not corner cases or incompetent organizations. They are normal organizations making normal decisions with normal cognitive equipment, under the influence of normal biases. The planning fallacy, the IKEA Effect, and the endowment effect are not bugs in human cognition. They are features. They exist because they served other purposes. They are just catastrophically expensive in the context of large software projects.
What Nobody Counts: Indirect Costs
The indirect costs are where the real damage accumulates, because they are harder to quantify and easier to ignore at decision time.
Opportunity cost of delay. Custom software typically takes 18 to 36 months to break even relative to a bought solution. In that window, competitors who chose to buy are already capturing value. A company needing analytics capabilities worth $20,000 per week in efficiency gains that builds over 20 weeks versus buys over 4 weeks has paid $320,000 in delayed value before a single line of code is compared on its merits. That $320,000 never appears on any project dashboard. It is invisible. But it is real.
The productivity J-curve. Any new internal system creates a temporary productivity decline before gains materialize. Users resist change, workflows must be adapted, training takes time, and bugs must be worked through in production. In large organizations, this productivity dip can represent months of reduced output. It is almost never included in the business case.
Security liability. Custom software means 100% internal security responsibility. Regular penetration testing, vulnerability assessments, compliance certifications, and incident response planning are ongoing costs without end. Legacy custom systems are disproportionately represented in major breaches: the 2017 Equifax breach, partly caused by unpatched internal code that had been running for years, cost the company approximately $1.3B in settlements, remediation, and reputational damage. That is the cost of a single security failure in a single legacy system.
Scalability creep. As usage grows, custom applications need architectural changes that were not anticipated at design time. These changes require deep knowledge of the original codebase, which may or may not still exist in the organization. Every growth event becomes a potential engineering crisis.
Integration maintenance. Custom systems must maintain connections to other systems as those systems change. This integration maintenance has no natural endpoint. Every upstream change - a new version of an external API, a regulatory update, a third-party data format change - creates a cascade of internal rework that a commercial vendor would have absorbed on behalf of all its customers.
The Maintenance Iceberg
The most consistently underestimated cost in any build decision is ongoing maintenance. Research shows that 78% of lifetime software Total Cost of Ownership occurs after launch. Maintenance runs 15 to 20% of original development cost every year. For a $5M custom build, that is between $750,000 and $1M per year, every year, indefinitely. After 5 years, the maintenance cost alone has exceeded the original build cost. Most build decisions do not include this calculation.
78% of lifetime software cost occurs after launch. The day you ship is not the end of the expense. It is the beginning of a much larger one.
Gartner estimates that technical debt now represents 20 to 40% of the total value of technology estates across enterprise organizations. McKinsey found that companies with fragmented legacy systems were 30% more likely to experience delays implementing AI and new technologies. Nearly two-thirds of businesses spend more than $2M annually just maintaining and upgrading legacy systems - money that produces nothing new, creates no competitive advantage, and captures no new value. It simply keeps the lights on.
The pattern is predictable enough to name. Year 1: the system works as designed, maintenance is manageable, the team is proud of what they built. Year 2: the first integration issues emerge, workarounds appear, and the backlog of enhancement requests begins to grow faster than the team can address it. Year 3: a key developer who knows the system deeply leaves for another job. Their replacement take 3 months just understanding what exists. Year 4: a regulatory change requires significant rework. The team that built the original system is largely gone. Year 5: a technology leader arrives with a mandate to modernize. They discover the system everyone thought was 5 years old is effectively legacy.
Engineers who have worked in legacy-heavy environments consistently report that delivery velocity on aging internal systems is 4x slower than on modern greenfield architecture. A feature that takes 2 weeks to build on a clean, well-documented system can take 8 weeks when engineers must navigate decade-old dependencies, undocumented workarounds, and deployment procedures designed to prevent cascading failures in fragile infrastructure. The 4x multiplier is what engineers report across organizations, industries, and continents.
- 78% of lifetime software TCO occurs after launch (research consensus)
- 70-80% of IT budgets consumed by maintenance in legacy-heavy organizations
- $2.41T annual US cost of technical debt and operational failures (AEI / CISQ)
- 4x slower delivery velocity on legacy systems versus modern architecture
- $2M+ spent annually on legacy maintenance by two-thirds of enterprises (CIO Dive)
CHAPTER # 3 - "The Uncomfortable Truth"
Here is what the polite version of this conversation always omits.
A significant portion of build decisions, in my experience, a majority of them - have less to do with the business case and more to do with the careers of the people making the recommendation. The logic of building serves individual interests in ways that buying never can. And because those interests are structural, not personal, they operate even in organizations full of smart, well-intentioned people.
I want to be clear about what I am saying and what I am not saying. I am not claiming that the engineers and architects who advocate for building are dishonest. Most of them genuinely believe what they are saying. The biases are real; the overconfidence is sincere; the uniqueness conviction is felt, not fabricated. But the structural incentives around the decision systematically reward one answer over the other. And that matters.
IT Empire Building
In 1971, economist William Niskanen published his theory of bureaucratic behavior, arguing that government agencies systematically expand their budgets and headcount not because expansion serves the public interest, but because larger organizations give their leaders more power, higher salaries, greater job security, and more influence. The mechanism is not corruption. It is rational self-interest operating within a structure that does not adequately constrain it.
This dynamic operates identically in corporate IT departments. Building custom software requires hiring developers, architects, project managers, QA engineers, DevOps specialists, and security professionals. Buying software requires a procurement process and a few administrators. The headcount math creates a gravitational pull toward building that has nothing to do with business value. A CTO who recommends buying a $2M SaaS solution runs a procurement process. A CTO who recommends building has a $5M budget, a team of 20 engineers, and a multi-year program that justifies their organizational footprint and elevates their status in the organization.
I have never met a CTO who framed their recommendation in these terms. I have met many CTOs whose recommendations, examined carefully, followed exactly this pattern. The bias does not announce itself. It disguises itself as technical judgment. And because technical judgment is genuinely difficult to evaluate from the outside, it is rarely challenged.
Résumé-Driven Development: The Career That Needs a Build
In 2021, researchers Fritzsch, Wyrich, Bogner, and Wagner published the first peer-reviewed study of what they named Résumé-Driven Development: the practice of making technology choices to enhance career prospects rather than to serve business needs. Their survey of 591 software professionals found that 82% believed using trending technologies would make them more attractive to future employers, and 60% of hiring managers confirmed that technology trends directly influence their job advertisements.
The implications for build decisions are direct. A developer who spends 2 years maintaining a smoothly running commercial SaaS implementation has, from a career perspective, almost nothing to show for it. A developer who spends 2 years building a custom AI-powered automation system on a modern architecture stack has an impressive line on their resume, conference talk material, and the kind of deep technical experience that commands a premium salary. Both developers may have served their organizations equally well. Only one of them has a compelling career story.
This is not a cynical observation about developer ethics. It is an observation about incentive structures. The job market rewards builders. Organizations that ask their engineers to procure wisely rather than build ambitiously are asking those engineers to accept a career penalty. Without structural countermeasures, the incentive to build will persistently win.
I have never heard a good articulated justification for building, beyond: we want to do it ourselves. No logical explanation. In many cases, what I was actually hearing was a career plan dressed up as a technology strategy.
The Principal-Agent Trap
Jensen and Meckling’s foundational 1976 paper on the principal-agent problem formalized a structural reality that most organizations experience but rarely name: when the person making a decision (the agent) has different interests from the person bearing the consequences (the principal), decisions systematically diverge from what is optimal for the organization. The gap between agent interests and principal interests is the source of enormous waste in corporate life.
In build-vs-buy decisions, the misalignment is particularly severe. The CTO or VP of Engineering who recommends building benefits directly from a larger team, a more complex project, and a more prominent role in the organization. The CEO and board who bear the financial consequences of that recommendation often lack the technical knowledge to evaluate it independently. Peter Weill and Jeanne Ross’s MIT research across 250 enterprises found that on average only 1 in 3 senior managers knew how IT was governed at their company. The people paying for the decision cannot adequately evaluate it. The people making the recommendation benefit from one outcome. There is no natural buy constituency in most IT organizations - no career incentive for wise procurement, no conference talk about the vendor you chose wisely, no promotion for the administrator who runs a commercial platform efficiently.
Who Bears the Consequences?
Perhaps the most revealing question to ask about any build decision is this: what happens to the people who made it when things go wrong?
In the documented history of large build failures, the answer is: not much, and not for long. The FBI cycled through 5 CIOs in 4 years during the Virtual Case File disaster. Each departure was painful. But the institutional momentum that produced the decision in the first place - the desire to build, the empire-building incentives, the career calculus - survived largely intact. The NHS IT programme destroyed the careers of several executives. The programme itself continued for 9 years and cost £12.7B before anyone finally stopped it.
Barry Staw’s research on escalation of commitment showed precisely why this happens. People who are personally responsible for a failing course of action increase their commitment to it rather than cutting their losses, because abandoning the project means admitting that the original decision was wrong. The psychological cost of that admission is experienced as larger than the financial cost of continued investment. Staw called this being “knee-deep in the big muddy”: the further in you are, the harder it is to turn around, because turning around means acknowledging that you walked into the mud in the first place.
Mark Keil’s IT-specific research found that 81% of IS audit professionals had witnessed escalation in recent projects. In one documented case, a system foundered for more than 10 years before anyone could stop it. Deescalation was almost never initiated by the project team itself. It required external intervention - a new CIO, an auditor, a financial crisis, a dramatic public failure. The people inside the project were constitutionally unable to pull the plug. Not because they lacked intelligence. Because the psychology of commitment had fused with the psychology of identity. Stopping the project felt like destroying a part of themselves.
CHAPTER # 4 - "The Psychology of “We Need Control”
Strip away every technical argument in any build debate, and you almost always arrive at the same place: a felt need for control. The language varies: flexibility, customization, ownership, strategic independence - but the emotional core is the same. We need to control this ourselves.
This instinct is not irrational in the abstract. Control over critical systems is a legitimate concern. The problem is that the psychology of control systematically detaches itself from the evidence about whether control actually produces better outcomes. The feeling of control becomes more important than the reality of it. And organizations pay an enormous premium for a feeling.
The Illusion of Control
In 1975, Ellen Langer at Yale conducted a series of experiments that established one of the most important findings in behavioral psychology: people develop a strong expectancy of personal success probability that is inappropriately higher than the objective probability would warrant, whenever the situation contains skill cues like personal choice, involvement, or competition - even when those cues are entirely irrelevant to the outcome.
In one experiment, participants playing a lottery were significantly less willing to trade their self-chosen lottery tickets for tickets with objectively better odds. In another, participants facing a task involving pure chance behaved as if skill were involved if they were allowed to practice first, were permitted to choose their materials, or were working against a visibly incompetent competitor. The feeling of control persisted even when control was objectively absent.
Software development is saturated with skill cues. Developers make architectural choices, select frameworks, write code, approve designs, decide on testing strategies. The entire activity is structured around individual and team agency. This creates a powerful, persistent, and largely false sense that project outcomes are within human control. The evidence from Flyvbjerg, from the Standish Group, from McKinsey - is that large software project outcomes are closer to a lottery than a controlled process. The feeling of control is vivid. The control itself is largely illusory.
The Dunning-Kruger effect adds another layer. Dunning and Kruger’s 1999 Cornell study found that people who performed in the bottom quartile of a test rated their skills far above average - those in the 12th percentile self-rated their expertise at the 62nd percentile on average. The effect operates because incompetence in a domain also impairs the ability to recognize incompetence. Applied to software teams evaluating a build decision: the teams with the least relevant experience building enterprise-scale software are precisely the teams most likely to be overconfident about their ability to deliver it. The teams best positioned to know how hard it actually is are, by the nature of the Dunning-Kruger effect, the most likely to express caution. The caution gets interpreted as pessimism. The overconfidence gets interpreted as drive.
The Control Premium
Owens, Grossman, and Fackler quantified the cost of the control bias in a 2014 paper published in the American Economic Journal. Their experiments showed that participants sacrificed 8 to 15% of expected earnings in order to retain decision-making autonomy, even in situations where delegating would have produced demonstrably better outcomes. Critically, only 66% of this premium was explained by overconfident beliefs about their own ability. The remaining 34% was a pure preference for autonomy, independent of any belief that autonomy would produce a better result. People paid to be in control even when they knew control was not helping them.
This is the control premium at the individual level. At the organizational level, the same premium plays out at vastly higher scale. When companies pay 2 to 5 times more to build custom software rather than buy, a substantial portion of that premium is not purchasing better outcomes - it is purchasing the feeling of being in control. That feeling is not worthless. But it is dramatically overpriced.
Loss Aversion and the Vendor Lock-in Illusion
The most powerful emotional argument against buying is not positive - it is negative. It is not “building gives us control” so much as “buying creates dependency.” The fear of vendor lock-in is real, widely felt, and systematically overestimated.
Kahneman and Tversky’s prospect theory, which earned Kahneman the Nobel Prize in Economics in 2002, established that losses are psychologically weighted approximately 2x as heavily as equivalent gains. When teams evaluate buying software, they mentally frame vendor dependency as a loss of autonomy. Prospect theory predicts this perceived loss will feel roughly 2x as large as the gain of faster deployment, lower cost, and proven functionality. The comparison is psychologically unfair before it begins.
What the loss aversion framing obscures is the actual lock-in structure of the 2 options. Building custom software creates its own, far more durable form of lock-in. You are locked to the architectural decisions made years ago by engineers who may have long since left. You are locked to the maintenance burden of code that accumulates debt whether or not you tend to it. You are locked to the talent market, dependent on finding engineers who understand your specific codebase and want to work on it. And you are locked to the compounding costs of keeping that system current, secure, and functional as the technology landscape around it changes.
Switching vendors is a one-time cost, typically manageable with planning. The lock-in from your own legacy code is perpetual, structural, and gets worse with every year. The fear of vendor dependency is a fear of a cost that is real but bounded. The reality of internal legacy dependency is a cost that is also real but unbounded. When organizations choose building to “avoid lock-in,” they are choosing the more expensive form of dependency while feeling virtuous about avoiding the cheaper one.
The Uniqueness Illusion: Your Business Is Not That Special
The final refuge of the build argument, when financial and technical objections have been addressed, is uniqueness. “Our processes are too specific for any off-the-shelf solution. Our business is genuinely different.”
Bent Flyvbjerg’s 2024 research on uniqueness bias studied 219 IT projects where planners rated their project’s uniqueness on a scale of 1 to 10. Among the 23 projects scoring 9 or 10 on uniqueness - projects their sponsors described as genuinely unprecedented - the researchers found precedence for similar projects in the same organization or industry for every single one. Not most. All of them. 5 were regulatory compliance initiatives at banks where every competitor was implementing the same regulation simultaneously.
Flyvbjerg quotes a veteran megaproject manager who has now spent 4 decades in the field: “The first 20 years, I saw uniqueness in each project; the next 20 years, similarities.” The experienced perspective is almost always the same: what feels unique from the inside looks like a variation on a familiar pattern from the outside.
The mechanism is a psychological phenomenon researchers call uniqueness bias: the tendency to overestimate how different our situation is from comparable situations, and to underestimate how much we can learn from others’ experience. Uniqueness bias is partly motivated if we are truly unique, we cannot be criticized for outcomes that differ from the norm and partly cognitive. We have more access to information about our own situation than about others’ situations, so our own situation naturally seems richer and more complex.
Your business is not as unique as you think. Your processes are not as differentiated as you believe. And the software market has already solved most of your ‘unique’ problems for someone else, at higher quality and lower cost than you can replicate internally.
Not Invented Here: The Tribal Defense of Internal Work
Ralph Katz and Thomas Allen’s landmark 1982 MIT study of 50 R&D project groups established the Not Invented Here syndrome with a finding that has held up through 4 decades of subsequent research: teams reject outside knowledge not because it is inferior, but because accepting it threatens group identity and cohesion. The external solution represents an implicit challenge: someone else solved your problem, and they may have solved it better.
More troublingly, Katz and Allen found that the highest-performing teams were the most susceptible to NIH syndrome. High performance breeds confidence. Confidence breeds a willingness to dismiss external perspectives. Success makes teams more likely to believe they are uniquely capable of solving their own problems, even when the evidence does not support that belief.
In practice, NIH syndrome manifests as a detailed, technical-sounding critique of every commercial alternative considered. The vendor’s API is not flexible enough. The data model does not match internal conventions. The performance benchmarks do not apply to the company’s scale. The UI is wrong for the user base. Each objection may contain a grain of truth. Collectively, they function as a rationalization for a conclusion that was reached before the evaluation began.
CHAPTER # 5 - "How Good Teams Make Bad Decisions Together"
The previous chapter examined the psychology of individual decision-makers. But build decisions are almost never made by individuals. They emerge from groups: steering committees, architecture reviews, leadership discussions, project planning sessions. And groups have their own pathologies that compound individual biases rather than correcting them.
The assumption that group deliberation improves individual decision-making is intuitive but often wrong. Under the right conditions, groups systematically amplify the biases of their members, suppress dissenting voices, and converge on confidence they have not earned. In technology decisions especially, the social dynamics of the room can matter more than the quality of the analysis.
Groupthink in the Architecture Review
Irving Janis first described groupthink in 1972, studying why groups of intelligent, informed people make catastrophically bad decisions together. He identified 8 symptoms, and technology committees exhibit all of them with textbook regularity.
The illusion of invulnerability: the team believes its collective capability will allow it to overcome any technical challenge, producing excessive optimism about the build timeline and scope. Collective rationalization: the team develops shared narratives that dismiss concerns raised by outsiders or skeptics - the vendor does not understand the business, the commercial product is not enterprise-grade, the reference customer’s implementation was simpler. Stereotyping of outsiders: vendors become “just sales people,” consultants become “people who have never built anything,” and the skeptical executive becomes someone who “does not understand the technical realities.”
Self-censorship is perhaps the most damaging symptom. In any room where the build decision has momentum where the CTO has expressed enthusiasm, where the engineering team has started designing, where a budget has been tentatively approved the person who privately thinks “maybe we should just buy this” will almost never say so. The social cost of raising that voice feels higher than the benefit of being heard. The silence that follows is interpreted as consensus. The consensus becomes commitment. The commitment becomes a program. The program becomes a legacy.
NASA engineer Bob Ebeling, in the days before the Challenger disaster, expressed this dynamic with painful clarity: “I had the distinct feeling that we were in the position of having to prove it was unsafe, rather than being asked to prove it was safe.” The burden of proof had been reversed by group momentum. The same reversal happens in technology decisions: the burden of proof falls on those questioning the build, never on those proposing it.
Identity Fusion and the Project That Becomes Personal
Henri Tajfel and John Turner’s social identity theory, developed in the 1970s and 1980s, explains why criticizing a build project feels so much like a personal attack on the people in the room. When a group assembles around a shared project, their professional identities fuse with it. The build becomes “us.” The alternative of buying from a vendor becomes the out-group, representing not just a different technology choice but an implicit judgment that outsiders can do what the team cannot.
Research on identity threat shows that information challenging a group’s value triggers a predictable defensive sequence: derogation of the source, doubling down on group identity, active avoidance of the threatening information, and in-group solidarity against the perceived attack. When a consultant presents a vendor comparison showing that commercial solutions outperform the internal build on cost and functionality, this is experienced not as useful data but as an attack on the people who proposed the build. The more senior and respected those people are, the stronger the defensive response.
This is why build reviews, once started, rarely reverse themselves on technical grounds. The social and identity dynamics of the room make reversal feel like betrayal. Only an external shock a budget crisis, a new leadership team, a catastrophic failure is usually sufficient to overcome the group’s investment in the path it has already chosen.
Escalation of Commitment: The Deeper You Go, The Harder It Is to Stop
Barry Staw’s landmark 1976 study, “Knee Deep in the Big Muddy,” examined what happens when people become personally responsible for a failing course of action. His finding was counterintuitive and consistent: rather than cutting losses, people increase commitment to the failing choice. The psychological mechanism is self-justification: to admit the project should be stopped is to admit that the original decision was wrong. For most people, the shame of that admission feels larger than the financial benefit of stopping.
Staw found that personal responsibility amplified escalation dramatically. People who had made the original decision committed far more to a failing course of action than people who had inherited the decision. In software projects, this means that the architects and leaders who proposed the build are precisely the people least able to call for its termination. The project’s champions are constitutionally compromised.
Mark Keil’s decade of research on IT project escalation at Georgia State University quantified the phenomenon in the specific context of software. He found that 30 to 40% of all information systems projects exhibit some degree of escalation. In a 1997 survey, 81% of IS audit professionals reported witnessing escalation in recent projects. The completion effect was particularly powerful: projects that appeared close to completion were the hardest to stop, because the team could see the finish line and believed one more push would get them there. Given the Ninety-Ninety Rule of software development — projects feel almost done for approximately half their total duration - the completion illusion can persist for years.
In one documented case in Keil’s research, a failing internal system was kept running for more than 10 years before deescalation finally occurred. The cost to the organization over that decade was multiples of what termination would have cost in year 2. Keil’s most consistent finding: deescalation of failing projects was almost never initiated by the people running them. It required external intervention. The project team was, by definition, unable to make the rational decision. They were too far in.
CHAPTER # 6 - "The AI Paradox"
In November 2022, OpenAI released ChatGPT to the public. Within 2 months, it had accumulated 100M users - the fastest adoption of any consumer technology in history. Within 6 months, every enterprise boardroom in the world had an AI agenda item. Within a year, the question was no longer whether to invest in AI but how much and how fast.
And with that question came a new and seductive argument for the wrong answer in the build-vs-buy debate.
The argument goes like this. Generative AI has dramatically lowered the barrier to building software. AI coding assistants can generate in hours what used to take days. LLM APIs allow any developer to add intelligent processing to any application. The marginal cost of a working prototype has fallen toward zero. Therefore, the case for buying rather than building has weakened, because the argument that building is too expensive or too slow no longer holds the way it once did.
This argument is logically coherent at the level of a single prototype. It collapses at the level of an enterprise-grade production system. And it is driving a wave of expensive, poorly governed internal AI projects whose failure modes are entirely predictable to anyone who has been watching the broader build-vs-buy story for the last 30 years.
The Numbers Behind the AI Build Wave
The scale of enterprise AI investment is genuinely unprecedented. Enterprise generative AI spending grew from $1.7B in 2023 to $37B in 2025. Enterprise AI adoption reached 78% in 2024, up from 55% in 2023. The financial commitment signals a genuine strategic shift: AI has moved from innovation budgets to core operational spending at most large companies.
And then the paradox: McKinsey found that while 78% of enterprises have deployed generative AI in at least one function, 80% report no significant improvement in productivity, cost, or revenue. An MIT study found that 95% of enterprise generative AI initiatives fail to reach P&L impact. The organizations spending tens of billions of dollars on AI are, in 4 cases out of 5, seeing essentially nothing on the bottom line.
The pattern should be familiar. New technology creates excitement. Excitement creates projects. Projects produce demos. Demos look like progress. Demos are mistaken for production systems. Production systems fail to materialize, or they materialize without the governance, data quality, integration, and operational infrastructure needed to function at scale. The planning fallacy, the IKEA Effect, and the Dunning-Kruger effect do not care whether the technology involved is enterprise resource planning software from 2005 or an agentic AI claims automation system from 2025. The cognitive equipment is the same. The outcomes follow the same distribution.
- $37B enterprise generative AI spend in 2025, up 22x from 2023 (Menlo Ventures)
- 78% of enterprises deploying GenAI in at least one function (McKinsey 2025)
- 80% of those enterprises reporting no meaningful bottom-line impact
- 95% of enterprise GenAI initiatives failing to reach P&L impact (MIT 2025)
Why AI Makes the Build Bias Worse, Not Better
AI coding tools are genuinely transformative for software development. More and more CTOs report that nearly 90% of code is now AI-generated - up from 10 to 15% in the last 4 years. The speed at which working prototypes can be created has changed fundamentally. This is real.
But it has a dangerous side effect on the build-vs-buy calculation that is almost never acknowledged. When anyone can generate a working prototype in an afternoon, the natural organizational response is: “Why wouldn’t we just build it?” The IKEA Effect applies with even greater force when the labor required to create the thing has been dramatically reduced. If assembling an IKEA cabinet makes you value it 63% more, what does having an AI generate your codebase in 48 hours do to your attachment to the result?
The answer, empirically, is that it makes the attachment worse, not better. The speed of prototyping creates an illusion that the hard part is done. It is not. Building a working demo is the cheapest part of the journey. The expensive parts are making it enterprise-grade: secure, observable, maintainable, integrated with production systems, compliant with regulatory requirements, documented for the engineers who will maintain it, and capable of handling the edge cases that only appear in production. AI tools accelerate the first day. They do not accelerate the next five years.
Technology Magazine’s 2026 analysis of the build-vs-buy landscape put it plainly: “Just because you can build it faster does not mean you should build it. The maintenance burden of custom code remains entirely unchanged. Without strict governance, the result is shadow IT at industrial scale: a patchwork of unmaintainable applications, each one someone’s weekend project that is now a business-critical dependency.”
The Dunning-Kruger Effect in the Age of AI
There is a specific and particularly ironic version of the Dunning-Kruger effect operating in enterprise AI decisions today. A 2025 study published in Computers in Human Behavior found that AI tools improve performance on logical reasoning tasks but lead to highly biased self-assessments. Users consistently overestimated how well they had done. More troublingly, participants with higher AI literacy were less accurate in their self-assessments than those with lower AI literacy. The more someone knows about AI, the more they overestimate what they can do with it.
Apply this to the teams evaluating whether to build custom AI systems. The engineers most capable of evaluating the technical feasibility of a build - the ones who understand LLM APIs, RAG architectures, agent orchestration frameworks are precisely the engineers most susceptible to overestimating what they can deliver. Their technical sophistication is real. Their operational experience deploying AI at enterprise scale is, in most cases, not.
Agentic AI: The Stakes Get Higher Again
If generative AI represented one escalation of the build temptation, agentic AI, systems that do not merely generate content but autonomously take actions across connected production systems, represents another order of magnitude. By 2025, 58% of companies had integrated some form of AI agent into their operations. Gartner predicts that by 2028, 33% of all enterprise software applications will include agentic AI capabilities.
The build temptation is acute with agentic systems because they touch proprietary workflows, proprietary data, and proprietary decision logic in ways that seem genuinely unique to each organization. And the stakes of getting the build decision wrong are categorically higher. With generative AI, a poor output is a bad draft that a human reads and does not send. With agentic AI, a poor decision is an action: a wrong transaction processed, a wrong communication sent, a wrong record updated in a production database. The errors scale automatically. The damage is real before anyone notices.
Gartner’s 2025 warning was unambiguous: over 40% of agentic AI projects will be cancelled by end of 2027 due to escalating costs, unclear business value, or inadequate risk controls. Most current projects, Gartner found, are hype-driven proof-of-concepts that stall before reaching production, precisely because organizations built demonstrations without the data governance, audit infrastructure, and operational discipline that production agentic systems require.
Research from Dimensional Research found that 88% of companies building in-house AI solutions needed 6 months or longer to get a single solution operating. Writer’s analysis of enterprise AI infrastructure costs found that a mid-sized enterprise processing 200,000 queries per month against a reasonable knowledge base could face RAG infrastructure costs alone exceeding $190,000 per month. And a RAG system that can answer questions is not an agent. It is a knowledgeable system. The agentic layer that allows it to act — with orchestration, security guardrails, human-in-the-loop checkpoints, audit trails, and rollback capabilities — is a separate and substantially more complex engineering challenge.
Over 40% of agentic AI projects will be cancelled by end of 2027. Most are hype-driven proof-of-concepts that stall before production — because organizations built demonstrations without the infrastructure production requires.
The One Way AI Genuinely Changes the Calculus
There is one dimension in which the rise of AI genuinely changes the build-vs-buy argument, and it makes the case for buying stronger, not weaker.
A focused AI vendor deploying solutions across hundreds or thousands of organizations learns from all of them simultaneously. Every edge case becomes training data. Every integration challenge informs the next deployment. Every regulatory change is absorbed and distributed to all clients. An internal team building an AI system from scratch accumulates experience from exactly one client: themselves. The knowledge asymmetry between a specialized vendor and an internal team has never been wider than it is today, precisely because the underlying models are advancing faster than any single organization can track.
In 2023, there was serious discussion in enterprise circles about building custom foundation models. BloombergGPT was the flagship example. By 2025, that conversation had essentially ended. Even the most technically sophisticated enterprises had largely abandoned custom model development and moved to commercial foundation model APIs. The reason was not ideology. It was economics. Foundation model development requires billions of dollars of compute, teams of hundreds of researchers, and access to training data at scales no individual enterprise can assemble. Buying access to the best foundation models costs a fraction of building a mediocre one.
The same logic applies, with proportionally smaller numbers, to every layer of the AI stack below foundation models: orchestration frameworks, RAG infrastructure, agent tooling, evaluation systems. These are commodity problems that hundreds of specialized vendors are solving at scale. Building them internally is the Redundancy Tax: paying to rediscover, at full cost, solutions that already exist and are improving every quarter.
CHAPTER # 7 - "Insurance: A Mirror"
The insurance industry is worth examining specifically, because it is an industry where the conditions for chronic build bias are structurally embedded, where the consequences of legacy accumulation have been playing out for decades, and where the AI build-vs-buy decision is being made right now, at scale, by every major carrier.
It is also an industry I know well. And what I see in it is a mirror of every dynamic described in this article, concentrated and made visible by the specific characteristics of insurance operations.
Why Insurance Feels Unique (And Why It Is Less Unique Than It Thinks)
Insurance has genuine technical complexity. Regulatory requirements vary by jurisdiction, line of business, and product type in ways that are genuinely intricate. Claims workflows involve document processing, liability assessment, coverage interpretation, fraud detection, and regulatory reporting in sequences that are non-trivial to automate. Underwriting requires actuarial models, external data integration, and risk classification logic that can involve hundreds of variables across dozens of data sources.
This complexity creates a legitimate-sounding justification for building: “Our processes are too specific for any off-the-shelf solution.” And because the processes are genuinely complex, the justification is harder to challenge from the outside. The technical vocabulary of insurance - FNOL, LAE, combined ratio, cession, bordereaux - creates a barrier that reinforces the uniqueness illusion. If you do not know what a bordereau is, how can you evaluate whether a vendor handles it correctly?
But complexity is not the same as uniqueness. A P&C insurer’s claims workflow - first notice of loss, assignment, investigation, coverage determination, settlement, closure - follows the same structural logic across virtually every carrier in every market. The parameters differ. The edge cases differ. The regulatory context differs by jurisdiction. But the structure is the same. Every dollar spent building that structure from scratch is a dollar spent rediscovering something that has already been built, refined, stress-tested, and regulatory-approved across dozens of deployments by vendors who do nothing else.
The uniqueness illusion in insurance is also self-perpetuating. Because carriers historically built custom systems, their operations adapted to those systems. The processes became customized around the technology rather than the other way around. Now the processes really are unique not because they represent best practices, but because they represent the accumulated workarounds of a system that was built internally and never forced to conform to external standards. The uniqueness is a scar, not a feature.
The Legacy Debt Is Deepest Here
Insurance technology spending reached $185B in 2024 and is forecast to grow to $420B by 2033. A significant and largely unacknowledged portion of that spending is not innovation. It is maintenance. It is the compound interest payment on decades of build decisions.
Forrester found that legacy systems, lack of AI talent, and difficulty integrating AI into existing processes are the three primary barriers to AI value realization in insurance. Fewer than 5% of insurers are expected to see direct, tangible AI gains in 2025. IBM research found that more than 40% of insurers do not have adequate internal skills and expertise to modernize. McKinsey notes that companies with fragmented legacy systems are 30% more likely to experience AI implementation delays.
The trap is self-reinforcing in a way that is almost diabolical. Past build decisions created legacy debt. Legacy debt now consumes the engineering capacity that would be needed to modernize. The modernization can’t happen because the people who would do it are maintaining the legacy. The legacy gets older. The debt compounds. And meanwhile, every year, the board approves another set of “transformation initiatives” that stall at the legacy integration layer.
The carriers that are moving fastest on AI automation today are, almost without exception, the ones that bought modern platforms rather than building them. They have clean data architectures. They have vendor-maintained infrastructure. They have engineering capacity available for genuine differentiation because they are not spending that capacity on maintenance. The competitive gap between these carriers and their legacy-laden competitors is already visible in combined ratios, in claims cycle times, in customer satisfaction scores. It will widen.
AI in Insurance: The Decision Is Being Made Right Now
Insurance IT spending on AI has reached 36% of total IT budgets, the highest priority among insurers. 78% of insurance leaders are expanding technology budgets in 2025. The AI build-vs-buy decision is not hypothetical in this industry. It is happening now, in real time, at every major carrier.
Some are building. Internal teams are constructing AI models for claims automation, underwriting augmentation, fraud detection, and customer service. Some of those efforts will succeed. The ones that succeed will almost certainly be those focused on the narrow layer where truly proprietary data and domain expertise create genuine differentiation - not on the infrastructure layer that every carrier needs and no carrier has a unique reason to build.
The carriers buying specialized vertical AI solutions are seeing value in months, not years. Implementations that automate end-to-end claims workflows in under 90 days are not theoretical. They are happening. The speed differential between carriers who buy purpose-built solutions and those who build internal AI stacks will compound over the next three years in ways that will be very visible in loss ratios, expense ratios, and market share.
Forrester’s 2025 prediction for the insurance industry is worth quoting in full: “As insurers prioritize agility and quicker time to value, they will cut back on new multiyear and complex transformation programs. More iterative peripheral development will grow building APIs and microservices and hollowing out complex business rules by constructing an abstraction layer around legacy backends.” Translation: the era of the grand internal build is ending, even in insurance. The question is how many more years and how many more billions will be spent on legacy maintenance before the industry fully accepts what the data has been saying for decades.
CHAPTER # 8 - " The Projects That Ended Arguments"
Everything discussed so far is theory, data, and structural analysis. But the most clarifying evidence is historical. The projects that played out at scale, in public, with documented costs and documented consequences, show the same pattern so consistently that it becomes difficult to attribute to coincidence. They are not cautionary tales about exceptional circumstances. They are representative examples of what happens when normal organizations make build decisions with normal psychology.
The FBI Virtual Case File: $170 Million, Total Loss
In the aftermath of September 11, 2001, the FBI faced a legitimate and urgent problem. Its case management system was paper-based in an era that demanded digital information sharing. The Virtual Case File project was initiated to solve this problem with a modern, integrated system. The budget was $170M. The timeline was 3 years.
The project was led internally, with the FBI’s own IT organization driving architectural decisions. The scope expanded continuously. Requirements changed as the organization changed. By 2005, the project had consumed its budget and produced 700,000 lines of code that FBI agents who tested it described as so defect-ridden it was unusable. The Department of Justice Inspector General’s report identified the cause with uncomfortable specificity: the FBI had tried to build too much, too fast, with inadequate project management, continuously changing requirements, and no effective mechanism for stopping a failing project.
The system was not restructured. It was not salvaged. It was abandoned entirely. Every dollar spent, all $170M, was written off as a total loss. The FBI then initiated a replacement project called Sentinel, which eventually succeeded by doing the opposite of almost everything the Virtual Case File project had done: it adopted commercial off-the-shelf components wherever possible, used agile development methods, and reduced the internal team from 400 to 45 people. Sentinel delivered. The custom build did not.
The FBI cycled through 5 CIOs in 4 years during the project’s life. Each departure was publicly painful. The institutional pressure to continue the project persisted through all of them, because the sunk cost was real, the organizational identity was invested, and stopping felt like admitting failure at a moment when the FBI could not afford to admit failure. Keil’s research on escalation of commitment plays out in perfect sequence in this case: personal responsibility, identity investment, completion illusion, inability to deescalate from within.
NHS NPfIT: £12.7B, 9 Years, Abandoned
In 2002, Prime Minister Tony Blair announced the National Programme for IT: an initiative to create integrated electronic patient records across the English NHS, connecting hospitals, GPs, and specialists in a unified digital health record accessible wherever a patient received care. The vision was genuinely valuable. The execution was a masterclass in every failure mode this article has described.
The original cost estimate was £2.3B. Over 9 years, the programme consumed £12.7B. The National Audit Office’s assessment was that it “did not deliver key benefits.” Parliament’s Public Accounts Committee called it one of the worst and most expensive contracting fiascos in the history of the UK public sector. Only 22 of the planned NHS trusts adopted the system. Major vendors withdrew. Directors resigned. The system was abandoned in 2011.
The post-mortem analysis identified causes that map directly to the psychology described in this article: a top-down decision driven by political ambition rather than operational need, inadequate engagement with the end users who would actually use the system, a rigidly centralized model that could not accommodate the genuine diversity of NHS operations, and, crucially, an organizational culture that made it impossible to stop the programme even as evidence of failure mounted year after year. The programme had become an institutional identity. Killing it felt like admitting that the entire digital health agenda had been wrong.
The Computer Weekly analysis of the programme’s failure identified a pattern that Keil’s research would predict precisely: the project was a top-down decision made for political reasons, and top-down projects inspired by political or vanity motivations are, in the historical record, disproportionately likely to fail. The commitment to the decision, rather than to the outcome, drove the programme long past the point where rational analysis would have stopped it.
Southwest Airlines: $800M, One Holiday Season
The Southwest Airlines case has been described in the previous chapter, but it deserves fuller treatment here as a cross-industry proof point. Southwest was not a struggling airline. It was, for most of its history, one of the best-run airlines in the world, profitable for 47 consecutive years, beloved by customers and employees alike, a case study in operational excellence. The catastrophic failure of December 2022 was not the result of poor management in general. It was the result of one specific, structural failure: the accumulated cost of decades of deferred investment in internal technology.
SkySolver, the crew scheduling system at the center of the failure, was built in the 1990s for an airline one-tenth of Southwest’s current size. As the airline grew, the system was patched and extended rather than replaced, because replacement was expensive, disruptive, and would require someone to own the cost and risk. No single leader in any single year found it rational to absorb that cost. The cumulative cost of not replacing it, absorbed invisibly over decades in the form of periodic smaller failures, maintenance overhead, and constrained operational capability, was never aggregated and presented as a single number. If it had been, the decision to replace would have been obvious.
This is the sunk cost fallacy operating at its most destructive. Each year, the cost of maintaining SkySolver appeared manageable. Each year, the cost of replacing it appeared large. The comparison was made annually, on a one-year basis, rather than cumulatively across the system’s useful life. The annual comparison consistently favored maintenance. The lifetime comparison, visible only after the catastrophe, clearly favored replacement by an enormous margin. The $800M single-event loss plus the $1.3B replacement program that followed was the compound interest on thirty years of individually rational, collectively catastrophic decisions.
CHAPTER # 9 - "When Building Actually Makes Sense"
Any article that argues this forcefully against building owes its readers an honest account of when building is the right answer. The case for buying is strong. It is not absolute. And the credibility of everything else argued here depends on acknowledging that.
The exceptions are real. They are also rarer than most organizations believe, and they require a substantially higher bar of justification than most build decisions actually clear. Here is an honest attempt to define them.
The 4 Legitimate Cases
Your software IS your product. Google’s search algorithm, Netflix’s recommendation engine, Stripe’s payment processing core: these are not tools these companies use. They are the companies. The software creates competitive differentiation that cannot be purchased, cannot be replicated by a competitor who buys the same commercial platform, and is the primary source of customer value. When custom software is the reason customers choose you over your competitors, build is not just justified - it is necessary.
The test is simple and important: if a direct competitor bought the same commercial solution you would buy, would they be able to replicate your core value proposition? If the answer is yes, you probably should not build. Your operational efficiency and unit economics depend on being a better buyer than your competitors, not a better builder. Building is justified only when it creates advantages that buying cannot.
No adequate commercial solution exists. In some genuinely novel domains, the market has not yet produced a viable product. When the capability you need does not exist at any price from any vendor, building is not a preference - it is a necessity. This condition is less common than build advocates claim. Most ‘“no adequate solution exists” arguments turn out to be “no perfect solution exists,” which is a different claim. Perfect solutions almost never exist. Adequate ones usually do. The question is whether the gap between adequate and perfect is worth the full cost of building.
Deep proprietary data creates genuine, defensible advantage. Organizations with unique data assets that took decades to accumulate and cannot be replicated by a vendor may have legitimate reason to build models and systems that exploit those assets in ways a generic vendor cannot. This is real but narrow. Most enterprise data is less unique than the people sitting on it believe. The test is whether the data genuinely cannot be accessed or replicated by a focused external vendor who specializes in your domain. In most cases, the answer is that it can be.
Regulatory requirements genuinely preclude commercial solutions. In specific circumstances - classified government environments, specific data sovereignty requirements, regulated financial activities in jurisdictions with unusual technology constraints - commercial solutions may genuinely be unavailable. This is a legitimate constraint, not a preference. It is also a constraint that applies to a much smaller proportion of build decisions than the organizations invoking it typically suggest.
A Decision Framework That Removes the Bias
The problem with every framework for evaluating build-vs-buy decisions is that they are typically conducted by the people who have the most to gain from one answer. A framework is only as good as the integrity of the process in which it is applied. With that caveat, three questions should be answered before any build decision is approved:
What does the reference class say? How have similar projects at similar organizations, with similar scale, similar technical complexity, and similar team experience actually performed? Not this project’s best-case scenario. The actual historical base rate. Flyvbjerg’s research shows this question is almost never asked. The Standish CHAOS data, project post-mortems from comparable organizations, and the published failure literature exist for exactly this purpose. Using them systematically before committing to a build is the single most effective correction for planning fallacy and optimism bias.
What is the honest five-year TCO comparison? Not the first-year cost. The full total cost of ownership over 5 years, including all maintenance, all talent costs at realistic market rates, all security and compliance obligations, and the opportunity cost of engineering capacity consumed by internal maintenance. Set beside the equivalent number for the best available commercial alternative. In most honest calculations, this comparison favors buying by a substantial margin. The calculation must be made by someone who does not benefit from the build answer.
Who benefits from this recommendation? Are the people recommending the build decision the same people whose teams, budgets, and career trajectories grow as a result? If yes, that conflict of interest should be named and addressed, not ignored. The recommendation should be validated by someone external to the decision - a finance leader, an independent technical advisor, or a board member with relevant experience. This is not a criticism of the people making the recommendation. It is an acknowledgment of structural reality.
Peter Drucker’s formulation remains the clearest available summary of the underlying principle: “Do what you do best, and outsource the rest.” The “best” is narrower than most organizations admit. The “rest” is broader than most organizations are comfortable acknowledging. Getting that boundary right is one of the most consequential strategic decisions any technology leader makes, and it is one of the decisions most systematically distorted by the biases described in this article.
CHAPTER # 10 - "A Path Forward"
This article has made a strong argument. Now let me be honest about its limits.
Changing the build-vs-buy calculus inside most organizations is not primarily a knowledge problem. The people in those rooms are not deciding to build because they have not read the Standish CHAOS report. They are deciding to build because the incentive structures, the psychological dynamics, and the organizational politics of the room push them in that direction regardless of what the data says. Giving people more data rarely changes behavior when behavior is driven by incentives and identity rather than ignorance.
What can change is the structure within which those decisions are made. And there are specific, practical interventions that reduce the influence of the biases without requiring individuals to overcome their own psychology through sheer willpower.
Shift the Burden of Proof
The default in most organizations is that buy requires justification if it significantly exceeds the immediate cost of building. This frames the wrong option as the safe default. Build decisions carry higher risk, higher TCO, and a documented history of underperformance relative to buying for non-differentiating capabilities. The burden of proof should be proportionate to the risk, which means build should require a higher standard of justification than buy.
Concretely: any proposal to build custom software for a capability that is not a direct source of competitive differentiation should be required to provide reference class forecasting (how have comparable projects actually performed?), an honest five-year TCO comparison, and a specific articulation of what would be impossible to achieve by buying. The absence of any of these should be sufficient grounds for returning the proposal for more work.
Separate the Decision From the Deciders
The principal-agent problem in build-vs-buy decisions is structural, not personal. The people who benefit most from the build answer are also the people with the most technical credibility to evaluate the options. Asking them to evaluate both options honestly is asking them to overcome a structural conflict with personal integrity alone. Most people cannot do this consistently, even with the best intentions.
The structural solution is to separate the recommendation from the evaluation. Some organizations require that build-vs-buy analyses be conducted or validated by finance or strategy teams, not IT. Others use external advisors for the comparison phase. Some boards have required independent technical assessments for large technology decisions. The mechanism matters less than the principle: the evaluation of options should not be performed exclusively by the people whose organizational standing depends on the outcome.
Value Speed to Value Above Speed to Prototype
One of the most consistent findings across the research is that bought solutions deliver business value in 3 to 6 months, while built solutions take 18 to 36 months to break even. In industries where the technology landscape is evolving rapidly - and in the current AI environment, every industry qualifies - this gap is not just a financial calculation. It is a competitive one.
Organizations that deployed modern claims automation platforms in 2022 and 2023 are running agentic AI on mature, validated infrastructure today. Organizations that chose to build their own automation infrastructure in those same years are still maintaining it, still debugging it, still trying to get it to production quality. The gap compounds every quarter. Every month of delay in realizing value is not neutral. It is a loss that accumulates interest.
The organizational countermeasure is to evaluate technology investments on time-to-value, not time-to-prototype. A team that can demonstrate a working demo in two weeks is impressive. A team that can demonstrate documented business impact in 90 days is rare and valuable. Those are different capabilities. Organizations that reward the second are more likely to buy wisely than organizations that celebrate the first.
For AI: Buy the Infrastructure, Build the Differentiation
The most practical framework for AI-era technology decisions is not binary. The question is not “build or buy.” It is: at which layer of the stack does genuine differentiation reside?
For almost every organization, the answer is: not at the infrastructure layer. Foundation models, orchestration frameworks, document processing pipelines, RAG infrastructure, vector databases, observability tooling, compliance frameworks - these are solved problems. Hundreds of specialized vendors are competing intensely to solve them. The marginal cost of buying access to excellent solutions in these categories is low and falling. Building them internally is the Redundancy Tax at its most expensive.
The narrow layer where building may be justified is the layer of genuine proprietary advantage: the specific risk assessment logic developed from decades of claims data, the pricing model built on loss history that competitors cannot replicate, the workflow optimization developed from domain expertise that a generic vendor cannot encode. This layer is real. It is also thinner than most organizations believe.
The practical implementation: buy the platform, configure it deeply for your domain, and build narrowly, with rigorous governance only the components that represent genuine intellectual property. Use the engineering capacity freed by buying the infrastructure for the work that is actually yours.
Buy the infrastructure. Build the genuine differentiation. The hard and necessary part is being honest about which is which - and resisting the structural forces that make everything feel like differentiation.
CONCLUSION - "The Argument, Ended"
The data is clear. The behavioral science is clear. The organizational dynamics are clear. The historical record is, if anything, even clearer.
Most organizations that build software they could buy do so not because the business case supports it, but because the people making the decision want to build. Because building grows teams and advances careers. Because the IKEA Effect and the endowment effect make internally created software feel more valuable than equivalent software built by someone else. Because the planning fallacy makes the cost feel smaller than it will be. Because uniqueness bias makes the business feel more special than it is. Because loss aversion makes vendor dependency feel riskier than internal dependency. Because groupthink suppresses the voices that might say otherwise. Because escalation of commitment makes it almost impossible to stop once started.
And because the people making the recommendation are structurally rewarded for one answer and structurally shielded from the consequences of that answer when it goes wrong.
None of this means organizations should never build. The exceptions exist. They are narrow, and the bar for justifying them should be high. Tesla builds its own battery management software because battery performance is the core of its competitive advantage. Netflix builds its own recommendation system because recommendation is the core of its customer experience. Google builds its own search infrastructure because search is its business. These are not examples that support the general case for building. They are exceptions that prove the general rule: build when the software is the differentiation, buy when the software enables the differentiation.
The paradox of the build-vs-buy decision is this: the choice that feels like taking control is almost always the choice that surrenders it. By building, you surrender the control of your roadmap to your own internal team’s limited capacity. You surrender the control of your costs to the compounding weight of maintenance and technical debt. You surrender the control of your timeline to the planning fallacy and the Iron Law. And you surrender the control of your future to a legacy system that becomes progressively more expensive to change and increasingly less capable of supporting the technology you will need next year and the year after.
Buying from a focused vendor who does nothing else - who iterates continuously, incorporates feedback from hundreds of clients, maintains regulatory compliance, and updates the platform as the underlying technology evolves - is not a loss of control. It is an intelligent transfer of the work to the people best positioned to do it. That is not a compromise. It is strategy.
Do what you do best, and outsource the rest.
Artem Gonchakov,
CEO, Simplifai











