What to assess

Assessing architecture and scalability

Architecture is the way a software system is built and organised, and scalability is whether that design keeps working under the growth the deal is paying for. Assessing both in technical due diligence means asking the team to explain the system as it is today, finding the specific parts that will break first as users, data, or traffic multiply, reading the incident history for whether root causes get fixed, and checking whether running costs rise faster than revenue. Most scaling problems are repairs with a price. The warning sign is a team that cannot answer the questions.

Published July 27, 2026. Updated September 17, 2026. Editorial.

Key takeaways

  • Architecture is how the system is put together; scalability is whether it keeps working as load grows. The deal question is whether it supports the growth in the thesis.
  • Ask what breaks first and at what scale. A strong team knows the answer; a team that says it will just scale has not thought about it or is not being honest.
  • The AWS Well-Architected Framework's six review areas and Google's service level objective model give a reviewer a shared vocabulary for reliability, cost and performance.
  • Most scaling problems are priced work for the first year. An architecture the team does not understand well enough to explain is the finding that changes a deal.

Architecture is how a system is put together: how the pieces are split, how they communicate with each other, and where the data lives. Scalability is whether that structure keeps working as the company grows. For an investor, these matter because almost every deal thesis assumes growth, and growth is exactly what breaks systems that were built for today's size.

Why is architecture a business question?

Architecture is a business question because the limits of a system cannot be seen until the growth in the plan happens, and that is the most expensive moment to discover them. A system that runs well for ten thousand users can stop working at a hundred thousand. A database design that is fast with a gigabyte of data can become slow at a terabyte. A company can look healthy until the platform cannot handle the load its own growth created. Then it must fix the problem under pressure, with unhappy customers, on a deadline it does not control.

So the question in diligence is "will this architecture support the growth the thesis assumes, and if not, what breaks first and what does it cost to fix". That connects directly to the deal, which is the approach of the main guide. If the thesis depends on scaling the organisation as much as the system, the scaling engineering teams guide is the related guide, because a system and the team behind it scale together.

Reveneau's architecture assessment starts with a live explanation by the engineers rather than a diagram, because a diagram shows the system as someone once drew it and a live explanation shows the system as the engineers understand it today; the difference between the two is itself a finding.

What should the team be able to explain?

The team should be able to explain the main components, where user data lives, what happens when a request comes in, and which parts everything else depends on. You are listening for two things: whether they understand their own system, and whether the design has single points that cannot be duplicated.

The twelve-factor methodology, published by Adam Wiggins and used as a reference for software delivered as a service, gives a reviewer a short list of things to ask about: one codebase tracked in version control, explicit dependencies, configuration kept out of the code, backing services such as databases treated as attachable resources, a strict separation of build, release and run, stateless processes, and development and production kept as similar as possible. A team that can say which of those it follows and which it does not has thought about its architecture.

There is also a pattern worth knowing when you compare the org chart with the diagram. Conway's law, stated by Melvin Conway in 1968, observes that organisations that design systems are constrained to produce designs that are copies of their own communication structures. If three teams do not talk, the system will have three parts that do not talk either. When the architecture looks disorganised, ask how the teams are organised before concluding the engineers are careless.

What breaks first as the company grows?

Ask the direct question: what breaks first as we grow, and at what scale. A strong team knows the answer, because they think about it. They can tell you the database will need attention around a certain size, or a particular service will need to be split. A team that says "it will just scale" either has not thought about it or is not being honest with you, and both are a concern.

Common limit How it shows up Usual repair Cost range
Database cannot keep up Slow pages as data grows, one large table everything reads Indexing, splitting data, read replicas, a different store for one workload Engineering weeks to months
Tightly coupled system Every change touches many parts, one failure stops everything Separate the parts that change at different rates Partial rewrite, months
Single server or single process No way to add capacity except a bigger machine Stateless processes behind a load balancer Known fix, weeks
Every call waits for a reply Slow third-party calls block users Queues and background jobs Weeks per workflow
Manual scaling Someone adds capacity by hand when it is slow Automated scaling rules Weeks, plus the cost work below
No separation of environments Testing happens in production Staging that matches production Weeks, and a process change

These are prices rather than verdicts. A system with real scaling work ahead of it is normal, especially for a fast-growing company that built for speed first. As we have argued, execution speed is the moat, and teams that moved fast often have unfinished scaling work as a result. The job is to price that work against the thesis, which the assessing technical debt page covers in full.

How do you read reliability from the incident history?

Read it for whether the team fixes the underlying causes of failures or only their visible effects. Every real system has incidents. The incident history is a finding only when the same failure repeats, when nobody wrote anything down, or when the team cannot say how long recovery took.

Google's Site Reliability Engineering book gives a reviewer the vocabulary. A service level indicator is "a carefully defined quantitative measure of some aspect of the level of service", such as the share of requests that succeed. A service level objective is "a target value or range of values for a service level that is measured by an SLI". A service level agreement is the contract with users that carries consequences for missing the objective. The book also states that insisting on meeting objectives 100 percent of the time is "both unrealistic and undesirable", because it slows innovation and forces expensive, over-conservative designs, and it recommends an error budget, a rate at which the objective can be missed. A team that can name its objectives and its error budget is a team that has decided how reliable it needs to be. A team with no objectives is guessing.

The same book describes a postmortem as "a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause(s), and the follow-up actions to prevent the incident from recurring", and lists triggers for writing one: user-visible downtime beyond a threshold, any data loss, on-call intervention, and a resolution time beyond a threshold. Ask for three recent postmortems. Read them for the root cause and the follow-up, and check whether the follow-up was completed. DORA's failed deployment recovery time, the time to recover from a deployment that needs immediate intervention, is the number to ask for alongside them; the assessing code quality page explains the other four DORA metrics.

What framework can a reviewer use to structure the assessment?

The AWS Well-Architected Framework is the most widely published one, and its structure works for systems on any cloud. AWS describes it as key concepts, design principles and best practices for designing and running workloads in the cloud, organised into six areas it calls pillars.

Pillar AWS's focus The diligence question
Operational excellence Running and monitoring systems, improving processes Who is on call, what do they see, how do they learn from incidents?
Security Protecting information and systems Covered on the security and compliance page
Reliability Performing intended functions and recovering from failure What are the objectives, what is the error budget, what broke last quarter?
Performance efficiency Allocating computing resources in a structured way What breaks first as load grows, and at what scale?
Cost optimisation Avoiding unnecessary costs How does running cost change with users and revenue?
Sustainability Minimising environmental impact Rarely a deal question, unless a customer contract requires it

A reviewer does not need to run the full AWS review. Using the six headings to organise the live explanation is enough to make sure nothing is skipped and to give the deal team a report structure they will see again from other advisers.

How does infrastructure cost reveal a scaling problem?

Infrastructure cost reveals a scaling problem when it rises faster than revenue, because that means each new customer costs more to serve than the last. The FinOps Foundation defines its discipline as an operational framework and cultural practice that creates financial accountability through collaboration between engineering, finance and business teams, and the first thing that practice asks for is a bill broken down by product, customer and environment. A target that can produce that breakdown can answer the question; a target that has one undivided monthly bill cannot.

Ask what the infrastructure costs today, how that cost has moved over twelve months, and how it moves with each new customer. Ask whether anyone is responsible for the bill. Ask what happens to the cost if traffic doubles. Moving between cloud platforms is a project of its own (our note on migrating an app from AWS to GCP sets out what it involves), so a plan that depends on a migration needs its own line in the budget. The cloud cost due diligence page covers rebuilding the bill in detail, and the SaaS metrics versus engineering reality page covers what happens when the gross margin in the model and the infrastructure bill disagree.

How do you put it in the deal?

Name the gap between where the architecture is and where the growth in the thesis needs it to be, estimate the work to close it, and raise that in the deal negotiation. A clean architecture with room to grow counts in favour of the deal. An architecture with known limits is a priced task for the first year, and the technology after the deal guide covers how that task becomes part of the hundred-day plan. An architecture the team does not understand well enough to answer questions about is the warning, and the warning signs page covers why. The technical due diligence checklist lists the architecture questions in a form you can take into the meeting.

Best for

  • Growth-equity and buyout deals where the thesis assumes several times today's load
  • Any target whose infrastructure bill is a meaningful share of cost of revenue
  • Reviews where the diagram in the data room is older than the last major release

Avoid if

  • The product is a prototype with no production traffic yet, where the team page matters more
  • You need the cloud bill rebuilt line by line, which has its own page

Check before you decide

  • Confirm the team can say what breaks first and at what scale, without vague answers
  • Confirm three recent postmortems exist and their follow-up actions were completed
  • Confirm infrastructure cost per customer over twelve months, from the bill rather than the model

Common questions

How do you assess scalability in technical due diligence?

Ask what breaks first as the company grows and at what scale, read the incident history for whether root causes get fixed, and check whether infrastructure cost rises faster than revenue. Use the AWS Well-Architected Framework's six review areas (operational excellence, security, reliability, performance efficiency, cost optimisation, sustainability), read on 17 September 2026, to structure the review so nothing is skipped.

Is a system that cannot scale a reason to stop the deal?

Usually not. Most scaling problems, such as a database that cannot keep up or a tightly coupled design, are fixable with engineering time, so you price the work against the deal. The finding that changes a deal is a team that does not understand its own architecture well enough to answer questions about it. Conway's law (1968) is a reminder to check the org chart before blaming the engineers for a disorganised design.

What questions reveal architecture risk?

Ask the team to explain the architecture live and say where data lives, ask what breaks first as load grows, and ask for an account of a recent outage. Google's SRE book, read on 17 September 2026, adds three more: what are your service level objectives, what is your error budget, and can you show a postmortem with a root cause and a follow-up that was completed. Clear, specific answers show a team that understands its system.

What is the difference between architecture and scalability?

Architecture is how a system is put together: how the pieces are split, how they communicate with each other, and where the data lives. Scalability is whether that structure keeps working as the company grows. A system can have a clean architecture today and still fail at ten times the load if it was never designed for it. The twelve-factor methodology, read on 17 September 2026, gives a short list of architectural practices that keep a service scalable.

How much does fixing a scaling problem usually cost?

The cost depends on the type of problem. A database that cannot keep up is usually fixable with engineering weeks to months: indexing, splitting the data, or a different store for one workload. A tightly coupled system where everything depends on everything is closer to a partial rewrite and costs more. Moving to a different cloud platform is a separate project with its own cost, so price it separately when the plan depends on one.

Is a history of outages a warning sign in technical due diligence?

No. Every real system has incidents. What matters is whether the team learned from them and fixed the underlying cause. Google's SRE book defines a postmortem as a written record of an incident, its impact, the actions taken, the root causes and the follow-up actions, and lists any data loss and user-visible downtime beyond a threshold as triggers for writing one. Ask for three and check whether the follow-ups were completed.

How does infrastructure cost show a scaling problem?

A system whose running cost rises faster than its revenue has a scaling problem that shows in the cost figures, because each new customer costs more to serve than the last. The FinOps Foundation defines its practice as creating financial accountability through collaboration between engineering, finance and business teams, and the first step is a bill broken down by product, customer and environment. A target with one undivided monthly bill cannot answer the question yet.

What is a service level objective and why does it matter in diligence?

A service level objective is, in Google's SRE book's words, a target value or range of values for a service level measured by an indicator such as the share of successful requests. It matters in diligence because a team that can name its objectives and its error budget has decided how reliable the system needs to be, and the book states that insisting on 100 percent is both unrealistic and undesirable. A team with no objectives is guessing.

Should a bad architecture finding stop a deal?

Rarely on its own. Most scaling problems have a price, and a system with real scaling work ahead is normal for a fast-growing company that built for speed first. Put the work into the first-year plan and, where it is large, into the price. The AWS Well-Architected Framework's reliability area, which focuses on performing intended functions and recovering quickly from failure, is the heading most of those findings belong to.

What does the twelve-factor methodology have to do with due diligence?

The twelve factors are a published checklist of practices for software delivered as a service: one codebase in version control, explicit dependencies, configuration outside the code, backing services as attachable resources, separated build, release and run stages, stateless processes, and development kept close to production. Asking a team which factors it follows, and which it does not, is a fast way to learn whether it has thought about its own architecture.

More in What to assess

How to assess code quality

Code quality in technical due diligence means one thing: whether the software is safe and cheap to change. If small changes are slow and risky, the roadmap in the deal deck will not be delivered on time. The most reliable evidence comes from how the team releases software today, measured with the five DORA delivery metrics, a live run of the test suite, and the pull request history, rather than from how the code looks in a screenshot. This page explains what to measure, what to ask, and how to judge quality when you cannot read the code yourself.

Assessing the engineering team

Assessing the engineering team in technical due diligence means judging whether the people can produce every future version of the software the deal is paying for. Code shows the system at one moment; the team is what turns it into the next release. The review measures seniority, the bus factor (how many people would have to leave before the project stops), tenure and turnover, how the team explains its own decisions, and how it works with AI tools. A strong senior team can fix weak code. Excellent code with too few engineers stops improving as soon as the work gets hard or a key person leaves.

Assessing security and compliance

Assessing security and compliance in technical due diligence means working out what liabilities come with the company and what it costs to reduce them. A breach or a regulatory gap can remove the value of a deal, and some of these problems take months to fix. The review does not need a penetration test to find the largest risks. It checks basic security practices (who can access production, where secrets are stored, whether systems are patched), compares the target's controls with the OWASP Top 10 and the NIST Cybersecurity Framework, reads the incident history, and confirms which certifications are current rather than planned.

Assessing technical debt

Technical debt is the extra work created by quick solutions a company chose to move fast, plus the ongoing cost they add later as slower, riskier changes. Assessing it in technical due diligence means telling acceptable debt from debt that will stop the roadmap, measuring it well enough to price, and checking whether the team has a plan to fix it. Every real codebase has debt, so the finding is never that debt exists. The finding is where it is, whether the team can name it, and whether releases are getting slower because of it.