Guides for building software well
Practical, opinionated guides from a team that builds and releases software. Every guide is written to be used.
How to choose a software development partner
Most of the cost of building software comes from choosing who builds it with you, and from getting that decision right before the first line of code is written.
Read the guide- What is a software development partner?
- In-house vs outsourced development: which fits your product?
- Development agency vs staff augmentation
- Onshore, nearshore, or offshore development
- Questions to ask a software development partner before you hire
- How to evaluate a development team's engineering quality
How to build an AI product that reaches production
The hard part of an AI product is everything between a demo that impresses and a product people trust every day. This guide is about that work.
10 chapters03 · GuideHow to launch a product: from idea to MVP
An MVP is not a smaller version of the product you imagine. It is the fastest honest test of whether the product is worth building at all. Get that framing right and everything after it (scope, stack, team, launch) gets simpler and cheaper.
9 chapters04 · GuideTechnical due diligence: a complete guide for investors
Technical due diligence is a focused review of a company's software and the people who build it, run before an investment or acquisition, to check whether the technology can support the plan being paid for. It covers code quality, architecture and scalability, the engineering team, security and compliance, and technical debt, and it ends in a priced list of risks the deal team can act on. This guide explains what the review is, how it runs on a deal timeline, what it costs, and which findings change the price or the terms.
19 chapters05 · GuideAI startup due diligence: how investors check what is real
AI startup due diligence is the work of checking that an AI company's product does what the pitch says, that the code behind it can be maintained and sold, and that the margin stays positive after the model-provider invoices. It differs from ordinary technical diligence in three places: the claim can be tested live instead of inspected, the codebase was probably written by a model, and the cost of goods is a variable bill from a vendor who can change the price. This guide covers all three for VC associates and principals, private equity deal teams, and angels in AI deals.
12 chapters06 · GuideAngel investor due diligence on technology: judging a startup's software
Angel investor due diligence on technology is the part of the checklist that the published angel checklists leave out. The Angel Capital Association's 2007 best-practice guide has no section on software or code, and the UK Business Angels Association's 2020 checklist asks eight technical questions, all of them about positioning and intellectual property. This guide covers that missing part for angels, angel group members and syndicate leads, most of whom do not write code. It covers what to check in a screen share, what an MVP claim should mean, how to see a single-developer risk, how to judge a technical founder, and when to pay a firm to look properly.
11 chapters07 · GuideTechnology after the deal: the engineering work that follows an investment
Technology after the deal is the engineering work that starts when the investment closes: the technology workstream in the 100-day plan, merging or separating codebases, sizing technical debt across a portfolio, choosing who leads engineering, and reporting progress to the fund in measures a board can read. This guide is for private equity operating partners, venture platform leads and portfolio company board members. It covers what to do in each phase, what to check, the mistakes that repeat, and the timelines and costs that third parties have measured.
10 chapters08 · GuideHow to prepare for technical due diligence: a founder's guide
Preparing for technical due diligence means assembling, before an investor asks, the evidence that your software does what your deck says, that your team can keep building it, and that nothing in the code, the licences or the state of your security will surprise anyone after the investment is made. This guide is written for founders and CTOs raising a round or selling the company. It covers what a technical review asks for, what to put in the engineering data room, how to present your architecture and roadmap, how the review changes from seed to Series B and at a sale, what changes when the code was written by AI, and a day-by-day checklist for the two weeks before the review starts.
9 chapters09 · GuideHow to scale your engineering team
Scaling an engineering team is not about hiring faster. It is about adding the right capacity at the right time, in a way that makes the team quicker rather than slower. Most teams get the timing and the method wrong, and pay for it in speed.
9 chapters10 · GuideCustom software: build vs buy, done right
Most companies do not need custom software for everything. They need it for the few things that make them different, and they need to buy the rest. This guide is about telling those apart, then building the custom part in a way you will not regret.
9 chapters11 · GuideWhat software costs and how to budget it
The price of software is not one number. It is the result of choices you make about scope, seniority, and how right the product has to be. This guide explains what actually drives cost so you can build a budget that stays accurate, instead of relying on a quote that does not.
8 chapters12 · GuideFrom AI-generated code to production software
AI can write most of a codebase in an afternoon. It cannot tell you whether that codebase is safe to run. This guide is about the difference between the two, and how to deal with it.
10 chapters13 · GuideEval-driven development: how to prove AI-written code works
AI can write a feature in minutes. Proving that the feature does what you asked still takes real work, and that work is now the only thing that stops a fast team from releasing a broken product.
13 chapters14 · GuideWhy AI evals matter: the test that shows whether an AI product works
An AI product can give different output on different days and for different inputs, so a demo, a vendor's promise and a public benchmark score are weak evidence that it works for your customers. An AI eval is the stronger evidence: a written, repeatable test with an input, an expectation and a scoring rule. This guide is written for the people who approve budgets and releases. It explains what an eval is, what public records show about wrong AI answers and advertising claims, how to read a pass rate and use it to decide a release, what the work costs and who owns it, and what the texts of regulators say about testing.
8 chapters15 · GuideLLM evals: how to measure whether an AI product works
LLM evals measure whether a product built on a large language model does its job. The method has a fixed order: learn the three kinds of eval, read real outputs and name the failure types, build a test set from real usage, size the set for the difference you need to detect, choose a score for each failure type, and keep measuring after release. A 90 percent pass rate on 100 cases has a 95 percent confidence interval of 84.1 to 95.9 percent, by the formula in Evan Miller's 2024 paper, so a small set cannot confirm a small improvement. Each step in this guide is tied to a named paper or vendor guide.
8 chapters16 · GuideAI agent evals: how to test an agent that takes actions
An AI agent is a language model working in a loop: it picks a step, calls a tool such as a search or a database, reads the result and repeats until the task is done. Because it takes many steps and acts on real systems, a test of its final answer covers a small part of what can go wrong. This guide is the method. Grade the end state and the steps, test each tool call, repeat every case because one pass proves little, set limits on cost, time and steps, run the tests in a closed copy of the system, test for attacks, and turn production runs into new test cases.
8 chapters17 · GuideAI benchmarks vs your own evals: how to read a model score
A benchmark score answers one question: how did this model do on that public test, under the settings chosen by whoever ran it. A buyer needs the answer to a different question: how will the model do on my work. This guide explains the benchmarks vendors quote, including SWE-bench, MMLU, GPQA, ARC-AGI and the agent tests, with each one's size and method from its own paper. It covers three ways a score misleads: the questions were in the training data, every model scores near the top, and the vendor chose the settings. It ends with a seven-step method for choosing a model on your own cases.
8 chapters18 · GuideFaster evals with Jev: how we grade AI-written code
An eval suite for AI-written code has two kinds of check. Some have an exact answer, and code decides them. The rest used to need a language model to read the change and give an opinion, which was the slowest and least repeatable step in the suite. Reveneau moved that step to Jev, a decision model that answers fixed questions with probabilities instead of writing text, and on our own suite the run now finishes ten times faster than it did with a language-model grader. This guide explains how the grader works, how to design its questions, and where it stops being useful.
11 chapters19 · GuideJev and System One models: a plain guide
Jev is a hosted model from TypeSafe AI that answers fixed questions about a piece of text with calibrated probabilities instead of writing a reply. TypeSafe calls this a System One model. It returns a choice, a score or a yes/no probability in well under a second, and TypeSafe prices it at $0.042 per million input tokens with output free. This guide explains what that buys you, where the model is known to fail, and how to decide between Jev and a language model for each decision in your software.
10 chapters20 · GuideJev in production: guardrails, routing and classification
A decision model answers a fixed question about a piece of text and returns a probability instead of a paragraph. Inside an application that makes it useful in four places: routing a request to a handler, gating an action on how sure the model is, checking a value that code already extracted, and screening what goes into and comes out of a language model. This guide covers each of those with Jev, TypeSafe's System One model, pattern by pattern, with the vendor's own figures marked as the vendor's and the limits stated at the start.
9 chapters21 · GuideShould your product use a decision model? A guide for founders and CTOs
A decision model answers a fixed question about a piece of text with a probability, in under half a second, at a price TypeSafe lists as $0.042 per million input tokens. It writes nothing. That makes it the right tool for a narrow set of product features, the wrong tool for most others, and a governance question either way. This guide is written for the person who approves the decision: how to sort features, how to run a two-week pilot, what to ask a vendor or partner, and how to keep the model auditable once it is in production.
9 chapters22 · GuideHow to deliver high-quality software with AI-native development
AI now writes code fast, so writing code is no longer the slow step. The slow steps are the work around it: deciding what to build, proving it works, releasing it safely, and running it afterwards. This guide describes how a team should organise all four, at the speed AI now makes possible.
15 chapters23 · GuideTaking AI agents from prototype to production
An agent that works in a demo and an agent that works in production are different systems. The demo has to succeed once, with someone friendly at the keyboard. Production has to succeed repeatedly, for people who did not read the instructions, with real money and real data affected by each call.
7 chapters24 · GuideProduct discovery and UX design: a buyer's guide
Most failed software was built competently. The team delivered what was asked for, on a reasonable timeline, and it still did not matter, because what was asked for was never tested against what users actually needed. Discovery and design are how you buy that test before you pay for the build.
8 chapters25 · GuideBuilding compliant software in financial services
Financial regulation is already written as rules. That is the part most teams miss. A rule that says a record must be recreatable after deletion is a specification, and a specification can be turned into a check that runs on every change. This guide covers how to do that, which obligations actually apply to your code, and why compliance gaps in a working app are so easy to miss.
11 chapters26 · GuideBuilding compliant software in healthcare
HIPAA is unusual among regulations in that it tells you what to achieve and mostly declines to tell you how. That flexibility is why so many healthcare applications are non-compliant while holding a completed risk assessment: the rule left the decision to you, and nobody wrote the decision down. This guide covers what the Security Rule requires from software, and how to turn it into something a system enforces.
8 chapters27 · GuideBuilding compliant software in real estate
Real estate software has a compliance problem the other regulated industries do not. In finance and healthcare, the rules mostly govern how you handle data. In housing, the rules govern the decision your software makes, which means the model's output is itself the regulated act. That changes what you have to build and what you have to test.
8 chapters28 · GuideBuilding compliant software in insurance
Insurance regulation is state by state, which most teams treat as a reason to guess. It is really a reason to build the rule into configuration rather than into code, so the same application can satisfy Iowa and Ohio without a special case for each. This guide covers what the NAIC model laws actually require, which parts apply to your application, and why a working system so often looks compliant to its own team while missing the one thing an examiner asks for.
8 chapters29 · GuideBuilding compliant software in legal technology
Legal technology carries a rule most other software does not: the product itself can end up practicing law, even when nobody intended it to. This guide covers where that boundary is, what the ethics rules that govern lawyers actually require of the tools they use, which e-discovery obligations affect your data model, and how to turn each of those into something a build can be tested against rather than a policy nobody rereads.
8 chapters30 · GuideSoftware for compliance-heavy industries
Regulated software has always needed exhaustive verification and rarely received it, because writing the checks cost more than anyone would approve. That constraint is gone. This guide covers what that changes, the obligations that apply regardless of sector, and how to tell a partner who verifies their work from one who says they do.
6 chapters31 · GuideForward deployed engineering: the complete guide
A forward deployed engineer is a software engineer who works inside the customer's environment, on the customer's real data, and stays responsible until the system runs in production. The role exists because software that works in a demo and software that keeps working inside a real company are two different pieces of work, and somebody has to do the second one.
46 chapters