What we learn building software and AI.

What to do with the code nobody understands
Every system has a part everyone avoids. The instinct is to rewrite it or ignore it, and both are wrong for the same reason: you cannot tell yet which of its behaviour is deliberate.
Read
What a buyer should do now that OpenAI has stopped reporting SWE-bench Verified
On 23 February 2026 OpenAI said it had stopped reporting SWE-bench Verified, and on 8 July 2026 it withdrew its recommendation of SWE-Bench Pro. This post gives OpenAI's figures with their limits, then shows how to choose a coding model on 100 of your own tasks.

How to plan a migration you cannot pause
The single, all-at-once cutover is popular because it is easy to describe and easy to schedule. It is also the version where you find out whether it worked at the moment you can least afford to be wrong.

An AI agent that passes a test once can fail it on the next run
The same agent on the same task passes on some runs and fails on others, and the tau-bench paper of 2024 measured by how much. Run each test case several times and report the share of tasks that passed every run.

Why flaky tests are worse when agents write them
A flaky test used to cost a rerun and some irritation. When an agent is reading the result to decide whether it is finished, an unreliable check becomes an unreliable instruction.

What a 90 percent pass rate on 50 test cases tells you
A pass rate of 90 percent measured on 50 test cases fits a true rate anywhere from 81.7 to 98.3 percent, so a result of 92 percent on the same 50 cases cannot be told apart from it. This post shows the arithmetic step by step and lists what to report next to every pass rate.

What the Air Canada chatbot ruling means for a company with an AI assistant
A British Columbia tribunal treated a chatbot's answer as the airline's own statement and ordered a payment of $812.02. This post sets out what the decision says, what it leaves out, and the written test that compares an assistant's answers with the policy text.

How to test a feature you cannot fully specify
Some features have no single correct output. A summary, a ranking, a suggested reply. You cannot check those against one exact expected answer, and the usual conclusion, that they cannot be tested, is wrong.

What a good handover document actually contains
Most handover documents describe the architecture, which is the one thing the next team can work out for themselves. The valuable parts are the decisions, the unknowns, and the things only one person knew.

How to negotiate a software contract you can actually verify
Most build contracts describe effort, timeline, and payment, and leave the one hard question unanswered: on what evidence do you agree the thing is finished?

Forward deployed engineer vs software engineer: which one to hire next
Both write production code to the same standard. Three questions about your next problem tell you which of the two engineers it needs.

Why observability matters more when a machine wrote the code
When a person writes a system, someone carries a mental model of it. When a model writes it, nobody does, and production becomes the only place where you can see what the system actually does.

What we checked before grading with a week-old model
Jev launched on 15 September 2026 and we put it in charge of grading our eval suite the same month. Here is the order of checks we ran first, written as a method you can repeat, and the one outcome we are willing to state.

The real price of an LLM judge is per decision, and it multiplies
A grader's price per call looks small. Multiply it by every check on every change, or every message through a guardrail, and the public figures from LangChain, Openlayer and TypeSafe show where a cheaper decision model pays for its own migration and where it does not.

How to decommission a feature safely
Removal is the only thing that actually reduces maintenance load, and it is the only engineering work nobody ever schedules. Here is how to find what nobody uses, spot the dependencies you cannot see, and take it out in steps you can reverse.

What changes when your junior engineers use agents
The way to become senior was writing a lot of code and having it criticised in detail. If the agent writes it, that cycle stops, and nobody notices for about eighteen months.
From AI News
Short briefs on AI and software engineering, published as things happen.
SCM searches every photo and every video frame on a Mac from a plain-language description, offline
jianying-headless turns a JSON plan into an editable Jianying Pro draft, 2,963 stars in 19 days
yomiyasu is an agent skill that rewrites AI-generated Japanese to sound like a person wrote it, with 7 rules and a built-in linter
live-panel is a Claude Code skill that turns one JSON file into an animated architecture diagram you can post as a video
BeefTV is a local-first node canvas for AI video work, with text, image, video and audio generation on one project surface
Guides
In-depth, practical guides from a team that builds and releases software.
How to choose a software development partner
Most of the cost of building software comes from choosing who builds it with you, and from getting that decision right before the first line of code is written.
How to build an AI product that reaches production
The hard part of an AI product is everything between a demo that impresses and a product people trust every day. This guide is about that work.
How to launch a product: from idea to MVP
An MVP is not a smaller version of the product you imagine. It is the fastest honest test of whether the product is worth building at all. Get that framing right and everything after it (scope, stack, team, launch) gets simpler and cheaper.
Technical due diligence: a complete guide for investors
Technical due diligence is a focused review of a company's software and the people who build it, run before an investment or acquisition, to check whether the technology can support the plan being paid for. It covers code quality, architecture and scalability, the engineering team, security and compliance, and technical debt, and it ends in a priced list of risks the deal team can act on. This guide explains what the review is, how it runs on a deal timeline, what it costs, and which findings change the price or the terms.
AI startup due diligence: how investors check what is real
AI startup due diligence is the work of checking that an AI company's product does what the pitch says, that the code behind it can be maintained and sold, and that the margin stays positive after the model-provider invoices. It differs from ordinary technical diligence in three places: the claim can be tested live instead of inspected, the codebase was probably written by a model, and the cost of goods is a variable bill from a vendor who can change the price. This guide covers all three for VC associates and principals, private equity deal teams, and angels in AI deals.
Angel investor due diligence on technology: judging a startup's software
Angel investor due diligence on technology is the part of the checklist that the published angel checklists leave out. The Angel Capital Association's 2007 best-practice guide has no section on software or code, and the UK Business Angels Association's 2020 checklist asks eight technical questions, all of them about positioning and intellectual property. This guide covers that missing part for angels, angel group members and syndicate leads, most of whom do not write code. It covers what to check in a screen share, what an MVP claim should mean, how to see a single-developer risk, how to judge a technical founder, and when to pay a firm to look properly.
Technology after the deal: the engineering work that follows an investment
Technology after the deal is the engineering work that starts when the investment closes: the technology workstream in the 100-day plan, merging or separating codebases, sizing technical debt across a portfolio, choosing who leads engineering, and reporting progress to the fund in measures a board can read. This guide is for private equity operating partners, venture platform leads and portfolio company board members. It covers what to do in each phase, what to check, the mistakes that repeat, and the timelines and costs that third parties have measured.
How to prepare for technical due diligence: a founder's guide
Preparing for technical due diligence means assembling, before an investor asks, the evidence that your software does what your deck says, that your team can keep building it, and that nothing in the code, the licences or the state of your security will surprise anyone after the investment is made. This guide is written for founders and CTOs raising a round or selling the company. It covers what a technical review asks for, what to put in the engineering data room, how to present your architecture and roadmap, how the review changes from seed to Series B and at a sale, what changes when the code was written by AI, and a day-by-day checklist for the two weeks before the review starts.
How to scale your engineering team
Scaling an engineering team is not about hiring faster. It is about adding the right capacity at the right time, in a way that makes the team quicker rather than slower. Most teams get the timing and the method wrong, and pay for it in speed.
Custom software: build vs buy, done right
Most companies do not need custom software for everything. They need it for the few things that make them different, and they need to buy the rest. This guide is about telling those apart, then building the custom part in a way you will not regret.
What software costs and how to budget it
The price of software is not one number. It is the result of choices you make about scope, seniority, and how right the product has to be. This guide explains what actually drives cost so you can build a budget that stays accurate, instead of relying on a quote that does not.
From AI-generated code to production software
AI can write most of a codebase in an afternoon. It cannot tell you whether that codebase is safe to run. This guide is about the difference between the two, and how to deal with it.
Eval-driven development: how to prove AI-written code works
AI can write a feature in minutes. Proving that the feature does what you asked still takes real work, and that work is now the only thing that stops a fast team from releasing a broken product.
Why AI evals matter: the test that shows whether an AI product works
An AI product can give different output on different days and for different inputs, so a demo, a vendor's promise and a public benchmark score are weak evidence that it works for your customers. An AI eval is the stronger evidence: a written, repeatable test with an input, an expectation and a scoring rule. This guide is written for the people who approve budgets and releases. It explains what an eval is, what public records show about wrong AI answers and advertising claims, how to read a pass rate and use it to decide a release, what the work costs and who owns it, and what the texts of regulators say about testing.
LLM evals: how to measure whether an AI product works
LLM evals measure whether a product built on a large language model does its job. The method has a fixed order: learn the three kinds of eval, read real outputs and name the failure types, build a test set from real usage, size the set for the difference you need to detect, choose a score for each failure type, and keep measuring after release. A 90 percent pass rate on 100 cases has a 95 percent confidence interval of 84.1 to 95.9 percent, by the formula in Evan Miller's 2024 paper, so a small set cannot confirm a small improvement. Each step in this guide is tied to a named paper or vendor guide.
AI agent evals: how to test an agent that takes actions
An AI agent is a language model working in a loop: it picks a step, calls a tool such as a search or a database, reads the result and repeats until the task is done. Because it takes many steps and acts on real systems, a test of its final answer covers a small part of what can go wrong. This guide is the method. Grade the end state and the steps, test each tool call, repeat every case because one pass proves little, set limits on cost, time and steps, run the tests in a closed copy of the system, test for attacks, and turn production runs into new test cases.
AI benchmarks vs your own evals: how to read a model score
A benchmark score answers one question: how did this model do on that public test, under the settings chosen by whoever ran it. A buyer needs the answer to a different question: how will the model do on my work. This guide explains the benchmarks vendors quote, including SWE-bench, MMLU, GPQA, ARC-AGI and the agent tests, with each one's size and method from its own paper. It covers three ways a score misleads: the questions were in the training data, every model scores near the top, and the vendor chose the settings. It ends with a seven-step method for choosing a model on your own cases.
Faster evals with Jev: how we grade AI-written code
An eval suite for AI-written code has two kinds of check. Some have an exact answer, and code decides them. The rest used to need a language model to read the change and give an opinion, which was the slowest and least repeatable step in the suite. Reveneau moved that step to Jev, a decision model that answers fixed questions with probabilities instead of writing text, and on our own suite the run now finishes ten times faster than it did with a language-model grader. This guide explains how the grader works, how to design its questions, and where it stops being useful.
Jev and System One models: a plain guide
Jev is a hosted model from TypeSafe AI that answers fixed questions about a piece of text with calibrated probabilities instead of writing a reply. TypeSafe calls this a System One model. It returns a choice, a score or a yes/no probability in well under a second, and TypeSafe prices it at $0.042 per million input tokens with output free. This guide explains what that buys you, where the model is known to fail, and how to decide between Jev and a language model for each decision in your software.
Jev in production: guardrails, routing and classification
A decision model answers a fixed question about a piece of text and returns a probability instead of a paragraph. Inside an application that makes it useful in four places: routing a request to a handler, gating an action on how sure the model is, checking a value that code already extracted, and screening what goes into and comes out of a language model. This guide covers each of those with Jev, TypeSafe's System One model, pattern by pattern, with the vendor's own figures marked as the vendor's and the limits stated at the start.
Should your product use a decision model? A guide for founders and CTOs
A decision model answers a fixed question about a piece of text with a probability, in under half a second, at a price TypeSafe lists as $0.042 per million input tokens. It writes nothing. That makes it the right tool for a narrow set of product features, the wrong tool for most others, and a governance question either way. This guide is written for the person who approves the decision: how to sort features, how to run a two-week pilot, what to ask a vendor or partner, and how to keep the model auditable once it is in production.
How to deliver high-quality software with AI-native development
AI now writes code fast, so writing code is no longer the slow step. The slow steps are the work around it: deciding what to build, proving it works, releasing it safely, and running it afterwards. This guide describes how a team should organise all four, at the speed AI now makes possible.
Taking AI agents from prototype to production
An agent that works in a demo and an agent that works in production are different systems. The demo has to succeed once, with someone friendly at the keyboard. Production has to succeed repeatedly, for people who did not read the instructions, with real money and real data affected by each call.
Product discovery and UX design: a buyer's guide
Most failed software was built competently. The team delivered what was asked for, on a reasonable timeline, and it still did not matter, because what was asked for was never tested against what users actually needed. Discovery and design are how you buy that test before you pay for the build.
Building compliant software in financial services
Financial regulation is already written as rules. That is the part most teams miss. A rule that says a record must be recreatable after deletion is a specification, and a specification can be turned into a check that runs on every change. This guide covers how to do that, which obligations actually apply to your code, and why compliance gaps in a working app are so easy to miss.
Building compliant software in healthcare
HIPAA is unusual among regulations in that it tells you what to achieve and mostly declines to tell you how. That flexibility is why so many healthcare applications are non-compliant while holding a completed risk assessment: the rule left the decision to you, and nobody wrote the decision down. This guide covers what the Security Rule requires from software, and how to turn it into something a system enforces.
Building compliant software in real estate
Real estate software has a compliance problem the other regulated industries do not. In finance and healthcare, the rules mostly govern how you handle data. In housing, the rules govern the decision your software makes, which means the model's output is itself the regulated act. That changes what you have to build and what you have to test.
Building compliant software in insurance
Insurance regulation is state by state, which most teams treat as a reason to guess. It is really a reason to build the rule into configuration rather than into code, so the same application can satisfy Iowa and Ohio without a special case for each. This guide covers what the NAIC model laws actually require, which parts apply to your application, and why a working system so often looks compliant to its own team while missing the one thing an examiner asks for.
Building compliant software in legal technology
Legal technology carries a rule most other software does not: the product itself can end up practicing law, even when nobody intended it to. This guide covers where that boundary is, what the ethics rules that govern lawyers actually require of the tools they use, which e-discovery obligations affect your data model, and how to turn each of those into something a build can be tested against rather than a policy nobody rereads.
Software for compliance-heavy industries
Regulated software has always needed exhaustive verification and rarely received it, because writing the checks cost more than anyone would approve. That constraint is gone. This guide covers what that changes, the obligations that apply regardless of sector, and how to tell a partner who verifies their work from one who says they do.
Forward deployed engineering: the complete guide
A forward deployed engineer is a software engineer who works inside the customer's environment, on the customer's real data, and stays responsible until the system runs in production. The role exists because software that works in a demo and software that keeps working inside a real company are two different pieces of work, and somebody has to do the second one.