Strategy
Perspectives from our team on strategy.

What a buyer should do now that OpenAI has stopped reporting SWE-bench Verified
On 23 February 2026 OpenAI said it had stopped reporting SWE-bench Verified, and on 8 July 2026 it withdrew its recommendation of SWE-Bench Pro. This post gives OpenAI's figures with their limits, then shows how to choose a coding model on 100 of your own tasks.

What the Air Canada chatbot ruling means for a company with an AI assistant
A British Columbia tribunal treated a chatbot's answer as the airline's own statement and ordered a payment of $812.02. This post sets out what the decision says, what it leaves out, and the written test that compares an assistant's answers with the policy text.

How to negotiate a software contract you can actually verify
Most build contracts describe effort, timeline, and payment, and leave the one hard question unanswered: on what evidence do you agree the thing is finished?

Forward deployed engineer vs software engineer: which one to hire next
Both write production code to the same standard. Three questions about your next problem tell you which of the two engineers it needs.

The real price of an LLM judge is per decision, and it multiplies
A grader's price per call looks small. Multiply it by every check on every change, or every message through a guardrail, and the public figures from LangChain, Openlayer and TypeSafe show where a cheaper decision model pays for its own migration and where it does not.

How to size a team when generation is cheap
Headcount planning still assumes writing code is the expensive part. It is not anymore, and that changes which roles are actually scarce and how many people a project can take on before it slows down.

Why the second AI project is harder than the first
The first one had no users, no legacy data, and no opinions to satisfy. The second one meets all three at once, and the team reads the slowdown as their own failure rather than a change in the problem.

Who pays for the last mile in enterprise AI?
Wonderful's 52 percent gross margin puts a price on getting AI into production for the first time at scale, and there are only three ways that cost ever gets paid.

Ask an AI vendor for its time to production, and skip the demo
Each of Wonderful's customer stories starts with the time it took to reach production. That time is the better proof, because it measures the vendor, while a demo only measures the model.

The Wonderful story: how a 50-person Hebrew voice-agent team became a $5 billion company in 20 months, and what everyone says about it
The fullest public account we could assemble of Wonderful, from a Tel Aviv seed round to a $5 billion valuation, with what its investors, the press, its critics and its own engineers say about it, and every claim sourced.

The Mechanical Orchard story: how a company built on pair programming came to argue against code review
Rob Mee built Pivotal Labs on pair programming and human review. The company he founded in 2022 now publishes an essay saying code must not be reviewed by humans, and the public record shows how it got there.

The Fluxon story: a bootstrapped product studio built by ex-Googlers
The fullest public account we could assemble of Fluxon, the San Francisco product development company that took no outside money, grew to 150 people in five years by its own account, and now presents itself as the build partner of the top AI labs, with every claim sourced and every self-declared figure labelled.

Wonderful went from zero to a $5 billion valuation in 20 months. Here is what it proves about enterprise AI
The fastest-growing enterprise AI company of the past two years sells deployment, and its investors paid for the final work of getting AI into production, which tells you where the value in AI now is.

How to choose what not to automate
Most automation decisions are made by asking whether a task can be automated. That question has been answered yes for almost everything, which means it has stopped being useful.

The EU AI Act deadline that moved, and what to do with the time
Teams spent two years planning around August 2026. The high-risk obligations were deferred shortly before it arrived. The useful response is to notice which parts of the plan were only ever about the date.

How to budget for AI coding tools without guessing
Seat licences are the small number. The real budget line is the review capacity you need to handle what the tools produce, and almost nobody puts that on the spreadsheet.

When speed is the wrong goal
We sell speed. It is still the wrong target on four kinds of work, and knowing which four is worth more than another week removed from a timeline.

Procurement for AI-built software
Standard software contracts were written for an era when a person typed every line. Six clauses need rewriting when a machine writes most of them, and one of them is about who owns the output.

The questions to ask before approving an AI build
You are being asked to approve a build where most of the code will be generated. You do not need to read the code. You need nine questions and the confidence to keep asking until you get a specific answer.

The real cost of shipping unverified code
The cost of unverified code does not arrive as a bug report. It arrives as a codebase nobody will touch, a review queue that never empties, and a team that has stopped trusting its own pipeline.

How to prepare for technical due diligence before a raise or sale
Technical due diligence is where a deal can quietly fail. Here is what investors' technical reviewers actually look at, how to prepare before they do, and the warning signs that worry them.

Who signs off on AI-written code?
A test can only enforce a rule somebody already thought of. When AI writes most of the code, the question that decides whether a codebase stays trustworthy is who is accountable for the rule that was missing.

How to estimate a software project honestly
A single-number estimate on new work is a promise nobody can keep. The honest version is a range that reflects real uncertainty, and the biggest cause of a bad estimate is risk nobody planned for.

Stop paying for headcount you no longer need
An hourly quote is a price for people and months. When most of the implementation is generated, that is a price for an input that has largely gone, and it quietly rewards your supplier for being slow.

In-house vs outsourced engineering: what to keep and what to give to others
The real question is not whether to build in-house or outsource. It is which parts belong to your own team forever, and which parts an outside team can do faster and better right now.

How to handle a software project that has failed
A failing software project is common, not rare. What separates a costly loss from a cheap lesson is how honestly you run the post-mortem and how clearly you decide to fix, restart, or stop.

Development agency vs. freelance marketplace: what you are actually buying
A curated freelance marketplace sells you access to individual people. An agency sells you a team that is accountable for the outcome. They solve different problems, and choosing on price alone leads to the wrong choice.

Staff augmentation vs a managed team that owns the outcome
Adding individual engineers to your team and hiring a team that owns an outcome look similar on an invoice. They are not the same thing, and picking the wrong one hides a large management cost you did not budget for.

How to manage an external development team so it feels like your own
Managing an outside team goes wrong when you count hours instead of outcomes. Here is how to guide a team that delivers, instead of closely supervising one that makes no progress.

Offshore vs nearshore vs onshore: why the cheapest hourly rate costs the most
People pick a development team by hourly rate and location, then wonder why the cheap option cost the most. What actually drives the outcome is seniority, communication, time-zone overlap, and ownership.

The early warning signs your software project is in trouble
A software project rarely fails in one dramatic moment. It falls behind slowly, and the early signs are easy to explain away. Here are the ones to watch and what to do about each before it is too late.

How to write a software development RFP that gets honest bids
Most software RFPs ask for a fixed price on a problem nobody has scoped yet, so the bids come back either set too high on purpose or dishonest. Here is how to write one that gets real answers.

Technical due diligence for VC portfolio companies
Before you invest, you need a clear understanding of the code, the team, and the risk behind it. Here is what a real technical due diligence review covers.

Software development partner vs. vendor: what is the real difference?
A vendor builds what you ask for. A partner tells you when what you asked for is wrong. The difference shows up months after the contract is signed.

Software consulting firm vs. development agency: what's the actual difference?
A consulting firm sells you a recommendation. A development agency sells you working software. Confusing the two can cost you three months you did not plan to spend.

Red flags when hiring a software development agency
The wrong agency costs you months, not just money. Here are the warning signs worth checking before you sign a contract, not after.

Fractional CTO vs. software development agency: which do you need?
One owns your technical decisions. The other builds what you already decided. Most founders need to know which problem they actually have before they hire either.

Speed of execution is the moat now
Being first used to be an advantage you could protect. When any capable team can build the same thing in a week, the advantage goes to the team that releases, learns, and releases again fastest.

Build vs buy: when custom software is actually worth it
We build custom software for a living, and we still tell most people to buy the tool. Here is how to know when building is worth it and when it is a costly mistake that shows up slowly.

Knowing what to build is now the most important decision
AI made writing software cheap. That moved the hard part earlier, to deciding what deserves to be built and having the discipline to cut the rest.

Can you trust an AI agent with real work yet?
An agent that answers a question and an agent that takes an action are not the same risk. Here is how we decide where an agent is ready to act, and where it is not.

An AI demo is not a product
A convincing AI demo takes an afternoon. Turning it into something people trust in production takes most of the work, and most failures happen at that stage.

Why a small senior team now outbuilds a big one
Adding people used to be how you went faster. With modern tools, a small team of senior engineers often releases more work, with fewer problems, than a large mixed one.

You cannot measure engineers by how much they produce
Lines of code, tickets closed, and hours logged all measure activity rather than progress. Here is how we think about engineering output without numbers that only look good.

What custom software actually costs when writing the code is cheap
Software used to be priced by how much of it there was. That was never a good measure and it is now a bad one. The cost of a build now depends on two things: how clearly you can say what you want, and how hard it is to prove you got it.

How to choose the right external development team
Hiring an outside team is a decision with serious consequences. The wrong partner costs you time and progress you cannot get back.

How to launch a product with minimal resources
A tight budget forces good decisions. It makes you focus on the value that actually matters and cut everything that does not.