Engineering
Perspectives from our team on engineering.

What to do with the code nobody understands
Every system has a part everyone avoids. The instinct is to rewrite it or ignore it, and both are wrong for the same reason: you cannot tell yet which of its behaviour is deliberate.

How to plan a migration you cannot pause
The single, all-at-once cutover is popular because it is easy to describe and easy to schedule. It is also the version where you find out whether it worked at the moment you can least afford to be wrong.

An AI agent that passes a test once can fail it on the next run
The same agent on the same task passes on some runs and fails on others, and the tau-bench paper of 2024 measured by how much. Run each test case several times and report the share of tasks that passed every run.

Why flaky tests are worse when agents write them
A flaky test used to cost a rerun and some irritation. When an agent is reading the result to decide whether it is finished, an unreliable check becomes an unreliable instruction.

What a 90 percent pass rate on 50 test cases tells you
A pass rate of 90 percent measured on 50 test cases fits a true rate anywhere from 81.7 to 98.3 percent, so a result of 92 percent on the same 50 cases cannot be told apart from it. This post shows the arithmetic step by step and lists what to report next to every pass rate.

How to test a feature you cannot fully specify
Some features have no single correct output. A summary, a ranking, a suggested reply. You cannot check those against one exact expected answer, and the usual conclusion, that they cannot be tested, is wrong.

What a good handover document actually contains
Most handover documents describe the architecture, which is the one thing the next team can work out for themselves. The valuable parts are the decisions, the unknowns, and the things only one person knew.

Why observability matters more when a machine wrote the code
When a person writes a system, someone carries a mental model of it. When a model writes it, nobody does, and production becomes the only place where you can see what the system actually does.

What we checked before grading with a week-old model
Jev launched on 15 September 2026 and we put it in charge of grading our eval suite the same month. Here is the order of checks we ran first, written as a method you can repeat, and the one outcome we are willing to state.

How to decommission a feature safely
Removal is the only thing that actually reduces maintenance load, and it is the only engineering work nobody ever schedules. Here is how to find what nobody uses, spot the dependencies you cannot see, and take it out in steps you can reverse.

How to handle a security finding in generated code
A scanner reports a problem in code no human wrote. The first instinct is to fix that line, and that is the one response almost guaranteed to leave the same defect in nine other places.

What to do with a grader that gives no reason
A Jev grade is a probability with no paragraph. That is enough for a check you wrote and calibrated yourself, and wrong for a new kind of failure, a design judgment, or anything a regulator wants explained in words.

The review queue is the new bottleneck
When code got faster to write, the slowest step became the one place nobody measures. A review queue is invisible for about a quarter, and by the time you can see it the fastest-moving work is also the least read.

Our eval suite now runs ten times faster. Here is what we changed.
This month we replaced the language model that graded the judged checks in our eval suite with Jev. The deterministic checks did not move, every rubric is now written inside the check, and the suite runs ten times faster.

What belongs in a runbook nobody reads
Most runbooks are written for a calm reader who has time. The person who opens one is tired, frightened, and has about ninety seconds, and almost nothing in a normal runbook is useful to that person.

What a rollback plan looks like for AI-written features
Reverting the commit is the easy half. The half that causes real problems is the data the feature already wrote, and nobody plans for that until they have to.

A $5 billion AI company banned manual coding. Here is what it had to build first
Wonderful's chief architect published what happened after the company banned hand-written code, and the useful part is the tooling that had to exist before the speed appeared.

The cost of a dependency you did not choose
A coding agent adds a package in four seconds and the decision is never discussed. Someone will be responsible for that package for the next five years, and it will not be the agent.

How to write a bug report an agent can act on
A bug report written for a colleague relies on everything that colleague already knows. Give the same three lines to an agent and it will fix something, confidently, that nobody asked it to change.

How to run a postmortem when AI wrote the code
The incident review reaches the line that caused the outage, and the answer to "why is it like that" is that a model produced it and a person approved it in four minutes. Most postmortem formats have no place to record that.

What to do when the AI writes the wrong abstraction
A wrong abstraction does not fail a test or break a build. It makes the next twenty changes harder without anyone noticing, and by the time someone does, removing it has become a project of its own.

Why AI-generated code fails code review in a different way
Hand-written code fails review because it looks wrong. Generated code fails because it looks right, and human review is a process designed to catch the first kind.

How to take over a codebase you did not write
Someone hands you a working system and leaves. The instinct is to read it. The better first step is to find out what it guarantees, because the code will tell you what it does and never what it was supposed to do.

Why your staging environment is lying to you
Staging passed. Production broke. That is not bad luck, it is a predictable consequence of the five ways staging differs from the place your users actually are.

What an AI-native team actually looks like
An AI-native team is not a normal team with a licence for a coding assistant. The roles change, reviewing becomes the slowest step, and the job that grows is the one nobody has a title for yet.

How to write an acceptance test a machine can run
Most acceptance criteria are written for a human reader who will add the missing details. A machine adds nothing. Here is how to write the sentence so an automated check can enforce it.

Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.

Mobile web, native, or cross-platform: how to decide
Mobile web, native, or cross-platform is a real decision with real trade-offs. But "we need an app" is often the wrong place to start. Here is how to decide honestly.

How many tests does AI-generated code need?
The honest answer is not a number, and coverage percentages are the wrong unit. Here is the unit we use instead, and why a suite of fifteen checks can be worth more than two thousand.

A practical pre-launch security review for a small team
You do not need perfect security to launch. You need to check the few basics that find most real problems, and to know when the risk is big enough to bring in a specialist.

AI is an amplifier, not a fix
The 2025 DORA data says AI raises throughput and hurts stability at the same time. Which of those two you get is decided by what has to pass before a change is released, and by nothing else.

Adding eval tests was the best decision we made
We generate every line of code we release. The single change that made that safe was writing the check before the code, and it turned out to change how we review, how we spec, and how fast we can work.

Monolith vs microservices for an early-stage startup
For almost every early-stage startup, one well-built monolith is the right first choice. Microservices solve problems you do not have yet, and they add cost you cannot afford to pay early.

How to set engineering OKRs that measure outcomes, not activity
"Release five features this quarter" is a to-do list, not a goal. Good engineering OKRs measure whether users and the business are better off, which is a much harder and more useful thing to write down.

How to choose a tech stack for a startup without costly mistakes
Most startups pick their stack for the wrong reasons. Boring and proven beats new and exciting almost every time. Here is how to choose for speed and hiring, not for a resume.

What technical debt really is, and when to pay it back
Technical debt is not messy code. It is a deliberate choice to give up some quality for speed. Here is how to tell smart debt from reckless debt, and when to pay each one back.

What a good technical spec looks like when a model writes the code
The old advice was to keep specs short and stop where writing the code is faster. That advice assumed a person was reading it. When a model writes the code, the cost of an unanswered question changes, and so does the right length of a spec.

When to rewrite vs refactor legacy code (and why the big rewrite is usually wrong)
The full rewrite feels clean and honest, and it is almost always the wrong instinct. Here is when a rewrite is genuinely justified and how to replace old code piece by piece instead.

How to get a working AI prototype in weeks, not quarters
Most AI ideas end during the planning stage. Here is how to show something real to users fast enough to know if the idea is worth the full build.

What good code review looks like when nobody wrote the code
With human code, the author is the first check and review is the second. With generated code, review is the only check. That one change alters most of what a reviewer should be doing.

How to move fast without breaking the product
Speed and quality are usually described as a trade-off. In practice, the teams that work fastest over a long time are the ones that made quality cheap to keep.

How to build mobile apps in today's dynamic environment
Mobile engineering has its own set of parts to manage: feature flags, release cycles, and performance budgets that change as your app grows.

Migrating an app from AWS to GCP
A cloud migration is rarely just a lift and shift, meaning a copy of the system to the new cloud with no changes. The decision affects cost, reliability, and the day-to-day experience of your engineers.