Latest

Engineering

Perspectives from our team on engineering.

What to do with the code nobody understands
EngineeringOct 4, 2026

What to do with the code nobody understands

Every system has a part everyone avoids. The instinct is to rewrite it or ignore it, and both are wrong for the same reason: you cannot tell yet which of its behaviour is deliberate.

How to plan a migration you cannot pause
EngineeringOct 3, 2026

How to plan a migration you cannot pause

The single, all-at-once cutover is popular because it is easy to describe and easy to schedule. It is also the version where you find out whether it worked at the moment you can least afford to be wrong.

An AI agent that passes a test once can fail it on the next run
EngineeringOct 3, 2026

An AI agent that passes a test once can fail it on the next run

The same agent on the same task passes on some runs and fails on others, and the tau-bench paper of 2024 measured by how much. Run each test case several times and report the share of tasks that passed every run.

Why flaky tests are worse when agents write them
EngineeringOct 2, 2026

Why flaky tests are worse when agents write them

A flaky test used to cost a rerun and some irritation. When an agent is reading the result to decide whether it is finished, an unreliable check becomes an unreliable instruction.

What a 90 percent pass rate on 50 test cases tells you
EngineeringOct 2, 2026

What a 90 percent pass rate on 50 test cases tells you

A pass rate of 90 percent measured on 50 test cases fits a true rate anywhere from 81.7 to 98.3 percent, so a result of 92 percent on the same 50 cases cannot be told apart from it. This post shows the arithmetic step by step and lists what to report next to every pass rate.

How to test a feature you cannot fully specify
EngineeringOct 1, 2026

How to test a feature you cannot fully specify

Some features have no single correct output. A summary, a ranking, a suggested reply. You cannot check those against one exact expected answer, and the usual conclusion, that they cannot be tested, is wrong.

What a good handover document actually contains
EngineeringSep 30, 2026

What a good handover document actually contains

Most handover documents describe the architecture, which is the one thing the next team can work out for themselves. The valuable parts are the decisions, the unknowns, and the things only one person knew.

Why observability matters more when a machine wrote the code
EngineeringSep 28, 2026

Why observability matters more when a machine wrote the code

When a person writes a system, someone carries a mental model of it. When a model writes it, nobody does, and production becomes the only place where you can see what the system actually does.

What we checked before grading with a week-old model
EngineeringSep 28, 2026

What we checked before grading with a week-old model

Jev launched on 15 September 2026 and we put it in charge of grading our eval suite the same month. Here is the order of checks we ran first, written as a method you can repeat, and the one outcome we are willing to state.

How to decommission a feature safely
EngineeringSep 27, 2026

How to decommission a feature safely

Removal is the only thing that actually reduces maintenance load, and it is the only engineering work nobody ever schedules. Here is how to find what nobody uses, spot the dependencies you cannot see, and take it out in steps you can reverse.

How to handle a security finding in generated code
EngineeringSep 25, 2026

How to handle a security finding in generated code

A scanner reports a problem in code no human wrote. The first instinct is to fix that line, and that is the one response almost guaranteed to leave the same defect in nine other places.

What to do with a grader that gives no reason
EngineeringSep 25, 2026

What to do with a grader that gives no reason

A Jev grade is a probability with no paragraph. That is enough for a check you wrote and calibrated yourself, and wrong for a new kind of failure, a design judgment, or anything a regulator wants explained in words.

The review queue is the new bottleneck
EngineeringSep 24, 2026

The review queue is the new bottleneck

When code got faster to write, the slowest step became the one place nobody measures. A review queue is invisible for about a quarter, and by the time you can see it the fastest-moving work is also the least read.

Our eval suite now runs ten times faster. Here is what we changed.
EngineeringSep 23, 2026

Our eval suite now runs ten times faster. Here is what we changed.

This month we replaced the language model that graded the judged checks in our eval suite with Jev. The deterministic checks did not move, every rubric is now written inside the check, and the suite runs ten times faster.

What belongs in a runbook nobody reads
EngineeringSep 22, 2026

What belongs in a runbook nobody reads

Most runbooks are written for a calm reader who has time. The person who opens one is tired, frightened, and has about ninety seconds, and almost nothing in a normal runbook is useful to that person.

What a rollback plan looks like for AI-written features
EngineeringSep 18, 2026

What a rollback plan looks like for AI-written features

Reverting the commit is the easy half. The half that causes real problems is the data the feature already wrote, and nobody plans for that until they have to.

A $5 billion AI company banned manual coding. Here is what it had to build first
EngineeringSep 17, 2026

A $5 billion AI company banned manual coding. Here is what it had to build first

Wonderful's chief architect published what happened after the company banned hand-written code, and the useful part is the tooling that had to exist before the speed appeared.

The cost of a dependency you did not choose
EngineeringSep 16, 2026

The cost of a dependency you did not choose

A coding agent adds a package in four seconds and the decision is never discussed. Someone will be responsible for that package for the next five years, and it will not be the agent.

How to write a bug report an agent can act on
EngineeringSep 15, 2026

How to write a bug report an agent can act on

A bug report written for a colleague relies on everything that colleague already knows. Give the same three lines to an agent and it will fix something, confidently, that nobody asked it to change.

How to run a postmortem when AI wrote the code
EngineeringSep 11, 2026

How to run a postmortem when AI wrote the code

The incident review reaches the line that caused the outage, and the answer to "why is it like that" is that a model produced it and a person approved it in four minutes. Most postmortem formats have no place to record that.

What to do when the AI writes the wrong abstraction
EngineeringSep 10, 2026

What to do when the AI writes the wrong abstraction

A wrong abstraction does not fail a test or break a build. It makes the next twenty changes harder without anyone noticing, and by the time someone does, removing it has become a project of its own.

Why AI-generated code fails code review in a different way
EngineeringSep 9, 2026

Why AI-generated code fails code review in a different way

Hand-written code fails review because it looks wrong. Generated code fails because it looks right, and human review is a process designed to catch the first kind.

How to take over a codebase you did not write
EngineeringSep 7, 2026

How to take over a codebase you did not write

Someone hands you a working system and leaves. The instinct is to read it. The better first step is to find out what it guarantees, because the code will tell you what it does and never what it was supposed to do.

Why your staging environment is lying to you
EngineeringSep 4, 2026

Why your staging environment is lying to you

Staging passed. Production broke. That is not bad luck, it is a predictable consequence of the five ways staging differs from the place your users actually are.

What an AI-native team actually looks like
EngineeringSep 2, 2026

What an AI-native team actually looks like

An AI-native team is not a normal team with a licence for a coding assistant. The roles change, reviewing becomes the slowest step, and the job that grows is the one nobody has a title for yet.

How to write an acceptance test a machine can run
EngineeringSep 1, 2026

How to write an acceptance test a machine can run

Most acceptance criteria are written for a human reader who will add the missing details. A machine adds nothing. Here is how to write the sentence so an automated check can enforce it.

Never let the model grade its own work
EngineeringAug 25, 2026

Never let the model grade its own work

When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.

Mobile web, native, or cross-platform: how to decide
EngineeringAug 24, 2026

Mobile web, native, or cross-platform: how to decide

Mobile web, native, or cross-platform is a real decision with real trade-offs. But "we need an app" is often the wrong place to start. Here is how to decide honestly.

How many tests does AI-generated code need?
EngineeringAug 22, 2026

How many tests does AI-generated code need?

The honest answer is not a number, and coverage percentages are the wrong unit. Here is the unit we use instead, and why a suite of fifteen checks can be worth more than two thousand.

A practical pre-launch security review for a small team
EngineeringAug 20, 2026

A practical pre-launch security review for a small team

You do not need perfect security to launch. You need to check the few basics that find most real problems, and to know when the risk is big enough to bring in a specialist.

AI is an amplifier, not a fix
EngineeringAug 20, 2026

AI is an amplifier, not a fix

The 2025 DORA data says AI raises throughput and hurts stability at the same time. Which of those two you get is decided by what has to pass before a change is released, and by nothing else.

Adding eval tests was the best decision we made
EngineeringAug 20, 2026

Adding eval tests was the best decision we made

We generate every line of code we release. The single change that made that safe was writing the check before the code, and it turned out to change how we review, how we spec, and how fast we can work.

Monolith vs microservices for an early-stage startup
EngineeringAug 19, 2026

Monolith vs microservices for an early-stage startup

For almost every early-stage startup, one well-built monolith is the right first choice. Microservices solve problems you do not have yet, and they add cost you cannot afford to pay early.

How to set engineering OKRs that measure outcomes, not activity
EngineeringAug 18, 2026

How to set engineering OKRs that measure outcomes, not activity

"Release five features this quarter" is a to-do list, not a goal. Good engineering OKRs measure whether users and the business are better off, which is a much harder and more useful thing to write down.

How to choose a tech stack for a startup without costly mistakes
EngineeringAug 6, 2026

How to choose a tech stack for a startup without costly mistakes

Most startups pick their stack for the wrong reasons. Boring and proven beats new and exciting almost every time. Here is how to choose for speed and hiring, not for a resume.

What technical debt really is, and when to pay it back
EngineeringAug 5, 2026

What technical debt really is, and when to pay it back

Technical debt is not messy code. It is a deliberate choice to give up some quality for speed. Here is how to tell smart debt from reckless debt, and when to pay each one back.

What a good technical spec looks like when a model writes the code
EngineeringAug 2, 2026

What a good technical spec looks like when a model writes the code

The old advice was to keep specs short and stop where writing the code is faster. That advice assumed a person was reading it. When a model writes the code, the cost of an unanswered question changes, and so does the right length of a spec.

When to rewrite vs refactor legacy code (and why the big rewrite is usually wrong)
EngineeringJul 30, 2026

When to rewrite vs refactor legacy code (and why the big rewrite is usually wrong)

The full rewrite feels clean and honest, and it is almost always the wrong instinct. Here is when a rewrite is genuinely justified and how to replace old code piece by piece instead.

How to get a working AI prototype in weeks, not quarters
EngineeringJul 27, 2026

How to get a working AI prototype in weeks, not quarters

Most AI ideas end during the planning stage. Here is how to show something real to users fast enough to know if the idea is worth the full build.

What good code review looks like when nobody wrote the code
EngineeringJul 19, 2026

What good code review looks like when nobody wrote the code

With human code, the author is the first check and review is the second. With generated code, review is the only check. That one change alters most of what a reviewer should be doing.

How to move fast without breaking the product
EngineeringJul 9, 2026

How to move fast without breaking the product

Speed and quality are usually described as a trade-off. In practice, the teams that work fastest over a long time are the ones that made quality cheap to keep.

How to build mobile apps in today's dynamic environment
EngineeringMay 29, 2026

How to build mobile apps in today's dynamic environment

Mobile engineering has its own set of parts to manage: feature flags, release cycles, and performance budgets that change as your app grows.

Migrating an app from AWS to GCP
EngineeringMay 7, 2026

Migrating an app from AWS to GCP

A cloud migration is rarely just a lift and shift, meaning a copy of the system to the new cloud with no changes. The decision affects cost, reliability, and the day-to-day experience of your engineers.