Asana cut its browser agent's model cost 76 times using Codex
OpenAI's case study on Asana reports that GPT-6 Astra inside Codex ran a 144-run experiment on Asana's StackAI browser agent, and that the resulting workflow on GPT-6.1 Sol averaged $0.47 per run and ran 5 times faster than the production baseline.

Image: OpenAI
Why it mattersA browser agent that resends its growing history on every request pays full price for the same screenshots many times, and a controlled experiment on cache policy is the kind of change one engineer plus a coding model can run in a week.
A browser agent that keeps gathering page text and screenshots pays for the same material on every single request, because a coding model treats a changing history as new input and bills for it at full price. Asana's StackAI team asked GPT-6 Astra, running inside OpenAI's Codex, to find out whether that was happening to them, and OpenAI published the result on 9 October.
The published numbers are 76 times cheaper and 5 times faster than the version Asana was running in production. The optimised workflow, on GPT-6.1 Sol, averaged $0.47 in model fees per task, OpenAI said. Frank Hidalgo, CTO of StackAI (the browser-automation platform Asana acquired), told OpenAI that work he expected to take one to two months by hand took one week.
What the experiment did
GPT-6 Astra read the StackAI code and found that the agent cached its fixed instructions and tool definitions, but never cached the growing record of page text and screenshots it collected. Every request resent that history at full price. The agent also trimmed old screenshots and shortened text at nearly every step, so caching the history on its own would not have worked.
Hidalgo picked three fixes from GPT-6 Astra's list: extend caching to the browsing history, raise the cap on how much text the agent could keep, and remove screenshots in batches instead of at every step. GPT-6 Astra then refactored the code to let many settings run in parallel against one task.
The 144-run study
The task was the same for every run: collect six fields for each of 32 books from a public demo catalog, representative of what Asana customers run in StackAI. The team tested two history budgets (120,000 and 480,000 characters) and six combinations of cache and screenshot policy. Each setting ran three times on each of four models, GPT-6.1 Sol and three unnamed competitor models called Model A, B and C.
The winning policy let screenshots build up to 20 before trimming back to the most recent one. Paired with the larger history budget, it kept the earlier history unchanged for long enough that each request found the previous one in cache instead of billing it again. OpenAI says Model B, the model Asana ran in production before this, is priced the same as GPT-6.1 Sol.
The case study is OpenAI's, and the numbers come from OpenAI running tests on its own model with its own tooling. Reveneau has not reproduced them. What the method does show is one way to run the test: pick one task, keep it fixed, change one piece of context policy at a time, and record every run. An engineer who sets the question and lets a coding model run the experiments can measure how much caching saves on their own agent, instead of guessing.
Source
Asana cuts model costs 76x in browser tests with GPT-6.1 Sol, OpenAI.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


