Asana says its browser agent runs 76x cheaper on GPT-6.1 Sol

Asana says its browser agent runs 76x cheaper on GPT-6.1 Sol

Asana helps customers automate work across business applications through StackAI, a platform it acquired. Customers use it to build no-code workflows that navigate websites, fill out forms and gather information. At Asana's scale, small inefficiencies in those workflows add up, so StackAI CTO Frank Hidalgo, PhD, set out to make the browser agent faster and cheaper to run. He directed GPT-6 Astra in Codex to investigate the agent, test improvements and compare the results. Work he estimates would have taken one to two months by hand took about a week.

The headline result comes from a 144-run study that tested GPT-6.1 Sol and three other frontier models, anonymized as Models A, B and C. The optimized workflow on GPT-6.1 Sol averaged $0.47 in estimated model cost and about four minutes per run. Asana states that is 76x cheaper and 5x faster than the original production setup on Model B.

The diagnosis came first. Hidalgo had GPT-6 Astra map the codebase and explain how the agent built each model request. It found that the agent cached its fixed instructions and tool definitions but not the growing history of page text and screenshots, so every request resent that history at full price. The agent also dropped older screenshots and trimmed text at nearly every step. Each edit altered the history, so caching the history alone would not have helped, and losing those facts could force the agent to revisit pages it had already read.

Hidalgo reviewed the proposed fixes and picked three to test: extend caching to the browsing history, raise the amount of text the agent can retain, and remove screenshots in batches rather than at every step. Because the code was not designed for controlled experiments, GPT-6 Astra first ran quick tests to find which variables mattered, then refactored the code so one frontend and backend could support many workflows in parallel, each with its own settings.

The full study used history budgets of 120,000 and 480,000 characters and six caching and screenshot policies, each tested three times on each of the four models. The best policy let screenshots accumulate to 20 before cutting back to the most recent one, which kept earlier history unchanged for longer stretches. Combined with the larger history budget, it became the optimized workflow. Every configuration did the same task: collecting six fields for each of 32 books from a public demo catalog, which Asana says is representative of what some of its customers run in StackAI.

The numbers break down further. For Model B, optimization cut estimated cost from at least $36.21 per run (some original runs hit the step limit before finishing) to $1.24, a 29x reduction. The optimized GPT-6.1 Sol workflow was 2.6x cheaper still, at $0.47. On GPT-6.1 Sol alone, with the larger history budget, the new caching and screenshot policy cut cost 4x, from $1.97 to $0.47 per run. Each call was about 3x cheaper because 89% of the input came from cache at 5% of the uncached price. Run time fell from at least 22.5 minutes on the original Model B setup to roughly four minutes. Every run in the optimized workflow completed the task and returned the correct answer.

History management also decided whether the agent produced an answer at all. With the smaller history budget, three of 18 GPT-6.1 Sol runs produced an answer; with the larger budget, all 18 did, each correct.

GPT-6 Astra ran the workflows, examined requests, usage records and outputs, and separate model sessions reviewed the work. Every session's requests, data traces and results were recorded in Command, Asana's software delivery platform, and the findings were turned into tickets, then pull requests, then production changes. Asana has released the changes to browser navigation in StackAI and is building tools to make such experiments easier to repeat. Over time it plans to fold this testing into the platform's evaluations so customers and internal teams can compare cost, runtime and answer quality when configuring agents. Asana also now uses GPT-6 Astra in Codex to test product features before release: Astra navigates the platform, tries different inputs and reports bugs for human QA reviewers. The complete study is on the Asana and StackAI blogs.

Key facts

  • In a 144-run study across GPT-6.1 Sol and three anonymized models, the optimized GPT-6.1 Sol workflow averaged $0.47 in estimated model cost and about four minutes per run, which Asana states is 76x cheaper and 5x faster than the original production setup on Model B.
  • GPT-6 Astra in Codex found that the agent cached fixed instructions and tool definitions but not the growing page and screenshot history, so every request resent that history at full price.
  • On GPT-6.1 Sol alone the new caching and screenshot policy cut cost 4x, from $1.97 to $0.47 per run, with 89% of input served from cache at 5% of the uncached price.
  • Frank Hidalgo, Asana's StackAI CTO, estimates the work would have taken one to two months by hand; it took about a week.
  • Asana has released the browser navigation changes in StackAI and plans to build this kind of testing into the platform's evaluations.

Why it matters

Asana says cost used to limit which models it could offer customers for these workloads, so making the agent cheaper widens the choice. The case is also a worked example of a human directing a coding agent through a full optimization loop: diagnosis, controlled experiments, review, tickets, pull requests and production. The cost lever it found is concrete. Resending an uncached, constantly edited history on every request was the expense, and changing when screenshots are removed let the cache do its job.

Who it affects

Most directly, StackAI customers who build no-code workflows that navigate websites, fill out forms and gather information, and the Asana engineers who maintain the browser agent. More broadly it speaks to any team running a long-lived browser agent whose history of page text and screenshots grows with every step. The source gives no figure for how many customers are affected.

How to use it

There is nothing to buy here; the post describes a method and a shipped change. The method: map how each model request is assembled, check what is cached and what is rewritten between steps, then test history budgets and caching and screenshot policies under controlled, repeated runs. Asana's best policy let screenshots accumulate to 20 before cutting back to the most recent one, paired with a 480,000-character history budget. The changes to browser navigation are already released in StackAI. Asana is developing tools to make such experiments repeatable and plans to include cost, runtime and answer-quality comparisons in the platform's evaluations, with no timeline given. The complete study is on the Asana and StackAI blogs.

How solid is it

The figures are Asana's own, reported in the source article, and all costs are estimated model costs rather than billed invoices. The study is 144 runs, with each configuration tested three times per model, on one task: six fields for 32 books from a public demo catalog. Every session was recorded in Command so the team could review the full study afterward, and separate model sessions reviewed the work. The headline 76x compares GPT-6.1 Sol optimized against Model B original, and the identities of Models A, B and C are not disclosed.

Risks and caveats

The 76x figure mixes a model change with a workflow change. Within the same model, Model B improved 29x and GPT-6.1 Sol alone improved 4x. The original Model B cost ($36.21) and run time (22.5 minutes) are lower bounds, because some runs hit the step limit. The task was a single benchmark representative of what some customers run, and no production-wide savings figure is given. The quoted passages in the post are not attributed to a named speaker. Plans to build the testing into platform evaluations are plans, with no date. The sensitivity to history budget is itself a warning: with the smaller budget only three of 18 GPT-6.1 Sol runs produced an answer.

“Cost used to limit which models we could offer customers for these workloads.”

— Quoted in the source article on Asana's StackAI browser agent