AI NewsDev toolsAnnouncement

ethanplusai released astra-flash-orchestrator, a Codex skill that keeps GPT-6 Astra on planning and sends the code writing to DeepSeek V4.1 Flash

ethanplusai released astra-flash-orchestrator as an open-source Codex skill on 17 September 2026, under MIT. The repository reached 723 stars and 63 forks in two weeks. In one field build the author ran, keeping GPT-6 Astra on planning and sending implementation to DeepSeek V4.1 Flash cut total compute per 1,000 implementation lines from $11.32 to between $0.26 and $0.34.

AI News

Editorial3 min read

LinkedInX

Why it mattersThe expensive-model-plans-and-reviews, cheap-model-writes-the-code pattern is one every Codex and Claude Code user can copy, and this is the first popular open-source template that writes it down as a skill with a separate reviewer agent.

The cheapest way to spend less on a coding agent is to stop asking the expensive model to write a for loop. A Codex skill published on GitHub on 17 September 2026, astra-flash-orchestrator, writes that idea down as a repeatable workflow: GPT-6 Astra plans and reviews, DeepSeek V4.1 Flash writes the code, Astra accepts or rejects the patch in one batched pass at the end. By 1 October the repository had 723 stars and 63 forks under an MIT licence.

What the skill does

The author, who publishes under the handle ethanplusai, calls it "a personal Codex skill designed to preserve Astra usage without giving up Astra's judgment." Astra is responsible for scope, design, the task brief and the final review. Flash takes over for repository discovery, implementation, testing, debugging and routine verification. The developer who starts the task has to approve the plan Astra drafts, and then review Astra's acceptance at the end.

The structure matters as much as the model choice. The skill installs a named astra_flash_builder role in a Codex session, so a bot that is only running under the orchestrator talks to Flash through a fixed route; "unrelated subagents keep their existing defaults," as the installer puts it. The workflow has one rule the README states twice: a separate review agent checks the patch before Astra accepts, so the agent that wrote the code never grades its own work.

The numbers the author published

A table in the repository compares the two shapes on one measured field build. All-Astra used 8.56 million Astra input tokens per 1,000 implementation lines and $11.32 of total compute per 1,000 lines. The Astra plus Flash workflow used 95,900 Astra input tokens per 1,000 lines, 98.9 percent less, and $0.26 to $0.34 of total compute, between 97.0 and 97.7 percent less. The measured phase produced 39 percent more implementation and test lines.

Those are the author's own figures from one build, and the README is honest about what that means. The status line says the results describe that one run and come with no guarantee. Astra has no public API price today, so the per-token comparison uses API-equivalent estimates against the published DeepSeek rates, and the author labels every Astra figure as an estimate derived from those rates. The per-token difference drives most of the saving: Astra's estimated $10 per million uncached input against Flash's $0.15 to $0.30, which the table marks as a 33 to 67 times premium.

What a reader needs before trying it

The requirements list is short. A Codex client with native subagent support, GPT-6 Astra as the root model, Python 3.11 or newer with no third-party dependencies, an installed Codex Router configured for one reviewed DeepSeek V4.1 Flash route, and a local Codex model catalog that advertises that route with multi_agent_version: "v2". If any piece is missing, the installer stops and reports the gap rather than picking a substitute. Supported Flash providers include DeepSeek direct, OpenRouter, opencode Go, Command Code, Nous Research and Ollama Cloud.

The author also states one honest limit about the project itself: it is an early release, and the measured phase above is one local field build. 723 stars in two weeks is real traction for a solo author's first popular release. A team adopting it on a work machine should expect bugs in the first week, and should check which provider and model the first real task sends work to before trusting the routing.

Source

SourceGitHub

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX