AI NewsOpen sourceAnnouncement
CodeAF is a new open-source coding harness that targets open models and claims first place on DeepSWE
AgentField AI released CodeAF on 1 October 2026, an Apache-2.0 coding harness written in Go that targets cheap open models like DeepSeek, GLM, Kimi and Qwen, and published a benchmark putting it ahead of nine other coding agents on 113 DeepSWE tasks.
Image: GitHub
Why it mattersA team paying monthly for Claude Code or Cursor now has a self-hosted alternative with a published benchmark saying it solved more issues at a lower cost per solve on the same model.
A team wanting to drive its coding agent with a cheap open model instead of Claude has had no scoreboard to point at.
CodeAF, from AgentField AI, is an Apache-2.0 coding harness that launched on Product Hunt on 1 October 2026 and had 261 GitHub stars by 3 October. It is written in Go and ships as one 53 MB binary. The project targets open models through built-in connectors for DeepSeek, GLM, Kimi, MiniMax, Qwen, Ollama and any OpenAI-compatible endpoint, with Codex reachable through a ChatGPT plan. The repository also publishes a head-to-head benchmark against nine other coding agents.
What AgentField's benchmark says
The benchmark ran 113 real GitHub issues from the DeepSWE set through ten coding harnesses on the same model, DeepSeek V4 Flash, with each harness getting one attempt per task and a three-hour budget. Grades came from the official DeepSWE verifiers.
AgentField reports that its own senior-dev subharness solved 62 of 113 tasks, or 54.9 percent, at 22 cents per task. The next-highest harness in its table, mini-swe-agent, solved 56 at 38 cents. Codex solved 51 at 37 cents. Claude Code solved 16 at 19 cents, which the authors describe as 3.4 times the cost per solved issue. Nine harnesses were run on 11 September 2026 and senior-dev on 12 September.
AgentField states its own limits plainly: one seed per harness, so the gap between senior-dev and mini-swe-agent is not a statistically resolved difference; senior-dev departed from the sampling contract the other nine followed by not sending a temperature or top-p; and five tasks across four harnesses produced no verifier outcome and were counted as unsolved.
What the tool actually is
CodeAF describes itself as a "factory" rather than one more agent in one more terminal. One window holds every project on the machine. A conversation can hand out several tasks at once, each on its own branch in its own copy of the repository, with the harness picking a model crew of worker, planner and checker for each task from what the authors call Pareto Crewing. Tasks that pass their check land on a branch, never on main. Tasks that cannot be checked wait under an unread list for a one-key accept or reject.
The project also publishes a size and startup benchmark against the same rivals, measured by AgentField: 53 MB on disk against Claude Code's 224 MB, 27 MB of RAM per added session against Claude Code's 216 MB, and a 145-millisecond resume on a 50-turn session against Claude Code's 392 milliseconds. These are the vendor's own measurements on its own methodology, so treat them as the vendor's claims.
AgentField's homepage lists JPMorgan, ServiceNow, NIH, SAP and Microsoft among companies it says fork and run its stack in production, and the GitHub README calls CodeAF an early preview. Neither claim has been checked here.
Source
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.

