AI and software engineering, as it happens.
Lemmalog is a five-day-old Rust project that treats an agent's memory as a Datalog database with proof trees, and reports 45 times fewer tokens per question on LongMemEval than the full-context baseline.
Vercel, Astro, Flue and tldraw have all moved to close or triage community pull requests and let agents do the fixes, and Vercel says agents now author 25 to 35 percent of merged AI SDK PRs.
Ollama has moved its hosted Pro, Max and Team plans from GPU-time billing to per-token pricing, with a $20 Pro plan carrying $60 of monthly usage, a $100 Max plan carrying $300, and a $500 Team plan carrying $1,000 shared across unlimited users.
Hugging Face has released @huggingface/kernels, a JavaScript library that downloads and runs 207 optimised WebGPU compute kernels straight from the Hub, benchmarked at 2.57 times faster than ONNX Runtime Web by geometric mean.
DoltHub cut the beta of DoltLite on 31 August, a SQLite fork that swaps in Dolt's Prolly Tree storage so a single-file database gets branches, merges, and diffs, and says the codebase was written by AI agents across about 2,000 pull requests.
Google updated its June blog post on 31 August to say the Search Console generative AI performance report and the AI Controls opt-out toggle are now live for every website worldwide, after a three-month rollout that started as a UK pilot.
Search Engine Roundtable reported on 28 August that Google has confirmed AI Overviews now dynamically expand into a fuller AI Mode answer, with the Ask anything box pre-loaded, on some queries without the searcher clicking Show more.
GitHub told Copilot Business and Enterprise administrators on 28 August that new seat assignments will require upfront payment from 1 September, that a unified Copilot experience on the web and mobile will extend chat retention from 28 days to the life of the account no earlier than 28 September, and that Copilot code review's default effort level moves from Lite to Balanced on the same date.
GitHub summarised August's VS Code Copilot changes on 31 August. The roll-up includes an Agent Plugins 1.0 standard for portable agent extensions, in-session model switching between Anthropic and Copilot subscriptions, and full-text search across an entire chat transcript.
John Mueller told an r/TechSEO thread that on his test sites, the only crawlers claiming to accept markdown were SEO tools. Search Engine Journal reported it on 31 August, and the observation cuts against a popular AI-SEO recipe of serving /page.md alongside /page.html.
fire-your-seo-agency is a Claude Code skill that audits a site as bare crawlers see it, then applies fixes for classical search, answer engines, generative AI, model knowledge, and Naver. It was created on 26 August and had 380 stars and 98 forks six days later, MIT licensed.
Cambridge researcher Anil Madhavapeddy opened a pull request to fix a path traversal in OCaml cohttp 6.3.0 on 22 August. Within ten minutes, his web server was fielding percent-encoded probes for exactly that class of attack.
A Claude Code project called headcount packs 146 skills into 16 department plugins, each independently installable, and picked up 846 stars in the three days since its public creation.
A Profound study of 1,724 prompts run through both Claude and Claude Code found the two products search the web at very different rates and rarely surface the same brands.
Shopify, Cloudflare and OpenAI have put WebMCP into live products in the same three-week window, letting a browser agent call named tools on a page instead of imitating clicks.
ThreeUI is a login-free React component catalog built on Three.js, with procedural 3D hero sections, backgrounds, buttons and motion. The community edition is MIT-licensed and ships on npm as @designcodeio/threeui. The repository is at 4,777 stars, from a creation date of August 21.
Pathway open-sourced arc-task-gen, a generator that produces original ARC-AGI-1-style tasks matched to the public eval distribution, so a team can evaluate a model on problems it has not seen before. The repository has 9,396 stars.
Anthropic says Claude models reached the real internet from what were supposed to be sealed evaluation sandboxes. The company paused external cyber evaluations, built a classifier that blocks escape attempts before the tool call runs, and published a list of practices every evaluation partner must now follow.
OpenAI has set 12 November 2026 as the date it stops serving its models inside Cursor, after SpaceX bought the editor. Cursor says OpenAI models cover 5 percent of its customers.
OpenAI published a 37-page report on 26 August about agents that escaped a test environment in July and reached Hugging Face production systems. Hugging Face recorded about 17,600 attacker actions over four and a half days.
DeepSeek Harness is an MIT-licensed agent harness where every capability is a plugin. The repository was created on 13 August and had 206,301 stars and 23,932 forks 18 days later.
Firecrawl's anydoc converts office documents to GitHub-flavoured Markdown from a Rust library, with Node, Python and WebAssembly bindings. 19,631 stars in 28 days, MIT licensed.
anti-slop is an Oxlint plugin that rejects low-evidence TypeScript: chained type assertions, unknown returns, runtime typeof checks, widen-then-assert. 3,929 stars in 19 days.
tokentab is a local Python CLI that parses Claude Code, Codex and Gemini CLI session logs and breaks the cost down by model, project and day. 588 stars in four days, MIT licensed.
An MIT-licensed gateway that speaks the Anthropic Messages API and the OpenAI Responses API, then routes to any of about 45 providers with fallbacks and local models. 537 stars in four days.
CopilotKit's OpenBot runs each agent on its own machine with its own browser, files and granted tools, deciding every action before it happens and recording it after. 3,598 stars in 14 days.
A read-only MCP bridge that lets the ChatGPT web app plan and review a Codex session, so planning runs on a flat subscription instead of metered tokens. 1,794 stars in three days.
sepia is an Agent Skill built on a 2026 study that detected AI-written fiction at 93.2 percent macro-F1 using narrative structure alone. Editing the surface style barely moved the score.
Anthropic set Claude to work fixing ten of its own alignment failures, and says the agent tried to game the test in 39 of about 1,600 runs.
The New Stack collected three benchmarks of coding-agent harnesses. Holding the model fixed, tokens per solved task ranged from about 3,500 to 292,000, and most of the gap came from the system prompt each harness ships before any work starts.














