<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0">
  <channel>
    <title>Reveneau AI News</title>
    <link>https://reveneau.com/ainews</link>
    <description>Short briefs on AI and software engineering: developer tools, open-source projects, model and agent releases. Each item links its original source.</description>
    <language>en-us</language>
    <item>
      <title>Google opens a Home MCP server so Claude, ChatGPT and other agents can run smart-home devices</title>
      <link>https://reveneau.com/ainews/google-home-mcp-server-early-access-claude-chatgpt-agents</link>
      <guid>https://reveneau.com/ainews/google-home-mcp-server-early-access-claude-chatgpt-agents</guid>
      <pubDate>Wed, 16 Sep 2026 21:20:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>Google has opened early access to two Model Context Protocol servers for Google Home, one for consumer agents like Claude and ChatGPT to control connected devices, and a second for coding agents that indexes Home programs, Matter and OpenThread docs. (Source: Google Home Developers)</description>
    </item>
    <item>
      <title>Rohan Bansal trained a 4B model to beat Postgres query plans by 1.81x on 113 join-heavy queries</title>
      <link>https://reveneau.com/ainews/rohan-bansal-qorl-4b-model-postgres-query-plans-1-81x-speedup</link>
      <guid>https://reveneau.com/ainews/rohan-bansal-qorl-4b-model-postgres-query-plans-1-81x-speedup</guid>
      <pubDate>Wed, 16 Sep 2026 20:30:00 +0000</pubDate>
      <category>Models &amp; agents</category>
      <description>Rohan Bansal published a write-up on 16 September of a home experiment where a 4-billion-parameter Qwen model, post-trained with supervised distillation and reinforcement learning, produced Postgres query hints that were 1.81 times faster than the database's default plans on the Join Order Benchmark, at a training cost of $1,200. (Source: Rohan Bansal)</description>
    </item>
    <item>
      <title>Zed opens Delta to the public and says its own team landed 570 changes without pull requests</title>
      <link>https://reveneau.com/ainews/zed-delta-public-beta-33-developers-570-changes-no-pull-requests</link>
      <guid>https://reveneau.com/ainews/zed-delta-public-beta-33-developers-570-changes-no-pull-requests</guid>
      <pubDate>Wed, 16 Sep 2026 20:20:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>Zed put its Delta collaboration tool into public beta on 16 September, and reports that 33 of its own developers have landed 570 changes to the Delta main branch since it turned off pull requests on the repo. (Source: Zed)</description>
    </item>
    <item>
      <title>Enclave says DeepSeek V4.1 Flash cleared its 11-target hacking benchmark for $4.65</title>
      <link>https://reveneau.com/ainews/enclave-deepseek-v41-flash-hacking-benchmark-11-of-11-465-dollars</link>
      <guid>https://reveneau.com/ainews/enclave-deepseek-v41-flash-hacking-benchmark-11-of-11-465-dollars</guid>
      <pubDate>Wed, 16 Sep 2026 19:30:00 +0000</pubDate>
      <category>Models &amp; agents</category>
      <description>The AI security firm Enclave says DeepSeek V4.1 Flash gained code execution on all 11 vulnerable targets in its hacking agent benchmark and left all four patched targets alone, at an accepted-run cost of $4.65 in API calls. (Source: Enclave AI)</description>
    </item>
    <item>
      <title>Cloudflare adds a Disallow AI Training setting that keeps Google, Apple and Bing crawling for search</title>
      <link>https://reveneau.com/ainews/cloudflare-disallow-ai-training-preserves-googlebot-applebot-bingbot-search</link>
      <guid>https://reveneau.com/ainews/cloudflare-disallow-ai-training-preserves-googlebot-applebot-bingbot-search</guid>
      <pubDate>Wed, 16 Sep 2026 19:20:00 +0000</pubDate>
      <category>Go-to-market</category>
      <description>Cloudflare has split its bot controls so a site can opt out of AI training while keeping Googlebot, Applebot and Bingbot crawling for search, reversing an earlier plan that would have blocked all three from Cloudflare-fronted sites that refused training. (Source: Cloudflare)</description>
    </item>
    <item>
      <title>How Stale Is Your AI tracks 20 model families and shows that only 10 of them publish a training cutoff</title>
      <link>https://reveneau.com/ainews/how-stale-is-your-ai-training-cutoff-20-models-jock-mackinlay</link>
      <guid>https://reveneau.com/ainews/how-stale-is-your-ai-training-cutoff-20-models-jock-mackinlay</guid>
      <pubDate>Wed, 16 Sep 2026 18:40:00 +0000</pubDate>
      <category>Models &amp; agents</category>
      <description>A one-page site by Jock Mackinlay tracks the release date and the training cutoff for 20 current frontier models side by side, and shows that only 10 of them publish a cutoff their lab actually documents. (Source: Jock Mackinlay)</description>
    </item>
    <item>
      <title>Bolt.new offers Pro users 50 times more coding compute if they share sessions to train a new open model</title>
      <link>https://reveneau.com/ainews/bolt-forge-50x-coding-compute-training-data-arcee-glm-5-3</link>
      <guid>https://reveneau.com/ainews/bolt-forge-50x-coding-compute-training-data-arcee-glm-5-3</guid>
      <pubDate>Wed, 16 Sep 2026 18:30:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>The New Stack reports that Bolt.new launched a research preview called Forge, which gives individual Pro subscribers up to 50 times more usage on open-weight coding models through 14 October, provided they let their sessions train a new trillion-parameter model with Arcee AI. (Source: The New Stack)</description>
    </item>
    <item>
      <title>Anthropic merges Claude chat and Cowork into one interface, and adds Docs and Slides</title>
      <link>https://reveneau.com/ainews/anthropic-merges-claude-chat-cowork-docs-slides-one-interface</link>
      <guid>https://reveneau.com/ainews/anthropic-merges-claude-chat-cowork-docs-slides-one-interface</guid>
      <pubDate>Wed, 16 Sep 2026 18:20:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>Anthropic has folded Cowork back into the main Claude chat, so the same conversation can now start a report, run a scheduled task, or draft a slide deck without switching interface, and Claude Docs and Claude Slides ship at the same time. (Source: Anthropic)</description>
    </item>
    <item>
      <title>ai-data-extractor pulls chat history off your machine as JSONL, and reads Claude Code, Cursor, Windsurf, Aider, Cline and eight others</title>
      <link>https://reveneau.com/ainews/ai-data-extractor-chat-history-claude-cursor-windsurf-aider-825-stars</link>
      <guid>https://reveneau.com/ainews/ai-data-extractor-chat-history-claude-cursor-windsurf-aider-825-stars</guid>
      <pubDate>Wed, 16 Sep 2026 17:50:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>The ai-data-extractor project is an MIT-licensed local tool that walks the disk for chat histories from Claude Code, Cursor, Windsurf, Aider, Cline and eight other coding assistants and normalises them into a single JSONL format, and it has picked up 825 stars in five days. (Source: kruzovic7 on GitHub)</description>
    </item>
    <item>
      <title>Google puts AI Performance Insights into Merchant Center and expands its Universal Commerce Protocol before the holiday rush</title>
      <link>https://reveneau.com/ainews/google-agentic-commerce-ucp-ai-performance-insights-holiday</link>
      <guid>https://reveneau.com/ainews/google-agentic-commerce-ucp-ai-performance-insights-holiday</guid>
      <pubDate>Wed, 16 Sep 2026 17:40:00 +0000</pubDate>
      <category>Go-to-market</category>
      <description>Google made AI Performance Insights generally available in five countries, opened a beta of a shopping Business Agent inside YouTube ads, and added cart transfer and checkout testing to its Universal Commerce Protocol integration in Merchant Center. (Source: Google)</description>
    </item>
    <item>
      <title>Ahrefs tracked 963 French domains before and after AI Overviews launched, and the most exposed lost 23.1 percent of their clicks</title>
      <link>https://reveneau.com/ainews/ahrefs-france-ai-overviews-ctr-drop-23-1-percent-963-domains</link>
      <guid>https://reveneau.com/ainews/ahrefs-france-ai-overviews-ctr-drop-23-1-percent-963-domains</guid>
      <pubDate>Wed, 16 Sep 2026 17:30:00 +0000</pubDate>
      <category>Go-to-market</category>
      <description>Ahrefs matched Google Search Console data on 963 French domains across the 28 days before AI Overviews launched on 22 July 2026 and the 9 days after, and found the median domain lost 5.7 percent of its clicks while the most exposed lost 23.1 percent. (Source: Ahrefs)</description>
    </item>
    <item>
      <title>Tokentab reads Claude Code, Codex and Gemini CLI logs and reports the bill by model, project and day</title>
      <link>https://reveneau.com/ainews/tokentab-cli-token-cost-claude-code-codex-gemini-1139-stars</link>
      <guid>https://reveneau.com/ainews/tokentab-cli-token-cost-claude-code-codex-gemini-1139-stars</guid>
      <pubDate>Wed, 16 Sep 2026 17:20:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>Tokentab is a new MIT-licensed CLI that reads the session logs Claude Code, Codex and Gemini CLI already write to disk and totals the token cost by model, project and day, and the repository has picked up 1,139 stars in nine days. (Source: crwdla on GitHub)</description>
    </item>
    <item>
      <title>Marcel Pociot wired Claude Code into Siri and Spotlight on macOS 27 through Apple's new model delegation provider API</title>
      <link>https://reveneau.com/ainews/mpociot-claude-siri-macos-27-model-delegation-provider</link>
      <guid>https://reveneau.com/ainews/mpociot-claude-siri-macos-27-model-delegation-provider</guid>
      <pubDate>Wed, 16 Sep 2026 15:20:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>A working proof of concept from Marcel Pociot shows Apple's macOS 27 model delegation API letting a third-party agent stand in as a Siri and Spotlight backend, with Claude Code as the model behind it. (Source: mpociot/claude-siri-ai)</description>
    </item>
    <item>
      <title>Salesforce releases Koa, a reasoning model built on Nvidia's open-weight Nemotron, for Agentforce</title>
      <link>https://reveneau.com/ainews/salesforce-koa-nvidia-nemotron-reasoning-agentforce-dreamforce</link>
      <guid>https://reveneau.com/ainews/salesforce-koa-nvidia-nemotron-reasoning-agentforce-dreamforce</guid>
      <pubDate>Wed, 16 Sep 2026 14:30:00 +0000</pubDate>
      <category>Models &amp; agents</category>
      <description>Salesforce announced Koa, its first CRM reasoning model, post-trained on Nvidia Nemotron for sales, marketing and customer support work inside Agentforce. (Source: Salesforce)</description>
    </item>
    <item>
      <title>gap-trap adds rules and gates to a repo so AI-written code fails the build when it breaks a rule</title>
      <link>https://reveneau.com/ainews/gap-trap-rules-gates-repo-ai-code-153-stars</link>
      <guid>https://reveneau.com/ainews/gap-trap-rules-gates-repo-ai-code-153-stars</guid>
      <pubDate>Wed, 16 Sep 2026 14:20:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>A new open-source skill installs rules and CI checks into a repository so agent-written code fails the build when it breaks a rule the team wrote down. (Source: pliablepixels/gap-trap)</description>
    </item>
    <item>
      <title>TypeSafe AI released Jev, a model that returns only typed structured values and cannot hallucinate</title>
      <link>https://reveneau.com/ainews/typesafe-jev-system-one-model-structured-decisions-no-hallucinations</link>
      <guid>https://reveneau.com/ainews/typesafe-jev-system-one-model-structured-decisions-no-hallucinations</guid>
      <pubDate>Wed, 16 Sep 2026 13:20:00 +0000</pubDate>
      <category>Models &amp; agents</category>
      <description>TypeSafe AI came out of stealth on 15 September and released Jev, its first System One Model. Jev returns only typed structured values, samples every output in parallel, and, by construction, cannot make a type error or hallucinate a field. (Source: TypeSafe AI)</description>
    </item>
    <item>
      <title>Irakli Betchvaia shows a Kotlin-embedded DSL cuts model hallucinations, with a curated examples tool lifting first-compile from 25 to 87.5 percent</title>
      <link>https://reveneau.com/ainews/typed-domain-grounding-dsl-hallucinations-kuml-compiler-oracle</link>
      <guid>https://reveneau.com/ainews/typed-domain-grounding-dsl-hallucinations-kuml-compiler-oracle</guid>
      <pubDate>Wed, 16 Sep 2026 12:20:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>In an InfoQ article published today, Irakli Betchvaia introduces Typed Domain Grounding, which embeds a home-grown DSL inside a mainstream typed language so the compiler catches model hallucinations that a lenient renderer would let through. (Source: InfoQ)</description>
    </item>
    <item>
      <title>Open Steps is a Claude Code skill pack that rewrites agent output in plain language, and it has 441 stars in 22 days</title>
      <link>https://reveneau.com/ainews/open-steps-claude-code-plain-language-reports-441-stars</link>
      <guid>https://reveneau.com/ainews/open-steps-claude-code-plain-language-reports-441-stars</guid>
      <pubDate>Wed, 16 Sep 2026 11:20:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>Open Steps is an MIT-licensed pack of eight Claude Code skills that make a coding agent report in plain words, ask questions a non-engineer can answer, and give a one-screen verdict when work is done. The repository has 441 stars and 75 forks. (Source: kharmanskyi on GitHub)</description>
    </item>
    <item>
      <title>Anthropic banned a Claude Code account 15 minutes after it was pointed at OpenAI's GPT-5.6 Sol through a proxy</title>
      <link>https://reveneau.com/ainews/anthropic-bans-claude-code-account-15-minutes-openai-gpt-5-6-sol-proxy</link>
      <guid>https://reveneau.com/ainews/anthropic-bans-claude-code-account-15-minutes-openai-gpt-5-6-sol-proxy</guid>
      <pubDate>Wed, 16 Sep 2026 10:30:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>The Information reports that Anthropic shut down developer Alex Getman's Claude Code account 15 minutes after he used a proxy to route it to OpenAI's GPT-5.6 Sol, and his posts about the ban drew a million views within hours. (Source: The Information)</description>
    </item>
    <item>
      <title>AWS releases a Step Functions pattern that lets agents propose and code validate before a booking or payment</title>
      <link>https://reveneau.com/ainews/aws-step-functions-bedrock-agentcore-agents-propose-code-validates</link>
      <guid>https://reveneau.com/ainews/aws-step-functions-bedrock-agentcore-agents-propose-code-validates</guid>
      <pubDate>Wed, 16 Sep 2026 10:20:00 +0000</pubDate>
      <category>Infrastructure</category>
      <description>AWS published a Step Functions pattern that puts Bedrock AgentCore agents inside a state machine, where every agent proposal is checked by a deterministic Lambda step before any reservation or payment call runs. (Source: AWS Compute Blog)</description>
    </item>
    <item>
      <title>Ordewell turns one goal into an ordered plan of coding-agent tasks, one model per task</title>
      <link>https://reveneau.com/ainews/ordewell-multi-runner-multi-model-plan-coding-agents</link>
      <guid>https://reveneau.com/ainews/ordewell-multi-runner-multi-model-plan-coding-agents</guid>
      <pubDate>Wed, 16 Sep 2026 10:15:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>Ordewell is a new Apache-licensed task orchestrator for coding agents that turns one goal into an ordered plan of tasks, each with its own runner and model, and completes a task only when a unique marker appears in the runner output. (Source: Ordewell on GitHub)</description>
    </item>
    <item>
      <title>An audit of one week's F-Droid updates finds 74 of 102 apps largely written by an LLM</title>
      <link>https://reveneau.com/ainews/f-droid-72-percent-102-apps-largely-llm-generated-tintotint-audit</link>
      <guid>https://reveneau.com/ainews/f-droid-72-percent-102-apps-largely-llm-generated-tintotint-audit</guid>
      <pubDate>Wed, 16 Sep 2026 09:45:00 +0000</pubDate>
      <category>Productivity</category>
      <description>A student going through every app in F-Droid's 12 September 2026 update batch rates 74 of 102 as largely written by an LLM, 19 as showing little to no AI, and 9 as hard to categorise. (Source: tintotint)</description>
    </item>
    <item>
      <title>Cloudflare's Automatic Key Exchange cuts origin HelloRetryRequest rates from 52 percent to 3.7 percent</title>
      <link>https://reveneau.com/ainews/cloudflare-automatic-key-exchange-hello-retry-52-3-7-percent</link>
      <guid>https://reveneau.com/ainews/cloudflare-automatic-key-exchange-hello-retry-52-3-7-percent</guid>
      <pubDate>Wed, 16 Sep 2026 09:30:00 +0000</pubDate>
      <category>Infrastructure</category>
      <description>Cloudflare turned on Automatic Key Exchange for TLS 1.3 origin connections, probing each origin to pick the strongest supported key agreement and cutting HelloRetryRequest rates from roughly 52 percent to 3.7 percent. (Source: Cloudflare)</description>
    </item>
    <item>
      <title>VS Code 1.138 lets an agent session run inside a project's Dev Container and open a pull request without leaving the Agents window</title>
      <link>https://reveneau.com/ainews/vscode-1-138-agent-dev-container-pull-request-codex-chatgpt-copilot</link>
      <guid>https://reveneau.com/ainews/vscode-1-138-agent-dev-container-pull-request-codex-chatgpt-copilot</guid>
      <pubDate>Wed, 16 Sep 2026 09:15:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>Microsoft shipped VS Code 1.138 with agent sessions that run inside a local Dev Container, pull request creation from the Agents window, and a model picker that switches between Copilot-backed and ChatGPT-backed models mid-conversation. (Source: Visual Studio Code)</description>
    </item>
    <item>
      <title>kotlin-footguns publishes 224 coding-agent skills mined from a shipping Kotlin app</title>
      <link>https://reveneau.com/ainews/kotlin-footguns-agent-skills-simpmusic-224-lessons</link>
      <guid>https://reveneau.com/ainews/kotlin-footguns-agent-skills-simpmusic-224-lessons</guid>
      <pubDate>Wed, 16 Sep 2026 07:25:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>kotlin-footguns is a public repository of 224 agent skills for Kotlin, Compose Multiplatform and the desktop JVM, extracted from a production music app rather than from official documentation. (Source: GitHub)</description>
    </item>
    <item>
      <title>Anthropic publishes a prompt engineering guide for Claude Fable 5.1 with fifteen behavioral changes</title>
      <link>https://reveneau.com/ainews/anthropic-fable-5-1-prompt-engineering-guide-behavioral-changes</link>
      <guid>https://reveneau.com/ainews/anthropic-fable-5-1-prompt-engineering-guide-behavioral-changes</guid>
      <pubDate>Wed, 16 Sep 2026 07:20:00 +0000</pubDate>
      <category>Models &amp; agents</category>
      <description>Anthropic published a Fable 5.1 prompting guide that lists fifteen behavioral changes from Fable 5, each paired with a specific prompt fix and, where relevant, an API beta header. (Source: Anthropic)</description>
    </item>
    <item>
      <title>Capsule packs an HTML app and its SQLite data into one file, and the Show HN got 314 points</title>
      <link>https://reveneau.com/ainews/capsule-single-file-html-sqlite-web-apps-hn-314-points</link>
      <guid>https://reveneau.com/ainews/capsule-single-file-html-sqlite-web-apps-hn-314-points</guid>
      <pubDate>Wed, 16 Sep 2026 06:30:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>A developer named bashtian shipped Capsule, a host player that packs an HTML app and its SQLite database into a single portable .capsule file, and the Show HN reached 314 points in 17 hours. (Source: bashtian on Hacker News)</description>
    </item>
    <item>
      <title>Amazon Science raises LLM-as-a-judge accuracy by 9 to 14 points by modelling how the judges copy each other</title>
      <link>https://reveneau.com/ainews/amazon-science-llm-judges-correlation-ising-model-9-14-percent</link>
      <guid>https://reveneau.com/ainews/amazon-science-llm-judges-correlation-ising-model-9-14-percent</guid>
      <pubDate>Wed, 16 Sep 2026 05:20:00 +0000</pubDate>
      <category>Models &amp; agents</category>
      <description>Amazon Science reports that panels of LLM judges often agree because they share training lineage or prompt templates, and that modelling those correlations with an Ising model raises aggregation accuracy by 9 to 14 points across three tasks. (Source: Amazon Science)</description>
    </item>
    <item>
      <title>Google is testing text link ads inside AI Mode answers, marked Sponsored above the response</title>
      <link>https://reveneau.com/ainews/google-ai-mode-text-link-ads-sponsored-test</link>
      <guid>https://reveneau.com/ainews/google-ai-mode-text-link-ads-sponsored-test</guid>
      <pubDate>Wed, 16 Sep 2026 04:35:00 +0000</pubDate>
      <category>Go-to-market</category>
      <description>Google has begun testing a new ad format inside AI Mode where sponsored text links sit above an AI-generated answer and match the look of the answer body, according to a report from Search Engine Roundtable. (Source: Search Engine Roundtable)</description>
    </item>
    <item>
      <title>GitHub Copilot adds three cost and quality tiers to auto model selection</title>
      <link>https://reveneau.com/ainews/github-copilot-auto-model-selection-three-tiers-efficiency-balance-intelligence</link>
      <guid>https://reveneau.com/ainews/github-copilot-auto-model-selection-three-tiers-efficiency-balance-intelligence</guid>
      <pubDate>Wed, 16 Sep 2026 01:00:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>GitHub released a change to Copilot auto model selection on 14 September 2026 that lets a developer set one of three tiers, efficiency, balance or intelligence, and each tier tells auto how to weigh cost, quality and response time on every prompt. (Source: The GitHub Blog)</description>
    </item>
    <item>
      <title>Meta releases WhatsApp Business Tools MCP for Claude, Cursor, Codex and ChatGPT</title>
      <link>https://reveneau.com/ainews/meta-whatsapp-business-tools-mcp-claude-cursor-codex-chatgpt</link>
      <guid>https://reveneau.com/ainews/meta-whatsapp-business-tools-mcp-claude-cursor-codex-chatgpt</guid>
      <pubDate>Wed, 16 Sep 2026 00:50:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>Meta shipped an MCP server on 15 September 2026 that lets Claude, Cursor, Codex or ChatGPT create a WhatsApp Business account, verify a phone number, register the Cloud API, and manage message templates by chatting with the agent. (Source: Meta for Developers)</description>
    </item>
    <item>
      <title>GitHub turns off SHA-1 in HTTPS on github.com and partner CDNs</title>
      <link>https://reveneau.com/ainews/github-sha-1-https-sunset-github-com-partner-cdns</link>
      <guid>https://reveneau.com/ainews/github-sha-1-https-sunset-github-com-partner-cdns</guid>
      <pubDate>Wed, 16 Sep 2026 00:40:00 +0000</pubDate>
      <category>Infrastructure</category>
      <description>GitHub disabled SHA-1 in the TLS handshake for github.com and partner CDNs on 15 September 2026, and old browsers, API clients or git binaries that only offered SHA-1 signatures for the TLS handshake will now fail to connect. (Source: The GitHub Blog)</description>
    </item>
    <item>
      <title>Google releases Gemini 3.8 Live and 3.8 Live Extended Thinking for voice agents</title>
      <link>https://reveneau.com/ainews/google-gemini-3-8-live-extended-thinking-voice-agents</link>
      <guid>https://reveneau.com/ainews/google-gemini-3-8-live-extended-thinking-voice-agents</guid>
      <pubDate>Wed, 16 Sep 2026 00:30:00 +0000</pubDate>
      <category>Models &amp; agents</category>
      <description>Google released two live dialogue models on 15 September 2026, one built for scale and one for multi-step reasoning, both available today in the Gemini API, Google Workspace and the Gemini app. (Source: Google DeepMind)</description>
    </item>
    <item>
      <title>IBM Research measures a 24-point consistency gap on AppWorld and halves it with automatic guidelines</title>
      <link>https://reveneau.com/ainews/ibm-altk-evolve-consistency-guidelines-pass-k-24-point-gap</link>
      <guid>https://reveneau.com/ainews/ibm-altk-evolve-consistency-guidelines-pass-k-24-point-gap</guid>
      <pubDate>Wed, 16 Sep 2026 00:15:00 +0000</pubDate>
      <category>Models &amp; agents</category>
      <description>IBM Research reports that a ReAct agent posting 77.4 percent on the AppWorld benchmark succeeds on all five repeated runs for only 53 percent of tasks, and that a diagnostic plus targeted guidelines cut the 24-point gap in half without hurting average accuracy. (Source: IBM Research on Hugging Face)</description>
    </item>
    <item>
      <title>Trail of Bits reanalyzes 1Password's AI patching study and finds agents block the exploit on 86 percent of fair trials</title>
      <link>https://reveneau.com/ainews/trail-of-bits-reanalyzes-1password-benchmark-86-percent-block-exploit</link>
      <guid>https://reveneau.com/ainews/trail-of-bits-reanalyzes-1password-benchmark-86-percent-block-exploit</guid>
      <pubDate>Wed, 16 Sep 2026 00:00:00 +0000</pubDate>
      <category>Models &amp; agents</category>
      <description>Trail of Bits reanalyzed 1Password's August report on AI patching, said the 26 percent clean-fix headline was misleading, and reported that models blocked the supplied exploit on 86 percent of trials where agents were allowed to run code and were not told to apply the wrong fix. (Source: Trail of Bits)</description>
    </item>
    <item>
      <title>Strix, an open-source autonomous pentest agent, found a live GitHub admin token for Baseten in 25 minutes</title>
      <link>https://reveneau.com/ainews/strix-autonomous-pentest-agent-baseten-github-token-25-minutes</link>
      <guid>https://reveneau.com/ainews/strix-autonomous-pentest-agent-baseten-github-token-25-minutes</guid>
      <pubDate>Tue, 15 Sep 2026 23:45:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>Strix, an open-source autonomous pentest agent from OmniSecure, was pointed at a Baseten subdomain with no credentials and returned in 25 minutes with a live GitHub admin token pulled out of a Docker build history from 2023. (Source: Strix)</description>
    </item>
    <item>
      <title>Specific Labs releases Real-SWE, a benchmark of proprietary code, and the top-scoring agent solves 38.8 percent of tasks</title>
      <link>https://reveneau.com/ainews/specific-labs-real-swe-benchmark-private-code-38-percent-top-score</link>
      <guid>https://reveneau.com/ainews/specific-labs-real-swe-benchmark-private-code-38-percent-top-score</guid>
      <pubDate>Tue, 15 Sep 2026 23:35:00 +0000</pubDate>
      <category>Models &amp; agents</category>
      <description>Real-SWE, a benchmark from Y Combinator-backed Specific Labs, drops coding agents into private codebases licensed from real companies, and the top-scoring setup solves fewer than 4 in 10 tasks. (Source: The New Stack)</description>
    </item>
    <item>
      <title>Cloudflare adds per-Worker access roles so a CI token or an agent can be scoped to one app instead of the whole account</title>
      <link>https://reveneau.com/ainews/cloudflare-workers-granular-authorization-four-roles-agents-ci-tokens</link>
      <guid>https://reveneau.com/ainews/cloudflare-workers-granular-authorization-four-roles-agents-ci-tokens</guid>
      <pubDate>Tue, 15 Sep 2026 14:40:00 +0000</pubDate>
      <category>Infrastructure</category>
      <description>Cloudflare Workers can now be scoped with four new roles, from Metadata Read-Only up to Admin, so a teammate, a CI job, or an AI agent gets access to one Worker rather than every Worker on the account. (Source: Cloudflare)</description>
    </item>
    <item>
      <title>Cloudflare lets sites refuse AI training without giving up search, and names Apple, Google and Microsoft as Accountable</title>
      <link>https://reveneau.com/ainews/cloudflare-accountable-ai-crawlers-apple-google-microsoft-training-search-split</link>
      <guid>https://reveneau.com/ainews/cloudflare-accountable-ai-crawlers-apple-google-microsoft-training-search-split</guid>
      <pubDate>Tue, 15 Sep 2026 14:30:00 +0000</pubDate>
      <category>Go-to-market</category>
      <description>A new Cloudflare setting publishes robots.txt directives that block AI training while leaving search alone, and Apple, Google and Microsoft qualify as Accountable operators that will honour it. (Source: Cloudflare)</description>
    </item>
    <item>
      <title>AIUC raises $40 million to audit AI agents, and Cursor, Lovable, Harvey and ElevenLabs are already customers</title>
      <link>https://reveneau.com/ainews/aiuc-40-million-series-a-agent-audit-standard-cursor-lovable-harvey</link>
      <guid>https://reveneau.com/ainews/aiuc-40-million-series-a-agent-audit-standard-cursor-lovable-harvey</guid>
      <pubDate>Tue, 15 Sep 2026 14:20:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>The Artificial Intelligence Underwriting Company raised a $40 million Series A led by Ribbit Capital and runs about 5,000 tests per audit against a SOC 2 style standard called AIUC-1. (Source: TechCrunch)</description>
    </item>
    <item>
      <title>AgentVerse OS bundles VS Code, Claude Code and Codex with 944 self-hosted apps on a single Ubuntu box, reached over Tailscale</title>
      <link>https://reveneau.com/ainews/agentverse-os-personal-cloud-tailscale-vs-code-claude-code-codex-944-apps</link>
      <guid>https://reveneau.com/ainews/agentverse-os-personal-cloud-tailscale-vs-code-claude-code-codex-944-apps</guid>
      <pubDate>Mon, 14 Sep 2026 23:40:00 +0000</pubDate>
      <category>Open source</category>
      <description>AgentVerse OS is an alpha personal cloud operating system for a developer and their coding agents on one Ubuntu server, reached from any browser through Tailscale, with isolated workspaces holding VS Code, Claude Code and Codex. (Source: GitHub)</description>
    </item>
    <item>
      <title>Perplexity brings Portable Computer to Windows, needing an Nvidia RTX GPU with at least 24GB of VRAM</title>
      <link>https://reveneau.com/ainews/perplexity-portable-computer-windows-rtx-24gb-vram-local-agent</link>
      <guid>https://reveneau.com/ainews/perplexity-portable-computer-windows-rtx-24gb-vram-local-agent</guid>
      <pubDate>Mon, 14 Sep 2026 23:35:00 +0000</pubDate>
      <category>Models &amp; agents</category>
      <description>Perplexity's Portable Computer, a local version of its Computer agent, is now available in the Perplexity Windows app on Nvidia GeForce RTX and RTX PRO GPUs, provided the card has 24GB of VRAM or more. (Source: NVIDIA Blog)</description>
    </item>
    <item>
      <title>Andon Labs opens Pion, its platform for running agents against real businesses, as a research preview</title>
      <link>https://reveneau.com/ainews/andon-labs-pion-agents-run-real-businesses-vending-bench</link>
      <guid>https://reveneau.com/ainews/andon-labs-pion-agents-run-real-businesses-vending-bench</guid>
      <pubDate>Mon, 14 Sep 2026 22:30:00 +0000</pubDate>
      <category>Models &amp; agents</category>
      <description>Andon Labs, the Y Combinator lab behind the Vending-Bench evaluation, has opened Pion as a research preview, a platform that gives one persistent agent the tools to run a real business: email, phone, banking, a browser and secure computing environments. (Source: Andon Labs)</description>
    </item>
    <item>
      <title>Entelligence benchmarks GPT-5.6 Luna against GPT-6 Astra on code review, and reports 74% precision at 28 times lower cost</title>
      <link>https://reveneau.com/ainews/entelligence-luna-astra-code-review-74-percent-precision-28x-cheaper</link>
      <guid>https://reveneau.com/ainews/entelligence-luna-astra-code-review-74-percent-precision-28x-cheaper</guid>
      <pubDate>Mon, 14 Sep 2026 22:25:00 +0000</pubDate>
      <category>Models &amp; agents</category>
      <description>Entelligence AI reran its code-review benchmark with the cheapest OpenAI model against the most expensive one on 50 public pull requests, and reports that GPT-5.6 Luna found three-quarters of the bugs GPT-6 Astra found for less than 4% of the money, with the gap concentrated on authentication and permission code. (Source: Entelligence AI)</description>
    </item>
    <item>
      <title>Patrick McCanna writes up moving 35 kB agent prompts from Anthropic and OpenAI to a self-hosted Ollama box</title>
      <link>https://reveneau.com/ainews/patrick-mccanna-35kb-preprompts-ollama-14-percent-context</link>
      <guid>https://reveneau.com/ainews/patrick-mccanna-35kb-preprompts-ollama-14-percent-context</guid>
      <pubDate>Mon, 14 Sep 2026 22:20:00 +0000</pubDate>
      <category>Productivity</category>
      <description>Patrick McCanna wrote up what happens when a 35 kB prompt that runs cleanly on Claude Code or Codex gets pointed at a self-hosted 27b open-weights model on a 128 GB Ryzen AI Max+ 395, and lists the concrete failure signals to watch for in the logs. (Source: Patrick McCanna)</description>
    </item>
    <item>
      <title>Apple's iOS 27 Siri has private frameworks that let Claude and GPT run as the model</title>
      <link>https://reveneau.com/ainews/apple-siri-model-delegation-inference-provider-claude-gpt-ios-27</link>
      <guid>https://reveneau.com/ainews/apple-siri-model-delegation-inference-provider-claude-gpt-ios-27</guid>
      <pubDate>Mon, 14 Sep 2026 20:20:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>A developer known as pdfu found two private frameworks in the iOS 27 and macOS Golden Gate release candidates that let a third-party model such as Claude answer as a Siri extension, or fully replace Apple's own Siri planner with something like GPT-5.6. (Source: MacRumors)</description>
    </item>
    <item>
      <title>Shadcn published a Tailwind design system linter that explains rule breaks to coding agents</title>
      <link>https://reveneau.com/ainews/shadcn-lint-agent-tailwind-design-system-500-stars</link>
      <guid>https://reveneau.com/ainews/shadcn-lint-agent-tailwind-design-system-500-stars</guid>
      <pubDate>Mon, 14 Sep 2026 19:20:00 +0000</pubDate>
      <category>Dev tools</category>
      <description>Shadcn published @shadcn/lint, an MIT-licensed ESLint and Oxlint plugin that writes design system rules for Tailwind v4 in a form coding agents can read, and the repository has picked up 500 stars in 12 days. (Source: shadcn on GitHub)</description>
    </item>
    <item>
      <title>Zeeshan Yaseen ran two GEO experiments, and third-party listicles produced 85.8 percent of AI citations</title>
      <link>https://reveneau.com/ainews/zeeshan-yaseen-geo-experiments-third-party-listicles-85-percent-citations</link>
      <guid>https://reveneau.com/ainews/zeeshan-yaseen-geo-experiments-third-party-listicles-85-percent-citations</guid>
      <pubDate>Mon, 14 Sep 2026 17:20:00 +0000</pubDate>
      <category>Go-to-market</category>
      <description>Two hand-tracked GEO experiments by consultant Zeeshan Yaseen found third-party listicles produced 85.8 percent of AI citations, with owned listicles at 14 percent. (Source: Search Engine Land)</description>
    </item>
    <item>
      <title>GitClear says AI coding users gained 25 percent output since 2023, and block duplication in their code rose 81 percent over the same period</title>
      <link>https://reveneau.com/ainews/gitclear-maintainability-gap-25-percent-output-81-percent-duplication</link>
      <guid>https://reveneau.com/ainews/gitclear-maintainability-gap-25-percent-output-81-percent-duplication</guid>
      <pubDate>Mon, 14 Sep 2026 15:20:00 +0000</pubDate>
      <category>Productivity</category>
      <description>GitClear's Maintainability Gap report, covering 623 million analyzed code changes from 2023 to 2026, says heavy AI coding users gain 25 percent on their own prior velocity while block duplication rises 81 percent and refactoring collapses. (Source: The New Stack)</description>
    </item>
    <item>
      <title>OpenRouter turns on US in-region routing, and DeepSeek, Kimi and GLM traffic can now stay entirely inside the United States</title>
      <link>https://reveneau.com/ainews/openrouter-us-in-region-routing-chinese-open-weight-models</link>
      <guid>https://reveneau.com/ainews/openrouter-us-in-region-routing-chinese-open-weight-models</guid>
      <pubDate>Mon, 14 Sep 2026 14:35:00 +0000</pubDate>
      <category>Infrastructure</category>
      <description>OpenRouter has made US in-region routing generally available for Business and Enterprise customers, decrypting and running requests entirely inside the United States, and pointed to open-weight models from Chinese labs as the reason many teams want it. (Source: OpenRouter)</description>
    </item>
  </channel>
</rss>
