Models & agents

Enclave says DeepSeek V4.1 Flash cleared its 11-target hacking benchmark for $4.65

September 16, 2026 at 12:30 PM PT

Enclave benchmark chart showing DeepSeek V4.1 Flash reaching code execution on 11 vulnerable Grafana, Jenkins and Nextcloud targets

Image: Enclave AI

Why it mattersA frontier model scoring 11 out of 11 on Grafana, Jenkins and Nextcloud exploitation for under five dollars sets a new low bar for what a bring-your-own-key hacking agent costs, and any red team pricing an in-house tool against a service should redo the maths.

Enclave AI, which runs a public hacking-agent benchmark, published a post on 16 September saying DeepSeek V4.1 Flash reached code execution on all 11 vulnerable targets in the suite and left all four fixed targets alone. Enclave says the accepted runs cost $4.65 in API charges, and the full cost including failed attempts was $5.14. The suite uses isolated copies of Grafana, Jenkins and Nextcloud.

The post is by Yanir Tsarimi, Enclave's co-founder and CPO. Enclave says the review that follows the raw score is more useful than the score itself. Six of the eleven successful runs used the planned exploit path. The other five used routes that the scoring system did not distinguish from the planned solutions, and that the audit only caught by reading every command the model sent.

The numbers Enclave published

Enclave reports 2,349 Bash commands issued across the full benchmark and two hours and 38 minutes of active model time. The median successful run took 4 minutes 38 seconds. The provider reported 268.3 million input tokens and about 2 million output tokens, and Enclave says 266.2 million of the input tokens were served from cache, which is what makes the run cheap.

The individual challenge times are in the post. Grafana fell in 52, 64 and 90 seconds across three runs, each time by dropping executable files into a temporary plugin folder and asking Grafana to load it. The first Jenkins challenge, on how the server reads command options from files, gave DeepSeek a route to read a private controller credential and then run code through Jenkins' built-in script tool. It completed the full attack in three out of three runs.

Where the audit changed the reading

The audit is where the story gets interesting. Enclave says three of the Grafana runs and two of the Jenkins runs used shorter routes that were only present because the vulnerable test versions carried extra weaknesses beyond the one the challenge was designed to test. Those five routes belong to Enclave's benchmark environment and, Enclave says, do not represent new security holes in the upstream Grafana or Jenkins projects.

Enclave has since patched the shorter routes and released new source versions of the challenges. Leaderboard entries will use results from matching benchmark versions, so models tested against the earlier suite need fresh runs before they rank against DeepSeek's score.

Why outcome-only scoring hides half the story

The Enclave post spells out where the leaderboard number gets you and where the audit takes over. A hacking agent searches for the fastest working route, and it has no reason to follow the path the test author intended. DeepSeek finding five alternates that a human auditor only spotted after the fact is a finding about the benchmark first, and about the model second.

The write-up is a self-scored benchmark by a vendor that sells AI security work, so treat the leaderboard placement accordingly. The commands, the timings, the cache-hit ratio and the money numbers all come from the same source. The methodological point in the post, that outcome-only scoring lets shortcuts through, is the piece any team writing its own agent evaluation should copy.

Source

Source: Enclave AI

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

More from AI News

CodeRabbit measured GPT-6 Astra catching 61.3 percent of labelled bugs in code review, at 2.5 times the token price of Sol

CodeRabbit published an early evaluation putting GPT-6 Astra at 61.3 percent actionable bug coverage against 59.0 for GPT-5.6 Sol, with the gap widening to 57.1 against 47.6 on cross-file reviews that span more than one file.

Source: Hacker NewsModels & agents

The best model in a new benchmark steered a coding agent through a full task 24.69% of the time

LoopArena tests how well a model can direct a separate coding agent through a long task, and the top score on complete tasks was 24.69%, with five models measured against the same worker.

Source: GitHubModels & agents

Latent Space burned 20 billion tokens on GPT-6 Astra and measured the running cost at under $6 an hour

Latent Space spent more than 20 billion tokens on GPT-6 Astra during early access and reports a sustained running cost under $6 an hour, with the real spending risk coming from how many agents the model starts in parallel.

Source: PressModels & agents