Skip to main content

AWS Benchmark Aims to Reduce Number of False Positives Found by AI Vulnerability Scanners

Amazon Web Services (AWS) has developed a benchmark that can be used to test whether a model can distinguish real vulnerabilities from code that looks risky but is actually safe.

The Deception Benchmark was created following an evaluation of the capabilities of 12 models from five different providers. In all, the benchmark includes 14,822 samples of code built using 16 different languages spanning more than 70 Common Weakness Enumeration (CWE) categories. Each sample is run through an adversarial loop to generate code, test it against frontier models, harden, repeat. If a model gets it right easily, the sample is removed.

According to the benchmark, every AI model has the same fundamental issue. While they identify up to 95% of real vulnerabilities, they also flag 41 to 99% of safe code. Proof-of-exploit prompting can cut false positives by 17 to 74 percentage points but misses 7% to 44% of real vulnerabilities. The environment-gated challenges are worse: models flag the code and ignore the Kubernetes Network Policy next to it. No tested configuration keeps both false positives and false negatives below 10%, according to the benchmark.

To provide those assessments, the Deception Benchmark creates two distinct types of challenges for AI models. Code-level challenges present vulnerable and safe variants that differ by a subtle fix. Both look suspicious, but only one is exploitable. Environment-gated challenges go a step further by using that same code to test specific IT environments to determine, for example, if a Kubernetes network policy blocks a server-side request forgery (SSRF) path. That foundation, for example, then makes it possible for CyberGym agents to test on more than 1,500 realistic tasks while another AI agent invokes the CYBENCH framework to evaluate capture the flag (CTF) challenges.

Additionally, the benchmark is designed to treat labeling as a convergent audit loop rather than a one-time step. Every label is re-examined independently by multiple reviewers who do not see one another’s assessments or the original reasoning behind the label. Disagreements escalate to direct adjudication, where the original reasoning is evaluated against the challenge. Unresolved cases are then passed on for human review.

Neha Rungta, director of applied science at AWS, said ultimately the goal is to identify the AI models that generate the fewest number of false positives when used to scan for vulnerabilities. That issue is becoming problematic because as AI models are used to scan for vulnerabilities, many DevSecOps teams are now wasting more time than ever investigating vulnerabilities that turn out to be false alarms.

Preventing those false positives requires AI models that reason not just about the existing code, but also how a remediation might impact the surrounding environment, noted Rungta. The goal is to provide the context needed to enable AI models to better understand what good should actually look like when fixing a vulnerability, she added.

Mitch Ashley, vice president and practice lead for software lifecycle engineering at the Futurum Group, said false positives are verification debt that ultimately undermines confidence in AI findings. AI scanners that read code without the environment around it cannot tell an exploitable path from a blocked one, he added. A model that flags safe code as vulnerable is only making more work for DevSecOps teams, noted Ashley.

Just how much work AI models are creating for DevSecOps teams is unclear, but other reports suggest that the more complex the environment, the less helpful AI becomes. Hopefully, as AI continues to evolve, these tools and platforms will not just be used to discover vulnerabilities but also fix them in a way that requires as little intervention from a DevSecOps team as possible.



from DevOps.com https://ift.tt/MUZxdDN

Comments

Popular posts from this blog

Rochelle Walensky on the Rocky Road to Normal

Rochelle Walensky on the Rocky Road to Normal By David Wallace-Wells from NYT Opinion https://ift.tt/0A5Wx6r internal-sub-only-nl, Coronavirus (2019-nCoV), Vaccination and Immunization, Rumors and Misinformation, Centers for Disease Control and Prevention, Walensky, Rochelle

LocalStack Acquires WonderTwin AI to Gain SaaS App Emulation Platform

LocalStack this week revealed it has acquired WonderTwin AI , a provider of an emulator of software-as-a-service (SaaS) applications that is used to build custom applications for those platforms. Colin Neagle, vice president of marketing for LocalStack, said the emulators WonderTwin AI has developed will be integrated into the company’s namesake emulation platform that application development teams currently rely on to emulate cloud services provided by Amazon Web Services (AWS). LocalStack and WonderTwin AI make it possible for application developers working on a local machine to build applications that are designed to be deployed on some type of external cloud platform using a local sandbox to test and validate integrations without having to connect to a service or build against a live application programming interface (API). That issue has been especially critical in an era where more code will soon be generated by AI coding agents that may for one reason or another circumvent the...

Microsoft Brings the Azure SDK for Rust to General Availability

Microsoft has moved the Azure SDK for Rust out of beta and into general availability, giving Rust developers a stable, production-ready way to connect to core Azure services. The release covers Core, Identity, Key Vault (Secrets, Keys, and Certificates), and Storage (Blobs and Queues), built around the same design patterns already used in the .NET, Java, JavaScript, Python, Go, and C++ SDKs. The announcement came as part of Microsoft’s May 2026 Azure SDK release, and was detailed separately in a post from Ronnie Geraghty, product manager for the Azure SDK. He framed the milestone with a simple scenario: a Rust service that signs in with Microsoft Entra ID, retrieves a signing key from Key Vault, pulls work items from a Storage Queue, and writes the results to Blob Storage. Every piece of that chain is now stable. That stability matters more than it might sound. A beta SDK is fine for experimentation, but most engineering teams won’t put it in front of production traffic. W...