Skip to content
back to the archive page
#AI

Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks (Category AI)

Image 1 of 3 for “Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks”

Image 2 of 3 for “Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks”

Image 3 of 3 for “Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks”

An interesting study on vibe-coding security came out on 2 December, by Songwen Zhao, Danqing Wang, Kexun Zhang, Jiaxuan Luo, Zhuo Li, and Lei Li. The main research group is associated with Carnegie Mellon University.

In simple terms, the authors describe vibe coding as an approach where a human programmer states a task in natural language and an agent powered by a large language model (LLM) performs complex coding work with minimal manual intervention. But is the resulting code secure?

To answer that, they built SUSVIBES, a benchmark of 200 realistic feature-development tasks for existing open-source projects. Each task comes from real development history: when people originally implemented these features, they unknowingly introduced security vulnerabilities. These tasks are far more complex than the isolated examples in earlier security benchmarks. They require changes to about 180 lines of code on average, across several files in a large repository, rather than trivial edits within one function or file. Together, they cover 77 vulnerability categories in the CWE classification.

Running the benchmark on popular agent systems (SWE-Agent, OpenHands, and Claude Code) with popular LLMs (Claude 4 Sonnet, Kimi K2, and Gemini 2.5 Pro) produced worrying results. The best system, SWE-Agent with Anthropic’s Claude 4 Sonnet, successfully solved 61% of the tasks, but only 10.5% of those solutions were actually secure, with no vulnerabilities.

Simple interventions, such as adding hints about possible vulnerabilities to the task description, did not help much. Agents more often failed the functional part of the task, while overall success at producing secure solutions barely improved.

At this stage of the technology, then, vibe coding speeds up feature creation but can quietly introduce security problems. Without substantial improvements to agent architecture or model training, this code cannot be trusted in critical systems. The study calls on the industry to work on those improvements.

P.S. I particularly liked the pipeline for assembling the benchmark and running the tests :)

The pipeline is automated and built around real repositories with known vulnerabilities. For each vulnerability in a project’s history, the researchers prepared a feature-development task whose earlier implementation had led to that security problem. Each task includes the code state before the vulnerability was fixed, a feature request, and tests. There are two test groups: functional tests and security tests that detect the specific original vulnerability. The requirements for a secure solution are therefore known in advance.

Agents tackled each task in an isolated environment, such as a Docker container, with access to the project’s code. Given a feature request, an agent could read and edit files, compile and run the project, run tests, and repeat these steps over several iterations, approximating a real development process. It then produced a patch implementing the requested functionality. That patch was automatically run through both test suites. The success metrics were:

  • Func Pass: the percentage of tasks whose patches passed all functional tests and implemented the task correctly.
  • Sec Pass: the percentage whose patches also passed all security tests, without introducing vulnerabilities.

#Engineering #Software #Processes #Productivity #Economics #Security