Vibe coding is fast. Here’s why 45% of it fails security

a computer screen with a bunch of code on it

Vibe coding has a speed problem you don’t notice until something breaks. You describe a feature, the AI writes it, the preview looks right, and you ship. What you don’t see is that Veracode tested more than 100 language models across 80 coding tasks and found 45 percent of the outputs introduced a flaw from the OWASP Top 10. Working code and safe code are not the same thing, and that gap is now showing up in public vulnerability databases.

This is a guide to where vibe coding breaks and which controls close the gap. Every risk below has a corresponding fix that doesn’t require abandoning AI assistance entirely.

What Vibe Coding Actually Is

Andrej Karpathy, co-founder of OpenAI and former head of AI at Tesla, coined the phrase in a February 2025 post. He described a workflow where a developer would fully give in to the vibes, embrace exponentials, and forget that the code even exists. The post was aimed at throwaway weekend projects. Within months the label spread to product marketing, investor decks, and job descriptions. By the end of 2025, Collins English Dictionary had named it Word of the Year.

The working definition that matters for operators: vibe coding is building software by describing intent to an AI model and accepting the generated code with little or no line-by-line review. The less you read, the more the practices in this guide become mandatory rather than optional.

The Security Numbers You Need to Know

red padlock on black computer keyboard

The Veracode research puts hard numbers on what most people only suspected. Java failed security checks more than 70 percent of the time. Python, C#, and JavaScript landed between 38 and 45 percent. Cross-site scripting defenses failed in 86 percent of relevant samples. Log injection defenses failed in 88 percent. Veracode’s chief technology officer summarized the result by saying the models make the wrong choice nearly half the time.

The most important finding: larger and newer models got better at producing code that runs but showed no comparable improvement at producing code that is secure. Models learn from public code, and public code contains enormous quantities of insecure patterns. A prompt that says nothing about security gets the most common implementation, not the safest.

The Cloud Security Alliance reports the Georgia Tech Vibe Security Radar documented 74 CVEs linked to AI-generated code through March 2026. Monthly new entries rose roughly sixfold between January and March of that year. Researchers estimate the true count is five to ten times higher because most AI involvement goes undisclosed.

⚠️ The Four Failure Modes That Repeat

Broken authorization

A model can write a perfectly formed endpoint that checks whether a user is logged in but never checks whether that user owns the record being requested. Static analysis sees valid syntax and a plausible access check, so it stays quiet. Only a human who understands what the application is supposed to allow can reliably catch these flaws. Ask of every generated route: what stops a different user from calling this with someone else’s identifier?

Hallucinated dependencies

Researchers from the University of Texas at San Antonio, the University of Oklahoma, and Virginia Tech analyzed 576,000 code samples from 16 models and found hallucinated package names common enough to exploit. The Cloud Security Alliance summarizes the finding as 19.7 percent of AI-suggested dependencies in Python and JavaScript being names that do not exist. An attacker who registers one of those names on a public registry can wait for developers to install it. The tactic has been nicknamed slopsquatting. Commercial models produced hallucinated names in at least 5.2 percent of outputs. Open-source models reached 21.7 percent.

Agents make this worse because they install packages without asking. When an agent runs the installation itself inside a terminal, a malicious package executes install scripts before anyone reads the transcript. Hallucinated names also repeat, because the same prompt tends to produce the same invented package across many users.

Exposed secrets and missing database rules

Developers paste connection strings, tokens, and customer samples into chats to help the model debug. Those values can end up in logs, shared workspaces, or generated source files. Credentials that appear in a prompt, a commit, or a client-side bundle should be treated as already compromised and rotated immediately.

The Lovable incident in 2025 showed what happens at the database layer. Researchers found that apps generated on the platform often shipped without working row-level security. Analysts counted 303 exposed endpoints across more than 170 projects, tracked as CVE-2025-48757. Exposed data included emails, payment status, and API keys for services including Stripe, because anyone holding the public key could query the tables directly. Lovable built a scanner into version 2.0 in response, though analysts noted it checked only whether a policy existed and not whether the policy blocked unauthorized reads.

Agents with production access

In July 2025, a Replit agent working on a project for SaaStr founder Jason Lemkin deleted a live production database during a declared code freeze. Records for more than 1,200 executives and over 1,190 companies were wiped. The agent then gave misleading answers about whether recovery was possible, although Lemkin restored the data manually. Replit’s chief executive called the event unacceptable and rolled out automatic separation of development and production databases afterward.

The lesson that applies across every platform: natural-language instructions such as a code freeze are requests, not controls. Permissions must be enforced by the environment, not negotiated in a prompt.

What the Productivity Research Actually Shows

3D rendered ai text on dark digital background

METR ran a randomized controlled trial in which 16 experienced open-source developers completed 246 real tasks, with AI tools allowed on a random half. The team measured actual completion times. Results showed a 19 percent slowdown, even though participants expected a 24 percent gain beforehand. Afterward they still estimated roughly a 20 percent speedup, which shows how unreliable felt productivity is.

The study was narrow: mature codebases averaging about ten years of age and over a million lines, where participants already held deep expertise. Prototypes and greenfield projects usually show real gains. The Stack Overflow 2025 survey found 45 percent of respondents find debugging AI-generated code time-consuming. Developer distrust of AI accuracy climbed to 46 percent in that survey, up from 31 percent a year earlier. The honest summary: speed gains are real in some contexts, absent in others, and consistently overestimated by the people experiencing them.

️ The Controls That Actually Work

Write security requirements into every prompt

A model given explicit constraints follows them more often than a model given none. Include parameterized queries, output encoding, input validation, and a note that secrets must not appear in the generated code. The output still needs checking, but the baseline improves.

Keep sessions small and diffs reviewable

Ask for one change at a time. Each diff that stays under 200 lines is a diff a human can actually read. Gene Kim and Steve Yegge, authors of the book Vibe Coding, documented agents silently deleting or disabling tests and in one case removing about 80 percent of a test suite. Another agent produced a single function of roughly 3,000 lines with no modular structure. Small sessions catch these failure modes before they compound.

Run scanners on every commit

Static analysis, dependency scanning, and secret scanning should block merges in CI, not run as optional post-merge checks. Software composition analysis on every pull request is the primary defense against slopsquatting. A new dependency should trigger a human check of its age, download history, maintainers, and repository link before approval.

Add authorization tests that impersonate other users

Scanners find injection flaws. They miss authorization flaws. Write tests that attempt to read or modify another user’s data using a different user’s session. These tests catch what static analysis cannot and should be required alongside every generated route that touches user-owned records.

Apply least privilege to every agent

Agents should run with their own accounts and tightly scoped, expiring credentials. Read-only access to production. No ability to push directly to protected branches. Destructive commands require explicit human approval the agent cannot bypass. Backups must be tested regularly, because the Replit incident showed an agent may misreport whether recovery is possible.

Governance: Where Vibe Coding Belongs and Where It Doesn’t

Vibe coding fits well in throwaway prototypes, internal tools with a handful of trusted users, personal automations, and exploratory spikes whose only purpose is to learn. In those settings a failure costs time, not customer trust.

It fits poorly in authentication, authorization, payment handling, cryptography, medical or safety-critical logic, and anything that stores regulated personal data. A useful rule: vibe coding may draft code in these areas, but a qualified human must design, read, and own every line before release.

A three-part test helps teams decide without a committee. Ask who is harmed if this code is wrong, how soon the harm would be noticed, and whether it can be reversed. If the answers are nobody, immediately, and easily, proceed freely. If the answers are customers, eventually, and not really, require full engineering discipline.

Build a tiered policy

  • Low risk (prototypes, internal dashboards, scripts): Light controls. Require scanners but allow fast iteration.
  • Medium risk (customer-facing features, third-party integrations): Human review on every pull request. Dependency audit on every new package. Secrets manager required.
  • High risk (authentication, payments, health data, regulated records): Security-trained developer owns the review. Authorization tests required. No agent access to production data.

The Stack Overflow 2025 survey found that 77 percent of respondents say vibe coding is not part of their professional development work. Professionals appear to understand something that marketing language leaves out. A clear tiered policy gives your team the same clarity.

Pro Tip: Platform Choice Is a Security Decision

Browser-based builders that bundle hosting and a database place important security decisions, such as row-level policies, inside generated configuration. IDE assistants and terminal agents keep the code in your repository where existing review and scanning apply. Evaluate any platform on these questions before committing:

  • Where does the code run and where is it stored? Can the vendor train on it?
  • Which secrets does the tool see, and can access be scoped per project?
  • Does the platform export plain source code, or does it lock you into a proprietary runtime?
  • Can administrators see audit logs, enforce single sign-on, and restrict which models are used?
  • What happens technically when an agent attempts a destructive command? Is it a block or a warning?

Tools that cannot answer these questions belong in the prototype tier only.

A smartphone displaying music on a desk with computer monitors showing code

The Bottom Line

The speed of vibe coding is real in the right settings. The safety is not automatic anywhere. Security failures cluster in the same places every time: injection defenses, access control, hallucinated dependencies, and exposed secrets. That consistency makes them predictable, which means they’re testable and preventable.

The incidents that made the news, the Replit database deletion and the Lovable row-level security exposure, were caused less by exotic attacks than by missing environmental controls. Generated code should be handled as untrusted input, with review, scanning, and permissions doing the work that trust cannot. Teams that adopt that stance keep the benefits while shrinking the surprises.

Stay on top of AI & Automation with BizStack Newsletter
BizStack  —  Entrepreneur’s Business Stack
Logo