If you are building with AI coding assistants and counting on a security prompt to keep your app safe, this study is a useful reality check.
Developer Nazmul Hasan ran 24 builds across Claude Code and Codex variants to test how well different models handle security by default. The headline finding: which model you choose matters more than whether you tell it to write secure code.
What the Builds Found
The results split cleanly by model tier. Claude Sonnet 5.5 shipped all 12 tested security protections without any explicit instruction to do so. It applied them as part of its default output.
Claude Haiku 4.5 did the opposite. It hardcoded a signing key into the build, even after being told not to. That is the kind of vulnerability that ends up in a breach report.
Where Semgrep Comes In
Semgrep, the static analysis tool, flagged the hardcoded key and identified the correct fix. That is the expected use case for Semgrep: catching what the model missed and pointing toward remediation. It worked as advertised here.
The Operator Takeaway
There are two practical conclusions for solo builders and small dev teams:
- Default model behavior is not uniform. A cheaper or faster model may skip security defaults that a flagship model applies without prompting. The cost difference between tiers can be small compared to the exposure difference.
- Static analysis is still required. Even if your model of choice generally handles security well, running a tool like Semgrep after each build catches the cases where it does not. One hardcoded secret in production negates a lot of prompt engineering.
The full study covers all 24 builds and the full list of 12 security protections tested. It is worth reading if you are shipping anything that touches user data or handles authentication. The source article is published on Level Up Coding by Nazmul Hasan.
