Small agencies and store owners now routinely let AI write code: a theme tweak, a custom shortcode, an order export, a shipping integration. Claude Code, Codex, Cursor and GitHub Copilot handle it fast, and the result usually works on the first try. The catch is that working code and secure code are two different things. A site that runs fine can still leave the door open for anyone who knows where to look.
So the real question is not whether a human or a model wrote the code. It is whether your process catches mistakes regardless of who made them. The Google developer team recently covered this in a short video built on DORA research. Below I apply its four pillars to WordPress and WooCommerce and add what the video left out: what an agent should never touch, how to handle credentials, and a workflow a small team can run. In short, the AI coding security WordPress teams can put in place this week.
The short version: four pillars of secure AI coding
- Small batches. One small change, one test, one review. Never a whole plugin in one go.
- Context and rules. The agent gets the security rules of the project in a file it reads on every task.
- Limited access and a real sandbox. The agent can do only what the task strictly requires. Production and customer data stay out of reach.
- External verification. Code is checked by tools and by a human, not only by the agent that wrote it.
Why AI-generated code is not secure by default
In its GenAI Code Security Report, Veracode tested more than 100 large language models. In 45% of cases the generated code failed security tests and introduced a vulnerability from the OWASP Top 10 categories. Against cross-site scripting (XSS), the models failed in 86% of relevant samples. Crucially, newer and larger models write more functional code, but their security stays flat.
For WordPress, that is bad news. According to the Patchstack report State of WordPress Security in 2025, the WordPress ecosystem saw 7,966 new vulnerabilities in 2024. A full 96% were in plugins, 43% could be exploited without authentication, and the most common type was XSS. Code an agent writes for you is, in practice, a new plugin or a theme modification, which is exactly where most holes appear.
Then there are credentials. In State of Secrets Sprawl 2025, GitGuardian reports that public repositories using Copilot had a secret leak rate of 6.4%, roughly 40% higher than the average across public repositories.
In the DORA 2025 announcement, Google notes that 90% of respondents use AI at work and more than 80% feel more productive. Yet 30% trust AI-generated code little or not at all. The central conclusion is that AI works as an amplifier. If you upload changes over FTP straight to a live site, an agent will just help you do that faster.
Pillar one: small batches and tests
In its description of working in small batches, DORA states that AI adoption often increases software delivery instability and that small batches are the main counterweight. Nobody reviews or tests a huge change properly, however fast the agent produced it.
So break the request down. Instead of asking for a loyalty program plugin, ask first for point calculation from an order, then for displaying points in the customer account, then for redeeming a discount. Each step is a separate commit you can read in a few minutes.
Tests belong with small batches. Test-driven development works well with an agent: first a test that describes the expected behavior, then code that passes it, then refactoring. The test narrows the task and protects you when a later fix touches the same code. WordPress has an official PHPUnit testing setup, so this is nothing exotic.
Review the tests the agent writes, too. An agent that misunderstood the task will write a test confirming the wrong behavior. For payment and login features, write the test scenarios yourself, even if only in plain words.
Pillar two: project context and rules
Every major tool now reads a project rules file. Gemini CLI loads GEMINI.md files hierarchically, from global files down to individual subfolders. Claude Code has CLAUDE.md, while Codex, Cursor and the Copilot coding agent support the open AGENTS.md format, where the file closest to the edited code takes precedence.
That hierarchy is useful. General rules go in the project root, stricter rules for the payment gateway API go in the payment integration folder. For WordPress, I recommend turning the principles from the official WordPress developer security handbook into rules:
- never trust user input, and validate and sanitize every input,
- escape output as late as possible, right before it is printed,
- verify a nonce on every form and AJAX request,
- check user capabilities before every action,
- build SQL queries only with prepared statements through the wpdb object, never by concatenating strings,
- never add a new plugin or library without explicit approval.
Keep the file short. The longer it gets, the more tokens every task burns and the easier it is for a key rule to get lost.
Pillar three: access and sandboxing
This is where feeling safe and being safe differ most. A rule in the context is a request, not a control. The agent can overlook it, or an injected instruction can override it, for example text hidden in a product review the agent is processing. OWASP lists this as prompt injection in its Top 10 for LLM applications and right next to it describes Excessive Agency, caused by too much functionality, too many permissions and too little human oversight. OWASP recommends enforcing permissions in the system the agent connects to, not relying on the model to judge what it may do.
A similar misconception applies to sandboxing. Many tools call a plain Docker container a sandbox, yet a container shares the kernel of the host, and a kernel vulnerability can let an attacker escape. Solutions such as gVisor put their own application kernel between the app and the host, so system calls never reach the host directly.
For a small team, configuring what the tools already offer is enough:
- Codex, according to its approvals and security documentation, runs without network access by default and offers read-only, workspace-write and full access modes. Do not use the last one.
- Claude Code lets you block reading files such as .env or whole secrets folders in its permission settings. These rules do not stop a script that opens the file itself, so for a hard boundary turn on sandboxing at the operating system level.
- Cursor and Copilot run terminal commands inside the editor. Keep automatic execution without confirmation off for anything that is not a pure read.
What an agent should never be allowed to do
- access the production store database, not even read-only,
- hold login credentials for hosting, FTP, SSH or the live site admin,
- use live API keys for the payment gateway, shipping carrier or accounting system,
- deploy to production without human approval,
- install plugins and libraries from unverified sources,
- work with real customer data from orders.
The last point is not only technical. Orders contain personal data. In the EU, pasting that data into an AI tool means sharing it with another processor, and Article 28 of the GDPR requires sufficient guarantees and a binding contract for that. In the US, the FTC guide Protecting Personal Information makes the same practical point: keep only the data you actually need. For development, a database copy with anonymized customers is enough.
Pillar four: external verification
If the same agent writes the code and reviews it, you still have one perspective. Independent tools belong in the process. Static application security testing (SAST) scans source code for risky patterns. Software composition analysis (SCA) checks whether your libraries have known vulnerabilities. Both are deterministic: the same code gives the same result, which is what you want from a check.
They flag potential problems, not proven exploits, and untuned they produce many false positives. Hand a finding to the agent, let it propose a fix, and let your tests confirm nothing broke.
Static analysis will not find business logic flaws, such as a reusable discount code or an order visible to another customer after changing a number in the URL. For these, let the agent play attacker and look for ways to abuse your store. In its description of pervasive security, DORA stresses that security belongs in daily work and automated testing, not in a final check before launch.
A secure workflow for WordPress and WooCommerce with an AI agent
1. Set up a separate development environment
The agent works on a local copy or on staging, never on the live site. Anonymize the database before copying it. Keep staging out of search results with noindex, and make sure that setting never travels back to production with a deploy. One wrong flag can wipe a store from the index, which is why technical SEO for AI search starts with a clean deployment process.
2. Put the project under Git
Every agent change must be visible as a diff. Without version control you do not know what changed and have nothing to roll back to.
3. Separate secrets from code
Keys and passwords belong in a .env file or environment variables listed in .gitignore and invisible to the agent. Protect wp-config.php following the official WordPress hardening guide, which also recommends disabling the admin file editor with the DISALLOW_FILE_EDIT constant. Use separate test keys with minimal permissions on staging. WooCommerce REST API keys can be set to read-only. If a key ever ends up in a chat or a repository, rotate it immediately. Deleting it is not enough.
4. Write the rules file
Create AGENTS.md or CLAUDE.md with the principles from pillar two, a description of the project structure and a list of forbidden actions.
5. Assign work in small steps
One function, one test, one commit. Read the diff at every step before you approve it. If you do not understand the diff, the change does not move forward.
6. Run automated checks
The minimum for WordPress is PHP_CodeSniffer with the WordPress Coding Standards ruleset, which includes security checks for escaping, sanitization and nonces. Add dependency checks with composer audit or npm audit, and turn on secret scanning in the repository.
7. Have a human review the code
At the very least for anything that touches payments, login, orders and personal data. The reviewer need not be a senior developer, just someone who asks, for every input, where it came from and who can see it.
8. Deploy with a way back
Back up before deployment and check the logs after it. The server access log often shows the first signs of bugs the tests missed. Then keep agent-written code updated with the same discipline as every other plugin and WordPress core.
Overview: what an agent can do and under what conditions
| Task | Agent allowed? | Condition |
|---|---|---|
| Theme tweak, CSS, shortcode | Yes | Staging, small commit, diff review |
| New custom plugin | Yes | Tests, SAST, human review |
| Shipping carrier or payment gateway API integration | With care | Test keys only, folder-level rules |
| Installing a plugin from the internet | Not alone | A human selects and approves |
| Working with the production database | No | Anonymized copy only |
| Deploying to the live site | Not alone | Backup and human approval |
My take: what the video did not cover
The video targets teams with a CI pipeline, a security department and in-house developers. Most small WooCommerce stores and many small agencies in the EU and the US have none of that, often not even staging. For them, the biggest risk is the order of the steps, not a missing SAST tool. As long as an agent works on production with administrator access, no scanner will save you.
The second gap is the supply chain. An agent is happy to suggest installing a plugin or library that solves the task. Who decides what gets into your site is the same question as with nulled plugins, in new packaging. OWASP lists supply chain as a separate risk, and in WordPress I consider it more critical than the quality of the generated code itself.
Third, not every task needs an agent with terminal access. A recurring order export or a product feed update is often better handled by a fixed automated workflow that has nothing to break. Choosing the smallest tool that does the job is itself a security decision.
Summary
Secure AI coding is not about whether to trust the agent. It is about setting things up so you do not have to. Small batches with tests, clear project rules, limited access with a real sandbox and independent review work whether a human, Claude or Copilot wrote the code. For WordPress and WooCommerce, start with staging, Git, secrets out of the code and no agent on production. Those steps take a few hours and remove most of the risk. Advanced scanners and attacker agents come after that.
If you are not sure your store has these basics in place, it is worth checking before the agent builds the next feature. WooAcademy, run by Milan Fraňo, helps companies in the EU and the US with this remotely, from setting up a secure AI-assisted development workflow and WooCommerce development to a technical SEO audit that makes sure a deployment never quietly hurts your visibility.
FAQ
Is code written by AI secure?
Not by default. Veracode testing found security flaws in 45% of AI-generated code, and newer models performed no better. Review it as strictly as code from a new colleague, ideally with automated tools and a human who understands WordPress security principles.
Can I give an AI agent access to my hosting or FTP?
I do not recommend it. The agent should work on a local copy or staging, and a human should approve every deployment. It should never see production credentials, not even in project files, because anything it can read it can also leak.
Is it enough to tell the agent not to open the file with passwords?
No. An instruction in the context is a request the model can overlook or injected text can override. Real protection means secrets are not in the agent working folder at all and the tool runs in a sandbox with limited file and network access.
Is a Docker container a sufficient sandbox?
It isolates a development environment, but it is not a security boundary on its own, because it shares the kernel of the host. Stronger isolation comes from gVisor or virtual machines. For a small team, the built-in sandbox of the tool with network access off is a reasonable minimum.
Which tool is the most secure: Claude Code, Codex, Cursor or Copilot?
Configuration matters more than the brand. All four support restricted permissions and command approval. Security depends on staging, secrets kept out of code, small changes and a review step before production.
Do I need a developer to use AI coding tools safely?
Not for simple theme tweaks. Code that handles payments, orders, login or personal data should be reviewed by someone who understands WordPress security. AI speeds up the work, but responsibility for the site stays with you.
How do I know the agent did not add something malicious to the code?
Review every diff before approving it, run automated checks on code and dependencies, and block new packages without your consent. If something looks suspicious, revert it in Git and reassign the task in smaller steps.

