I keep one rule when I read a vendor’s AI coding post: the percentage of code that came from an agent means nothing by itself. The number that matters is the percentage that ships without a human reading it. JetBrains (a software development toolmaker that runs an annual survey of more than 15,000 professional developers) put the first figure at 47 percent for May through July 2026. The second figure is barely tracked at all. That gap is the whole story.
I want to walk through what the JetBrains survey actually says, what it does not say, and where the verification math falls apart for the average team. By the end I will tell you what I would measure on my own team if I were standing one up tomorrow, because the right answer is not “use less AI” or “use more AI.” It is “decide what gets read, and measure it.”
What the survey actually counts
The headline number comes from a bucket-midpoint methodology. JetBrains asks developers to place their own output into rough bands: how much of the code they shipped last month was hand-written, how much was AI-assisted (the human wrote it with autocomplete suggestions), and how much was fully agent-written (the human typed a prompt and accepted the output). The percentages are then calculated from the midpoint of each band.
A consequence of that methodology is that the three figures can add up to more than 100 percent. That is not a math error, it is the structure of the question. The survey is also self-reported, which means the developers who answered it are the kind of developers who answer JetBrains surveys. That is not a representative cross-section of the industry, but it is a useful pulse.
The author reports a separate, smaller number that matters more for the verification story: about 56 percent of AI-generated code passes basic security tests on the first try. That is the rate for code that has been through the agent, through a SAST (static application security testing) tool, and out the other side without flagging. It is not the rate for code that has been read by a human reviewer. The two numbers measure different things, and the gap between them is where the risk lives.
Why the 47 percent number is not the crisis
The panic reading of the headline is that AI is shipping half your code without a human ever looking at it. That panic reading is wrong. The 47 percent is a measure of typing, not a measure of trust. Most professional teams have at least a pull request review, a CI (continuous integration) test run, and a deploy gate before code reaches production. AI-generated code still flows through that pipeline.
The story that should worry you is more specific. About 22 percent of developers report that more than 80 percent of their code comes from agents. That is a real chunk of the industry running close to an AI-first development model. For a senior engineer who already knows the codebase intimately and is using agents to type boilerplate they would have typed anyway, the risk is small. For a junior engineer who is using an agent to generate code in a language they do not yet understand, the risk is enormous. The 47 percent average hides that distribution.
Domain split is the other hidden variable. A 47 percent agent share on a CRUD (create, read, update, delete) web form is not the same risk profile as a 47 percent agent share on a payments handler or an authentication path. JetBrains does not break this out. Most teams do not measure it.
Where the verification math actually breaks
The cleanest way to think about the verification gap is to do the math. Take a team that writes 100 lines of code per week, with 47 of those lines coming from an agent. The team has a senior engineer who can review about 800 lines of code per week without burning out. That is generous, and it does not include meetings, design reviews, or the other half of the senior engineer’s job.
If the senior engineer reviews 800 lines a week and the team writes 1,000 lines a week, the review backlog grows by 200 lines a week. By the end of a quarter the backlog is 2,400 lines, which is about a month of unpaid review debt. That is the math that nobody publishes in a vendor blog post, and it is the math that explains why teams ship AI code faster than they can read it.
The compounding factor is that AI-generated code is harder to review than hand-written code. The agent writes code that looks right at a glance, with consistent style and correct imports, so a reviewer skimming a pull request is more likely to approve it without reading every line. That is the opposite of what review is supposed to do.
What the productivity numbers actually cover
The other set of numbers worth naming is the productivity story. Vendors like to claim that AI coding tools make developers two times faster, sometimes more. The honest version of that claim is narrower. AI coding tools make developers faster on specific kinds of work: boilerplate, test scaffolding, and the kind of glue code that nobody wants to write by hand. They make developers slower, or no faster, on the kind of work that requires understanding domain logic, debugging a flaky test, or untangling a legacy abstraction.
The JetBrains survey does ask about productivity, and the answers skew positive for the easy end of the work and neutral for the hard end. The vendor narrative usually only quotes the easy end.
Here are the four numbers I would track on any team that ships AI-generated code to production:
- The percentage of AI-generated code that gets reviewed by a human before merge. If this number is below 80 percent, you have a verification problem whether you know it or not.
- The first-pass security scan rate on AI-generated commits. Anything below the team baseline for hand-written commits means your security team is doing the agent’s job a second time.
- The median time from commit to first human review. Past 24 hours, the review debt starts compounding. Past a week, the diff is no longer small enough to review cleanly.
- The percentage of merged AI-generated commits that get reverted or hot-fixed within 30 days. High numbers here are an early signal that the review pipeline is rubber-stamping AI code.
Trade-offs
The 47 percent agent share is not free in time. There are real costs that the headline number hides. The first cost is review debt. If your team writes more code than it can read, the unread code accumulates, and the unread code is the code that breaks in production at 3 a.m. The second cost is onboarding. Junior engineers who learn from AI-generated code learn the style of the model, not the style of your codebase. The third cost is security pass rate. A 56 percent first-pass rate on AI code means your security team is reviewing roughly 44 percent of agent-written code a second time. That is real work, and most teams do not budget for it.
The honest reason to track the verification gap is that AI code is not unsafe by default. It is unsafe by default when nobody reads it. A team that ships 100 percent AI-generated code with a strict review pipeline and a measured security pass rate is in a better position than a team that ships 50 percent AI-generated code with no review pipeline. The percentage of code that came from an agent is a means, not the goal.
In our case, the team I would stand up tomorrow would set three numbers in the first week and never lose sight of them: the percentage of AI-generated code that gets reviewed by a human before merge, the security pass rate on AI-generated code, and the time from commit to first human review. The lines-of-code metric can move wherever it wants.
Bottom line
If you remember three things from this piece, make them these. The 47 percent figure measures typing, not trust, and the trust figure is the one that matters. The verification gap is a math problem, not an attitude problem, and the math problem has a solution that looks like a stricter pipeline rather than a stricter policy. The productivity story is real for boilerplate, neutral for the hard work, and the vendor narrative usually only quotes the first half. Treat AI coding agents like a fast junior developer who learned from the internet. Useful for the parts that are easy. Dangerous on their own for the parts that are hard. Strict review pipeline either way.
If I could send a message back to the version of me that shipped his first AI-generated pull request without reading it twice, I would say one thing. The agent is faster than you. That is not permission to stop reading. It is permission to read more carefully, because the code that looks the safest is usually the code that breaks the loudest.