Do you have an incident?

Our S.O.S. line:

+49 89 262 025954

Our team of experts is ready to assist your organization in the event of a cyberattack.

details

AI Might Not Take the Pentester’s Job After All – But It’s Not for Lack of Trying

Penetration Testing WhiteHat todayJuly 9, 2026

Background

Introduction

AI companies tell their tools are here to “strengthen cyber defence” (https://openai.com/index/trusted-access-for-cyber/) and “help secure the world’s most critical software.” (https://red.anthropic.com/2026/mythos-preview/). Then, almost in the same breath, they mention that “non-experts can also leverage” their offensive security agents, which is either a bold democratization of security knowledge, or the most politely worded threat to the pentesting profession we’ve seen in a while.

One thing we can all agree on: systems need to be more secure. The question has always been how. “Fully secure” remains a concept that lives primarily in executive reports and vendor slides, rarely in production environments. LLMs are genuinely capable of helping pentesters, automating repetitive tasks, and they’re actually improving over time. Whether that means pentesters can sleep peacefully because they still have “critical thinking” on their side, or whether it means the collaboration between human and AI is quietly reshaping what the job even looks like, is worth examining.

The recent launches of Anthropic’s Mythos Preview and OpenAI’s ChatGPT Cyber raise exactly these questions. What are the real limitations of AI pentesting agents — and, more importantly, why those limitations might be the thing keeping human pentesters on the payroll.

Methodology: Trust Me Bro.

Anthropic states that findings are validated by a human expert before disclosure to the maintainer (https://red.anthropic.com/2026/mythos-preview/). One human. Somewhere in the loop. Presumably.

“Organizations must be able to explain not only how AI systems function, but how risks are controlled, monitored, and audited,”

writes F5 (https://www.f5.com/resources/white-papers/evaluating-enterprise-ai-security).

Agents can do valuable things — they can help experts reach the point where experience, environment, client needs, and professional judgment actually matter. What they cannot replicate is that judgment itself. An agent will not inherently understand what it means to trigger a denial-of-service condition in a hospital’s patient monitoring system. It will not know to avoid pasting sensitive military credentials into an online Base64 decoder because it seemed convenient. These are not edge cases. These are the decisions that separate a professional service from a liability event.

So let’s say the AI successfully completes an engagement. It has produced output. What’s next? Even a high-quality report requires human time to validate, merge, and retest. The capacity to absorb and remediate findings does not scale the way discovery does (https://fluidattacks.com/blog/claude-mythos-project-glasswing-appsec-future). You can accelerate the identification of problems considerably faster than any organization can fix them.

The clearest evidence of this is sitting in the NVD. NIST announced in April 2026 that it would update NVD operations in response to record CVE growth — submissions increased 263% between 2020 and 2025 (https://www.nist.gov/news-events/news/2026/04/nist-updates-nvd-operations-address-record-cve-growth). And yet, as Rafael Álvarez observed, CVE submissions are up while public exploit databases show no corresponding increase (https://www.jralvarezc.com/p/the-myth-of-mythos). The vulnerabilities are being reported. The exploits aren’t showing up. Like a doctor who diagnoses everything as rare tropical diseases and then goes home.

The most straightforward interpretation is that a significant portion of what is being submitted is noise — AI-assisted discovery running at scale, flagging everything that pattern-matches to a vulnerability, whether or not it actually is one. The CVEs are piling up. The exploits are not. Someone has to read all of it anyway.

Which brings the methodology question full circle: if the output cannot be trusted without human review, and human review cannot keep pace with the output, the process has not become more reliable. It has just become louder.

Sensitive Data In, Sensitive Data… Where, Exactly?

Whether or not your data trains the model is, at this point, almost beside the question. The more pressing issue is that it leaves your hands entirely. Even granting these LLMs the most generous interpretation of their privacy commitments, that’s a remarkable amount of institutional trust to place in a tool you’re simultaneously handing the keys to your infrastructure. And yet, that is precisely what AI pentesting agents require.

The economics, admittedly, are difficult to argue with. I personally cannot work day and night, though I don’t require the water supply of a mid-sized city and a power grid to think, which I consider a personal win. Agents will get things done faster and cheaper than any human. Based on the AI Security Institute’s research on “Measuring AI Agents’ Progress on Multi-Step Cyber Attack Scenarios” in a 32-step corporate network attack scenario in which the objective is to steal sensitive data from a protected internal database, the Opus 4.6 model from Anthropic destroyed the cyber range in just 6 hours for a dirt cheap $80 while a human expert would complete it in approximately 14 hours. In the United States, the annual entry-level salary of a pen tester is $70,200, which is about $33.75 per hour – stated by Cyber Security Jobs . It means that a penetration tester would get it done for $472,5 at more than twice the cost in time.(https://www.cybersecurityjobs.com/penetration-tester-salary/)(https://arxiv.org/pdf/2603.11214).

At what cost, though.

To use AI effectively in offensive security, giving it sensitive data isn’t optional. Consider a straightforward scenario: a company engages an LLM-based agent to assess their web application. Before the agent has touched a single endpoint, the client has already handed over access to the source code repository. The assessment hasn’t started. The exposure already has. Developer credentials, commit history, contributor identities, hardcoded secrets, passwords, API keys — all of it, sitting in context, before the first request is ever made.

If we return to that 32-step simulation, the full picture gets considerably more uncomfortable. Over the course of the attack chain, the AI collected default credentials, VPN configuration, plaintext domain user credentials, saved browser session credentials, internal web application access, and — the crown jewel — an entire password database. It impersonated privileged users, exfiltrated hidden environment variables, and planted multiple payloads. From a data governance perspective, this reads less like a penetration test and more like a corporate horror story.
The natural response is: no reasonable organization would grant that level of uncontrolled access and privilege to an autonomous agent. Which is a comforting thought, right up until you read the Fortra survey finding that 43% of respondents admitted to sharing work-sensitive information with AI models without telling their employers (https://www.fortra.com/blog/sensitive-work-data-ai-highlights-oh-behave-report). Apparently, 43% of employees didn’t get the memo. Or they got it, read it, and asked ChatGPT to summarize it anyway.

Your Threat Actor Just Got the Same Upgrade You Did

The same AI your vendor sold you to find vulnerabilities will find vulnerabilities. It just doesn’t particularly care who asked. Early 2026 reports confirm that AI chatbots were used to breach government systems, resulting in 195 million stolen records. The attackers manipulated AI models to generate exploit code and bypass safety guardrails — Los Angeles Times Reports (https://www.latimes.com/business/story/2026-03-05/how-our-ai-bots-are-ignoring-their-programming-giving-hackers-superpowers). It continues:

“hackers continuously prompted Claude in creative ways and were able to ‘jailbreak’ the chatbot to assist them. When they encountered problems with Claude, the hackers used OpenAI’s ChatGPT for data analysis.”

So the attacker’s entire operational toolkit was: a browser, two chat subscriptions, and enough creativity to rephrase the same question until the AI stopped protesting.
The Gambit Security report on the breach is well worth reading (https://cdn.prod.website-files.com/69944dd945f20ca4a27a7c47/69d8bb5aea59e31efb3b8a7f_Tech_Report_ai_breach_mex_gov.pdf). It states:

“Claude wrote s2_005_exploit.py — a 285-line Python script with proxy support, retry logic, and the full injection payload for remote command execution.”

Proxy support. Retry logic. Eight iterative edits. The AI didn’t just help, it debugged its own exploit until it worked. Which is either impressive or deeply inconvenient, depending on whether you’re reading the pentest report or the incident report.

No AI provider is handing their offensive agent directly to a state-sponsored APT. They don’t have to. Earlier this year, Bloomberg reported that Anthropic experienced unauthorized access to an early preview of its Claude Mythos model through a third-party vendor environment (https://www.bbc.com/news/articles/cy41zejp9pko) (https://www.bloomberg.com/news/articles/2026-04-21/anthropic-s-mythos-model-is-being-accessed-by-unauthorized-users). The model, that is part of the so-called “Project Glasswing” had been shared with a selected group of tech giants including Amazon, Apple, Google, NVIDIA, and Cisco (https://fortune.com/2026/04/07/anthropic-claude-mythos-model-project-glasswing-cybersecurity/). Access was reportedly leaked through an insider rather than a sophisticated attack, which is somehow both reassuring and more unsettling at the same time.

Conclusion

The math here is almost ironic: the more AI penetration testing agents are deployed, the more human pentesters are needed to validate what those agents actually did. AI can run an assessment in hours — it has never had to CC seven people just to get read access. But speed without thorough inspection is just a faster way to miss something.

“The only constant in life is change”

— and the security industry is hardly exempt. Penetration testing won’t disappear; it’ll just keep reinventing itself, as it always has. New attack surfaces emerge, methodologies evolve, and the humans who understand why something is a vulnerability will always matter more than the tool that found it.

Used with the right guardrails, AI accelerates the work that bogs experts down and frees them to focus on what genuinely requires human judgment. As Jensen Huang, CEO of NVIDIA, put it:

“You won’t lose your job to AI — you’ll lose your job to somebody who uses AI.”

At White Hat, we take that seriously. We stay current with the tools redefining the field — and we’re equally deliberate about the safeguards that ensure our clients’ data never becomes a casualty of moving fast.

Cybersecurity Services – White Hat IT Security

 

Written by: WhiteHat

Tagged as: , , , , , , , .

Previous post

Similar posts

Penetration Testing WhiteHat / August 21, 2026

From Auto-Login to Full RCE: Inside the IBM Langflow CVE

AI has been a hotspot for everything recently and that’s true for attacks too proven by the recently discovered and also exploited IBM Langflow OSS (versions 1.0.0 through 1.10.0) vulnerability CVE-2026-9198, our #CVEoftheWeek. The issue was addressed in the release of version 1.10.1 on 24th June, which contained many security updates, but known Proof-of-Concepts (PoC) ...

Read more trending_flat