Problem Threats Stories How it works Blogs Pricing
Security Testing

Automated vs manual penetration testing: what AI can test, and what still needs a human

AI agents now find real bugs at scale, but they still miss flaws that need context. Here is where each approach wins, and how to use both.

By Arxiis ResearchUpdated 21 min read

Key takeaways

  • AI pentesting is real: in 2025 an autonomous tool reached the top of HackerOne's US leaderboard after submitting nearly 1,060 vulnerability reports, though its team reviewed each one before submission.1
  • In DARPA's AI Cyber Challenge final (August 2025), AI systems found 86% of the planted flaws across 54 million lines of code and patched 68% of the ones they found.2
  • AI agents still struggle with long attack chains: in one 2026 study, 58% of agent failures came from poor planning and tracking, not missing tools.3
  • OWASP says automating business logic abuse cases is not possible and still relies on the skill of the tester.4
  • Indian rules still ask for people: RBI's 2026 NBFC Directions want VA and PT done by trained, independent experts, and SEBI's CSCRF names CERT-In empanelled auditors.56
  • A hybrid model runs automated testing every week or day between human-led pentests; with releases spread evenly, that cuts the average wait for a first test from about 182 days (annual pentest) to 3.5 days or less.

A developer ships a new login page on a Tuesday. Your pentest ran in March and the next one is booked for March next year. For the next eleven months, nobody checks whether that page lets someone skip the one-time password. Attackers do not wait that long: CrowdStrike puts the average eCrime breakout time at 29 minutes.7 So the real question is not "tool or human". It is what each one can check, and how often.

Last reviewed on 5 October 2026 against OWASP WSTG, DARPA AIxCC results, the Verizon 2026 DBIR, RBI's 2026 NBFC Directions and SEBI CSCRF.

What is manual penetration testing?

Manual penetration testing is a time-boxed attack on your systems run by skilled people. Testers use tools, but they choose the targets, read the responses, and chain small weaknesses into real impact. The output is a report with proof, risk ratings and fixes, usually delivered once or twice a year.

A manual test follows a plan. The team agrees on scope and rules, maps what is exposed, tries to get in, and then tries to go further. A good tester asks questions a tool does not think of. What happens if I change the order ID in this request? What if I pay for one item and ship ten? Can this help-desk account reset an admin password?

That judgement is the main strength. The main limits are time and cost. A test is a snapshot of the systems on the days it ran. If you want the basics of how a pentest differs from a scan, read our guide on vulnerability assessment vs penetration testing.

How often do companies run manual pentests?

Not often. Cobalt's 2025 survey found that quarterly is the most common cadence (30% of respondents), while 27% test only once a year.8 Meanwhile, the Verizon 2026 DBIR found that exploiting vulnerabilities is now the top way attackers get in, at 31% of breaches, up from 20%.9 Only 26% of known exploited vulnerabilities were fully fixed, with a median of 43 days to fix.9

29 minAverage eCrime breakout time in 2025CrowdStrike 2026
31%Breaches that began with an exploited vulnerabilityVerizon DBIR 2026
27%Organisations that pentest only once a yearCobalt 2025
59%Security teams reporting critical or significant skills gapsISC2 2025

What is automated (and AI) penetration testing?

Automated penetration testing uses software to run attack steps without a person driving each one. Older tools follow fixed scripts. Newer AI pentesting tools use language models to plan steps, read results and decide what to try next. Both can run often, on a schedule, and re-test fixes on demand.

Autonomous penetration testing
A form of automated testing where AI agents choose their own next steps inside set limits (scope, allowed techniques, stop rules), try to exploit what they find, and record proof. "Autonomous" describes who picks the next move. It does not mean "no human oversight".

It helps to separate three kinds of tool, because vendors blur them:

  • Vulnerability scanners check for known weaknesses and report them. They rarely prove that a flaw can be used.
  • Scripted attack tools (often called breach and attack simulation) replay known attacker techniques to see if your controls catch them.
  • AI pentesting agents try to exploit what they find, chain steps, and adapt when a step fails. Some work as a team of agents, for example one for recon and one for exploitation.

Gartner now groups the last two under adversarial exposure validation. Its market definition says these tools must offer automated scheduling so teams can test more often without human intervention.10 Hadrian's summary of Gartner's March 2026 Market Guide says this category replaces the older "breach and attack simulation" and "automated penetration testing and red teaming" labels.11 The same idea sits inside continuous threat exposure management, which we explain in our CTEM guide.

Is pentest as a service the same thing?

No. Pentest as a service (PTaaS) is a delivery model. You book tests through a platform, see findings as they land, and ask for retests online. The testing itself may be human, automated or both. Ask each provider which parts a person does.

Automated vs manual penetration testing: side by side

Automated vs manual penetration testing is a trade between frequency and judgement. Automated and AI testing wins on speed, repeatability and scale. Manual testing wins on business logic, complex chains and context. For compliance, most regulators still expect a qualified, independent person to own the test and the report.

Automated (AI) vs manual penetration testing across 11 factors
FactorManual testingAutomated and AI testing
SpeedDays to weeks per test, plus report writingHours per run; can run after every release
Cost per testHigh, priced by tester days (see VAPT cost in India)Low per run once licensed; cost is spread over many runs
CoverageLimited to the scope and days bookedWide; can sweep every internet-facing asset each run
DepthDeep on chosen targetsGood on known patterns; uneven on new or odd targets
Business logicStrong; needs human understanding of the processWeak; tools do not know what the business intends
Attack chainingStrong; testers plan multi-step pathsPartial; agents lose track on long chains
RepeatabilityVaries by tester and by dayHigh; the same checks run the same way each time
False positivesLow when testers confirm each findingVaries; AI can report flaws that do not exist unless it proves them
EvidenceScreenshots and narrative, written by handLogs, requests and proof captured automatically
Compliance acceptanceWidely accepted when done by qualified testersUseful as supporting evidence; rarely enough on its own
ScaleLimited by skilled peopleGrows with compute, not headcount

The table shows why this is not a "pick one" choice. The two approaches are strong in different rows. The skills side matters too: ISC2's 2025 study of 16,029 professionals found 95% of teams have at least one skills gap, and 59% call their gaps critical or significant, up from 44% in 2024.12 Skilled testers are scarce, so it makes sense to spend their hours where a tool cannot help.

What can AI pentesting do well in 2026?

AI pentesting in 2026 is good at fast, wide, repeatable work: mapping exposed assets, confirming known vulnerabilities, trying common credential attacks, and re-testing fixes. Public results from HackerOne, DARPA and peer-reviewed studies show real findings at scale. The best results still pair the AI with human review.

Evidence from the real world

  • Bug bounty results. XBOW, an autonomous pentesting tool, reached the top of HackerOne's US leaderboard in 2025. It submitted nearly 1,060 vulnerability reports; at the time of its write-up, 130 were resolved and 303 triaged, while 208 were duplicates and 209 were marked informative.1 The company says the findings were fully automated but its security team reviewed them before submission to meet HackerOne's policy.1
  • Code-level flaw finding. In the DARPA AI Cyber Challenge final, AI systems analysed 54 million lines of code, found 54 of 63 planted flaws (86%), and patched 68% of what they found. They also found 18 real, unplanted vulnerabilities. Patches took about 45 minutes on average, at about USD 152 per task.2 Note that this contest tested source code analysis, not network attacks.
  • Research tools. The PentestGPT paper (USENIX Security 2024) reported a 228.6% rise in task completion over the plain GPT-3.5 model on its benchmark.13 The CAI framework's authors report it was up to 3,600 times faster than humans on some tasks and 11 times faster on average, with a human in the loop.14

Where that speed pays off

  • Recon and exposure mapping across hundreds of hosts and apps, every run.
  • Known CVE exploitation, which matters because exploitation is now the top breach entry point.9
  • Credential attacks such as password spraying and reused passwords. Read how these play out in They Don't Break In. They Log In.
  • Retesting fixes the same day a patch lands, instead of waiting for the next engagement.
  • Consistent evidence: every request and response is logged, which makes findings easier to reproduce.

Read the fine print on AI claims. Many headline numbers are self-reported, measured on benchmarks, or involve human review before anything is submitted. HackerOne itself warns that AI-generated reports can include hallucinated vulnerabilities that look real, and says strong validation is needed to tell them apart.15 Ask any vendor for proof of exploit, not just a list of findings.

Where do human testers still win?

Human penetration testers still win on business logic flaws, long multi-step attack paths, complex login and role systems, social engineering, and judging real business impact. These tasks need context about how the organisation works. Research on AI agents shows their weak spot is planning over many steps, not tool use.

Capability matrix: who tests what well10 tasks · 3 approaches
Capability matrix: manual vs automated AI vs hybrid penetration testing Capability matrix rating ten penetration testing tasks for manual testing, automated AI testing and a hybrid of both. Recon and asset discovery: manual partial, automated AI strong, hybrid strong. Known CVE exploitation: manual partial, automated AI strong, hybrid strong. Credential attacks: manual strong, automated AI strong, hybrid strong. Attack path chaining: manual strong, automated AI partial, hybrid strong. Business logic flaws: manual strong, automated AI weak, hybrid partial. Complex auth flows: manual strong, automated AI partial, hybrid strong. Social engineering: manual strong, automated AI weak, hybrid partial. Retesting fixes: manual weak, automated AI strong, hybrid strong. Reporting and evidence: manual strong, automated AI partial, hybrid strong. Testing frequency: manual weak, automated AI strong, hybrid strong. Manual totals 6 strong, 2 partial, 2 weak; Automated AI totals 5 strong, 3 partial, 2 weak; Hybrid totals 8 strong, 2 partial, 0 weak. TESTING TASK MANUAL AUTOMATED AI HYBRID Recon and asset discovery finding what is exposed Partial Strong Strong Known CVE exploitation proving public flaws are real Partial Strong Strong Credential attacks spraying, weak and reused passwords Strong Strong Strong Attack path chaining linking 5+ steps to a crown jewel Strong Partial Strong Business logic flaws abusing how the app is meant to work Strong Weak Partial Complex auth flows SSO, MFA, multi-role access Strong Partial Strong Social engineering phishing and pretext calls Strong Weak Partial Retesting fixes confirming a patch closed the hole Weak Strong Strong Reporting and evidence proof, context, fix advice Strong Partial Strong Testing frequency how often each change gets tested Weak Strong Strong Count of strong ratings 6/ 10 5/ 10 8/ 10 StrongPartialWeak Hybrid = automated between human tests Capability matrix (mobile layout): manual vs automated AI vs hybrid Capability matrix rating ten penetration testing tasks for manual testing, automated AI testing and a hybrid of both. Recon and asset discovery: manual partial, automated AI strong, hybrid strong. Known CVE exploitation: manual partial, automated AI strong, hybrid strong. Credential attacks: manual strong, automated AI strong, hybrid strong. Attack path chaining: manual strong, automated AI partial, hybrid strong. Business logic flaws: manual strong, automated AI weak, hybrid partial. Complex auth flows: manual strong, automated AI partial, hybrid strong. Social engineering: manual strong, automated AI weak, hybrid partial. Retesting fixes: manual weak, automated AI strong, hybrid strong. Reporting and evidence: manual strong, automated AI partial, hybrid strong. Testing frequency: manual weak, automated AI strong, hybrid strong. Manual totals 6 strong, 2 partial, 2 weak; Automated AI totals 5 strong, 3 partial, 2 weak; Hybrid totals 8 strong, 2 partial, 0 weak. MANUAL AUTOMATED AI HYBRID Recon and asset discovery Partial Strong Strong Known CVE exploitation Partial Strong Strong Credential attacks Strong Strong Strong Attack path chaining Strong Partial Strong Business logic flaws Strong Weak Partial Complex auth flows Strong Partial Strong Social engineering Strong Weak Partial Retesting fixes Weak Strong Strong Reporting and evidence Strong Partial Strong Testing frequency Weak Strong Strong 6 of 10 strong 5 of 10 strong 8 of 10 strong StrongPartialWeak
Arxiis Research rating, October 2026, based on OWASP WSTG, USENIX Security 2024 (PentestGPT), arXiv 2602.17622, Cybench (ICLR 2025), DARPA AIxCC and HackerOne. Ratings are a judgement from public evidence, not a benchmark score.

Business logic

OWASP's testing guide is direct about this. It says business logic flaws "cannot be detected by a vulnerability scanner" and that automating abuse cases "is not possible and remains a manual art relying on the skills of the tester."4 AI agents can read an app better than a scanner can, but they still do not know your rules. Only a person who understands your loan approval flow will notice that an applicant can edit the sanctioned amount after approval.

Long attack chains

Real breaches often join five or more small steps. A February 2026 study of LLM pentest agents found that 58% of failures came from planning and state tracking problems that better tools did not fix. On tasks with five or more steps, 79% of failures were of this kind.3 The original PentestGPT work noted the same weakness: models struggle to keep the whole test scenario in mind.13 This is why internal network and Active Directory penetration testing still benefits from a human steering the path.

Hard problems take time

The Cybench benchmark (ICLR 2025) used 40 professional capture-the-flag tasks. Without extra guidance, the AI agents only solved tasks that took human teams up to 11 minutes. The hardest task took human teams 24 hours and 54 minutes.16 Models have improved since then, but the gap on hard, novel problems is still where human testers earn their fee.

Judgement and trust

  • Social engineering needs consent, care and a human voice.
  • Impact: a person can explain why a low-scoring flaw matters to your board.
  • Safety calls: a tester can stop before an exploit harms a fragile production system.
Automation buys you frequency. People buy you judgement. You need both, on different clocks.Arxiis Research

Is automated pentesting accepted for compliance?

Automated pentesting for compliance is usually accepted as supporting evidence, not as a full replacement. RBI's 2026 NBFC Directions ask for trained, independent experts, SEBI's CSCRF names CERT-In empanelled auditors, and PCI DSS asks for qualified, independent testers. Automation helps you test between those audits and prove fixes.

What key rules say about who tests and how often
RuleWhat it asksWhere automation helps
RBI NBFC Directions 2026 (RBI/DoS/2026-27/461)VA at least every 6 months and PT at least every 12 months for critical and customer-facing DMZ systems (clause 121); done by "appropriately trained and independent" experts or auditors (clause 122); VA/PT on production after go-live (clause 124)5Testing after each release, and between the yearly PT
SEBI CSCRF (20 August 2024)VAPT by a CERT-In empanelled IS auditing organisation; findings closed within 3 months of the report617Tracking and re-testing findings well before the 3-month deadline
PCI DSS v4.0.1Internal and external pentests at least once every 12 months and after significant changes (11.4.2, 11.4.3); qualified and organisationally independent tester (11.4.1); retest fixes (11.4.4)18Change-driven testing and fast retests
ISO 27001 (A.8.8) and SOC 2 (CC4.1)No fixed pentest frequency; you must show a risk-based processEvidence that testing is ongoing

The pattern is clear. Regulators care about who stands behind the test. A tool can do much of the work, but a qualified and independent person or firm usually has to own the scope, check the results and sign the report. For the Indian rules in detail, see our guides to RBI Cyber Security Directions 2026 and SEBI CSCRF VAPT requirements.

Automation still helps a lot with compliance. RBI's clause 124 asks for testing on production after an IT project or upgrade goes live.5 SEBI wants findings closed within 3 months.17 PCI DSS wants testing after significant changes.18 All three are easier when a test can run the day a change ships.

Not legal advice. This section summarises public texts as of October 2026. Rules differ by entity type and category, and regulators update them. Check the latest official text, and ask your auditor whether a given tool's output is acceptable for your audit.

How do you evaluate an AI pentesting tool?

Evaluating an AI pentesting tool comes down to control, proof and trust. Check that it stays in scope, exploits safely, proves every finding, re-tests fixes, maps results to MITRE ATT&CK, keeps your data where you need it, logs every action, and lets a human review before anything risky happens.

Use this checklist in a proof of concept. Run the tool against a test copy of a real app you know well, so you can judge what it missed.

  • Safe exploitation controls. Can you block destructive actions, set rate limits and stop a run at once?
  • Scope guardrails. Does it enforce an allow-list of hosts, domains and time windows, and refuse anything outside it?
  • Proof of exploit. Does each finding include the request, the response and the impact, so you can replay it?
  • False positive handling. What share of findings did your team confirm? Ask for the vendor's own validation method.
  • Retest on demand. Can you re-run one finding after a fix, without a full scan?
  • MITRE ATT&CK mapping. Are techniques mapped to the current matrix, which has 15 enterprise tactics since v19?
  • Data residency and on-premise option. Where do scan data, credentials and model prompts go? Can it run fully inside your network?
  • Audit logs. Is every agent action logged with time, target and result, and can you export it for auditors?
  • Model choice. Which AI models does it use, and can you choose or host them yourself?
  • Human review. Can a person approve high-risk steps and sign off the report?
  • Compliance mapping. Can reports map findings to the frameworks you answer to, such as RBI, SEBI CSCRF or PCI DSS?

A simple test. Plant two known flaws in a staging app: one common (an outdated library with a public exploit) and one business logic flaw (a price you can edit in the cart). A good tool should prove the first. Few will find the second. That tells you where your human hours should go.

The hybrid model that most teams land on

Hybrid penetration testing runs automated and AI tests often, weekly or after every release, and keeps human-led pentests for depth and compliance. Automation catches known and repeatable issues within days. People focus on business logic, chained attacks and new features. Each human finding becomes a check the tools repeat.

  1. Keep the human test your rules require. Book it with a qualified, independent tester at the frequency your regulator sets.
  2. Automate between tests. Run AI testing weekly on internet-facing apps, and after every production release.
  3. Demand proof and route it. Send only proven findings to the owning team's ticket queue, with a due date.
  4. Retest automatically. Close a finding only when a re-run confirms the fix.
  5. Aim people at the hard parts. Give testers the business logic, new features and long attack paths.
  6. Feed findings back. Turn each human finding into a repeatable automated check.

Why does the cadence matter so much? If releases are spread across the year and you test once a year, a change waits about six months, on average, for its first test. That is the gap we describe in The 363-Day Blind Spot. Use the calculator below to see your own numbers.

Interactive · Testing gap calculator

How long does a new change wait for its first test?

Enter your estate and your testing rhythm. See how many releases go untested for over a week, and how a hybrid model changes that.

Automated testing runs
Releases untested for over 7 days
98%
Average days a change waits
182.5
Releases per year
960

Bars show untested change-days per year (releases × average wait).

With 1 pentest a year and no automated testing, 960 releases a year wait about 182.5 days on average for their first test. Adding weekly automated testing would cut total untested change-days by about 98%.

Indicative only. Formula: days between tests T = 365 ÷ pentests per year, or 7 (weekly) or 1 (daily) when automated testing runs. Releases per year N = apps × releases per app × 12. Assuming releases land evenly, average wait = T ÷ 2, share waiting over 7 days = (T − 7) ÷ T (zero if T ≤ 7), and untested change-days = N × T ÷ 2. It assumes each test covers every live change. Automated tools cover only some flaw types (see the capability matrix), so hybrid figures are a best case for those types.

See what weekly testing would find →

What the numbers say

Moving from an annual test to quarterly still leaves the average change untested for about 46 days. That is longer than the 43-day median Verizon found for fixing known exploited flaws.9 Weekly automated runs bring the average wait down to 3.5 days. The human test still matters, because the automated run will not catch the logic flaw in your checkout flow.

Frequently asked questions

Can AI replace manual penetration testing?

No, not fully in 2026. AI tools now find real vulnerabilities at scale and re-test fixes quickly, but they still struggle with business logic flaws, long attack chains and judging business impact. Most regulators also expect a qualified, independent person to own the test. The practical answer is a hybrid: AI testing runs often, and human testers focus on the parts that need context.

Is automated penetration testing accurate?

It depends on whether the tool proves what it reports. Tools that exploit a flaw and capture the request and response give reliable findings. Tools that only guess can produce false positives, and HackerOne has warned that AI-generated reports may include hallucinated vulnerabilities that look real. Ask for proof of exploit on every finding and check a sample during your trial.

How often should you run automated penetration testing?

Run it at least weekly on internet-facing systems, and after every production release if you can. Attackers move fast, with CrowdStrike reporting a 29-minute average eCrime breakout time. Weekly runs cut the average wait for a first test to about 3.5 days, compared with about 182 days for a single annual pentest. Keep your human-led tests at the frequency your regulator requires.

Does automated pentesting count for RBI, SEBI or PCI DSS compliance?

Usually only as supporting evidence. RBI's 2026 NBFC Directions ask for VA and PT by trained, independent experts. SEBI's CSCRF names CERT-In empanelled auditing organisations for VAPT. PCI DSS asks for a qualified, organisationally independent tester. Automated testing helps you test after changes, close findings faster and prove fixes, but confirm with your auditor what they accept.

Is autonomous penetration testing safe to run in production?

It can be, if the tool has strong guardrails. Look for a strict scope allow-list, blocks on destructive actions, rate limits, a kill switch and a full audit log of every action. Start in staging, then move to production in a maintenance window. RBI's 2026 NBFC Directions ask for testing on production after go-live, so safe production testing is a real need.

What is the difference between a vulnerability scan and automated penetration testing?

A vulnerability scan lists weaknesses that might exist, based on versions and settings. Automated penetration testing goes further and tries to exploit those weaknesses, then records proof of what an attacker could reach. AI pentesting tools can also chain steps and change tactics when one fails. A scan tells you which doors look weak. A pentest tries the handles.

What is pentest as a service (PTaaS)?

Pentest as a service is a way of buying penetration tests through an online platform instead of a one-off project. You can book tests, see findings as they are found, talk to testers and request retests in one place. The testing may be manual, automated or both, so ask each provider exactly which steps a human performs and which a tool performs.

Where Arxiis fits

Automated testing every day, with humans on the hard parts

The gap in most programmes is not skill. It is time between tests. Arxiis is an autonomous AI red teaming and penetration testing platform built to fill that gap, so your team can run real attack tests often and save human hours for the work only people can do.

  • A multi-agent AI crew covers OSINT, exploitation, lateral movement and reporting across 26 security modules and 6 attack vectors: ransomware, Active Directory, cloud, web applications, containers and credentials.
  • Findings are CVSS-scored and mapped to MITRE ATT&CK, with compliance overlays for 11 frameworks, including RBI, CERT-In, SEBI CSCRF, PCI DSS and DPDP 2023.
  • Fully on-premise deployment, so your data never leaves your environment, with an MIT-licensed open-source core you can inspect.
  • A pentest report in hours instead of weeks, so you can retest a fix the day it ships.

Your company gets defended every day.

Sources

  1. XBOW, "How XBOW Ranked #1 in Autonomous Penetration Testing (The road to Top 1)", 24 June 2025. https://xbow.com/blog/top-1-how-xbow-did-it
  2. DARPA, "AI Cyber Challenge marks pivotal inflection point for cyber defense", 8 August 2025. https://www.darpa.mil/news/2025/aixcc-results
  3. arXiv 2602.17622, "What Makes a Good LLM Agent for Real-world Penetration Testing?", 19 February 2026. https://arxiv.org/html/2602.17622v1
  4. OWASP, "Web Security Testing Guide: Introduction to Business Logic Testing", stable edition, accessed October 2026. https://owasp.org/www-project-web-security-testing-guide/stable/4-Web_Application_Security_Testing/10-Business_Logic_Testing/00-Introduction_to_Business_Logic
  5. Reserve Bank of India (text as reproduced by TaxGuru), "Non-Banking Financial Companies: Cybersecurity, Technology Risk, Resilience and Assurance Framework Directions, 2026 (RBI/DoS/2026-27/461), clauses 121, 122 and 124", 31 July 2026. https://taxguru.in/rbi/rbi-issues-nbfc-cybersecurity-technology-risk-directions-2026-governance-framework.html
  6. SEBI, "Cybersecurity and Cyber Resilience Framework (CSCRF) for SEBI Regulated Entities, circular SEBI/HO/ITD-1/ITD_CSC_EXT/P/CIR/2024/113", 20 August 2024. https://www.sebi.gov.in/sebi_data/attachdocs/aug-2024/1724326790365.pdf
  7. CrowdStrike, "2026 CrowdStrike Global Threat Report: AI Accelerates Adversaries and Reshapes the Attack Surface", 24 February 2026. https://ir.crowdstrike.com/news-releases/news-release-details/2026-crowdstrike-global-threat-report-ai-accelerates-adversaries
  8. Cobalt, "State of Pentesting Report 2025", April 2025. https://resource.cobalt.io/state-of-pentesting-2025
  9. Verizon, "2026 Data Breach Investigations Report", May 2026. https://www.verizon.com/business/resources/reports/dbir/
  10. Gartner Peer Insights, "Adversarial Exposure Validation: market definition", accessed October 2026. https://www.gartner.com/reviews/market/adversarial-exposure-validation
  11. Hadrian, "What the 2026 Gartner Market Guide for Adversarial Exposure Validation means for offensive security", 2026. https://hadrian.io/blog/what-the-2026-gartner-r-market-guide-for-adversarial-exposure-validation-means-for-offensive-security
  12. ISC2, "2025 ISC2 Cybersecurity Workforce Study", 4 December 2025. https://www.isc2.org/Insights/2025/12/2025-ISC2-Cybersecurity-Workforce-Study
  13. USENIX Security 2024 (Deng et al.), "PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration Testing", August 2024. https://www.usenix.org/conference/usenixsecurity24/presentation/deng
  14. arXiv 2504.06017 (Alias Robotics and co-authors), "CAI: An Open, Bug Bounty-Ready Cybersecurity AI", April 2025. https://arxiv.org/abs/2504.06017v2
  15. HackerOne, "What We've Learned from 5 Months of Hackbot Activity", 26 June 2025. https://www.hackerone.com/blog/ai-hackbots-security-testing-update
  16. Zhang et al., ICLR 2025, "Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models", August 2024. https://arxiv.org/abs/2408.08926
  17. SEBI, "Frequently Asked Questions on Cybersecurity and Cyber Resilience Framework (CSCRF), Q.15", June 2025. https://www.sebi.gov.in/sebi_data/faqfiles/jun-2025/1749647139924.pdf
  18. PCI Security Standards Council, "Just Published: PCI DSS v4.0.1 (requirements 11.4.1 to 11.4.4)", 11 June 2024. https://blog.pcisecuritystandards.org/just-published-pci-dss-v4-0-1
Arxiis Research

Written by the Arxiis research team. Facts checked against primary sources on 5 October 2026. Not legal advice.