Introducing AI Swarm Pentest: From One AI Pentester to an Orchestrated Team of Attack Agents

How specialized attack agents share context, challenge hypotheses, and independently verify attack paths inside the Plexicus Proof-Driven AppSec workflow.

José Palanco José Palanco
Last Updated:
25 min read
Share
Introducing AI Swarm Pentest: From One AI Pentester to an Orchestrated Team of Attack Agents

Beyond ASPM

Proof-Driven AppSec for teams building with AI

Plexicus uses AI Swarm Pentest to explore authorized application paths, validate what is exploitable, and give teams evidence they can use to prioritize remediation.

Explore AI Swarm Pentest

Software development is becoming agentic.

Code can now be generated, reviewed, refactored, tested, and deployed by AI systems. Applications are changing faster, APIs are multiplying, and security teams are being asked to validate more software without proportionally larger teams.

Penetration testing is beginning to change for the same reason.

The first wave of AI security tools largely helped humans work faster: explain a vulnerability, generate a payload, summarize scanner output, or suggest the next pentesting step.

The next wave became more autonomous. Instead of waiting for a human to issue every command, an AI agent could explore an application, interact with endpoints, test hypotheses, analyze responses, and decide what to try next.

But autonomy creates another problem:

A penetration test is not one task.

It is a long chain of decisions.

An attacker may need to discover an endpoint, understand authentication, identify an object reference, test authorization boundaries, obtain additional context, attempt several payload variants, correlate the result with another weakness, and finally prove that the complete attack path actually works.

That is where AI Swarm Pentest comes in.

Rather than treating penetration testing as one long conversation between a model and a target, AI Swarm Pentest treats it as an orchestrated security investigation in which specialized capabilities can explore, reason, share context, verify one another’s work, and turn attack hypotheses into evidence.

The goal is not simply:

Find more vulnerabilities.

The more useful question is:

Which attack paths are real, what can an attacker actually reach, and what evidence does the engineering team need to act?

That distinction is at the center of Plexicus AI Swarm Pentest. Plexicus describes AI Swarm Pentest as validating exploitable paths across authorized applications, APIs, and source code, with supporting evidence and independent verification before a finding is presented as verified.


Watch the AI Swarm Pentest motion graphic on YouTube.

First: What Does “AI Pentesting” Actually Mean?

There is no single universally accepted definition of AI penetration testing.

A vulnerability scanner with an LLM-generated explanation may be marketed as AI security.

A pentester using Claude or another model as a copilot may also call the workflow AI pentesting.

At the other end of the spectrum, autonomous systems can interact with applications, choose tools, generate payloads, analyze responses, change strategy, exploit vulnerabilities, and produce evidence with limited human intervention.

These systems are not equivalent.

Even vendors building autonomous pentesting platforms acknowledge the distinction. XBOW, for example, describes AI pentesting as a spectrum that can include an LLM wrapped around a scanner, human-led testing augmented by AI, and autonomous systems coordinating specialized agents.

So before discussing AI Swarm Pentest, it helps to separate four concepts.

ApproachMain Decision MakerTypical BehaviorMain Output
Traditional DAST / scannerRules and signaturesSends predetermined probes and matches known patternsPotential vulnerabilities
AI-assisted pentestingHuman pentesterAI explains, recommends, writes commands, or analyzes resultsHuman-validated findings
Autonomous AI pentestingAI agent or agent systemExplores targets, chooses actions, adapts to resultsAutonomous findings and attack evidence
AI Swarm PentestingOrchestrated agent systemMultiple capabilities explore, correlate, share context, challenge hypotheses, and verify attack pathsVerified attack paths with evidence

The boundaries are not absolute.

A sophisticated autonomous pentesting platform may already employ multiple agents internally. Therefore, “swarm” should not be interpreted as merely “more than one AI agent.”

The meaningful difference is in how those agents are organized.


What Is AI Swarm Pentest?

AI Swarm Pentest is an agentic penetration-testing approach in which specialized security capabilities operate as an orchestrated system instead of relying on one general-purpose AI agent to perform the entire engagement sequentially.

Think of the difference between asking one highly capable security engineer to investigate an application alone and giving a coordinated security team the same mission.

One person might enumerate the attack surface.

Another focuses on authentication.

Another investigates authorization.

Another examines APIs.

Another tries to turn an interesting behavior into a reproducible exploit.

Another independently attempts to reproduce the finding before the team reports it.

The important part is not simply that several people are working.

They need a shared understanding of the target.

If one agent discovers that:

/api/invoices/{id}

is reachable after authentication, another agent should not necessarily have to rediscover everything from zero.

If the first agent also discovers that invoice IDs are sequential, the authorization-testing capability can use that information as a new hypothesis.

If another action reveals a tenant identifier, that observation may become relevant elsewhere.

The pentest becomes a continuously evolving attack graph rather than a collection of isolated prompts.

The description of José Ramón Palanco’s upcoming APIAddicts presentation on October 15, 2026 proposes specialized skills, MCP, and graph-oriented attack memory. This is a proof-of-concept architecture proposed for the talk.

That is a much closer analogy to a real offensive-security engagement than:

Prompt → AI → vulnerability report.


Why One AI Agent Is Not Always Enough

Large language models are remarkably capable security assistants.

But penetration testing exposes several of their weaknesses at once.

A pentest is:

long-running,

stateful,

tool-heavy,

uncertain,

adversarial,

and dependent on evidence collected over many previous actions.

The PentestGPT paper documented this problem early. Researchers found LLMs useful for individual pentesting subtasks such as operating tools, interpreting output, and proposing subsequent actions, but maintaining an integrated understanding of the entire engagement remained difficult. PentestGPT therefore separated responsibilities into interacting modules rather than asking one model instance to manage everything.

Later research continued to identify challenges across enumeration, exploitation, privilege escalation, context management, and end-to-end autonomous reasoning.

More recent multi-agent research points in the same direction.

The ARTEMIS research system, for example, uses dynamic sub-agents and automated vulnerability triage. In one live university-network study, researchers reported advantages in systematic enumeration and parallel exploitation while also observing important limitations, including false positives and difficulty with some GUI-heavy workflows. These results should be evaluated within that study’s targets, scoring rules, and budget.

The lesson is not that “multiple agents automatically solve pentesting.”

They do not.

The lesson is that complex offensive-security work benefits from decomposing responsibility while maintaining shared state and verification.


From a Single Agent to a Swarm

A simplified autonomous AI pentester might look like this:

Target
  ↓
AI Agent
  ↓
Recon
  ↓
Hypothesis
  ↓
Exploit
  ↓
Report

This can work remarkably well.

But everything depends on the same reasoning loop preserving context, deciding what deserves attention, avoiding dead ends, handling tool output, validating results, and documenting the engagement.

A swarm-oriented architecture looks conceptually different:

                         ┌────────────────────┐
                         │ Authorized Scope   │
                         │ + Rules of         │
                         │ Engagement         │
                         └─────────┬──────────┘
                                   │
                                   ▼
                         ┌────────────────────┐
                         │ Swarm Orchestrator │
                         └─────────┬──────────┘
                                   │
                ┌──────────────────┼──────────────────┐
                ▼                  ▼                  ▼
       ┌────────────────┐ ┌────────────────┐ ┌────────────────┐
       │ Surface / Recon│ │ Auth / Logic   │ │ Exploitation   │
       │ Capability     │ │ Capability     │ │ Capability     │
       └───────┬────────┘ └───────┬────────┘ └───────┬────────┘
               │                  │                  │
               └──────────────────┼──────────────────┘
                                  ▼
                       ┌──────────────────────┐
                       │ Shared Attack Context│
                       │ / Graph / Evidence   │
                       └──────────┬───────────┘
                                  │
                                  ▼
                       ┌──────────────────────┐
                       │ Independent          │
                       │ Verification         │
                       └──────────┬───────────┘
                                  │
                                  ▼
                       ┌──────────────────────┐
                       │ Verified Finding     │
                       │ + Evidence           │
                       └──────────┬───────────┘
                                  │
                                  ▼
                       ┌──────────────────────┐
                       │ Human Review +       │
                       │ Remediation          │
                       └──────────────────────┘

The architectural shift is subtle but important.

The unit of work is no longer just the prompt.

It becomes the mission.


How AI Swarm Pentest Works

At a high level, an AI Swarm Pentest can be understood as seven connected stages:

  1. Define the authorized mission. The target, environment, credentials if required, limits, time boundaries, and rules of engagement are defined before testing begins. In Plexicus, scope and rules are agreed before the engagement, and the system is designed to stay within that authorized scope.

  2. Map the application surface. The system explores reachable applications, APIs, routes, parameters, authentication states, forms, and other useful attack surfaces rather than simply firing a fixed collection of signatures. Plexicus documentation distinguishes this behavior from DAST: its AI pentest agents can interact with applications, fill forms, use authenticated sessions, and pursue more complex attack chains.

  3. Generate attack hypotheses. Observations are transformed into questions. Can this identifier be changed? Can this internal endpoint be reached externally? Can a role boundary be bypassed? Can two individually moderate weaknesses be combined?

  4. Delegate exploration. Relevant capabilities investigate those hypotheses. Some paths die quickly. Others generate new information that becomes useful to the rest of the swarm.

  5. Connect evidence into attack paths. Instead of treating findings as isolated alerts, observations can become relationships: user → endpoint → object → permission → vulnerable operation → sensitive resource.

  6. Independently verify the finding. Discovery is not automatically equivalent to proof. Plexicus states that findings presented as verified undergo separate reproduction and that unsupported hypotheses can be discarded rather than included simply to make the report longer.

  7. Hand the evidence to humans. The result includes impact context, technical evidence, prioritization, and the information required for remediation. Production changes are not automatically applied; human teams retain the final decision.

This last stage matters.

The purpose of autonomous pentesting should not be autonomous vulnerability generation.

It should be autonomous investigation with reviewable proof.

Plexicus AI Swarm Pentest assessment overview showing configuration and captured target screenshots

The assessment overview shows the configuration and captured target screenshots while the pentest is running.


AI Swarm Pentest vs. Traditional DAST

DAST remains useful.

But DAST and AI Swarm Pentest answer different questions.

A conventional dynamic scanner is extremely good at repeatedly checking large numbers of targets for known classes of weaknesses.

It may ask:

Does this parameter react to a known SQL injection probe?

An agentic pentester can ask something closer to:

What does this endpoint do, what assumptions does the application make about the caller, and can I manipulate those assumptions to reach something I should not?

Plexicus documentation illustrates this difference directly. Its DAST mode sends automated probes for issues including HTTP security misconfigurations and known exploit patterns, while AI Pentest interactively explores applications, supports authenticated flows, and attempts complex attack chains and business-logic vulnerabilities.

That distinction matters because many real application vulnerabilities are contextual.

The problem may not be:

Parameter X accepts payload Y.

It may instead be:

A low-privilege user can obtain identifier X from workflow A, pass it into API B, bypass ownership validation, and access another tenant's resource.

No individual request necessarily looks extraordinary.

The relationship between the requests is the vulnerability.

That is exactly where attack-state reasoning becomes valuable.


AI Swarm Pentest vs. a “Normal” AI Pentest

This is the most important distinction—and also the easiest one to oversimplify.

An ordinary autonomous AI pentest often follows an agent loop:

observe → reason → act → observe → repeat.

That approach can be powerful.

But as the engagement becomes larger, the agent must simultaneously remember everything it has learned, decide priorities, operate tools, maintain authentication state, investigate many hypotheses, abandon unproductive paths, and distinguish genuine success from misleading output.

A swarm architecture can distribute those responsibilities.

DimensionTypical Single-Agent AI PentestAI Swarm Pentest
ReasoningPrimarily one reasoning loopOrchestrated specialized capabilities
ExplorationOften sequentialCan explore multiple hypotheses concurrently
ContextAgent conversation / working memoryShared mission context and attack relationships
SpecializationGeneral-purpose pentesting agentCapabilities can specialize around parts of the mission
Failed pathsManaged by the same agentCan be isolated without losing the entire investigation
Attack chainsAgent must preserve long-horizon contextFindings can be connected through shared state
ValidationMay be performed by the discovering agentIndependent reproduction can be a separate stage
OutputFinding/reportEvidence-backed attack path and handoff
Human roleDepends on implementationScope, guardrails, review, and final decision remain explicit

However, this table describes architectural patterns—not hard industry categories.

Some modern autonomous pentesting products already use specialized agents and validators. XBOW, for example, publicly describes coordinating specialized agents and using validator agents to confirm exploitability.

So the important question for buyers should never be:

“Does it have multiple agents?”

A better question is:

“How does the system coordinate them, preserve state, verify findings, control scope, and turn results into evidence my team can trust?”


The Swarm Is About Coordination, Not Agent Count

Imagine deploying 100 AI agents against an application.

If each one independently scans the same endpoints and then produces its own report, you technically have many agents.

You do not necessarily have a useful swarm.

A swarm becomes useful when knowledge discovered by one capability can influence the decisions of another.

Suppose an exploration capability discovers:

POST /api/projects/{project_id}/export

The authenticated user appears to belong to:

tenant_A

Another capability discovers a project ID belonging to:

tenant_B

An authorization-testing capability combines the observations.

It substitutes the second identifier into the first request.

The response unexpectedly succeeds.

A verification capability then reproduces the sequence using a clean state.

Now the system has something much more useful than:

Possible IDOR detected.

It has an attack chain:

Low-privilege user
        ↓
Authenticated project export endpoint
        ↓
Attacker-controlled project identifier
        ↓
Missing tenant ownership validation
        ↓
Cross-tenant project export
        ↓
Sensitive data exposure

And each step can carry evidence.

That is the difference between detecting an anomaly and demonstrating an attack path.

Plexicus AI Swarm Pentest embedded War room live board connecting to the stream

The embedded War room live board is shown while connecting to the stream.


Shared Context Changes What AI Can Investigate

Long-running penetration tests generate enormous amounts of temporary knowledge.

An agent may learn that:

an endpoint requires a specific token,

a user belongs to one role,

a parameter controls a backend object,

a service exposes an internal hostname,

a failed payload still revealed a framework version,

or one authentication flow creates credentials useful somewhere else.

If those observations remain trapped in the local context of the agent that discovered them, their value is limited.

A shared attack representation allows information to accumulate.

The description of the upcoming October 15 APIAddicts Days 2026 talk proposes a Plexicus proof of concept combining specialized skills, MCP, and a graph-oriented database for attack memory. The proposed representation connects assets, vulnerabilities, and exploitation paths.

Graph-like representations make intuitive sense for offensive security because attacks themselves are graphs.

A node might represent:

User
Endpoint
Credential
Repository
Service
Role
Vulnerability
Asset
Secret

An edge might represent:

CAN_ACCESS
AUTHENTICATES_TO
CALLS
OWNS
EXPOSES
DEPENDS_ON
BYPASSES
LEADS_TO

The interesting security question then becomes:

What new path exists through these relationships?

That is considerably closer to how experienced attackers reason.


Verification Is More Important Than Generation

Generative AI introduces an obvious problem into offensive security:

It can be confidently wrong.

A model can misinterpret a response.

A command can return an unexpected success code.

An application may behave inconsistently.

A payload may appear to work while the actual impact is different from what the agent concludes.

This is why an AI pentesting system should not equate:

“The model thinks it found a vulnerability”

with:

“A vulnerability has been verified.”

The offensive-AI industry is increasingly converging on this idea.

XBOW publicly describes using validator agents and requiring proof before validated findings are reported.

Plexicus applies the same broader principle inside its Proof-Driven AppSec model: findings presented as verified require reproducible evidence and separate confirmation. If a second pass cannot reproduce the result, Plexicus states that the finding remains unverified or is discarded.

That changes the optimization target.

A weak security AI system optimizes for:

How many vulnerabilities can we generate?

A proof-driven system should optimize for:

How many meaningful attack paths can we demonstrate with defensible evidence?

Expanded Plexicus AI Swarm Pentest board with four hunter agents, one skeptic, agent activity and narrative feed

The expanded board shows four hunter agents and one skeptic, per-agent activity, and the narrative feed. This snapshot shows zero findings and blocked attempts; it does not demonstrate a verified exploit.


Why Attack Chains Matter More Than Alert Counts

Security teams rarely suffer from a shortage of alerts.

A modern organization may already have results from:

SAST,

SCA,

DAST,

cloud scanners,

container scanners,

secret scanners,

CSPM,

dependency intelligence,

bug bounty,

and manual pentesting.

The harder problem is deciding what actually matters.

Consider three isolated findings:

A publicly reachable endpoint exposes a service identifier.

An internal service accepts an unvalidated redirect target.

A privileged endpoint trusts requests originating from that service.

Independently, none may appear catastrophic.

Together:

External attacker
      ↓
Public endpoint
      ↓
Internal service discovery
      ↓
SSRF / routing primitive
      ↓
Trusted internal request
      ↓
Privileged endpoint

The risk emerges from the path, not merely the findings.

This is why Plexicus positions AI Swarm Pentest around validated attack paths instead of an endless list of potential issues.


But Why Call It a “Swarm”?

The word can sound like marketing jargon.

It should not mean “we launched lots of agents.”

In a useful security swarm, capabilities should contribute toward a common mission.

The mental model is similar to a coordinated pentesting team:

                     Mission
                        │
           ┌────────────┼────────────┐
           │            │            │
        Explore       Reason       Exploit
           │            │            │
           └────────────┼────────────┘
                        │
                   Shared Context
                        │
           ┌────────────┼────────────┐
           │                         │
        Challenge                  Verify
           │                         │
           └────────────┬────────────┘
                        │
                     Evidence

Different capabilities can pursue different parts of the attack.

But they are not independent researchers writing unrelated notes.

They contribute to the same evolving representation of the target.

The intelligence of the system therefore comes not only from the model, but also from the orchestration around it.

That point is particularly important as open-weight models improve.

The upcoming APIAddicts talk description frames a related question: how orchestration, specialized skills, MCP tools, and persistent attack memory can help an open-weight model investigate an authorized target.


AI Swarm Pentest Inside Proof-Driven AppSec

Pentesting is useful.

But discovering an exploitable vulnerability is still only the beginning of the engineering workflow.

Someone must understand the affected code.

Someone must assess reachability and business impact.

Someone must determine the safest fix.

Someone must implement it.

Someone must review the change.

And someone should verify that the original exploit no longer works.

This is why Plexicus positions AI Swarm Pentest as one component of a broader Proof-Driven AppSec workflow.

The model can be summarized as:

VALIDATE
AI Swarm Pentest
      │
      │ exploitable path + evidence
      ▼
UNDERSTAND
Deep Code Analysis
      │
      │ code context + impact
      ▼
REMEDIATE
Reviewed Remediation
      │
      │ proposed fix
      ▼
RETEST
Does the attack still reproduce?

Plexicus’s current platform describes this as Validate → Understand → Remediate: AI Swarm Pentest explores authorized attack paths, Deep Code Analysis adds code-level context, and remediation can produce reviewer-ready changes followed by re-testing.

This connection is important.

Without it, autonomous pentesting risks producing a new version of an old security problem:

more findings than engineering teams can fix.


What AI Swarm Pentest Is Not

AI Swarm Pentest is not just a vulnerability scanner with AI-generated descriptions.

It is not a chatbot that tells you how an attacker might theoretically exploit something.

It is not permission for an autonomous agent to attack arbitrary infrastructure.

It is not a guarantee that every vulnerability will be discovered.

And it is not intended to eliminate human judgment from security decisions.

Plexicus explicitly describes AI Swarm Pentest as a scoped engagement with agreed rules of engagement, human decision points, independent verification, and no automatic modification of production systems.

Those guardrails are not limitations that should be removed.

They are part of what makes autonomous offensive security usable in a professional environment.


Can AI Swarm Pentest Replace Human Pentesters?

Not completely.

And that should not be the immediate goal.

Human pentesters remain particularly valuable when testing requires unusual creativity, deep organizational context, social reasoning, physical access, ambiguous business processes, or judgment about consequences that cannot be safely delegated to software.

AI has different advantages.

Machines can systematically enumerate large surfaces.

They can repeat tasks without fatigue.

They can pursue many hypotheses.

They can retest after changes.

And they can perform certain forms of parallel exploration at a scale that would be expensive to reproduce with human labor.

Research comparing AI agents and human cybersecurity professionals already shows this combination of strengths and limitations. Multi-agent systems can perform well in systematic enumeration and parallel exploitation, while still struggling in areas such as some GUI-heavy tasks and producing more false positives without sufficient verification.

So the more realistic future is not:

AI versus pentester

It is:

AI handles scalable exploration
        +
verification reduces noise
        +
humans control scope and judgment
        +
security experts investigate the hard edges

Plexicus similarly states that AI Swarm Pentest automates repetitive exploration, correlation, and reproduction without removing human context and final decision-making.


Where AI Swarm Pentest Can Be Especially Useful

The approach becomes particularly interesting for organizations where the software surface changes faster than periodic manual testing can reasonably follow.

A team may ship application changes daily.

APIs may appear and disappear between pentest cycles.

AI-generated code may significantly increase development throughput.

New dependencies and services may be introduced continuously.

A traditional annual pentest remains valuable, but it represents one snapshot.

Agentic testing makes it possible to move toward more frequent offensive validation.

Not merely:

“Did our scanner find something?”

But:

“Can an attacker still reproduce this attack path after the latest change?”

That is an important shift from periodic vulnerability discovery toward continuous security validation.


A Practical Example

Imagine an AI-generated SaaS application.

The application contains:

Web frontend
Authentication service
Billing API
Project API
File storage
Administrative API

A traditional scanner detects a few security headers and an outdated dependency.

Useful information—but not necessarily the highest-risk issue.

During an AI Swarm Pentest, the exploration stage discovers:

GET /api/projects/{project_id}

A normal user can retrieve their own project.

The swarm records the relationship:

USER_A → OWNS → PROJECT_123

Another observation reveals:

PROJECT_456 → OWNED_BY → USER_B

An authorization hypothesis is created.

The relevant capability requests:

GET /api/projects/456

while authenticated as User A.

The server returns metadata.

Interesting—but not yet enough.

The investigation continues.

The returned metadata contains:

export_id: EXP-8821

Another capability discovers:

GET /api/exports/{export_id}/download

The request succeeds.

The downloaded archive includes User B’s private project data.

The attack path now becomes:

User A
  ↓
Manipulate project identifier
  ↓
Read User B project metadata
  ↓
Discover export identifier
  ↓
Download export
  ↓
Cross-tenant data exposure

Now the verification stage repeats the chain independently.

If the exploit reproduces, the system can attach evidence to each step.

Deep code analysis can then trace the missing authorization control.

A remediation workflow can propose tenant-scoped validation.

After review, the original exploit can be replayed.

If it fails:

Attack reproduced before fix: YES
Attack reproduced after fix: NO

That is a much stronger security artifact than:

Potential IDOR — High

The Most Important Output Is Proof

AI will make generating security hypotheses extremely cheap.

That is both powerful and dangerous.

A model can generate hundreds of plausible explanations for why an application might be vulnerable.

But AppSec teams do not need hundreds of plausible explanations.

They need confidence.

They need to know:

What did you reach?

How did you reach it?

Can it be reproduced?

What is the impact?

Where does the vulnerable behavior originate?

What should we change?

Did the fix actually stop the attack?

That is why evidence becomes more important—not less—as AI becomes more capable.

The more autonomous security becomes, the stronger its verification layer needs to be.


The Future of Pentesting Is Not One Bigger Model

For several years, progress in AI was frequently described through model size.

A smarter model produced better answers.

Pentesting is showing why that mental model is incomplete.

Offensive security is not merely a reasoning benchmark.

It is a systems problem.

The AI needs tools.

It needs memory.

It needs permissions.

It needs state.

It needs a representation of the environment.

It needs to know when to explore.

It needs to know when to stop.

It needs to distinguish a promising hypothesis from a demonstrated exploit.

And it needs a mechanism for another process to challenge its conclusions.

That means the future of autonomous penetration testing may depend as much on architecture and orchestration as on raw model intelligence.

The progression looks something like this:

Security Scanner
      ↓
LLM Security Assistant
      ↓
Autonomous Pentest Agent
      ↓
Multi-Agent Pentesting
      ↓
Orchestrated AI Swarm
      ↓
Proof-Driven AppSec

Each stage does not necessarily replace the one before it.

Scanners remain useful.

Human pentesters remain useful.

Single-agent systems remain useful.

The question is which architecture best matches the complexity and frequency of the security problem being tested.


AI Swarm Pentest at Plexicus

Plexicus AI Swarm Pentest is designed around a simple principle:

Do not stop at detecting what might be vulnerable. Demonstrate what an attacker can actually reach.

Within an authorized scope, Plexicus explores application and API behavior, forms and tests attack hypotheses, connects relevant observations, and keeps supporting evidence attached to findings.

Findings presented as verified undergo an independent reproduction step.

The result is then connected to the rest of the Plexicus Proof-Driven AppSec workflow, where teams can add code context, understand impact, review remediation, and verify the result again.

The objective is not to produce the largest pentest report.

It is to give security and engineering teams something more useful:

A validated attack path.
Evidence showing why it is real.
Context explaining what it affects.
And a clear path toward fixing it.

Frequently Asked Questions

Is AI Swarm Pentest the same as automated vulnerability scanning?

No.

A scanner usually executes predetermined checks or probes and reports matching behavior. AI Swarm Pentest uses agentic reasoning to explore an authorized target, form hypotheses, interact with application behavior, and investigate attack paths.

Plexicus itself offers both DAST and AI Pentest modes and documents them separately.

Is AI Swarm Pentest just several LLMs running simultaneously?

No.

Parallelism alone does not create useful swarm behavior.

The critical elements are orchestration, specialized responsibilities, shared attack context, controlled execution, and verification.

Does every agent need a different AI model?

No.

Agent specialization and model specialization are different concepts.

Several agents can use the same underlying model while receiving different missions, tools, permissions, context, or validation responsibilities.

Conversely, an orchestration layer could choose different models for different tasks.

Is the “swarm” concept unique to Plexicus?

Multi-agent security systems are not unique to Plexicus.

Academic systems and commercial autonomous pentesting platforms also use multi-agent architectures, specialized agents, or validator agents.

Plexicus uses AI Swarm Pentest to describe its implementation of this approach within its broader Proof-Driven AppSec workflow, emphasizing shared attack context, evidence, independent verification, code analysis, and remediation.

Can AI Swarm Pentest find business-logic vulnerabilities?

Agentic testing is particularly relevant to business-logic and multi-step vulnerabilities because it can interact with workflows rather than relying solely on fixed signatures.

Plexicus documentation states that its AI Pentest can test authenticated flows and discover complex attack chains and business-logic flaws.

Does AI Swarm Pentest automatically attack production?

Testing must remain within an explicitly authorized scope and agreed rules of engagement.

Plexicus states that its engagement is scope-controlled and does not automatically modify production systems.

Does it eliminate false positives?

No autonomous security system should promise that.

The objective is to reduce uncertainty through reproducibility and evidence.

With Plexicus, a finding must survive a separate verification step before being presented as verified.

Does AI Swarm Pentest replace manual pentesting?

No.

It changes which parts of penetration testing can be automated and repeated at machine scale.

Human expertise remains important for scope, interpretation, specialized edge cases, and final decisions.


From Finding Vulnerabilities to Proving Attack Paths

AI is making it easier to generate code.

It is also making it easier to generate attacks.

Security validation has to evolve accordingly.

The answer cannot simply be another scanner producing another queue of alerts.

And adding an LLM to that scanner does not automatically solve the problem.

What changes the workflow is the ability to investigate dynamically:

Explore.
Form a hypothesis.
Test it.
Share what was learned.
Connect it to another observation.
Challenge the conclusion.
Reproduce the attack.
Keep the evidence.

That is the idea behind AI Swarm Pentest.

Not one AI pretending to be an entire red team.

An orchestrated system working toward one authorized security mission.

And most importantly:

No reproducible proof, no verified finding.


See What an Attacker Can Actually Reach

Plexicus AI Swarm Pentest explores authorized application and API attack paths, validates what is exploitable, and gives your team evidence it can review and act on.

Validate the path. Understand the impact. Fix it with evidence.

Explore AI Swarm Pentest →

Written by
José Palanco
José Palanco
José Ramón Palanco is the CEO/CTO of Plexicus, a pioneering company in ASPM (Application Security Posture Management) launched in 2024, offering AI-powered remediation capabilities. Previously, he founded Dinoflux in 2014, a Threat Intelligence startup that was acquired by Telefonica, and has been working with 11paths since 2018. His experience includes roles at Ericsson`s R&D department and Optenet (Allot). He holds a Telecommunications Engineering degree from the University of Alcala de Henares and a Master`s in IT Governance from the University of Deusto. As a recognized cybersecurity expert, he has been a speaker at various prestigious conferences including OWASP, ROOTEDCON, ROOTCON, MALCON, and FAQin. His contributions to the cybersecurity field include multiple CVE publications and the development of various open source tools such as nmap-scada, ProtocolDetector, escan, pma, EKanalyzer, SCADA IDS, and more.
Read More from José
More to read

Related posts

Ready to validate what matters?

Ready to validate what matters?

Plexicus is Proof-Driven AppSec: validated findings, contextual understanding, and reviewed remediation — anchored in evidence, scoped with you.

Qualification

Check whether AI Swarm Pentest fits your environment.

Share the minimum context. We will review the scope and tell you the next commercial step.

Before submitting — verify you fit

0 / 280

No commitment. If you don't fit, we'll tell you.

SAMPLE HANDOVER · ILLUSTRATIVE

Sample evidence handover

A trimmed view of what your team receives at the end of an AI Swarm Pentest engagement. Real engagements include full technical evidence, executive narrative, and a remediation plan.

VALIDATED FINDING Evidence attached

Server-Side Request Forgery in webhooks/receiver

demo-project/sample-app · src/webhooks/receiver.py:42

SeverityHigh CVSS 3.18.6 Priority79 Confirmedvia replay

Untrusted caller-supplied URLs reach an internal egress without an allowlist. Replayed in a sandbox against a fresh authorized target — the same control was validated to fail twice.

REVIEWER-READY REMEDIATION Merge-ready PR

Validate the target URL against an allowlist of permitted hostnames. Reject private/internal IP ranges. Enforce HTTPS only.

plexicus/remediation/webhooks-ssrf 3 changed · 0 new files
42resp = requests.get(target_url)
42+if not is_allowed_host(target_url):
43+  raise WebhookRejected(target_url)
44+resp = requests.get(target_url, timeout=5)
Every engagement hands over:
  • Executive briefing
  • Validated findings list
  • Merge-ready PRs
  • Compliance mapping (NIS2 · DORA · CRA)
Private Round For investors