AI Agent & MCP Security Audit

AI Agent & MCP Audit — An AI agent and MCP audit reviews systems where a language model can take actions — calling tools, moving funds, editing data — and establishes what an attacker achieves by controlling any text the model reads: prompt injection, tool poisoning, RAG poisoning, privilege escalation and unbounded autonomous action.

What a ai agent & mcp audit covers

An agent with tools is a remote code execution surface with a natural-language parser in front of it. Anything the model reads — a webpage, a document, a transaction memo, a GitHub issue, an MCP tool description — is untrusted input that can redirect its behaviour. The security question is not whether the model can be tricked; it is what the model is permitted to do once it has been.

We audit the blast radius first: every tool the agent can call, every credential it holds, every irreversible action it can take, and every boundary that is enforced by a prompt rather than by code. Then we attack it — direct and indirect injection, tool description poisoning, retrieved-context poisoning, confused-deputy chains across MCP servers — and report what survives.

Vulnerability classes we look for

Direct and indirect prompt injection

Instructions embedded in retrieved documents, web pages, file names, transaction memos or issue comments that redirect the agent while looking like data.

Tool poisoning and description injection

MCP tool descriptions and schemas carrying instructions to the model, including tools that redefine the behaviour of other tools after approval.

Excessive agency and irreversible actions

Agents able to transfer funds, delete data, deploy code or send messages without a human gate, and confirmation flows the model can satisfy itself.

RAG and memory poisoning

Untrusted content written into a vector store or long-term memory, persisting across sessions and influencing every later run.

Credential and secret exposure

Long-lived API keys and private keys inside the agent context, secrets returned in tool output, and credentials shared across tools with different trust levels.

MCP server authorisation and isolation

Servers running with more filesystem or network access than needed, missing origin and session validation, and cross-server confused-deputy chains.

Output handling and downstream injection

Model output executed as code, rendered as HTML, or passed to a shell, database or transaction builder without validation.

On-chain agent risk

Agents holding keys or signing transactions: spending limits, allowlists, slippage bounds, and whether a hostile prompt can produce a valid signed transaction.

In scope

Not in scope unless agreed

How the engagement runs

  1. Scoping and threat modelling

    We fix a commit hash, agree the in-scope contracts and read your architecture docs, then build a threat model: who the actors are, what the trust boundaries are, and which invariants must never break. Nothing is reviewed against assumptions we have not written down.

  2. Manual review

    Line-by-line review by at least two auditors working independently, focused on authorisation, accounting, upgrade paths, external integrations and the gap between what the code does and what the documentation claims it does. Most critical findings come from this phase, not from tooling.

  3. Static and dynamic analysis

    Static analysers appropriate to the language, plus property-based fuzzing and invariant testing to push the system into states no unit test covers. Tooling is used to widen coverage, never to replace the manual pass.

  4. Exploit-path simulation

    Candidate findings are proven on a forked network with a working proof of concept. We report what an attacker can actually do and what it costs them, not a theoretical severity label.

  5. Reporting

    Every finding gets a severity rating, reproduction steps, the affected code, the impact in concrete terms and a specific remediation. You get a draft for discussion before anything is finalised.

  6. Fix review and re-test

    We re-test every remediation against the original proof of concept and check that the fix has not opened a new path. The final report is yours to publish.

What you receive

How we rate severity

SeverityWhat it means
CriticalDirect loss of funds or permanent freezing of assets, exploitable by any actor.
HighLoss of funds or protocol insolvency under realistic conditions, or requiring a privileged actor to misbehave.
MediumBroken protocol behaviour, denial of service, or value leakage that does not directly drain the contract.
LowEdge-case incorrectness with limited impact, or an issue requiring implausible preconditions.
InformationalCode quality, gas efficiency, documentation mismatch and defence-in-depth suggestions.

Pricing

Single token contract: starts from $999, report in 24–48 hours. dApp, GameFi or RWA project: starts from $2,999. DeFi protocol, L2 / rollup, Bridge, ZK circuit, AI agent / MCP: scoped per project after we have seen the code.

AI Agent & MCP Audit: frequently asked questions

What is an MCP security audit?

A review of Model Context Protocol servers and the agents that use them: what each server exposes, how it authenticates callers, whether tool descriptions can inject instructions, whether one server can be used to reach another, and what an attacker achieves by controlling any content the model reads.

Can prompt injection be fixed?

Not eliminated by prompting — it is an input-trust problem, not a wording problem. It is contained by architecture: least-privilege tools, human gates on irreversible actions, provenance-aware context, output validation and spend limits. We audit whether those controls exist and hold.

Do you audit agents that control funds?

Yes, and those get the strictest treatment: signing authority, spend caps, allowlists, slippage bounds and the exact sequence by which a hostile prompt could produce a signed transaction. This is the highest-risk agent category in Web3.

What do you test that a normal pentest does not?

The model boundary. We treat every input the model reads as attacker-controlled and test whether tool descriptions, retrieved documents, memory entries and inter-server calls can redirect behaviour — categories a traditional application pentest does not cover.

Do you review the MCP server implementation itself?

Yes: transport and origin validation, session handling, filesystem and network scope, secret storage, and what a malicious client or a malicious server on the same host can reach.

What do you need from us to start an audit?

A repository or contract address, a commit hash to freeze the scope, whatever architecture or spec documentation exists, and a point of contact who can answer design questions. If documentation is thin we will write our understanding of the system back to you and ask you to confirm it — that step alone catches design-level bugs.

How long does an audit take?

A single token contract is 24–48 hours. A typical dApp or mid-sized protocol runs one to two weeks. Large DeFi systems, L2s, bridges and ZK circuits are scoped per project after we have seen the code. We will give you a fixed timeline with the quote, not an estimate that moves.

Is a re-test included after we fix the issues?

Yes. Fix review is part of the engagement, not an upsell. We re-run the original proof of concept against your patched code and confirm the fix has not introduced a new path.

Related security services

Get a fixed quote in 24 hours

Send the repository and a commit hash through the contact form, message @bugtester25 on Telegram, or book a 30-minute scoping call. 200+ protocols audited · $4B+ secured · 0 hacks post-audit. Prefer email? info@safeedges.in.