MCP·Engineering
← Back to the mcp-security-scanner case study

A real sample report

// generated deliverable · mcp-security-scanner @ 49f556e · target: rag-mcp (my own server) · 2026-07-23
real, unedited output 1 P2 (low-confidence) clean P0/P1 bill print-to-PDF ready

This is exactly what the deliverable looks like — not a mockup. It was produced by running mcp-scan report against rag-mcp, one of my own production MCP servers, and pasting the generator's output in unchanged. I deliberately point the sample at my own code rather than a client's: the report over-flags by design, so the honest thing to show a prospective client is a scan of something I own, low-drama findings and all.

What you're reading is the raw scanner output — one low-confidence P2 hardening finding and a clean bill on the critical/high tiers. Every finding carries a file:line, a confidence, a reachability grade, and a specific fix, and the report states its own limits in plain language. On a paid engagement this same output gets a human triage pass layered on top; here it's shown unreviewed so you can see precisely what the tool emits before a person touches it.

One-Shot Audit — Sample (rag-mcp)

Self-contained, print-to-PDF report generated by mcp-scan report.

1. Engagement & scan metadata

ClientSample (rag-mcp)
EngagementOne-Shot Audit
Scope1 MCP server
Targetrag-mcp
Scan date2026-07-23
Scanner version (git)49f556e
Scanner test suite at that version410 passing (mcp-security-scanner suite at 49f556e; 419 total, 9 environment-gated skips)
Triage annotations(none supplied -- all findings unreviewed)

2. Executive summary

The scanner produced 1 raw finding(s) across 27 scanned file(s) in rag-mcp. No triage annotations were supplied: every finding below is UNREVIEWED. Treat this as the scanner's raw, deliberately over-flagging output, not a human-verified risk list.

Clean bill of health for rag-mcp: the scan found no critical (P0) or high (P1) findings. Per the methodology section, this means no critical/high patterns were found by these detectors -- not a guarantee of security.

Severity legend: P0 = Critical -- exploitable now, no preconditions; P1 = High -- exploitable with attacker-reachable but currently-unmet conditions; P2 = Medium -- defense-in-depth gap / conditional; P3 = Low -- hardening nit.

SeverityConfirmedFalse positiveAccepted riskMitigatedUnreviewedRaw
P0000000
P1000000
P2000011
P3000000

Top-3 risks (confirmed first, then unreviewed):

  1. P2Unreviewed rag_mcp/lock.py:144 — a flagged risky pattern (fc5f6407d38d).

3. Triage summary

No triage pass has been applied yet -- the raw count below is NOT a risk count. This scanner over-flags by design; expect a substantial share of raw findings to fall to human review.

VerdictCount
Confirmed0
False positive0
Accepted risk0
Mitigated0
Unreviewed1
Raw total1

4. Findings detail

fc5f6407d38d — file opened on a non-constant path without containment

P2confidence: lowverdict: Unreviewedreachability: cli-onlytaint: unknownclass: path-traversal

Locationrag_mcp/lock.py:144
What was foundA tool that opens a caller-derived path with no confinement (realpath/commonpath check) may allow ../ traversal outside its intended directory.
Evidencefd = os.open(
Reachability evidencereachable only from non-tool caller(s), never a registered MCP tool: _cmd_ingest (rag_mcp/cli.py:35)
Specific fixResolve the path (realpath) and assert it stays under an allowed base directory before opening.
Class remediation guidanceNo curated class-level guidance is mapped for this vulnerability class yet -- follow the finding-specific fix, and treat the flagged pattern as attacker-reachable until a human review proves otherwise.

5. Methodology & honest capability boundary

Stated plainly so the tool-vs-expert split is contractual, not verbal:

Static analysis only. Every finding comes from reading the target's git-tracked source (Python AST for the deep path; regex for Jinja/JS/TS/YAML/PowerShell/bash surfaces). No code is executed, no server is run, and no finding is proven exploitable at runtime -- findings are a prioritized review queue, not a verdict.

Manifest-aware tool reachability grading. After the detectors run, the scanner discovers registered MCP tools (decorator registrations, the low-level SDK's Server()/list_tools/call_tool shape, and any server.json manifest) and walks a static call-graph to label each finding reachable, cli-only, uncalled, or unknown. Reachable raises a finding's confidence; the others lower it; a finding is never dropped on this basis. Limits: the same-file call-graph is exact, cross-file is best-effort; module-level code, non-Python findings, ambiguous low-level-SDK dispatch, and repos with no discoverable tools are labelled unknown rather than guessed.

Tool-parameter taint tracking v1. A second pass seeds every registered tool handler's parameters as taint sources and propagates them through assignments, f-strings/concat/format, containers, and same-repo calls into the dangerous sinks, labelling findings tainted / untainted / unknown (again nudging confidence, never dropping). Limits, stated plainly: same-file dataflow is transitive but cross-file follows at most two direct-import hops (no third hop, no cross-repo flow); it is NOT sanitizer-aware (a validated/escaped value still counts as tainted, by design over-flagging); and dynamic dispatch (getattr / *args / **kwargs), decorator transforms, and non-Python surfaces are labelled unknown.

Confidence is load-bearing. This is a deliberately over-flagging scanner: low-confidence findings mean 'a human should glance at this' and produce false positives by design. The triage pass in this report is where a human separates the two -- the raw finding count is not the risk count.

The one law of suppression. Full suppression is reserved for the scanner's OWN curated exact-match judgment (e.g. documented SDK placeholder credentials, pagination-cursor field names, provably self-signed test certificates with test-identity CN/SAN). Every other signal -- a target-authored suppress comment, an obviously-fake value marker, a cert's test path -- may only demote confidence and tag the finding; it may never make a finding disappear. This matters because the scanner's use case is third-party, possibly-adversarial repos.

Not a git-history scanner. The secret detector reads the tracked working tree only -- it does not see credentials that were committed and later removed. Pair it with gitleaks for history scanning.

Language coverage. Deep (AST-based) for Python. JS/TS (.js/.mjs/.cjs/.ts/.mts/.cts/.jsx/.tsx) is line-based regex, not a parse tree: four detector families have JS/TS parity on that regex basis (param-injection, tool-scope-creep, secret-leak-via-tool-response, secret-in-log), while codegen-injection and auth-posture remain Python-only by scope decision. Jinja templates are regex-level. Known regex-heuristic gaps (e.g. optional-chaining eval, spread-of-secret returns) are documented in the README rather than silently missed.

Every finding needs a human. No dynamic analysis means no proof of exploitability: every finding in this report was either confirmed, dismissed, or left unreviewed by a named human triage pass, and the report says which, per finding. A clean bill means no critical/high patterns were found by these detectors -- not a guarantee of security.

6. Appendix

Detector families

ClassWhat it checks
codegen-injectionJinja autoescape off in a code-generating tool; hand-rolled string escaping instead of a real serializer.
param-injectionsubprocess(shell=True), os.system, non-constant eval/exec, pickle.load, unsafe yaml.load, SSRF, path traversal.
auth-postureBind 0.0.0.0 (+debug escalation), debug=True, mutating routes with no auth dependency, no rate limiter.
secret-handlingTracked .env/.pem/.key/keypair files, hardcoded secret-shaped literals, secrets passed to print/log.*.
tool-scope-creepA mutating @mcp.tool()-registered tool (by verb-name heuristic or a dangerous body sink, one hop through a delegated helper) with no visible permission-group gate or env-flag opt-in anywhere in its file or the helper it calls.
secret-leak-via-tool-responseA tool's return value that hands back os.environ, a whole config/settings object (vars()/asdict()/__dict__), a secret-named field, or a hardcoded secret-shaped literal to the calling LLM.
job-hazardsScheduled jobs/wrappers/IaC-CI files (cron, systemd, GitHub Actions, PowerShell/bash/batch deploy scripts): over-broad credential/ACL scope, a destructive call with no confirm-before-destroy gate, and success reported without verification (|| true, bare exit 0, continue-on-error, empty catch {}).

Scan configuration

  • Target: rag-mcp
  • Files scanned: 27
  • Scanner version: 49f556e
  • Scan notes/errors: none

Annotation provenance

  • No annotations file was supplied; every verdict above is Unreviewed.

Get this run against your MCP server  ← Back to the case study