<< All versions

Skill v1.0.0

currentAutomated scan95/100
vulpy-io/vulpy-commerce-core/security-audit
──Details
PublishedOctober 1, 2026 at 03:59 AM
Content Hashsha256:f251695863124ab1...
Git SHA55bb96e12cfc
──Files
Files (1 file, 15.4 KB)
SKILL.md15.4 KBactive
SKILL.md · 75 lines · 15.4 KB

version: "1.0.0" name: security-audit description: Adversarial security audits of implementations/services before merge or release — webhook fail-closed checks, auth token hygiene, session cookies, provisioning scope, metering/money math, secrets, SSRF, input validation + XSS/CSP, dependency risk. Use when asked to "audit" a codebase, service, branch, or diff for security (security gate/profile reviews, audit briefs with attack-surface checklists, "verify claims before reporting" tasks). Read-only discipline, verify against real source, structured verdict report.


Security audit (adversarial, read-only)

When to use

  • "Audit X for security", adversarial/security-gate review of a branch or diff before merge.
  • Briefs with an attack-surface checklist and a mandated report format (Verdict + severity tiers + non-issues checked). Follow the brief exactly — scope, constraints, report shape.

Ground rules

  1. Read the brief FIRST; it defines scope, report format, and constraints.
  2. Read-only: no edits, commits, docker, or live external calls. Running tests (npm test, npm run check) and read-only tooling (npm audit) IS allowed — do run them; a green suite is evidence for the report.
  3. Verify every claim against real source before reporting. No "probably": read the file, check the line, run the command. If tool output looks corrupted, confirm with raw bytes before claiming a bug.
  4. Report absolute paths + line numbers. Separate "found safe" (non-issues checked) from findings so the operator sees coverage.
  5. Issue trackers are hints, not evidence of remediation. When rechecking an open/closed ticket, inspect the live issue body and comments, then verify the current target tree (and the exact release ref) with source, tests, and history. An unchecked issue may be stale; a closed issue may still be incomplete. Report FIXED, PARTIAL, UNFIXED, or SUPERSEDED separately from GitHub state, and never change issue status as part of an audit unless explicitly asked.
  6. Audit the release ref, not just the working tree. If local HEAD is ahead of origin/main, compare both and name the SHA used for the verdict. Separate local-only fixes, unpushed fixes, and fixes present on the release branch. A green focused test on a local tree does not clear a red CI run on the release ref.

Attack-surface checklist (audit ALL of these)

  1. Webhooks: signature fail-closed (verbatim raw body, 401 + nothing processed on failure), idempotency (UNIQUE constraints + guarded status transitions inside a transaction; replay/duplicate/concurrent delivery), event-type handling (completed vs expired vs async_payment_succeeded/failed — and payment_status check), API version pinning, no secret leakage in logs. When the webhook handler is in a third-party plugin/package NOT in the diff (e.g. Medusa's payment-stripe plugin, Stripe SDK), flag it as a medium finding: verify STRIPE_WEBHOOK_SECRET deployment-time and confirm signature validation is active.
  2. Magic-link / token auth: entropy (≥32 random bytes), hash-only storage (SHA-256 of a random token is fine — slow hashes unnecessary), TTL, one-time use (TOCTOU: guarded UPDATE ... WHERE used=0 + changes===1 check), rate limiting (keyed on what? X-Forwarded-For spoofing?), enumeration (generic responses AND response-timing side channels), timing-safe compare.
  3. Session cookies: HMAC signing, HttpOnly/Secure/SameSite, expiry, server-side session validation per request, revocation on logout, fixation (fresh token per login).
  4. Provisioning: key scope (model allowlist, team_id, budget), team isolation, key delivery (one-time? emailed plaintext? plaintext at rest?), admin endpoints (bearer compare timing-safety), idempotency of side effects.
  5. Metering / money math: SQLi (parameterized everywhere), math (ceil direction, overflow → Infinity, negative spend), checkpoint integrity, cross-team attribution, dedupe (INSERT OR IGNORE + deduct-only-on-insert → no double charge under overlap/concurrent runs).
  6. SQLite/DB: WAL, parameterized queries (grep for string-interpolated SQL), file permissions (chmod 0600 for files holding secrets), path traversal on env-configured paths (usually operator-controlled → non-issue, say so).
  7. Secrets: no hardcoded (grep key/secret/private-key patterns), compose uses ${ENV} interpolation only, .env.example placeholders only, logs never carry keys/tokens/bodies (CHECK DEFAULT LOG SERIALIZERS — see pitfall below). Also check for preflight/contract scripts that refuse live keys and redact credentials from output.
  8. SSRF: outbound hosts allowlisted (Stripe SDK, env-configured services), no user-controlled URLs.
  9. Input validation + XSS: validation on all routes, body size limits, HTML escaping in server-rendered output AND client-side innerHTML sinks, CSP/security headers (Referrer-Policy, X-Content-Type-Options, X-Frame-Options). In React/Next.js apps, grep for innerHTML and dangerouslySetInnerHTML — zero occurrences is the expected bar.
  10. Dependency risk: npm audit (distinguish audit-flagged from exploitable-in-this-code-path), unused deps dragging vulnerable transitive deps, lockfile present, non-root runtime user. For ecommerce stacks (Medusa, Payload, Shopify), distinguish: (a) critical direct deps like swiper — these are exploitable in the code path and must be upgraded; (b) deep transitive deps through AWS SDK, ORM, or CMS — flag them but note reachability; (c) E2E test infra — confirm Playwright config disables traces/screenshots/video so checkout secrets aren't recorded.
  11. Ecommerce cart integrity: cart totals are server-side computed (Medusa, Shop API, etc.), not client-trusted. Promotion codes are validated server-side. No client-side discount manipulation possible. Cart ID comes from an httpOnly cookie, never from URL params or localStorage. Verify the getCartId() / setCartId() pattern: httpOnly, sameSite, secure, path, maxAge.
  12. Payment redirect / polling: When the payment flow redirects to a third-party (Stripe 3DS, APM) and back, the return handler must: (a) read cart ID from cookie only, never from URL params; (b) poll with a hard upper bound (max attempts, max total time, Fibonacci-ish delays); (c) never claim "payment failed" on timeout — the message should be "your payment may already be complete, check your email"; (d) use server-side order completion (idempotent, deduped) not client-side state.
  13. Framework CSRF: Next.js Server Actions, Remix actions, and similar frameworks include built-in CSRF protection (POST-only enforcement, crypto tokens). Verify the framework version supports this and that no custom GET endpoints mutate state. If the app uses a custom API route layer, check for explicit CSRF tokens or cookie-based SameSite protection.

Remediation re-audits (finding-only scope)

When re-auditing a named prior finding rather than the whole branch:

  1. Diff exactly the supplied base and remediation commits; list changed paths and confirm the branch/worktree before drawing conclusions.
  2. Trace the original exploit through two or more deliveries, identifying the durable guard (UNIQUE constraint, guarded update, or persisted flag), not just a local boolean.
  3. Verify that first-time-only side effects (emails/token issuance) are separated from retry-required convergence work (provisioning, budgets, state repair). A replay fix must not gate the latter and recreate a post-crash recovery failure.
  4. Inspect and run a focused regression test that delivers the webhook twice and asserts the relevant side-effect count; run typecheck/full tests when practical.
  5. Produce only the requested review artifact; do not modify product source or tests. Use the brief's verdict vocabulary (for example, CLOSED, PARTIAL, or REOPENED) and include file/line and test-count evidence.

Report format

## Verdict: PASS | PASS-WITH-FIXES | FAIL
### Critical / High / Medium / Low
- file:line — issue — suggested fix
### Non-issues checked
- (what you checked and found safe, so operator sees coverage)

Pitfalls (learned the hard way)

  • `raise X from None` suppresses the traceback but does NOT clear `__context__` (verified 2026-08-19). When a handler does except Exception: raise Marker() from None, Python sets __cause__ = None and __suppress_context__ = True — so the raw exception is absent from the formatted traceback — but __context__ still holds the original exception object, including any malicious/raw value. If that marker is then handed to a failure hook, error serializer, APM agent, or logging integration that traverses exc.__context__, the raw value can still leak. Two consequences for audits:
  1. A reviewer claiming "from None clears __context__" is WRONG — verify empirically (e.__suppress_context__ is True but e.__context__ is not None). The formatted traceback is clean; direct __context__ traversal is not.
  2. If the requirement is genuinely "no raw value reachable anywhere on the exception object" (not merely "not in the formatted traceback"), from None is insufficient. Raise the marker OUTSIDE the except block (defer via a flag so no exception is active when the raise executes) — then __context__ is genuinely None. Or, if the marker must be raised inside an except handler, accept the residual __context__ and instead ensure the downstream hook never serializes it.
  • Hermes tool output redacts secret-like text: read_file/grep/terminal output may render identifiers such as key_value, string;, ${POSTGRES_PASSWORD} as ... / ***. This is a display artifact, NOT file corruption. Before flagging a "syntax error" or "hardcoded secret", verify raw bytes with od -c <file> (xxd is often not installed). A redacted-looking line can be perfectly valid code — in one audit apiKey: afterKey.key_value rendered as apiKey: afterK...lue, which looked like a compile error until od showed the real bytes.
  • Fastify logs query strings by default: node_modules/fastify/lib/logger-pino.js serializes url: req.url — the full URL including query string. Magic-link tokens / API keys passed as ?token=... land verbatim in INFO logs (and any log shipping). Check the logger serializers; strip/redact query strings or move secrets to POST bodies. Also set Referrer-Policy: no-referrer when tokens travel in URLs — the token leaks into the Referer header of the redirect target.
  • Stripe `checkout.session.completed` ≠ paid: fires pre-settlement for async payment methods (SEPA, CashApp, Klarna, bank transfer). Require session.payment_status === 'paid' (or handle async_payment_succeeded/async_payment_failed) before granting value; cross-check amount_total vs expected price as defense-in-depth.
  • Server-side escaping doesn't cover client-side sinks: tables rendered via innerHTML from JSON APIs are XSS sinks even when every server-rendered string is escaped. Grep for innerHTML and trace which DB/API fields feed it (model names, names, free-text). Escaping <>& covers text context but quote-handling matters in attribute context.
  • Idempotency ≠ concurrency: UNIQUE + guarded UPDATE protects the credit path, but provisioning side-effects (create team/key, send email) executed OUTSIDE the transaction can orphan resources or double-fire on concurrent deliveries — report even if low severity.
  • Audit-flagged vs exploitable: an npm audit HIGH is a finding, but state whether the vulnerable feature is reachable in this code path (e.g. nodemailer raw-option SSRF advisory is inert if the mailer never passes raw/attachments; still upgrade).
  • Rate limiters keyed only by email/user stop per-target abuse but not flood/DB-growth abuse — check unauthenticated create endpoints (e.g. checkout) for missing per-IP limits and unbounded row creation.
  • Opaque webhook handlers: When the webhook handler lives in a third-party plugin (e.g. Medusa's payment-stripe, Stripe SDK's built-in endpoint) NOT in the diff, you cannot verify signature validation, idempotency, or event-type handling. Report it as a medium finding: "Must verify STRIPE_WEBHOOK_SECRET is set in production and Stripe Dashboard endpoint is configured with the correct signing secret." Do not downgrade to "non-issue" — the handler is unverified.
  • Tier npm audit findings by reachability: (a) direct deps like swiper — exploitable in the code path, must be upgraded before merge; (b) deep transitive deps through AWS SDK, ORM, or CMS — flag them but note reachability context; (c) E2E test infra deps — usually non-issue, but check Playwright config for trace/screenshot/video settings that could leak checkout secrets. Never report a flat list of 100+ advisories without stratification.
  • New side-effects in an idempotent webhook handler fire on every replay unless gated: when a diff ADDS an email/token/provisioning call to an existing handler, trace the replay path, not just the happy path. Handlers often deliberately fall through to convergence code on replay (skip the grant, still provision), so anything there not behind a one-time flag (e.g. welcome_sent) or the first-time-grant branch double-fires. Compare each side-effect's guard against its siblings — if the welcome email is flagged exactly-once but the newly added magic-link email is not, that asymmetry IS the finding. (billing-ux-signup audit: post-signup magic link issued + emailed on every Stripe retry of the setup-mode session — medium.)
  • Bounded download links require two idempotency checks: minting must be unique per purchase + product even when an email/webhook retries; hash-only storage can use a deterministic HMAC bearer token keyed by a server secret and purchase/product/version to regenerate the identical link without persisting plaintext. Require UNIQUE(purchase_id, product_id) plus conflict-safe insert. Separately, retrieve/validate an upstream release asset before the guarded decrement so an outage or missing asset does not reduce entitlement without delivering a file. Test SMTP retry and upstream-fetch failure paths.
  • Bearer tokens in path parameters are logged by Fastify too: query-string redaction is insufficient for /commerce/download/:token. Redact the path segment in request serialization, retain Referrer-Policy: no-referrer, and assert serializer output. Treat deploy-configured release versions as untrusted input: validate a conservative tag grammar before composing attachment filenames/GitHub paths, and bound download limits/expiry to positive sane values at config load.
  • read_file dedup can wall off the unread tail of a truncated file: re-reading a path returns "unchanged — refer to the earlier result" even when that earlier result was truncated at line N and you asked for offset N+1; repeated attempts hard-block. Use sed -n 'A,Bp' <file> via terminal for line-range reads instead. Same redaction caveat as the first pitfall applies to grep: grep -o "sk_test_[A-Za-z0-9]*" prints *** and grep -c counts matches on redacted-looking lines — confirm raw bytes with od -c before claiming anything about file contents.

Support files

  • references/billing-bridge-audit.md — worked example: full finding list + non-issues from the billing-bridge audit (Node/Fastify money-moving service: Stripe webhook, magic-link auth, LiteLLM provisioning, metering, SQLite, portal).
  • references/checkout-epic-audit-worked-example.md — worked example: Medusa/Next.js/Stripe ecommerce checkout audit (Server Action CSRF, payment redirect polling, cookie-based cart ID, E2E test infra hygiene, deep transitive dep triage).
All versions