
Leaked secrets—API keys, access tokens, database credentials, SSH keys, signing keys—turn routine mistakes into high-impact security incidents. A single exposed key can enable data exfiltration, cryptomining, supply-chain tampering, or privilege escalation across cloud and SaaS environments. The difference between a small scare and a breach is usually response speed and coordination.
This guide provides an incident response playbook for leaked keys that you can adapt into a runbook: how to confirm the exposure, contain damage, rotate safely, preserve evidence, and prevent recurrence—without breaking production.
What counts as a “leaked key” incident?
A leaked key is any secret that becomes accessible to an unauthorized party. Common exposure paths include:
- Source control: committed to Git, pushed to public repos, copied into issues or pull requests
- CI/CD logs: secrets printed by debug output or misconfigured masking
- Client-side apps: mobile/web bundles containing server-side keys
- Chat, tickets, docs: pasted into Slack, email, wikis, support portals
- Compromised endpoints: developer laptops, build agents, or shared servers
Not every leak results in abuse, but you should assume exposure implies compromise until you can prove otherwise.
Incident goals and guiding principles
- Minimize blast radius: quickly restrict access tied to the leaked secret.
- Preserve evidence: keep logs and context so you can assess impact and meet compliance obligations.
- Restore safely: rotate credentials in a way that avoids outages and prevents re-leakage.
- Prevent recurrence: harden pipelines and adopt stronger secret lifecycle controls.
Rule of thumb: Revoke first when you can, rotate fast when you must, and investigate in parallel.
Roles and ownership (define before you need it)
Your incident response playbook for leaked keys should clearly define who does what. A simple model:
- Incident Commander (IC): coordinates actions, timeline, approvals, and comms.
- Security Lead: containment strategy, forensics, threat assessment.
- Service Owner: executes rotations, validates application behavior, manages rollback.
- Cloud/SaaS Admin: revokes tokens, adjusts IAM policies, reviews audit logs.
- Comms/Legal/Privacy: customer/internal notifications, regulatory assessment.
Step 1: Triage and confirm exposure
Start with a structured triage checklist to avoid wasting time on false positives while still acting quickly.
Triage checklist
- Identify the secret type: cloud access key, OAuth token, database password, webhook secret, signing key, etc.
- Locate the exposure point: commit hash, CI job URL, log line, shared doc link, ticket ID.
- Determine scope: which environments (dev/stage/prod), which services, which accounts.
- Check time window: when it first appeared and whether it was accessible publicly.
- Assess privileges: what can the key do (read-only vs admin)? What data/systems are reachable?
Decision point: If the key grants production access or wide permissions, treat as high severity and move immediately to containment while triage continues.
Step 2: Immediate containment (stop the bleeding)
Containment is about reducing the chance of ongoing misuse. Options (often combined):
- Revoke/disable the credential (preferred when safe).
- Rotate to a new secret and invalidate the old one after cutover.
- Constrain permissions: temporarily reduce IAM scope, restrict IPs, enforce MFA/session policies.
- Block suspicious traffic: WAF rules, rate limits, egress restrictions, or temporary lockdown.
Containment priorities by secret type
| Secret type | Typical risk | Containment action | Target timeline |
|---|---|---|---|
| Cloud access key (IAM) | Account takeover, data access | Disable key; review policies; rotate | < 15 minutes |
| API key to paid/critical API | Abuse, data scraping, billing fraud | Revoke key; create replacement; add quotas | < 30 minutes |
| Database credentials | Data exfiltration/modification | Rotate; invalidate old; review connections | < 60 minutes |
| OAuth refresh token | Persistent SaaS access | Revoke token/app session; rotate client secret | < 60 minutes |
| Signing key (JWT/code signing) | Forgery, supply-chain risk | Initiate key rollover; update verifiers; revoke | Same day |
Quick containment example (AWS-style access key)
Use your cloud provider tooling to disable the credential immediately, then investigate usage. For example:
# Disable an access key (example pattern; adapt to your environment)
aws iam update-access-key \
--user-name <user> \
--access-key-id <AKIA...> \
--status Inactive
# Then look for suspicious usage in audit logs (e.g., CloudTrail)
# Filter by accessKeyId, eventName, sourceIPAddress, userAgent
Tip: If disabling will cause an outage, use a two-step rotation: introduce a new key, deploy it, verify, then disable the old key.
Step 3: Eradication and safe rotation (without breaking production)
Rotation is where many teams get stuck: you need speed, but you also need reliability. Use a repeatable method:
Rotation runbook
- Create replacement secret with least privilege and an owner tag/label.
- Update secret at the source of truth (secret manager, CI variables, runtime config).
- Deploy to workloads using a controlled rollout (canary, blue/green, staged deploy).
- Validate: health checks, key-specific smoke tests, and error budget monitoring.
- Invalidate old secret: revoke key, drop DB user/password, revoke OAuth token.
- Confirm no dependencies remain: search code/configs/logs for the old value or identifier.
Common rotation pitfalls
- Multiple consumers: the same key used by many services. Fix by issuing per-service credentials.
- Hidden copies: secrets duplicated in old Helm charts, Terraform state, or wiki snippets.
- Long-lived tokens: refresh tokens or service account keys that stay valid for months.
- No rollback plan: rotation fails and teams “temporarily” re-enable the leaked secret.
Step 4: Forensics and impact assessment
While containment/rotation is happening, run investigation in parallel. Your goal is to answer: Was it used, what did it access, and what data/actions occurred?
What to collect (minimum viable evidence)
- Exposure artifact: commit URL/hash, log snippet, ticket link, timestamp
- Identity details: key ID, token identifier, associated user/service account
- Access logs: cloud audit logs, API gateway logs, DB audit, SaaS admin logs
- Network indicators: source IPs, user agents, geo anomalies, request patterns
- Change events: privilege changes, new users/keys, altered policies
Investigation questions
- Did the key perform unusual actions (new IAM policies, new tokens, privilege escalation)?
- Was there data access (object storage reads, database exports, bulk API reads)?
- Was the key used from unexpected networks (new ASN, country, TOR)?
- Were there follow-on implants (backdoor users, webhook changes, CI runner modifications)?
If evidence suggests unauthorized access to regulated data, escalate to your legal/privacy process and preserve logs under retention/hold policies.
Step 5: Communications and reporting
Communication is part of the playbook, not an afterthought. Keep it factual and time-stamped.
Internal communications
- Create a dedicated incident channel and a single source of truth (doc or ticket).
- Post regular updates: what happened, what’s contained, what’s pending, who owns next steps.
- Avoid pasting secrets or sensitive indicators in chat—use secure attachments/links.
External communications
Only notify customers/partners if required or if there is credible risk to their data or integrations. Coordinate wording with legal/privacy and include practical steps (e.g., rotate webhook secrets on their side) if applicable.
Step 6: Recovery and validation
Recovery means returning to normal operations with confidence:
- Verify services: error rates, auth failures, throughput, latency, and key-dependent workflows.
- Confirm revocation: ensure old secrets no longer authenticate anywhere.
- Hunt for recurrence: run secret scanning across repos and artifact stores.
- Close the loop: ensure incident ticket includes timeline, decisions, and evidence links.
Hardening: Prevent the next leaked key
The best incident response playbook for leaked keys includes preventive controls that reduce the likelihood and impact of future leaks.
High-impact controls
- Central secrets management: store secrets in a vault/secret manager, not in code or plaintext configs.
- Short-lived credentials: prefer dynamic secrets, OIDC federation, and expiring tokens over static keys.
- Least privilege by default: per-service identities, scoped permissions, and separation by environment.
- Secret scanning: pre-commit hooks, CI checks, and repository scanners for known formats.
- Masked logging: ensure CI/CD and app logs redact sensitive values.
- Automated rotation: rotate on schedule and on-demand, with safe rollout patterns.
- Audit logging + alerting: alerts on anomalous key usage (new geos, new IPs, high volume, admin actions).
Practical policy snippet (what “good” looks like)
Policy: Secrets must never be stored in source control.
- All production secrets are retrieved at runtime from an approved secrets manager.
- Static access keys are forbidden when federation (OIDC/SAML) is supported.
- Any suspected leak triggers: revoke/rotate <60 minutes, investigation within 24 hours.
- Keys must be per-service, per-environment, and least-privileged.
Post-incident review (PIR): Turn a fire drill into resilience
Within 3–5 business days, run a blameless PIR focused on system improvements:
- Root cause: why was the secret created, stored, and exposed that way?
- Detection gaps: how could you have found it earlier (or prevented the commit/logging)?
- Response friction: what slowed rotation (unknown owners, shared keys, missing runbooks)?
- Action items: assign owners, deadlines, and measurable outcomes.
Printable checklist (copy into your runbook)
- Classify secret + severity; identify exposure source and time window
- Contain: revoke/disable or restrict permissions; block suspicious traffic if needed
- Rotate safely: create new secret; deploy; validate; invalidate old secret
- Investigate: audit logs, data access, privilege changes, follow-on persistence
- Communicate: internal updates; assess external notification obligations
- Recover: confirm stability; confirm old secret is dead everywhere
- Harden: scanning, least privilege, short-lived credentials, automation
- PIR: root cause + tracked remediation plan
Final note
A leaked key incident is stressful, but it’s also one of the most improvable failure modes in modern security. With a rehearsed incident response playbook for leaked keys, you can reduce downtime, limit impact, and build a stronger secrets lifecycle. If you’re standardizing on centralized secrets storage and automated rotation, platforms like Vaulify can help operationalize those practices across teams.