Rotate an API or DB Secret With No Downtime: Correct Order of Operations
For developers rotating API keys, database passwords, or service credentials in production. This shows the exact sequence that avoids outages: add the new secret first, deploy code that accepts both, switch consumers, verify, then revoke the old secret.
TL;DR — Zero-downtime secret rotation is an order-of-operations problem: create the new secret first, update applications to accept/use both old and new, switch traffic or workers to the new one, verify, then revoke the old secret last. The most common outage is deleting or replacing the old secret before every running process, job worker, and connection pool has reloaded. Reading time: ~5 min
Goal
When you finish, all production workloads will authenticate successfully using the new secret, no user-visible errors will occur during the cutover, and the old secret will be revoked without breaking running apps, workers, cron jobs, or connection pools.
Prerequisites
- Access to the system that issues the secret and supports at least one overlap period or multiple valid credentials at once.
- Access to your application runtime and deployment system: SSH, Kubernetes, Docker, systemd, or your CI/CD pipeline.
- A way to update secrets in the runtime environment: Kubernetes Secret,
.env, systemdEnvironmentFile, cloud secret manager, or CI variables. - A way to restart or reload every consumer of the secret: web app, worker, scheduler, background jobs.
curlandjqinstalled. Check with:
curl --version
jq --version
- If rotating a database password, the database CLI installed, for example:
psql --version
mysql --version
- The list of all consumers of the secret: app pods, worker deployments, cron jobs, CI jobs, integration webhooks, sidecars, and external systems.
- The current secret identifier and the new secret identifier if your provider exposes IDs or key names.
Steps
Step 1: Inventory every consumer before touching the secret
Run commands that show where the secret is referenced. Do not rotate until this list is complete.
# Kubernetes: find env var references and mounted secrets
kubectl get deploy,statefulset,daemonset,cronjob -A -o yaml | grep -nE 'secretKeyRef|secretName|DB_PASSWORD|API_KEY|TOKEN'
# systemd hosts: find environment files and unit references
sudo grep -RniE 'DB_PASSWORD|API_KEY|TOKEN|EnvironmentFile' /etc/systemd /opt /srv /var/www 2>/dev/null
# repo search: find code paths and job definitions
git grep -nE 'DB_PASSWORD|API_KEY|TOKEN|Authorization|password='
Success looks like: you have a concrete list of every deployment, worker, cron job, and external integration that uses the secret.
Step 2: Create the new secret without deleting the old one
Use your issuer's API, CLI, or SQL to add a second valid credential. The exact command varies by system; here are generic patterns.
For a PostgreSQL user password rotation where overlap is possible through a second role:
psql "$ADMIN_DATABASE_URL" -v ON_ERROR_STOP=1 <<'SQL'
CREATE ROLE app_user_next LOGIN PASSWORD 'REPLACE_WITH_NEW_LONG_RANDOM_SECRET';
GRANT app_role TO app_user_next;
SQL
For an HTTP API that issues keys:
curl -sS -X POST https://issuer.example.com/api/keys \
-H "Authorization: Bearer $ADMIN_TOKEN" \
-H "Content-Type: application/json" \
-d '{"name":"app-prod-2026-10-rotation"}' | jq .
Example success output shape:
{
"id": "key_7f3c2d",
"name": "app-prod-2026-10-rotation",
"secret": "sk_live_...",
"created_at": "2026-10-01T12:34:56Z"
}
Success looks like: the issuer shows both old and new credentials as valid at the same time.
Step 3: Update the runtime to carry both old and new values during the transition
Do not overwrite the old value yet if your app can support fallback. Add a second variable or secret entry.
Kubernetes example:
kubectl -n prod create secret generic app-secrets-v2 \
--from-literal=API_KEY_OLD="$API_KEY_OLD" \
--from-literal=API_KEY_NEW="$API_KEY_NEW" \
--dry-run=client -o yaml | kubectl apply -f -
Systemd environment file example:
sudo install -m 600 /dev/null /etc/myapp/myapp.env
sudo sh -c 'cat > /etc/myapp/myapp.env <<"EOF"
API_KEY_OLD="REPLACE_WITH_OLD"
API_KEY_NEW="REPLACE_WITH_NEW"
EOF'
Success looks like: the new secret is present in the runtime secret store without removing the old one.
Step 4: Deploy code or config that prefers the new secret and falls back to the old one
If your app already supports a primary and fallback secret, set literal variable names and deploy. If it does not, add that support before rotation day.
Kubernetes deployment patch example:
kubectl -n prod set env deployment/myapp \
--from=secret/app-secrets-v2
kubectl -n prod rollout status deployment/myapp --timeout=180s
Systemd reload example:
sudo systemctl daemon-reload
sudo systemctl restart myapp.service
sudo systemctl status myapp.service --no-pager -n 20
Success looks like: the new release is running and health checks stay green.
Step 5: Switch all consumers to use the new secret
Now point every app, worker, and job at the new value as primary. Restart anything that caches credentials or keeps long-lived connections.
Kubernetes restart examples:
kubectl -n prod rollout restart deployment/myapp deployment/myapp-worker
kubectl -n prod rollout status deployment/myapp --timeout=180s
kubectl -n prod rollout status deployment/myapp-worker --timeout=180s
For database clients with pooled connections, force pool refresh by restarting the process that owns the pool.
sudo systemctl restart myapp.service myapp-worker.service
Success looks like: new requests, jobs, and fresh DB sessions authenticate with the new secret.
⚠️ Revoking the old secret before this step is fully complete causes intermittent failures: old pods, cron jobs, warm workers, and pooled DB connections may continue using cached credentials for minutes or hours.
Step 6: Verify the new secret is actually in use
Check logs, issuer audit events, or the database session list. Do not revoke the old secret based only on a successful deploy.
HTTP issuer audit example:
curl -sS https://issuer.example.com/api/audit?service=myapp \
-H "Authorization: Bearer $ADMIN_TOKEN" | jq '.events[0:10] | map({credential_id, status, ts})'
PostgreSQL session check example:
psql "$ADMIN_DATABASE_URL" -x -c "SELECT usename, application_name, client_addr, state FROM pg_stat_activity WHERE usename IN ('app_user','app_user_next') ORDER BY usename;"
Success looks like: only the new credential ID or new DB role appears for fresh activity.
Step 7: Revoke the old secret last
After you have observed only new-secret usage for at least one full worker/job cycle, revoke the old credential.
PostgreSQL example:
psql "$ADMIN_DATABASE_URL" -v ON_ERROR_STOP=1 <<'SQL'
REVOKE app_role FROM app_user;
ALTER ROLE app_user NOLOGIN;
SQL
HTTP API example:
curl -sS -X DELETE https://issuer.example.com/api/keys/$OLD_KEY_ID \
-H "Authorization: Bearer $ADMIN_TOKEN" -i
Example success output shape:
HTTP/1.1 204 No Content
Date: Thu, 01 Oct 2026 12:48:10 GMT
Success looks like: the old credential is invalid, and production traffic still succeeds.
Step 8: Remove fallback references and clean up
Delete the old variable names, old secret objects, and any temporary dual-secret logic after the revocation has been stable.
Kubernetes cleanup example:
kubectl -n prod create secret generic app-secrets \
--from-literal=API_KEY="$API_KEY_NEW" \
--dry-run=client -o yaml | kubectl apply -f -
kubectl -n prod set env deployment/myapp deployment/myapp-worker --keys API_KEY --from=secret/app-secrets
kubectl -n prod rollout restart deployment/myapp deployment/myapp-worker
kubectl -n prod delete secret app-secrets-v2
Success looks like: only the new secret remains in config, and no code path references the old one.
Verify it works
Run an end-to-end check from outside and from the runtime.
# App health
curl -sS -o /dev/null -w '%{http_code}\n' https://app.example.com/health
# Kubernetes: confirm all pods are on the new ReplicaSet and ready
kubectl -n prod get pods -o wide
# Logs: no auth failures after revocation
kubectl -n prod logs deploy/myapp --since=10m | grep -E '401|403|auth|password|token' || true
kubectl -n prod logs deploy/myapp-worker --since=10m | grep -E '401|403|auth|password|token' || true
Expected results:
200
And no repeating errors like these:
FATAL: password authentication failed for user "app_user"
HTTP 401 Unauthorized
invalid_token
Access denied for user 'app_user'@'10.0.12.34'
If you rotated a DB credential, also verify fresh connections use the new identity:
psql "$ADMIN_DATABASE_URL" -x -c "SELECT usename, count(*) FROM pg_stat_activity GROUP BY usename ORDER BY usename;"
Expected: the old DB user count drops to zero after restarts and normal connection churn.
Common pitfalls
Revoking the old secret first
Mistake: deleting the old key or changing the only DB password before all consumers reload. Symptom: intermittent 401 Unauthorized, invalid_token, or DB auth failures only on some pods or workers. Fix: recreate or re-enable the old credential immediately, then repeat the rotation with overlap.
Restarting the web app but not workers or cron jobs
Mistake: only the HTTP deployment is rolled. Symptom: the site works, but background jobs start failing minutes later with auth errors. Fix: restart every consumer explicitly, including worker deployments, CronJobs, queue runners, and scheduled tasks.
Assuming env var changes update running processes
Mistake: updating a secret object without restarting processes that read env vars only at startup. Symptom: kubectl get secret shows the new value, but old pods still authenticate with the old secret. Fix: run kubectl rollout restart ... or restart the service/process that owns the env vars.
Forgetting long-lived connection pools
Mistake: rotating a DB password and expecting existing pooled connections to switch automatically. Symptom: old sessions keep working until the pool reconnects, then new sessions fail after revocation. Fix: restart the app and worker processes to force new DB connections before revoking the old credential.
Missing one external integration
Mistake: rotating the secret used by a webhook sender, CI job, or third-party integration that was not in the inventory. Symptom: production looks healthy, but one integration starts returning 403 Forbidden or webhook retries spike. Fix: search audit logs for old credential usage and update that specific integration before revoking again.
This article was written by an AI system and published pending human review. Verify anything you intend to act on.
Have a project in mind?
Get an instant AI price estimate for it, or talk directly to our team.
One email a month on what we learn building with AI