Why 'it worked yesterday': reading a failed sign-in at 3am
A failed sign-in at 3am is rarely a single bug. It is usually a chain: token drift, conditional access, device posture, DNS, or a quietly expired certificate. This post shows how to read the evidence fast, separate signal from noise, and turn one failed login into a fix you can ship before sunrise.
Nesqual Tech AI
The 3am failure is usually lying to you
At 3:07am, the alert says: failed sign-in. By 3:10am, someone has already blamed the IdP, and by 3:20am, the incident channel is full of screenshots with no root cause. In practice, the login that "worked yesterday" often failed because something else changed: a policy, a token, a device claim, a DNS record, or a certificate that expired at 02:58.
Here is the uncomfortable truth: in many enterprise environments, failed sign-in is not the problem. It is the symptom. Teams that treat it as a one-line auth issue usually waste 30-90 minutes chasing the wrong layer and extend user impact from a 2-minute blip to a 2-hour outage.
Start with the evidence, not the guess
A good 3am read starts with a narrow question: what changed between the last successful sign-in and the first failure? If you cannot answer that in five minutes, you are not debugging auth; you are doing archaeology.
Build a timeline from three sources
Use these sources in this order:
- Identity logs: Entra ID sign-in logs, Okta System Log, Ping logs, or your IdP equivalent.
- Policy logs: Conditional Access, device compliance, MFA, risk, and session control changes.
- Infrastructure logs: DNS, proxy, certificate validation, WAF, VPN, and endpoint telemetry.
A useful pattern is to line up timestamps to the second. In one enterprise case, the failure window was 3:06:11 to 3:08:40. The sign-in log showed MFA required, but the real trigger was a device compliance refresh that failed at 3:05:58 because the MDM service could not reach its CRL endpoint.
03:05:58 MDM compliance refresh failed: CRL timeout
03:06:11 Token refresh requested
03:06:12 Conditional Access evaluated: compliant device = false
03:06:12 MFA required
03:06:13 User denied: app session blocked
That is why failed sign-in reading is a correlation exercise, not a single-log lookup.
Look for the last good token, not just the failed request
If the user says, "It worked yesterday," ask: which token, on which device, for which app? A browser session may have silently refreshed at 11:42pm, while a mobile client failed at 3:07am because its refresh token was invalidated after a password reset.
Common 2026 patterns include:
- refresh token rotation breaking legacy clients
- CAE-driven session revocation after a risk event
- device-bound tokens failing after TPM or certificate changes
- app-specific scopes expiring while the SSO portal still works
Read the sign-in like a packet trace
A failed sign-in has layers. If you skip layers, you misdiagnose. Treat the request path like a packet trace: client, network, identity provider, policy engine, app.
The five questions that cut through noise
Ask these in order:
- Did the request reach the IdP?
- Did authentication succeed?
- Did policy allow the session?
- Did the app accept the assertion or token?
- Did the client fail to store or renew the session?
If the answer to 1 is no, the issue is usually DNS, proxy, firewall, or client reachability. If 2 is yes but 3 is no, you are in policy territory. If 4 fails, inspect claims, audience, signing certs, or clock skew.
A realistic example: an engineering org on Entra ID saw a failed sign-in spike after enabling stricter device filters. The IdP authenticated users successfully, but the app rejected the token because the deviceid claim was missing on noncompliant Linux endpoints. The fix was not "disable MFA"; it was updating the app registration and adding a fallback access path for unmanaged devices.
Use a decision tree, not intuition
Failed sign-in
├── No request at IdP?
│ ├── DNS/proxy/VPN issue
│ └── Client TLS/cert trust issue
├── Auth failed?
│ ├── Password, MFA, risk, lockout
│ └── Federation or upstream IdP issue
├── Policy denied?
│ ├── Conditional Access / device compliance
│ └── Session or risk policy
└── App rejected token?
├── Claims mismatch / clock skew
└── Signing cert rotation / audience mismatch
This tree saves time because it converts a vague failed sign-in into a bounded search space.
The usual culprits in 2026, ranked by pain
By 2026, the most expensive failures are rarely "password wrong." They are policy, token, and trust failures that hide behind a generic failed sign-in label.
1. Conditional Access and device posture drift
A laptop can be "compliant" at 9pm and noncompliant at 3am after an MDM check-in fails. If your CA policy requires a compliant device, the user sees a sign-in failure even though credentials are correct.
Concrete example:
- Intune compliance refresh interval: 8 hours
- CRL fetch timeout: 15 seconds
- Wi-Fi captive portal on hotel network: 302 redirect loop
- Result: token refresh fails, CA blocks access
In one rollout, tightening CA reduced unauthorized access attempts by 41%, but it also increased help desk tickets by 18% until the team added better device-state messaging.
2. Certificate and signing-key rotation
Many teams still underestimate how often signing material changes. If you rotate SAML signing certs, OIDC keys, or mTLS client certs without overlapping validity windows, the app can reject otherwise valid logins.
A practical benchmark: a safe rotation window is usually 7-14 days of overlap for enterprise SSO, with automated metadata refresh every 24 hours. If your app caches metadata for 72 hours, a midnight cert swap can create a wave of failed sign-in events by breakfast.
3. Clock skew and token lifetime mismatch
A 90-second clock skew can break JWT validation in strict systems. In distributed environments, the failure often appears random because only some nodes drift.
Example fix:
- NTP drift alert at >30 seconds
- app token
nbfandexpvalidation with 60-second leeway - synchronized time sources across Kubernetes nodes and VM hosts
4. Federation and upstream outages
If you federate to another IdP, the failure may be upstream. The local sign-in page still looks healthy, but the assertion never returns or returns late enough to time out.
A realistic metric: users start seeing failures when the end-to-end auth flow exceeds 2.5-3.0 seconds on interactive apps. Even if the upstream eventually recovers, the client may have already abandoned the flow.
What to inspect before you touch production
You do not need to guess. You need a repeatable triage checklist that works at 3am and survives handoff at 9am.
A fast triage checklist
- Confirm the exact user, app, device, and timestamp.
- Pull the sign-in event and note the failure stage.
- Check whether the same user can sign in from another device or network.
- Compare the failed session with the last successful one.
- Inspect policy deltas in the last 24 hours.
- Validate certificate, DNS, and NTP health.
- Check for upstream IdP or federation incidents.
# Example: quick triage on a Linux jump host
export USER_EMAIL="user@company.com"
export START="2026-10-01T02:45:00Z"
export END="2026-10-01T03:15:00Z"
curl -s "https://log-api.example.com/signins?user=${USER_EMAIL}&start=${START}&end=${END}" \
-H "Authorization: Bearer $TOKEN" | jq '.events[] | {time, app, status, failureStage, policy, device, ip}'
That kind of query turns failed sign-in from a guess into a table you can reason about.
Make the failure reproducible
If you cannot reproduce it, you do not understand it. Reproduce from:
- the same device posture
- the same network path
- the same browser or client version
- the same policy set
In one case, a user could sign in on corporate Wi-Fi but not on home fiber. The cause was not the IdP. A split-tunnel VPN rule bypassed a private DNS zone, so the app discovery endpoint resolved to a stale IP. The fix was a DNS policy update, not an auth change.
Common Pitfalls
Blaming the IdP too early
The sign-in page is visible, so teams assume the IdP is guilty. Often the IdP authenticated successfully and the app or policy layer failed later. Always check the failure stage before assigning blame.
Ignoring the network path
A proxy auth loop, TLS inspection box, or stale DNS cache can produce the same user-facing symptom as a password failure. If the request never reaches the IdP cleanly, identity logs alone will mislead you.
Overlooking policy drift
A small CA change can affect thousands of users. Teams often forget to diff policy changes against the first failure timestamp. Keep a change log with exact rollout times and blast radius.
Treating all clients the same
Browser, mobile, thick client, and CLI auth flows fail differently. A browser may recover through interactive MFA, while a daemon app fails because its client secret expired. Segment by client type before you start remediation.
Missing time sync problems
If one node is 2 minutes behind, JWT validation, Kerberos, and signed assertions can all fail intermittently. Monitor NTP drift like you monitor CPU.
Turn one failed sign-in into a durable control
The best response to a failed sign-in is not just a fix. It is a control that makes the next one cheaper to diagnose.
Instrument the path end to end
You want observability across identity, policy, and app layers. At minimum, emit:
- authentication result
- policy decision
- token issuance time
- token validation outcome
- client network context
- device compliance state
A healthy setup usually gets you to under 5 minutes mean time to isolate the layer, even if mean time to repair is longer.
observability:
identity:
sign_in_logs: enabled
correlation_id: required
policy:
conditional_access_audit: enabled
change_window_alerts: enabled
app:
jwt_validation_logs: enabled
clock_skew_threshold_seconds: 30
network:
dns_latency_ms_p95: 50
tls_handshake_ms_p95: 200
Add guardrails that prevent recurrence
Practical controls that pay off quickly:
- stagger cert rotations with overlap
- alert on CA policy changes outside maintenance windows
- enforce NTP drift thresholds on all auth-related hosts
- keep a break-glass path for critical admin access
- test sign-in flows from at least two networks and two device states before rollout
A team that added these controls after repeated failed sign-in incidents cut after-hours auth escalations by 62% in one quarter.
Key Takeaways
- Start with the failure stage: authenticate, policy, app, or client.
- Compare the last good sign-in to the first failed one; the delta usually reveals the cause.
- Treat Conditional Access, device posture, certificates, and time sync as first-class auth dependencies.
- Reproduce the issue with the same device, network, and client type before changing production.
- Add correlation IDs and layer-specific logs so the next failed sign-in takes minutes, not hours, to isolate.
- Use change windows, cert overlap, and NTP alerts to prevent the same incident from repeating.
This article was written by an AI system and published pending human review. Verify anything you intend to act on.
Written by
Nesqual Tech AI
Nesqual Tech
Have a project in mind?
Get an instant AI price estimate for it, or talk directly to our team.
One email a month on what we learn building with AI