How a Monitoring Checker Decides Something Is Actually a Problem
This guide is for non-engineers who need to understand why a monitoring system raises an alert, stays quiet, or marks something as unhealthy. You will learn the practical decision logic behind checks, thresholds, retries, and time windows so you can choose sensible settings and avoid noisy or missed alerts.
TL;DR — A checker usually does not alert on a single bad event. It decides something is a real problem by comparing a measured result to a rule, often over a time window, with retries or a required number of failures to avoid false alarms. The single most useful fix is usually to tune the rule from “1 failure = alert” to “N failures in M minutes” and define what “bad” means in business terms, not just technical terms. Reading time: ~7 min
What it is and where it sits
A "checker" is the part of a monitoring or health system that asks: "Is this result normal, or is it a problem?" It sits between raw signals and human action.
Raw signals can be things like:
- an HTTP response from your website
- a database connection test
- CPU or memory metrics
- an error rate from application logs
- a failed background job
The checker does not usually create the signal itself. Something else collects or emits the signal:
- a probe (a small process that tests a URL or port)
- an agent (software installed on a server to report metrics)
- the application itself sending logs or metrics
- a cloud service reporting status
Then the checker applies rules. If the rule says the condition is bad enough, long enough, or frequent enough, the checker marks it as unhealthy and may trigger an alert.
In a typical setup, it replaces ad-hoc human judgment like "I saw one timeout in the logs, maybe the site is down" with repeatable logic like "alert only if 3 of the last 5 checks failed."
Here is the usual flow:
User request / system activity
|
v
Signal source (probe, agent, app metric, log)
|
v
Collector / monitoring platform
|
v
Checker evaluates rule
- threshold
- time window
- retry count
- missing-data policy
|
v
State change
OK -> Warning -> Critical
|
v
Alert / dashboard / ticket / page
In a website request flow, the checker is usually off to the side, not in the path of the customer request itself. Your customer loads https://example.com; separately, every 1 minute, a probe requests the same URL and the checker decides whether the result means "healthy" or "problem."
How it actually works
At a practical level, a checker usually combines five decisions:
- What to test — for example
GET /healthon your website. - What counts as bad — for example HTTP status 500, or response time above 2 seconds.
- How many bad results are needed — for example 3 failures.
- Over what period — for example within 5 minutes.
- What to do with missing data — for example treat no response as failure, or ignore it.
One realistic end-to-end example
Let’s use a common case: an uptime checker for an online store.
You want to know when the storefront is really unavailable, but you do not want a pager alert because of one brief network hiccup.
Your team sets this rule:
- Check
https://shop.example.com/health - Every 1 minute
- A check is bad if:
- HTTP status is not
200, or - total response time is above
3000 ms
- HTTP status is not
- Mark as a problem only if 3 of the last 5 checks are bad
- Resolve the incident after 2 good checks in a row
Now walk through what happens.
Step 1: The probe runs the test
At 10:00, the monitoring probe requests the URL.
It records:
- DNS lookup succeeded
- TLS handshake succeeded (TLS = the encryption used by HTTPS)
- HTTP status
200 - response time
420 ms
The checker compares that result to the rule. Status is good, response time is below 3000 ms, so this check is good.
Step 2: A temporary issue appears
At 10:01, the app server is restarting after a deployment.
The probe gets HTTP 502.
The checker marks this single check as bad, but it does not alert yet. Why? Because the rule is not "1 bad check = problem." The checker stores the result in the recent history window.
Recent 5-check window now looks like:
- 09:57 good
- 09:58 good
- 09:59 good
- 10:00 good
- 10:01 bad
That is 1 bad out of 5. Below threshold. State stays OK.
Step 3: The issue continues
At 10:02, another 502.
At 10:03, another 502.
Now the 5-check window contains 3 bad checks. The checker evaluates the rule again:
- bad count = 3
- threshold = 3
- result = threshold met
The checker changes state from OK to Critical and triggers the alert action: email, Slack, PagerDuty, ticket, or whatever your workflow uses.
This is the key idea: the checker is usually stateful (it remembers recent results), not just reactive to one event.
Step 4: Why this avoids noise
If the 10:01 failure had been a one-off network blip, the next checks would have been good and the checker would never have crossed the threshold. That is how retries and windows reduce false positives (alerts for things that are not real incidents).
Step 5: Recovery logic matters too
At 10:04 and 10:05, the service returns 200 again.
If your recovery rule is "1 good check resolves the incident," the alert may flap (flip between bad and good repeatedly) during an unstable recovery. Instead, the checker waits for 2 good checks in a row before changing state back to OK.
That small choice often matters as much as the alert threshold.
What else a checker may consider
Depending on the system, the checker may also apply:
- Severity levels — warning at 2 slow checks, critical at 3 failed checks
- Percent-based rules — alert if error rate is above 5% over 10 minutes
- Baseline rules — alert if today is far outside normal behavior, even if no fixed threshold was crossed
- Dependency suppression — if the whole region is down, suppress hundreds of child alerts
- Maintenance windows — ignore expected failures during planned work
When to use it (and when not to)
Use a checker when you need a repeatable, explainable rule for deciding whether a signal should become an incident.
You probably do not need a sophisticated checker for every metric. Many teams over-monitor and then stop trusting alerts.
| Scenario | Recommendation |
|---|---|
| Public website or API uptime matters to customers | Use a checker with retries and a short time window |
| Background jobs must finish on time | Use a checker on job age, queue depth, or failure count |
| You only want a dashboard, not alerts | Collect metrics, but keep checker rules minimal |
| A metric naturally spikes during normal business hours | Use percentiles or time-window rules, not a single hard threshold |
| A one-off failure is harmless | Do not alert on single events |
| You have no clear action when alerted | Do not create the check yet; define the action first |
| You are still learning what “normal” looks like | Start by graphing data for 1-2 weeks before setting strict thresholds |
You probably don't need this if...
- you only check the system manually once in a while
- there is no agreed response when an alert fires
- the service is internal, low-risk, and temporary
- a simple built-in health check from your hosting platform already answers the question you care about
A good test: if the checker says "problem," can someone do something useful within minutes? If not, the rule is probably premature.
Trade-offs
Every benefit comes with a cost.
| Benefit | What it costs |
|---|---|
| Fewer false alarms with retries and windows | Slower detection; you may wait 3-5 minutes before alerting |
| Clear thresholds | Someone must choose and maintain them as the system changes |
| Richer checks like response time, status code, and content matching | More complexity and more ways to misconfigure the rule |
| More alerts caught | More operational burden; someone must own the alert response |
| External probes catch real customer-visible outages | Usually more money than basic internal metrics alone |
| Vendor-managed monitoring is quick to adopt | Some lock-in around alert rules, dashboards, and incident workflows |
| Aggressive sensitivity catches subtle issues | More noise and alert fatigue |
| Conservative sensitivity reduces noise | Greater chance of missing short but important incidents |
The practical balancing act is this: every checker sits somewhere between "too twitchy" and "too sleepy."
In practice
Below are two examples you can adapt today: one application health endpoint, and one checker rule in a common monitoring style.
Example 1: A simple health endpoint
If your agency is building your app, ask for a dedicated health URL like /health that returns a plain success when the app can serve requests.
{
"status": "ok",
"checks": {
"app": "ok",
"database": "ok"
}
}
This is the kind of JSON a checker can request every minute. The gotcha: do not put expensive work in this endpoint. If /health runs a long report or checks every downstream service, the check itself can create load and false alarms.
If you are testing it manually in a browser, visit https://your-domain.example/health. If you have shell access, you can test it with:
curl -i https://your-domain.example/health
This shows the HTTP status and body. The gotcha: a health endpoint that always returns 200 even when the database is broken is not useful; agree in advance what dependencies it should include.
Example 2: An alert rule with retries and a window
In many monitoring systems, the rule concept looks roughly like this:
name: storefront-uptime
type: http_check
target: https://shop.example.com/health
interval: 60s
timeout: 5s
conditions:
status_code: 200
max_response_time_ms: 3000
alert_when:
failures_in_last: 5
failures_needed: 3
resolve_when:
consecutive_successes: 2
This says: check once a minute, fail on non-200 or slower than 3 seconds, alert after 3 bad checks in the last 5, and resolve after 2 good checks. The gotcha: if your interval is 60 seconds, "3 bad checks" means you may not alert for about 3 minutes. That delay is intentional, but it must match the business impact.
Where to set this in a dashboard
Exact screens vary by provider, but the path is usually similar to:
- Monitoring or Observability → Uptime / Synthetic Checks → New Check
- Enter the URL
- Set interval to
1 minute - Under failure conditions, set expected status
200 - Add a response-time threshold like
3000 ms - Under alerting, choose something like
3 failures within 5 minutes - Under recovery, choose
2 successful checks
If your provider also offers maintenance windows, use:
- Monitoring → Alerting → Maintenance Windows → New Window
That prevents planned deployments from generating noise.
⚠️ If you change an existing production alert from a strict rule to a looser one, you can create a blind spot and miss a real outage. Before saving, write down the old rule, the new rule, and who approved the change.
Example 3: Nginx health endpoint
If your app sits behind Nginx, a minimal endpoint can be served directly by Nginx.
location = /health {
access_log off;
add_header Content-Type text/plain;
return 200 'ok';
}
This returns a fast 200 ok without touching the app. The gotcha: this only proves Nginx is alive, not that your application or database works. Use it for load balancers or basic liveness, not full application readiness.
Further reading
- Google SRE Book — the chapter "Service Level Objectives"
- Prometheus documentation — "Alerting rules"
- MDN Web Docs — the "HTTP response status codes" reference
- RFC 9110 — HTTP Semantics
- Martin Fowler — "Monitoring"
This article was written by an AI system and published pending human review. Verify anything you intend to act on.
Have a project in mind?
Get an instant AI price estimate for it, or talk directly to our team.
One email a month on what we learn building with AI