The Complete Guide to Centreon Alert Management for Network Engineers
Centreon can detect every failure in your network. The harder problem is turning that fire hose of alerts into an actionable, low-noise signal your NOC can act on. This guide covers alert configuration, severity levels, escalation policies, and the pitfalls that cause most teams to drown in false positives.
Centreon Alert States: The Foundation
Centreon uses four native states for both hosts and services. Understanding the state machine is essential before configuring alerts and escalations:
| State | Host Meaning | Service Meaning | Notification Default |
|---|---|---|---|
| ● OK | Host is reachable and UP | Service is healthy | Notify on recovery |
| ● WARNING | Host is UP but metrics are abnormal | Service is degraded but functional | Notify immediately |
| ● CRITICAL | Host is DOWN or unreachable | Service is failing | Notify immediately |
| ● UNKNOWN | Check plugin cannot determine state | Check failed to return valid data | Notify (configurable) |
Soft vs. Hard states
This is one of the most misunderstood aspects of Centreon (and Nagios-derived) monitoring. A state change doesn't immediately trigger a notification. Centreon first checks whether the state is soft or hard:
- Soft state: The check returned a non-OK result, but Centreon hasn't confirmed it yet.
It will re-check up to
max_check_attemptstimes before declaring a hard state. - Hard state: The non-OK result has been confirmed over multiple consecutive checks. Notifications are sent at this point.
The max_check_attempts parameter controls how many soft checks occur before a hard notification.
Default is typically 3. Increasing this value reduces false positive notifications from transient
glitches but delays notification on real failures. Set it based on the criticality of each host type:
| Host Type | Recommended max_check_attempts | Rationale |
|---|---|---|
| Core network (switches, routers) | 2–3 | Fast notification critical; core devices rarely have transient failures |
| Servers (production) | 3 | Balance between false positives and notification speed |
| Applications / services | 3–5 | Services can flap; reduce noise from brief restarts |
| Non-critical endpoints | 5 | High tolerance for transients; reduce alert fatigue |
Centreon Severity Levels: Beyond OK/WARNING/CRITICAL
Native Centreon states (OK/WARNING/CRITICAL) are binary. Real NOC environments need finer-grained prioritization. Centreon solves this with custom severity levels — a 1–5 ranking you assign to hosts and services.
Severity levels allow you to answer: "Given 20 simultaneous CRITICAL alerts, which do we fix first?"
Configuring severity levels
Navigate to: Configuration → Hosts → Severity (or Service Severity for services).
A sensible baseline severity model for network infrastructure:
| Level | Name | Example Hosts |
|---|---|---|
| 1 | Mission Critical | Core switches, border routers, primary DNS/DHCP |
| 2 | High | Distribution switches, production servers, firewalls |
| 3 | Medium | Access switches, application servers, databases |
| 4 | Low | Non-critical services, monitoring infrastructure, dev servers |
| 5 | Informational | Printers, IoT devices, conference room equipment |
Once assigned, severity levels appear in the Centreon Resources Status view, allowing engineers to sort and filter by priority during an incident. This is how you prevent engineers from wasting time on a CRITICAL printer alert while a core switch is also down.
Configuring Alert Thresholds
Default plugin thresholds often don't match your environment. Tuning them is one of the highest-ROI configuration tasks available — poorly tuned thresholds are the primary cause of alert fatigue.
Service check threshold format
Centreon (and Nagios plugins) use a standard threshold format:
# Format: [min:]max
# WARNING if value is outside this range
# CRITICAL if value is outside this range
# CPU check example:
-w 80 -c 95 # WARNING at 80%, CRITICAL at 95%
# Disk check example:
-w 75% -c 90% # WARNING at 75% full, CRITICAL at 90% full
# Response time example:
-w 200 -c 500 # WARNING above 200ms, CRITICAL above 500ms
Common threshold tuning mistakes
- CPU thresholds too low — Setting WARNING at 50% on a normally busy application server creates constant false positives. Tune based on observed baseline, not theoretical maximums.
- Disk thresholds as percentages only — 90% full on a 100GB drive (10GB free) is very different from 90% full on a 10TB drive (1TB free). Consider adding absolute free space thresholds for large storage systems.
- No hysteresis on volatile metrics — Metrics that oscillate around a threshold create flapping alerts. Use Centreon's flap detection or add hysteresis to check scripts.
Alert fatigue warning: The fastest path to NOC engineers ignoring alerts is too many false positives. Before adding new alert checks, ask: "What will an engineer actually do when this fires?" If the answer is "check and clear it without taking action," the threshold is wrong — or the check shouldn't send notifications at all.
Escalation Policies: Ensuring Nothing Slips Through
Escalation policies define what happens when a CRITICAL alert isn't acknowledged within a defined timeframe. Without escalation, a P1 incident at 3 AM that the on-call engineer misses gets no further attention until someone notices at 7 AM.
Setting up escalations in Centreon
Navigate to: Configuration → Notifications → Escalations
Each escalation rule defines:
- Hosts / Host groups — which systems this escalation applies to
- First notification — at which notification number the escalation triggers (e.g., after 3 unacknowledged notifications)
- Last notification — when the escalation stops (set to 0 for unlimited)
- Notification interval — how frequently to re-notify during the escalation
- Time period — when this escalation is active (24x7, business hours, etc.)
- Contact groups — who to notify at this escalation level
A practical escalation chain
| Escalation Level | Trigger | Notify | Interval |
|---|---|---|---|
| Level 1 (initial) | Notification 1 | NOC on-call engineer | 10 min |
| Level 2 | Notification 4 (30 min unacked) | NOC team lead | 15 min |
| Level 3 | Notification 7 (75 min unacked) | Network manager | 30 min |
| Level 4 (P1 only) | Notification 10 (120 min unacked) | IT Director | 60 min |
Parent-Child Host Relationships: The Most Under-Used Feature
When a core switch fails, all servers behind it go DOWN. Centreon correctly alerts on all of them. Without parent-child configuration, your NOC sees 40 CRITICAL alerts instead of 1.
Parent-child relationships solve this. When a parent host goes DOWN, Centreon marks child hosts as UNREACHABLE instead of DOWN, and — by default — suppresses notifications for UNREACHABLE hosts.
Configuring host parents
Navigate to: Configuration → Hosts → Hosts, edit a host, and set the Parent hosts field. You can also configure this in bulk via host templates.
For a typical tiered network:
# Topology hierarchy:
Border Router (no parent)
└── Core Switch (parent: border router)
├── Distribution Switch A (parent: core switch)
│ ├── Server 01 (parent: distribution switch A)
│ └── Server 02 (parent: distribution switch A)
└── Distribution Switch B (parent: core switch)
└── Server 03 (parent: distribution switch B)
When the core switch goes DOWN, Distribution Switches A and B go UNREACHABLE (no notification). All servers go UNREACHABLE (no notification). One alert instead of seven. This scales to hundreds of hosts.
Implementation tip: Start with your tier-1 infrastructure (core switches, border routers) and work down. Even configuring parents for the top two tiers typically reduces alert storm severity by 60–80%. Full topology mapping is ideal but not required to see immediate benefits.
Skip the manual triage — try AlertLens free
Even with perfect Centreon configuration, active incidents generate alert storms. AlertLens analyzes your alert screenshot and delivers root cause analysis, severity ranking, and CLI commands in 60 seconds.
Try AlertLens free →Notification Contacts and Contact Groups
Centreon notifications go to contacts, which can be grouped into contact groups. Proper contact organization is essential for clean escalation chains.
Best practices for contact configuration
- Use contact groups, not individual contacts — Assign contact groups to hosts/services. Adding or removing an engineer only requires updating the group, not hundreds of host configurations.
- Set time period restrictions — Each contact has a notification time period. The on-call engineer's time period should match their shift. Out-of-hours notifications route to the on-call group automatically.
- Define notification options explicitly — For each contact, specify which states trigger notifications (d for DOWN, r for RECOVERY, u for UNKNOWN, w for WARNING, c for CRITICAL). Don't notify on UNKNOWN unless your environment generates meaningful UNKNOWN states.
- Use service-specific contact groups — Network events should go to the network team. Application alerts to the app team. Avoid routing everything to a single NOC group for large environments.
Common Centreon Alert Configuration Pitfalls
Pitfall 1: UNKNOWN state notifications enabled by default
UNKNOWN means the check plugin couldn't get a result — often due to connectivity issues, timeouts, or SNMP community mismatches. In most environments, UNKNOWN alerts are infrastructure noise, not real incidents. Disable UNKNOWN notifications for non-critical hosts and review why UNKNOWN events occur before enabling them.
Pitfall 2: Re-notification interval set to 0
A re-notification interval of 0 means Centreon re-notifies every time the check runs and the problem persists.
For a 1-minute check interval, that's 60 notifications per hour for a single unacknowledged problem.
Set notification_interval to something reasonable (30–60 minutes for non-P1 alerts).
Pitfall 3: No maintenance windows for planned work
Scheduled downtime in Centreon (Monitoring → Downtime) suppresses notifications during planned maintenance windows. Teams that don't use downtime generate alert storms during every patching window — training NOC engineers to ignore notifications, which defeats the entire purpose.
Pitfall 4: Alert acknowledgment without notes
When an engineer acknowledges an alert in Centreon without adding a comment, the next person who looks at the acknowledged alert has no context. Enforce a policy: acknowledgment requires a note with the initial diagnosis. This also feeds post-incident reviews and runbook updates.
Pitfall 5: Monitoring the monitoring platform itself
If Centreon goes down, you stop receiving alerts for everything. Configure external health monitoring for your Centreon server — a simple ping check from an external system — and alert via a separate channel (email, SMS) if it becomes unreachable. This is the monitoring blind spot most teams discover only during an actual outage.
Taking Centreon Alert Analysis Further with AI
Even a perfectly configured Centreon environment generates alert storms during complex incidents. Configuration handles the predictable cases. AI handles the rest.
When your Centreon dashboard shows 30 simultaneous alerts during a major incident, AlertLens delivers instant analysis from a screenshot:
- Which alerts are root causes vs. cascade effects
- Recommended priority order for investigation
- CLI diagnostic commands for the most likely root cause
- Executive-ready incident summary for stakeholders
Works with Centreon, Nagios, Zabbix, and 20+ other monitoring platforms — see our platform comparison guide for a full breakdown. Also see our complete MTTR reduction guide for the full NOC optimization playbook.