[v10.0.x] Docs: recreates setup for oss alerting (#70231)
Docs: recreates setup for oss alerting (#70156)
* recreates setup for oss alerting
* fixes relrefs
* fix relrefs
* fix relref
* adds other setup topics
* adds other setup topics 2
* adds other setup topics 3
* adds cloud setup
* adds cloud setup from legacy
* fixes links
* removing link
* gilles feedback
(cherry picked from commit 5fc6bf5c69)
Co-authored-by: brendamuir <100768211+brendamuir@users.noreply.github.com>
This commit is contained in:
co-authored by
brendamuir
parent
2209deb0bf
commit
15ae80757b
@@ -1,15 +1,61 @@
|
||||
---
|
||||
aliases:
|
||||
- unified-alerting/set-up/
|
||||
description: How to set up additional alerting features and integrations
|
||||
title: Additional setup
|
||||
weight: 600
|
||||
labels:
|
||||
products:
|
||||
- oss
|
||||
description: How to set up alerting features and integrations
|
||||
title: Set up Alerting
|
||||
weight: 107
|
||||
---
|
||||
|
||||
# Additional setup
|
||||
# Set up Alerting
|
||||
|
||||
Alerting supports a plethora of configurations, from [configuring external Alertmanagers]({{< relref "./configure-alertmanager" >}}) to routing Grafana Managed Alerts outside of Grafana, to defining your alerting setup [as-code using Provisioning]({{< relref "./provision-alerting-resources" >}}).
|
||||
Configure the features and integrations that you need to create and manage your alerts.
|
||||
|
||||
Setup Grafana Alerting high-availability mode by [following this guide]({{< relref "./configure-high-availability" >}}).
|
||||
**Note:**
|
||||
|
||||
Connect your alerting setup to Grafana OnCall by [following this guide](/docs/oncall/latest/integrations/available-integrations/add-grafana-alerting/).
|
||||
These are set up instructions for Grafana Alerting Open Source.
|
||||
|
||||
To set up Grafana Alerting for Cloud, see ({{< relref "../set-up-cloud" >}})
|
||||
|
||||
## Before you begin
|
||||
|
||||
- Configure your [data sources]({{< relref "../../administration/data-source-management" >}})
|
||||
- Check which data sources are compatible with and supported by [Grafana Alerting]({{< relref "../fundamentals/data-source-alerting" >}})
|
||||
|
||||
## Set up Alerting
|
||||
|
||||
To set up Alerting, you need to:
|
||||
|
||||
1. Configure alert rules
|
||||
|
||||
- Create Grafana-managed or Mimir/Loki-managed alert rules and recording rules
|
||||
|
||||
1. Configure contact points
|
||||
|
||||
- Check the default contact point and update the email address
|
||||
|
||||
- [Optional] Add new contact points and integrations
|
||||
|
||||
1. Configure notification policies
|
||||
|
||||
- Check the default notification policy
|
||||
|
||||
- [Optional] Add additional nested policies
|
||||
|
||||
- [Optional] Add labels and label matchers to control alert routing
|
||||
|
||||
1. [Optional] Integrate with [Grafana OnCall]
|
||||
(/docs/oncall/latest/integrations/grafana-alerting/)
|
||||
|
||||
## Advanced set-up options
|
||||
|
||||
Grafana Alerting supports many additional configuration options, from configuring external Alertmanagers to routing Grafana-managed alerts outside of Grafana, to defining your alerting setup as code.
|
||||
|
||||
The following topics provide you with advanced configuration options for Grafana Alerting.
|
||||
|
||||
- [Provision alert rules using file provisioning]({{< relref "../set-up/provision-alerting-resources/file-provisioning" >}})
|
||||
- [Provision alert rules using Terraform]({{< relref "../set-up/provision-alerting-resources/terraform-provisioning" >}})
|
||||
- [Add an external Alertmanager]({{< relref "../set-up/configure-alertmanager" >}})
|
||||
- [Configure high availability]({{< relref "../set-up/configure-high-availability" >}})
|
||||
|
||||
@@ -9,7 +9,7 @@ keywords:
|
||||
- configure
|
||||
- external Alertmanager
|
||||
title: Add an external Alertmanager
|
||||
weight: 100
|
||||
weight: 200
|
||||
---
|
||||
|
||||
# Add an external Alertmanager
|
||||
|
||||
@@ -10,7 +10,7 @@ keywords:
|
||||
- ha
|
||||
- high availability
|
||||
title: Enable alerting high availability
|
||||
weight: 300
|
||||
weight: 400
|
||||
---
|
||||
|
||||
# Enable alerting high availability
|
||||
|
||||
@@ -0,0 +1,147 @@
|
||||
---
|
||||
aliases:
|
||||
- meta-monitoring/
|
||||
description: Meta monitoring
|
||||
keywords:
|
||||
- grafana
|
||||
- alerting
|
||||
- meta-monitoring
|
||||
title: Meta monitoring
|
||||
weight: 500
|
||||
---
|
||||
|
||||
# Meta monitoring
|
||||
|
||||
Meta monitoring is the process of monitoring your monitoring, and alerting when your monitoring is not working as it should. Whether you use Grafana Managed Alerts or Mimir, meta monitoring is possible both on-premise and in Grafana Cloud.
|
||||
|
||||
## Grafana Managed Alerts
|
||||
|
||||
Meta monitoring of Grafana Managed Alerts requires having a Prometheus server, or other metrics database, collecting and storing metrics exported by Grafana. For example, if using Prometheus you should add a `scrape_config` to Prometheus to scrape metrics from your Grafana server.
|
||||
|
||||
Here is an example of how this might look:
|
||||
|
||||
```
|
||||
- job_name: grafana
|
||||
honor_timestamps: true
|
||||
scrape_interval: 15s
|
||||
scrape_timeout: 10s
|
||||
metrics_path: /metrics
|
||||
scheme: http
|
||||
follow_redirects: true
|
||||
static_configs:
|
||||
- targets:
|
||||
- grafana:3000
|
||||
```
|
||||
|
||||
The Grafana ruler, which is responsible for evaluating alert rules, and the Grafana Alertmanager, which is responsible for sending notifications of firing and resolved alerts, provide a number of metrics that let you observe them.
|
||||
|
||||
#### grafana_alerting_alerts
|
||||
|
||||
This metric is a counter that shows you the number of `normal`, `pending`, `alerting`, `nodata` and `error` alerts. For example, you might want to create an alert that fires when `grafana_alerting_alerts{state="error"}` is greater than 0.
|
||||
|
||||
#### grafana_alerting_schedule_alert_rules
|
||||
|
||||
This metric is a gauge that shows you the number of alert rules scheduled. An alert rule is scheduled unless it is paused, and the value of this metric should match the total number of non-paused alert rules in Grafana.
|
||||
|
||||
#### grafana_alerting_schedule_periodic_duration_seconds_bucket
|
||||
|
||||
This metric is a histogram that shows you the time it takes to process an individual tick in the scheduler that evaluates alert rules. If the scheduler takes longer than 10 seconds to process a tick then pending evaluations will start to accumulate such that alert rules might later than expected.
|
||||
|
||||
#### grafana_alerting_schedule_query_alert_rules_duration_seconds_bucket
|
||||
|
||||
This metric is a histogram that shows you how long it takes the scheduler to fetch the latest rules from the database. If this metric is elevated then so will `schedule_periodic_duration_seconds`.
|
||||
|
||||
#### grafana_alerting_scheduler_behind_seconds
|
||||
|
||||
This metric is a gauge that shows you the number of seconds that the scheduler is behind where it should be. This number will increase if `schedule_periodic_duration_seconds` is longer than 10 seconds, and decrease when it is less than 10 seconds. The smallest possible value of this metric is 0.
|
||||
|
||||
#### grafana_alerting_notification_latency_seconds_bucket
|
||||
|
||||
This metric is a histogram that shows you the number of seconds taken to send notifications for firing and resolved alerts. This metric will let you observe slow or over-utilized integrations, such as an SMTP server that is being given emails faster than it can send them.
|
||||
|
||||
> These metrics are not available at present in Grafana Cloud.
|
||||
|
||||
## Grafana Mimir
|
||||
|
||||
Meta monitoring in Grafana Mimir requires having a Prometheus/Mimir server, or other metrics database, collecting and storing metrics exported by the Mimir ruler.
|
||||
|
||||
#### cortex_prometheus_rule_evaluation_failures_total
|
||||
|
||||
This metric is a counter that shows you the total number of rule evaluation failures.
|
||||
|
||||
## Alertmanager
|
||||
|
||||
Meta monitoring in Alertmanager also requires having a Prometheus/Mimir server, or other metrics database, collecting and storing metrics exported by Alertmanager. For example, if using Prometheus you should add a `scrape_config` to Prometheus to scrape metrics from your Alertmanager.
|
||||
|
||||
Here is an example of how this might look:
|
||||
|
||||
```
|
||||
- job_name: alertmanager
|
||||
honor_timestamps: true
|
||||
scrape_interval: 15s
|
||||
scrape_timeout: 10s
|
||||
metrics_path: /metrics
|
||||
scheme: http
|
||||
follow_redirects: true
|
||||
static_configs:
|
||||
- targets:
|
||||
- alertmanager:9093
|
||||
```
|
||||
|
||||
#### alertmanager_alerts
|
||||
|
||||
This metric is a counter that shows you the number of active, suppressed and unprocessed alerts in Alertmanager. Suppressed alerts are silenced alerts, and unprocessed alerts are alerts that have been sent to the Alertmanager but have not been processed.
|
||||
|
||||
#### alertmanager_alerts_invalid_total
|
||||
|
||||
This metric is a counter that shows you the number of invalid alerts that were sent to Alertmanager. This counter should not exceed 0, and so in most cases you will want to create an alert that fires if whenever this metric increases.
|
||||
|
||||
#### alertmanager_notifications_total
|
||||
|
||||
This metric is a counter that shows you how many notifications have been sent by Alertmanager. The metric uses a label "integration" to show the number of notifications sent by integration, such as email.
|
||||
|
||||
#### alertmanager_notifications_failed_total
|
||||
|
||||
This metric is a counter that shows you how many notifications have failed in total. This metric also uses a label "integration" to show the number of failed notifications by integration, such as failed emails. In most cases you will want to use the `rate` function to understand how often notifications are failing to be sent.
|
||||
|
||||
#### alertmanager_notification_latency_seconds_bucket
|
||||
|
||||
This metric is a histogram that shows you the amount of time it takes Alertmanager to send notifications and for those notifications to be accepted by the receiving service. This metric uses a label "integration" to show the amount of time by integration. For example, you can use this metric to show the 95th percentile latency of sending emails.
|
||||
|
||||
> In Grafana Cloud some of these metrics are available via the Prometheus usage datasource that is provisioned for all Grafana Cloud customers.
|
||||
|
||||
## Alertmanager in high availability mode
|
||||
|
||||
If using Alertmanager in high availability mode there are a number of additional metrics that you might want to create alerts for.
|
||||
|
||||
#### alertmanager_cluster_members
|
||||
|
||||
This metric is a gauge that shows you the current number of members in the cluster. The value of this gauge should be the same across all Alertmanagers. If different Alertmanagers are showing different numbers of members then this is indicative of an issue with your Alertmanager cluster. You should look at the metrics and logs from your Alertmanagers to better understand what might be going wrong.
|
||||
|
||||
#### alertmanager_cluster_failed_peers
|
||||
|
||||
This metric is a gauge that shows you the current number of failed peers.
|
||||
|
||||
#### alertmanager_cluster_health_score
|
||||
|
||||
This metric is a gauge showing the health score of the Alertmanager. Lower values are better, and zero means the Alertmanager is healthy.
|
||||
|
||||
#### alertmanager_cluster_peer_info
|
||||
|
||||
This metric is a gauge. It has a constant value `1`, and contains a label called "peer" containing the Peer ID of each known peer.
|
||||
|
||||
#### alertmanager_cluster_reconnections_failed_total
|
||||
|
||||
This metric is a counter that shows you the number of failed peer connection attempts. In most cases you will want to use the `rate` function to understand how often reconnections fail as this may be indicative of an issue or instability in your network.
|
||||
|
||||
> These metrics are not available in Grafana Cloud as it uses a different high availability strategy than on-premise Alertmanagers.
|
||||
|
||||
<!---
|
||||
#### cortex_prometheus_rule_group_last_evaluation_timestamp_seconds
|
||||
|
||||
#### cortex_prometheus_rule_group_rules
|
||||
|
||||
This metric is a counter that shows
|
||||
|
||||
> In Grafana Cloud these metrics are available via the Prometheus usage datasource that is provisioned for all Grafana Cloud customers.
|
||||
-->
|
||||
@@ -0,0 +1,50 @@
|
||||
---
|
||||
aliases:
|
||||
- alerting-limitations/
|
||||
description: Performance considerations and limitations
|
||||
keywords:
|
||||
- grafana
|
||||
- alerting
|
||||
- performance
|
||||
- limitations
|
||||
title: Performance considerations and limitations
|
||||
weight: 600
|
||||
---
|
||||
|
||||
# Performance considerations and limitations
|
||||
|
||||
Grafana Alerting supports multi-dimensional alerting, where one alert rule can generate many alerts. For example, you can configure an alert rule to fire an alert every time the CPU of individual VMs max out. This topic discusses performance considerations resulting from multi-dimensional alerting.
|
||||
|
||||
Evaluating alerting rules consumes RAM and CPU to compute the output of an alerting query, and network resources to send alert notifications and write the results to the Grafana SQL database. The configuration of individual alert rules affects the resource consumption and, therefore, the maximum number of rules a given configuration can support.
|
||||
|
||||
The following section provides a list of alerting performance considerations.
|
||||
|
||||
- Frequency of rule evaluation consideration. The "Evaluate Every" property of an alert rule controls the frequency of rule evaluation. We recommend using the lowest acceptable evaluation frequency to support more concurrent rules.
|
||||
- Cardinality of the rule's result set. For example, suppose you are monitoring API response errors for every API path, on every VM in your fleet. This set has a cardinality of _n_ number of paths multiplied by _v_ number of VMs. You can reduce the cardinality of a result set - perhaps by monitoring errors-per-VM instead of for each path per VM.
|
||||
- Complexity of the alerting query consideration. Queries that data sources can process and respond to quickly consume fewer resources. Although this consideration is less important than the other considerations listed above, if you have reduced those as much as possible, looking at individual query performance could make a difference.
|
||||
|
||||
Each evaluation of an alert rule generates a set of alert instances; one for each member of the result set. The state of all the instances is written to the `alert_instance` table in Grafana's SQL database. This number of write-heavy operations can cause issues when using SQLite.
|
||||
|
||||
Grafana Alerting exposes a metric, `grafana_alerting_rule_evaluations_total` that counts the number of alert rule evaluations. To get a feel for the influence of rule evaluations on your Grafana instance, you can observe the rate of evaluations and compare it with resource consumption. In a Prometheus-compatible database, you can use the query `rate(grafana_alerting_rule_evaluations_total[5m])` to compute the rate over 5 minute windows of time. It's important to remember that this isn't the full picture of rule evaluation. For example, the load will be unevenly distributed if you have some rules that evaluate every 10 seconds, and others every 30 minutes.
|
||||
|
||||
These factors all affect the load on the Grafana instance, but you should also be aware of the performance impact that evaluating these rules has on your data sources. Alerting queries are often the vast majority of queries handled by monitoring databases, so the same load factors that affect the Grafana instance affect them as well.
|
||||
|
||||
## Limited rule sources support
|
||||
|
||||
Grafana Alerting can retrieve alerting and recording rules **stored** in most available Prometheus, Loki, Mimir, and Alertmanager compatible data sources.
|
||||
|
||||
It does not support reading or writing alerting rules from any other data sources but the ones previously mentioned at this time.
|
||||
|
||||
## Prometheus version support
|
||||
|
||||
We support the latest two minor versions of both Prometheus and Alertmanager. We cannot guarantee that older versions will work.
|
||||
|
||||
As an example, if the current Prometheus version is `2.31.1`, we support >= `2.29.0`.
|
||||
|
||||
## The Grafana Alertmanager can only receive Grafana managed alerts
|
||||
|
||||
Grafana cannot be used to receive external alerts. You can only send alerts to the Grafana Alertmanager using Grafana managed alerts.
|
||||
|
||||
You have the option to send Grafana managed alerts to an external Alertmanager, you can find this option in the admin tab on the Alerting page.
|
||||
|
||||
For more information, refer to [this GitHub discussion](https://github.com/grafana/grafana/discussions/45773).
|
||||
@@ -9,7 +9,7 @@ keywords:
|
||||
- configure
|
||||
- provisioning
|
||||
title: Provision Grafana Alerting resources
|
||||
weight: 200
|
||||
weight: 300
|
||||
---
|
||||
|
||||
# Provision Grafana Alerting resources
|
||||
|
||||
@@ -0,0 +1,243 @@
|
||||
---
|
||||
aliases:
|
||||
- /docs/grafana-cloud/alerts/alerts-rules/
|
||||
- /docs/grafana-cloud/how-do-i/alerts/alerts-rules/
|
||||
- /docs/grafana-cloud/legacy-alerting/alerts-rules/
|
||||
- /docs/grafana-cloud/metrics/prometheus/alerts_rules/
|
||||
- /docs/hosted-metrics/prometheus/alerts_rules/
|
||||
- /docs/grafana-cloud/alerts/grafana-cloud-alerting/
|
||||
- /docs/grafana-cloud/how-do-i/grafana-cloud-alerting/
|
||||
- /docs/grafana-cloud/legacy-alerting/grafana-cloud-alerting/
|
||||
labels:
|
||||
products:
|
||||
- cloud
|
||||
description: How to set up Alerting for Cloud
|
||||
title: Set up Alerting for Cloud
|
||||
weight: 100
|
||||
---
|
||||
|
||||
# Set up Alerting for Cloud
|
||||
|
||||
Configure the features and integrations that you need to create and manage your alerts.
|
||||
|
||||
Grafana Cloud alerts are directly tied to metrics and log data.
|
||||
|
||||
They can be configured either using the UI or by uploading files containing Prometheus and Loki alert rules with mimirtool.
|
||||
|
||||
Grafana Cloud Alerting's Prometheus-style alerts are built by querying directly from the data source itself.
|
||||
|
||||
**Note:**
|
||||
|
||||
These are set up instructions for Grafana Alerting Cloud.
|
||||
|
||||
To set up Grafana Alerting for Open Source, see ({{< relref "../set-up" >}})
|
||||
|
||||
To set up Alerting, you need to:
|
||||
|
||||
1. Configure alert rules
|
||||
|
||||
- Create Mimir/Loki-managed alert rules and recording rules
|
||||
|
||||
2. Configure contact points
|
||||
- Check the default contact point and update the email address
|
||||
- [Optional] Add new contact points and integrations
|
||||
3. Configure notification policies
|
||||
|
||||
- Check the default notification policy
|
||||
- [Optional] Add additional nested policies
|
||||
- [Optional] Add labels and label matchers to control alert routing
|
||||
|
||||
4. [Optional] Integrate with [Grafana OnCall](/docs/oncall/latest/integrations/grafana-alerting/) and [Grafana Incident](/docs/grafana-cloud/incident/set-up/)
|
||||
|
||||
The following topics provide you with advanced configuration options for Grafana Alerting for Cloud.
|
||||
|
||||
## Provision alert rules using mimirtool
|
||||
|
||||
Use `mimirtool` to create and upload alert and recording rules to your Grafana Cloud instance.
|
||||
|
||||
Once created, you can view these alert and recordiing rules from within the Grafana Cloud Alerting page in the UI.
|
||||
|
||||
{{% admonition type="note" %}}
|
||||
`mimirtool` does _not_ support Loki.
|
||||
{{% /admonition %}}
|
||||
|
||||
Prometheus-style alerting is driven by your Grafana Cloud Metrics, Grafana Cloud Logs, and Grafana Cloud Alerts instances. The Metrics and Logs instance holds the rules definition, while the Alerts instance is in charge of routing and managing the alerts that fire from the Metrics and Logs instance. These are separate systems that must be individually configured in order for alerting to work correctly.
|
||||
|
||||
The following sections cover all of these concepts:
|
||||
|
||||
- How to upload alerting and recording rules definition to your Grafana Cloud Metrics instance
|
||||
- How to upload alerting rules definition to your Grafana Cloud Logs instance
|
||||
- How to configure an Alertmanager for your Grafana Cloud Alerts instance, giving you access to the Alertmanager UI.
|
||||
|
||||
**Note:** You need an API key with proper permissions. You can use the same API key for your Metric, Log, and Alerting instances.
|
||||
|
||||
### Download and install mimirtool
|
||||
|
||||
`mimirtool` is a powerful command-line tool for interacting with Grafana Mimir, which powers Grafana Cloud Metrics and Alerts. Use `mimirtool` to upload your metric and log rules definition and the Alertmanager configuration using YAML files.
|
||||
|
||||
For more information, including installation instructions, see [Grafana Mimirtool](/docs/mimir/latest/operators-guide/tools/mimirtool).
|
||||
|
||||
{{% admonition type="note" %}}
|
||||
For `mimirtool` to interact with Grafana Cloud, you must set the correct configuration variables. Set them using either environment variables or a command line flags.
|
||||
{{% /admonition %}}
|
||||
|
||||
### Upload rules definition to your Grafana Cloud Metrics and Logs instance
|
||||
|
||||
First, you'll need to upload your alerting and recording rules to your Metrics and Logs instance. You'll need the instance ID and the URL. These should be part of /orgs/`<yourOrgName>`/.
|
||||
|
||||
**Metrics instance**
|
||||
|
||||
Your Metrics instance is likely to be in the `us-central1` region. Its address would be in the form of [https://prometheus-us-central1.grafana.net](https://prometheus-us-central1.grafana.net).
|
||||
|
||||
**Logs instance**
|
||||
|
||||
Your Logs instance is likely to be in the `us-central1` region. Its address would be in the form of [https://logs-prod-us-central1.grafana.net](https://logs-prod-us-central1.grafana.net).
|
||||
|
||||
### Use mimirtool
|
||||
|
||||
With your instance ID, URL, and API key you're now ready to upload your rules to your metrics instance. Use the following commands and files as a reference.
|
||||
|
||||
Below is an example alert and rule definition YAML file. Take note of the namespace key which replaces the concept of "files" in this context given each instance only supports 1 configuration file.
|
||||
|
||||
```yaml
|
||||
# first_rules.yml
|
||||
namespace: 'first_rules'
|
||||
groups:
|
||||
- name: 'shopping_service_rules_and_alerts'
|
||||
rules:
|
||||
- alert: 'PromScrapeFailed'
|
||||
annotations:
|
||||
message: 'Prometheus failed to scrape a target {{ $labels.job }} / {{ $labels.instance }}'
|
||||
expr: 'up != 1'
|
||||
for: '1m'
|
||||
labels:
|
||||
'severity': 'critical'
|
||||
- record: 'job:up:sum'
|
||||
expr: 'sum by(job) (up)'
|
||||
```
|
||||
|
||||
Although both recording and alerting rules are defined under the key `rules` the difference between a rule and and alert is _generally_ (as there are others) whenever the key `record` or `alert` is defined.
|
||||
|
||||
With this file, you can run the following commands to upload your rules file in your Metrics or Logs instance. Keep in mind that these are example commands for your Metrics instance, and they use placeholders and command line flags. Follow a similar pattern for your Logs instances by switching the address to the correct one. The examples also assume that files are located in the same directory.
|
||||
|
||||
```bash
|
||||
$ mimirtool rules load first_rules.yml \
|
||||
--address=https://prometheus-us-central1.grafana.net \
|
||||
--id=<yourID> \
|
||||
--key=<yourKey>
|
||||
```
|
||||
|
||||
Next, confirm that the rules were uploaded correctly by running:
|
||||
|
||||
```bash
|
||||
$ mimirtool rules list \
|
||||
--address=https://prometheus-us-central1.grafana.net \
|
||||
--id=<yourID> \
|
||||
--key=<yourKey>
|
||||
```
|
||||
|
||||
Output is a list that shows you all the namespaces and rule groups for your instance ID:
|
||||
|
||||
```bash
|
||||
Namespace | Rule Group
|
||||
first_rules | shopping_service_rules_and_alerts
|
||||
```
|
||||
|
||||
You can also print the rules:
|
||||
|
||||
```bash
|
||||
$ mimirtool rules print \
|
||||
--address=https://prometheus-us-central1.grafana.net \
|
||||
--id=<yourID> \
|
||||
--key=<yourKey>
|
||||
```
|
||||
|
||||
Output from the print command should look like this:
|
||||
|
||||
```yaml
|
||||
first_rules:
|
||||
- name: shopping_service_rules_and_alerts
|
||||
interval: 0s
|
||||
rules:
|
||||
- alert: PromScrapeFailed
|
||||
expr: up != 1
|
||||
for: 1m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
message: Prometheus failed to scrape a target {{ $labels.job }} / {{ $labels.instance }}
|
||||
- record: job:up:sum
|
||||
expr: sum by(job) (up)
|
||||
```
|
||||
|
||||
## Add an external Alertmanager using mimirtool
|
||||
|
||||
To receive alerts you need to upload your Alertmanager configuration to your Grafana Cloud Alerts instance. Similar to the previous step, you need the corresponding instance ID, URL and API key. These should be part of /orgs/`<yourOrgName>`/.
|
||||
|
||||
Your Alerts instance is likely to be in the `us-central1` region. Its address would be in the form of [https://alertmanager-us-central1.grafana.net](https://alertmanager-us-central1.grafana.net).
|
||||
|
||||
### Use mimirtool
|
||||
|
||||
With your instance ID, URL, and API key you're now ready to upload your Alertmanager configuration to your Alerts instance. Use the following commands and files as a reference.
|
||||
|
||||
Ultimately, you'll need to [write your own](https://prometheus.io/docs/alerting/latest/configuration/) or adapt an [example config file](https://github.com/prometheus/alertmanager/blob/master/doc/examples/simple.yml) for alerts to be delivered.
|
||||
|
||||
Below is an example Alertmanager configuration. Please take that this not a working configuration, your alerts won't be delivered with the following configuration but your Alertmanager UI will be accessible.
|
||||
|
||||
```yaml
|
||||
# alertmanager.yml
|
||||
global:
|
||||
smtp_smarthost: 'localhost:25'
|
||||
smtp_from: 'youraddress@example.org'
|
||||
route:
|
||||
receiver: example-email
|
||||
receivers:
|
||||
- name: example-email
|
||||
email_configs:
|
||||
- to: 'youraddress@example.org'
|
||||
```
|
||||
|
||||
With this file, you can run the following commands to upload your Alertmanager configuration in your Alerts instance.
|
||||
|
||||
```bash
|
||||
$ mimirtool alertmanager load alertmanager.yml \
|
||||
--address=https://alertmanager-us-central1.grafana.net \
|
||||
--id=<yourID> \
|
||||
--key=<yourKey>
|
||||
```
|
||||
|
||||
Then, confirm that the rules were uploaded correctly by running:
|
||||
|
||||
```bash
|
||||
$ mimirtool alertmanager get \
|
||||
--address=https://alertmanager-us-central1.grafana.net \
|
||||
--id=<yourID> \
|
||||
--key=<yourKey>
|
||||
```
|
||||
|
||||
You should see output similar to the following:
|
||||
|
||||
```bash
|
||||
global:
|
||||
smtp_smarthost: 'localhost:25'
|
||||
smtp_from: 'youraddress@example.org'
|
||||
route:
|
||||
receiver: example-email
|
||||
receivers:
|
||||
- name: example-email
|
||||
email_configs:
|
||||
- to: 'youraddress@example.org'
|
||||
```
|
||||
|
||||
Finally, you can delete the configuration with:
|
||||
|
||||
```bash
|
||||
$ mimirtool alertmanager delete \
|
||||
--address=https://alertmanager-us-central1.grafana.net \
|
||||
--id=<yourID> \
|
||||
--key=<yourKey>
|
||||
```
|
||||
|
||||
### UI access
|
||||
|
||||
After you upload a working Alertmanager configuration file, you can access the Alertmanager UI at: https://alertmanager-us-central1.grafana.net/alertmanager.
|
||||
Reference in New Issue
Block a user