mirror of
https://github.com/rancher/rancher-docs.git
synced 2026-09-29 06:29:34 +00:00
Compare commits
66
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
5ff3581630 | ||
|
|
37b8db8823 | ||
|
|
c15221a43f | ||
|
|
67c0f1c3c4 | ||
|
|
75c792d297 | ||
|
|
64369895cb | ||
|
|
28130a9cd3 | ||
|
|
84b151ba01 | ||
|
|
9f2239b355 | ||
|
|
e348d8ae96 | ||
|
|
f2b53f0297 | ||
|
|
b15636ea3b | ||
|
|
18facc20dc | ||
|
|
0ba3bf937e | ||
|
|
4537a4fe0b | ||
|
|
01bfd00de6 | ||
|
|
97c7ee878e | ||
|
|
24f3481a30 | ||
|
|
76d839d8d9 | ||
|
|
fa45b2448c | ||
|
|
e9f87ecbd6 | ||
|
|
45bd2e660b | ||
|
|
7a818a8616 | ||
|
|
5b880d38b1 | ||
|
|
e5adc9708d | ||
|
|
9ae475be33 | ||
|
|
0e72872376 | ||
|
|
99d72ac6e0 | ||
|
|
184053bc1a | ||
|
|
9d5ae85d8d | ||
|
|
82ea5770cc | ||
|
|
68c68e5715 | ||
|
|
71f3b829d6 | ||
|
|
46c874f466 | ||
|
|
5493a494d8 | ||
|
|
2360f9a7b9 | ||
|
|
8798ed2016 | ||
|
|
7cc72a01c8 | ||
|
|
f25f4a4040 | ||
|
|
9381d6f714 | ||
|
|
78f2f5160e | ||
|
|
2b8743287d | ||
|
|
c5bc0a85a8 | ||
|
|
457d9d8114 | ||
|
|
7bece923f8 | ||
|
|
9858b7a3e5 | ||
|
|
6230e57cae | ||
|
|
0476c4aebe | ||
|
|
bac3129453 | ||
|
|
ed0b4be092 | ||
|
|
5bb0484e29 | ||
|
|
a4bf2e1aea | ||
|
|
897ff93fd6 | ||
|
|
e913bab46b | ||
|
|
a1aab226d2 | ||
|
|
b2be054e7a | ||
|
|
b65bedd4e8 | ||
|
|
b86c4eaedb | ||
|
|
d3cc52f288 | ||
|
|
b15041642a | ||
|
|
d2815b8707 | ||
|
|
f62e904391 | ||
|
|
a19ff57088 | ||
|
|
0c0d5992e0 | ||
|
|
a321c9fa41 | ||
|
|
64269bea22 |
@@ -12,6 +12,7 @@ Rancher will publish deprecated features as part of the [release notes](https://
|
||||
|
||||
| Patch Version | Release Date |
|
||||
|---------------|---------------|
|
||||
| [2.14.2](https://github.com/rancher/rancher/releases/tag/v2.14.2) | May 28, 2026 |
|
||||
| [2.14.1](https://github.com/rancher/rancher/releases/tag/v2.14.1) | April 30, 2026 |
|
||||
| [2.14.0](https://github.com/rancher/rancher/releases/tag/v2.14.0) | March 25, 2026 |
|
||||
|
||||
|
||||
+34
-9
@@ -26,6 +26,31 @@ Review the list of known issues for each Rancher version, which can be found in
|
||||
|
||||
Note that upgrades _to_ or _from_ any chart in the [rancher-alpha repository](../resources/choose-a-rancher-version.md#helm-chart-repositories) aren't supported.
|
||||
|
||||
### Upgrade Path
|
||||
|
||||
:::important
|
||||
|
||||
**Important:** The only tested and supported Rancher upgrade path between minor versions (e.g. v2.13.x to v2.14.x) is to upgrade from the latest available patch version of your current running minor release to the latest available patch version of the next minor release.
|
||||
|
||||
Before initiating a minor version upgrade, verify that you are running the most recent patch release of your current version.
|
||||
|
||||
You can query the available chart versions with the Helm CLI:
|
||||
|
||||
1. Update your local Helm repo cache.
|
||||
|
||||
```
|
||||
helm repo update
|
||||
```
|
||||
|
||||
1. Search for available versions in your [specific repository](../resources/choose-a-rancher-version.md#helm-chart-repositories) (e.g., rancher-stable):
|
||||
|
||||
```
|
||||
helm search repo rancher-<CHART_REPO>/rancher --versions
|
||||
```
|
||||
|
||||
If your installation is not on the latest patch version of the current minor release, you must upgrade to that version before proceeding to the next minor version.
|
||||
:::
|
||||
|
||||
### Helm Version
|
||||
|
||||
:::important
|
||||
@@ -170,17 +195,17 @@ There will be more values that are listed with this command. This is just an exa
|
||||
|
||||
:::
|
||||
|
||||
:::tip
|
||||
:::tip
|
||||
|
||||
Your deployment name may vary; for example, if you're deploying Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
Your deployment name may vary; for example, if you're deploying Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
```
|
||||
helm get values rancher-stable -n cattle-system
|
||||
|
||||
hostname: rancher.my.org
|
||||
```
|
||||
|
||||
:::
|
||||
:::
|
||||
|
||||
If you are upgrading cert-manager to the latest version from v1.5 or below, follow the [cert-manager upgrade docs](../resources/upgrade-cert-manager.md#option-c-upgrade-cert-manager-from-versions-15-and-below) to learn how to upgrade cert-manager without needing to perform an uninstall or reinstall of Rancher. Otherwise, follow the [steps to upgrade Rancher](#steps-to-upgrade-rancher) below.
|
||||
|
||||
@@ -203,17 +228,17 @@ The above is an example, there may be more values from the previous step that ne
|
||||
|
||||
:::
|
||||
|
||||
:::tip
|
||||
:::tip
|
||||
|
||||
If you deploy Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
If you deploy Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
```
|
||||
helm upgrade rancher-stable rancher-<CHART_REPO>/rancher \
|
||||
--namespace cattle-system \
|
||||
--set hostname=rancher.my.org
|
||||
```
|
||||
|
||||
:::
|
||||
:::
|
||||
|
||||
Alternatively, it's possible to export the current values to a file and reference that file during upgrade. For example, to only change the Rancher version:
|
||||
|
||||
@@ -223,7 +248,7 @@ Alternatively, it's possible to export the current values to a file and referenc
|
||||
```
|
||||
1. Update only the Rancher version:
|
||||
|
||||
|
||||
|
||||
```
|
||||
helm upgrade rancher rancher-<CHART_REPO>/rancher \
|
||||
--namespace cattle-system \
|
||||
|
||||
+2
-2
@@ -33,7 +33,7 @@ To use a premade dashboard, go to [https://grafana.com/grafana/dashboards](https
|
||||
To use your own dashboard:
|
||||
|
||||
1. Click on the link to open Grafana. On the cluster detail page, click **Monitoring**.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin/prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
:::note
|
||||
|
||||
@@ -113,7 +113,7 @@ Note that the RBAC roles exposed by the Monitoring chart to add Grafana Dashboar
|
||||
1. On the **Clusters** page, go to the cluster where you want to configure the Grafana namespace and click **Explore**.
|
||||
1. In the left navigation bar, click **Monitoring**.
|
||||
1. Click **Grafana**.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin/prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
:::note
|
||||
|
||||
|
||||
+1
-1
@@ -20,7 +20,7 @@ To see the links to the external monitoring UIs, including Grafana dashboards, y
|
||||
1. In the left navigation menu, click **Monitoring.**
|
||||
1. Click **Grafana.** The Grafana dashboard should open in a new tab.
|
||||
1. Go to the log in icon in the lower left corner and click **Sign In.**
|
||||
1. Log in to Grafana. The default Admin username and password for the Grafana instance is `admin/prom-operator`. (Regardless of who has the password, cluster administrator permission in Rancher is still required access the Grafana instance.) Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. (Regardless of who has the password, cluster administrator permission in Rancher is still required access the Grafana instance.) Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
|
||||
### Getting the PromQL Query Powering a Grafana Panel
|
||||
|
||||
+8
-2
@@ -31,11 +31,17 @@ _Cluster roles_ are roles that you can assign to users, granting them access to
|
||||
|
||||
- **Cluster Owner:**
|
||||
|
||||
These users have full control over the cluster and all resources in it.
|
||||
These users have full control over the cluster and all resources in it.
|
||||
|
||||
- **Cluster Member:**
|
||||
|
||||
These users can view most cluster level resources and create new projects.
|
||||
These users can view most cluster level resources and create new projects.
|
||||
|
||||
:::warning
|
||||
|
||||
When a Cluster Member creates a project, the user is automatically assigned [Project Owner privileges](#project-roles). This grants them comprehensive control over the project and its associated resources, including permissions to deploy workloads. Without enforced [Pod Security Standards (PSS) and Pod Security Admission (PSA)](../pod-security-standards.md), a Cluster Member is able to execute privileged containers in the cluster.
|
||||
|
||||
:::
|
||||
|
||||
#### Custom Cluster Roles
|
||||
|
||||
|
||||
+1
-1
@@ -234,7 +234,7 @@ docker stop <original-rancher-container>
|
||||
|
||||
:::note
|
||||
|
||||
If you wish to keep the original Rancher environment running, you can also restart the cattle-cluster-agent pods on each cluster connected to your Rancher environment.
|
||||
If clusters do not automatically reconnect to the new environment after you have redirected traffic, for example if there is a delay in scaling down the original Rancher instance, you can also restart the cattle-cluster-agent pods on each cluster connected to your Rancher environment.
|
||||
|
||||
```bash
|
||||
kubectl rollout restart deployment cattle-cluster-agent -n cattle-system
|
||||
|
||||
+1
-1
@@ -24,7 +24,7 @@ The following steps can also be performed using the `kubectl` command line tool.
|
||||
:::
|
||||
|
||||
1. Click **☰ > Cluster Management**.
|
||||
1. Choose the cluster you want to provide vSphere storage to and click **Exlpore**.
|
||||
1. Choose the cluster you want to provide vSphere storage to and click **Explore**.
|
||||
1. In the left navigation bar, select **Storage > StorageClasses**.
|
||||
1. Click **Create**.
|
||||
3. Enter a **Name** for the StorageClass.
|
||||
|
||||
+1
@@ -19,6 +19,7 @@ In order to deploy and run the adapter successfully, you need to ensure its vers
|
||||
|
||||
| Rancher Version | Adapter Version |
|
||||
|-----------------|------------------|
|
||||
| v2.14.2 | 109.0.0+up9.0.0 |
|
||||
| v2.14.1 | 109.0.0+up9.0.0 |
|
||||
| v2.14.0 | 109.0.0+up9.0.0 |
|
||||
|
||||
|
||||
@@ -116,6 +116,19 @@ By default, Rancher collects logs for control plane components and node componen
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Resource Exhaustion of `inotify` Watchers and File Descriptors
|
||||
|
||||
When enabling the **Logging** app on Linux systems that heavily monitor the filesystem, you may encounter `Too many open files` or `CrashLoopBackOff` failures related to applications that leverage `inotify` to watch for file changes.
|
||||
|
||||
This happens because the Linux kernel caps the number of files a user can open and the number of directory paths a subsystem can watch simultaneously. To resolve this, you must explicitly increase your `inotify` system limits.
|
||||
|
||||
Below are example commands an admin user can run to increase system limits for `inotify` user instances and watches:
|
||||
|
||||
```shell
|
||||
sysctl -w fs.inotify.max_user_instances=8192
|
||||
sysctl -w fs.inotify.max_user_watches=524288
|
||||
```
|
||||
|
||||
### The Logging Buffer Overloads Pods
|
||||
|
||||
Depending on your configuration, the default buffer size may be too large and cause pod failures. One way to reduce the load is to lower the logger's flush interval. This prevents logs from overfilling the buffer. You can also add more flush threads to handle moments when many logs are attempting to fill the buffer at once.
|
||||
|
||||
@@ -14,64 +14,74 @@ Examples of built-in Rancher extensions are Fleet, Explorer, and Harvester. Exam
|
||||
|
||||
## Prerequisites
|
||||
|
||||
> You must log in as an admin in order to view and interact with the extensions management page.
|
||||
> 1. You must log in as an administrator to view and interact with the extensions management page.
|
||||
> 2. You must enable extension support.
|
||||
|
||||
## Enabling Extension Support in Rancher
|
||||
|
||||
Rancher v2.9.0 and later includes extension support.
|
||||
|
||||
You can confirm if extension support is enabled by checking the `uiextension` feature flag. Any changes to this feature flag cause the Rancher pod to restart.
|
||||
|
||||
When you enable extension support for the first time, it creates resources, such as the CRDs, so that UI Extensions can work. When extension support is disabled, it disables the endpoints and does not cache any files. However, it does not remove any CRs or delete any extensions that were installed before. If re-enabled, it exposes the required endpoints again and creates CRDs as needed. The extensions that were already installed load after the Rancher pod restarts.
|
||||
|
||||
### Caching Extension Files
|
||||
|
||||
By default, Rancher caches every extension file in the file system. You can change that behavior by setting `plugin.noCache` to `true`.
|
||||
|
||||
Rancher does have a cached file size limit of 30MB. If an extension has a file bigger than that, the cache is disabled and `plugin.noCache` is set to `true`, regardless of user input.
|
||||
|
||||
|
||||
## Installing Extensions
|
||||
|
||||
1. Click **☰ > Extensions** under **Configuration**.
|
||||
|
||||
2. If not already installed in **Apps**, you must enable the extension operator by clicking the **Enable** button.
|
||||
2. On the **Extensions** page, select the **Available** tab to choose the extensions that you want to install.
|
||||
|
||||
- Click **OK** to add the Rancher extension repository if your installation is not air-gapped. Otherwise, uncheck the box to do so and click **OK**.
|
||||
3. If no extensions are listed as available, you can manually add the repos:
|
||||
|
||||

|
||||
3.1. On the upper right, click **⋮ > Manage Repositories > Create**.
|
||||
|
||||
3. On the **Extensions** page, click on the **Available** tab to select which extensions you want to install.
|
||||
3.2. Add the desired repo name, making sure to also specify the Git repo URL and the Git branch.
|
||||
|
||||
4. If no extensions are showing as available, you may manually add repos as follows:
|
||||
3.3. Click **Create** in the lower right again to complete.
|
||||
|
||||
4.1. On the upper right of screen, click on **⋮ > Manage Repositories > Create**.
|
||||

|
||||
|
||||
4.2. Add the desired repo name, making sure to also specify the Git Repo URL and the Git Branch.
|
||||
4. Under the **Available** tab, click **Install** on the desired extension and version, as in the example below. You can also update your extension from this screen. The button to **Update** appears on the extension card if an update is available.
|
||||
|
||||
4.3. Click **Create** in the lower right again to complete.
|
||||

|
||||
|
||||

|
||||
5. Click **Reload** after your extension successfully installs to check its status. Updates to the UI aren't visible until you reload the page.
|
||||
|
||||
5. Under the **Available** tab, click **Install** on the desired extension and version as in the example below. You can also update your extension from this screen, as the button to **Update** will appear on the extension if one is available.
|
||||
|
||||

|
||||
|
||||
6. Click the **Reload** page button that will appear after your extension successfully installs. Note that a logged-in user who has just installed an extension will not see a change to the UI **unless** they reload the page.
|
||||
|
||||

|
||||

|
||||
|
||||
## Updating and Upgrading Extensions
|
||||
|
||||
1. Click **☰ > Extensions** under **Configuration**.
|
||||
1. Select the **Updates** tab.
|
||||
1. Click **Update**.
|
||||
2. Select the **Updates** tab.
|
||||
3. Click **Update**.
|
||||
|
||||
If there is a new version of the extension, there will also be an **Update** button visible on the associated card for the extension in the **Available** tab.
|
||||
|
||||
## Deleting Extensions
|
||||
|
||||
1. Click **☰**, then click on the name of your local cluster.
|
||||
1. From the sidebar, select **Apps > Installed Apps**.
|
||||
1. Find the name of the chart you want to delete and select the checkbox next to it.
|
||||
1. Click **Delete**.
|
||||
2. From the sidebar, select **Apps > Installed Apps**.
|
||||
3. Find the name of the chart you want to delete and select the checkbox next to it.
|
||||
4. Click **Delete**.
|
||||
|
||||
## Deleting Extension Repositories
|
||||
|
||||
1. Click **☰ > Extensions** under **Configuration**.
|
||||
1. On the top right, click **⋮ > Manage Repositories**.
|
||||
1. Find the name of the extension repository you want to delete. Select the checkbox next to the repository name, then click **Delete**.
|
||||
2. On the top right, click **⋮ > Manage Repositories**.
|
||||
3. Find the name of the extension repository you want to delete. Select the checkbox next to the repository name, then click **Delete**.
|
||||
|
||||
## Deleting Extension Repository Container Images
|
||||
|
||||
1. Click **☰**, then select **Extensions**, under **Configuration**.
|
||||
1. On the top right, click **⋮ > Manage Extension Catalogs**.
|
||||
1. Find the name of the container image you want to delete, then click **⋮ > Uninstall**.
|
||||
2. On the top right, click **⋮ > Manage Extension Catalogs**.
|
||||
3. Find the name of the container image you want to delete, then click **⋮ > Uninstall**.
|
||||
|
||||
## Uninstalling Extensions
|
||||
|
||||
@@ -79,11 +89,11 @@ There are two ways to uninstall or disable an extension:
|
||||
|
||||
1. Under the **Installed** tab, click the **Uninstall** button on the extension you wish to remove.
|
||||
|
||||

|
||||

|
||||
|
||||
1. On the extensions management page, click **⋮ > Disable Extension Support**. This will disable all installed extensions.
|
||||
2. On the extensions management page, click **⋮ > Disable Extension Support**. This will disable all installed extensions.
|
||||
|
||||

|
||||

|
||||
|
||||
:::caution
|
||||
|
||||
@@ -91,6 +101,12 @@ You must reload the page after disabling extensions or display issues may occur.
|
||||
|
||||
:::
|
||||
|
||||
## Enabling Unauthenticated Access to an Extension
|
||||
|
||||
In Rancher v2.9.0 and later, you can allow unauthenticated access to an extension. You may want to enable unauthenticated access if the extension enables a new locale or adds custom branding. By default, all extensions require user authentication to load.
|
||||
|
||||
To enable unauthenticated access to an extension, set `plugin.noAuth` to `true` in the CR used by the extension.
|
||||
|
||||
## Developing Extensions
|
||||
|
||||
To learn how to develop your own extensions, refer to the official [Getting Started](https://rancher.github.io/dashboard/extensions/extensions-getting-started) guide.
|
||||
@@ -173,8 +189,8 @@ After you successfully set up these resources, you can install the extensions fr
|
||||
1. Click **☰**, then select **Extensions**, under **Configuration**.
|
||||
1. On the top right, click **⋮ > Manage Extension Catalogs**.
|
||||
1. Select the **Import Extension Catalog** button.
|
||||
1. Enter the image address in the **Catalog Image Reference** field.
|
||||
* **(Optional)** If the container image is private, select the secret you just created from the **Pull Secrets** drop-down menu.
|
||||
1. Enter the image address in the **Catalog Image Reference** field.
|
||||
- **(Optional)** If the container image is private, select the secret you just created from the **Pull Secrets** drop-down menu.
|
||||
1. Click **Load**. The extension will now be **Pending**.
|
||||
1. Return to the **Extensions** page.
|
||||
1. Select the **Available** tab, and click **Reload** to make sure that the list of extensions is up to date.
|
||||
@@ -191,7 +207,7 @@ After you mirror the latest changes, follow these steps:
|
||||
1. Click **☰ > Local**.
|
||||
1. From the sidebar, select **Workloads > Deployments**.
|
||||
1. From the namespaces dropdown menu, select **cattle-ui-plugin-system**.
|
||||
1. Find the **cattle-ui-plugin-system** namespace.
|
||||
1. Find the **cattle-ui-plugin-system** namespace.
|
||||
1. Select the `ui-plugin-catalog` deployment.
|
||||
1. Click **⋮ > Edit config**.
|
||||
1. Update the **Container Image** field within the deployment's container with the latest image.
|
||||
|
||||
@@ -20,7 +20,8 @@ Each Rancher version is designed to be compatible with a single version of the w
|
||||
|
||||
| Rancher Version | Webhook Version | Availability in Prime | Availability in Community |
|
||||
|-----------------|-----------------|-----------------------|---------------------------|
|
||||
| v2.14.1 | v0.10.1 | ✓ | ✓ |
|
||||
| v2.14.2 | v0.10.5 | ✓ | ✓ |
|
||||
| v2.14.1 | v0.10.4 | ✓ | ✓ |
|
||||
| v2.14.0 | v0.10.0 | ✗ | ✓ |
|
||||
|
||||
## Why Do We Need It?
|
||||
|
||||
@@ -6,29 +6,63 @@ title: Troubleshooting etcd Nodes
|
||||
<link rel="canonical" href="https://ranchermanager.docs.rancher.com/troubleshooting/kubernetes-components/troubleshooting-etcd-nodes"/>
|
||||
</head>
|
||||
|
||||
This section contains commands and tips for troubleshooting nodes with the `etcd` role.
|
||||
This section contains commands and tips for troubleshooting nodes with the `etcd` role in RKE2 and K3s clusters.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
As RKE2 and K3s rely on `containerd` as the container runtime, `crictl` replaces Docker for container management. Before proceeding with the troubleshooting commands, configure your environment by exporting the following variables:
|
||||
|
||||
### RKE2
|
||||
|
||||
```bash
|
||||
export PATH=$PATH:/var/lib/rancher/rke2/bin/
|
||||
export CRI_CONFIG_FILE=/var/lib/rancher/rke2/agent/etc/crictl.yaml
|
||||
etcdcontainer=$(crictl ps --name etcd --quiet)
|
||||
```
|
||||
|
||||
### K3s
|
||||
|
||||
> ### ⚠️ **Warning**
|
||||
> K3s does not include `etcdctl` in the system PATH. If you need to perform etcd troubleshooting on a K3s cluster, you may need to install it or locate it within the K3s data directory.
|
||||
|
||||
```bash
|
||||
export PATH=$PATH:/usr/local/bin
|
||||
export CRI_CONFIG_FILE=/var/lib/rancher/k3s/agent/etc/crictl.yaml
|
||||
```
|
||||
|
||||
|
||||
## Checking if the etcd Container is Running
|
||||
|
||||
The container for etcd should have status **Up**. The duration shown after **Up** is the time the container has been running.
|
||||
**RKE2**: The container for etcd should be in the **Running** state.
|
||||
|
||||
```
|
||||
docker ps -a -f=name=etcd$
|
||||
```bash
|
||||
crictl ps --name etcd
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES
|
||||
d26adbd23643 rancher/mirrored-coreos-etcd:v3.5.7 "/usr/local/bin/etcd…" 30 minutes ago Up 30 minutes etcd
|
||||
CONTAINER IMAGE CREATED STATE NAME ATTEMPT POD ID POD NAMESPACE
|
||||
f1e289d202ed0 11ad16872a9cf 58 minutes ago Running etcd 0 7b56aab8204ea etcd-cluster1 kube-system
|
||||
```
|
||||
|
||||
## etcd Container Logging
|
||||
|
||||
The logging of the container can contain information on what the problem could be.
|
||||
**K3s**: Etcd runs as an embedded process in the K3s service. Check the service status:
|
||||
|
||||
```bash
|
||||
systemctl status k3s
|
||||
```
|
||||
docker logs etcd
|
||||
|
||||
## etcd Logging
|
||||
|
||||
The logs can contain information on what the problem could be.
|
||||
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl logs $etcdcontainer
|
||||
```
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
journalctl -u k3s | grep -i etcd
|
||||
```
|
||||
| Log | Explanation |
|
||||
|-----|------------------|
|
||||
@@ -46,18 +80,43 @@ The address where etcd is listening depends on the address configuration of the
|
||||
|
||||
Output should contain all the nodes with the `etcd` role and the output should be identical on all nodes.
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
Run the command inside the etcd container.
|
||||
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl member list \
|
||||
--cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt \
|
||||
--key /var/lib/rancher/rke2/server/tls/etcd/server-client.key \
|
||||
--cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec etcd etcdctl member list
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl member list \
|
||||
--cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt \
|
||||
--key /var/lib/rancher/k3s/server/tls/etcd/server-client.key \
|
||||
--cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
1c424074df86e854, started, cluster-node1-f289ac71, https://IP:2380, https://IP:2379, false
|
||||
45c68c44c5a792ff, started, cluster-node2-67e3cf6f, https://IP:2380, https://IP:2379, false
|
||||
7c584f77c5180258, started, cluster-node3-e976bc00, https://IP:2380, https://IP:2379, false
|
||||
```
|
||||
|
||||
### Check Endpoint Status
|
||||
|
||||
The values for `RAFT TERM` should be equal and `RAFT INDEX` should be not be too far apart from each other.
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl endpoint status --write-out table --endpoints=$(crictl exec $etcdcontainer etcdctl member list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec -e ETCDCTL_ENDPOINTS=$(docker exec etcd etcdctl member list | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') etcd etcdctl endpoint status --write-out table
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl endpoint status --write-out table --endpoints=$(etcdctl member list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
@@ -65,17 +124,22 @@ Example output:
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
| ENDPOINT | ID | VERSION | DB SIZE | IS LEADER | RAFT TERM | RAFT INDEX |
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
| https://IP:2379 | 333ef673fc4add56 | 3.5.7 | 24 MB | false | 72 | 66887 |
|
||||
| https://IP:2379 | 5feed52d940ce4cf | 3.5.7 | 24 MB | true | 72 | 66887 |
|
||||
| https://IP:2379 | db6b3bdb559a848d | 3.5.7 | 25 MB | false | 72 | 66887 |
|
||||
| https://IP:2379 | 333ef673fc4add56 | 3.6.7 | 24 MB | false | 72 | 66887 |
|
||||
| https://IP:2379 | 5feed52d940ce4cf | 3.6.7 | 24 MB | true | 72 | 66887 |
|
||||
| https://IP:2379 | db6b3bdb559a848d | 3.6.7 | 25 MB | false | 72 | 66887 |
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
```
|
||||
|
||||
### Check Endpoint Health
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl endpoint health --endpoints=$(crictl exec $etcdcontainer etcdctl member list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec -e ETCDCTL_ENDPOINTS=$(docker exec etcd etcdctl member list | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') etcd etcdctl endpoint health
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl endpoint health --endpoints=$(etcdctl member list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
@@ -84,54 +148,104 @@ https://IP:2379 is healthy: successfully committed proposal: took = 2.113189ms
|
||||
https://IP:2379 is healthy: successfully committed proposal: took = 2.649963ms
|
||||
https://IP:2379 is healthy: successfully committed proposal: took = 2.451201ms
|
||||
```
|
||||
### Check Connectivity on etcd Ports
|
||||
|
||||
### Check Connectivity on Port TCP/2379
|
||||
> In modern versions of Kubernetes, the etcd database (versions 3.5 and newer) introduced significant architectural changes regarding network traffic handling. Previously, etcd permitted standard HTTP REST requests on its primary client port (`2379`). However, to enhance performance and security, etcd 3.5+ strictly enforces the gRPC protocol on this port.<br />
|
||||
If you attempt to use standard HTTP tools like `curl` to test connectivity on port `2379`, the etcd server will automatically terminate the connection or return an error. This behavior often leads administrators to misinterpret the result as a closed port or a node failure.
|
||||
|
||||
Command:
|
||||
Since standard HTTP clients can no longer probe the primary etcd ports, the transport layer must be utilized for network troubleshooting. Using `openssl s_client` instead of `curl` bypasses the gRPC application requirement, allowing the raw TCP and TLS handshake to be tested directly.
|
||||
|
||||
These script isolate the network and security infrastructure from the database application. A successful `Verify return code: 0 (ok)` explicitly confirms four critical infrastructure components:
|
||||
|
||||
* **Network Path:** Routing is functional, and firewalls permit traffic on TCP port `2379` or `2380`.
|
||||
* **Process Availability:** The etcd service is running and actively listening on the designated port.
|
||||
* **Certificate Validity:** The TLS certificates are active, correctly formatted, and have not expired.
|
||||
* **Mutual Authentication (mTLS):** The node successfully authenticates against the cluster's specific Certificate Authority (CA).
|
||||
|
||||
**How these tests differ from the `etcdctl endpoint health` test**:
|
||||
|
||||
If `etcdctl endpoint health` test is failing, run these Connectivity Ports test scripts. If the scripts succeed, your network and certificates are intact, and the issue is likely confined to the etcd database itself. If these scripts fail, the issue is related to a firewall/network restriction, or certificate expiration.
|
||||
|
||||
#### Port TCP/2379
|
||||
|
||||
**RKE2**:
|
||||
```bash
|
||||
for endpoint in $(crictl exec $etcdcontainer etcdctl member list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | cut -d, -f5); do
|
||||
echo "Validating connection to ${endpoint} (Client)";
|
||||
echo | openssl s_client -connect ${endpoint#https://} \
|
||||
-CAfile /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt \
|
||||
-cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt \
|
||||
-key /var/lib/rancher/rke2/server/tls/etcd/server-client.key 2>/dev/null | grep -E 'Verify return code' || echo "Connection Failed/Timeout"
|
||||
done
|
||||
```
|
||||
for endpoint in $(docker exec etcd etcdctl member list | cut -d, -f5); do
|
||||
echo "Validating connection to ${endpoint}/health"
|
||||
docker run --net=host -v $(docker inspect kubelet --format '{{ range .Mounts }}{{ if eq .Destination "/etc/kubernetes" }}{{ .Source }}{{ end }}{{ end }}')/ssl:/etc/kubernetes/ssl:ro appropriate/curl -s -w "\n" --cacert $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_CACERT" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) --cert $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_CERT" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) --key $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_KEY" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) "${endpoint}/health"
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
for endpoint in $(etcdctl member list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | cut -d, -f5); do
|
||||
echo "Validating connection to ${endpoint} (Client)";
|
||||
echo | openssl s_client -connect ${endpoint#https://} \
|
||||
-CAfile /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt \
|
||||
-cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt \
|
||||
-key /var/lib/rancher/k3s/server/tls/etcd/server-client.key 2>/dev/null | grep -E 'Verify return code' || echo "Connection Failed/Timeout"
|
||||
done
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
Validating connection to https://IP:2379/health
|
||||
{"health": "true"}
|
||||
Validating connection to https://IP:2379/health
|
||||
{"health": "true"}
|
||||
Validating connection to https://IP:2379/health
|
||||
{"health": "true"}
|
||||
Validating connection to https://IP:2379/health (Client)
|
||||
Verify return code: 0 (ok)
|
||||
Validating connection to https://IP:2379/health (Client)
|
||||
Verify return code: 0 (ok)
|
||||
Validating connection to https://IP:2379/health (Client)
|
||||
Verify return code: 0 (ok)
|
||||
```
|
||||
|
||||
### Check Connectivity on Port TCP/2380
|
||||
#### Port TCP/2380
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
for endpoint in $(crictl exec $etcdcontainer etcdctl member list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | cut -d, -f4); do
|
||||
echo "Validating connection to ${endpoint} (Peer)";
|
||||
echo | openssl s_client -connect ${endpoint#https://} \
|
||||
-CAfile /var/lib/rancher/rke2/server/tls/etcd/peer-ca.crt \
|
||||
-cert /var/lib/rancher/rke2/server/tls/etcd/peer-server-client.crt \
|
||||
-key /var/lib/rancher/rke2/server/tls/etcd/peer-server-client.key 2>/dev/null | grep -E 'Verify return code' || echo "Connection Failed/Timeout"
|
||||
done
|
||||
```
|
||||
for endpoint in $(docker exec etcd etcdctl member list | cut -d, -f4); do
|
||||
echo "Validating connection to ${endpoint}/version";
|
||||
docker run --net=host -v $(docker inspect kubelet --format '{{ range .Mounts }}{{ if eq .Destination "/etc/kubernetes" }}{{ .Source }}{{ end }}{{ end }}')/ssl:/etc/kubernetes/ssl:ro appropriate/curl --http1.1 -s -w "\n" --cacert $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_CACERT" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) --cert $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_CERT" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) --key $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_KEY" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) "${endpoint}/version"
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
for endpoint in $(etcdctl member list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | cut -d, -f4); do
|
||||
echo "Validating connection to ${endpoint} (Peer)";
|
||||
echo | openssl s_client -connect ${endpoint#https://} \
|
||||
-CAfile /var/lib/rancher/k3s/server/tls/etcd/peer-ca.crt \
|
||||
-cert /var/lib/rancher/k3s/server/tls/etcd/peer-server-client.crt \
|
||||
-key /var/lib/rancher/k3s/server/tls/etcd/peer-server-client.key 2>/dev/null | grep -E 'Verify return code' || echo "Connection Failed/Timeout"
|
||||
done
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
Validating connection to https://IP:2380/version
|
||||
{"etcdserver":"3.5.7","etcdcluster":"3.5.0"}
|
||||
Validating connection to https://IP:2380/version
|
||||
{"etcdserver":"3.5.7","etcdcluster":"3.5.0"}
|
||||
Validating connection to https://IP:2380/version
|
||||
{"etcdserver":"3.5.7","etcdcluster":"3.5.0"}
|
||||
Validating connection to https://IP:2380/version (Peer)
|
||||
Verify return code: 0 (ok)
|
||||
Validating connection to https://IP:2380/version (Peer)
|
||||
Verify return code: 0 (ok)
|
||||
Validating connection to https://IP:2380/version (Peer)
|
||||
Verify return code: 0 (ok)
|
||||
```
|
||||
|
||||
## etcd Alarms
|
||||
|
||||
etcd will trigger alarms, for instance when it runs out of space.
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl alarm list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec etcd etcdctl alarm list
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl alarm list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output when NOSPACE alarm is triggered:
|
||||
@@ -154,10 +268,16 @@ Resolutions:
|
||||
|
||||
### Compact the Keyspace
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
rev=$(crictl exec $etcdcontainer etcdctl endpoint status --write-out json --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | egrep -o '"revision":[0-9]*' | egrep -o '[0-9]*' | head -1)
|
||||
crictl exec $etcdcontainer etcdctl compact "$rev" --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
rev=$(docker exec etcd etcdctl endpoint status --write-out json | egrep -o '"revision":[0-9]*' | egrep -o '[0-9]*')
|
||||
docker exec etcd etcdctl compact "$rev"
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
rev=$(etcdctl endpoint status --write-out json --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | egrep -o '"revision":[0-9]*' | egrep -o '[0-9]*' | head -1)
|
||||
etcdctl compact "$rev" --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
@@ -167,55 +287,39 @@ compacted revision xxx
|
||||
|
||||
### Defrag All etcd Members
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl defrag --endpoints=$(crictl exec $etcdcontainer etcdctl member list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec -e ETCDCTL_ENDPOINTS=$(docker exec etcd etcdctl member list | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') etcd etcdctl defrag
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl defrag --endpoints=$(etcdctl member list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
Finished defragmenting etcd member[https://IP:2379]
|
||||
Finished defragmenting etcd member[https://IP:2379]
|
||||
Finished defragmenting etcd member[https://IP:2379]
|
||||
```
|
||||
|
||||
### Check Endpoint Status
|
||||
|
||||
Command:
|
||||
```
|
||||
docker exec -e ETCDCTL_ENDPOINTS=$(docker exec etcd etcdctl member list | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') etcd etcdctl endpoint status --write-out table
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
| ENDPOINT | ID | VERSION | DB SIZE | IS LEADER | RAFT TERM | RAFT INDEX |
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
| https://IP:2379 | e973e4419737125 | 3.5.7 | 553 kB | false | 32 | 2449410 |
|
||||
| https://IP:2379 | 4a509c997b26c206 | 3.5.7 | 553 kB | false | 32 | 2449410 |
|
||||
| https://IP:2379 | b217e736575e9dd3 | 3.5.7 | 553 kB | true | 32 | 2449410 |
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
Finished defragmenting etcd member[https://IP:2379]. took xx.xxxxxxms
|
||||
Finished defragmenting etcd member[https://IP:2379]. took xx.xxxxxxms
|
||||
Finished defragmenting etcd member[https://IP:2379]. took xx.xxxxxxms
|
||||
```
|
||||
|
||||
### Disarm Alarm
|
||||
|
||||
After verifying that the DB size went down after compaction and defragmenting, the alarm needs to be disarmed for etcd to allow writes again.
|
||||
|
||||
Command:
|
||||
```
|
||||
docker exec etcd etcdctl alarm list
|
||||
docker exec etcd etcdctl alarm disarm
|
||||
docker exec etcd etcdctl alarm list
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl alarm list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
crictl exec $etcdcontainer etcdctl alarm disarm --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
crictl exec $etcdcontainer etcdctl alarm list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
docker exec etcd etcdctl alarm list
|
||||
memberID:x alarm:NOSPACE
|
||||
memberID:x alarm:NOSPACE
|
||||
memberID:x alarm:NOSPACE
|
||||
docker exec etcd etcdctl alarm disarm
|
||||
docker exec etcd etcdctl alarm list
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl alarm list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
etcdctl alarm disarm --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
etcdctl alarm list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
## Configure Log Level
|
||||
@@ -228,7 +332,7 @@ You can no longer dynamically change the log level in etcd v3.5 or later.
|
||||
|
||||
### etcd v3.5 And Later
|
||||
|
||||
To configure the log level for etcd, edit the cluster YAML:
|
||||
To configure the log level for etcd, edit the cluster configuration YAML:
|
||||
|
||||
```
|
||||
services:
|
||||
@@ -237,20 +341,7 @@ services:
|
||||
log-level: "debug"
|
||||
```
|
||||
|
||||
### etcd v3.4 And Earlier
|
||||
|
||||
In earlier etcd versions, you can use the API to dynamically change the log level. Configure debug logging using the commands below:
|
||||
|
||||
```
|
||||
docker run --net=host -v $(docker inspect kubelet --format '{{ range .Mounts }}{{ if eq .Destination "/etc/kubernetes" }}{{ .Source }}{{ end }}{{ end }}')/ssl:/etc/kubernetes/ssl:ro appropriate/curl -s -XPUT -d '{"Level":"DEBUG"}' --cacert $(docker exec etcd printenv ETCDCTL_CACERT) --cert $(docker exec etcd printenv ETCDCTL_CERT) --key $(docker exec etcd printenv ETCDCTL_KEY) $(docker exec etcd printenv ETCDCTL_ENDPOINTS)/config/local/log
|
||||
```
|
||||
|
||||
To reset the log level back to the default (`INFO`), you can use the following command.
|
||||
|
||||
Command:
|
||||
```
|
||||
docker run --net=host -v $(docker inspect kubelet --format '{{ range .Mounts }}{{ if eq .Destination "/etc/kubernetes" }}{{ .Source }}{{ end }}{{ end }}')/ssl:/etc/kubernetes/ssl:ro appropriate/curl -s -XPUT -d '{"Level":"INFO"}' --cacert $(docker exec etcd printenv ETCDCTL_CACERT) --cert $(docker exec etcd printenv ETCDCTL_CERT) --key $(docker exec etcd printenv ETCDCTL_KEY) $(docker exec etcd printenv ETCDCTL_ENDPOINTS)/config/local/log
|
||||
```
|
||||
After modifying the configuration, restart the service (`systemctl restart rke2-server` or `systemctl restart k3s`) if you are configuring a stand-alone cluster.
|
||||
|
||||
## etcd Content
|
||||
|
||||
@@ -258,24 +349,40 @@ If you want to investigate the contents of your etcd, you can either watch strea
|
||||
|
||||
### Watch Streaming Events
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl watch --prefix /registry --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec etcd etcdctl watch --prefix /registry
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl watch --prefix /registry --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
If you only want to see the affected keys (and not the binary data), you can append `| grep -a ^/registry` to the command to filter for keys only.
|
||||
|
||||
### Query etcd Directly
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl get /registry --prefix=true --keys-only --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec etcd etcdctl get /registry --prefix=true --keys-only
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl get /registry --prefix=true --keys-only --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
You can process the data to get a summary of count per key, using the command below:
|
||||
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl get /registry --prefix=true --keys-only --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | grep -v ^$ | awk -F'/' '{ if ($3 ~ /cattle.io/) {h[$3"/"$4]++} else { h[$3]++ }} END { for(k in h) print h[k], k }' | sort -nr
|
||||
```
|
||||
docker exec etcd etcdctl get /registry --prefix=true --keys-only | grep -v ^$ | awk -F'/' '{ if ($3 ~ /cattle.io/) {h[$3"/"$4]++} else { h[$3]++ }} END { for(k in h) print h[k], k }' | sort -nr
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl get /registry --prefix=true --keys-only --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | grep -v ^$ | awk -F'/' '{ if ($3 ~ /cattle.io/) {h[$3"/"$4]++} else { h[$3]++ }} END { for(k in h) print h[k], k }' | sort -nr
|
||||
```
|
||||
|
||||
## Replacing Unhealthy etcd Nodes
|
||||
|
||||
+68
-14
@@ -8,31 +8,85 @@ title: Troubleshooting Worker Nodes and Generic Components
|
||||
|
||||
This section applies to every node as it includes components that run on nodes with any role.
|
||||
|
||||
## Check if the Containers are Running
|
||||
## Prerequisites
|
||||
|
||||
There are two specific containers launched on nodes with the `worker` role:
|
||||
Since RKE2 and K3s utilize `containerd` as the container runtime, `crictl` serves as the primary tool for container management, replacing the Docker CLI. To allow `crictl` to communicate with `containerd`, you must configure your environment by exporting the following variables:
|
||||
|
||||
* kubelet
|
||||
* kube-proxy
|
||||
|
||||
The containers should have status `Up`. The duration shown after `Up` is the time the container has been running.
|
||||
### RKE2
|
||||
|
||||
```bash
|
||||
export PATH=$PATH:/var/lib/rancher/rke2/bin/
|
||||
export CRI_CONFIG_FILE=/var/lib/rancher/rke2/agent/etc/crictl.yaml
|
||||
```
|
||||
docker ps -a -f=name='kubelet|kube-proxy'
|
||||
|
||||
### K3s
|
||||
|
||||
```bash
|
||||
export PATH=$PATH:/usr/local/bin
|
||||
export CRI_CONFIG_FILE=/var/lib/rancher/k3s/agent/etc/crictl.yaml
|
||||
```
|
||||
|
||||
## Check if the Components are Running
|
||||
|
||||
There are two specific components launched on nodes with the `worker` role:
|
||||
|
||||
* `kubelet`
|
||||
* `kube-proxy`
|
||||
|
||||
### RKE2
|
||||
|
||||
The `kubelet` runs natively as part of the `rke2-agent` (or `rke2-server`) systemd process, while `kube-proxy` runs as a Static Pod managed by `containerd`.
|
||||
|
||||
Check the status of the `kubelet` via the agent service:
|
||||
```bash
|
||||
systemctl status rke2-agent
|
||||
```
|
||||
:::note
|
||||
|
||||
If you are checking a controlplane node, use `systemctl status rke2-server` instead.
|
||||
|
||||
:::
|
||||
|
||||
Check the status of `kube-proxy` using `crictl`:
|
||||
```bash
|
||||
crictl ps --name kube-proxy
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES
|
||||
158d0dcc33a5 rancher/hyperkube:v1.11.5-rancher1 "/opt/rke-tools/en..." 3 hours ago Up 3 hours kube-proxy
|
||||
a30717ecfb55 rancher/hyperkube:v1.11.5-rancher1 "/opt/rke-tools/en..." 3 hours ago Up 3 hours kubelet
|
||||
CONTAINER IMAGE CREATED STATE NAME ATTEMPT POD ID
|
||||
26c7159abbcc rancher/hardened-kubernetes:v1.28.8-rke2r1-build20240404 3 hours ago Running kube-proxy 0 1a2b3c4d5e6f7
|
||||
```
|
||||
|
||||
## Container Logging
|
||||
### K3s
|
||||
|
||||
The logging of the containers can contain information on what the problem could be.
|
||||
Both `kubelet` and `kube-proxy` run as embedded processes inside the `k3s-agent` (or `k3s` server) systemd service. There are no separate containers for them.
|
||||
|
||||
Check their status by checking the K3s service:
|
||||
```bash
|
||||
systemctl status k3s-agent
|
||||
```
|
||||
docker logs kubelet
|
||||
docker logs kube-proxy
|
||||
:::note
|
||||
If you are checking a controlplane node, use `systemctl status k3s` instead.
|
||||
:::
|
||||
|
||||
## Component Logging
|
||||
|
||||
The logging of the components can contain information on what the problem could be.
|
||||
|
||||
### RKE2
|
||||
|
||||
```bash
|
||||
# kubelet logs are part of the systemd service
|
||||
journalctl -u rke2-agent -f | grep -i "kubelet"
|
||||
|
||||
# kube-proxy logs are retrieved from containerd
|
||||
crictl logs $(crictl ps --name kube-proxy -q)
|
||||
```
|
||||
|
||||
### K3s
|
||||
|
||||
```bash
|
||||
# Both components log to the systemd service
|
||||
journalctl -u k3s-agent -f | grep -iE "kubelet|kube-proxy"
|
||||
```
|
||||
|
||||
@@ -78,66 +78,137 @@ kubectl -n kube-system get endpoints kube-scheduler -o jsonpath='{.metadata.anno
|
||||
|
||||
## Ingress Controller
|
||||
|
||||
The default Ingress Controller is Traefik and is deployed as a DaemonSet in the `traefik` namespace. The pods are only scheduled to nodes with the `worker` role.
|
||||
The default Ingress Controller is Traefik and is deployed as a DaemonSet in the `kube-system` namespace. The pods are only scheduled to nodes with the `worker` role.
|
||||
|
||||
Check if the pods are running on all nodes:
|
||||
|
||||
```
|
||||
kubectl -n traefik get pods -o wide
|
||||
kubectl -n kube-system get pods -o wide
|
||||
```
|
||||
|
||||
Example output:
|
||||
Example RKE2 output:
|
||||
|
||||
```
|
||||
kubectl -n traefik get pods -o wide
|
||||
kubectl -n kube-system get pods -o wide
|
||||
NAME READY STATUS RESTARTS AGE IP NODE
|
||||
default-http-backend-797c5bc547-kwwlq 1/1 Running 0 17m x.x.x.x worker-1
|
||||
traefik-4qd64 1/1 Running 0 14m x.x.x.x worker-1
|
||||
traefik-8wxhm 1/1 Running 0 13m x.x.x.x worker-0
|
||||
local-path-provisioner-xxxxxxxxxx-xxxxx 1/1 Running 0 17m x.x.x.x worker-1
|
||||
rke2-traefik-xxxxxxxxxx-xxxxx 1/1 Running 0 14m x.x.x.x worker-1
|
||||
svclb-rke2-traefik-xxxxxxxx-xxxxx 1/1 Running 0 13m x.x.x.x worker-0
|
||||
...
|
||||
```
|
||||
|
||||
Example K3s output:
|
||||
|
||||
```
|
||||
kubectl -n kube-system get pods -o wide
|
||||
NAME READY STATUS RESTARTS AGE IP NODE
|
||||
local-path-provisioner-xxxxxxxxxx-xxxxx 1/1 Running 0 17m x.x.x.x worker-1
|
||||
traefik-xxxxxxxxxx-xxxxx 1/1 Running 0 14m x.x.x.x worker-1
|
||||
svclb-traefik-xxxxxxxx-xxxxx 1/1 Running 0 13m x.x.x.x worker-0
|
||||
...
|
||||
```
|
||||
|
||||
If a pod is unable to run (Status is not **Running**, Ready status is not showing `1/1` or you see a high count of Restarts), check the pod details, logs and namespace events.
|
||||
|
||||
### Pod details
|
||||
|
||||
RKE2 example:
|
||||
|
||||
```
|
||||
kubectl -n traefik describe pods -l app=traefik
|
||||
kubectl -n kube-system describe pods -l app.kubernetes.io/name=rke2-traefik
|
||||
```
|
||||
|
||||
K3s example:
|
||||
|
||||
```
|
||||
kubectl -n kube-system describe pods -l app.kubernetes.io/name=traefik
|
||||
```
|
||||
|
||||
### Pod container logs
|
||||
|
||||
The below command can show the logs of all the pods labeled "app=traefik", but it will display only 10 lines of log because of the restrictions of the `kubectl logs` command. Refer to `--tail` of `kubectl logs -h` for more information.
|
||||
The below command can show the logs of all the pods labeled "app.kubernetes.io/name=rke2-traefik" if using RKE2 or "app.kubernetes.io/name=traefik" if using K3s, but it will display only 10 lines of log because of the restrictions of the `kubectl logs` command. Refer to `--tail` of `kubectl logs -h` for more information.
|
||||
|
||||
RKE2 example:
|
||||
|
||||
```
|
||||
kubectl -n traefik logs -l app=traefik
|
||||
kubectl -n kube-system logs -l app.kubernetes.io/name=rke2-traefik
|
||||
```
|
||||
|
||||
K3s example:
|
||||
|
||||
```
|
||||
kubectl -n kube-system logs -l app.kubernetes.io/name=traefik
|
||||
```
|
||||
|
||||
If the full log is needed, specify the pod name in the trailing command:
|
||||
|
||||
```
|
||||
kubectl -n traefik logs <pod name>
|
||||
kubectl -n kube-system logs <pod name>
|
||||
```
|
||||
|
||||
### Namespace events
|
||||
|
||||
```
|
||||
kubectl -n traefik get events
|
||||
kubectl -n kube-system get events
|
||||
```
|
||||
|
||||
### Debug logging
|
||||
|
||||
To enable debug logging:
|
||||
|
||||
RKE2 example:
|
||||
|
||||
```
|
||||
kubectl -n traefik patch ds traefik --type='json' -p='[{"op": "add", "path": "/spec/template/spec/containers/0/args/-", "value": "--v=5"}]'
|
||||
cat <<EOF | kubectl apply -f -
|
||||
apiVersion: helm.cattle.io/v1
|
||||
kind: HelmChartConfig
|
||||
metadata:
|
||||
name: rke2-traefik
|
||||
namespace: kube-system
|
||||
spec:
|
||||
valuesContent: |-
|
||||
additionalArguments:
|
||||
- "--log.level=DEBUG"
|
||||
EOF
|
||||
```
|
||||
|
||||
K3s example:
|
||||
|
||||
```
|
||||
cat <<EOF | kubectl apply -f -
|
||||
apiVersion: helm.cattle.io/v1
|
||||
kind: HelmChartConfig
|
||||
metadata:
|
||||
name: traefik
|
||||
namespace: kube-system
|
||||
spec:
|
||||
valuesContent: |-
|
||||
logs:
|
||||
general:
|
||||
level: "DEBUG"
|
||||
EOF
|
||||
```
|
||||
|
||||
### Check configuration
|
||||
|
||||
Retrieve generated configuration in each pod:
|
||||
|
||||
RKE2 example for manual Traefik configuration file check:
|
||||
|
||||
```
|
||||
kubectl -n traefik get pods -l app=traefik --no-headers -o custom-columns=.NAME:.metadata.name | while read pod; do kubectl -n traefik exec $pod -- cat /etc/nginx/nginx.conf; done
|
||||
kubectl exec -n kube-system pod/traefik-xxxxxxxxx-xxxxx -- cat /var/lib/rancher/rke2/server/manifests/rke2-traefik-config.yaml
|
||||
```
|
||||
|
||||
K3s example for manual Traefik configuration file check:
|
||||
|
||||
```
|
||||
kubectl exec -n kube-system pod/traefik-xxxxxxxxx-xxxxx -- cat /var/lib/rancher/k3s/server/manifests/k3s-traefik-config.yaml
|
||||
```
|
||||
|
||||
RKE2/K3s example for Traefik CLI argument configuration check:
|
||||
|
||||
```
|
||||
kubectl get pod traefik-xxxxxxxxx-xxxxx -n kube-system -o jsonpath='{.spec.containers[0].args}'
|
||||
```
|
||||
|
||||
## Rancher agents
|
||||
|
||||
@@ -14,6 +14,13 @@ Make sure you configured the correct kubeconfig (for example, `export KUBECONFIG
|
||||
|
||||
Double check if all the [required ports](../../how-to-guides/new-user-guides/kubernetes-clusters-in-rancher-setup/node-requirements-for-rancher-managed-clusters.md#networking-requirements) are opened in your (host) firewall. The overlay network uses UDP in comparison to all other required ports which are TCP.
|
||||
|
||||
## Check if your downstream node can communicate to Rancher Manager
|
||||
|
||||
Rancher components with HTTP endpoints generally contain a `ping` liveness probe which you can use to test connectivity. Replace the `$RANCHER_URL` as appropriate and run the following from a node to check that it has connectivity to Rancher Manager's servers in the `local` cluster. If successful, it should return `pong`.
|
||||
|
||||
```
|
||||
curl -k https://$RANCHER_URL/ping
|
||||
```
|
||||
|
||||
## Check if Overlay Network is Functioning Correctly
|
||||
|
||||
|
||||
@@ -12,6 +12,7 @@ Rancher 将在 GitHub 上发布的 Rancher 的[发版说明](https://github.com/
|
||||
|
||||
| Patch 版本 | 发布时间 |
|
||||
| ----------------------------------------------------------------- | ------------------ |
|
||||
| [2.14.2](https://github.com/rancher/rancher/releases/tag/v2.14.2) | 2026 年 05 月 28 日 |
|
||||
| [2.14.1](https://github.com/rancher/rancher/releases/tag/v2.14.1) | 2026 年 04 月 30 日 |
|
||||
| [2.14.0](https://github.com/rancher/rancher/releases/tag/v2.14.0) | 2026 年 03 月 25 日 |
|
||||
|
||||
|
||||
+1
@@ -15,6 +15,7 @@ title: 安装 Adapter
|
||||
|
||||
| Rancher 版本 | Adapter 版本 |
|
||||
|-----------------|------------------|
|
||||
| v2.14.2 | 109.0.0+up9.0.0 |
|
||||
| v2.14.1 | 109.0.0+up9.0.0 |
|
||||
| v2.14.0 | 109.0.0+up9.0.0 |
|
||||
|
||||
|
||||
+13
@@ -79,6 +79,19 @@ Rancher Logging 有两个角色,分别是 `logging-admin` 和 `logging-view`
|
||||
|
||||
## 故障排除
|
||||
|
||||
### Resource Exhaustion of `inotify` Watchers and File Descriptors
|
||||
|
||||
When enabling the **Logging** app on Linux systems that heavily monitor the filesystem, you may encounter `Too many open files` or `CrashLoopBackOff` failures related to applications that leverage `inotify` to watch for file changes.
|
||||
|
||||
This happens because the Linux kernel caps the number of files a user can open and the number of directory paths a subsystem can watch simultaneously. To resolve this, you must explicitly increase your `inotify` system limits.
|
||||
|
||||
Below are example commands an admin user can run to increase system limits for `inotify` user instances and watches:
|
||||
|
||||
```shell
|
||||
sysctl -w fs.inotify.max_user_instances=8192
|
||||
sysctl -w fs.inotify.max_user_watches=524288
|
||||
```
|
||||
|
||||
### 日志缓冲区导致 Pod 过载
|
||||
|
||||
根据你的配置,默认缓冲区大小可能太大并导致 Pod 故障。减少负载的一种方法是降低记录器的刷新间隔。这可以防止日志溢出缓冲区。你还可以添加更多刷新线程来处理大量日志试图同时填充缓冲区的情况。
|
||||
|
||||
@@ -20,7 +20,8 @@ Rancher 将 Rancher-Webhook 作为单独的 deployment 和服务部署在 local
|
||||
|
||||
| Rancher Version | Webhook Version | Availability in Prime | Availability in Community |
|
||||
|-----------------|-----------------|-----------------------|---------------------------|
|
||||
| v2.14.1 | v0.10.1 | ✓ | ✓ |
|
||||
| v2.14.2 | v0.10.5 | ✓ | ✓ |
|
||||
| v2.14.1 | v0.10.4 | ✓ | ✓ |
|
||||
| v2.14.0 | v0.10.0 | ✗ | ✓ |
|
||||
|
||||
## 为什么我们需要它?
|
||||
|
||||
@@ -12,6 +12,7 @@ Rancher 将在 GitHub 上发布的 Rancher 的[发版说明](https://github.com/
|
||||
|
||||
| Patch 版本 | 发布时间 |
|
||||
| --------------------------------------------------------------- | -------------------- |
|
||||
| [2.10.12](https://github.com/rancher/rancher/releases/tag/v2.10.12) | 2026 年 05 月 27 日 |
|
||||
| [2.10.11](https://github.com/rancher/rancher/releases/tag/v2.10.11) | 2026 年 01 月 29 日 |
|
||||
| [2.10.10](https://github.com/rancher/rancher/releases/tag/v2.10.10) | 2025 年 9 月 25 日 |
|
||||
| [2.10.9](https://github.com/rancher/rancher/releases/tag/v2.10.9) | 2025 年 8 月 27 日 |
|
||||
|
||||
+1
@@ -15,6 +15,7 @@ title: 安装 Adapter
|
||||
|
||||
| Rancher 版本 | Adapter 版本 |
|
||||
|-----------------|:----------------:|
|
||||
| v2.10.12 | v105.0.0+up5.0.1 |
|
||||
| v2.10.11 | v105.0.0+up5.0.1 |
|
||||
| v2.10.10 | v105.0.0+up5.0.1 |
|
||||
| v2.10.9 | v105.0.0+up5.0.1 |
|
||||
|
||||
+13
@@ -79,6 +79,19 @@ Rancher Logging 有两个角色,分别是 `logging-admin` 和 `logging-view`
|
||||
|
||||
## 故障排除
|
||||
|
||||
### Resource Exhaustion of `inotify` Watchers and File Descriptors
|
||||
|
||||
When enabling the **Logging** app on Linux systems that heavily monitor the filesystem, you may encounter `Too many open files` or `CrashLoopBackOff` failures related to applications that leverage `inotify` to watch for file changes.
|
||||
|
||||
This happens because the Linux kernel caps the number of files a user can open and the number of directory paths a subsystem can watch simultaneously. To resolve this, you must explicitly increase your `inotify` system limits.
|
||||
|
||||
Below are example commands an admin user can run to increase system limits for `inotify` user instances and watches:
|
||||
|
||||
```shell
|
||||
sysctl -w fs.inotify.max_user_instances=8192
|
||||
sysctl -w fs.inotify.max_user_watches=524288
|
||||
```
|
||||
|
||||
### 日志缓冲区导致 Pod 过载
|
||||
|
||||
根据你的配置,默认缓冲区大小可能太大并导致 Pod 故障。减少负载的一种方法是降低记录器的刷新间隔。这可以防止日志溢出缓冲区。你还可以添加更多刷新线程来处理大量日志试图同时填充缓冲区的情况。
|
||||
|
||||
+1
@@ -20,6 +20,7 @@ Rancher 将 Rancher-Webhook 作为单独的 deployment 和服务部署在 local
|
||||
|
||||
| Rancher Version | Webhook Version | Availability in Prime | Availability in Community |
|
||||
| --------------- | --------------- | --------------------- | ------------------------- |
|
||||
| v2.10.12 | v0.6.12 | ✓ | ✗ |
|
||||
| v2.10.11 | v0.6.12 | ✓ | ✗ |
|
||||
| v2.10.10 | v0.6.11 | ✓ | ✗ |
|
||||
| v2.10.9 | v0.6.10 | ✓ | ✗ |
|
||||
|
||||
@@ -12,6 +12,7 @@ Rancher 将在 GitHub 上发布的 Rancher 的[发版说明](https://github.com/
|
||||
|
||||
| Patch 版本 | 发布时间 |
|
||||
| --------------------------------------------------------------- | ------------------ |
|
||||
| [2.11.14](https://github.com/rancher/rancher/releases/tag/v2.11.14) | 2026 年 05 月 27 日 |
|
||||
| [2.11.13](https://github.com/rancher/rancher/releases/tag/v2.11.13) | 2026 年 04 月 30 日 |
|
||||
| [2.11.12](https://github.com/rancher/rancher/releases/tag/v2.11.12) | 2026 年 03 月 25 日 |
|
||||
| [2.11.11](https://github.com/rancher/rancher/releases/tag/v2.11.11) | 2026 年 02 月 25 日 |
|
||||
|
||||
+1
@@ -15,6 +15,7 @@ title: 安装 Adapter
|
||||
|
||||
| Rancher 版本 | Adapter 版本 |
|
||||
|-----------------|:----------------:|
|
||||
| v2.11.14 | v106.0.1+up6.0.1 |
|
||||
| v2.11.13 | v106.0.0+up6.0.0 |
|
||||
| v2.11.12 | v106.0.0+up6.0.0 |
|
||||
| v2.11.11 | v106.0.0+up6.0.0 |
|
||||
|
||||
+13
@@ -79,6 +79,19 @@ Rancher Logging 有两个角色,分别是 `logging-admin` 和 `logging-view`
|
||||
|
||||
## 故障排除
|
||||
|
||||
### Resource Exhaustion of `inotify` Watchers and File Descriptors
|
||||
|
||||
When enabling the **Logging** app on Linux systems that heavily monitor the filesystem, you may encounter `Too many open files` or `CrashLoopBackOff` failures related to applications that leverage `inotify` to watch for file changes.
|
||||
|
||||
This happens because the Linux kernel caps the number of files a user can open and the number of directory paths a subsystem can watch simultaneously. To resolve this, you must explicitly increase your `inotify` system limits.
|
||||
|
||||
Below are example commands an admin user can run to increase system limits for `inotify` user instances and watches:
|
||||
|
||||
```shell
|
||||
sysctl -w fs.inotify.max_user_instances=8192
|
||||
sysctl -w fs.inotify.max_user_watches=524288
|
||||
```
|
||||
|
||||
### 日志缓冲区导致 Pod 过载
|
||||
|
||||
根据你的配置,默认缓冲区大小可能太大并导致 Pod 故障。减少负载的一种方法是降低记录器的刷新间隔。这可以防止日志溢出缓冲区。你还可以添加更多刷新线程来处理大量日志试图同时填充缓冲区的情况。
|
||||
|
||||
+1
@@ -20,6 +20,7 @@ Rancher 将 Rancher-Webhook 作为单独的 deployment 和服务部署在 local
|
||||
|
||||
| Rancher Version | Webhook Version | Availability in Prime | Availability in Community |
|
||||
|-----------------|-----------------|-----------------------|---------------------------|
|
||||
| v2.11.14 | v0.7.9 | ✓ | ✗ |
|
||||
| v2.11.13 | v0.7.8 | ✓ | ✗ |
|
||||
| v2.11.12 | v0.7.8 | ✓ | ✗ |
|
||||
| v2.11.11 | v0.7.8 | ✓ | ✗ |
|
||||
|
||||
@@ -12,6 +12,7 @@ Rancher 将在 GitHub 上发布的 Rancher 的[发版说明](https://github.com/
|
||||
|
||||
| Patch 版本 | 发布时间 |
|
||||
| ----------------------------------------------------------------- | ------------------ |
|
||||
| [2.12.10](https://github.com/rancher/rancher/releases/tag/v2.12.10) | 2026 年 05 月 27 日 |
|
||||
| [2.12.9](https://github.com/rancher/rancher/releases/tag/v2.12.9) | 2026 年 04 月 30 日 |
|
||||
| [2.12.8](https://github.com/rancher/rancher/releases/tag/v2.12.8) | 2026 年 03 月 25 日 |
|
||||
| [2.12.7](https://github.com/rancher/rancher/releases/tag/v2.12.7) | 2026 年 02 月 25 日 |
|
||||
|
||||
+1
@@ -15,6 +15,7 @@ title: 安装 Adapter
|
||||
|
||||
| Rancher 版本 | Adapter 版本 |
|
||||
|-----------------|:----------------:|
|
||||
| v2.12.10 | 107.0.0+up7.0.0 |
|
||||
| v2.12.9 | 107.0.0+up7.0.0 |
|
||||
| v2.12.8 | 107.0.0+up7.0.0 |
|
||||
| v2.12.7 | 107.0.0+up7.0.0 |
|
||||
|
||||
+13
@@ -79,6 +79,19 @@ Rancher Logging 有两个角色,分别是 `logging-admin` 和 `logging-view`
|
||||
|
||||
## 故障排除
|
||||
|
||||
### Resource Exhaustion of `inotify` Watchers and File Descriptors
|
||||
|
||||
When enabling the **Logging** app on Linux systems that heavily monitor the filesystem, you may encounter `Too many open files` or `CrashLoopBackOff` failures related to applications that leverage `inotify` to watch for file changes.
|
||||
|
||||
This happens because the Linux kernel caps the number of files a user can open and the number of directory paths a subsystem can watch simultaneously. To resolve this, you must explicitly increase your `inotify` system limits.
|
||||
|
||||
Below are example commands an admin user can run to increase system limits for `inotify` user instances and watches:
|
||||
|
||||
```shell
|
||||
sysctl -w fs.inotify.max_user_instances=8192
|
||||
sysctl -w fs.inotify.max_user_watches=524288
|
||||
```
|
||||
|
||||
### 日志缓冲区导致 Pod 过载
|
||||
|
||||
根据你的配置,默认缓冲区大小可能太大并导致 Pod 故障。减少负载的一种方法是降低记录器的刷新间隔。这可以防止日志溢出缓冲区。你还可以添加更多刷新线程来处理大量日志试图同时填充缓冲区的情况。
|
||||
|
||||
+1
@@ -20,6 +20,7 @@ Rancher 将 Rancher-Webhook 作为单独的 deployment 和服务部署在 local
|
||||
|
||||
| Rancher Version | Webhook Version | Availability in Prime | Availability in Community |
|
||||
|-----------------|-----------------|-----------------------|---------------------------|
|
||||
| v2.12.10 | v0.8.6 | ✓ | ✗ |
|
||||
| v2.12.9 | v0.8.5 | ✓ | ✗ |
|
||||
| v2.12.8 | v0.8.5 | ✓ | ✗ |
|
||||
| v2.12.7 | v0.8.5 | ✓ | ✗ |
|
||||
|
||||
@@ -16,6 +16,7 @@ Rancher 将在 GitHub 上发布的 Rancher 的[发版说明](https://github.com/
|
||||
|
||||
| Patch 版本 | 发布时间 |
|
||||
| ----------------------------------------------------------------- | ------------------ |
|
||||
| [2.13.6](https://github.com/rancher/rancher/releases/tag/v2.13.6) | 2026 年 05 月 27 日 |
|
||||
| [2.13.5](https://github.com/rancher/rancher/releases/tag/v2.13.5) | 2026 年 04 月 30 日 |
|
||||
| [2.13.4](https://github.com/rancher/rancher/releases/tag/v2.13.4) | 2026 年 03 月 25 日 |
|
||||
| [2.13.3](https://github.com/rancher/rancher/releases/tag/v2.13.3) | 2026 年 02 月 25 日 |
|
||||
|
||||
+1
@@ -15,6 +15,7 @@ title: 安装 Adapter
|
||||
|
||||
| Rancher 版本 | Adapter 版本 |
|
||||
|-----------------|------------------|
|
||||
| v2.13.6 | 108.0.0+up8.0.0 |
|
||||
| v2.13.5 | 108.0.0+up8.0.0 |
|
||||
| v2.13.4 | 108.0.0+up8.0.0 |
|
||||
| v2.13.3 | 108.0.0+up8.0.0 |
|
||||
|
||||
+13
@@ -79,6 +79,19 @@ Rancher Logging 有两个角色,分别是 `logging-admin` 和 `logging-view`
|
||||
|
||||
## 故障排除
|
||||
|
||||
### Resource Exhaustion of `inotify` Watchers and File Descriptors
|
||||
|
||||
When enabling the **Logging** app on Linux systems that heavily monitor the filesystem, you may encounter `Too many open files` or `CrashLoopBackOff` failures related to applications that leverage `inotify` to watch for file changes.
|
||||
|
||||
This happens because the Linux kernel caps the number of files a user can open and the number of directory paths a subsystem can watch simultaneously. To resolve this, you must explicitly increase your `inotify` system limits.
|
||||
|
||||
Below are example commands an admin user can run to increase system limits for `inotify` user instances and watches:
|
||||
|
||||
```shell
|
||||
sysctl -w fs.inotify.max_user_instances=8192
|
||||
sysctl -w fs.inotify.max_user_watches=524288
|
||||
```
|
||||
|
||||
### 日志缓冲区导致 Pod 过载
|
||||
|
||||
根据你的配置,默认缓冲区大小可能太大并导致 Pod 故障。减少负载的一种方法是降低记录器的刷新间隔。这可以防止日志溢出缓冲区。你还可以添加更多刷新线程来处理大量日志试图同时填充缓冲区的情况。
|
||||
|
||||
+2
-1
@@ -20,7 +20,8 @@ Rancher 将 Rancher-Webhook 作为单独的 deployment 和服务部署在 local
|
||||
|
||||
| Rancher Version | Webhook Version | Availability in Prime | Availability in Community |
|
||||
|-----------------|-----------------|-----------------------|---------------------------|
|
||||
| v2.13.5 | v0.9.3 | ✓ | ✗ |
|
||||
| v2.13.6 | v0.9.5 | ✓ | ✗ |
|
||||
| v2.13.5 | v0.9.4 | ✓ | ✗ |
|
||||
| v2.13.4 | v0.9.3 | ✓ | ✗ |
|
||||
| v2.13.3 | v0.9.3 | ✓ | ✓ |
|
||||
| v2.13.2 | v0.9.2 | ✓ | ✓ |
|
||||
|
||||
@@ -12,6 +12,7 @@ Rancher 将在 GitHub 上发布的 Rancher 的[发版说明](https://github.com/
|
||||
|
||||
| Patch 版本 | 发布时间 |
|
||||
| ----------------------------------------------------------------- | ------------------ |
|
||||
| [2.14.2](https://github.com/rancher/rancher/releases/tag/v2.14.2) | 2026 年 05 月 28 日 |
|
||||
| [2.14.1](https://github.com/rancher/rancher/releases/tag/v2.14.1) | 2026 年 04 月 30 日 |
|
||||
| [2.14.0](https://github.com/rancher/rancher/releases/tag/v2.14.0) | 2026 年 03 月 25 日 |
|
||||
|
||||
|
||||
+1
@@ -15,6 +15,7 @@ title: 安装 Adapter
|
||||
|
||||
| Rancher 版本 | Adapter 版本 |
|
||||
|-----------------|------------------|
|
||||
| v2.14.2 | 109.0.0+up9.0.0 |
|
||||
| v2.14.1 | 109.0.0+up9.0.0 |
|
||||
| v2.14.0 | 109.0.0+up9.0.0 |
|
||||
|
||||
|
||||
+13
@@ -79,6 +79,19 @@ Rancher Logging 有两个角色,分别是 `logging-admin` 和 `logging-view`
|
||||
|
||||
## 故障排除
|
||||
|
||||
### Resource Exhaustion of `inotify` Watchers and File Descriptors
|
||||
|
||||
When enabling the **Logging** app on Linux systems that heavily monitor the filesystem, you may encounter `Too many open files` or `CrashLoopBackOff` failures related to applications that leverage `inotify` to watch for file changes.
|
||||
|
||||
This happens because the Linux kernel caps the number of files a user can open and the number of directory paths a subsystem can watch simultaneously. To resolve this, you must explicitly increase your `inotify` system limits.
|
||||
|
||||
Below are example commands an admin user can run to increase system limits for `inotify` user instances and watches:
|
||||
|
||||
```shell
|
||||
sysctl -w fs.inotify.max_user_instances=8192
|
||||
sysctl -w fs.inotify.max_user_watches=524288
|
||||
```
|
||||
|
||||
### 日志缓冲区导致 Pod 过载
|
||||
|
||||
根据你的配置,默认缓冲区大小可能太大并导致 Pod 故障。减少负载的一种方法是降低记录器的刷新间隔。这可以防止日志溢出缓冲区。你还可以添加更多刷新线程来处理大量日志试图同时填充缓冲区的情况。
|
||||
|
||||
+2
-1
@@ -20,7 +20,8 @@ Rancher 将 Rancher-Webhook 作为单独的 deployment 和服务部署在 local
|
||||
|
||||
| Rancher Version | Webhook Version | Availability in Prime | Availability in Community |
|
||||
|-----------------|-----------------|-----------------------|---------------------------|
|
||||
| v2.14.1 | v0.10.1 | ✓ | ✓ |
|
||||
| v2.14.2 | v0.10.5 | ✓ | ✓ |
|
||||
| v2.14.1 | v0.10.4 | ✓ | ✓ |
|
||||
| v2.14.0 | v0.10.0 | ✗ | ✓ |
|
||||
|
||||
## 为什么我们需要它?
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
<!-- releaseTask -->
|
||||
The following table summarizes different GitHub metrics to give you an idea of each project's popularity and activity levels. This data was collected in December 2025.
|
||||
The following table summarizes different GitHub metrics to give you an idea of each project's popularity and activity levels. This data was collected in May 2026.
|
||||
|
||||
| Provider | Project | Stars | Forks | Contributors |
|
||||
| ---- | ---- | ---- | ---- | ---- |
|
||||
| Canal | https://github.com/projectcalico/canal | 722 | 97 | 20 |
|
||||
| Flannel | https://github.com/flannel-io/flannel | 9.4k | 2.9k | 248 |
|
||||
| Canal | https://github.com/projectcalico/canal | 723 | 97 | 20 |
|
||||
| Flannel | https://github.com/flannel-io/flannel | 9.5k | 2.9k | 249 |
|
||||
| Calico | https://github.com/projectcalico/calico | 7.2k | 1.6k | 412 |
|
||||
| Weave | https://github.com/weaveworks/weave | 6.6k | 675 | 82 |
|
||||
| Cilium | https://github.com/cilium/cilium | 24.2k | 3.7k | 1067 |
|
||||
| Cilium | https://github.com/cilium/cilium | 24.4k | 3.8k | 1085 |
|
||||
|
||||
+51
-11
@@ -17,9 +17,9 @@ Here you can find links to supporting documentation for the current released ver
|
||||
<th>Community</th>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>v2.14.1</b></td>
|
||||
<td><b>v2.14.2</b></td>
|
||||
<td><a href="https://ranchermanager.docs.rancher.com/v2.14">Documentation</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.14.1">Release Notes</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.14.2">Release Notes</a></td>
|
||||
<td><center>N/A</center></td>
|
||||
<td><center>✓</center></td>
|
||||
<td><center>✓</center></td>
|
||||
@@ -38,9 +38,9 @@ Here you can find links to supporting documentation for the current released ver
|
||||
<th>Community</th>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>v2.13.5</b></td>
|
||||
<td><b>v2.13.6</b></td>
|
||||
<td><a href="https://ranchermanager.docs.rancher.com/v2.13">Documentation</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.13.5">Release Notes</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.13.6">Release Notes</a></td>
|
||||
<td><center>N/A</center></td>
|
||||
<td><center>✓</center></td>
|
||||
<td><center>N/A</center></td>
|
||||
@@ -59,9 +59,9 @@ Here you can find links to supporting documentation for the current released ver
|
||||
<th>Community</th>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>v2.12.9</b></td>
|
||||
<td><b>v2.12.10</b></td>
|
||||
<td><a href="https://ranchermanager.docs.rancher.com/v2.12">Documentation</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.12.9">Release Notes</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.12.10">Release Notes</a></td>
|
||||
<td><center>N/A</center></td>
|
||||
<td><center>✓</center></td>
|
||||
<td><center>N/A</center></td>
|
||||
@@ -80,9 +80,9 @@ Here you can find links to supporting documentation for the current released ver
|
||||
<th>Community</th>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>v2.11.13</b></td>
|
||||
<td><b>v2.11.14</b></td>
|
||||
<td><a href="https://ranchermanager.docs.rancher.com/v2.11">Documentation</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.11.13">Release Notes</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.11.14">Release Notes</a></td>
|
||||
<td><center>N/A</center></td>
|
||||
<td><center>✓</center></td>
|
||||
<td><center>NA</center></td>
|
||||
@@ -101,9 +101,9 @@ Here you can find links to supporting documentation for the current released ver
|
||||
<th>Community</th>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>v2.10.11</b></td>
|
||||
<td><b>v2.10.12</b></td>
|
||||
<td><a href="https://ranchermanager.docs.rancher.com/v2.10">Documentation</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.10.11">Release Notes</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.10.12">Release Notes</a></td>
|
||||
<td><center>N/A</center></td>
|
||||
<td><center>✓</center></td>
|
||||
<td><center>N/A</center></td>
|
||||
@@ -123,6 +123,14 @@ Here you can find links to supporting documentation for previous versions of Ran
|
||||
<th>Prime</th>
|
||||
<th>Community</th>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>v2.14.1</b></td>
|
||||
<td><a href="https://ranchermanager.docs.rancher.com/v2.14">Documentation</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.14.1">Release Notes</a></td>
|
||||
<td><center><a href="https://www.suse.com/suse-rancher/support-matrix/all-supported-versions/rancher-v2-14-1/">Support Matrix</a></center></td>
|
||||
<td><center>✓</center></td>
|
||||
<td><center>✓</center></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>v2.14.0</b></td>
|
||||
<td><a href="https://ranchermanager.docs.rancher.com/v2.14">Documentation</a></td>
|
||||
@@ -144,6 +152,14 @@ Here you can find links to supporting documentation for previous versions of Ran
|
||||
<th>Prime</th>
|
||||
<th>Community</th>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>v2.13.5</b></td>
|
||||
<td><a href="https://ranchermanager.docs.rancher.com/v2.13">Documentation</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.13.5">Release Notes</a></td>
|
||||
<td><center><a href="https://www.suse.com/suse-rancher/support-matrix/all-supported-versions/rancher-v2-13-5/">Support Matrix</a></center></td>
|
||||
<td><center>✓</center></td>
|
||||
<td><center>N/A</center></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>v2.13.4</b></td>
|
||||
<td><a href="https://ranchermanager.docs.rancher.com/v2.13">Documentation</a></td>
|
||||
@@ -197,6 +213,14 @@ Here you can find links to supporting documentation for previous versions of Ran
|
||||
<th>Prime</th>
|
||||
<th>Community</th>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>v2.12.9</b></td>
|
||||
<td><a href="https://ranchermanager.docs.rancher.com/v2.12">Documentation</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.12.9">Release Notes</a></td>
|
||||
<td><center><a href="https://www.suse.com/suse-rancher/support-matrix/all-supported-versions/rancher-v2-12-9/">Support Matrix</a></center></td>
|
||||
<td><center>✓</center></td>
|
||||
<td><center>N/A</center></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>v2.12.8</b></td>
|
||||
<td><a href="https://ranchermanager.docs.rancher.com/v2.12">Documentation</a></td>
|
||||
@@ -282,6 +306,14 @@ Here you can find links to supporting documentation for previous versions of Ran
|
||||
<th>Prime</th>
|
||||
<th>Community</th>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>v2.11.13</b></td>
|
||||
<td><a href="https://ranchermanager.docs.rancher.com/v2.11">Documentation</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.11.13">Release Notes</a></td>
|
||||
<td><center><a href="https://www.suse.com/suse-rancher/support-matrix/all-supported-versions/rancher-v2-11-13/">Support Matrix</a></center></td>
|
||||
<td><center>✓</center></td>
|
||||
<td><center>NA</center></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>v2.11.12</b></td>
|
||||
<td><a href="https://ranchermanager.docs.rancher.com/v2.11">Documentation</a></td>
|
||||
@@ -399,7 +431,15 @@ Here you can find links to supporting documentation for previous versions of Ran
|
||||
<th>Prime</th>
|
||||
<th>Community</th>
|
||||
</tr>
|
||||
<tr>
|
||||
<tr>
|
||||
<td><b>v2.10.11</b></td>
|
||||
<td><a href="https://ranchermanager.docs.rancher.com/v2.10">Documentation</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.10.11">Release Notes</a></td>
|
||||
<td><center><a href="https://www.suse.com/suse-rancher/support-matrix/all-supported-versions/rancher-v2-10-11/">Support Matrix</a></center></td>
|
||||
<td><center>✓</center></td>
|
||||
<td><center>N/A</center></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><b>v2.10.10</b></td>
|
||||
<td><a href="https://ranchermanager.docs.rancher.com/v2.10">Documentation</a></td>
|
||||
<td><a href="https://github.com/rancher/rancher/releases/tag/v2.10.10">Release Notes</a></td>
|
||||
|
||||
@@ -12,6 +12,7 @@ Rancher will publish deprecated features as part of the [release notes](https://
|
||||
|
||||
| Patch Version | Release Date |
|
||||
|---------------|---------------|
|
||||
| [2.10.12](https://github.com/rancher/rancher/releases/tag/v2.10.12) | May 27, 2026 |
|
||||
| [2.10.11](https://github.com/rancher/rancher/releases/tag/v2.10.11) | January 29, 2026 |
|
||||
| [2.10.10](https://github.com/rancher/rancher/releases/tag/v2.10.10) | September 25, 2025 |
|
||||
| [2.10.9](https://github.com/rancher/rancher/releases/tag/v2.10.9) | August 27, 2025 |
|
||||
|
||||
+34
-9
@@ -29,6 +29,31 @@ Review the list of known issues for each Rancher version, which can be found in
|
||||
|
||||
Note that upgrades _to_ or _from_ any chart in the [rancher-alpha repository](../resources/choose-a-rancher-version.md#helm-chart-repositories) aren't supported.
|
||||
|
||||
### Upgrade Path
|
||||
|
||||
:::important
|
||||
|
||||
**Important:** The only tested and supported Rancher upgrade path between minor versions (e.g. v2.9.x to v2.10.x) is to upgrade from the latest available patch version of your current running minor release to the latest available patch version of the next minor release.
|
||||
|
||||
Before initiating a minor version upgrade, verify that you are running the most recent patch release of your current version.
|
||||
|
||||
You can query the available chart versions with the Helm CLI:
|
||||
|
||||
1. Update your local Helm repo cache.
|
||||
|
||||
```
|
||||
helm repo update
|
||||
```
|
||||
|
||||
1. Search for available versions in your [specific repository](../resources/choose-a-rancher-version.md#helm-chart-repositories) (e.g., rancher-stable):
|
||||
|
||||
```
|
||||
helm search repo rancher-<CHART_REPO>/rancher --versions
|
||||
```
|
||||
|
||||
If your installation is not on the latest patch version of the current minor release, you must upgrade to that version before proceeding to the next minor version.
|
||||
:::
|
||||
|
||||
### Helm Version
|
||||
|
||||
The upgrade instructions assume you are using Helm 3.
|
||||
@@ -141,17 +166,17 @@ There will be more values that are listed with this command. This is just an exa
|
||||
|
||||
:::
|
||||
|
||||
:::tip
|
||||
:::tip
|
||||
|
||||
Your deployment name may vary; for example, if you're deploying Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
Your deployment name may vary; for example, if you're deploying Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
```
|
||||
helm get values rancher-stable -n cattle-system
|
||||
|
||||
hostname: rancher.my.org
|
||||
```
|
||||
|
||||
:::
|
||||
:::
|
||||
|
||||
If you are upgrading cert-manager to the latest version from v1.5 or below, follow the [cert-manager upgrade docs](../resources/upgrade-cert-manager.md#option-c-upgrade-cert-manager-from-versions-15-and-below) to learn how to upgrade cert-manager without needing to perform an uninstall or reinstall of Rancher. Otherwise, follow the [steps to upgrade Rancher](#steps-to-upgrade-rancher) below.
|
||||
|
||||
@@ -174,17 +199,17 @@ The above is an example, there may be more values from the previous step that ne
|
||||
|
||||
:::
|
||||
|
||||
:::tip
|
||||
:::tip
|
||||
|
||||
If you deploy Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
If you deploy Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
```
|
||||
helm upgrade rancher-stable rancher-<CHART_REPO>/rancher \
|
||||
--namespace cattle-system \
|
||||
--set hostname=rancher.my.org
|
||||
```
|
||||
|
||||
:::
|
||||
:::
|
||||
|
||||
Alternatively, it's possible to export the current values to a file and reference that file during upgrade. For example, to only change the Rancher version:
|
||||
|
||||
@@ -194,7 +219,7 @@ Alternatively, it's possible to export the current values to a file and referenc
|
||||
```
|
||||
1. Update only the Rancher version:
|
||||
|
||||
|
||||
|
||||
```
|
||||
helm upgrade rancher rancher-<CHART_REPO>/rancher \
|
||||
--namespace cattle-system \
|
||||
|
||||
+2
-2
@@ -33,7 +33,7 @@ To use a premade dashboard, go to [https://grafana.com/grafana/dashboards](https
|
||||
To use your own dashboard:
|
||||
|
||||
1. Click on the link to open Grafana. On the cluster detail page, click **Monitoring**.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin/prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
:::note
|
||||
|
||||
@@ -113,7 +113,7 @@ Note that the RBAC roles exposed by the Monitoring chart to add Grafana Dashboar
|
||||
1. On the **Clusters** page, go to the cluster where you want to configure the Grafana namespace and click **Explore**.
|
||||
1. In the left navigation bar, click **Monitoring**.
|
||||
1. Click **Grafana**.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin/prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
:::note
|
||||
|
||||
|
||||
+1
-1
@@ -20,7 +20,7 @@ To see the links to the external monitoring UIs, including Grafana dashboards, y
|
||||
1. In the left navigation menu, click **Monitoring.**
|
||||
1. Click **Grafana.** The Grafana dashboard should open in a new tab.
|
||||
1. Go to the log in icon in the lower left corner and click **Sign In.**
|
||||
1. Log in to Grafana. The default Admin username and password for the Grafana instance is `admin/prom-operator`. (Regardless of who has the password, cluster administrator permission in Rancher is still required access the Grafana instance.) Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. (Regardless of who has the password, cluster administrator permission in Rancher is still required access the Grafana instance.) Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
|
||||
### Getting the PromQL Query Powering a Grafana Panel
|
||||
|
||||
+8
-2
@@ -31,11 +31,17 @@ _Cluster roles_ are roles that you can assign to users, granting them access to
|
||||
|
||||
- **Cluster Owner:**
|
||||
|
||||
These users have full control over the cluster and all resources in it.
|
||||
These users have full control over the cluster and all resources in it.
|
||||
|
||||
- **Cluster Member:**
|
||||
|
||||
These users can view most cluster level resources and create new projects.
|
||||
These users can view most cluster level resources and create new projects.
|
||||
|
||||
:::warning
|
||||
|
||||
When a Cluster Member creates a project, the user is automatically assigned [Project Owner privileges](#project-roles). This grants them comprehensive control over the project and its associated resources, including permissions to deploy workloads. Without enforced [Pod Security Standards (PSS) and Pod Security Admission (PSA)](../pod-security-standards.md), a Cluster Member is able to execute privileged containers in the cluster.
|
||||
|
||||
:::
|
||||
|
||||
#### Custom Cluster Roles
|
||||
|
||||
|
||||
+1
-1
@@ -234,7 +234,7 @@ docker stop <original-rancher-container>
|
||||
|
||||
:::note
|
||||
|
||||
If you wish to keep the original Rancher environment running, you can also restart the cattle-cluster-agent pods on each cluster connected to your Rancher environment.
|
||||
If clusters do not automatically reconnect to the new environment after you have redirected traffic, for example if there is a delay in scaling down the original Rancher instance, you can also restart the cattle-cluster-agent pods on each cluster connected to your Rancher environment.
|
||||
|
||||
```bash
|
||||
kubectl rollout restart deployment cattle-cluster-agent -n cattle-system
|
||||
|
||||
+1
-1
@@ -24,7 +24,7 @@ The following steps can also be performed using the `kubectl` command line tool.
|
||||
:::
|
||||
|
||||
1. Click **☰ > Cluster Management**.
|
||||
1. Choose the cluster you want to provide vSphere storage to and click **Exlpore**.
|
||||
1. Choose the cluster you want to provide vSphere storage to and click **Explore**.
|
||||
1. In the left navigation bar, select **Storage > StorageClasses**.
|
||||
1. Click **Create**.
|
||||
3. Enter a **Name** for the StorageClass.
|
||||
|
||||
+1
@@ -19,6 +19,7 @@ In order to deploy and run the adapter successfully, you need to ensure its vers
|
||||
|
||||
| Rancher Version | Adapter Version |
|
||||
|-----------------|------------------|
|
||||
| v2.10.12 | v105.0.0+up5.0.1 |
|
||||
| v2.10.11 | v105.0.0+up5.0.1 |
|
||||
| v2.10.10 | v105.0.0+up5.0.1 |
|
||||
| v2.10.9 | v105.0.0+up5.0.1 |
|
||||
|
||||
@@ -116,6 +116,19 @@ By default, Rancher collects logs for control plane components and node componen
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Resource Exhaustion of `inotify` Watchers and File Descriptors
|
||||
|
||||
When enabling the **Logging** app on Linux systems that heavily monitor the filesystem, you may encounter `Too many open files` or `CrashLoopBackOff` failures related to applications that leverage `inotify` to watch for file changes.
|
||||
|
||||
This happens because the Linux kernel caps the number of files a user can open and the number of directory paths a subsystem can watch simultaneously. To resolve this, you must explicitly increase your `inotify` system limits.
|
||||
|
||||
Below are example commands an admin user can run to increase system limits for `inotify` user instances and watches:
|
||||
|
||||
```shell
|
||||
sysctl -w fs.inotify.max_user_instances=8192
|
||||
sysctl -w fs.inotify.max_user_watches=524288
|
||||
```
|
||||
|
||||
### The Logging Buffer Overloads Pods
|
||||
|
||||
Depending on your configuration, the default buffer size may be too large and cause pod failures. One way to reduce the load is to lower the logger's flush interval. This prevents logs from overfilling the buffer. You can also add more flush threads to handle moments when many logs are attempting to fill the buffer at once.
|
||||
|
||||
@@ -14,64 +14,74 @@ Examples of built-in Rancher extensions are Fleet, Explorer, and Harvester. Exam
|
||||
|
||||
## Prerequisites
|
||||
|
||||
> You must log in as an admin in order to view and interact with the extensions management page.
|
||||
> 1. You must log in as an administrator to view and interact with the extensions management page.
|
||||
> 2. You must enable extension support.
|
||||
|
||||
## Enabling Extension Support in Rancher
|
||||
|
||||
Rancher v2.9.0 and later includes extension support.
|
||||
|
||||
You can confirm if extension support is enabled by checking the `uiextension` feature flag. Any changes to this feature flag cause the Rancher pod to restart.
|
||||
|
||||
When you enable extension support for the first time, it creates resources, such as the CRDs, so that UI Extensions can work. When extension support is disabled, it disables the endpoints and does not cache any files. However, it does not remove any CRs or delete any extensions that were installed before. If re-enabled, it exposes the required endpoints again and creates CRDs as needed. The extensions that were already installed load after the Rancher pod restarts.
|
||||
|
||||
### Caching Extension Files
|
||||
|
||||
By default, Rancher caches every extension file in the file system. You can change that behavior by setting `plugin.noCache` to `true`.
|
||||
|
||||
Rancher does have a cached file size limit of 30MB. If an extension has a file bigger than that, the cache is disabled and `plugin.noCache` is set to `true`, regardless of user input.
|
||||
|
||||
|
||||
## Installing Extensions
|
||||
|
||||
1. Click **☰ > Extensions** under **Configuration**.
|
||||
|
||||
2. If not already installed in **Apps**, you must enable the extension operator by clicking the **Enable** button.
|
||||
2. On the **Extensions** page, select the **Available** tab to choose the extensions that you want to install.
|
||||
|
||||
- Click **OK** to add the Rancher extension repository if your installation is not air-gapped. Otherwise, uncheck the box to do so and click **OK**.
|
||||
3. If no extensions are listed as available, you can manually add the repos:
|
||||
|
||||

|
||||
3.1. On the upper right, click **⋮ > Manage Repositories > Create**.
|
||||
|
||||
3. On the **Extensions** page, click on the **Available** tab to select which extensions you want to install.
|
||||
3.2. Add the desired repo name, making sure to also specify the Git repo URL and the Git branch.
|
||||
|
||||
4. If no extensions are showing as available, you may manually add repos as follows:
|
||||
3.3. Click **Create** in the lower right again to complete.
|
||||
|
||||
4.1. On the upper right of screen, click on **⋮ > Manage Repositories > Create**.
|
||||

|
||||
|
||||
4.2. Add the desired repo name, making sure to also specify the Git Repo URL and the Git Branch.
|
||||
4. Under the **Available** tab, click **Install** on the desired extension and version, as in the example below. You can also update your extension from this screen. The button to **Update** appears on the extension card if an update is available.
|
||||
|
||||
4.3. Click **Create** in the lower right again to complete.
|
||||

|
||||
|
||||

|
||||
5. Click **Reload** after your extension successfully installs to check its status. Updates to the UI aren't visible until you reload the page.
|
||||
|
||||
5. Under the **Available** tab, click **Install** on the desired extension and version as in the example below. You can also update your extension from this screen, as the button to **Update** will appear on the extension if one is available.
|
||||
|
||||

|
||||
|
||||
6. Click the **Reload** page button that will appear after your extension successfully installs. Note that a logged-in user who has just installed an extension will not see a change to the UI **unless** they reload the page.
|
||||
|
||||

|
||||

|
||||
|
||||
## Updating and Upgrading Extensions
|
||||
|
||||
1. Click **☰ > Extensions** under **Configuration**.
|
||||
1. Select the **Updates** tab.
|
||||
1. Click **Update**.
|
||||
2. Select the **Updates** tab.
|
||||
3. Click **Update**.
|
||||
|
||||
If there is a new version of the extension, there will also be an **Update** button visible on the associated card for the extension in the **Available** tab.
|
||||
|
||||
## Deleting Extensions
|
||||
|
||||
1. Click **☰**, then click on the name of your local cluster.
|
||||
1. From the sidebar, select **Apps > Installed Apps**.
|
||||
1. Find the name of the chart you want to delete and select the checkbox next to it.
|
||||
1. Click **Delete**.
|
||||
2. From the sidebar, select **Apps > Installed Apps**.
|
||||
3. Find the name of the chart you want to delete and select the checkbox next to it.
|
||||
4. Click **Delete**.
|
||||
|
||||
## Deleting Extension Repositories
|
||||
|
||||
1. Click **☰ > Extensions** under **Configuration**.
|
||||
1. On the top right, click **⋮ > Manage Repositories**.
|
||||
1. Find the name of the extension repository you want to delete. Select the checkbox next to the repository name, then click **Delete**.
|
||||
2. On the top right, click **⋮ > Manage Repositories**.
|
||||
3. Find the name of the extension repository you want to delete. Select the checkbox next to the repository name, then click **Delete**.
|
||||
|
||||
## Deleting Extension Repository Container Images
|
||||
|
||||
1. Click **☰**, then select **Extensions**, under **Configuration**.
|
||||
1. On the top right, click **⋮ > Manage Extension Catalogs**.
|
||||
1. Find the name of the container image you want to delete, then click **⋮ > Uninstall**.
|
||||
2. On the top right, click **⋮ > Manage Extension Catalogs**.
|
||||
3. Find the name of the container image you want to delete, then click **⋮ > Uninstall**.
|
||||
|
||||
## Uninstalling Extensions
|
||||
|
||||
@@ -79,11 +89,11 @@ There are two ways to uninstall or disable an extension:
|
||||
|
||||
1. Under the **Installed** tab, click the **Uninstall** button on the extension you wish to remove.
|
||||
|
||||

|
||||

|
||||
|
||||
1. On the extensions management page, click **⋮ > Disable Extension Support**. This will disable all installed extensions.
|
||||
2. On the extensions management page, click **⋮ > Disable Extension Support**. This will disable all installed extensions.
|
||||
|
||||

|
||||

|
||||
|
||||
:::caution
|
||||
|
||||
@@ -91,6 +101,12 @@ You must reload the page after disabling extensions or display issues may occur.
|
||||
|
||||
:::
|
||||
|
||||
## Enabling Unauthenticated Access to an Extension
|
||||
|
||||
In Rancher v2.9.0 and later, you can allow unauthenticated access to an extension. You may want to enable unauthenticated access if the extension enables a new locale or adds custom branding. By default, all extensions require user authentication to load.
|
||||
|
||||
To enable unauthenticated access to an extension, set `plugin.noAuth` to `true` in the CR used by the extension.
|
||||
|
||||
## Developing Extensions
|
||||
|
||||
To learn how to develop your own extensions, refer to the official [Getting Started](https://rancher.github.io/dashboard/extensions/extensions-getting-started) guide.
|
||||
@@ -173,8 +189,8 @@ After you successfully set up these resources, you can install the extensions fr
|
||||
1. Click **☰**, then select **Extensions**, under **Configuration**.
|
||||
1. On the top right, click **⋮ > Manage Extension Catalogs**.
|
||||
1. Select the **Import Extension Catalog** button.
|
||||
1. Enter the image address in the **Catalog Image Reference** field.
|
||||
* **(Optional)** If the container image is private, select the secret you just created from the **Pull Secrets** drop-down menu.
|
||||
1. Enter the image address in the **Catalog Image Reference** field.
|
||||
- **(Optional)** If the container image is private, select the secret you just created from the **Pull Secrets** drop-down menu.
|
||||
1. Click **Load**. The extension will now be **Pending**.
|
||||
1. Return to the **Extensions** page.
|
||||
1. Select the **Available** tab, and click **Reload** to make sure that the list of extensions is up to date.
|
||||
@@ -191,7 +207,7 @@ After you mirror the latest changes, follow these steps:
|
||||
1. Click **☰ > Local**.
|
||||
1. From the sidebar, select **Workloads > Deployments**.
|
||||
1. From the namespaces dropdown menu, select **cattle-ui-plugin-system**.
|
||||
1. Find the **cattle-ui-plugin-system** namespace.
|
||||
1. Find the **cattle-ui-plugin-system** namespace.
|
||||
1. Select the `ui-plugin-catalog` deployment.
|
||||
1. Click **⋮ > Edit config**.
|
||||
1. Update the **Container Image** field within the deployment's container with the latest image.
|
||||
|
||||
@@ -20,6 +20,7 @@ Each Rancher version is designed to be compatible with a single version of the w
|
||||
|
||||
| Rancher Version | Webhook Version | Availability in Prime | Availability in Community |
|
||||
|-----------------|-----------------|-----------------------|---------------------------|
|
||||
| v2.10.12 | v0.6.12 | ✓ | ✗ |
|
||||
| v2.10.11 | v0.6.12 | ✓ | ✗ |
|
||||
| v2.10.10 | v0.6.11 | ✓ | ✗ |
|
||||
| v2.10.9 | v0.6.10 | ✓ | ✗ |
|
||||
|
||||
+85
-14
@@ -78,66 +78,137 @@ kubectl -n kube-system get endpoints kube-scheduler -o jsonpath='{.metadata.anno
|
||||
|
||||
## Ingress Controller
|
||||
|
||||
The default Ingress Controller is Traefik and is deployed as a DaemonSet in the `traefik` namespace. The pods are only scheduled to nodes with the `worker` role.
|
||||
The default Ingress Controller is Traefik and is deployed as a DaemonSet in the `kube-system` namespace. The pods are only scheduled to nodes with the `worker` role.
|
||||
|
||||
Check if the pods are running on all nodes:
|
||||
|
||||
```
|
||||
kubectl -n traefik get pods -o wide
|
||||
kubectl -n kube-system get pods -o wide
|
||||
```
|
||||
|
||||
Example output:
|
||||
Example RKE2 output:
|
||||
|
||||
```
|
||||
kubectl -n traefik get pods -o wide
|
||||
kubectl -n kube-system get pods -o wide
|
||||
NAME READY STATUS RESTARTS AGE IP NODE
|
||||
default-http-backend-797c5bc547-kwwlq 1/1 Running 0 17m x.x.x.x worker-1
|
||||
traefik-4qd64 1/1 Running 0 14m x.x.x.x worker-1
|
||||
traefik-8wxhm 1/1 Running 0 13m x.x.x.x worker-0
|
||||
local-path-provisioner-xxxxxxxxxx-xxxxx 1/1 Running 0 17m x.x.x.x worker-1
|
||||
rke2-traefik-xxxxxxxxxx-xxxxx 1/1 Running 0 14m x.x.x.x worker-1
|
||||
svclb-rke2-traefik-xxxxxxxx-xxxxx 1/1 Running 0 13m x.x.x.x worker-0
|
||||
...
|
||||
```
|
||||
|
||||
Example K3s output:
|
||||
|
||||
```
|
||||
kubectl -n kube-system get pods -o wide
|
||||
NAME READY STATUS RESTARTS AGE IP NODE
|
||||
local-path-provisioner-xxxxxxxxxx-xxxxx 1/1 Running 0 17m x.x.x.x worker-1
|
||||
traefik-xxxxxxxxxx-xxxxx 1/1 Running 0 14m x.x.x.x worker-1
|
||||
svclb-traefik-xxxxxxxx-xxxxx 1/1 Running 0 13m x.x.x.x worker-0
|
||||
...
|
||||
```
|
||||
|
||||
If a pod is unable to run (Status is not **Running**, Ready status is not showing `1/1` or you see a high count of Restarts), check the pod details, logs and namespace events.
|
||||
|
||||
### Pod details
|
||||
|
||||
RKE2 example:
|
||||
|
||||
```
|
||||
kubectl -n traefik describe pods -l app=traefik
|
||||
kubectl -n kube-system describe pods -l app.kubernetes.io/name=rke2-traefik
|
||||
```
|
||||
|
||||
K3s example:
|
||||
|
||||
```
|
||||
kubectl -n kube-system describe pods -l app.kubernetes.io/name=traefik
|
||||
```
|
||||
|
||||
### Pod container logs
|
||||
|
||||
The below command can show the logs of all the pods labeled "app=traefik", but it will display only 10 lines of log because of the restrictions of the `kubectl logs` command. Refer to `--tail` of `kubectl logs -h` for more information.
|
||||
The below command can show the logs of all the pods labeled "app.kubernetes.io/name=rke2-traefik" if using RKE2 or "app.kubernetes.io/name=traefik" if using K3s, but it will display only 10 lines of log because of the restrictions of the `kubectl logs` command. Refer to `--tail` of `kubectl logs -h` for more information.
|
||||
|
||||
RKE2 example:
|
||||
|
||||
```
|
||||
kubectl -n traefik logs -l app=traefik
|
||||
kubectl -n kube-system logs -l app.kubernetes.io/name=rke2-traefik
|
||||
```
|
||||
|
||||
K3s example:
|
||||
|
||||
```
|
||||
kubectl -n kube-system logs -l app.kubernetes.io/name=traefik
|
||||
```
|
||||
|
||||
If the full log is needed, specify the pod name in the trailing command:
|
||||
|
||||
```
|
||||
kubectl -n traefik logs <pod name>
|
||||
kubectl -n kube-system logs <pod name>
|
||||
```
|
||||
|
||||
### Namespace events
|
||||
|
||||
```
|
||||
kubectl -n traefik get events
|
||||
kubectl -n kube-system get events
|
||||
```
|
||||
|
||||
### Debug logging
|
||||
|
||||
To enable debug logging:
|
||||
|
||||
RKE2 example:
|
||||
|
||||
```
|
||||
kubectl -n traefik patch ds traefik --type='json' -p='[{"op": "add", "path": "/spec/template/spec/containers/0/args/-", "value": "--v=5"}]'
|
||||
cat <<EOF | kubectl apply -f -
|
||||
apiVersion: helm.cattle.io/v1
|
||||
kind: HelmChartConfig
|
||||
metadata:
|
||||
name: rke2-traefik
|
||||
namespace: kube-system
|
||||
spec:
|
||||
valuesContent: |-
|
||||
additionalArguments:
|
||||
- "--log.level=DEBUG"
|
||||
EOF
|
||||
```
|
||||
|
||||
K3s example:
|
||||
|
||||
```
|
||||
cat <<EOF | kubectl apply -f -
|
||||
apiVersion: helm.cattle.io/v1
|
||||
kind: HelmChartConfig
|
||||
metadata:
|
||||
name: traefik
|
||||
namespace: kube-system
|
||||
spec:
|
||||
valuesContent: |-
|
||||
logs:
|
||||
general:
|
||||
level: "DEBUG"
|
||||
EOF
|
||||
```
|
||||
|
||||
### Check configuration
|
||||
|
||||
Retrieve generated configuration in each pod:
|
||||
|
||||
RKE2 example for manual Traefik configuration file check:
|
||||
|
||||
```
|
||||
kubectl -n traefik get pods -l app=traefik --no-headers -o custom-columns=.NAME:.metadata.name | while read pod; do kubectl -n traefik exec $pod -- cat /etc/nginx/nginx.conf; done
|
||||
kubectl exec -n kube-system pod/traefik-xxxxxxxxx-xxxxx -- cat /var/lib/rancher/rke2/server/manifests/rke2-traefik-config.yaml
|
||||
```
|
||||
|
||||
K3s example for manual Traefik configuration file check:
|
||||
|
||||
```
|
||||
kubectl exec -n kube-system pod/traefik-xxxxxxxxx-xxxxx -- cat /var/lib/rancher/k3s/server/manifests/k3s-traefik-config.yaml
|
||||
```
|
||||
|
||||
RKE2/K3s example for Traefik CLI argument configuration check:
|
||||
|
||||
```
|
||||
kubectl get pod traefik-xxxxxxxxx-xxxxx -n kube-system -o jsonpath='{.spec.containers[0].args}'
|
||||
```
|
||||
|
||||
## Rancher agents
|
||||
|
||||
@@ -15,6 +15,14 @@ Make sure you configured the correct kubeconfig (for example, `export KUBECONFIG
|
||||
Double check if all the [required ports](../../how-to-guides/new-user-guides/kubernetes-clusters-in-rancher-setup/node-requirements-for-rancher-managed-clusters.md#networking-requirements) are opened in your (host) firewall. The overlay network uses UDP in comparison to all other required ports which are TCP.
|
||||
|
||||
|
||||
## Check if your downstream node can communicate to Rancher Manager
|
||||
|
||||
Rancher components with HTTP endpoints generally contain a `ping` liveness probe which you can use to test connectivity. Replace the `$RANCHER_URL` as appropriate and run the following from a node to check that it has connectivity to Rancher Manager's servers in the `local` cluster. If successful, it should return `pong`.
|
||||
|
||||
```
|
||||
curl -k https://$RANCHER_URL/ping
|
||||
```
|
||||
|
||||
## Check if Overlay Network is Functioning Correctly
|
||||
|
||||
The pod can be scheduled to any of the hosts you used for your cluster, but that means that the NGINX ingress controller needs to be able to route the request from `NODE_1` to `NODE_2`. This happens over the overlay network. If the overlay network is not functioning, you will experience intermittent TCP/HTTP connection failures due to the NGINX ingress controller not being able to route to the pod.
|
||||
|
||||
@@ -12,6 +12,7 @@ Rancher will publish deprecated features as part of the [release notes](https://
|
||||
|
||||
| Patch Version | Release Date |
|
||||
|---------------|---------------|
|
||||
| [2.11.14](https://github.com/rancher/rancher/releases/tag/v2.11.14) | May 27, 2026 |
|
||||
| [2.11.13](https://github.com/rancher/rancher/releases/tag/v2.11.13) | April 30, 2026 |
|
||||
| [2.11.12](https://github.com/rancher/rancher/releases/tag/v2.11.12) | March 25, 2026 |
|
||||
| [2.11.11](https://github.com/rancher/rancher/releases/tag/v2.11.11) | February 25, 2026 |
|
||||
|
||||
+34
-9
@@ -29,6 +29,31 @@ Review the list of known issues for each Rancher version, which can be found in
|
||||
|
||||
Note that upgrades _to_ or _from_ any chart in the [rancher-alpha repository](../resources/choose-a-rancher-version.md#helm-chart-repositories) aren't supported.
|
||||
|
||||
### Upgrade Path
|
||||
|
||||
:::important
|
||||
|
||||
**Important:** The only tested and supported Rancher upgrade path between minor versions (e.g. v2.10.x to v2.11.x) is to upgrade from the latest available patch version of your current running minor release to the latest available patch version of the next minor release.
|
||||
|
||||
Before initiating a minor version upgrade, verify that you are running the most recent patch release of your current version.
|
||||
|
||||
You can query the available chart versions with the Helm CLI:
|
||||
|
||||
1. Update your local Helm repo cache.
|
||||
|
||||
```
|
||||
helm repo update
|
||||
```
|
||||
|
||||
1. Search for available versions in your [specific repository](../resources/choose-a-rancher-version.md#helm-chart-repositories) (e.g., rancher-stable):
|
||||
|
||||
```
|
||||
helm search repo rancher-<CHART_REPO>/rancher --versions
|
||||
```
|
||||
|
||||
If your installation is not on the latest patch version of the current minor release, you must upgrade to that version before proceeding to the next minor version.
|
||||
:::
|
||||
|
||||
### Helm Version
|
||||
|
||||
The upgrade instructions assume you are using Helm 3.
|
||||
@@ -141,17 +166,17 @@ There will be more values that are listed with this command. This is just an exa
|
||||
|
||||
:::
|
||||
|
||||
:::tip
|
||||
:::tip
|
||||
|
||||
Your deployment name may vary; for example, if you're deploying Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
Your deployment name may vary; for example, if you're deploying Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
```
|
||||
helm get values rancher-stable -n cattle-system
|
||||
|
||||
hostname: rancher.my.org
|
||||
```
|
||||
|
||||
:::
|
||||
:::
|
||||
|
||||
If you are upgrading cert-manager to the latest version from v1.5 or below, follow the [cert-manager upgrade docs](../resources/upgrade-cert-manager.md#option-c-upgrade-cert-manager-from-versions-15-and-below) to learn how to upgrade cert-manager without needing to perform an uninstall or reinstall of Rancher. Otherwise, follow the [steps to upgrade Rancher](#steps-to-upgrade-rancher) below.
|
||||
|
||||
@@ -174,17 +199,17 @@ The above is an example, there may be more values from the previous step that ne
|
||||
|
||||
:::
|
||||
|
||||
:::tip
|
||||
:::tip
|
||||
|
||||
If you deploy Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
If you deploy Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
```
|
||||
helm upgrade rancher-stable rancher-<CHART_REPO>/rancher \
|
||||
--namespace cattle-system \
|
||||
--set hostname=rancher.my.org
|
||||
```
|
||||
|
||||
:::
|
||||
:::
|
||||
|
||||
Alternatively, it's possible to export the current values to a file and reference that file during upgrade. For example, to only change the Rancher version:
|
||||
|
||||
@@ -194,7 +219,7 @@ Alternatively, it's possible to export the current values to a file and referenc
|
||||
```
|
||||
1. Update only the Rancher version:
|
||||
|
||||
|
||||
|
||||
```
|
||||
helm upgrade rancher rancher-<CHART_REPO>/rancher \
|
||||
--namespace cattle-system \
|
||||
|
||||
+2
-2
@@ -33,7 +33,7 @@ To use a premade dashboard, go to [https://grafana.com/grafana/dashboards](https
|
||||
To use your own dashboard:
|
||||
|
||||
1. Click on the link to open Grafana. On the cluster detail page, click **Monitoring**.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin/prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
:::note
|
||||
|
||||
@@ -113,7 +113,7 @@ Note that the RBAC roles exposed by the Monitoring chart to add Grafana Dashboar
|
||||
1. On the **Clusters** page, go to the cluster where you want to configure the Grafana namespace and click **Explore**.
|
||||
1. In the left navigation bar, click **Monitoring**.
|
||||
1. Click **Grafana**.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin/prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
:::note
|
||||
|
||||
|
||||
+1
-1
@@ -20,7 +20,7 @@ To see the links to the external monitoring UIs, including Grafana dashboards, y
|
||||
1. In the left navigation menu, click **Monitoring.**
|
||||
1. Click **Grafana.** The Grafana dashboard should open in a new tab.
|
||||
1. Go to the log in icon in the lower left corner and click **Sign In.**
|
||||
1. Log in to Grafana. The default Admin username and password for the Grafana instance is `admin/prom-operator`. (Regardless of who has the password, cluster administrator permission in Rancher is still required access the Grafana instance.) Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. (Regardless of who has the password, cluster administrator permission in Rancher is still required access the Grafana instance.) Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
|
||||
### Getting the PromQL Query Powering a Grafana Panel
|
||||
|
||||
+8
-2
@@ -31,11 +31,17 @@ _Cluster roles_ are roles that you can assign to users, granting them access to
|
||||
|
||||
- **Cluster Owner:**
|
||||
|
||||
These users have full control over the cluster and all resources in it.
|
||||
These users have full control over the cluster and all resources in it.
|
||||
|
||||
- **Cluster Member:**
|
||||
|
||||
These users can view most cluster level resources and create new projects.
|
||||
These users can view most cluster level resources and create new projects.
|
||||
|
||||
:::warning
|
||||
|
||||
When a Cluster Member creates a project, the user is automatically assigned [Project Owner privileges](#project-roles). This grants them comprehensive control over the project and its associated resources, including permissions to deploy workloads. Without enforced [Pod Security Standards (PSS) and Pod Security Admission (PSA)](../pod-security-standards.md), a Cluster Member is able to execute privileged containers in the cluster.
|
||||
|
||||
:::
|
||||
|
||||
#### Custom Cluster Roles
|
||||
|
||||
|
||||
+1
-1
@@ -234,7 +234,7 @@ docker stop <original-rancher-container>
|
||||
|
||||
:::note
|
||||
|
||||
If you wish to keep the original Rancher environment running, you can also restart the cattle-cluster-agent pods on each cluster connected to your Rancher environment.
|
||||
If clusters do not automatically reconnect to the new environment after you have redirected traffic, for example if there is a delay in scaling down the original Rancher instance, you can also restart the cattle-cluster-agent pods on each cluster connected to your Rancher environment.
|
||||
|
||||
```bash
|
||||
kubectl rollout restart deployment cattle-cluster-agent -n cattle-system
|
||||
|
||||
+1
-1
@@ -24,7 +24,7 @@ The following steps can also be performed using the `kubectl` command line tool.
|
||||
:::
|
||||
|
||||
1. Click **☰ > Cluster Management**.
|
||||
1. Choose the cluster you want to provide vSphere storage to and click **Exlpore**.
|
||||
1. Choose the cluster you want to provide vSphere storage to and click **Explore**.
|
||||
1. In the left navigation bar, select **Storage > StorageClasses**.
|
||||
1. Click **Create**.
|
||||
3. Enter a **Name** for the StorageClass.
|
||||
|
||||
+1
@@ -19,6 +19,7 @@ In order to deploy and run the adapter successfully, you need to ensure its vers
|
||||
|
||||
| Rancher Version | Adapter Version |
|
||||
|-----------------|------------------|
|
||||
| v2.11.14 | v106.0.1+up6.0.1 |
|
||||
| v2.11.13 | v106.0.0+up6.0.0 |
|
||||
| v2.11.12 | v106.0.0+up6.0.0 |
|
||||
| v2.11.11 | v106.0.0+up6.0.0 |
|
||||
|
||||
@@ -116,6 +116,19 @@ By default, Rancher collects logs for control plane components and node componen
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Resource Exhaustion of `inotify` Watchers and File Descriptors
|
||||
|
||||
When enabling the **Logging** app on Linux systems that heavily monitor the filesystem, you may encounter `Too many open files` or `CrashLoopBackOff` failures related to applications that leverage `inotify` to watch for file changes.
|
||||
|
||||
This happens because the Linux kernel caps the number of files a user can open and the number of directory paths a subsystem can watch simultaneously. To resolve this, you must explicitly increase your `inotify` system limits.
|
||||
|
||||
Below are example commands an admin user can run to increase system limits for `inotify` user instances and watches:
|
||||
|
||||
```shell
|
||||
sysctl -w fs.inotify.max_user_instances=8192
|
||||
sysctl -w fs.inotify.max_user_watches=524288
|
||||
```
|
||||
|
||||
### The Logging Buffer Overloads Pods
|
||||
|
||||
Depending on your configuration, the default buffer size may be too large and cause pod failures. One way to reduce the load is to lower the logger's flush interval. This prevents logs from overfilling the buffer. You can also add more flush threads to handle moments when many logs are attempting to fill the buffer at once.
|
||||
|
||||
@@ -14,64 +14,74 @@ Examples of built-in Rancher extensions are Fleet, Explorer, and Harvester. Exam
|
||||
|
||||
## Prerequisites
|
||||
|
||||
> You must log in as an admin in order to view and interact with the extensions management page.
|
||||
> 1. You must log in as an administrator to view and interact with the extensions management page.
|
||||
> 2. You must enable extension support.
|
||||
|
||||
## Enabling Extension Support in Rancher
|
||||
|
||||
Rancher v2.9.0 and later includes extension support.
|
||||
|
||||
You can confirm if extension support is enabled by checking the `uiextension` feature flag. Any changes to this feature flag cause the Rancher pod to restart.
|
||||
|
||||
When you enable extension support for the first time, it creates resources, such as the CRDs, so that UI Extensions can work. When extension support is disabled, it disables the endpoints and does not cache any files. However, it does not remove any CRs or delete any extensions that were installed before. If re-enabled, it exposes the required endpoints again and creates CRDs as needed. The extensions that were already installed load after the Rancher pod restarts.
|
||||
|
||||
### Caching Extension Files
|
||||
|
||||
By default, Rancher caches every extension file in the file system. You can change that behavior by setting `plugin.noCache` to `true`.
|
||||
|
||||
Rancher does have a cached file size limit of 30MB. If an extension has a file bigger than that, the cache is disabled and `plugin.noCache` is set to `true`, regardless of user input.
|
||||
|
||||
|
||||
## Installing Extensions
|
||||
|
||||
1. Click **☰ > Extensions** under **Configuration**.
|
||||
|
||||
2. If not already installed in **Apps**, you must enable the extension operator by clicking the **Enable** button.
|
||||
2. On the **Extensions** page, select the **Available** tab to choose the extensions that you want to install.
|
||||
|
||||
- Click **OK** to add the Rancher extension repository if your installation is not air-gapped. Otherwise, uncheck the box to do so and click **OK**.
|
||||
3. If no extensions are listed as available, you can manually add the repos:
|
||||
|
||||

|
||||
3.1. On the upper right, click **⋮ > Manage Repositories > Create**.
|
||||
|
||||
3. On the **Extensions** page, click on the **Available** tab to select which extensions you want to install.
|
||||
3.2. Add the desired repo name, making sure to also specify the Git repo URL and the Git branch.
|
||||
|
||||
4. If no extensions are showing as available, you may manually add repos as follows:
|
||||
3.3. Click **Create** in the lower right again to complete.
|
||||
|
||||
4.1. On the upper right of screen, click on **⋮ > Manage Repositories > Create**.
|
||||

|
||||
|
||||
4.2. Add the desired repo name, making sure to also specify the Git Repo URL and the Git Branch.
|
||||
4. Under the **Available** tab, click **Install** on the desired extension and version, as in the example below. You can also update your extension from this screen. The button to **Update** appears on the extension card if an update is available.
|
||||
|
||||
4.3. Click **Create** in the lower right again to complete.
|
||||

|
||||
|
||||

|
||||
5. Click **Reload** after your extension successfully installs to check its status. Updates to the UI aren't visible until you reload the page.
|
||||
|
||||
5. Under the **Available** tab, click **Install** on the desired extension and version as in the example below. You can also update your extension from this screen, as the button to **Update** will appear on the extension if one is available.
|
||||
|
||||

|
||||
|
||||
6. Click the **Reload** page button that will appear after your extension successfully installs. Note that a logged-in user who has just installed an extension will not see a change to the UI **unless** they reload the page.
|
||||
|
||||

|
||||

|
||||
|
||||
## Updating and Upgrading Extensions
|
||||
|
||||
1. Click **☰ > Extensions** under **Configuration**.
|
||||
1. Select the **Updates** tab.
|
||||
1. Click **Update**.
|
||||
2. Select the **Updates** tab.
|
||||
3. Click **Update**.
|
||||
|
||||
If there is a new version of the extension, there will also be an **Update** button visible on the associated card for the extension in the **Available** tab.
|
||||
|
||||
## Deleting Extensions
|
||||
|
||||
1. Click **☰**, then click on the name of your local cluster.
|
||||
1. From the sidebar, select **Apps > Installed Apps**.
|
||||
1. Find the name of the chart you want to delete and select the checkbox next to it.
|
||||
1. Click **Delete**.
|
||||
2. From the sidebar, select **Apps > Installed Apps**.
|
||||
3. Find the name of the chart you want to delete and select the checkbox next to it.
|
||||
4. Click **Delete**.
|
||||
|
||||
## Deleting Extension Repositories
|
||||
|
||||
1. Click **☰ > Extensions** under **Configuration**.
|
||||
1. On the top right, click **⋮ > Manage Repositories**.
|
||||
1. Find the name of the extension repository you want to delete. Select the checkbox next to the repository name, then click **Delete**.
|
||||
2. On the top right, click **⋮ > Manage Repositories**.
|
||||
3. Find the name of the extension repository you want to delete. Select the checkbox next to the repository name, then click **Delete**.
|
||||
|
||||
## Deleting Extension Repository Container Images
|
||||
|
||||
1. Click **☰**, then select **Extensions**, under **Configuration**.
|
||||
1. On the top right, click **⋮ > Manage Extension Catalogs**.
|
||||
1. Find the name of the container image you want to delete, then click **⋮ > Uninstall**.
|
||||
2. On the top right, click **⋮ > Manage Extension Catalogs**.
|
||||
3. Find the name of the container image you want to delete, then click **⋮ > Uninstall**.
|
||||
|
||||
## Uninstalling Extensions
|
||||
|
||||
@@ -79,11 +89,11 @@ There are two ways to uninstall or disable an extension:
|
||||
|
||||
1. Under the **Installed** tab, click the **Uninstall** button on the extension you wish to remove.
|
||||
|
||||

|
||||

|
||||
|
||||
1. On the extensions management page, click **⋮ > Disable Extension Support**. This will disable all installed extensions.
|
||||
2. On the extensions management page, click **⋮ > Disable Extension Support**. This will disable all installed extensions.
|
||||
|
||||

|
||||

|
||||
|
||||
:::caution
|
||||
|
||||
@@ -91,6 +101,12 @@ You must reload the page after disabling extensions or display issues may occur.
|
||||
|
||||
:::
|
||||
|
||||
## Enabling Unauthenticated Access to an Extension
|
||||
|
||||
In Rancher v2.9.0 and later, you can allow unauthenticated access to an extension. You may want to enable unauthenticated access if the extension enables a new locale or adds custom branding. By default, all extensions require user authentication to load.
|
||||
|
||||
To enable unauthenticated access to an extension, set `plugin.noAuth` to `true` in the CR used by the extension.
|
||||
|
||||
## Developing Extensions
|
||||
|
||||
To learn how to develop your own extensions, refer to the official [Getting Started](https://rancher.github.io/dashboard/extensions/extensions-getting-started) guide.
|
||||
@@ -173,8 +189,8 @@ After you successfully set up these resources, you can install the extensions fr
|
||||
1. Click **☰**, then select **Extensions**, under **Configuration**.
|
||||
1. On the top right, click **⋮ > Manage Extension Catalogs**.
|
||||
1. Select the **Import Extension Catalog** button.
|
||||
1. Enter the image address in the **Catalog Image Reference** field.
|
||||
* **(Optional)** If the container image is private, select the secret you just created from the **Pull Secrets** drop-down menu.
|
||||
1. Enter the image address in the **Catalog Image Reference** field.
|
||||
- **(Optional)** If the container image is private, select the secret you just created from the **Pull Secrets** drop-down menu.
|
||||
1. Click **Load**. The extension will now be **Pending**.
|
||||
1. Return to the **Extensions** page.
|
||||
1. Select the **Available** tab, and click **Reload** to make sure that the list of extensions is up to date.
|
||||
@@ -191,7 +207,7 @@ After you mirror the latest changes, follow these steps:
|
||||
1. Click **☰ > Local**.
|
||||
1. From the sidebar, select **Workloads > Deployments**.
|
||||
1. From the namespaces dropdown menu, select **cattle-ui-plugin-system**.
|
||||
1. Find the **cattle-ui-plugin-system** namespace.
|
||||
1. Find the **cattle-ui-plugin-system** namespace.
|
||||
1. Select the `ui-plugin-catalog` deployment.
|
||||
1. Click **⋮ > Edit config**.
|
||||
1. Update the **Container Image** field within the deployment's container with the latest image.
|
||||
|
||||
@@ -20,6 +20,7 @@ Each Rancher version is designed to be compatible with a single version of the w
|
||||
|
||||
| Rancher Version | Webhook Version | Availability in Prime | Availability in Community |
|
||||
|-----------------|-----------------|-----------------------|---------------------------|
|
||||
| v2.11.14 | v0.7.9 | ✓ | ✗ |
|
||||
| v2.11.13 | v0.7.8 | ✓ | ✗ |
|
||||
| v2.11.12 | v0.7.8 | ✓ | ✗ |
|
||||
| v2.11.11 | v0.7.8 | ✓ | ✗ |
|
||||
|
||||
+85
-14
@@ -78,66 +78,137 @@ kubectl -n kube-system get endpoints kube-scheduler -o jsonpath='{.metadata.anno
|
||||
|
||||
## Ingress Controller
|
||||
|
||||
The default Ingress Controller is Traefik and is deployed as a DaemonSet in the `traefik` namespace. The pods are only scheduled to nodes with the `worker` role.
|
||||
The default Ingress Controller is Traefik and is deployed as a DaemonSet in the `kube-system` namespace. The pods are only scheduled to nodes with the `worker` role.
|
||||
|
||||
Check if the pods are running on all nodes:
|
||||
|
||||
```
|
||||
kubectl -n traefik get pods -o wide
|
||||
kubectl -n kube-system get pods -o wide
|
||||
```
|
||||
|
||||
Example output:
|
||||
Example RKE2 output:
|
||||
|
||||
```
|
||||
kubectl -n traefik get pods -o wide
|
||||
kubectl -n kube-system get pods -o wide
|
||||
NAME READY STATUS RESTARTS AGE IP NODE
|
||||
default-http-backend-797c5bc547-kwwlq 1/1 Running 0 17m x.x.x.x worker-1
|
||||
traefik-4qd64 1/1 Running 0 14m x.x.x.x worker-1
|
||||
traefik-8wxhm 1/1 Running 0 13m x.x.x.x worker-0
|
||||
local-path-provisioner-xxxxxxxxxx-xxxxx 1/1 Running 0 17m x.x.x.x worker-1
|
||||
rke2-traefik-xxxxxxxxxx-xxxxx 1/1 Running 0 14m x.x.x.x worker-1
|
||||
svclb-rke2-traefik-xxxxxxxx-xxxxx 1/1 Running 0 13m x.x.x.x worker-0
|
||||
...
|
||||
```
|
||||
|
||||
Example K3s output:
|
||||
|
||||
```
|
||||
kubectl -n kube-system get pods -o wide
|
||||
NAME READY STATUS RESTARTS AGE IP NODE
|
||||
local-path-provisioner-xxxxxxxxxx-xxxxx 1/1 Running 0 17m x.x.x.x worker-1
|
||||
traefik-xxxxxxxxxx-xxxxx 1/1 Running 0 14m x.x.x.x worker-1
|
||||
svclb-traefik-xxxxxxxx-xxxxx 1/1 Running 0 13m x.x.x.x worker-0
|
||||
...
|
||||
```
|
||||
|
||||
If a pod is unable to run (Status is not **Running**, Ready status is not showing `1/1` or you see a high count of Restarts), check the pod details, logs and namespace events.
|
||||
|
||||
### Pod details
|
||||
|
||||
RKE2 example:
|
||||
|
||||
```
|
||||
kubectl -n traefik describe pods -l app=traefik
|
||||
kubectl -n kube-system describe pods -l app.kubernetes.io/name=rke2-traefik
|
||||
```
|
||||
|
||||
K3s example:
|
||||
|
||||
```
|
||||
kubectl -n kube-system describe pods -l app.kubernetes.io/name=traefik
|
||||
```
|
||||
|
||||
### Pod container logs
|
||||
|
||||
The below command can show the logs of all the pods labeled "app=traefik", but it will display only 10 lines of log because of the restrictions of the `kubectl logs` command. Refer to `--tail` of `kubectl logs -h` for more information.
|
||||
The below command can show the logs of all the pods labeled "app.kubernetes.io/name=rke2-traefik" if using RKE2 or "app.kubernetes.io/name=traefik" if using K3s, but it will display only 10 lines of log because of the restrictions of the `kubectl logs` command. Refer to `--tail` of `kubectl logs -h` for more information.
|
||||
|
||||
RKE2 example:
|
||||
|
||||
```
|
||||
kubectl -n traefik logs -l app=traefik
|
||||
kubectl -n kube-system logs -l app.kubernetes.io/name=rke2-traefik
|
||||
```
|
||||
|
||||
K3s example:
|
||||
|
||||
```
|
||||
kubectl -n kube-system logs -l app.kubernetes.io/name=traefik
|
||||
```
|
||||
|
||||
If the full log is needed, specify the pod name in the trailing command:
|
||||
|
||||
```
|
||||
kubectl -n traefik logs <pod name>
|
||||
kubectl -n kube-system logs <pod name>
|
||||
```
|
||||
|
||||
### Namespace events
|
||||
|
||||
```
|
||||
kubectl -n traefik get events
|
||||
kubectl -n kube-system get events
|
||||
```
|
||||
|
||||
### Debug logging
|
||||
|
||||
To enable debug logging:
|
||||
|
||||
RKE2 example:
|
||||
|
||||
```
|
||||
kubectl -n traefik patch ds traefik --type='json' -p='[{"op": "add", "path": "/spec/template/spec/containers/0/args/-", "value": "--v=5"}]'
|
||||
cat <<EOF | kubectl apply -f -
|
||||
apiVersion: helm.cattle.io/v1
|
||||
kind: HelmChartConfig
|
||||
metadata:
|
||||
name: rke2-traefik
|
||||
namespace: kube-system
|
||||
spec:
|
||||
valuesContent: |-
|
||||
additionalArguments:
|
||||
- "--log.level=DEBUG"
|
||||
EOF
|
||||
```
|
||||
|
||||
K3s example:
|
||||
|
||||
```
|
||||
cat <<EOF | kubectl apply -f -
|
||||
apiVersion: helm.cattle.io/v1
|
||||
kind: HelmChartConfig
|
||||
metadata:
|
||||
name: traefik
|
||||
namespace: kube-system
|
||||
spec:
|
||||
valuesContent: |-
|
||||
logs:
|
||||
general:
|
||||
level: "DEBUG"
|
||||
EOF
|
||||
```
|
||||
|
||||
### Check configuration
|
||||
|
||||
Retrieve generated configuration in each pod:
|
||||
|
||||
RKE2 example for manual Traefik configuration file check:
|
||||
|
||||
```
|
||||
kubectl -n traefik get pods -l app=traefik --no-headers -o custom-columns=.NAME:.metadata.name | while read pod; do kubectl -n traefik exec $pod -- cat /etc/nginx/nginx.conf; done
|
||||
kubectl exec -n kube-system pod/traefik-xxxxxxxxx-xxxxx -- cat /var/lib/rancher/rke2/server/manifests/rke2-traefik-config.yaml
|
||||
```
|
||||
|
||||
K3s example for manual Traefik configuration file check:
|
||||
|
||||
```
|
||||
kubectl exec -n kube-system pod/traefik-xxxxxxxxx-xxxxx -- cat /var/lib/rancher/k3s/server/manifests/k3s-traefik-config.yaml
|
||||
```
|
||||
|
||||
RKE2/K3s example for Traefik CLI argument configuration check:
|
||||
|
||||
```
|
||||
kubectl get pod traefik-xxxxxxxxx-xxxxx -n kube-system -o jsonpath='{.spec.containers[0].args}'
|
||||
```
|
||||
|
||||
## Rancher agents
|
||||
|
||||
@@ -15,6 +15,14 @@ Make sure you configured the correct kubeconfig (for example, `export KUBECONFIG
|
||||
Double check if all the [required ports](../../how-to-guides/new-user-guides/kubernetes-clusters-in-rancher-setup/node-requirements-for-rancher-managed-clusters.md#networking-requirements) are opened in your (host) firewall. The overlay network uses UDP in comparison to all other required ports which are TCP.
|
||||
|
||||
|
||||
## Check if your downstream node can communicate to Rancher Manager
|
||||
|
||||
Rancher components with HTTP endpoints generally contain a `ping` liveness probe which you can use to test connectivity. Replace the `$RANCHER_URL` as appropriate and run the following from a node to check that it has connectivity to Rancher Manager's servers in the `local` cluster. If successful, it should return `pong`.
|
||||
|
||||
```
|
||||
curl -k https://$RANCHER_URL/ping
|
||||
```
|
||||
|
||||
## Check if Overlay Network is Functioning Correctly
|
||||
|
||||
The pod can be scheduled to any of the hosts you used for your cluster, but that means that the NGINX ingress controller needs to be able to route the request from `NODE_1` to `NODE_2`. This happens over the overlay network. If the overlay network is not functioning, you will experience intermittent TCP/HTTP connection failures due to the NGINX ingress controller not being able to route to the pod.
|
||||
|
||||
@@ -12,6 +12,7 @@ Rancher will publish deprecated features as part of the [release notes](https://
|
||||
|
||||
| Patch Version | Release Date |
|
||||
|---------------|---------------|
|
||||
| [2.12.10](https://github.com/rancher/rancher/releases/tag/v2.12.10) | May 27, 2026 |
|
||||
| [2.12.9](https://github.com/rancher/rancher/releases/tag/v2.12.9) | April 30, 2026 |
|
||||
| [2.12.8](https://github.com/rancher/rancher/releases/tag/v2.12.8) | March 25, 2026 |
|
||||
| [2.12.7](https://github.com/rancher/rancher/releases/tag/v2.12.7) | February 25, 2026 |
|
||||
|
||||
+34
-9
@@ -26,6 +26,31 @@ Review the list of known issues for each Rancher version, which can be found in
|
||||
|
||||
Note that upgrades _to_ or _from_ any chart in the [rancher-alpha repository](../resources/choose-a-rancher-version.md#helm-chart-repositories) aren't supported.
|
||||
|
||||
### Upgrade Path
|
||||
|
||||
:::important
|
||||
|
||||
**Important:** The only tested and supported Rancher upgrade path between minor versions (e.g. v2.11.x to v2.12.x) is to upgrade from the latest available patch version of your current running minor release to the latest available patch version of the next minor release.
|
||||
|
||||
Before initiating a minor version upgrade, verify that you are running the most recent patch release of your current version.
|
||||
|
||||
You can query the available chart versions with the Helm CLI:
|
||||
|
||||
1. Update your local Helm repo cache.
|
||||
|
||||
```
|
||||
helm repo update
|
||||
```
|
||||
|
||||
1. Search for available versions in your [specific repository](../resources/choose-a-rancher-version.md#helm-chart-repositories) (e.g., rancher-stable):
|
||||
|
||||
```
|
||||
helm search repo rancher-<CHART_REPO>/rancher --versions
|
||||
```
|
||||
|
||||
If your installation is not on the latest patch version of the current minor release, you must upgrade to that version before proceeding to the next minor version.
|
||||
:::
|
||||
|
||||
### Helm Version
|
||||
|
||||
The upgrade instructions assume you are using Helm 3.
|
||||
@@ -138,17 +163,17 @@ There will be more values that are listed with this command. This is just an exa
|
||||
|
||||
:::
|
||||
|
||||
:::tip
|
||||
:::tip
|
||||
|
||||
Your deployment name may vary; for example, if you're deploying Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
Your deployment name may vary; for example, if you're deploying Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
```
|
||||
helm get values rancher-stable -n cattle-system
|
||||
|
||||
hostname: rancher.my.org
|
||||
```
|
||||
|
||||
:::
|
||||
:::
|
||||
|
||||
If you are upgrading cert-manager to the latest version from v1.5 or below, follow the [cert-manager upgrade docs](../resources/upgrade-cert-manager.md#option-c-upgrade-cert-manager-from-versions-15-and-below) to learn how to upgrade cert-manager without needing to perform an uninstall or reinstall of Rancher. Otherwise, follow the [steps to upgrade Rancher](#steps-to-upgrade-rancher) below.
|
||||
|
||||
@@ -171,17 +196,17 @@ The above is an example, there may be more values from the previous step that ne
|
||||
|
||||
:::
|
||||
|
||||
:::tip
|
||||
:::tip
|
||||
|
||||
If you deploy Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
If you deploy Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
```
|
||||
helm upgrade rancher-stable rancher-<CHART_REPO>/rancher \
|
||||
--namespace cattle-system \
|
||||
--set hostname=rancher.my.org
|
||||
```
|
||||
|
||||
:::
|
||||
:::
|
||||
|
||||
Alternatively, it's possible to export the current values to a file and reference that file during upgrade. For example, to only change the Rancher version:
|
||||
|
||||
@@ -191,7 +216,7 @@ Alternatively, it's possible to export the current values to a file and referenc
|
||||
```
|
||||
1. Update only the Rancher version:
|
||||
|
||||
|
||||
|
||||
```
|
||||
helm upgrade rancher rancher-<CHART_REPO>/rancher \
|
||||
--namespace cattle-system \
|
||||
|
||||
+2
-2
@@ -33,7 +33,7 @@ To use a premade dashboard, go to [https://grafana.com/grafana/dashboards](https
|
||||
To use your own dashboard:
|
||||
|
||||
1. Click on the link to open Grafana. On the cluster detail page, click **Monitoring**.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin/prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
:::note
|
||||
|
||||
@@ -113,7 +113,7 @@ Note that the RBAC roles exposed by the Monitoring chart to add Grafana Dashboar
|
||||
1. On the **Clusters** page, go to the cluster where you want to configure the Grafana namespace and click **Explore**.
|
||||
1. In the left navigation bar, click **Monitoring**.
|
||||
1. Click **Grafana**.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin/prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
:::note
|
||||
|
||||
|
||||
+1
-1
@@ -20,7 +20,7 @@ To see the links to the external monitoring UIs, including Grafana dashboards, y
|
||||
1. In the left navigation menu, click **Monitoring.**
|
||||
1. Click **Grafana.** The Grafana dashboard should open in a new tab.
|
||||
1. Go to the log in icon in the lower left corner and click **Sign In.**
|
||||
1. Log in to Grafana. The default Admin username and password for the Grafana instance is `admin/prom-operator`. (Regardless of who has the password, cluster administrator permission in Rancher is still required access the Grafana instance.) Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. (Regardless of who has the password, cluster administrator permission in Rancher is still required access the Grafana instance.) Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
|
||||
### Getting the PromQL Query Powering a Grafana Panel
|
||||
|
||||
+8
-2
@@ -31,11 +31,17 @@ _Cluster roles_ are roles that you can assign to users, granting them access to
|
||||
|
||||
- **Cluster Owner:**
|
||||
|
||||
These users have full control over the cluster and all resources in it.
|
||||
These users have full control over the cluster and all resources in it.
|
||||
|
||||
- **Cluster Member:**
|
||||
|
||||
These users can view most cluster level resources and create new projects.
|
||||
These users can view most cluster level resources and create new projects.
|
||||
|
||||
:::warning
|
||||
|
||||
When a Cluster Member creates a project, the user is automatically assigned [Project Owner privileges](#project-roles). This grants them comprehensive control over the project and its associated resources, including permissions to deploy workloads. Without enforced [Pod Security Standards (PSS) and Pod Security Admission (PSA)](../pod-security-standards.md), a Cluster Member is able to execute privileged containers in the cluster.
|
||||
|
||||
:::
|
||||
|
||||
#### Custom Cluster Roles
|
||||
|
||||
|
||||
+1
-1
@@ -234,7 +234,7 @@ docker stop <original-rancher-container>
|
||||
|
||||
:::note
|
||||
|
||||
If you wish to keep the original Rancher environment running, you can also restart the cattle-cluster-agent pods on each cluster connected to your Rancher environment.
|
||||
If clusters do not automatically reconnect to the new environment after you have redirected traffic, for example if there is a delay in scaling down the original Rancher instance, you can also restart the cattle-cluster-agent pods on each cluster connected to your Rancher environment.
|
||||
|
||||
```bash
|
||||
kubectl rollout restart deployment cattle-cluster-agent -n cattle-system
|
||||
|
||||
+1
-1
@@ -24,7 +24,7 @@ The following steps can also be performed using the `kubectl` command line tool.
|
||||
:::
|
||||
|
||||
1. Click **☰ > Cluster Management**.
|
||||
1. Choose the cluster you want to provide vSphere storage to and click **Exlpore**.
|
||||
1. Choose the cluster you want to provide vSphere storage to and click **Explore**.
|
||||
1. In the left navigation bar, select **Storage > StorageClasses**.
|
||||
1. Click **Create**.
|
||||
3. Enter a **Name** for the StorageClass.
|
||||
|
||||
+1
@@ -19,6 +19,7 @@ In order to deploy and run the adapter successfully, you need to ensure its vers
|
||||
|
||||
| Rancher Version | Adapter Version |
|
||||
|-----------------|------------------|
|
||||
| v2.12.10 | 107.0.0+up7.0.0 |
|
||||
| v2.12.9 | 107.0.0+up7.0.0 |
|
||||
| v2.12.8 | 107.0.0+up7.0.0 |
|
||||
| v2.12.7 | 107.0.0+up7.0.0 |
|
||||
|
||||
@@ -116,6 +116,19 @@ By default, Rancher collects logs for control plane components and node componen
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Resource Exhaustion of `inotify` Watchers and File Descriptors
|
||||
|
||||
When enabling the **Logging** app on Linux systems that heavily monitor the filesystem, you may encounter `Too many open files` or `CrashLoopBackOff` failures related to applications that leverage `inotify` to watch for file changes.
|
||||
|
||||
This happens because the Linux kernel caps the number of files a user can open and the number of directory paths a subsystem can watch simultaneously. To resolve this, you must explicitly increase your `inotify` system limits.
|
||||
|
||||
Below are example commands an admin user can run to increase system limits for `inotify` user instances and watches:
|
||||
|
||||
```shell
|
||||
sysctl -w fs.inotify.max_user_instances=8192
|
||||
sysctl -w fs.inotify.max_user_watches=524288
|
||||
```
|
||||
|
||||
### The Logging Buffer Overloads Pods
|
||||
|
||||
Depending on your configuration, the default buffer size may be too large and cause pod failures. One way to reduce the load is to lower the logger's flush interval. This prevents logs from overfilling the buffer. You can also add more flush threads to handle moments when many logs are attempting to fill the buffer at once.
|
||||
|
||||
@@ -14,64 +14,74 @@ Examples of built-in Rancher extensions are Fleet, Explorer, and Harvester. Exam
|
||||
|
||||
## Prerequisites
|
||||
|
||||
> You must log in as an admin in order to view and interact with the extensions management page.
|
||||
> 1. You must log in as an administrator to view and interact with the extensions management page.
|
||||
> 2. You must enable extension support.
|
||||
|
||||
## Enabling Extension Support in Rancher
|
||||
|
||||
Rancher v2.9.0 and later includes extension support.
|
||||
|
||||
You can confirm if extension support is enabled by checking the `uiextension` feature flag. Any changes to this feature flag cause the Rancher pod to restart.
|
||||
|
||||
When you enable extension support for the first time, it creates resources, such as the CRDs, so that UI Extensions can work. When extension support is disabled, it disables the endpoints and does not cache any files. However, it does not remove any CRs or delete any extensions that were installed before. If re-enabled, it exposes the required endpoints again and creates CRDs as needed. The extensions that were already installed load after the Rancher pod restarts.
|
||||
|
||||
### Caching Extension Files
|
||||
|
||||
By default, Rancher caches every extension file in the file system. You can change that behavior by setting `plugin.noCache` to `true`.
|
||||
|
||||
Rancher does have a cached file size limit of 30MB. If an extension has a file bigger than that, the cache is disabled and `plugin.noCache` is set to `true`, regardless of user input.
|
||||
|
||||
|
||||
## Installing Extensions
|
||||
|
||||
1. Click **☰ > Extensions** under **Configuration**.
|
||||
|
||||
2. If not already installed in **Apps**, you must enable the extension operator by clicking the **Enable** button.
|
||||
2. On the **Extensions** page, select the **Available** tab to choose the extensions that you want to install.
|
||||
|
||||
- Click **OK** to add the Rancher extension repository if your installation is not air-gapped. Otherwise, uncheck the box to do so and click **OK**.
|
||||
3. If no extensions are listed as available, you can manually add the repos:
|
||||
|
||||

|
||||
3.1. On the upper right, click **⋮ > Manage Repositories > Create**.
|
||||
|
||||
3. On the **Extensions** page, click on the **Available** tab to select which extensions you want to install.
|
||||
3.2. Add the desired repo name, making sure to also specify the Git repo URL and the Git branch.
|
||||
|
||||
4. If no extensions are showing as available, you may manually add repos as follows:
|
||||
3.3. Click **Create** in the lower right again to complete.
|
||||
|
||||
4.1. On the upper right of screen, click on **⋮ > Manage Repositories > Create**.
|
||||

|
||||
|
||||
4.2. Add the desired repo name, making sure to also specify the Git Repo URL and the Git Branch.
|
||||
4. Under the **Available** tab, click **Install** on the desired extension and version, as in the example below. You can also update your extension from this screen. The button to **Update** appears on the extension card if an update is available.
|
||||
|
||||
4.3. Click **Create** in the lower right again to complete.
|
||||

|
||||
|
||||

|
||||
5. Click **Reload** after your extension successfully installs to check its status. Updates to the UI aren't visible until you reload the page.
|
||||
|
||||
5. Under the **Available** tab, click **Install** on the desired extension and version as in the example below. You can also update your extension from this screen, as the button to **Update** will appear on the extension if one is available.
|
||||
|
||||

|
||||
|
||||
6. Click the **Reload** page button that will appear after your extension successfully installs. Note that a logged-in user who has just installed an extension will not see a change to the UI **unless** they reload the page.
|
||||
|
||||

|
||||

|
||||
|
||||
## Updating and Upgrading Extensions
|
||||
|
||||
1. Click **☰ > Extensions** under **Configuration**.
|
||||
1. Select the **Updates** tab.
|
||||
1. Click **Update**.
|
||||
2. Select the **Updates** tab.
|
||||
3. Click **Update**.
|
||||
|
||||
If there is a new version of the extension, there will also be an **Update** button visible on the associated card for the extension in the **Available** tab.
|
||||
|
||||
## Deleting Extensions
|
||||
|
||||
1. Click **☰**, then click on the name of your local cluster.
|
||||
1. From the sidebar, select **Apps > Installed Apps**.
|
||||
1. Find the name of the chart you want to delete and select the checkbox next to it.
|
||||
1. Click **Delete**.
|
||||
2. From the sidebar, select **Apps > Installed Apps**.
|
||||
3. Find the name of the chart you want to delete and select the checkbox next to it.
|
||||
4. Click **Delete**.
|
||||
|
||||
## Deleting Extension Repositories
|
||||
|
||||
1. Click **☰ > Extensions** under **Configuration**.
|
||||
1. On the top right, click **⋮ > Manage Repositories**.
|
||||
1. Find the name of the extension repository you want to delete. Select the checkbox next to the repository name, then click **Delete**.
|
||||
2. On the top right, click **⋮ > Manage Repositories**.
|
||||
3. Find the name of the extension repository you want to delete. Select the checkbox next to the repository name, then click **Delete**.
|
||||
|
||||
## Deleting Extension Repository Container Images
|
||||
|
||||
1. Click **☰**, then select **Extensions**, under **Configuration**.
|
||||
1. On the top right, click **⋮ > Manage Extension Catalogs**.
|
||||
1. Find the name of the container image you want to delete, then click **⋮ > Uninstall**.
|
||||
2. On the top right, click **⋮ > Manage Extension Catalogs**.
|
||||
3. Find the name of the container image you want to delete, then click **⋮ > Uninstall**.
|
||||
|
||||
## Uninstalling Extensions
|
||||
|
||||
@@ -79,11 +89,11 @@ There are two ways to uninstall or disable an extension:
|
||||
|
||||
1. Under the **Installed** tab, click the **Uninstall** button on the extension you wish to remove.
|
||||
|
||||

|
||||

|
||||
|
||||
1. On the extensions management page, click **⋮ > Disable Extension Support**. This will disable all installed extensions.
|
||||
2. On the extensions management page, click **⋮ > Disable Extension Support**. This will disable all installed extensions.
|
||||
|
||||

|
||||

|
||||
|
||||
:::caution
|
||||
|
||||
@@ -91,6 +101,12 @@ You must reload the page after disabling extensions or display issues may occur.
|
||||
|
||||
:::
|
||||
|
||||
## Enabling Unauthenticated Access to an Extension
|
||||
|
||||
In Rancher v2.9.0 and later, you can allow unauthenticated access to an extension. You may want to enable unauthenticated access if the extension enables a new locale or adds custom branding. By default, all extensions require user authentication to load.
|
||||
|
||||
To enable unauthenticated access to an extension, set `plugin.noAuth` to `true` in the CR used by the extension.
|
||||
|
||||
## Developing Extensions
|
||||
|
||||
To learn how to develop your own extensions, refer to the official [Getting Started](https://rancher.github.io/dashboard/extensions/extensions-getting-started) guide.
|
||||
@@ -173,8 +189,8 @@ After you successfully set up these resources, you can install the extensions fr
|
||||
1. Click **☰**, then select **Extensions**, under **Configuration**.
|
||||
1. On the top right, click **⋮ > Manage Extension Catalogs**.
|
||||
1. Select the **Import Extension Catalog** button.
|
||||
1. Enter the image address in the **Catalog Image Reference** field.
|
||||
* **(Optional)** If the container image is private, select the secret you just created from the **Pull Secrets** drop-down menu.
|
||||
1. Enter the image address in the **Catalog Image Reference** field.
|
||||
- **(Optional)** If the container image is private, select the secret you just created from the **Pull Secrets** drop-down menu.
|
||||
1. Click **Load**. The extension will now be **Pending**.
|
||||
1. Return to the **Extensions** page.
|
||||
1. Select the **Available** tab, and click **Reload** to make sure that the list of extensions is up to date.
|
||||
@@ -191,7 +207,7 @@ After you mirror the latest changes, follow these steps:
|
||||
1. Click **☰ > Local**.
|
||||
1. From the sidebar, select **Workloads > Deployments**.
|
||||
1. From the namespaces dropdown menu, select **cattle-ui-plugin-system**.
|
||||
1. Find the **cattle-ui-plugin-system** namespace.
|
||||
1. Find the **cattle-ui-plugin-system** namespace.
|
||||
1. Select the `ui-plugin-catalog` deployment.
|
||||
1. Click **⋮ > Edit config**.
|
||||
1. Update the **Container Image** field within the deployment's container with the latest image.
|
||||
|
||||
@@ -20,6 +20,7 @@ Each Rancher version is designed to be compatible with a single version of the w
|
||||
|
||||
| Rancher Version | Webhook Version | Availability in Prime | Availability in Community |
|
||||
|-----------------|-----------------|-----------------------|---------------------------|
|
||||
| v2.12.10 | v0.8.6 | ✓ | ✗ |
|
||||
| v2.12.9 | v0.8.5 | ✓ | ✗ |
|
||||
| v2.12.8 | v0.8.5 | ✓ | ✗ |
|
||||
| v2.12.7 | v0.8.5 | ✓ | ✗ |
|
||||
|
||||
+209
-102
@@ -6,29 +6,63 @@ title: Troubleshooting etcd Nodes
|
||||
<link rel="canonical" href="https://ranchermanager.docs.rancher.com/troubleshooting/kubernetes-components/troubleshooting-etcd-nodes"/>
|
||||
</head>
|
||||
|
||||
This section contains commands and tips for troubleshooting nodes with the `etcd` role.
|
||||
This section contains commands and tips for troubleshooting nodes with the `etcd` role in RKE2 and K3s clusters.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
As RKE2 and K3s rely on `containerd` as the container runtime, `crictl` replaces Docker for container management. Before proceeding with the troubleshooting commands, configure your environment by exporting the following variables:
|
||||
|
||||
### RKE2
|
||||
|
||||
```bash
|
||||
export PATH=$PATH:/var/lib/rancher/rke2/bin/
|
||||
export CRI_CONFIG_FILE=/var/lib/rancher/rke2/agent/etc/crictl.yaml
|
||||
etcdcontainer=$(crictl ps --name etcd --quiet)
|
||||
```
|
||||
|
||||
### K3s
|
||||
|
||||
> ### ⚠️ **Warning**
|
||||
> K3s does not include `etcdctl` in the system PATH. If you need to perform etcd troubleshooting on a K3s cluster, you may need to install it or locate it within the K3s data directory.
|
||||
|
||||
```bash
|
||||
export PATH=$PATH:/usr/local/bin
|
||||
export CRI_CONFIG_FILE=/var/lib/rancher/k3s/agent/etc/crictl.yaml
|
||||
```
|
||||
|
||||
|
||||
## Checking if the etcd Container is Running
|
||||
|
||||
The container for etcd should have status **Up**. The duration shown after **Up** is the time the container has been running.
|
||||
**RKE2**: The container for etcd should be in the **Running** state.
|
||||
|
||||
```
|
||||
docker ps -a -f=name=etcd$
|
||||
```bash
|
||||
crictl ps --name etcd
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES
|
||||
d26adbd23643 rancher/mirrored-coreos-etcd:v3.5.7 "/usr/local/bin/etcd…" 30 minutes ago Up 30 minutes etcd
|
||||
CONTAINER IMAGE CREATED STATE NAME ATTEMPT POD ID POD NAMESPACE
|
||||
f1e289d202ed0 11ad16872a9cf 58 minutes ago Running etcd 0 7b56aab8204ea etcd-cluster1 kube-system
|
||||
```
|
||||
|
||||
## etcd Container Logging
|
||||
|
||||
The logging of the container can contain information on what the problem could be.
|
||||
**K3s**: Etcd runs as an embedded process in the K3s service. Check the service status:
|
||||
|
||||
```bash
|
||||
systemctl status k3s
|
||||
```
|
||||
docker logs etcd
|
||||
|
||||
## etcd Logging
|
||||
|
||||
The logs can contain information on what the problem could be.
|
||||
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl logs $etcdcontainer
|
||||
```
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
journalctl -u k3s | grep -i etcd
|
||||
```
|
||||
| Log | Explanation |
|
||||
|-----|------------------|
|
||||
@@ -46,18 +80,43 @@ The address where etcd is listening depends on the address configuration of the
|
||||
|
||||
Output should contain all the nodes with the `etcd` role and the output should be identical on all nodes.
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
Run the command inside the etcd container.
|
||||
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl member list \
|
||||
--cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt \
|
||||
--key /var/lib/rancher/rke2/server/tls/etcd/server-client.key \
|
||||
--cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec etcd etcdctl member list
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl member list \
|
||||
--cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt \
|
||||
--key /var/lib/rancher/k3s/server/tls/etcd/server-client.key \
|
||||
--cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
1c424074df86e854, started, cluster-node1-f289ac71, https://IP:2380, https://IP:2379, false
|
||||
45c68c44c5a792ff, started, cluster-node2-67e3cf6f, https://IP:2380, https://IP:2379, false
|
||||
7c584f77c5180258, started, cluster-node3-e976bc00, https://IP:2380, https://IP:2379, false
|
||||
```
|
||||
|
||||
### Check Endpoint Status
|
||||
|
||||
The values for `RAFT TERM` should be equal and `RAFT INDEX` should be not be too far apart from each other.
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl endpoint status --write-out table --endpoints=$(crictl exec $etcdcontainer etcdctl member list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec -e ETCDCTL_ENDPOINTS=$(docker exec etcd etcdctl member list | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') etcd etcdctl endpoint status --write-out table
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl endpoint status --write-out table --endpoints=$(etcdctl member list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
@@ -65,17 +124,22 @@ Example output:
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
| ENDPOINT | ID | VERSION | DB SIZE | IS LEADER | RAFT TERM | RAFT INDEX |
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
| https://IP:2379 | 333ef673fc4add56 | 3.5.7 | 24 MB | false | 72 | 66887 |
|
||||
| https://IP:2379 | 5feed52d940ce4cf | 3.5.7 | 24 MB | true | 72 | 66887 |
|
||||
| https://IP:2379 | db6b3bdb559a848d | 3.5.7 | 25 MB | false | 72 | 66887 |
|
||||
| https://IP:2379 | 333ef673fc4add56 | 3.6.7 | 24 MB | false | 72 | 66887 |
|
||||
| https://IP:2379 | 5feed52d940ce4cf | 3.6.7 | 24 MB | true | 72 | 66887 |
|
||||
| https://IP:2379 | db6b3bdb559a848d | 3.6.7 | 25 MB | false | 72 | 66887 |
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
```
|
||||
|
||||
### Check Endpoint Health
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl endpoint health --endpoints=$(crictl exec $etcdcontainer etcdctl member list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec -e ETCDCTL_ENDPOINTS=$(docker exec etcd etcdctl member list | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') etcd etcdctl endpoint health
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl endpoint health --endpoints=$(etcdctl member list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
@@ -84,54 +148,104 @@ https://IP:2379 is healthy: successfully committed proposal: took = 2.113189ms
|
||||
https://IP:2379 is healthy: successfully committed proposal: took = 2.649963ms
|
||||
https://IP:2379 is healthy: successfully committed proposal: took = 2.451201ms
|
||||
```
|
||||
### Check Connectivity on etcd Ports
|
||||
|
||||
### Check Connectivity on Port TCP/2379
|
||||
> In modern versions of Kubernetes, the etcd database (versions 3.5 and newer) introduced significant architectural changes regarding network traffic handling. Previously, etcd permitted standard HTTP REST requests on its primary client port (`2379`). However, to enhance performance and security, etcd 3.5+ strictly enforces the gRPC protocol on this port.<br />
|
||||
If you attempt to use standard HTTP tools like `curl` to test connectivity on port `2379`, the etcd server will automatically terminate the connection or return an error. This behavior often leads administrators to misinterpret the result as a closed port or a node failure.
|
||||
|
||||
Command:
|
||||
Since standard HTTP clients can no longer probe the primary etcd ports, the transport layer must be utilized for network troubleshooting. Using `openssl s_client` instead of `curl` bypasses the gRPC application requirement, allowing the raw TCP and TLS handshake to be tested directly.
|
||||
|
||||
These script isolate the network and security infrastructure from the database application. A successful `Verify return code: 0 (ok)` explicitly confirms four critical infrastructure components:
|
||||
|
||||
* **Network Path:** Routing is functional, and firewalls permit traffic on TCP port `2379` or `2380`.
|
||||
* **Process Availability:** The etcd service is running and actively listening on the designated port.
|
||||
* **Certificate Validity:** The TLS certificates are active, correctly formatted, and have not expired.
|
||||
* **Mutual Authentication (mTLS):** The node successfully authenticates against the cluster's specific Certificate Authority (CA).
|
||||
|
||||
**How these tests differ from the `etcdctl endpoint health` test**:
|
||||
|
||||
If `etcdctl endpoint health` test is failing, run these Connectivity Ports test scripts. If the scripts succeed, your network and certificates are intact, and the issue is likely confined to the etcd database itself. If these scripts fail, the issue is related to a firewall/network restriction, or certificate expiration.
|
||||
|
||||
#### Port TCP/2379
|
||||
|
||||
**RKE2**:
|
||||
```bash
|
||||
for endpoint in $(crictl exec $etcdcontainer etcdctl member list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | cut -d, -f5); do
|
||||
echo "Validating connection to ${endpoint} (Client)";
|
||||
echo | openssl s_client -connect ${endpoint#https://} \
|
||||
-CAfile /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt \
|
||||
-cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt \
|
||||
-key /var/lib/rancher/rke2/server/tls/etcd/server-client.key 2>/dev/null | grep -E 'Verify return code' || echo "Connection Failed/Timeout"
|
||||
done
|
||||
```
|
||||
for endpoint in $(docker exec etcd etcdctl member list | cut -d, -f5); do
|
||||
echo "Validating connection to ${endpoint}/health"
|
||||
docker run --net=host -v $(docker inspect kubelet --format '{{ range .Mounts }}{{ if eq .Destination "/etc/kubernetes" }}{{ .Source }}{{ end }}{{ end }}')/ssl:/etc/kubernetes/ssl:ro appropriate/curl -s -w "\n" --cacert $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_CACERT" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) --cert $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_CERT" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) --key $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_KEY" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) "${endpoint}/health"
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
for endpoint in $(etcdctl member list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | cut -d, -f5); do
|
||||
echo "Validating connection to ${endpoint} (Client)";
|
||||
echo | openssl s_client -connect ${endpoint#https://} \
|
||||
-CAfile /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt \
|
||||
-cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt \
|
||||
-key /var/lib/rancher/k3s/server/tls/etcd/server-client.key 2>/dev/null | grep -E 'Verify return code' || echo "Connection Failed/Timeout"
|
||||
done
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
Validating connection to https://IP:2379/health
|
||||
{"health": "true"}
|
||||
Validating connection to https://IP:2379/health
|
||||
{"health": "true"}
|
||||
Validating connection to https://IP:2379/health
|
||||
{"health": "true"}
|
||||
Validating connection to https://IP:2379/health (Client)
|
||||
Verify return code: 0 (ok)
|
||||
Validating connection to https://IP:2379/health (Client)
|
||||
Verify return code: 0 (ok)
|
||||
Validating connection to https://IP:2379/health (Client)
|
||||
Verify return code: 0 (ok)
|
||||
```
|
||||
|
||||
### Check Connectivity on Port TCP/2380
|
||||
#### Port TCP/2380
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
for endpoint in $(crictl exec $etcdcontainer etcdctl member list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | cut -d, -f4); do
|
||||
echo "Validating connection to ${endpoint} (Peer)";
|
||||
echo | openssl s_client -connect ${endpoint#https://} \
|
||||
-CAfile /var/lib/rancher/rke2/server/tls/etcd/peer-ca.crt \
|
||||
-cert /var/lib/rancher/rke2/server/tls/etcd/peer-server-client.crt \
|
||||
-key /var/lib/rancher/rke2/server/tls/etcd/peer-server-client.key 2>/dev/null | grep -E 'Verify return code' || echo "Connection Failed/Timeout"
|
||||
done
|
||||
```
|
||||
for endpoint in $(docker exec etcd etcdctl member list | cut -d, -f4); do
|
||||
echo "Validating connection to ${endpoint}/version";
|
||||
docker run --net=host -v $(docker inspect kubelet --format '{{ range .Mounts }}{{ if eq .Destination "/etc/kubernetes" }}{{ .Source }}{{ end }}{{ end }}')/ssl:/etc/kubernetes/ssl:ro appropriate/curl --http1.1 -s -w "\n" --cacert $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_CACERT" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) --cert $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_CERT" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) --key $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_KEY" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) "${endpoint}/version"
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
for endpoint in $(etcdctl member list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | cut -d, -f4); do
|
||||
echo "Validating connection to ${endpoint} (Peer)";
|
||||
echo | openssl s_client -connect ${endpoint#https://} \
|
||||
-CAfile /var/lib/rancher/k3s/server/tls/etcd/peer-ca.crt \
|
||||
-cert /var/lib/rancher/k3s/server/tls/etcd/peer-server-client.crt \
|
||||
-key /var/lib/rancher/k3s/server/tls/etcd/peer-server-client.key 2>/dev/null | grep -E 'Verify return code' || echo "Connection Failed/Timeout"
|
||||
done
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
Validating connection to https://IP:2380/version
|
||||
{"etcdserver":"3.5.7","etcdcluster":"3.5.0"}
|
||||
Validating connection to https://IP:2380/version
|
||||
{"etcdserver":"3.5.7","etcdcluster":"3.5.0"}
|
||||
Validating connection to https://IP:2380/version
|
||||
{"etcdserver":"3.5.7","etcdcluster":"3.5.0"}
|
||||
Validating connection to https://IP:2380/version (Peer)
|
||||
Verify return code: 0 (ok)
|
||||
Validating connection to https://IP:2380/version (Peer)
|
||||
Verify return code: 0 (ok)
|
||||
Validating connection to https://IP:2380/version (Peer)
|
||||
Verify return code: 0 (ok)
|
||||
```
|
||||
|
||||
## etcd Alarms
|
||||
|
||||
etcd will trigger alarms, for instance when it runs out of space.
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl alarm list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec etcd etcdctl alarm list
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl alarm list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output when NOSPACE alarm is triggered:
|
||||
@@ -154,10 +268,16 @@ Resolutions:
|
||||
|
||||
### Compact the Keyspace
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
rev=$(crictl exec $etcdcontainer etcdctl endpoint status --write-out json --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | egrep -o '"revision":[0-9]*' | egrep -o '[0-9]*' | head -1)
|
||||
crictl exec $etcdcontainer etcdctl compact "$rev" --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
rev=$(docker exec etcd etcdctl endpoint status --write-out json | egrep -o '"revision":[0-9]*' | egrep -o '[0-9]*')
|
||||
docker exec etcd etcdctl compact "$rev"
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
rev=$(etcdctl endpoint status --write-out json --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | egrep -o '"revision":[0-9]*' | egrep -o '[0-9]*' | head -1)
|
||||
etcdctl compact "$rev" --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
@@ -167,55 +287,39 @@ compacted revision xxx
|
||||
|
||||
### Defrag All etcd Members
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl defrag --endpoints=$(crictl exec $etcdcontainer etcdctl member list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec -e ETCDCTL_ENDPOINTS=$(docker exec etcd etcdctl member list | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') etcd etcdctl defrag
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl defrag --endpoints=$(etcdctl member list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
Finished defragmenting etcd member[https://IP:2379]
|
||||
Finished defragmenting etcd member[https://IP:2379]
|
||||
Finished defragmenting etcd member[https://IP:2379]
|
||||
```
|
||||
|
||||
### Check Endpoint Status
|
||||
|
||||
Command:
|
||||
```
|
||||
docker exec -e ETCDCTL_ENDPOINTS=$(docker exec etcd etcdctl member list | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') etcd etcdctl endpoint status --write-out table
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
| ENDPOINT | ID | VERSION | DB SIZE | IS LEADER | RAFT TERM | RAFT INDEX |
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
| https://IP:2379 | e973e4419737125 | 3.5.7 | 553 kB | false | 32 | 2449410 |
|
||||
| https://IP:2379 | 4a509c997b26c206 | 3.5.7 | 553 kB | false | 32 | 2449410 |
|
||||
| https://IP:2379 | b217e736575e9dd3 | 3.5.7 | 553 kB | true | 32 | 2449410 |
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
Finished defragmenting etcd member[https://IP:2379]. took xx.xxxxxxms
|
||||
Finished defragmenting etcd member[https://IP:2379]. took xx.xxxxxxms
|
||||
Finished defragmenting etcd member[https://IP:2379]. took xx.xxxxxxms
|
||||
```
|
||||
|
||||
### Disarm Alarm
|
||||
|
||||
After verifying that the DB size went down after compaction and defragmenting, the alarm needs to be disarmed for etcd to allow writes again.
|
||||
|
||||
Command:
|
||||
```
|
||||
docker exec etcd etcdctl alarm list
|
||||
docker exec etcd etcdctl alarm disarm
|
||||
docker exec etcd etcdctl alarm list
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl alarm list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
crictl exec $etcdcontainer etcdctl alarm disarm --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
crictl exec $etcdcontainer etcdctl alarm list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
docker exec etcd etcdctl alarm list
|
||||
memberID:x alarm:NOSPACE
|
||||
memberID:x alarm:NOSPACE
|
||||
memberID:x alarm:NOSPACE
|
||||
docker exec etcd etcdctl alarm disarm
|
||||
docker exec etcd etcdctl alarm list
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl alarm list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
etcdctl alarm disarm --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
etcdctl alarm list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
## Configure Log Level
|
||||
@@ -228,7 +332,7 @@ You can no longer dynamically change the log level in etcd v3.5 or later.
|
||||
|
||||
### etcd v3.5 And Later
|
||||
|
||||
To configure the log level for etcd, edit the cluster YAML:
|
||||
To configure the log level for etcd, edit the cluster configuration YAML:
|
||||
|
||||
```
|
||||
services:
|
||||
@@ -237,20 +341,7 @@ services:
|
||||
log-level: "debug"
|
||||
```
|
||||
|
||||
### etcd v3.4 And Earlier
|
||||
|
||||
In earlier etcd versions, you can use the API to dynamically change the log level. Configure debug logging using the commands below:
|
||||
|
||||
```
|
||||
docker run --net=host -v $(docker inspect kubelet --format '{{ range .Mounts }}{{ if eq .Destination "/etc/kubernetes" }}{{ .Source }}{{ end }}{{ end }}')/ssl:/etc/kubernetes/ssl:ro appropriate/curl -s -XPUT -d '{"Level":"DEBUG"}' --cacert $(docker exec etcd printenv ETCDCTL_CACERT) --cert $(docker exec etcd printenv ETCDCTL_CERT) --key $(docker exec etcd printenv ETCDCTL_KEY) $(docker exec etcd printenv ETCDCTL_ENDPOINTS)/config/local/log
|
||||
```
|
||||
|
||||
To reset the log level back to the default (`INFO`), you can use the following command.
|
||||
|
||||
Command:
|
||||
```
|
||||
docker run --net=host -v $(docker inspect kubelet --format '{{ range .Mounts }}{{ if eq .Destination "/etc/kubernetes" }}{{ .Source }}{{ end }}{{ end }}')/ssl:/etc/kubernetes/ssl:ro appropriate/curl -s -XPUT -d '{"Level":"INFO"}' --cacert $(docker exec etcd printenv ETCDCTL_CACERT) --cert $(docker exec etcd printenv ETCDCTL_CERT) --key $(docker exec etcd printenv ETCDCTL_KEY) $(docker exec etcd printenv ETCDCTL_ENDPOINTS)/config/local/log
|
||||
```
|
||||
After modifying the configuration, restart the service (`systemctl restart rke2-server` or `systemctl restart k3s`) if you are configuring a stand-alone cluster.
|
||||
|
||||
## etcd Content
|
||||
|
||||
@@ -258,24 +349,40 @@ If you want to investigate the contents of your etcd, you can either watch strea
|
||||
|
||||
### Watch Streaming Events
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl watch --prefix /registry --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec etcd etcdctl watch --prefix /registry
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl watch --prefix /registry --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
If you only want to see the affected keys (and not the binary data), you can append `| grep -a ^/registry` to the command to filter for keys only.
|
||||
|
||||
### Query etcd Directly
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl get /registry --prefix=true --keys-only --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec etcd etcdctl get /registry --prefix=true --keys-only
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl get /registry --prefix=true --keys-only --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
You can process the data to get a summary of count per key, using the command below:
|
||||
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl get /registry --prefix=true --keys-only --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | grep -v ^$ | awk -F'/' '{ if ($3 ~ /cattle.io/) {h[$3"/"$4]++} else { h[$3]++ }} END { for(k in h) print h[k], k }' | sort -nr
|
||||
```
|
||||
docker exec etcd etcdctl get /registry --prefix=true --keys-only | grep -v ^$ | awk -F'/' '{ if ($3 ~ /cattle.io/) {h[$3"/"$4]++} else { h[$3]++ }} END { for(k in h) print h[k], k }' | sort -nr
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl get /registry --prefix=true --keys-only --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | grep -v ^$ | awk -F'/' '{ if ($3 ~ /cattle.io/) {h[$3"/"$4]++} else { h[$3]++ }} END { for(k in h) print h[k], k }' | sort -nr
|
||||
```
|
||||
|
||||
## Replacing Unhealthy etcd Nodes
|
||||
|
||||
+68
-14
@@ -8,31 +8,85 @@ title: Troubleshooting Worker Nodes and Generic Components
|
||||
|
||||
This section applies to every node as it includes components that run on nodes with any role.
|
||||
|
||||
## Check if the Containers are Running
|
||||
## Prerequisites
|
||||
|
||||
There are two specific containers launched on nodes with the `worker` role:
|
||||
Since RKE2 and K3s utilize `containerd` as the container runtime, `crictl` serves as the primary tool for container management, replacing the Docker CLI. To allow `crictl` to communicate with `containerd`, you must configure your environment by exporting the following variables:
|
||||
|
||||
* kubelet
|
||||
* kube-proxy
|
||||
|
||||
The containers should have status `Up`. The duration shown after `Up` is the time the container has been running.
|
||||
### RKE2
|
||||
|
||||
```bash
|
||||
export PATH=$PATH:/var/lib/rancher/rke2/bin/
|
||||
export CRI_CONFIG_FILE=/var/lib/rancher/rke2/agent/etc/crictl.yaml
|
||||
```
|
||||
docker ps -a -f=name='kubelet|kube-proxy'
|
||||
|
||||
### K3s
|
||||
|
||||
```bash
|
||||
export PATH=$PATH:/usr/local/bin
|
||||
export CRI_CONFIG_FILE=/var/lib/rancher/k3s/agent/etc/crictl.yaml
|
||||
```
|
||||
|
||||
## Check if the Components are Running
|
||||
|
||||
There are two specific components launched on nodes with the `worker` role:
|
||||
|
||||
* `kubelet`
|
||||
* `kube-proxy`
|
||||
|
||||
### RKE2
|
||||
|
||||
The `kubelet` runs natively as part of the `rke2-agent` (or `rke2-server`) systemd process, while `kube-proxy` runs as a Static Pod managed by `containerd`.
|
||||
|
||||
Check the status of the `kubelet` via the agent service:
|
||||
```bash
|
||||
systemctl status rke2-agent
|
||||
```
|
||||
:::note
|
||||
|
||||
If you are checking a controlplane node, use `systemctl status rke2-server` instead.
|
||||
|
||||
:::
|
||||
|
||||
Check the status of `kube-proxy` using `crictl`:
|
||||
```bash
|
||||
crictl ps --name kube-proxy
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES
|
||||
158d0dcc33a5 rancher/hyperkube:v1.11.5-rancher1 "/opt/rke-tools/en..." 3 hours ago Up 3 hours kube-proxy
|
||||
a30717ecfb55 rancher/hyperkube:v1.11.5-rancher1 "/opt/rke-tools/en..." 3 hours ago Up 3 hours kubelet
|
||||
CONTAINER IMAGE CREATED STATE NAME ATTEMPT POD ID
|
||||
26c7159abbcc rancher/hardened-kubernetes:v1.28.8-rke2r1-build20240404 3 hours ago Running kube-proxy 0 1a2b3c4d5e6f7
|
||||
```
|
||||
|
||||
## Container Logging
|
||||
### K3s
|
||||
|
||||
The logging of the containers can contain information on what the problem could be.
|
||||
Both `kubelet` and `kube-proxy` run as embedded processes inside the `k3s-agent` (or `k3s` server) systemd service. There are no separate containers for them.
|
||||
|
||||
Check their status by checking the K3s service:
|
||||
```bash
|
||||
systemctl status k3s-agent
|
||||
```
|
||||
docker logs kubelet
|
||||
docker logs kube-proxy
|
||||
:::note
|
||||
If you are checking a controlplane node, use `systemctl status k3s` instead.
|
||||
:::
|
||||
|
||||
## Component Logging
|
||||
|
||||
The logging of the components can contain information on what the problem could be.
|
||||
|
||||
### RKE2
|
||||
|
||||
```bash
|
||||
# kubelet logs are part of the systemd service
|
||||
journalctl -u rke2-agent -f | grep -i "kubelet"
|
||||
|
||||
# kube-proxy logs are retrieved from containerd
|
||||
crictl logs $(crictl ps --name kube-proxy -q)
|
||||
```
|
||||
|
||||
### K3s
|
||||
|
||||
```bash
|
||||
# Both components log to the systemd service
|
||||
journalctl -u k3s-agent -f | grep -iE "kubelet|kube-proxy"
|
||||
```
|
||||
|
||||
+85
-14
@@ -78,66 +78,137 @@ kubectl -n kube-system get endpoints kube-scheduler -o jsonpath='{.metadata.anno
|
||||
|
||||
## Ingress Controller
|
||||
|
||||
The default Ingress Controller is Traefik and is deployed as a DaemonSet in the `traefik` namespace. The pods are only scheduled to nodes with the `worker` role.
|
||||
The default Ingress Controller is Traefik and is deployed as a DaemonSet in the `kube-system` namespace. The pods are only scheduled to nodes with the `worker` role.
|
||||
|
||||
Check if the pods are running on all nodes:
|
||||
|
||||
```
|
||||
kubectl -n traefik get pods -o wide
|
||||
kubectl -n kube-system get pods -o wide
|
||||
```
|
||||
|
||||
Example output:
|
||||
Example RKE2 output:
|
||||
|
||||
```
|
||||
kubectl -n traefik get pods -o wide
|
||||
kubectl -n kube-system get pods -o wide
|
||||
NAME READY STATUS RESTARTS AGE IP NODE
|
||||
default-http-backend-797c5bc547-kwwlq 1/1 Running 0 17m x.x.x.x worker-1
|
||||
traefik-4qd64 1/1 Running 0 14m x.x.x.x worker-1
|
||||
traefik-8wxhm 1/1 Running 0 13m x.x.x.x worker-0
|
||||
local-path-provisioner-xxxxxxxxxx-xxxxx 1/1 Running 0 17m x.x.x.x worker-1
|
||||
rke2-traefik-xxxxxxxxxx-xxxxx 1/1 Running 0 14m x.x.x.x worker-1
|
||||
svclb-rke2-traefik-xxxxxxxx-xxxxx 1/1 Running 0 13m x.x.x.x worker-0
|
||||
...
|
||||
```
|
||||
|
||||
Example K3s output:
|
||||
|
||||
```
|
||||
kubectl -n kube-system get pods -o wide
|
||||
NAME READY STATUS RESTARTS AGE IP NODE
|
||||
local-path-provisioner-xxxxxxxxxx-xxxxx 1/1 Running 0 17m x.x.x.x worker-1
|
||||
traefik-xxxxxxxxxx-xxxxx 1/1 Running 0 14m x.x.x.x worker-1
|
||||
svclb-traefik-xxxxxxxx-xxxxx 1/1 Running 0 13m x.x.x.x worker-0
|
||||
...
|
||||
```
|
||||
|
||||
If a pod is unable to run (Status is not **Running**, Ready status is not showing `1/1` or you see a high count of Restarts), check the pod details, logs and namespace events.
|
||||
|
||||
### Pod details
|
||||
|
||||
RKE2 example:
|
||||
|
||||
```
|
||||
kubectl -n traefik describe pods -l app=traefik
|
||||
kubectl -n kube-system describe pods -l app.kubernetes.io/name=rke2-traefik
|
||||
```
|
||||
|
||||
K3s example:
|
||||
|
||||
```
|
||||
kubectl -n kube-system describe pods -l app.kubernetes.io/name=traefik
|
||||
```
|
||||
|
||||
### Pod container logs
|
||||
|
||||
The below command can show the logs of all the pods labeled "app=traefik", but it will display only 10 lines of log because of the restrictions of the `kubectl logs` command. Refer to `--tail` of `kubectl logs -h` for more information.
|
||||
The below command can show the logs of all the pods labeled "app.kubernetes.io/name=rke2-traefik" if using RKE2 or "app.kubernetes.io/name=traefik" if using K3s, but it will display only 10 lines of log because of the restrictions of the `kubectl logs` command. Refer to `--tail` of `kubectl logs -h` for more information.
|
||||
|
||||
RKE2 example:
|
||||
|
||||
```
|
||||
kubectl -n traefik logs -l app=traefik
|
||||
kubectl -n kube-system logs -l app.kubernetes.io/name=rke2-traefik
|
||||
```
|
||||
|
||||
K3s example:
|
||||
|
||||
```
|
||||
kubectl -n kube-system logs -l app.kubernetes.io/name=traefik
|
||||
```
|
||||
|
||||
If the full log is needed, specify the pod name in the trailing command:
|
||||
|
||||
```
|
||||
kubectl -n traefik logs <pod name>
|
||||
kubectl -n kube-system logs <pod name>
|
||||
```
|
||||
|
||||
### Namespace events
|
||||
|
||||
```
|
||||
kubectl -n traefik get events
|
||||
kubectl -n kube-system get events
|
||||
```
|
||||
|
||||
### Debug logging
|
||||
|
||||
To enable debug logging:
|
||||
|
||||
RKE2 example:
|
||||
|
||||
```
|
||||
kubectl -n traefik patch ds traefik --type='json' -p='[{"op": "add", "path": "/spec/template/spec/containers/0/args/-", "value": "--v=5"}]'
|
||||
cat <<EOF | kubectl apply -f -
|
||||
apiVersion: helm.cattle.io/v1
|
||||
kind: HelmChartConfig
|
||||
metadata:
|
||||
name: rke2-traefik
|
||||
namespace: kube-system
|
||||
spec:
|
||||
valuesContent: |-
|
||||
additionalArguments:
|
||||
- "--log.level=DEBUG"
|
||||
EOF
|
||||
```
|
||||
|
||||
K3s example:
|
||||
|
||||
```
|
||||
cat <<EOF | kubectl apply -f -
|
||||
apiVersion: helm.cattle.io/v1
|
||||
kind: HelmChartConfig
|
||||
metadata:
|
||||
name: traefik
|
||||
namespace: kube-system
|
||||
spec:
|
||||
valuesContent: |-
|
||||
logs:
|
||||
general:
|
||||
level: "DEBUG"
|
||||
EOF
|
||||
```
|
||||
|
||||
### Check configuration
|
||||
|
||||
Retrieve generated configuration in each pod:
|
||||
|
||||
RKE2 example for manual Traefik configuration file check:
|
||||
|
||||
```
|
||||
kubectl -n traefik get pods -l app=traefik --no-headers -o custom-columns=.NAME:.metadata.name | while read pod; do kubectl -n traefik exec $pod -- cat /etc/nginx/nginx.conf; done
|
||||
kubectl exec -n kube-system pod/traefik-xxxxxxxxx-xxxxx -- cat /var/lib/rancher/rke2/server/manifests/rke2-traefik-config.yaml
|
||||
```
|
||||
|
||||
K3s example for manual Traefik configuration file check:
|
||||
|
||||
```
|
||||
kubectl exec -n kube-system pod/traefik-xxxxxxxxx-xxxxx -- cat /var/lib/rancher/k3s/server/manifests/k3s-traefik-config.yaml
|
||||
```
|
||||
|
||||
RKE2/K3s example for Traefik CLI argument configuration check:
|
||||
|
||||
```
|
||||
kubectl get pod traefik-xxxxxxxxx-xxxxx -n kube-system -o jsonpath='{.spec.containers[0].args}'
|
||||
```
|
||||
|
||||
## Rancher agents
|
||||
|
||||
@@ -15,6 +15,14 @@ Make sure you configured the correct kubeconfig (for example, `export KUBECONFIG
|
||||
Double check if all the [required ports](../../how-to-guides/new-user-guides/kubernetes-clusters-in-rancher-setup/node-requirements-for-rancher-managed-clusters.md#networking-requirements) are opened in your (host) firewall. The overlay network uses UDP in comparison to all other required ports which are TCP.
|
||||
|
||||
|
||||
## Check if your downstream node can communicate to Rancher Manager
|
||||
|
||||
Rancher components with HTTP endpoints generally contain a `ping` liveness probe which you can use to test connectivity. Replace the `$RANCHER_URL` as appropriate and run the following from a node to check that it has connectivity to Rancher Manager's servers in the `local` cluster. If successful, it should return `pong`.
|
||||
|
||||
```
|
||||
curl -k https://$RANCHER_URL/ping
|
||||
```
|
||||
|
||||
## Check if Overlay Network is Functioning Correctly
|
||||
|
||||
The pod can be scheduled to any of the hosts you used for your cluster, but that means that the NGINX ingress controller needs to be able to route the request from `NODE_1` to `NODE_2`. This happens over the overlay network. If the overlay network is not functioning, you will experience intermittent TCP/HTTP connection failures due to the NGINX ingress controller not being able to route to the pod.
|
||||
|
||||
@@ -12,6 +12,7 @@ Rancher will publish deprecated features as part of the [release notes](https://
|
||||
|
||||
| Patch Version | Release Date |
|
||||
|---------------|---------------|
|
||||
| [2.13.6](https://github.com/rancher/rancher/releases/tag/v2.13.6) | May 27, 2026 |
|
||||
| [2.13.5](https://github.com/rancher/rancher/releases/tag/v2.13.5) | April 30, 2026 |
|
||||
| [2.13.4](https://github.com/rancher/rancher/releases/tag/v2.13.4) | March 25, 2026 |
|
||||
| [2.13.3](https://github.com/rancher/rancher/releases/tag/v2.13.3) | February 25, 2026 |
|
||||
|
||||
+34
-9
@@ -26,6 +26,31 @@ Review the list of known issues for each Rancher version, which can be found in
|
||||
|
||||
Note that upgrades _to_ or _from_ any chart in the [rancher-alpha repository](../resources/choose-a-rancher-version.md#helm-chart-repositories) aren't supported.
|
||||
|
||||
### Upgrade Path
|
||||
|
||||
:::important
|
||||
|
||||
**Important:** The only tested and supported Rancher upgrade path between minor versions (e.g. v2.12.x to v2.13.x) is to upgrade from the latest available patch version of your current running minor release to the latest available patch version of the next minor release.
|
||||
|
||||
Before initiating a minor version upgrade, verify that you are running the most recent patch release of your current version.
|
||||
|
||||
You can query the available chart versions with the Helm CLI:
|
||||
|
||||
1. Update your local Helm repo cache.
|
||||
|
||||
```
|
||||
helm repo update
|
||||
```
|
||||
|
||||
1. Search for available versions in your [specific repository](../resources/choose-a-rancher-version.md#helm-chart-repositories) (e.g., rancher-stable):
|
||||
|
||||
```
|
||||
helm search repo rancher-<CHART_REPO>/rancher --versions
|
||||
```
|
||||
|
||||
If your installation is not on the latest patch version of the current minor release, you must upgrade to that version before proceeding to the next minor version.
|
||||
:::
|
||||
|
||||
### Helm Version
|
||||
|
||||
:::important
|
||||
@@ -170,17 +195,17 @@ There will be more values that are listed with this command. This is just an exa
|
||||
|
||||
:::
|
||||
|
||||
:::tip
|
||||
:::tip
|
||||
|
||||
Your deployment name may vary; for example, if you're deploying Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
Your deployment name may vary; for example, if you're deploying Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
```
|
||||
helm get values rancher-stable -n cattle-system
|
||||
|
||||
hostname: rancher.my.org
|
||||
```
|
||||
|
||||
:::
|
||||
:::
|
||||
|
||||
If you are upgrading cert-manager to the latest version from v1.5 or below, follow the [cert-manager upgrade docs](../resources/upgrade-cert-manager.md#option-c-upgrade-cert-manager-from-versions-15-and-below) to learn how to upgrade cert-manager without needing to perform an uninstall or reinstall of Rancher. Otherwise, follow the [steps to upgrade Rancher](#steps-to-upgrade-rancher) below.
|
||||
|
||||
@@ -203,17 +228,17 @@ The above is an example, there may be more values from the previous step that ne
|
||||
|
||||
:::
|
||||
|
||||
:::tip
|
||||
:::tip
|
||||
|
||||
If you deploy Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
If you deploy Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
```
|
||||
helm upgrade rancher-stable rancher-<CHART_REPO>/rancher \
|
||||
--namespace cattle-system \
|
||||
--set hostname=rancher.my.org
|
||||
```
|
||||
|
||||
:::
|
||||
:::
|
||||
|
||||
Alternatively, it's possible to export the current values to a file and reference that file during upgrade. For example, to only change the Rancher version:
|
||||
|
||||
@@ -223,7 +248,7 @@ Alternatively, it's possible to export the current values to a file and referenc
|
||||
```
|
||||
1. Update only the Rancher version:
|
||||
|
||||
|
||||
|
||||
```
|
||||
helm upgrade rancher rancher-<CHART_REPO>/rancher \
|
||||
--namespace cattle-system \
|
||||
|
||||
+2
-2
@@ -33,7 +33,7 @@ To use a premade dashboard, go to [https://grafana.com/grafana/dashboards](https
|
||||
To use your own dashboard:
|
||||
|
||||
1. Click on the link to open Grafana. On the cluster detail page, click **Monitoring**.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin/prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
:::note
|
||||
|
||||
@@ -113,7 +113,7 @@ Note that the RBAC roles exposed by the Monitoring chart to add Grafana Dashboar
|
||||
1. On the **Clusters** page, go to the cluster where you want to configure the Grafana namespace and click **Explore**.
|
||||
1. In the left navigation bar, click **Monitoring**.
|
||||
1. Click **Grafana**.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin/prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
:::note
|
||||
|
||||
|
||||
+1
-1
@@ -20,7 +20,7 @@ To see the links to the external monitoring UIs, including Grafana dashboards, y
|
||||
1. In the left navigation menu, click **Monitoring.**
|
||||
1. Click **Grafana.** The Grafana dashboard should open in a new tab.
|
||||
1. Go to the log in icon in the lower left corner and click **Sign In.**
|
||||
1. Log in to Grafana. The default Admin username and password for the Grafana instance is `admin/prom-operator`. (Regardless of who has the password, cluster administrator permission in Rancher is still required access the Grafana instance.) Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. (Regardless of who has the password, cluster administrator permission in Rancher is still required access the Grafana instance.) Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
|
||||
### Getting the PromQL Query Powering a Grafana Panel
|
||||
|
||||
+8
-2
@@ -31,11 +31,17 @@ _Cluster roles_ are roles that you can assign to users, granting them access to
|
||||
|
||||
- **Cluster Owner:**
|
||||
|
||||
These users have full control over the cluster and all resources in it.
|
||||
These users have full control over the cluster and all resources in it.
|
||||
|
||||
- **Cluster Member:**
|
||||
|
||||
These users can view most cluster level resources and create new projects.
|
||||
These users can view most cluster level resources and create new projects.
|
||||
|
||||
:::warning
|
||||
|
||||
When a Cluster Member creates a project, the user is automatically assigned [Project Owner privileges](#project-roles). This grants them comprehensive control over the project and its associated resources, including permissions to deploy workloads. Without enforced [Pod Security Standards (PSS) and Pod Security Admission (PSA)](../pod-security-standards.md), a Cluster Member is able to execute privileged containers in the cluster.
|
||||
|
||||
:::
|
||||
|
||||
#### Custom Cluster Roles
|
||||
|
||||
|
||||
+1
-1
@@ -234,7 +234,7 @@ docker stop <original-rancher-container>
|
||||
|
||||
:::note
|
||||
|
||||
If you wish to keep the original Rancher environment running, you can also restart the cattle-cluster-agent pods on each cluster connected to your Rancher environment.
|
||||
If clusters do not automatically reconnect to the new environment after you have redirected traffic, for example if there is a delay in scaling down the original Rancher instance, you can also restart the cattle-cluster-agent pods on each cluster connected to your Rancher environment.
|
||||
|
||||
```bash
|
||||
kubectl rollout restart deployment cattle-cluster-agent -n cattle-system
|
||||
|
||||
+1
-1
@@ -24,7 +24,7 @@ The following steps can also be performed using the `kubectl` command line tool.
|
||||
:::
|
||||
|
||||
1. Click **☰ > Cluster Management**.
|
||||
1. Choose the cluster you want to provide vSphere storage to and click **Exlpore**.
|
||||
1. Choose the cluster you want to provide vSphere storage to and click **Explore**.
|
||||
1. In the left navigation bar, select **Storage > StorageClasses**.
|
||||
1. Click **Create**.
|
||||
3. Enter a **Name** for the StorageClass.
|
||||
|
||||
+1
@@ -19,6 +19,7 @@ In order to deploy and run the adapter successfully, you need to ensure its vers
|
||||
|
||||
| Rancher Version | Adapter Version |
|
||||
|-----------------|------------------|
|
||||
| v2.13.6 | 108.0.0+up8.0.0 |
|
||||
| v2.13.5 | 108.0.0+up8.0.0 |
|
||||
| v2.13.4 | 108.0.0+up8.0.0 |
|
||||
| v2.13.3 | 108.0.0+up8.0.0 |
|
||||
|
||||
@@ -116,6 +116,19 @@ By default, Rancher collects logs for control plane components and node componen
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Resource Exhaustion of `inotify` Watchers and File Descriptors
|
||||
|
||||
When enabling the **Logging** app on Linux systems that heavily monitor the filesystem, you may encounter `Too many open files` or `CrashLoopBackOff` failures related to applications that leverage `inotify` to watch for file changes.
|
||||
|
||||
This happens because the Linux kernel caps the number of files a user can open and the number of directory paths a subsystem can watch simultaneously. To resolve this, you must explicitly increase your `inotify` system limits.
|
||||
|
||||
Below are example commands an admin user can run to increase system limits for `inotify` user instances and watches:
|
||||
|
||||
```shell
|
||||
sysctl -w fs.inotify.max_user_instances=8192
|
||||
sysctl -w fs.inotify.max_user_watches=524288
|
||||
```
|
||||
|
||||
### The Logging Buffer Overloads Pods
|
||||
|
||||
Depending on your configuration, the default buffer size may be too large and cause pod failures. One way to reduce the load is to lower the logger's flush interval. This prevents logs from overfilling the buffer. You can also add more flush threads to handle moments when many logs are attempting to fill the buffer at once.
|
||||
|
||||
@@ -14,64 +14,74 @@ Examples of built-in Rancher extensions are Fleet, Explorer, and Harvester. Exam
|
||||
|
||||
## Prerequisites
|
||||
|
||||
> You must log in as an admin in order to view and interact with the extensions management page.
|
||||
> 1. You must log in as an administrator to view and interact with the extensions management page.
|
||||
> 2. You must enable extension support.
|
||||
|
||||
## Enabling Extension Support in Rancher
|
||||
|
||||
Rancher v2.9.0 and later includes extension support.
|
||||
|
||||
You can confirm if extension support is enabled by checking the `uiextension` feature flag. Any changes to this feature flag cause the Rancher pod to restart.
|
||||
|
||||
When you enable extension support for the first time, it creates resources, such as the CRDs, so that UI Extensions can work. When extension support is disabled, it disables the endpoints and does not cache any files. However, it does not remove any CRs or delete any extensions that were installed before. If re-enabled, it exposes the required endpoints again and creates CRDs as needed. The extensions that were already installed load after the Rancher pod restarts.
|
||||
|
||||
### Caching Extension Files
|
||||
|
||||
By default, Rancher caches every extension file in the file system. You can change that behavior by setting `plugin.noCache` to `true`.
|
||||
|
||||
Rancher does have a cached file size limit of 30MB. If an extension has a file bigger than that, the cache is disabled and `plugin.noCache` is set to `true`, regardless of user input.
|
||||
|
||||
|
||||
## Installing Extensions
|
||||
|
||||
1. Click **☰ > Extensions** under **Configuration**.
|
||||
|
||||
2. If not already installed in **Apps**, you must enable the extension operator by clicking the **Enable** button.
|
||||
2. On the **Extensions** page, select the **Available** tab to choose the extensions that you want to install.
|
||||
|
||||
- Click **OK** to add the Rancher extension repository if your installation is not air-gapped. Otherwise, uncheck the box to do so and click **OK**.
|
||||
3. If no extensions are listed as available, you can manually add the repos:
|
||||
|
||||

|
||||
3.1. On the upper right, click **⋮ > Manage Repositories > Create**.
|
||||
|
||||
3. On the **Extensions** page, click on the **Available** tab to select which extensions you want to install.
|
||||
3.2. Add the desired repo name, making sure to also specify the Git repo URL and the Git branch.
|
||||
|
||||
4. If no extensions are showing as available, you may manually add repos as follows:
|
||||
3.3. Click **Create** in the lower right again to complete.
|
||||
|
||||
4.1. On the upper right of screen, click on **⋮ > Manage Repositories > Create**.
|
||||

|
||||
|
||||
4.2. Add the desired repo name, making sure to also specify the Git Repo URL and the Git Branch.
|
||||
4. Under the **Available** tab, click **Install** on the desired extension and version, as in the example below. You can also update your extension from this screen. The button to **Update** appears on the extension card if an update is available.
|
||||
|
||||
4.3. Click **Create** in the lower right again to complete.
|
||||

|
||||
|
||||

|
||||
5. Click **Reload** after your extension successfully installs to check its status. Updates to the UI aren't visible until you reload the page.
|
||||
|
||||
5. Under the **Available** tab, click **Install** on the desired extension and version as in the example below. You can also update your extension from this screen, as the button to **Update** will appear on the extension if one is available.
|
||||
|
||||

|
||||
|
||||
6. Click the **Reload** page button that will appear after your extension successfully installs. Note that a logged-in user who has just installed an extension will not see a change to the UI **unless** they reload the page.
|
||||
|
||||

|
||||

|
||||
|
||||
## Updating and Upgrading Extensions
|
||||
|
||||
1. Click **☰ > Extensions** under **Configuration**.
|
||||
1. Select the **Updates** tab.
|
||||
1. Click **Update**.
|
||||
2. Select the **Updates** tab.
|
||||
3. Click **Update**.
|
||||
|
||||
If there is a new version of the extension, there will also be an **Update** button visible on the associated card for the extension in the **Available** tab.
|
||||
|
||||
## Deleting Extensions
|
||||
|
||||
1. Click **☰**, then click on the name of your local cluster.
|
||||
1. From the sidebar, select **Apps > Installed Apps**.
|
||||
1. Find the name of the chart you want to delete and select the checkbox next to it.
|
||||
1. Click **Delete**.
|
||||
2. From the sidebar, select **Apps > Installed Apps**.
|
||||
3. Find the name of the chart you want to delete and select the checkbox next to it.
|
||||
4. Click **Delete**.
|
||||
|
||||
## Deleting Extension Repositories
|
||||
|
||||
1. Click **☰ > Extensions** under **Configuration**.
|
||||
1. On the top right, click **⋮ > Manage Repositories**.
|
||||
1. Find the name of the extension repository you want to delete. Select the checkbox next to the repository name, then click **Delete**.
|
||||
2. On the top right, click **⋮ > Manage Repositories**.
|
||||
3. Find the name of the extension repository you want to delete. Select the checkbox next to the repository name, then click **Delete**.
|
||||
|
||||
## Deleting Extension Repository Container Images
|
||||
|
||||
1. Click **☰**, then select **Extensions**, under **Configuration**.
|
||||
1. On the top right, click **⋮ > Manage Extension Catalogs**.
|
||||
1. Find the name of the container image you want to delete, then click **⋮ > Uninstall**.
|
||||
2. On the top right, click **⋮ > Manage Extension Catalogs**.
|
||||
3. Find the name of the container image you want to delete, then click **⋮ > Uninstall**.
|
||||
|
||||
## Uninstalling Extensions
|
||||
|
||||
@@ -79,11 +89,11 @@ There are two ways to uninstall or disable an extension:
|
||||
|
||||
1. Under the **Installed** tab, click the **Uninstall** button on the extension you wish to remove.
|
||||
|
||||

|
||||

|
||||
|
||||
1. On the extensions management page, click **⋮ > Disable Extension Support**. This will disable all installed extensions.
|
||||
2. On the extensions management page, click **⋮ > Disable Extension Support**. This will disable all installed extensions.
|
||||
|
||||

|
||||

|
||||
|
||||
:::caution
|
||||
|
||||
@@ -91,6 +101,12 @@ You must reload the page after disabling extensions or display issues may occur.
|
||||
|
||||
:::
|
||||
|
||||
## Enabling Unauthenticated Access to an Extension
|
||||
|
||||
In Rancher v2.9.0 and later, you can allow unauthenticated access to an extension. You may want to enable unauthenticated access if the extension enables a new locale or adds custom branding. By default, all extensions require user authentication to load.
|
||||
|
||||
To enable unauthenticated access to an extension, set `plugin.noAuth` to `true` in the CR used by the extension.
|
||||
|
||||
## Developing Extensions
|
||||
|
||||
To learn how to develop your own extensions, refer to the official [Getting Started](https://rancher.github.io/dashboard/extensions/extensions-getting-started) guide.
|
||||
@@ -173,8 +189,8 @@ After you successfully set up these resources, you can install the extensions fr
|
||||
1. Click **☰**, then select **Extensions**, under **Configuration**.
|
||||
1. On the top right, click **⋮ > Manage Extension Catalogs**.
|
||||
1. Select the **Import Extension Catalog** button.
|
||||
1. Enter the image address in the **Catalog Image Reference** field.
|
||||
* **(Optional)** If the container image is private, select the secret you just created from the **Pull Secrets** drop-down menu.
|
||||
1. Enter the image address in the **Catalog Image Reference** field.
|
||||
- **(Optional)** If the container image is private, select the secret you just created from the **Pull Secrets** drop-down menu.
|
||||
1. Click **Load**. The extension will now be **Pending**.
|
||||
1. Return to the **Extensions** page.
|
||||
1. Select the **Available** tab, and click **Reload** to make sure that the list of extensions is up to date.
|
||||
@@ -191,7 +207,7 @@ After you mirror the latest changes, follow these steps:
|
||||
1. Click **☰ > Local**.
|
||||
1. From the sidebar, select **Workloads > Deployments**.
|
||||
1. From the namespaces dropdown menu, select **cattle-ui-plugin-system**.
|
||||
1. Find the **cattle-ui-plugin-system** namespace.
|
||||
1. Find the **cattle-ui-plugin-system** namespace.
|
||||
1. Select the `ui-plugin-catalog` deployment.
|
||||
1. Click **⋮ > Edit config**.
|
||||
1. Update the **Container Image** field within the deployment's container with the latest image.
|
||||
|
||||
@@ -20,7 +20,8 @@ Each Rancher version is designed to be compatible with a single version of the w
|
||||
|
||||
| Rancher Version | Webhook Version | Availability in Prime | Availability in Community |
|
||||
|-----------------|-----------------|-----------------------|---------------------------|
|
||||
| v2.13.5 | v0.9.3 | ✓ | ✗ |
|
||||
| v2.13.6 | v0.9.5 | ✓ | ✗ |
|
||||
| v2.13.5 | v0.9.4 | ✓ | ✗ |
|
||||
| v2.13.4 | v0.9.3 | ✓ | ✗ |
|
||||
| v2.13.3 | v0.9.3 | ✓ | ✓ |
|
||||
| v2.13.2 | v0.9.2 | ✓ | ✓ |
|
||||
|
||||
+209
-102
@@ -6,29 +6,63 @@ title: Troubleshooting etcd Nodes
|
||||
<link rel="canonical" href="https://ranchermanager.docs.rancher.com/troubleshooting/kubernetes-components/troubleshooting-etcd-nodes"/>
|
||||
</head>
|
||||
|
||||
This section contains commands and tips for troubleshooting nodes with the `etcd` role.
|
||||
This section contains commands and tips for troubleshooting nodes with the `etcd` role in RKE2 and K3s clusters.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
As RKE2 and K3s rely on `containerd` as the container runtime, `crictl` replaces Docker for container management. Before proceeding with the troubleshooting commands, configure your environment by exporting the following variables:
|
||||
|
||||
### RKE2
|
||||
|
||||
```bash
|
||||
export PATH=$PATH:/var/lib/rancher/rke2/bin/
|
||||
export CRI_CONFIG_FILE=/var/lib/rancher/rke2/agent/etc/crictl.yaml
|
||||
etcdcontainer=$(crictl ps --name etcd --quiet)
|
||||
```
|
||||
|
||||
### K3s
|
||||
|
||||
> ### ⚠️ **Warning**
|
||||
> K3s does not include `etcdctl` in the system PATH. If you need to perform etcd troubleshooting on a K3s cluster, you may need to install it or locate it within the K3s data directory.
|
||||
|
||||
```bash
|
||||
export PATH=$PATH:/usr/local/bin
|
||||
export CRI_CONFIG_FILE=/var/lib/rancher/k3s/agent/etc/crictl.yaml
|
||||
```
|
||||
|
||||
|
||||
## Checking if the etcd Container is Running
|
||||
|
||||
The container for etcd should have status **Up**. The duration shown after **Up** is the time the container has been running.
|
||||
**RKE2**: The container for etcd should be in the **Running** state.
|
||||
|
||||
```
|
||||
docker ps -a -f=name=etcd$
|
||||
```bash
|
||||
crictl ps --name etcd
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES
|
||||
d26adbd23643 rancher/mirrored-coreos-etcd:v3.5.7 "/usr/local/bin/etcd…" 30 minutes ago Up 30 minutes etcd
|
||||
CONTAINER IMAGE CREATED STATE NAME ATTEMPT POD ID POD NAMESPACE
|
||||
f1e289d202ed0 11ad16872a9cf 58 minutes ago Running etcd 0 7b56aab8204ea etcd-cluster1 kube-system
|
||||
```
|
||||
|
||||
## etcd Container Logging
|
||||
|
||||
The logging of the container can contain information on what the problem could be.
|
||||
**K3s**: Etcd runs as an embedded process in the K3s service. Check the service status:
|
||||
|
||||
```bash
|
||||
systemctl status k3s
|
||||
```
|
||||
docker logs etcd
|
||||
|
||||
## etcd Logging
|
||||
|
||||
The logs can contain information on what the problem could be.
|
||||
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl logs $etcdcontainer
|
||||
```
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
journalctl -u k3s | grep -i etcd
|
||||
```
|
||||
| Log | Explanation |
|
||||
|-----|------------------|
|
||||
@@ -46,18 +80,43 @@ The address where etcd is listening depends on the address configuration of the
|
||||
|
||||
Output should contain all the nodes with the `etcd` role and the output should be identical on all nodes.
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
Run the command inside the etcd container.
|
||||
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl member list \
|
||||
--cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt \
|
||||
--key /var/lib/rancher/rke2/server/tls/etcd/server-client.key \
|
||||
--cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec etcd etcdctl member list
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl member list \
|
||||
--cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt \
|
||||
--key /var/lib/rancher/k3s/server/tls/etcd/server-client.key \
|
||||
--cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
1c424074df86e854, started, cluster-node1-f289ac71, https://IP:2380, https://IP:2379, false
|
||||
45c68c44c5a792ff, started, cluster-node2-67e3cf6f, https://IP:2380, https://IP:2379, false
|
||||
7c584f77c5180258, started, cluster-node3-e976bc00, https://IP:2380, https://IP:2379, false
|
||||
```
|
||||
|
||||
### Check Endpoint Status
|
||||
|
||||
The values for `RAFT TERM` should be equal and `RAFT INDEX` should be not be too far apart from each other.
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl endpoint status --write-out table --endpoints=$(crictl exec $etcdcontainer etcdctl member list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec -e ETCDCTL_ENDPOINTS=$(docker exec etcd etcdctl member list | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') etcd etcdctl endpoint status --write-out table
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl endpoint status --write-out table --endpoints=$(etcdctl member list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
@@ -65,17 +124,22 @@ Example output:
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
| ENDPOINT | ID | VERSION | DB SIZE | IS LEADER | RAFT TERM | RAFT INDEX |
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
| https://IP:2379 | 333ef673fc4add56 | 3.5.7 | 24 MB | false | 72 | 66887 |
|
||||
| https://IP:2379 | 5feed52d940ce4cf | 3.5.7 | 24 MB | true | 72 | 66887 |
|
||||
| https://IP:2379 | db6b3bdb559a848d | 3.5.7 | 25 MB | false | 72 | 66887 |
|
||||
| https://IP:2379 | 333ef673fc4add56 | 3.6.7 | 24 MB | false | 72 | 66887 |
|
||||
| https://IP:2379 | 5feed52d940ce4cf | 3.6.7 | 24 MB | true | 72 | 66887 |
|
||||
| https://IP:2379 | db6b3bdb559a848d | 3.6.7 | 25 MB | false | 72 | 66887 |
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
```
|
||||
|
||||
### Check Endpoint Health
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl endpoint health --endpoints=$(crictl exec $etcdcontainer etcdctl member list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec -e ETCDCTL_ENDPOINTS=$(docker exec etcd etcdctl member list | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') etcd etcdctl endpoint health
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl endpoint health --endpoints=$(etcdctl member list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
@@ -84,54 +148,104 @@ https://IP:2379 is healthy: successfully committed proposal: took = 2.113189ms
|
||||
https://IP:2379 is healthy: successfully committed proposal: took = 2.649963ms
|
||||
https://IP:2379 is healthy: successfully committed proposal: took = 2.451201ms
|
||||
```
|
||||
### Check Connectivity on etcd Ports
|
||||
|
||||
### Check Connectivity on Port TCP/2379
|
||||
> In modern versions of Kubernetes, the etcd database (versions 3.5 and newer) introduced significant architectural changes regarding network traffic handling. Previously, etcd permitted standard HTTP REST requests on its primary client port (`2379`). However, to enhance performance and security, etcd 3.5+ strictly enforces the gRPC protocol on this port.<br />
|
||||
If you attempt to use standard HTTP tools like `curl` to test connectivity on port `2379`, the etcd server will automatically terminate the connection or return an error. This behavior often leads administrators to misinterpret the result as a closed port or a node failure.
|
||||
|
||||
Command:
|
||||
Since standard HTTP clients can no longer probe the primary etcd ports, the transport layer must be utilized for network troubleshooting. Using `openssl s_client` instead of `curl` bypasses the gRPC application requirement, allowing the raw TCP and TLS handshake to be tested directly.
|
||||
|
||||
These script isolate the network and security infrastructure from the database application. A successful `Verify return code: 0 (ok)` explicitly confirms four critical infrastructure components:
|
||||
|
||||
* **Network Path:** Routing is functional, and firewalls permit traffic on TCP port `2379` or `2380`.
|
||||
* **Process Availability:** The etcd service is running and actively listening on the designated port.
|
||||
* **Certificate Validity:** The TLS certificates are active, correctly formatted, and have not expired.
|
||||
* **Mutual Authentication (mTLS):** The node successfully authenticates against the cluster's specific Certificate Authority (CA).
|
||||
|
||||
**How these tests differ from the `etcdctl endpoint health` test**:
|
||||
|
||||
If `etcdctl endpoint health` test is failing, run these Connectivity Ports test scripts. If the scripts succeed, your network and certificates are intact, and the issue is likely confined to the etcd database itself. If these scripts fail, the issue is related to a firewall/network restriction, or certificate expiration.
|
||||
|
||||
#### Port TCP/2379
|
||||
|
||||
**RKE2**:
|
||||
```bash
|
||||
for endpoint in $(crictl exec $etcdcontainer etcdctl member list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | cut -d, -f5); do
|
||||
echo "Validating connection to ${endpoint} (Client)";
|
||||
echo | openssl s_client -connect ${endpoint#https://} \
|
||||
-CAfile /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt \
|
||||
-cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt \
|
||||
-key /var/lib/rancher/rke2/server/tls/etcd/server-client.key 2>/dev/null | grep -E 'Verify return code' || echo "Connection Failed/Timeout"
|
||||
done
|
||||
```
|
||||
for endpoint in $(docker exec etcd etcdctl member list | cut -d, -f5); do
|
||||
echo "Validating connection to ${endpoint}/health"
|
||||
docker run --net=host -v $(docker inspect kubelet --format '{{ range .Mounts }}{{ if eq .Destination "/etc/kubernetes" }}{{ .Source }}{{ end }}{{ end }}')/ssl:/etc/kubernetes/ssl:ro appropriate/curl -s -w "\n" --cacert $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_CACERT" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) --cert $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_CERT" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) --key $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_KEY" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) "${endpoint}/health"
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
for endpoint in $(etcdctl member list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | cut -d, -f5); do
|
||||
echo "Validating connection to ${endpoint} (Client)";
|
||||
echo | openssl s_client -connect ${endpoint#https://} \
|
||||
-CAfile /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt \
|
||||
-cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt \
|
||||
-key /var/lib/rancher/k3s/server/tls/etcd/server-client.key 2>/dev/null | grep -E 'Verify return code' || echo "Connection Failed/Timeout"
|
||||
done
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
Validating connection to https://IP:2379/health
|
||||
{"health": "true"}
|
||||
Validating connection to https://IP:2379/health
|
||||
{"health": "true"}
|
||||
Validating connection to https://IP:2379/health
|
||||
{"health": "true"}
|
||||
Validating connection to https://IP:2379/health (Client)
|
||||
Verify return code: 0 (ok)
|
||||
Validating connection to https://IP:2379/health (Client)
|
||||
Verify return code: 0 (ok)
|
||||
Validating connection to https://IP:2379/health (Client)
|
||||
Verify return code: 0 (ok)
|
||||
```
|
||||
|
||||
### Check Connectivity on Port TCP/2380
|
||||
#### Port TCP/2380
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
for endpoint in $(crictl exec $etcdcontainer etcdctl member list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | cut -d, -f4); do
|
||||
echo "Validating connection to ${endpoint} (Peer)";
|
||||
echo | openssl s_client -connect ${endpoint#https://} \
|
||||
-CAfile /var/lib/rancher/rke2/server/tls/etcd/peer-ca.crt \
|
||||
-cert /var/lib/rancher/rke2/server/tls/etcd/peer-server-client.crt \
|
||||
-key /var/lib/rancher/rke2/server/tls/etcd/peer-server-client.key 2>/dev/null | grep -E 'Verify return code' || echo "Connection Failed/Timeout"
|
||||
done
|
||||
```
|
||||
for endpoint in $(docker exec etcd etcdctl member list | cut -d, -f4); do
|
||||
echo "Validating connection to ${endpoint}/version";
|
||||
docker run --net=host -v $(docker inspect kubelet --format '{{ range .Mounts }}{{ if eq .Destination "/etc/kubernetes" }}{{ .Source }}{{ end }}{{ end }}')/ssl:/etc/kubernetes/ssl:ro appropriate/curl --http1.1 -s -w "\n" --cacert $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_CACERT" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) --cert $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_CERT" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) --key $(docker inspect -f '{{range $index, $value := .Config.Env}}{{if eq (index (split $value "=") 0) "ETCDCTL_KEY" }}{{range $i, $part := (split $value "=")}}{{if gt $i 1}}{{print "="}}{{end}}{{if gt $i 0}}{{print $part}}{{end}}{{end}}{{end}}{{end}}' etcd) "${endpoint}/version"
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
for endpoint in $(etcdctl member list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | cut -d, -f4); do
|
||||
echo "Validating connection to ${endpoint} (Peer)";
|
||||
echo | openssl s_client -connect ${endpoint#https://} \
|
||||
-CAfile /var/lib/rancher/k3s/server/tls/etcd/peer-ca.crt \
|
||||
-cert /var/lib/rancher/k3s/server/tls/etcd/peer-server-client.crt \
|
||||
-key /var/lib/rancher/k3s/server/tls/etcd/peer-server-client.key 2>/dev/null | grep -E 'Verify return code' || echo "Connection Failed/Timeout"
|
||||
done
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
Validating connection to https://IP:2380/version
|
||||
{"etcdserver":"3.5.7","etcdcluster":"3.5.0"}
|
||||
Validating connection to https://IP:2380/version
|
||||
{"etcdserver":"3.5.7","etcdcluster":"3.5.0"}
|
||||
Validating connection to https://IP:2380/version
|
||||
{"etcdserver":"3.5.7","etcdcluster":"3.5.0"}
|
||||
Validating connection to https://IP:2380/version (Peer)
|
||||
Verify return code: 0 (ok)
|
||||
Validating connection to https://IP:2380/version (Peer)
|
||||
Verify return code: 0 (ok)
|
||||
Validating connection to https://IP:2380/version (Peer)
|
||||
Verify return code: 0 (ok)
|
||||
```
|
||||
|
||||
## etcd Alarms
|
||||
|
||||
etcd will trigger alarms, for instance when it runs out of space.
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl alarm list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec etcd etcdctl alarm list
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl alarm list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output when NOSPACE alarm is triggered:
|
||||
@@ -154,10 +268,16 @@ Resolutions:
|
||||
|
||||
### Compact the Keyspace
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
rev=$(crictl exec $etcdcontainer etcdctl endpoint status --write-out json --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | egrep -o '"revision":[0-9]*' | egrep -o '[0-9]*' | head -1)
|
||||
crictl exec $etcdcontainer etcdctl compact "$rev" --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
rev=$(docker exec etcd etcdctl endpoint status --write-out json | egrep -o '"revision":[0-9]*' | egrep -o '[0-9]*')
|
||||
docker exec etcd etcdctl compact "$rev"
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
rev=$(etcdctl endpoint status --write-out json --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | egrep -o '"revision":[0-9]*' | egrep -o '[0-9]*' | head -1)
|
||||
etcdctl compact "$rev" --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
@@ -167,55 +287,39 @@ compacted revision xxx
|
||||
|
||||
### Defrag All etcd Members
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl defrag --endpoints=$(crictl exec $etcdcontainer etcdctl member list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec -e ETCDCTL_ENDPOINTS=$(docker exec etcd etcdctl member list | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') etcd etcdctl defrag
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl defrag --endpoints=$(etcdctl member list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
Finished defragmenting etcd member[https://IP:2379]
|
||||
Finished defragmenting etcd member[https://IP:2379]
|
||||
Finished defragmenting etcd member[https://IP:2379]
|
||||
```
|
||||
|
||||
### Check Endpoint Status
|
||||
|
||||
Command:
|
||||
```
|
||||
docker exec -e ETCDCTL_ENDPOINTS=$(docker exec etcd etcdctl member list | cut -d, -f5 | sed -e 's/ //g' | paste -sd ',') etcd etcdctl endpoint status --write-out table
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
| ENDPOINT | ID | VERSION | DB SIZE | IS LEADER | RAFT TERM | RAFT INDEX |
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
| https://IP:2379 | e973e4419737125 | 3.5.7 | 553 kB | false | 32 | 2449410 |
|
||||
| https://IP:2379 | 4a509c997b26c206 | 3.5.7 | 553 kB | false | 32 | 2449410 |
|
||||
| https://IP:2379 | b217e736575e9dd3 | 3.5.7 | 553 kB | true | 32 | 2449410 |
|
||||
+-----------------+------------------+---------+---------+-----------+-----------+------------+
|
||||
Finished defragmenting etcd member[https://IP:2379]. took xx.xxxxxxms
|
||||
Finished defragmenting etcd member[https://IP:2379]. took xx.xxxxxxms
|
||||
Finished defragmenting etcd member[https://IP:2379]. took xx.xxxxxxms
|
||||
```
|
||||
|
||||
### Disarm Alarm
|
||||
|
||||
After verifying that the DB size went down after compaction and defragmenting, the alarm needs to be disarmed for etcd to allow writes again.
|
||||
|
||||
Command:
|
||||
```
|
||||
docker exec etcd etcdctl alarm list
|
||||
docker exec etcd etcdctl alarm disarm
|
||||
docker exec etcd etcdctl alarm list
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl alarm list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
crictl exec $etcdcontainer etcdctl alarm disarm --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
crictl exec $etcdcontainer etcdctl alarm list --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
docker exec etcd etcdctl alarm list
|
||||
memberID:x alarm:NOSPACE
|
||||
memberID:x alarm:NOSPACE
|
||||
memberID:x alarm:NOSPACE
|
||||
docker exec etcd etcdctl alarm disarm
|
||||
docker exec etcd etcdctl alarm list
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl alarm list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
etcdctl alarm disarm --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
etcdctl alarm list --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
## Configure Log Level
|
||||
@@ -228,7 +332,7 @@ You can no longer dynamically change the log level in etcd v3.5 or later.
|
||||
|
||||
### etcd v3.5 And Later
|
||||
|
||||
To configure the log level for etcd, edit the cluster YAML:
|
||||
To configure the log level for etcd, edit the cluster configuration YAML:
|
||||
|
||||
```
|
||||
services:
|
||||
@@ -237,20 +341,7 @@ services:
|
||||
log-level: "debug"
|
||||
```
|
||||
|
||||
### etcd v3.4 And Earlier
|
||||
|
||||
In earlier etcd versions, you can use the API to dynamically change the log level. Configure debug logging using the commands below:
|
||||
|
||||
```
|
||||
docker run --net=host -v $(docker inspect kubelet --format '{{ range .Mounts }}{{ if eq .Destination "/etc/kubernetes" }}{{ .Source }}{{ end }}{{ end }}')/ssl:/etc/kubernetes/ssl:ro appropriate/curl -s -XPUT -d '{"Level":"DEBUG"}' --cacert $(docker exec etcd printenv ETCDCTL_CACERT) --cert $(docker exec etcd printenv ETCDCTL_CERT) --key $(docker exec etcd printenv ETCDCTL_KEY) $(docker exec etcd printenv ETCDCTL_ENDPOINTS)/config/local/log
|
||||
```
|
||||
|
||||
To reset the log level back to the default (`INFO`), you can use the following command.
|
||||
|
||||
Command:
|
||||
```
|
||||
docker run --net=host -v $(docker inspect kubelet --format '{{ range .Mounts }}{{ if eq .Destination "/etc/kubernetes" }}{{ .Source }}{{ end }}{{ end }}')/ssl:/etc/kubernetes/ssl:ro appropriate/curl -s -XPUT -d '{"Level":"INFO"}' --cacert $(docker exec etcd printenv ETCDCTL_CACERT) --cert $(docker exec etcd printenv ETCDCTL_CERT) --key $(docker exec etcd printenv ETCDCTL_KEY) $(docker exec etcd printenv ETCDCTL_ENDPOINTS)/config/local/log
|
||||
```
|
||||
After modifying the configuration, restart the service (`systemctl restart rke2-server` or `systemctl restart k3s`) if you are configuring a stand-alone cluster.
|
||||
|
||||
## etcd Content
|
||||
|
||||
@@ -258,24 +349,40 @@ If you want to investigate the contents of your etcd, you can either watch strea
|
||||
|
||||
### Watch Streaming Events
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl watch --prefix /registry --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec etcd etcdctl watch --prefix /registry
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl watch --prefix /registry --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
If you only want to see the affected keys (and not the binary data), you can append `| grep -a ^/registry` to the command to filter for keys only.
|
||||
|
||||
### Query etcd Directly
|
||||
|
||||
Command:
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl get /registry --prefix=true --keys-only --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
docker exec etcd etcdctl get /registry --prefix=true --keys-only
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl get /registry --prefix=true --keys-only --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt
|
||||
```
|
||||
|
||||
You can process the data to get a summary of count per key, using the command below:
|
||||
|
||||
**RKE2**:
|
||||
```bash
|
||||
crictl exec $etcdcontainer etcdctl get /registry --prefix=true --keys-only --cert /var/lib/rancher/rke2/server/tls/etcd/server-client.crt --key /var/lib/rancher/rke2/server/tls/etcd/server-client.key --cacert /var/lib/rancher/rke2/server/tls/etcd/server-ca.crt | grep -v ^$ | awk -F'/' '{ if ($3 ~ /cattle.io/) {h[$3"/"$4]++} else { h[$3]++ }} END { for(k in h) print h[k], k }' | sort -nr
|
||||
```
|
||||
docker exec etcd etcdctl get /registry --prefix=true --keys-only | grep -v ^$ | awk -F'/' '{ if ($3 ~ /cattle.io/) {h[$3"/"$4]++} else { h[$3]++ }} END { for(k in h) print h[k], k }' | sort -nr
|
||||
|
||||
**K3s**:
|
||||
```bash
|
||||
etcdctl get /registry --prefix=true --keys-only --cert /var/lib/rancher/k3s/server/tls/etcd/server-client.crt --key /var/lib/rancher/k3s/server/tls/etcd/server-client.key --cacert /var/lib/rancher/k3s/server/tls/etcd/server-ca.crt | grep -v ^$ | awk -F'/' '{ if ($3 ~ /cattle.io/) {h[$3"/"$4]++} else { h[$3]++ }} END { for(k in h) print h[k], k }' | sort -nr
|
||||
```
|
||||
|
||||
## Replacing Unhealthy etcd Nodes
|
||||
|
||||
+68
-14
@@ -8,31 +8,85 @@ title: Troubleshooting Worker Nodes and Generic Components
|
||||
|
||||
This section applies to every node as it includes components that run on nodes with any role.
|
||||
|
||||
## Check if the Containers are Running
|
||||
## Prerequisites
|
||||
|
||||
There are two specific containers launched on nodes with the `worker` role:
|
||||
Since RKE2 and K3s utilize `containerd` as the container runtime, `crictl` serves as the primary tool for container management, replacing the Docker CLI. To allow `crictl` to communicate with `containerd`, you must configure your environment by exporting the following variables:
|
||||
|
||||
* kubelet
|
||||
* kube-proxy
|
||||
|
||||
The containers should have status `Up`. The duration shown after `Up` is the time the container has been running.
|
||||
### RKE2
|
||||
|
||||
```bash
|
||||
export PATH=$PATH:/var/lib/rancher/rke2/bin/
|
||||
export CRI_CONFIG_FILE=/var/lib/rancher/rke2/agent/etc/crictl.yaml
|
||||
```
|
||||
docker ps -a -f=name='kubelet|kube-proxy'
|
||||
|
||||
### K3s
|
||||
|
||||
```bash
|
||||
export PATH=$PATH:/usr/local/bin
|
||||
export CRI_CONFIG_FILE=/var/lib/rancher/k3s/agent/etc/crictl.yaml
|
||||
```
|
||||
|
||||
## Check if the Components are Running
|
||||
|
||||
There are two specific components launched on nodes with the `worker` role:
|
||||
|
||||
* `kubelet`
|
||||
* `kube-proxy`
|
||||
|
||||
### RKE2
|
||||
|
||||
The `kubelet` runs natively as part of the `rke2-agent` (or `rke2-server`) systemd process, while `kube-proxy` runs as a Static Pod managed by `containerd`.
|
||||
|
||||
Check the status of the `kubelet` via the agent service:
|
||||
```bash
|
||||
systemctl status rke2-agent
|
||||
```
|
||||
:::note
|
||||
|
||||
If you are checking a controlplane node, use `systemctl status rke2-server` instead.
|
||||
|
||||
:::
|
||||
|
||||
Check the status of `kube-proxy` using `crictl`:
|
||||
```bash
|
||||
crictl ps --name kube-proxy
|
||||
```
|
||||
|
||||
Example output:
|
||||
```
|
||||
CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES
|
||||
158d0dcc33a5 rancher/hyperkube:v1.11.5-rancher1 "/opt/rke-tools/en..." 3 hours ago Up 3 hours kube-proxy
|
||||
a30717ecfb55 rancher/hyperkube:v1.11.5-rancher1 "/opt/rke-tools/en..." 3 hours ago Up 3 hours kubelet
|
||||
CONTAINER IMAGE CREATED STATE NAME ATTEMPT POD ID
|
||||
26c7159abbcc rancher/hardened-kubernetes:v1.28.8-rke2r1-build20240404 3 hours ago Running kube-proxy 0 1a2b3c4d5e6f7
|
||||
```
|
||||
|
||||
## Container Logging
|
||||
### K3s
|
||||
|
||||
The logging of the containers can contain information on what the problem could be.
|
||||
Both `kubelet` and `kube-proxy` run as embedded processes inside the `k3s-agent` (or `k3s` server) systemd service. There are no separate containers for them.
|
||||
|
||||
Check their status by checking the K3s service:
|
||||
```bash
|
||||
systemctl status k3s-agent
|
||||
```
|
||||
docker logs kubelet
|
||||
docker logs kube-proxy
|
||||
:::note
|
||||
If you are checking a controlplane node, use `systemctl status k3s` instead.
|
||||
:::
|
||||
|
||||
## Component Logging
|
||||
|
||||
The logging of the components can contain information on what the problem could be.
|
||||
|
||||
### RKE2
|
||||
|
||||
```bash
|
||||
# kubelet logs are part of the systemd service
|
||||
journalctl -u rke2-agent -f | grep -i "kubelet"
|
||||
|
||||
# kube-proxy logs are retrieved from containerd
|
||||
crictl logs $(crictl ps --name kube-proxy -q)
|
||||
```
|
||||
|
||||
### K3s
|
||||
|
||||
```bash
|
||||
# Both components log to the systemd service
|
||||
journalctl -u k3s-agent -f | grep -iE "kubelet|kube-proxy"
|
||||
```
|
||||
|
||||
+85
-14
@@ -78,66 +78,137 @@ kubectl -n kube-system get endpoints kube-scheduler -o jsonpath='{.metadata.anno
|
||||
|
||||
## Ingress Controller
|
||||
|
||||
The default Ingress Controller is Traefik and is deployed as a DaemonSet in the `traefik` namespace. The pods are only scheduled to nodes with the `worker` role.
|
||||
The default Ingress Controller is Traefik and is deployed as a DaemonSet in the `kube-system` namespace. The pods are only scheduled to nodes with the `worker` role.
|
||||
|
||||
Check if the pods are running on all nodes:
|
||||
|
||||
```
|
||||
kubectl -n traefik get pods -o wide
|
||||
kubectl -n kube-system get pods -o wide
|
||||
```
|
||||
|
||||
Example output:
|
||||
Example RKE2 output:
|
||||
|
||||
```
|
||||
kubectl -n traefik get pods -o wide
|
||||
kubectl -n kube-system get pods -o wide
|
||||
NAME READY STATUS RESTARTS AGE IP NODE
|
||||
default-http-backend-797c5bc547-kwwlq 1/1 Running 0 17m x.x.x.x worker-1
|
||||
traefik-4qd64 1/1 Running 0 14m x.x.x.x worker-1
|
||||
traefik-8wxhm 1/1 Running 0 13m x.x.x.x worker-0
|
||||
local-path-provisioner-xxxxxxxxxx-xxxxx 1/1 Running 0 17m x.x.x.x worker-1
|
||||
rke2-traefik-xxxxxxxxxx-xxxxx 1/1 Running 0 14m x.x.x.x worker-1
|
||||
svclb-rke2-traefik-xxxxxxxx-xxxxx 1/1 Running 0 13m x.x.x.x worker-0
|
||||
...
|
||||
```
|
||||
|
||||
Example K3s output:
|
||||
|
||||
```
|
||||
kubectl -n kube-system get pods -o wide
|
||||
NAME READY STATUS RESTARTS AGE IP NODE
|
||||
local-path-provisioner-xxxxxxxxxx-xxxxx 1/1 Running 0 17m x.x.x.x worker-1
|
||||
traefik-xxxxxxxxxx-xxxxx 1/1 Running 0 14m x.x.x.x worker-1
|
||||
svclb-traefik-xxxxxxxx-xxxxx 1/1 Running 0 13m x.x.x.x worker-0
|
||||
...
|
||||
```
|
||||
|
||||
If a pod is unable to run (Status is not **Running**, Ready status is not showing `1/1` or you see a high count of Restarts), check the pod details, logs and namespace events.
|
||||
|
||||
### Pod details
|
||||
|
||||
RKE2 example:
|
||||
|
||||
```
|
||||
kubectl -n traefik describe pods -l app=traefik
|
||||
kubectl -n kube-system describe pods -l app.kubernetes.io/name=rke2-traefik
|
||||
```
|
||||
|
||||
K3s example:
|
||||
|
||||
```
|
||||
kubectl -n kube-system describe pods -l app.kubernetes.io/name=traefik
|
||||
```
|
||||
|
||||
### Pod container logs
|
||||
|
||||
The below command can show the logs of all the pods labeled "app=traefik", but it will display only 10 lines of log because of the restrictions of the `kubectl logs` command. Refer to `--tail` of `kubectl logs -h` for more information.
|
||||
The below command can show the logs of all the pods labeled "app.kubernetes.io/name=rke2-traefik" if using RKE2 or "app.kubernetes.io/name=traefik" if using K3s, but it will display only 10 lines of log because of the restrictions of the `kubectl logs` command. Refer to `--tail` of `kubectl logs -h` for more information.
|
||||
|
||||
RKE2 example:
|
||||
|
||||
```
|
||||
kubectl -n traefik logs -l app=traefik
|
||||
kubectl -n kube-system logs -l app.kubernetes.io/name=rke2-traefik
|
||||
```
|
||||
|
||||
K3s example:
|
||||
|
||||
```
|
||||
kubectl -n kube-system logs -l app.kubernetes.io/name=traefik
|
||||
```
|
||||
|
||||
If the full log is needed, specify the pod name in the trailing command:
|
||||
|
||||
```
|
||||
kubectl -n traefik logs <pod name>
|
||||
kubectl -n kube-system logs <pod name>
|
||||
```
|
||||
|
||||
### Namespace events
|
||||
|
||||
```
|
||||
kubectl -n traefik get events
|
||||
kubectl -n kube-system get events
|
||||
```
|
||||
|
||||
### Debug logging
|
||||
|
||||
To enable debug logging:
|
||||
|
||||
RKE2 example:
|
||||
|
||||
```
|
||||
kubectl -n traefik patch ds traefik --type='json' -p='[{"op": "add", "path": "/spec/template/spec/containers/0/args/-", "value": "--v=5"}]'
|
||||
cat <<EOF | kubectl apply -f -
|
||||
apiVersion: helm.cattle.io/v1
|
||||
kind: HelmChartConfig
|
||||
metadata:
|
||||
name: rke2-traefik
|
||||
namespace: kube-system
|
||||
spec:
|
||||
valuesContent: |-
|
||||
additionalArguments:
|
||||
- "--log.level=DEBUG"
|
||||
EOF
|
||||
```
|
||||
|
||||
K3s example:
|
||||
|
||||
```
|
||||
cat <<EOF | kubectl apply -f -
|
||||
apiVersion: helm.cattle.io/v1
|
||||
kind: HelmChartConfig
|
||||
metadata:
|
||||
name: traefik
|
||||
namespace: kube-system
|
||||
spec:
|
||||
valuesContent: |-
|
||||
logs:
|
||||
general:
|
||||
level: "DEBUG"
|
||||
EOF
|
||||
```
|
||||
|
||||
### Check configuration
|
||||
|
||||
Retrieve generated configuration in each pod:
|
||||
|
||||
RKE2 example for manual Traefik configuration file check:
|
||||
|
||||
```
|
||||
kubectl -n traefik get pods -l app=traefik --no-headers -o custom-columns=.NAME:.metadata.name | while read pod; do kubectl -n traefik exec $pod -- cat /etc/nginx/nginx.conf; done
|
||||
kubectl exec -n kube-system pod/traefik-xxxxxxxxx-xxxxx -- cat /var/lib/rancher/rke2/server/manifests/rke2-traefik-config.yaml
|
||||
```
|
||||
|
||||
K3s example for manual Traefik configuration file check:
|
||||
|
||||
```
|
||||
kubectl exec -n kube-system pod/traefik-xxxxxxxxx-xxxxx -- cat /var/lib/rancher/k3s/server/manifests/k3s-traefik-config.yaml
|
||||
```
|
||||
|
||||
RKE2/K3s example for Traefik CLI argument configuration check:
|
||||
|
||||
```
|
||||
kubectl get pod traefik-xxxxxxxxx-xxxxx -n kube-system -o jsonpath='{.spec.containers[0].args}'
|
||||
```
|
||||
|
||||
## Rancher agents
|
||||
|
||||
@@ -15,6 +15,14 @@ Make sure you configured the correct kubeconfig (for example, `export KUBECONFIG
|
||||
Double check if all the [required ports](../../how-to-guides/new-user-guides/kubernetes-clusters-in-rancher-setup/node-requirements-for-rancher-managed-clusters.md#networking-requirements) are opened in your (host) firewall. The overlay network uses UDP in comparison to all other required ports which are TCP.
|
||||
|
||||
|
||||
## Check if your downstream node can communicate to Rancher Manager
|
||||
|
||||
Rancher components with HTTP endpoints generally contain a `ping` liveness probe which you can use to test connectivity. Replace the `$RANCHER_URL` as appropriate and run the following from a node to check that it has connectivity to Rancher Manager's servers in the `local` cluster. If successful, it should return `pong`.
|
||||
|
||||
```
|
||||
curl -k https://$RANCHER_URL/ping
|
||||
```
|
||||
|
||||
## Check if Overlay Network is Functioning Correctly
|
||||
|
||||
The pod can be scheduled to any of the hosts you used for your cluster, but that means that the NGINX ingress controller needs to be able to route the request from `NODE_1` to `NODE_2`. This happens over the overlay network. If the overlay network is not functioning, you will experience intermittent TCP/HTTP connection failures due to the NGINX ingress controller not being able to route to the pod.
|
||||
|
||||
@@ -12,6 +12,7 @@ Rancher will publish deprecated features as part of the [release notes](https://
|
||||
|
||||
| Patch Version | Release Date |
|
||||
|---------------|---------------|
|
||||
| [2.14.2](https://github.com/rancher/rancher/releases/tag/v2.14.2) | May 28, 2026 |
|
||||
| [2.14.1](https://github.com/rancher/rancher/releases/tag/v2.14.1) | April 30, 2026 |
|
||||
| [2.14.0](https://github.com/rancher/rancher/releases/tag/v2.14.0) | March 25, 2026 |
|
||||
|
||||
|
||||
+34
-9
@@ -26,6 +26,31 @@ Review the list of known issues for each Rancher version, which can be found in
|
||||
|
||||
Note that upgrades _to_ or _from_ any chart in the [rancher-alpha repository](../resources/choose-a-rancher-version.md#helm-chart-repositories) aren't supported.
|
||||
|
||||
### Upgrade Path
|
||||
|
||||
:::important
|
||||
|
||||
**Important:** The only tested and supported Rancher upgrade path between minor versions (e.g. v2.13.x to v2.14.x) is to upgrade from the latest available patch version of your current running minor release to the latest available patch version of the next minor release.
|
||||
|
||||
Before initiating a minor version upgrade, verify that you are running the most recent patch release of your current version.
|
||||
|
||||
You can query the available chart versions with the Helm CLI:
|
||||
|
||||
1. Update your local Helm repo cache.
|
||||
|
||||
```
|
||||
helm repo update
|
||||
```
|
||||
|
||||
1. Search for available versions in your [specific repository](../resources/choose-a-rancher-version.md#helm-chart-repositories) (e.g., rancher-stable):
|
||||
|
||||
```
|
||||
helm search repo rancher-<CHART_REPO>/rancher --versions
|
||||
```
|
||||
|
||||
If your installation is not on the latest patch version of the current minor release, you must upgrade to that version before proceeding to the next minor version.
|
||||
:::
|
||||
|
||||
### Helm Version
|
||||
|
||||
:::important
|
||||
@@ -170,17 +195,17 @@ There will be more values that are listed with this command. This is just an exa
|
||||
|
||||
:::
|
||||
|
||||
:::tip
|
||||
:::tip
|
||||
|
||||
Your deployment name may vary; for example, if you're deploying Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
Your deployment name may vary; for example, if you're deploying Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
```
|
||||
helm get values rancher-stable -n cattle-system
|
||||
|
||||
hostname: rancher.my.org
|
||||
```
|
||||
|
||||
:::
|
||||
:::
|
||||
|
||||
If you are upgrading cert-manager to the latest version from v1.5 or below, follow the [cert-manager upgrade docs](../resources/upgrade-cert-manager.md#option-c-upgrade-cert-manager-from-versions-15-and-below) to learn how to upgrade cert-manager without needing to perform an uninstall or reinstall of Rancher. Otherwise, follow the [steps to upgrade Rancher](#steps-to-upgrade-rancher) below.
|
||||
|
||||
@@ -203,17 +228,17 @@ The above is an example, there may be more values from the previous step that ne
|
||||
|
||||
:::
|
||||
|
||||
:::tip
|
||||
:::tip
|
||||
|
||||
If you deploy Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
If you deploy Rancher through the AWS Marketplace, the deployment name is 'rancher-stable'.
|
||||
Thus:
|
||||
```
|
||||
helm upgrade rancher-stable rancher-<CHART_REPO>/rancher \
|
||||
--namespace cattle-system \
|
||||
--set hostname=rancher.my.org
|
||||
```
|
||||
|
||||
:::
|
||||
:::
|
||||
|
||||
Alternatively, it's possible to export the current values to a file and reference that file during upgrade. For example, to only change the Rancher version:
|
||||
|
||||
@@ -223,7 +248,7 @@ Alternatively, it's possible to export the current values to a file and referenc
|
||||
```
|
||||
1. Update only the Rancher version:
|
||||
|
||||
|
||||
|
||||
```
|
||||
helm upgrade rancher rancher-<CHART_REPO>/rancher \
|
||||
--namespace cattle-system \
|
||||
|
||||
+2
-2
@@ -33,7 +33,7 @@ To use a premade dashboard, go to [https://grafana.com/grafana/dashboards](https
|
||||
To use your own dashboard:
|
||||
|
||||
1. Click on the link to open Grafana. On the cluster detail page, click **Monitoring**.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin/prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
:::note
|
||||
|
||||
@@ -113,7 +113,7 @@ Note that the RBAC roles exposed by the Monitoring chart to add Grafana Dashboar
|
||||
1. On the **Clusters** page, go to the cluster where you want to configure the Grafana namespace and click **Explore**.
|
||||
1. In the left navigation bar, click **Monitoring**.
|
||||
1. Click **Grafana**.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin/prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
1. Log in to Grafana. Note: The default Admin username and password for the Grafana instance is `admin` and `prom-operator`. Alternative credentials can also be supplied on deploying or upgrading the chart.
|
||||
|
||||
:::note
|
||||
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user