news 2026/10/8 7:51:25

Prometheus Node Exporter Helm Chart 部署指南:基于 charts 仓库的 Kubernetes 节点监控实践

作者头像

张小明

前端开发工程师

1.2k 24
文章封面图
Prometheus Node Exporter Helm Chart 部署指南:基于 charts 仓库的 Kubernetes 节点监控实践

【免费下载链接】charts

⚠️(OBSOLETE) Curated applications for Kubernetes

项目地址:https://gitcode.com/gh_mirrors/chart/charts
点击查看免费下载

导读

本文以 charts 仓库中的 stable/prometheus-node-exporter/README.md 为核心文档,系统讲解如何在 Kubernetes 集群上通过 Helm 部署 Prometheus Node Exporter,覆盖安装/卸载、全部可配置参数、参数默认值,以及 DaemonSet、Service、ServiceMonitor、RBAC 等底层模板实现。读完本文,你将掌握用一条helm install命令拉起全节点指标采集、按需调整镜像与端口、打通 Prometheus 服务发现,并为外部已部署的 Node Exporter 补充 Endpoints 的完整实战方案。

一、背景:为什么用 Node Exporter 监控 Kubernetes 节点

Prometheus Node Exporter 是 Prometheus 生态中最常用的主机指标采集器,负责暴露节点的 CPU、内存、磁盘、网络、文件系统等操作系统级指标。在 Kubernetes 中,节点层面的指标(如node_cpu_seconds_total、node_memory_*、node_filesystem_*)都来自 Node Exporter,是集群可观测性建设的基础组件。

本仓库中的prometheus-node-exporter是一个 Helm 图表(chart),其元数据定义在 Chart.yaml:name: prometheus-node-exporter,version: 1.11.2,appVersion: 1.0.1(即默认打包的 Node Exporter 镜像版本为 v1.0.1)。需要特别说明的是,该图表在仓库中已被标记为deprecated: true,README 首行即声明 "DEPRECATED and moved to https://github.com/prometheus-community/helm-charts"——即该项目已停止维护并迁移到 prometheus-community 的 Helm Charts 仓库。对于新部署,建议优先使用迁移后的维护版本;本文仍以本仓库中的这份图表为准,讲解其配置与实现逻辑,供在既有环境中沿用此版本的用户参考。

二、快速开始:安装与卸载

TL;DR(最简安装)

$ helm install stable/prometheus-node-exporter

该命令会以默认配置在集群中部署 Node Exporter,默认每个节点运行一个 Pod(DaemonSet 形态),并创建对应的 Service、ServiceAccount 与 RBAC 资源。

指定 release 名称安装

$ helm install --name my-release stable/prometheus-node-exporter

安装完成后,release 名为my-release,集群中生成的资源名称遵循my-release-prometheus-node-exporter的命名规则(由 _helpers.tpl 中的fullname模板生成,超过 63 字符会被截断以符合 DNS 规范)。若 release 名已包含 chart 名,则直接使用 release 名,避免重复拼接。

卸载

$ helm delete my-release

该命令会删除该 release 关联的所有 Kubernetes 组件(DaemonSet、Service、ServiceAccount、PSP、ServiceMonitor 等)并删除 release 记录。

三、核心设计:DaemonSet + hostNetwork 的逐节点采集方案

Node Exporter 采集的是节点操作系统指标,因此天然需要“每个节点运行一个实例”。本图表的默认工作负载类型是DaemonSet,其定义见 daemonset.yaml,关键设计如下:

  • hostNetwork: true(默认值):Pod 直接使用宿主机网络栈,配合--web.listen-address=$(HOST_IP):9100监听地址,使指标可在节点 IP 的 9100 端口直接访问;
  • hostPID: true:Pod 共享宿主机 PID 命名空间,配合挂载的/host/proc读取宿主机全部进程信息;
  • hostPath 卷挂载:将宿主机的/proc与/sys分别只读挂载到容器内/host/proc与/host/sys,并通过默认启动参数--path.procfs=/host/proc、--path.sysfs=/host/sys指向它们,从而采集节点级 CPU、内存、磁盘与网络指标;
  • HOST_IP 环境变量:当service.listenOnAllInterfaces为true(默认)时,HOST_IP 固定为0.0.0.0,监听所有接口;否则从status.hostIP字段读取 Pod 被分配到的节点 IP,只监听该地址(见 daemonset.yaml);
  • 探针:容器配置了基于 HTTPGET /的 livenessProbe 与 readinessProbe,探针端口取service.port(默认 9100),便于滚动更新时健康检查;
  • updateStrategy:默认采用RollingUpdate,maxUnavailable: 1,即滚动升级时最多允许 1 个节点上的副本不可用,保证采集能力平滑过渡。

四、完整参数表与默认值

以下为 README 中给出的全部可配置参数,默认值均可在 values.yaml 中对应确认:

ParameterDescriptionDefault
image.repositoryImage repositoryquay.io/prometheus/node-exporter
image.tagImage tagv1.0.1
image.pullPolicyImage pull policyIfNotPresent
extraArgsAdditional container arguments[]
extraHostVolumeMountsAdditional host volume mounts[]
podAnnotationsAnnotations to be added to node exporter pods{}
podLabelsAdditional labels to be added to pods{}
rbac.createIf true, create & use RBAC resourcestrue
rbac.pspEnabledSpecifies whether a PodSecurityPolicy should be created.true
resourcesCPU/Memory resource requests/limits{}
service.typeService typeClusterIP
service.portThe service port9100
service.targetPortThe target port of the container9100
service.nodePortThe node port of the service
service.listenOnAllInterfacesIf true, listen on all interfaces using IP0.0.0.0. Else listen on the IP address pod has been assigned by Kubernetes.true
service.annotationsKubernetes service annotations{prometheus.io/scrape: "true"}
serviceAccount.createSpecifies whether a service account should be created.true
serviceAccount.nameService account to be used. If not set andserviceAccount.createistrue, a name is generated using the fullname template
serviceAccount.imagePullSecretsSpecify image pull secrets[]
securityContextSecurityContextSee values.yaml
affinityA group of affinity scheduling rules for pod assignment{}
nodeSelectorNode labels for pod assignment{}
tolerationsList of node taints to tolerate- effect: NoSchedule operator: Exists
priorityClassNameName of Priority Class to assign podsnil
endpointslist of addresses that have node exporter deployed outside of the cluster[]
hostNetworkWhether to expose the service to the host networktrue
prometheus.monitor.enabledSet this totrueto create ServiceMonitor for Prometheus operatorfalse
prometheus.monitor.additionalLabelsAdditional labels that can be used so ServiceMonitor will be discovered by Prometheus{}
prometheus.monitor.namespacenamespace where servicemonitor resource should be createdthe same namespace as prometheus node exporter
prometheus.monitor.relabelingsRelabelings that should be applied on the ServerMonitor{}
prometheus.monitor.scrapeTimeoutTimeout after which the scrape is ended10s
configmapsAllow mounting additional configmaps.[]
namespaceOverrideOverride the deployment namespace""(Release.Namespace)
updateStrategyConfigure a custom update strategy for the daemonsetRolling update with 1 max unavailable
sidecarsAdditional containers for export metrics to text file[]
sidecarVolumeMountVolume for sidecar containers[]

五、关键参数深入解读

1. 镜像与资源(image / resources)

image: repository: quay.io/prometheus/node-exporter tag: v1.0.1 pullPolicy: IfNotPresent

模板中镜像渲染为"{{ .Values.image.repository }}:{{ .Values.image.tag }}"(见 daemonset.yaml),即默认quay.io/prometheus/node-exporter:v1.0.1。若需使用镜像仓库镜像,可替换image.repository。

resources默认留空({})。values.yaml 中给出了推荐示例:limits: {cpu: 200m, memory: 50Mi}、requests: {cpu: 100m, memory: 30Mi}。按 charts 社区惯例,默认不设资源上限以兼容 Minikube 等小资源环境,生产环境建议显式配置。

2. Service 暴露方式(service.*)

service.yaml 生成的 Service 默认类型为ClusterIP,端口port与targetPort均为 9100,端口名metrics。要点:

  • 默认注解prometheus.io/scrape: "true"供 Prometheus 基于注解的服务发现直接抓取;
  • 当service.type为NodePort且设置了service.nodePort时,模板会为其追加nodePort字段(见 service.yaml);
  • service.targetPort与service.port可分别调整,仓库的 CI 用例 ci/port-values.yaml 演示了将两者同时改为 9102 的场景,用于验证自定义端口下模板仍可正常渲染。

3. 附加启动参数(extraArgs)

Node Exporter 的默认启动参数由模板固定注入:--path.procfs=/host/proc、--path.sysfs=/host/sys、--web.listen-address=$(HOST_IP):<service.port>。需要额外控制采集器时通过extraArgs追加,values.yaml 提供了两个高频示例:

extraArgs: - --collector.diskstats.ignored-devices=^(ram|loop|fd|(h|s|v)d[a-z]|nvme\d+n\d+p)\d+$ - --collector.textfile.directory=/run/prometheus

第一个示例忽略 ram、loop、fd 及 sd/hd/vd/nvme 分区等磁盘设备,减少无关指标噪声;第二个示例启用 textfile collector,从/run/prometheus目录读取自定义文本指标文件(通常由 cron 等任务写出的业务自定义指标),是扩展 Node Exporter 指标能力的最常见做法。

4. 主机目录挂载与 ConfigMap(extraHostVolumeMounts / configmaps)

extraHostVolumeMounts: - name: <mountName> hostPath: <hostPath> mountPath: <mountPath> readOnly: true|false mountPropagation: None|HostToContainer|Bidirectional

extraHostVolumeMounts用于把宿主机额外目录(如/run/prometheus文本指标目录)挂载进容器。模板会同时生成容器volumeMounts与对应的hostPath卷(见 daemonset.yaml),并支持mountPropagation传播模式。

configmaps则用于挂载 ConfigMap 文件:

configmaps: - name: <configMapName> mountPath: <mountPath>

5. 边车容器与文本指标卷(sidecars / sidecarVolumeMount)

这是一个值得展开的高级特性:借助textfile collector,可以在 Node Exporter 旁挂边车(sidecar)容器(如 GPU 厂商的nvidia/dcgm-exporter),由边车把指标写入共享的文本文件目录,再由 Node Exporter 的 textfile collector 统一暴露。配置方式:

sidecars: - name: nvidia-dcgm-exporter image: nvidia/dcgm-exporter:1.4.3 sidecarVolumeMount: - name: collector-textfiles mountPath: /run/prometheus readOnly: false

模板实现上,sidecar 卷以emptyDir(medium: Memory,内存盘)挂载,node-exporter 主容器将该卷挂到同一路径(只读),边车容器则读写共享(见 daemonset.yaml)。将--collector.textfile.directory=/run/prometheus加入extraArgs后即可打通整条链路。

6. 外部节点采集(endpoints)

当部分节点(如裸机服务器、虚拟机)无法运行 Kubernetes Pod,但已经单独部署了 Node Exporter 时,可通过endpoints参数把这些外部地址纳入同一 Service 的 Endpoints:

endpoints: - 10.0.0.1 - 10.0.0.2

endpoints.yaml 会生成同名的 Endpoints 资源,将列出的 IP 全部绑定到metrics端口(固定 9100,TCP),使 Prometheus 通过该 Service 即可同时抓取集群内 DaemonSet Pod 与集群外 Node Exporter。注意:启用endpoints时,集群内 Pod 的 Endpoints 由 DaemonSet 自动维护,二者以同名 Endpoints 共存。

7. 安全上下文与 RBAC / PSP

  • securityContext(values.yaml 默认):fsGroup: 65534、runAsGroup: 65534、runAsNonRoot: true、runAsUser: 65534,即以非 root 用户(nobody,UID 65534)运行,符合最小权限原则;
  • rbac.create / serviceAccount.create:均为true,模板会创建专用 ServiceAccount(名称为fullname模板生成,可用serviceAccount.name覆盖;imagePullSecrets可注入私有镜像拉取凭据,见 serviceaccount.yaml),并在 DaemonSet 中引用;
  • rbac.pspEnabled(默认true):当rbac.create为真时,会创建policy/v1beta1的 PodSecurityPolicy,允许hostPath卷、hostNetwork: true、hostPID: true及 0–65535 的 hostPort(见 psp.yaml),并配套创建 ClusterRole 与 ClusterRoleBinding 授予该 ServiceAccount 使用权限(见 psp-clusterrole.yaml 与 psp-clusterrolebinding.yaml)。PSP 仅对启用 PSP 准入控制的旧集群生效;在较新集群中该能力已被 Pod Security Admission 取代,可视集群版本将rbac.pspEnabled设为false。

8. 调度与更新策略(affinity / nodeSelector / tolerations / updateStrategy)

  • tolerations默认- effect: NoSchedule, operator: Exists,使 Node Exporter 能调度到所有节点(包括被打上污点的节点),这是“每节点一个实例”的兜底保证;
  • nodeSelector可限定在特定架构/系统节点运行(如beta.kubernetes.io/arch: amd64、beta.kubernetes.io/os: linux),混合架构集群常用;
  • affinity支持节点亲和等高级调度规则(values.yaml 附有nodeAffinity注释示例);
  • priorityClassName为空时不指定,可设置为高优先级保证采集任务不被驱逐;
  • updateStrategy默认RollingUpdate / maxUnavailable: 1,模板在 DaemonSet 上原样渲染(见 daemonset.yaml)。

9. 命名空间覆盖(namespaceOverride)

namespaceOverride默认空,此时所有资源部署在Release.Namespace。设置后,模板中所有资源统一使用覆盖后的命名空间(见 _helpers.tpl 中的namespace定义),便于在多命名空间组合部署(combined charts)场景下使用。

六、指定参数的方式:--set 与 values 文件

方式一:--set 逐项指定

$ helm install --name my-release \ --set serviceAccount.name=node-exporter \ stable/prometheus-node-exporter

--set支持key=value[,key=value]的多项逗号分隔语法,适合少量覆盖。

方式二:-f 指定 YAML 文件

$ helm install --name my-release -f values.yaml stable/prometheus-node-exporter

适合将完整的自定义配置(如上面的 sidecar、extraArgs、endpoints 等组合)写入values.yaml后统一注入。两种方式可混用,--set优先级更高。

七、与 Prometheus 集成:注解发现与 ServiceMonitor

图表提供了两条开箱即用的抓取路径:

  1. 注解自动发现:Service 默认带注解prometheus.io/scrape: "true",配合 Prometheus 的kubernetes_sd_configs(role: service/endpoints)+ 注解过滤即可自动纳入抓取,无需额外配置;
  2. ServiceMonitor(Prometheus Operator 模式):设置prometheus.monitor.enabled: true后,monitor.yaml 会生成monitoring.coreos.com/v1的 ServiceMonitor,按app与release标签选择 Service 的metrics端口,并支持:
    • additionalLabels:追加标签便于 Operator 的 ServiceMonitorSelector 发现;
    • scrapeTimeout:抓取超时,默认10s;
    • relabelings:自定义 relabel 规则(如过滤、重命名、追加指标标签);
    • prometheus.monitor.namespace:默认与 Node Exporter 同命名空间,可指定 ServiceMonitor 所在命名空间。

八、验证与访问指标

安装完成后,可按 Service 类型用 NOTES.txt(见 NOTES.txt)中的方式访问:

  • ClusterIP(默认):端口转发后访问本机 9100:
$ kubectl port-forward --namespace <ns> <pod-name> 9100
  • NodePort:
$ export NODE_PORT=$(kubectl get --namespace <ns> -o jsonpath="{.spec.ports[0].nodePort}" services <release>-prometheus-node-exporter) $ export NODE_IP=$(kubectl get nodes --namespace <ns> -o jsonpath="{.items[0].status.addresses[0].address}") $ echo http://$NODE_IP:$NODE_PORT
  • LoadBalancer:等待外部 IP 就绪后访问http://$SERVICE_IP:<port>。

访问后可用curl http://<endpoint>:9100/metrics验证指标输出(如node_cpu_seconds_total、node_memory_*、node_filesystem_*等),配合node_uname_info可确认节点身份。

九、注意事项与适用前提

  • 弃用状态:本仓库中的该图表已标记 deprecated(Chart.yaml中deprecated: true),并迁移至 prometheus-community/helm-charts;新环境建议使用迁移后的维护版本,本文参数与实现逻辑可供对照迁移;
  • hostNetwork 默认开启:Pod 复用宿主机网络,9100 端口暴露于节点网络,请结合网络策略/防火墙评估暴露面;如不希望如此,可设置hostNetwork: false并调整listenOnAllInterfaces;
  • PSP 兼容性:默认生成的 PodSecurityPolicy 属于policy/v1beta1,仅在集群启用了 PSP 准入时生效,新版集群请关闭rbac.pspEnabled;
  • 自定义端口:若通过service.port/service.targetPort更改端口(参考 ci/port-values.yaml),请同步调整 Prometheus 的抓取配置,勿遗漏--web.listen-address中的端口联动关系。

十、参考资源

  • 核心文档:stable/prometheus-node-exporter/README.md
  • 默认配置:values.yaml
  • Chart 元数据:Chart.yaml
  • 工作负载与网络模板:daemonset.yaml、service.yaml、endpoints.yaml
  • Prometheus 集成:monitor.yaml
  • 安全与 RBAC:psp.yaml、serviceaccount.yaml
  • CI 校验示例:ci/port-values.yaml
  • 命名与命名空间模板:templates/_helpers.tpl

【免费下载链接】charts

⚠️(OBSOLETE) Curated applications for Kubernetes

项目地址:https://gitcode.com/gh_mirrors/chart/charts
点击查看免费下载

相关推荐

上一篇:RamaLama命令行自动补全:提升开发效率的小技巧
下一篇:如何破解Charles代理工具:3分钟完成4.2.7版本激活

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

版权声明: 本文来自互联网用户投稿,该文观点仅代表作者本人,不代表本站立场。本站仅提供信息存储空间服务,不拥有所有权,不承担相关法律责任。如若内容造成侵权/违法违规/事实不符,请联系邮箱:809451989@qq.com进行投诉反馈,一经查实,立即删除!
网站建设 2026/10/8 7:51:23

Superpowers开发工具链:本地AI编程增强体系实战指南

1. 项目概述&#xff1a;Superpowers 不是超能力&#xff0c;而是开发者效率革命的代号 “Superpowers”这个词最近在开发者圈子里炸开了锅——它不是漫威电影里的变种人设定&#xff0c;也不是什么玄学概念&#xff0c;而是一套正在真实改变日常编码方式的智能开发工具链统称…

作者头像 李华
网站建设 2026/10/8 7:50:54

yfinance 取数:3 行代码拿数据,4 类异常一次修完的实用速查

yfinance 取数&#xff1a;3 行代码拿数据&#xff0c;4 类异常一次修完的实用速查 【免费下载链接】yfinance Download market data from Yahoo! Finances API 项目地址: https://gitcode.com/GitHub_Trending/yf/yfinance 写回测脚本&#xff0c;第一道坎往往是数据&a…

作者头像 李华
网站建设 2026/10/8 7:50:43

OpenShell 实战:浏览器交互式 Shell 的会话保持与部署避坑指南

1. 从一个空输入框说起&#xff1a;OpenShell 到底在解决什么问题第一次看到“OpenShell”这个词&#xff0c;是在一个终端工具讨论帖里。有人丢出一句话&#xff1a;“有没有那种能让我在浏览器里直接开一个像样的 shell&#xff0c;又不用装一堆东西的方案&#xff1f;”底下…

作者头像 李华
网站建设 2026/10/8 7:49:08

第116篇 Kotlin 代码评审清单:团队规范与静态检查

前面几节讲的是"怎么写",这一节讲"怎么保证团队写出来的代码跟你一样好"。这题的面试形态很特别——它不考语法也不考 API,考的是你有没有一套能落地的质量保障机制。答得浅的人说"团队 review 很严格",答得深的人会讲清"哪些交给机器、…

作者头像 李华
网站建设 2026/10/8 7:46:34

Meson Qt4 模块实战指南:moc/uic/rcc 工具链集成与 Qt4 项目构建配置

构建工具 【免费下载链接】meson The Meson Build System 项目地址&#xff1a; https://gitcode.com/gh_mirrors/me/meson 点击查看 免费下载 本文面向仍在维护 Qt4 遗留代码库、或需要理解 Meson Qt 模块统一抽象机制的开发者&#xff0c;系统讲解 Meson 中 qt4 模块的加载方…

作者头像 李华