【免费下载链接】charts
⚠️(OBSOLETE) Curated applications for Kubernetes
导读
本文以 charts 仓库中的 stable/prometheus-node-exporter/README.md 为核心文档,系统讲解如何在 Kubernetes 集群上通过 Helm 部署 Prometheus Node Exporter,覆盖安装/卸载、全部可配置参数、参数默认值,以及 DaemonSet、Service、ServiceMonitor、RBAC 等底层模板实现。读完本文,你将掌握用一条helm install命令拉起全节点指标采集、按需调整镜像与端口、打通 Prometheus 服务发现,并为外部已部署的 Node Exporter 补充 Endpoints 的完整实战方案。
一、背景:为什么用 Node Exporter 监控 Kubernetes 节点
Prometheus Node Exporter 是 Prometheus 生态中最常用的主机指标采集器,负责暴露节点的 CPU、内存、磁盘、网络、文件系统等操作系统级指标。在 Kubernetes 中,节点层面的指标(如node_cpu_seconds_total、node_memory_*、node_filesystem_*)都来自 Node Exporter,是集群可观测性建设的基础组件。
本仓库中的prometheus-node-exporter是一个 Helm 图表(chart),其元数据定义在 Chart.yaml:name: prometheus-node-exporter,version: 1.11.2,appVersion: 1.0.1(即默认打包的 Node Exporter 镜像版本为 v1.0.1)。需要特别说明的是,该图表在仓库中已被标记为deprecated: true,README 首行即声明 "DEPRECATED and moved to https://github.com/prometheus-community/helm-charts"——即该项目已停止维护并迁移到 prometheus-community 的 Helm Charts 仓库。对于新部署,建议优先使用迁移后的维护版本;本文仍以本仓库中的这份图表为准,讲解其配置与实现逻辑,供在既有环境中沿用此版本的用户参考。
二、快速开始:安装与卸载
TL;DR(最简安装)
$ helm install stable/prometheus-node-exporter该命令会以默认配置在集群中部署 Node Exporter,默认每个节点运行一个 Pod(DaemonSet 形态),并创建对应的 Service、ServiceAccount 与 RBAC 资源。
指定 release 名称安装
$ helm install --name my-release stable/prometheus-node-exporter安装完成后,release 名为my-release,集群中生成的资源名称遵循my-release-prometheus-node-exporter的命名规则(由 _helpers.tpl 中的fullname模板生成,超过 63 字符会被截断以符合 DNS 规范)。若 release 名已包含 chart 名,则直接使用 release 名,避免重复拼接。
卸载
$ helm delete my-release该命令会删除该 release 关联的所有 Kubernetes 组件(DaemonSet、Service、ServiceAccount、PSP、ServiceMonitor 等)并删除 release 记录。
三、核心设计:DaemonSet + hostNetwork 的逐节点采集方案
Node Exporter 采集的是节点操作系统指标,因此天然需要“每个节点运行一个实例”。本图表的默认工作负载类型是DaemonSet,其定义见 daemonset.yaml,关键设计如下:
- hostNetwork: true(默认值):Pod 直接使用宿主机网络栈,配合
--web.listen-address=$(HOST_IP):9100监听地址,使指标可在节点 IP 的 9100 端口直接访问; - hostPID: true:Pod 共享宿主机 PID 命名空间,配合挂载的
/host/proc读取宿主机全部进程信息; - hostPath 卷挂载:将宿主机的
/proc与/sys分别只读挂载到容器内/host/proc与/host/sys,并通过默认启动参数--path.procfs=/host/proc、--path.sysfs=/host/sys指向它们,从而采集节点级 CPU、内存、磁盘与网络指标; - HOST_IP 环境变量:当
service.listenOnAllInterfaces为true(默认)时,HOST_IP 固定为0.0.0.0,监听所有接口;否则从status.hostIP字段读取 Pod 被分配到的节点 IP,只监听该地址(见 daemonset.yaml); - 探针:容器配置了基于 HTTP
GET /的 livenessProbe 与 readinessProbe,探针端口取service.port(默认 9100),便于滚动更新时健康检查; - updateStrategy:默认采用
RollingUpdate,maxUnavailable: 1,即滚动升级时最多允许 1 个节点上的副本不可用,保证采集能力平滑过渡。
四、完整参数表与默认值
以下为 README 中给出的全部可配置参数,默认值均可在 values.yaml 中对应确认:
| Parameter | Description | Default |
|---|---|---|
image.repository | Image repository | quay.io/prometheus/node-exporter |
image.tag | Image tag | v1.0.1 |
image.pullPolicy | Image pull policy | IfNotPresent |
extraArgs | Additional container arguments | [] |
extraHostVolumeMounts | Additional host volume mounts | [] |
podAnnotations | Annotations to be added to node exporter pods | {} |
podLabels | Additional labels to be added to pods | {} |
rbac.create | If true, create & use RBAC resources | true |
rbac.pspEnabled | Specifies whether a PodSecurityPolicy should be created. | true |
resources | CPU/Memory resource requests/limits | {} |
service.type | Service type | ClusterIP |
service.port | The service port | 9100 |
service.targetPort | The target port of the container | 9100 |
service.nodePort | The node port of the service | |
service.listenOnAllInterfaces | If true, listen on all interfaces using IP0.0.0.0. Else listen on the IP address pod has been assigned by Kubernetes. | true |
service.annotations | Kubernetes service annotations | {prometheus.io/scrape: "true"} |
serviceAccount.create | Specifies whether a service account should be created. | true |
serviceAccount.name | Service account to be used. If not set andserviceAccount.createistrue, a name is generated using the fullname template | |
serviceAccount.imagePullSecrets | Specify image pull secrets | [] |
securityContext | SecurityContext | See values.yaml |
affinity | A group of affinity scheduling rules for pod assignment | {} |
nodeSelector | Node labels for pod assignment | {} |
tolerations | List of node taints to tolerate | - effect: NoSchedule operator: Exists |
priorityClassName | Name of Priority Class to assign pods | nil |
endpoints | list of addresses that have node exporter deployed outside of the cluster | [] |
hostNetwork | Whether to expose the service to the host network | true |
prometheus.monitor.enabled | Set this totrueto create ServiceMonitor for Prometheus operator | false |
prometheus.monitor.additionalLabels | Additional labels that can be used so ServiceMonitor will be discovered by Prometheus | {} |
prometheus.monitor.namespace | namespace where servicemonitor resource should be created | the same namespace as prometheus node exporter |
prometheus.monitor.relabelings | Relabelings that should be applied on the ServerMonitor | {} |
prometheus.monitor.scrapeTimeout | Timeout after which the scrape is ended | 10s |
configmaps | Allow mounting additional configmaps. | [] |
namespaceOverride | Override the deployment namespace | ""(Release.Namespace) |
updateStrategy | Configure a custom update strategy for the daemonset | Rolling update with 1 max unavailable |
sidecars | Additional containers for export metrics to text file | [] |
sidecarVolumeMount | Volume for sidecar containers | [] |
五、关键参数深入解读
1. 镜像与资源(image / resources)
image: repository: quay.io/prometheus/node-exporter tag: v1.0.1 pullPolicy: IfNotPresent模板中镜像渲染为"{{ .Values.image.repository }}:{{ .Values.image.tag }}"(见 daemonset.yaml),即默认quay.io/prometheus/node-exporter:v1.0.1。若需使用镜像仓库镜像,可替换image.repository。
resources默认留空({})。values.yaml 中给出了推荐示例:limits: {cpu: 200m, memory: 50Mi}、requests: {cpu: 100m, memory: 30Mi}。按 charts 社区惯例,默认不设资源上限以兼容 Minikube 等小资源环境,生产环境建议显式配置。
2. Service 暴露方式(service.*)
service.yaml 生成的 Service 默认类型为ClusterIP,端口port与targetPort均为 9100,端口名metrics。要点:
- 默认注解
prometheus.io/scrape: "true"供 Prometheus 基于注解的服务发现直接抓取; - 当
service.type为NodePort且设置了service.nodePort时,模板会为其追加nodePort字段(见 service.yaml); service.targetPort与service.port可分别调整,仓库的 CI 用例 ci/port-values.yaml 演示了将两者同时改为 9102 的场景,用于验证自定义端口下模板仍可正常渲染。
3. 附加启动参数(extraArgs)
Node Exporter 的默认启动参数由模板固定注入:--path.procfs=/host/proc、--path.sysfs=/host/sys、--web.listen-address=$(HOST_IP):<service.port>。需要额外控制采集器时通过extraArgs追加,values.yaml 提供了两个高频示例:
extraArgs: - --collector.diskstats.ignored-devices=^(ram|loop|fd|(h|s|v)d[a-z]|nvme\d+n\d+p)\d+$ - --collector.textfile.directory=/run/prometheus第一个示例忽略 ram、loop、fd 及 sd/hd/vd/nvme 分区等磁盘设备,减少无关指标噪声;第二个示例启用 textfile collector,从/run/prometheus目录读取自定义文本指标文件(通常由 cron 等任务写出的业务自定义指标),是扩展 Node Exporter 指标能力的最常见做法。
4. 主机目录挂载与 ConfigMap(extraHostVolumeMounts / configmaps)
extraHostVolumeMounts: - name: <mountName> hostPath: <hostPath> mountPath: <mountPath> readOnly: true|false mountPropagation: None|HostToContainer|BidirectionalextraHostVolumeMounts用于把宿主机额外目录(如/run/prometheus文本指标目录)挂载进容器。模板会同时生成容器volumeMounts与对应的hostPath卷(见 daemonset.yaml),并支持mountPropagation传播模式。
configmaps则用于挂载 ConfigMap 文件:
configmaps: - name: <configMapName> mountPath: <mountPath>5. 边车容器与文本指标卷(sidecars / sidecarVolumeMount)
这是一个值得展开的高级特性:借助textfile collector,可以在 Node Exporter 旁挂边车(sidecar)容器(如 GPU 厂商的nvidia/dcgm-exporter),由边车把指标写入共享的文本文件目录,再由 Node Exporter 的 textfile collector 统一暴露。配置方式:
sidecars: - name: nvidia-dcgm-exporter image: nvidia/dcgm-exporter:1.4.3 sidecarVolumeMount: - name: collector-textfiles mountPath: /run/prometheus readOnly: false模板实现上,sidecar 卷以emptyDir(medium: Memory,内存盘)挂载,node-exporter 主容器将该卷挂到同一路径(只读),边车容器则读写共享(见 daemonset.yaml)。将--collector.textfile.directory=/run/prometheus加入extraArgs后即可打通整条链路。
6. 外部节点采集(endpoints)
当部分节点(如裸机服务器、虚拟机)无法运行 Kubernetes Pod,但已经单独部署了 Node Exporter 时,可通过endpoints参数把这些外部地址纳入同一 Service 的 Endpoints:
endpoints: - 10.0.0.1 - 10.0.0.2endpoints.yaml 会生成同名的 Endpoints 资源,将列出的 IP 全部绑定到metrics端口(固定 9100,TCP),使 Prometheus 通过该 Service 即可同时抓取集群内 DaemonSet Pod 与集群外 Node Exporter。注意:启用endpoints时,集群内 Pod 的 Endpoints 由 DaemonSet 自动维护,二者以同名 Endpoints 共存。
7. 安全上下文与 RBAC / PSP
- securityContext(values.yaml 默认):
fsGroup: 65534、runAsGroup: 65534、runAsNonRoot: true、runAsUser: 65534,即以非 root 用户(nobody,UID 65534)运行,符合最小权限原则; - rbac.create / serviceAccount.create:均为
true,模板会创建专用 ServiceAccount(名称为fullname模板生成,可用serviceAccount.name覆盖;imagePullSecrets可注入私有镜像拉取凭据,见 serviceaccount.yaml),并在 DaemonSet 中引用; - rbac.pspEnabled(默认
true):当rbac.create为真时,会创建policy/v1beta1的 PodSecurityPolicy,允许hostPath卷、hostNetwork: true、hostPID: true及 0–65535 的 hostPort(见 psp.yaml),并配套创建 ClusterRole 与 ClusterRoleBinding 授予该 ServiceAccount 使用权限(见 psp-clusterrole.yaml 与 psp-clusterrolebinding.yaml)。PSP 仅对启用 PSP 准入控制的旧集群生效;在较新集群中该能力已被 Pod Security Admission 取代,可视集群版本将rbac.pspEnabled设为false。
8. 调度与更新策略(affinity / nodeSelector / tolerations / updateStrategy)
tolerations默认- effect: NoSchedule, operator: Exists,使 Node Exporter 能调度到所有节点(包括被打上污点的节点),这是“每节点一个实例”的兜底保证;nodeSelector可限定在特定架构/系统节点运行(如beta.kubernetes.io/arch: amd64、beta.kubernetes.io/os: linux),混合架构集群常用;affinity支持节点亲和等高级调度规则(values.yaml 附有nodeAffinity注释示例);priorityClassName为空时不指定,可设置为高优先级保证采集任务不被驱逐;updateStrategy默认RollingUpdate / maxUnavailable: 1,模板在 DaemonSet 上原样渲染(见 daemonset.yaml)。
9. 命名空间覆盖(namespaceOverride)
namespaceOverride默认空,此时所有资源部署在Release.Namespace。设置后,模板中所有资源统一使用覆盖后的命名空间(见 _helpers.tpl 中的namespace定义),便于在多命名空间组合部署(combined charts)场景下使用。
六、指定参数的方式:--set 与 values 文件
方式一:--set 逐项指定
$ helm install --name my-release \ --set serviceAccount.name=node-exporter \ stable/prometheus-node-exporter--set支持key=value[,key=value]的多项逗号分隔语法,适合少量覆盖。
方式二:-f 指定 YAML 文件
$ helm install --name my-release -f values.yaml stable/prometheus-node-exporter适合将完整的自定义配置(如上面的 sidecar、extraArgs、endpoints 等组合)写入values.yaml后统一注入。两种方式可混用,--set优先级更高。
七、与 Prometheus 集成:注解发现与 ServiceMonitor
图表提供了两条开箱即用的抓取路径:
- 注解自动发现:Service 默认带注解
prometheus.io/scrape: "true",配合 Prometheus 的kubernetes_sd_configs(role: service/endpoints)+ 注解过滤即可自动纳入抓取,无需额外配置; - ServiceMonitor(Prometheus Operator 模式):设置
prometheus.monitor.enabled: true后,monitor.yaml 会生成monitoring.coreos.com/v1的 ServiceMonitor,按app与release标签选择 Service 的metrics端口,并支持:additionalLabels:追加标签便于 Operator 的 ServiceMonitorSelector 发现;scrapeTimeout:抓取超时,默认10s;relabelings:自定义 relabel 规则(如过滤、重命名、追加指标标签);prometheus.monitor.namespace:默认与 Node Exporter 同命名空间,可指定 ServiceMonitor 所在命名空间。
八、验证与访问指标
安装完成后,可按 Service 类型用 NOTES.txt(见 NOTES.txt)中的方式访问:
- ClusterIP(默认):端口转发后访问本机 9100:
$ kubectl port-forward --namespace <ns> <pod-name> 9100- NodePort:
$ export NODE_PORT=$(kubectl get --namespace <ns> -o jsonpath="{.spec.ports[0].nodePort}" services <release>-prometheus-node-exporter) $ export NODE_IP=$(kubectl get nodes --namespace <ns> -o jsonpath="{.items[0].status.addresses[0].address}") $ echo http://$NODE_IP:$NODE_PORT- LoadBalancer:等待外部 IP 就绪后访问
http://$SERVICE_IP:<port>。
访问后可用curl http://<endpoint>:9100/metrics验证指标输出(如node_cpu_seconds_total、node_memory_*、node_filesystem_*等),配合node_uname_info可确认节点身份。
九、注意事项与适用前提
- 弃用状态:本仓库中的该图表已标记 deprecated(
Chart.yaml中deprecated: true),并迁移至 prometheus-community/helm-charts;新环境建议使用迁移后的维护版本,本文参数与实现逻辑可供对照迁移; - hostNetwork 默认开启:Pod 复用宿主机网络,9100 端口暴露于节点网络,请结合网络策略/防火墙评估暴露面;如不希望如此,可设置
hostNetwork: false并调整listenOnAllInterfaces; - PSP 兼容性:默认生成的 PodSecurityPolicy 属于
policy/v1beta1,仅在集群启用了 PSP 准入时生效,新版集群请关闭rbac.pspEnabled; - 自定义端口:若通过
service.port/service.targetPort更改端口(参考 ci/port-values.yaml),请同步调整 Prometheus 的抓取配置,勿遗漏--web.listen-address中的端口联动关系。
十、参考资源
- 核心文档:stable/prometheus-node-exporter/README.md
- 默认配置:values.yaml
- Chart 元数据:Chart.yaml
- 工作负载与网络模板:daemonset.yaml、service.yaml、endpoints.yaml
- Prometheus 集成:monitor.yaml
- 安全与 RBAC:psp.yaml、serviceaccount.yaml
- CI 校验示例:ci/port-values.yaml
- 命名与命名空间模板:templates/_helpers.tpl
【免费下载链接】charts
⚠️(OBSOLETE) Curated applications for Kubernetes
相关推荐
终极指南:如何使用werf实现Kubernetes应用的自动化测试全流程
终极指南:如何使用werf实现Kubernetes应用的自动化测试全流程 werf是一个强大的解决方案,旨在为Kubernetes实现高效且一致的软件交付流程,
人工智能AI 应用语音移动开发后端桌面应用智能硬件MCP 服务Alluxio Kubernetes 集群监控部署指南:基于 Prometheus + Grafana 的 Monitor Helm Chart 实践
Alluxio Kubernetes 集群监控部署指南:基于 Prometheus + Grafana 的 Monitor Helm Chart 实践 导读 本
存储分布式文件系统缓存大数据基于 helm/charts 仓库的 ChartMuseum 部署实战:在 Kubernetes 上自建私有 Helm Chart 仓库
基于 helm/charts 仓库的 ChartMuseum 部署实战:在 Kubernetes 上自建私有 Helm Chart 仓库 ChartMuseum
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考