- 网络安全
- 网络
- IDS
【免费下载链接】zeek
Zeek is a powerful network analysis framework that is much different from the typical IDS you may know.
导读
本文围绕 Zeek 框架自动生成的 API 参考文档 doc/scripts/base/bif/telemetry_consts.bif.zeek.rst 中声明的两个全局常量展开:Telemetry::callback_timeout(interval 类型)与Telemetry::civetweb_threads(count 类型)。它们共同控制 Zeek 内置 Prometheus 指标服务端(基于 CivetWeb)的线程规模与回调超时策略,是集群化部署中配置指标抓取并发与稳定性的关键旋钮。读完本文,你将掌握这两个常量的确切默认值、它们在 src/telemetry/Manager.cc 中的真实调用链,以及如何通过redef在local.zeek中按需调整。
一、常量出处:BIF 声明、脚本默认值与文档生成
Zeek 的遥测常量以 BIF(Built-In Function/Constant)形式声明于 C++ 侧,由 Zeek 的文档生成工具(zeekygen)自动渲染为 RST 参考页。三个相关文件构成完整的"声明—默认值—文档"链路:
| 环节 | 文件 | 内容 |
|---|---|---|
| 类型声明(C++ BIF) | src/telemetry/telemetry_consts.bif | 仅两行:const Telemetry::callback_timeout: interval;与const Telemetry::civetweb_threads: count; |
| 脚本默认值 | scripts/base/init-bare.zeek | 在module Telemetry; export { ... }块内给出默认值并标记&redef |
| 文档参考页 | doc/scripts/base/bif/telemetry_consts.bif.zeek.rst | 本文所依托的 API 参考页,位于base/bif/系列生成的文档中 |
注意参考页中的:tocdepth: 3与Summary / Detailed Interface结构是 zeekygen 对所有 BIF 参考页的统一排版模板,其核心信息正是文档正文所指出的两个常量及其命名空间归属(Telemetry::前缀表明二者位于Telemetry模块,尽管参考页顶部标有 GLOBAL 命名空间标记,实际以 BIF 声明为准)。
两个常量的脚本默认值及原始注释如下(scripts/base/init-bare.zeek):
module Telemetry; export { ## Maximum amount of time for CivetWeb HTTP threads to ## wait for metric callbacks to complete on the IO loop. const callback_timeout: interval = 5sec &redef; ## Number of CivetWeb threads to use. const civetweb_threads: count = 2 &redef; }两处关键语义:默认值分别为5sec与2;&redef属性允许用户在任何加载的脚本(如site/local.zeek)中安全地重新定义。
二、Telemetry::civetweb_threads:Prometheus 抓取服务的线程池
2.1 语义与默认值
该常量指定 Zeek 内置 Prometheus HTTP 服务所使用的 CivetWeb 线程数量,默认2。它直接决定能同时服务多少个并发 HTTP 抓取请求(scrape)。
2.2 源码调用链
在 src/telemetry/Manager.cc 的ListenPrometheus()中,该常量被传入prometheus::Exposer构造器:
std::string prometheus_url = util::fmt("%.*s:%u", static_cast<int>(metrics_address.size()), metrics_address.data(), metrics_port); try { prometheus_exposer = std::make_unique<prometheus::Exposer>(prometheus_url, BifConst::Telemetry::civetweb_threads, callbacks); // ... } catch ( const CivetException& exc ) { reporter->FatalError("Failed to setup Prometheus endpoint: %s. Attempted to bind to %s.", exc.what(), prometheus_url.c_str()); }也就是说:civetweb_threads的数值直接作为 CivetWeb 服务器线程数;若端口绑定失败(例如被占用),Zeek 会以FatalError终止并报告绑定地址。绑定成功后,ZeekCollectable与prometheus_registry按顺序注册为 collector,抓取时先更新指标值再渲染文本。
2.3 调优要点
- 默认
2适用于单节点、低频抓取(如每 15s 一次); - 若 Prometheus 以高频率(如 5s 间隔)并行抓取多个集群节点,或存在多个 scraper,可适当增大线程数,避免抓取请求排队;
- 增大线程数会带来额外的内存与调度开销,建议结合 doc/frameworks/telemetry.rst 中的集群服务发现方案,先控制抓取并发再决定线程数。
三、Telemetry::callback_timeout:指标回调的最大等待时间
3.1 语义与默认值
该常量定义 CivetWeb HTTP 线程在 IO 循环上等待指标回调完成的最大时间,默认5sec。它的作用场景是:Prometheus 抓取/metrics时,CivetWeb 线程必须等 Zeek 主事件循环完成指标采集回调后,才能拿到最新数据并返回响应。
3.2 源码调用链
在 src/telemetry/Manager.cc 中可以看到完整的"唤醒—等待—超时"机制:
void Manager::ProcessFd(int fd, int flags) { // 由 collector_flare 唤醒,采集并更新指标 collector_flare.Extinguish(); UpdateMetrics(); collector_response_idx = collector_request_idx; collector_cv.notify_all(); } void Manager::WaitForPrometheusCallbacks() { ++collector_request_idx; uint64_t expected_idx = collector_request_idx; collector_flare.Fire(); // 源码注释明确指出:正常情况下遍历全部回调不应花费 5 秒, // 设置超时只是为了避免死锁。 bool res = collector_cv.wait_for( lk, std::chrono::microseconds( static_cast<long>(BifConst::Telemetry::callback_timeout * 1000000)), [expected_idx]() { return telemetry_mgr->collector_response_idx >= expected_idx || zeek::run_state::terminating; }); if ( ! res ) fprintf(stderr, "Timeout waiting for prometheus callbacks\n"); }关键细节:
callback_timeout以interval(秒)为单位,在 C++ 侧乘以1000000转换为微秒后传给std::condition_variable::wait_for;- 内部通过
collector_flare(自管文件描述符)+ 条件变量实现跨线程唤醒,ProcessFd由 IO 循环驱动; - 若在超时时间内回调未完成,会向 stderr 输出
Timeout waiting for prometheus callbacks,但不会崩溃,仅返回过期数据——这是刻意设计的防死锁兜底。
3.3 何时需要调大
- 采集的指标族(metric family)数量极大,或存在重量级回调(例如每次抓取遍历海量表的 size 指标)时,5 秒可能不够;
- 集群规模较大且单节点指标繁多、磁盘 IO 抖动时,可适度增大到
10sec或15sec; - 若 stderr 中频繁出现上述超时消息,说明回调耗时已逼近上限,此时应优先排查单次回调的复杂度,而非盲目加大超时。
四、实战:在 local.zeek 中 redef 并验证
4.1 配置示例
与 Zeek 其他常量一样,这两个常量支持在站点脚本中重新定义,例如 scripts/site/local.zeek:
@load base/frameworks/telemetry # 放宽指标回调等待时间,适配指标量较大的部署 redef Telemetry::callback_timeout = 10sec; # 提升 CivetWeb 线程数,支持更多并发抓取 redef Telemetry::civetweb_threads = 4;4.2 先决条件:开启指标端口
仅调整上述两个常量不会自动开启 HTTP 服务,还需设置Telemetry::metrics_port。根据 doc/frameworks/telemetry.rst:其默认值为0/unknown(禁用),设置为具体 TCP 端口即启用。集群场景下Cluster::Node的 metrics 端口字段会自动覆盖该值,也可手工指定:
redef Telemetry::metrics_port = 9090/tcp;4.3 验证效果
服务开启后,用curl直接验证线程与超时配置生效:
curl -s http://<node>:9090/metrics响应中会包含exposer_transferred_bytes_total、zeek_event_handler_invocations_total等指标。若要验证超时路径,可临时将callback_timeout调小并制造高负载抓取,观察 stderr 是否出现Timeout waiting for prometheus callbacks。
4.4 集群场景的补充说明
在集群部署中,Zeek 7.0 起移除了向 manager 内置聚合遥测的功能,改为在 manager 节点暴露http://<manager>:<manager-metrics-port>/services.json服务发现端点(见 src/telemetry/Manager.cc 中begin_request回调对/services.json的处理),由 Prometheus 的http_sd_configs拉取全部节点端点后自行聚合。此时各节点的 CivetWeb 线程数与回调超时需要按"被多少 scraper 同时抓取"来统一规划。
五、与 MetricType 枚举的关系
文档参考页 doc/scripts/base/bif/telemetry_types.bif.zeek.rst 还展示了Telemetry::MetricType枚举(COUNTER、GAUGE、HISTOGRAM)。它与本文两个常量的关系在于:指标类型的多样性直接决定了抓取回调的工作量——例如HISTOGRAM需要按预定义桶(bounds)聚合观察值,回调成本高于简单COUNTER;当部署中包含大量直方图指标时,callback_timeout的默认 5 秒更可能被触及,这正是需要关注该常量的典型场景。
六、总结
Telemetry::callback_timeout与Telemetry::civetweb_threads虽只是两行常量声明(src/telemetry/telemetry_consts.bif),却是 Zeek 内置 Prometheus 指标服务可用性的两条生命线:
- civetweb_threads(默认 2)控制 HTTP 线程池规模,决定并发抓取能力;
- callback_timeout(默认 5sec)控制抓取等待回调的上限,是防死锁的兜底机制,超时时仅降级返回旧数据。
二者均在 scripts/base/init-bare.zeek 中以&redef提供默认值,运维者可在local.zeek中按集群规模与指标量级重新定义;底层行为由 src/telemetry/Manager.cc 的ListenPrometheus()与WaitForPrometheusCallbacks()实现,可通过/metrics抓取与 stderr 超时日志进行验证。掌握这两个常量,即可在 Zeek 集群化部署中精准控制指标采集的并发与稳定性。
- 网络安全
- 网络
- IDS
【免费下载链接】zeek
Zeek is a powerful network analysis framework that is much different from the typical IDS you may know.
相关推荐
axios 配置默认值详解:全局默认值、实例默认值与配置优先级(附源码解析)
axios 配置默认值详解:全局默认值、实例默认值与配置优先级(附源码解析) axios 允许为每个请求指定配置默认值,包括 baseURL 、 headers
网络后端前端VSCodium 彻底清除遥测(Telemetry)的完整指南:默认设置、源码替换与隐私验证
VSCodium 彻底清除遥测(Telemetry)的完整指南:默认设置、源码替换与隐私验证 本篇指南系统讲解 VSCodium(无微软品牌、遥测与许可限制的
开发工具Claude Code 插件遥测实践:内置 telemetry mod 与 `$.telemetry` 接口完全解析
Claude Code 插件遥测实践:内置 telemetry mod 与 $.telemetry 接口完全解析 本篇文章以 mods/telemetry/RE
AI 应用AI 技能/插件开发工具
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考