- 人工智能
- AI Agent
- Agent 框架
- 大模型
- 工具调用
- RAG
- 提示工程
- 强化学习
【免费下载链接】agent-core
openJiuwen agent-core可提供AI Agent开发、运行、调优与演进相关的全套SDK能力
导读
本文以 examples/security_rail_demo 为例,系统讲解 openJiuwen agent-core 中BaseSecurityRail安全护栏(Security Rail)框架的设计思想与完整实战方法:从 REJECT(拦截)、INTERRUPT(人工审批 HITL)、ALERT(放行但告警)三种决策模式,到目录/类命名硬性规范、扩展加载与验证、事件选择与优先级编排,再到直接可复用的五个示例 Rail 的源码级剖析。读完本文,你将能够独立编写、注册并测试自己的安全 Rail,在工具调用与模型调用两个层面阻止 API Key、Secret、Token 等敏感信息的泄露。
一、安全 Rail 是什么:拦截 Agent 操作的统一检查点
在 openJiuwen agent-core 中,Agent 的一次完整运行由若干生命周期事件串联(如模型调用、工具执行等)。BaseSecurityRail正是在这些事件点上提供的一套结构化安全检查模式,让开发者以声明式方式拦截 Agent 操作。其核心实现在 openjiuwen/harness/rails/security/base_security_rail.py,并作为SecurityRail体系由 openjiuwen/harness/rails/security/init.py 与 openjiuwen/harness/rails/init.py 对外导出(BaseSecurityRail、SecurityAllow、SecurityReject、SecurityInterrupt、SecurityCheckContext等均在__all__中)。
BaseSecurityRail提供三种基础决策类型:
| 决策类型 | 语义 | 典型用途 |
|---|---|---|
SecurityAllow | 放行,操作继续 | 检查未命中敏感内容 |
SecurityReject | 拦截,返回错误消息 | 工具结果泄露密钥时阻断 |
SecurityInterrupt | 暂停等待人工审批(HITL) | 执行前需用户确认 |
除上述三种外,源码中还提供了第四种SecurityAlert(base_security_rail.py#L84-L103):放行但向用户告警,支持info/warning/error/critical四个告警级别与popup/history/inline三种展示模式,这在不需要阻断业务流程、仅需监控的场景下非常有用。
从继承体系看,BaseSecurityRail派生自核心事件框架中的AgentRail(openjiuwen/core/single_agent/rail/base.py#L832),而 harness 层又在其上扩展出DeepAgentRail(openjiuwen/harness/rails/base.py#L28-L93),额外提供before_task_iteration/after_task_iteration两个外层任务循环钩子。因此自定义安全 Rail 既可以继承BaseSecurityRail实现纯安全检查,也可以继承DeepAgentRail参与任务循环生命周期。
二、快速安装:把示例 Rail 复制到 jiuwenclaw 扩展目录
examples/security_rail_demo/README.md 给出了最直接的接入方式——将示例文件夹复制到 jiuwenclaw 的 extensions 目录并重启:
# 复制 Reject 与 Interrupt 示例 cp -r examples/security_rail_demo/ApiKeyGuardReject ~/guardrail/extensions/ cp -r examples/security_rail_demo/ApiKeyGuardInterrupt ~/guardrail/extensions/ # 复制配置 cp examples/security_rail_demo/example_config.json ~/guardrail/extensions/extensions_config.json然后重启 jiuwenclaw,并在配置中一次只启用一个扩展,逐一测试每种模式。
需要说明的是,当前仓库 examples/security_rail_demo 目录下实际包含五个示例子目录:ApiKeyGuardAlert、ApiKeyGuardInterrupt、ModelCallGuard、SensitiveDataSanitize、ToolRejectExample。其中ApiKeyGuardAlert即为 REJECT 之外的四态示例中的 ALERT 模式。各示例的定位对比如下表:
| 示例文件夹 | 决策模式 | 类名 | 行为 |
|---|---|---|---|
ApiKeyGuardAlert | ALERT | ApikeyguardalertRail | 放行执行但向前端告警 |
ApiKeyGuardInterrupt | INTERRUPT | ApikeyguardinterruptRail | 需要人工审批 |
ModelCallGuard | REJECT | ModelcallguardRail | 阻断 LLM 输入/输出中的密钥 |
SensitiveDataSanitize | SANITIZE | SensitivedatasanitizeRail | 用[REDACTED]替换密钥但不阻断 |
ToolRejectExample | REJECT | ToolrejectexampleRail | 直接拒绝,演示 BEFORE/AFTER 行为差异 |
配套的 example_config.json 给出了配置格式示例,其中enabled字段控制扩展开关,class_name指定要实例化的 Rail 类,priority决定执行优先级:
{ "ApiKeyGuardReject": { "name": "ApiKeyGuardReject", "class_name": "ApikeyguardrejectRail", "enabled": false, "description": "Blocks API key/secret leakage completely (REJECT mode)", "priority": 89 }, "ApiKeyGuardInterrupt": { "name": "ApiKeyGuardInterrupt", "class_name": "ApikeyguardinterruptRail", "enabled": false, "description": "Requires human approval for API key operations (INTERRUPT mode)", "priority": 89 } }从 harness 的元加载机制看(openjiuwen/harness/manifest/meta_elements.py#L477-L493),harness.rail.file类型的资源会按params.file_path + params.class_name从指定 Python 文件动态加载类并实例化,实例化后还会校验其必须是AgentRail子类;扩展清单(如 openjiuwen/harness/resources/extension_loader.py 中的_build_rail_specs)负责把 manifest 中的 rail 条目解析成RailSpec。这就是"复制目录 + 写配置 + 重启自动加载"能够生效的底层原因。
三、目录结构与命名规范(RailManager 的硬性约定)
关键警告:jiuwenclaw 的 RailManager 对目录名、文件名、类名、配置键有一套严格的命名约定,写错任何一个都无法被加载。
目录结构
~/guardrail/extensions/ApiKeyGuardReject/ # 扩展目录(CapitalCamelCase) ├── rail.py # 强制要求:必须命名为 "rail.py" └── (可选: __init__.py) # 支持相对导入命名规则
| 元素 | 规则 | 示例 |
|---|---|---|
| 目录名 | CapitalCamelCase(大写驼峰) | ApiKeyGuardReject、MyCustom |
| 文件名 | 必须为rail.py | rail.py(不能是api_key_guard_rail.py) |
| 类名 | {目录首字母大写}{目录其余字母小写}Rail | ApiKeyGuardReject→ApikeyguardrejectRail |
| 配置键 | 必须与目录名一致 | "ApiKeyGuardReject" |
类命名模式
# 目录: ApiKeyGuardReject → 类: ApikeyguardrejectRail # 目录: MyCustom → 类: MycustomRail # 规则: 首字母大写,其余小写,末尾加 "Rail" class ApikeyguardrejectRail(BaseSecurityRail): ...必带 import(用于校验)
即使你的类直接继承自BaseSecurityRail,也必须在文件中保留下面这行 import:
from openjiuwen.harness.rails.base import DeepAgentRail # RailManager 校验所需示例代码中均以# noqa: F401标注该行(如 ApiKeyGuardAlert/rail.py#L17),说明它是"仅为校验而导入、虽未直接使用"的依赖,ruff 检查会跳过该行的未使用告警。
四、编写自定义安全 Rail:完整模板与关键方法
完整模板
examples/security_rail_demo/README.md 提供了可直接套用的模板:
from __future__ import annotations import re from typing import Set # 校验所需 import(即使未使用也必须保留) from openjiuwen.harness.rails.base import DeepAgentRail from openjiuwen.core.foundation.llm import ToolMessage from openjiuwen.core.single_agent.interrupt.response import InterruptRequest from openjiuwen.core.single_agent.rail.base import AgentCallbackEvent from openjiuwen.harness.rails.security.base_security_rail import ( BaseSecurityRail, SecurityAllow, SecurityReject, SecurityInterrupt, SecurityCheckContext, ) FILE_READING_TOOLS: Set[str] = {"read_file", "bash", "grep", "glob"} # 类名:首字母大写,其余小写,末尾加 "Rail" # 例如目录 "MyCustom" → 类 "MycustomRail" class MycustomRail(BaseSecurityRail): """自定义安全 rail。""" priority = 85 supported_events = {AgentCallbackEvent.AFTER_TOOL_CALL} async def run_security_check(self, security_ctx: SecurityCheckContext): ctx = security_ctx.callback_ctx inputs = ctx.inputs tool_result = getattr(inputs, "tool_result", None) if self._is_sensitive(tool_result): return self.reject(message="Sensitive data blocked") return self.allow() async def apply_security_decision(self, security_ctx, decision): if isinstance(decision, SecurityAllow): return if isinstance(decision, SecurityReject): ctx = security_ctx.callback_ctx inputs = ctx.inputs tool_call = getattr(inputs, "tool_call", None) inputs.tool_result = decision.message inputs.tool_msg = ToolMessage( content=decision.message, tool_call_id=tool_call.id if tool_call else "", ) return if isinstance(decision, SecurityInterrupt): ctx = security_ctx.callback_ctx user_input = security_ctx.user_input if user_input is None: self._raise_tool_interrupt( tool_name=ctx.inputs.tool_name, tool_call=ctx.inputs.tool_call, request=decision.request, ) approved = user_input.get("approved", False) if isinstance(user_input, dict) else False if not approved: ctx.inputs.tool_result = "Rejected by user" return await super().apply_security_decision(security_ctx, decision) def _is_sensitive(self, result) -> bool: return False __all__ = ["MycustomRail"]两个必须实现/理解的方法
自定义 Rail 的核心逻辑分布在两个方法中:
run_security_check(security_ctx):安全检查入口,返回一个SecurityDecision。检查上下文SecurityCheckContext包含callback_ctx(Agent 回调上下文)、event(当前事件)、user_input(用户审批输入)、auto_confirm_config(会话级自动确认配置)与subject_id(审批主体标识)。apply_security_decision(security_ctx, decision):根据决策执行落地动作。基类已提供默认实现(base_security_rail.py#L302-L340),模板中的写法是在默认行为上做定制,因此最后对未覆盖的决策类型调用super()。
基类默认落地行为(源码级)
BaseSecurityRail._run_and_apply(base_security_rail.py#L225-L252)负责统一调度:解析subject_id→ 从会话提取用户输入 → 构造SecurityCheckContext→ 调用run_security_check→ 将决策写入ctx.extra["_interrupt_decision"]→ 调用apply_security_decision。默认落地策略如下:
| 决策 | 事件 | 默认行为 |
|---|---|---|
SecurityAllow | 任意 | 直接返回,继续执行 |
SecurityAlert | 任意 | 按级别写日志 + 通过OutputSchema(type="message")流式推送前端,继续执行 |
SecurityReject | BEFORE_MODEL_CALL/AFTER_MODEL_CALL | request_force_finish强制结束 Agent |
SecurityReject | BEFORE_TOOL_CALL | _skip_tool跳过工具执行,Agent 继续(可换其他方案) |
SecurityReject | AFTER_TOOL_CALL | 改写tool_result/tool_msg为错误消息 +request_force_finish结束 Agent |
SecurityInterrupt | 模型类事件 | 自动转换为 Reject(模型事件不支持 HITL,见 base_security_rail.py#L241-L249) |
SecurityInterrupt | 工具类事件 | 无用户输入时抛出ToolInterruptException(经AbortError包装)等待审批 |
其中 REJECT 时_build_force_finish_result(base_security_rail.py#L468-L475)会构造{"output": ..., "result_type": "error"}作为 Agent 终止时的返回结果。
支持的事件与触发时机
| 事件 | 触发时机 | 是否支持 Interrupt |
|---|---|---|
BEFORE_INVOKE | Agent invoke 开始前 | 是 |
BEFORE_MODEL_CALL | LLM 调用前 | 否(自动转为 Reject) |
AFTER_MODEL_CALL | LLM 响应后 | 否(自动转为 Reject) |
BEFORE_TOOL_CALL | 工具执行前 | 是 |
AFTER_TOOL_CALL | 工具执行后 | 是 |
ON_MODEL_EXCEPTION | LLM 调用失败时 | 否(自动转为 Reject) |
ON_TOOL_EXCEPTION | 工具执行失败时 | 是 |
所有事件定义在 openjiuwen/core/single_agent/rail/base.py#L473-L516,事件到钩子方法名的映射由EVENT_METHOD_MAP(base.py#L812-L826)提供;BaseSecurityRail.get_callbacks会根据你声明的supported_events自动收集对应钩子并注册。
优先级指南
| 优先级 | 典型用途 |
|---|---|
| 100+ | 最高优先级,最终检查 |
| 85-95 | 安全 Rail |
| 50-70 | 处理型 Rail |
| 10-30 | 日志/遥测 Rail |
优先级语义在 AgentRail 基类文档中有明确说明:数值越大越先执行,同时决定init的先后顺序(因为init负责注册工具与提示词段落,先初始化后初始化的顺序与钩子执行顺序同源)。集成测试 tests/system_tests/rail/test_base_security_rail_integration.py 中的test_priority_ordering_high_runs_before_low用例验证了 priority 90 的 Rail 一定先于 priority 10 的 Rail 执行。
五、五个示例 Rail 的源码级剖析
5.1 ApiKeyGuardAlert —— ALERT 模式:告警但不阻断
ApiKeyGuardAlert/rail.py 演示了SecurityAlert决策。它监听AFTER_TOOL_CALL,当文件读取类工具(read_file、bash、grep、glob、read)的结果命中 API Key 正则时,返回self.alert(...),执行不被阻断。
if self._contains_api_key(content): return self.alert( message=f"API key/secret detected in {tool_name} result. Execution allowed but flagged.", level=self._alert_level, alert_type="api_key_leakage", display_mode=self._display_mode, )其中alert()帮助方法(base_security_rail.py#L148-L171)返回SecurityAlert(message, level, alert_type, display_mode)。基类_apply_alert(base_security_rail.py#L419-L466)会:按级别调用logger记录告警;通过ctx.session.write_stream(OutputSchema(type="message", ...))向前端推送系统消息,并在payload.metadata中携带is_security_alert=True、level、alert_type、display_mode、rail等字段,供前端做特殊样式(红框、警告图标等)。
display_mode是给前端的展示提示,三选一:
| 模式 | 前端行为 | metadata |
|---|---|---|
popup | Toast/弹窗通知 | display_mode: "popup" |
history | 插入聊天历史 | display_mode: "history" |
inline | 实时流式输出 | display_mode: "inline" |
构造参数可按需定制:
# 默认: popup 通知, WARNING 级别 rail = ApikeyguardalertRail() # 自定义展示模式 rail_inline = ApikeyguardalertRail(display_mode="inline") # 自定义告警级别 rail_critical = ApikeyguardalertRail( display_mode="popup", alert_level=SecurityAlertLevel.CRITICAL, )5.2 ApiKeyGuardInterrupt —— INTERRUPT 模式:人工审批(HITL)
ApiKeyGuardInterrupt/rail.py 同时监听BEFORE_TOOL_CALL与AFTER_TOOL_CALL,分别检查工具参数与工具结果:
- BEFORE_TOOL_CALL:参数中含密钥 →
InterruptRequest暂停,用户批准则执行,拒绝则跳过工具(Agent 继续,可换其他方案); - AFTER_TOOL_CALL:结果中含密钥 → 暂停,用户批准则继续,拒绝则强制结束 Agent(因为数据已泄露)。
InterruptRequest(openjiuwen/core/single_agent/interrupt/response.py)携带审批消息、payload_schema(审批表单结构)、auto_confirm_key与ui_options(前端按钮:Approve / Always Allow / Reject):
return self.interrupt( InterruptRequest( message="API key/secret detected in tool arguments. Approve execution?", payload_schema={ "type": "object", "properties": { "approved": {"type": "boolean"}, "feedback": {"type": "string"}, "auto_confirm": {"type": "boolean", "description": "Remember approval for this tool"}, }, "required": ["approved"], }, auto_confirm_key=auto_confirm_key, ui_options=[ {"label": "Approve", "description": "Execute this tool call", "value": "approve"}, {"label": "Always Allow", "description": "Remember approval for this tool", "value": "always_allow"}, {"label": "Reject", "description": "Skip tool execution", "value": "reject"}, ], ), subject_id=tool_call_id, )该 Rail 还演示了**自动确认(auto-confirm)**机制:auto_confirm_key格式为api_key_guard:{tool_name}:{event}(如api_key_guard:read_file:before),前后事件使用独立 key。_handle_interrupt_resume(base_security_rail.py#L691-L735)封装了标准续跑流程:先查会话级 auto-confirm 配置(INTERRUPT_AUTO_CONFIRM_KEY),命中则直接放行;否则解析用户输入{"approved": bool, "auto_confirm": bool};用户选择"Always Allow"时通过_store_auto_confirm把 key 写入会话状态,后续同类操作自动放行。流程示例(BEFORE 中断):
第 1 轮: User: "读取路径中包含 secret 的文件" LLM: ToolCall(args="path=sk-secret123...") Rail: 中断 "参数中检测到 secret,是否批准执行?" User: 拒绝 Rail: 跳过工具(Agent 继续) LLM: 尝试其他方案(如向用户索要安全路径)5.3 ModelCallGuard —— REJECT 模式:阻断 LLM 输入/输出中的密钥
ModelCallGuard/rail.py 监听BEFORE_MODEL_CALL与AFTER_MODEL_CALL:
- BEFORE_MODEL_CALL:扫描发给 LLM 的历史消息,发现密钥后先
_pop_matching_messages弹出含密钥的消息再 Reject(仅清理当轮新增的用户消息,历史"特赦"); - AFTER_MODEL_CALL:检查 LLM 响应内容与
tool_call.arguments,发现密钥则从整个历史中清除含密钥消息(彻底清理)并 Reject。
被拒后对话可以继续——用户看到错误消息后可重试。基类为其提供的辅助方法包括:_pop_matching_messages(按正则弹出历史中含密钥的消息,base_security_rail.py#L567-L602)、_pop_last_user_message(弹出当轮最后一条用户消息,base_security_rail.py#L545-L565)等。由于模型事件不支持 Interrupt,这里的拒绝直接走request_force_finish。
5.4 SensitiveDataSanitize —— SANITIZE 模式:脱敏但不阻断
SensitiveDataSanitize/rail.py 与 ModelCallGuard 形成互补:不弹消息、不拒绝,而是把对话历史与 LLM 响应中的密钥原地替换为[REDACTED]后放行(priority = 85)。核心依赖基类的_sanitize_matching_messages(base_security_rail.py#L648-L689),并支持自定义替换串:
rail = SensitivedatasanitizeRail(replacement="[REDACTED]") rail = SensitivedatasanitizeRail(replacement="<SECRET>")它可与拒绝类 Rail 组合使用(先脱敏、后检查)。适合日志审计、调试模式、只遮蔽不阻断的管道处理等场景。
5.5 ToolRejectExample —— REJECT 模式:BEFORE/AFTER 行为差异演示
ToolRejectExample/rail.py 不含 Interrupt,命中即直接 Reject,用于讲透工具事件拒绝行为的差异:
| 事件 | 拒绝后的动作 | 原因 |
|---|---|---|
BEFORE_TOOL_CALL | _skip_tool,Agent 继续 | 密钥还在参数里,工具尚未执行,Agent 可换其他方案 |
AFTER_TOOL_CALL | request_force_finish,Agent 终止 | 密钥已在结果中,数据已泄露,必须终止 |
基类_apply_reject正是按上述逻辑分流的(base_security_rail.py#L342-L390)。
5.6 共用的敏感信息检测模式
各 Rail 均使用类似的正则集合,可直接复用并按需扩展:
| 模式 | 匹配目标 |
|---|---|
sk-[a-zA-Z0-9_-]{20,} | OpenAI 风格密钥(20+ 字符,支持下划线/连字符) |
(?:api_key\|API_KEY\|...)\s*[=:]\s*["']?\S+["']? | 通用api_key=xxx/SECRET=xxx/token=xxx |
Bearer\s+[a-zA-Z0-9\-_]+ | Bearer Token |
AKIA[0-9A-Z]{16} | AWS Access Key |
值得注意的教训(见 ApiKeyGuardInterrupt/README.md):曾用过的r"\.env\b"模式因**过度宽泛(会匹配文件名)**被移除,编写正则时应优先聚焦密钥自身的结构特征而非文件名。
六、测试你的 Rail
仓库内置了两级测试来验证安全 Rail 框架本身:
# 在 agent-core 根目录下执行 PYTHONPATH=. uv run pytest \ tests/unit_tests/harness/rails/test_base_security_rail.py \ tests/system_tests/rail/test_base_security_rail_integration.py \ -v- tests/unit_tests/harness/rails/test_base_security_rail.py 验证:
supported_events只注册声明的事件;模型事件 Reject 会触发request_force_finish且结果包含result_type: "error";模型事件上的 Interrupt 会被自动转为 Reject; - tests/system_tests/rail/test_base_security_rail_integration.py 通过
MockLLMModel+ 真实工具走完整agent.invoke()流程,验证:AFTER_TOOL_CALL Reject 会替换工具结果并让 Agent 输出含 "blocked" 的内容;Allow 则原样放行;多事件 Rail 能收到全部声明事件;高优先级先执行;HITL 场景下resume_input注入{"approved": true/false}可分别走通放行与强制结束。
针对 Interrupt 模式的自动化用例,还可参考:
uv run pytest tests/unit_tests/harness/rails/test_model_call_guard.py::TestToolInterruptReject -v uv run pytest tests/unit_tests/harness/rails/test_model_call_guard.py::TestApiKeyGuardInterruptAutoConfirm -v七、实战要点小结
- 命名是硬约束:目录用 CapitalCamelCase、文件固定叫
rail.py、类名按{首字母大写}{其余小写}Rail生成、配置键与目录名一致,并保留from openjiuwen.harness.rails.base import DeepAgentRail这一校验 import; - 事件按需声明:通过
supported_events声明关注的事件;工具类事件(BEFORE/AFTER_TOOL_CALL、ON_TOOL_EXCEPTION)支持 HITL,模型类事件不支持 Interrupt(会自动转为 Reject); - 拒绝行为与事件强相关:BEFORE_TOOL_CALL 拒绝是"跳过工具、Agent 继续",AFTER_TOOL_CALL 拒绝是"强制结束",写业务逻辑前先想清楚数据泄露发生在哪个阶段;
- 合理设置优先级:安全 Rail 建议 85-95 区间,
priority越大越先执行;ALERT/SANITIZE 类低侵入 Rail 可与 REJECT 类 Rail 按优先级编排组合(如 SensitiveDataSanitize/README.md 中先priority=85脱敏、再priority=90检查的组合示例); - 模式检测要精准:优先匹配密钥结构特征(
sk-、AKIA、Bearer、key=value),避免用\.env这类会误伤文件名的宽泛模式; - 四态决策覆盖不同需求:REJECT(阻断)、INTERRUPT(人工审批)、ALERT(放行但告警)、SANITIZE(脱敏后放行),按业务安全等级与体验取舍选择。
如需深入底层,可继续阅读核心实现 openjiuwen/harness/rails/security/base_security_rail.py、事件框架 openjiuwen/core/single_agent/rail/base.py、扩展加载机制 openjiuwen/harness/manifest/meta_elements.py 与 openjiuwen/harness/resources/extension_loader.py,以及集成测试 tests/system_tests/rail/test_base_security_rail_integration.py,即可在自己的项目中定制出贴合业务的安全防线。
- 人工智能
- AI Agent
- Agent 框架
- 大模型
- 工具调用
- RAG
- 提示工程
- 强化学习
【免费下载链接】agent-core
openJiuwen agent-core可提供AI Agent开发、运行、调优与演进相关的全套SDK能力
相关推荐
基于 openJiuwen 快速构建自定义 Agent:从 AgentCard 配置到 LLM 调用与敏感词安全校验实战
基于 openJiuwen 快速构建自定义 Agent:从 AgentCard 配置到 LLM 调用与敏感词安全校验实战 本篇文章以 openJiuwen(ag
人工智能AI AgentAgent 框架大模型工具调用RAG提示工程强化学习openJiuwen agent-core ModelCallGuard 安全 Rail 实战:拦截 LLM 输入输出中的 API Key 与密钥
openJiuwen agent core ModelCallGuard 安全 Rail 实战:拦截 LLM 输入输出中的 API Key 与密钥 本篇技术指南
人工智能AI AgentAgent 框架大模型工具调用RAG提示工程强化学习openJiuwen agent-core Guardrail 安全护栏框架:基于事件回调的 Agent 安全检测与拦截
openJiuwen agent core Guardrail 安全护栏框架:基于事件回调的 Agent 安全检测与拦截 导读 :本文全面讲解 openJiuw
人工智能AI AgentAgent 框架大模型工具调用RAG提示工程强化学习
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考