1. 从一次训练崩溃说起:Reward Judging 到底在判什么
如果你正在读 OpenClaw-RL 的源码,大概率已经翻过了环境构建、动作采样、状态观测这几块。到了 reward_judging 这一层,很多人会卡住:代码能跑,但训练曲线像心电图,一会儿冲上去一会儿掉下来,甚至出现四足机器人原地趴着不动、靠“零惩罚”混日子的情况。这不是模型不行,而是奖励判定逻辑没吃透。
Reward Judging 在 Agentic RL 里扮演的角色,可以理解成训练循环里的“裁判”。它不产生动作,也不更新网络,只做一件事:拿到当前状态、动作和环境反馈后,输出一个标量奖励,告诉策略“这一步走得好还是差”。在 OpenClaw-RL 里,这个裁判由 RewardJudging 类实现,内部把前进、生存、姿态、能耗等多个子项加权求和。权重稍微偏一点,智能体的行为就会从“稳步前进”变成“原地抽搐”或者“躺平摆烂”。
这篇笔记聚焦三件事:拆开 reward_judging 模块的判定逻辑,给出可复制的 settings.json / config.toml 骨架,以及用 TaoToken 统一 Key 接入后验证 Reward Judging 是否真正生效的检查动作。适合已经跑通 OpenClaw-RL 基础环境、想深入调奖励函数的读者。下面所有配置和命令都可以在本地复现,不需要额外硬件。
2. TaoToken 前置:统一 Key 接入与配置入口
在讲 Reward Judging 的配置之前,先把模型调用这一层理顺。OpenClaw-RL 在训练过程中会调用大模型做轨迹评估、指令解析或者 reward model 打分,如果每个模块各自配一套 Key,调试时会非常乱。TaoToken 的作用就是把这些调用收敛到一个入口,用统一 Key 管理。
你可以先到官网 https://taotoken.net/?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= 了解整体能力,然后进控制台创建 Key。API 地址是 https://taotoken.net/api,注意这个地址不带 UTM 参数,直接填到配置文件里即可。控制台入口在 https://taotoken.net/console?utm_source=taotoken_aicg_blog_end&utm_content=console&utm_campaign=rewrite ,API Keys 管理页在 https://taotoken.net/api-keys?utm_source=taotoken_aicg_blog_end&utm_content=api-keys&utm_campaign=rewrite 。
拿到 Key 之后,不要急着写进代码。先在模型对话页面 https://taotoken.net/chat?utm_source=taotoken_aicg_blog_end&utm_content=model_chat&utm_campaign=rewrite 发一条测试消息,确认 Key 可用、额度正常。这一步能排掉大部分“配置写了但请求 401”的问题。如果你后续要做长期编码或者 Agent 任务,可以看 Coding Plan 页面 https://taotoken.net/coding-plan?utm_source=taotoken_aicg_blog_end&utm_content=coding_plan&utm_campaign=rewrite ,接入文档在 https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_content=doc&utm_campaign=rewrite 。
注意:TaoToken 是模型调用入口,不是编辑器替代品,也不要把生产数据库直连进去。配置时只放 Key 和 API 地址,不要塞敏感凭据。
3. 可复制配置:settings.json 与 config.toml 骨架
OpenClaw-RL 的配置分两层:一层是训练框架的 settings.json,管全局路径和日志;另一层是 reward_judging 的 config.toml,管奖励子项权重和判定阈值。下面给出可直接复制的骨架,字段名和源码里的 config.get 调用一一对应。
3.1 settings.json 骨架
{ "project_name": "openclaw_rl_reward_judging", "seed": 42, "device": "cuda:0", "log_dir": "./logs/reward_judging", "checkpoint_dir": "./checkpoints", "taotoken": { "api_base": "https://taotoken.net/api", "api_key_env": "TAOTOKEN_API_KEY", "timeout_sec": 30, "max_retries": 3 }, "reward_judging": { "config_path": "./configs/reward_judging.toml", "enable_debug_info": true, "log_sub_rewards": true } }这里把 Key 放在环境变量TAOTOKEN_API_KEY里,而不是硬编码进 JSON。这样提交代码时不会泄露,也方便在不同机器上切换。enable_debug_info打开后,compute_reward 会把各子项写进 info 字典,后面验证时直接读。
3.2 config.toml 骨架
[reward] forward_weight = 1.0 survival_weight = 1.0 orientation_weight = -0.5 energy_weight = -0.1 [reward.params] target_velocity = 0.5 max_tilt_angle = 0.4 forward_scale = 2.0 orientation_scale = 10.0 survival_bonus = 0.1 [reward.termination] fail_on_tilt = true fail_on_collision = true max_episode_steps = 1000对照源码,forward_weight对应self.w_forward,target_velocity对应self.target_velocity,max_tilt_angle对应self.max_tilt_angle。forward_scale是前进奖励指数衰减的系数,源码里写死成 2.0,这里提出来方便调。orientation_scale是姿态惩罚的二次项系数,默认 10.0,倾斜越界越多惩罚越重。
3.3 加载配置的代码片段
import json import os import toml def load_config(settings_path="./settings.json"): with open(settings_path, "r", encoding="utf-8") as f: settings = json.load(f) api_key = os.environ.get(settings["taotoken"]["api_key_env"]) if not api_key: raise RuntimeError("TAOTOKEN_API_KEY 未设置,请先导出环境变量") reward_cfg_path = settings["reward_judging"]["config_path"] with open(reward_cfg_path, "r", encoding="utf-8") as f: reward_cfg = toml.load(f) return settings, reward_cfg, api_key if __name__ == "__main__": settings, reward_cfg, api_key = load_config() print("API Base:", settings["taotoken"]["api_base"]) print("Forward Weight:", reward_cfg["reward"]["forward_weight"]) print("Target Velocity:", reward_cfg["reward"]["params"]["target_velocity"])运行前先导出 Key:
export TAOTOKEN_API_KEY="你的Key" python load_config.py输出里能看到 API Base 和权重,说明配置链路通了。这一步不涉及训练,纯配置校验,出错概率低。
4. 验证 Reward Judging 生效:三个检查动作
配置写完不代表奖励判定就对了。下面三个动作可以确认 RewardJudging 真的在按预期工作。
4.1 检查子项是否写入 info
在 compute_reward 返回前,源码会把reward_forward、reward_survival、reward_orientation、reward_energy写进 info。你可以在训练循环里加一行打印:
total_reward = judge.compute_reward(state, action, info) if step % 50 == 0: print(f"step={step} total={total_reward:.4f} " f"forward={info['reward_forward']:.4f} " f"survival={info['reward_survival']:.4f} " f"orientation={info['reward_orientation']:.4f} " f"energy={info['reward_energy']:.4f}")如果打印出来全是 0,说明 config.toml 没被正确加载,或者 compute_reward 没被调用。先查config_path路径,再查 step 函数里有没有真的调 judge。
4.2 用固定状态做单元测试
构造三个固定场景,看输出是否符合预期。理想状态速度接近目标、姿态正、动作平滑,总奖励应该最高;姿态倾斜过大时 orientation 子项应为负;终止状态下 survival 应为 0。
import numpy as np from reward_judging import RewardJudging config = { "reward_forward_weight": 1.0, "reward_survival_weight": 1.0, "reward_orientation_weight": -0.5, "reward_energy_weight": -0.1, "target_velocity": 0.5, "max_tilt_angle": 0.4, } judge = RewardJudging(config) state = {"base_linear_vel_x": 0.48, "base_roll": 0.1, "base_pitch": 0.05} info = {"terminated": False, "truncated": False} action = np.ones(12) * 0.5 reward = judge.compute_reward(state, action, info) print(f"理想状态 total={reward:.4f} orientation={info['reward_orientation']:.4f}") state2 = {"base_linear_vel_x": 0.3, "base_roll": 0.5, "base_pitch": 0.6} info2 = {"terminated": False, "truncated": False} reward2 = judge.compute_reward(state2, action, info2) print(f"倾斜状态 total={reward2:.4f} orientation={info2['reward_orientation']:.4f}") state3 = {"base_linear_vel_x": 0.0, "base_roll": 1.0, "base_pitch": 0.0} info3 = {"terminated": True, "truncated": False} reward3 = judge.compute_reward(state3, action, info3) print(f"终止状态 total={reward3:.4f} survival={info3['reward_survival']:.4f}")预期结果:理想状态 total 在 1.5 左右,倾斜状态 orientation 为负且 total 明显下降,终止状态 survival 为 0。如果倾斜状态的 orientation 还是 0,检查max_tilt_angle是不是设太大,或者 tilt_magnitude 计算用的 roll/pitch 单位是不是弧度。
4.3 通过 TaoToken 拉取评估结果做交叉验证
如果训练里接了模型评估,可以用 TaoToken 的模型对话接口发一条评估请求,确认外部调用链路正常。API 地址用 https://taotoken.net/api ,请求头带Authorization: Bearer $TAOTOKEN_API_KEY。这一步不是必须,但能帮你区分“奖励函数写错”和“模型调用失败”两类问题。
curl -s https://taotoken.net/api/v1/chat/completions \ -H "Authorization: Bearer $TAOTOKEN_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-4o-mini", "messages": [{"role": "user", "content": "返回一个 JSON: {\"ok\": true}"}], "temperature": 0 }' | head -c 500返回里有choices字段说明 Key 和网络都正常。如果这里报 401,先回 API Keys 页面确认 Key 状态,再检查环境变量有没有导出到当前 shell。
5. 本篇常见错排查
5.1 奖励曲线震荡,子项互相打架
前进奖励和能耗惩罚经常冲突。机器人为了拿前进奖励会剧烈摆动关节,能耗惩罚又把它压回去,结果策略在两种行为之间反复横跳。解决办法是先固定forward_weight=1.0、survival_weight=1.0,把energy_weight设成 0,只调orientation_weight,等姿态稳定后再逐步加能耗惩罚。每次只动一个权重,观察 200 个 episode 的曲线再决定下一步。
5.2 机器人学会“躺平”
如果orientation_weight绝对值太大,机器人发现不动就不会倾斜,于是选择趴着拿生存奖励。这时候看 info 里的reward_forward,如果长期接近 0,说明前进激励不够。把target_velocity从 0.5 降到 0.3,同时把survival_bonus从 0.1 降到 0.05,逼它动起来。
5.3 config.toml 字段名和源码对不上
源码里用的是config.get("reward_forward_weight", 1.0),而 config.toml 里写的是forward_weight,两者不一致时源码会走默认值,你改配置根本不生效。排查方法是在 RewardJudging 初始化后打印所有权重:
print(f"w_forward={judge.w_forward} w_survival={judge.w_survival} " f"w_orientation={judge.w_orientation} w_energy={judge.w_energy}")如果打印出来全是默认值,说明字段名映射错了。要么改 config.toml 的键名,要么在加载时做一层映射。
5.4 终止状态生存奖励没归零
_compute_survival_reward里判断的是info.get("terminated", False),如果你的环境返回的键是done而不是terminated,生存奖励就不会归零,机器人摔倒后还能继续拿奖励。检查环境 step 返回的 info 字典键名,统一成terminated和truncated。
5.5 TaoToken 请求超时导致评估卡住
训练循环里如果同步调用模型评估,网络抖动会让整个训练卡住。settings.json 里设了timeout_sec=30和max_retries=3,但代码里要真的用上。建议把模型评估放到独立线程或队列里,训练主循环只读缓存结果,避免 Reward Judging 被网络阻塞。
6. 继续深入:从 Reward Judging 到 Coding Plan
Reward Judging 调通之后,下一步通常是把它接到更长的 Agent 任务里,比如让模型根据奖励信号自动调整策略参数。这类长期编码和 Agent 场景,用 TaoToken 的 Coding Plan 会比较顺,入口在 https://taotoken.net/coding-plan?utm_source=taotoken_aicg_blog_end&utm_content=coding_plan&utm_campaign=rewrite 。接入文档里有完整的请求示例和参数说明,地址是 https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_content=doc&utm_campaign=rewrite 。
如果你在配置过程中遇到 Key 报错或者请求不通,先回 API Keys 页面 https://taotoken.net/api-keys?utm_source=taotoken_aicg_blog_end&utm_content=api-keys&utm_campaign=rewrite 确认 Key 状态,再对照接入文档检查请求头。模型对话页面 https://taotoken.net/chat?utm_source=taotoken_aicg_blog_end&utm_content=model_chat&utm_campaign=rewrite 可以用来快速验证模型是否正常响应,不用写代码。
奖励函数的设计没有标准答案,只有不断试出来的权重组合。我自己的习惯是每次只改一个参数,跑 200 个 episode,把子项曲线和总奖励画在一张图上,看哪个子项在拖后腿。这套流程跑顺之后,OpenClaw-RL 的训练效率会有明显提升。