原文:New in llama.cpp: Decision Models
发布日期:2026 年 10 月 2 日
作者:Xuan-Son Nguyen (ngxson)、Victor Mustar、ggml-org
llama.cpp 服务器现在通过/v1/systemone端点支持决策模型。你发送一个状态(文本、JSON、截图)和类型化问题,模型在单次前向传播中返回每个选项的概率。
API 遵循 TypeSafe 的 Jev 模型引入的 System One 格式,因此现有客户端只需更换 base URL 即可。实现细节见 PR #29818。
什么是决策模型?决策模型通过为你给定的选项打分来回答问题,而不是生成文本。聊天模型每个输出 token 需要一次前向传播,且输出仍需解析。决策模型只读取一次输入,其答案始终是你给出的选项之一,并附带概率。典型用途包括:路由请求、内容审核、检查 Agent 步骤是否成功、或选择下一步动作。
支持的模型
| 模型 | 大小 | 基于 | 语言 | 图片 | 许可证 | 速度* |
|---|---|---|---|---|---|---|
| Julia-1 | 144M | mmBERT-small | 50+ | 否 | Apache 2.0 | 3 ms |
| Laya | 421M | ModernBERT-large | 英语 | 否 | Apache 2.0 | 5 ms |
| Kev-4B | 4B | Qwen3.5-4B-Base | 英语 | 否 | Apache 2.0 | 12 ms |
| lev | 4B | Qwen3.5-4B | 英语 | 否 | Apache 2.0 | 36 ms |
| OpenJev | 27B | Qwen3.8-27B | en, de, fr, hi, zh, ja | 是 | CC BY-NC 4.0 | 43 ms |
| Clef | 27B | Qwen3.8-27B | 英语 | 是 | Apache 2.0 | — |
* 在单块 NVIDIA RTX PRO 6000 上回答一个问题的中位时间。
在 Decision models 合集 中查找这些模型,更多模型即将推出。社区 Decision Index 展示了它们的对比。
快速开始
从 llama.app 获取最新版 llama.cpp(或运行llama update),然后启动模型:
llama serve-hfggml-org/Kev-4B-GGUF一个请求包含一个状态和一个或多个问题。问题有三种类型:
| 类型 | 你发送 | 你得到 |
|---|---|---|
choice | 选项(含可选描述) | 最佳选项 + 每个选项的概率 |
score | 2 到 10 个等级(从低到高) | 期望等级(可以落在两个等级之间) |
noul | 是/否问题 | 是(yes)的概率 |
发送包含状态和问题的请求:
curlhttp://localhost:8080/v1/systemone\-H"Content-Type: application/json"\-d'{ "state": "Customer message: I was charged twice for my order last week and nobody has replied.", "questions": { "route": { "type": "choice", "instructions": "Which team should handle this?", "criteria": { "billing": "payments, charges, refunds, invoices", "shipping": "delivery, tracking, lost or late parcels", "technical": "bugs, errors, login problems" } }, "angry": { "type": "noul", "instructions": "Is the customer angry?" }, "urgency": { "type": "score", "instructions": "How urgent is this?", "criteria": ["can wait", "this week", "today", "right now"] } } }'响应(数值已四舍五入):
{"model":"ggml-org/Kev-4B-GGUF","answers":{"route":{"type":"choice","choice":"billing","probabilities":{"billing":0.9049,"shipping":0.0275,"technical":0.0676},"confidence":0.8574},"angry":{"type":"noul","noul":0.8208},"urgency":{"type":"score","score":2.2821,"legend":{"0":"can wait","1":"this week","2":"today","3":"right now"},"probabilities":{"0":0.036,"1":0.1937,"2":0.2225,"3":0.5478},"confidence":0.2821}},"usage":{"input_tokens":130,"output_tokens":0}}完整参考见 服务器文档。
图片
某些模型(目前为 OpenJev)还可以读取图片,如文档或截图。视觉投影器会自动下载:
llama serve-hfggml-org/OpenJev-GGUF例如,对上传的文档进行分类:
importbase64importrequestswithopen("document.png","rb")asf:image="data:image/png;base64,"+base64.b64encode(f.read()).decode()response=requests.post("http://localhost:8080/v1/systemone",json={"state":"A file uploaded by a customer.","images":[image],"questions":{"kind":{"type":"choice","instructions":"What kind of document is this?","criteria":{"invoice":None,"receipt":None,"contract":None,"other":None},},},})print(response.json()["answers"]["kind"]["choice"])# invoicestate也可以是聊天消息列表。任何image_url部分(data URL)都会被当作图片读取,与聊天补全相同。
多模型,单服务器
在路由模式下,模型按需加载,你可以在每个请求中选择一个:
llama servecurlhttp://localhost:8080/v1/systemone\-H"Content-Type: application/json"\-d'{"model": "ggml-org/Julia-1-GGUF:Q8_0", "state": "...", "questions": {...}}'/v1/models列出所有 id。当只加载了一个模型时,model字段会被忽略。
使用技巧
- 尝试不同大小的模型。小模型更快,大模型知识更丰富。Decision Index 对它们进行了比较。
- 描述你的选项。Julia-1 在仅有标签时将"我被扣了两次钱"路由到
shipping,而在每个选项都有描述时路由到billing(0.99)。 - 为每个模型选择置信度阈值。常见的模式是对有信心的答案采取行动,其余的发送给人工。正确的阈值取决于模型:一个模糊的工单(“Hi, quick question about my account”)在 Julia-1 上得分 0.25,而在 Kev-4B 上得分 0.80。在选择阈值之前,请在你自己的示例上进行测试。
- 批量处理问题。问题是独立回答的,Kev-4B、lev 和 OpenJev 只处理一次状态。
- 尝试不同的量化。与任何 GGUF 一样,这些模型有多种精度可选,例如
llama serve -hf ggml-org/Kev-4B-GGUF:Q8_0。