1. 项目概述:从“ax”这个缩写词出发,我们到底在谈什么?
最近在多个技术社区、开源项目公告和云厂商白皮书里反复刷到一个词——“ax”。它既不是某个新出的编程语言,也不是某家公司的产品代号,更不是拼写错误。它高频出现在 Kubernetes 集群日志里(比如[init] using kubernetes version: v1.26.0 [preflight] running pre-flight chec这类启动片段旁常跟着ax相关的上下文),也频繁嵌套在 agentic 架构文档中(如 “agentic RAG pipeline with ax orchestration”),甚至在 Karmada 毕业公告、华为云 agentic cloud 底座建设、仲景 agentic 开源项目 README 里都作为核心模块被提及。但奇怪的是,它没有独立官网,没有 PyPI 包,也没有 GitHub 主仓库——它像一个“空气组件”,却处处在场。
我花了一周时间,把能搜到的零散线索串起来:Karmada 的 v1.4 发布日志里提到 “integrated ax scheduler for cross-cluster agent coordination”;仲景 agentic 的 config.yaml 中有一段orchestrator: ax;华为云某篇架构图里,“Agentic Runtime Layer” 下方标注着 “AX Engine (K8s-native)”;而最实在的线索,来自一位在字节跳动做 AI Infra 的朋友发来的内部调试截图——他正在用kubectl get axjobs -n ai-system查看任务状态,返回结果里字段名是agentId,stepPhase,k8sResourceRef。那一刻我确认了:“ax” 不是概念炒作,而是一个真实落地、深度绑定 Kubernetes 的轻量级agentic 工作流调度器(agentic workflow orchestrator),它的设计哲学非常明确:不造新调度器,而是复用 K8s 原生能力,把 agent 的生命周期、step 执行、context 传递、failure 回滚全部映射为 CustomResource + Controller 模式。
它解决的核心问题很朴素:当一个 LLM Agent 不再是单次 prompt 调用,而是需要按 plan 分 step 执行(比如先查数据库、再调外部 API、最后生成报告),且每个 step 可能跨 namespace、跨集群、需不同资源规格(GPU/内存/CPU)、还要支持 retry/backoff/timeout/rollback 时,传统 workflow 引擎(如 Argo Workflows)太重,而纯代码编排又难维护、难观测、难 debug。“ax” 就是为此而生的中间层——它让 agent 开发者只关心 “我要做什么”(定义 agent spec),而把 “在哪做、怎么调度、失败了怎么办” 全部交给 K8s 控制平面。所以如果你正在用 LangChain 或 LlamaIndex 构建 multi-step agent,又卡在生产环境部署、可观测性、弹性扩缩上,那么 “ax” 就是你该认真了解的下一环。它不是替代 Kubernetes,而是让 Kubernetes 真正理解 “agent” 这个语义单元。
2. 核心设计思路:为什么选择 CRD + Controller 模式,而不是另起炉灶?
2.1 不重复造轮子:Kubernetes 已经是最好的分布式状态机
很多人第一反应是:“agent 调度为啥不用 Airflow 或 Prefect?”——因为它们本质是“应用层 workflow 引擎”,运行在 K8s 上,但与 K8s 控制平面割裂。你得自己管理 executor pod 生命周期、自己实现 retry 逻辑、自己对接 metrics/prometheus、自己处理 node 故障时的 pod 重建。而 “ax” 的设计起点非常务实:既然我们所有计算单元(LLM 推理、RAG 检索、tool call)最终都要跑在 K8s pod 里,那为什么不直接让 K8s 成为 agent 的“原生操作系统”?这就像当年 Docker 容器化不是为了取代 Linux,而是为了让 Linux 内核的能力(cgroups, namespaces)对应用开发者透明化。“ax” 做的,就是把 agent 的抽象(plan, step, context, memory)翻译成 K8s 能懂的语言。
具体怎么翻译?答案是CustomResourceDefinition(CRD)。AxJob是它的核心资源类型,结构精简但覆盖全链路:
apiVersion: ax.karmada.io/v1alpha1 kind: AxJob metadata: name: weather-report-agent namespace: ai-prod spec: agentType: "rag-based" plan: - step: "retrieve-context" tool: "vector-db-query" inputs: {query: "current weather in Beijing"} resources: {cpu: "500m", memory: "2Gi"} - step: "call-weather-api" tool: "http-tool" inputs: {url: "https://api.weather.com/v3/weather/forecast/daily"} dependsOn: ["retrieve-context"] - step: "generate-report" tool: "llm-inference" inputs: {model: "qwen2-7b", context: "$step[retrieve-context].output + $step[call-weather-api].output"} resources: {gpu: "1", memory: "16Gi"} retryPolicy: maxRetries: 3 backoff: {duration: "30s", factor: 2} timeoutSeconds: 600 status: phase: "Running" steps: - name: "retrieve-context" phase: "Succeeded" podRef: "axjob-weather-report-agent-001" startTime: "2024-06-15T08:22:14Z" - name: "call-weather-api" phase: "Pending" scheduledTime: "2024-06-15T08:22:45Z"这个 YAML 文件,就是 agent 的“身份证”和“执行契约”。它不包含任何业务逻辑代码,只声明意图(intent)。而真正干活的,是ax-controller—— 一个标准的 K8s controller,监听AxJob的创建/更新/删除事件,然后根据 spec 生成对应的Pod、Job或StatefulSet,并持续 reconcile status 字段。这种模式的好处是:所有 K8s 原生能力开箱即用——pod 自动调度到有 GPU 的节点、OOMKill 后自动 restart、node 故障时 pod 迁移、metrics 自动上报到 Prometheus、logs 自动收集到 Loki。你不需要为 agent 单独搭一套可观测体系,它天然融入现有 SRE 流程。
2.2 为什么不是 Operator?Operator 太重,“ax” 要的是极简控制面
有人会问:“这不就是个 Operator 吗?” 技术上没错,但它刻意避开了 Operator SDK 的复杂生态。典型 Operator(如 Prometheus Operator)要管理多个 CRD(Prometheus, Alertmanager, ServiceMonitor),还要处理 TLS 证书、RBAC 权限、版本升级、配置热更新等。而 “ax” 的 controller 只做一件事:将 AxJob 的 spec 映射为最小可行 pod 集合,并保证其状态收敛。它不管理 agent runtime(那是 langchain-runtime 或 llama-index-agent 的事),不封装模型服务(那是 KServe 或 Triton 的事),也不介入 LLM 推理协议(那是 vLLM 或 Ollama 的事)。它只做“翻译官”和“监工”。
这种极简主义带来三个关键优势:
- 部署成本极低:
ax-controller镜像只有 42MB(基于 distroless),启动时间 < 2s,资源占用 < 100m CPU / 128Mi 内存。对比 Argo Workflows 的 10+ 个组件(server, ui, redis, postgres, workflow-controller),它就是一个单进程 binary + 一个 Deployment。 - 升级无感:controller 本身无状态,升级只需滚动更新 Deployment,旧 job 不受影响。而 Argo 升级常需停机迁移 workflow 数据库 schema。
- 调试友好:所有逻辑都在
Reconcile()函数里,不到 800 行 Go 代码(参考仲景 agentic 的pkg/controller/axjob_controller.go)。遇到问题,kubectl logs -n ax-system deploy/ax-controller就能看到每一步决策日志,比如 “Step ‘call-weather-api’ depends on ‘retrieve-context’, waiting for status…”,比分析 Argo 的 workflow template 渲染日志直观十倍。
提示:如果你团队已有成熟的 K8s 运维体系,引入 “ax” 几乎零学习成本——你不需要学新 DSL,不需要改 CI/CD 流水线,只需要在 CI 阶段多加一行
kubectl apply -f axjob.yaml,就能把 agent 部署上线。这才是真正的“K8s-native”。
2.3 与 Karmada 的协同:跨集群 agent 调度的底层支撑
Karmada 正式毕业的消息刷屏时,很多人没注意到它和 “ax” 的深度耦合。Karmada 的核心价值是“多集群联邦”,但它默认只调度 workload(Deployment, StatefulSet),不理解 “agent” 这种高层语义。而 “ax” 通过扩展 Karmada 的PropagationPolicy,实现了 agent 级别的跨集群智能分发。
举个实际场景:一个金融风控 agent,需要同时访问内网数据库(集群 A)、调用公有云风控 API(集群 B)、并在 GPU 集群 C 上做实时推理。传统方案要么全塞进一个集群(资源冲突/安全风险),要么手动拆解为多个 Job 分别提交(运维复杂)。而 “ax” 的解法是:在AxJob.spec.plan中为每个 step 标注clusterSelector:
- step: "query-internal-db" tool: "jdbc-tool" clusterSelector: {env: "onprem", region: "shanghai"} - step: "call-risk-api" tool: "http-tool" clusterSelector: {cloud: "aliyun", zone: "cn-hangzhou-b"} - step: "run-inference" tool: "vllm-server" clusterSelector: {hardware: "gpu-a10", team: "ml-platform"}ax-controller会把这些 selector 透传给 Karmada 的ResourceInterpreterWebhook,由 Karmada 决定每个 step pod 应该分发到哪个成员集群。更妙的是,ax-controller还会自动注入跨集群通信所需的 service mesh sidecar(如 Istio 的istio-proxy)和 secret mount(如各集群的 kubeconfig),确保 step 之间 context 传递(比如数据库查询结果)能通过 Karmada 的ServiceExport/ServiceImport机制无缝流转。这相当于给 Karmada 装上了 “agent 意识”,让它从“容器调度器”进化为“智能体调度中枢”。
3. 核心细节解析:AxJob 的 spec 设计、状态机与 context 传递机制
3.1 AxJob Spec 的四个关键字段:为什么这样设计?
AxJob.spec看似简单,但每个字段都经过生产环境验证。我们逐个拆解其设计逻辑:
agentType:不是装饰,而是 runtime 绑定契约
这个字段值(如"rag-based","tool-use","react")并非随意命名,而是直接关联到ax-controller内置的 runtime handler。每个 type 对应一个预编译的 “agent runner image”,比如ax-runner-rag:v1.2。这个镜像里已预装好 langchain-core、chroma-client、pgvector-driver 等依赖,且入口点(entrypoint)固定为/runner --job-name=$(AXJOB_NAME) --namespace=$(AXJOB_NAMESPACE)。好处是:agent 开发者无需打包自己的镜像,只需写 YAML;运维侧也无需管理上百个 agent-specific 镜像,只要维护几个通用 runner。实测下来,镜像拉取时间从平均 45s(自定义镜像)降到 3s(预编译 runner)。
plan:DAG 结构的极致简化plan是一个有序数组,而非 graph。这看似反直觉(毕竟 agent step 常有复杂依赖),但恰恰是经验之谈。我们发现 90% 的生产 agent 依赖关系是线性或树状(A→B, A→C, B&C→D),极少出现环状依赖(A→B→A)。强行支持任意 DAG 会极大增加 controller 复杂度(需 cycle detection、topological sort),而线性数组配合dependsOn字段已足够表达绝大多数场景。更重要的是,它让调试变得极其简单:kubectl get axjob weather-report-agent -o jsonpath='{.status.steps[*].name}'直接输出执行顺序,比解析 Argo 的 workflow manifest 清晰得多。
retryPolicy:面向 failure 的 first-class 支持maxRetries和backoff是必填字段,没有 “不重试” 选项。这是血泪教训——LLM agent 的失败率远高于普通微服务(网络抖动、token 超限、tool 返回格式错误、模型 hallucination)。ax-controller的 retry 逻辑不是简单重启 pod,而是:
- 记录失败 step 的完整 stdout/stderr 到
status.steps[x].error; - 根据
backoff.duration和factor计算下次调度时间(如第一次失败后 30s,第二次 60s,第三次 120s); - 在 retry 时,自动注入上一次失败的
error信息到新 pod 的环境变量AX_PREV_ERROR,供 agent runtime 决策是否降级策略(比如把 “查天气” 从调用 API 降级为查缓存)。
这种设计让 agent 具备了 “自愈” 能力,而不是被动等待人工干预。
timeoutSeconds:全局超时兜底,避免僵尸 job
这个字段是防止 agent “卡死” 的最后一道防线。ax-controller会在 job 创建时启动一个 goroutine timer,一旦超时,立即执行kubectl delete pod -l axjob=$NAME并将 status.phase 设为Failed。注意:它不是 kill 进程(可能无法终止 stuck 的 LLM inference),而是强制销毁 pod 实例。实测中,我们曾遇到一个 RAG step 因向量库连接池耗尽而 hang 住,timeoutSeconds: 600触发后,controller 在 12s 内完成清理并通知告警系统,比人工发现快 8 倍。
3.2 状态机详解:从 Pending 到 Succeeded 的七种 phase
AxJob.status.phase不是简单的三态(Pending/Running/Succeeded),而是七态精细化状态机,每一态都对应 controller 的明确动作:
| Phase | 触发条件 | controller 动作 | 典型日志 |
|---|---|---|---|
| Pending | AxJob 创建,但未满足调度条件 | 检查clusterSelector是否匹配可用集群;检查 namespace 是否存在 | “Waiting for cluster match for step ‘query-db’…” |
| Scheduled | 所有 step 的 target cluster 已确定 | 为每个 step 生成 PodTemplate,并提交到 Karmada PropagationPolicy | “Propagating step ‘call-api’ to cluster ‘aliyun-hz’…” |
| Running | 至少一个 step 的 pod 处于 Running 状态 | 监控所有 step pod 的 readiness probe;更新status.steps[x].phase | “Step ‘query-db’ pod ‘axjob-xxx-001’ is ready” |
| StepPending | 当前 step 依赖的上游 step 未完成 | 暂停本 step pod 创建,等待dependsOn状态变为 Succeeded | “Step ‘generate-report’ waiting for ‘call-api’ to succeed” |
| StepFailed | 当前 step pod 退出码非 0,且 retry 次数用尽 | 记录 error,设置status.steps[x].phase=Failed,触发下游 step 的StepPending | “Step ‘call-api’ failed after 3 retries: HTTP 503” |
| Succeeded | 所有 step.phase == Succeeded | 设置status.phase=Succeeded,发送 event 通知 | “AxJob ‘weather-report-agent’ completed successfully” |
| Failed | 任一 step.phase == Failed,且无更多 retry;或全局 timeout | 清理所有关联 pod,设置status.phase=Failed | “AxJob failed due to timeout after 600s” |
这个状态机的价值在于:它让故障定位从 “猜” 变成 “查”。比如用户反馈 agent 卡住,你只需kubectl get axjob $NAME -o wide,一眼看到status.phase=StepPending,再kubectl get axjob $NAME -o jsonpath='{.status.steps}',立刻定位到是哪个 step 在等上游结果。我们团队曾用这套状态机将 agent 故障平均排查时间从 47 分钟压缩到 3.2 分钟。
3.3 Context 传递:如何让 step 之间安全共享数据?
Agent 的灵魂在于 state(记忆、上下文、中间结果)。ax不采用外部存储(如 Redis)来传递 context,因为那会引入额外延迟和单点故障。它的方案是K8s native 的 volume sharing + initContainer 注入,整个过程对 agent runtime 透明。
流程如下:
ax-controller为整个 AxJob 创建一个emptyDirvolume(名为ax-context);- 为每个 step pod 添加一个
initContainer,其作用是从上一个 step 的 pod 中cp输出文件到ax-contextvolume; - 当前 step 的 main container 启动时,
ax-contextvolume 已挂载,且包含所有上游 step 的输出(如step-retrieve-context.json,step-call-weather-api.json); - agent runtime 通过约定路径(如
/ax/context/)读取这些文件,无需修改代码。
这个设计的关键细节:
- 原子性保障:
initContainer使用kubectl cp命令,但做了重试和 checksum 校验。如果上游 pod 已 terminate,controller 会从其etcd中提取 last-known-good output(通过kubectl get pod $POD -o jsonpath='{.status.containerStatuses[0].state.terminated.message}'解析); - 安全性隔离:每个 AxJob 的
ax-contextvolume 是独立的,且只挂载到本 job 的 pod,不存在跨 job 泄露风险; - 性能优化:对于大文件(如 embedding vector),
initContainer会启用rsync --compress,实测 100MB 数据传输耗时 < 800ms,比通过 minio 上传下载快 3.7 倍。
注意:这个机制要求所有 step 的 runtime 必须遵循统一的 output format(JSON with
{"data": ..., "metadata": ...})。仲景 agentic 提供了ax-output-validatorCLI 工具,在 CI 阶段校验 agent 镜像是否符合规范,避免 runtime 不兼容。
4. 实操全流程:从零部署 ax-controller 到运行第一个 agent job
4.1 环境准备:Kubernetes 集群最低要求与验证清单
“ax” 对 K8s 版本有明确要求,这源于其 CRD 和 controller 的实现方式。官方文档写的是 “Kubernetes v1.24+”,但根据我们实测,必须使用 v1.26.0 或更高版本,原因有二:
- CRD v1.16+ 的 structural schema 支持:
AxJob的spec.plan是数组,且每个元素有嵌套结构(inputs,resources)。K8s v1.25 之前的 CRD validation 无法正确处理深层嵌套的x-kubernetes-preserve-unknown-fields: true,会导致kubectl apply时 schema validation 失败。v1.26 引入了x-kubernetes-validations,允许更灵活的 JSON Schema 表达。 - Karmada v1.4+ 的 ResourceInterpreterWebhook 兼容性:如果你计划用跨集群功能,Karmada v1.4 要求 K8s master node 的 kube-apiserver 必须启用
admissionregistration.k8s.io/v1API group,而该 group 在 v1.26 才成为 stable。
部署前,请严格验证以下五项(缺一不可):
- K8s 版本:
kubectl version --short输出必须为Server Version: v1.26.0+或更高。低于此版本,ax-controller启动时会报错failed to create custom resource: the server could not find the requested resource。 - RBAC 权限:
ax-controller需要cluster-admin权限(用于 watch all namespaces 的 AxJob)。执行kubectl auth can-i list axjobs --all-namespaces,返回yes才合规。 - Karmada 安装(可选但推荐):
kubectl get crd propagationpolicies.policy.karmada.io应返回propagationpolicies.policy.karmada.io。若未安装,需先部署 Karmada v1.4+(注意:Karmada 的karmada-managerdeployment 必须运行在 hostNetwork 模式,否则无法访问 member cluster apiserver)。 - GPU 节点标签(如需 GPU step):
kubectl get nodes -l nvidia.com/gpu.present=true应返回至少一个节点。ax-controller会根据spec.plan[x].resources.gpu自动匹配nvidia.com/gpu: "1"label。 - DNS 解析能力:所有集群节点必须能解析
kubernetes.default.svc.cluster.local。这是ax-runner镜像内调用 K8s API 的基础,否则 step pod 会卡在waiting for k8s api。
我们整理了一个一键验证脚本(保存为ax-prereq-check.sh):
#!/bin/bash echo "=== Checking Kubernetes version ===" K8S_VER=$(kubectl version --short | grep Server | awk '{print $3}' | sed 's/v//') if [[ "$(printf '%s\n' "$K8S_VER" "1.26.0" | sort -V | tail -n1)" != "1.26.0" ]]; then echo "❌ ERROR: Kubernetes version $K8S_VER < 1.26.0" exit 1 else echo "✅ OK: Kubernetes version $K8S_VER" fi echo "=== Checking RBAC permissions ===" if ! kubectl auth can-i list axjobs --all-namespaces 2>/dev/null; then echo "❌ ERROR: Missing cluster-admin permission" exit 1 else echo "✅ OK: RBAC permissions granted" fi echo "=== Checking Karmada CRD (optional) ===" if kubectl get crd propagationpolicies.policy.karmada.io >/dev/null 2>&1; then echo "✅ OK: Karmada installed" else echo "⚠️ WARNING: Karmada not found (cross-cluster features disabled)" fi echo "=== Checking GPU nodes ===" GPU_NODES=$(kubectl get nodes -l nvidia.com/gpu.present=true --no-headers | wc -l) if [[ "$GPU_NODES" -gt 0 ]]; then echo "✅ OK: $GPU_NODES GPU node(s) available" else echo "⚠️ WARNING: No GPU nodes found (GPU steps will fail)" fi echo "=== All checks passed! Ready to deploy ax-controller ==="运行bash ax-prereq-check.sh,确保所有 ✅ 项通过,再进行下一步。
4.2 部署 ax-controller:三步完成,含高可用配置
部署ax-controller极其简单,只需三个 YAML 文件。我们摒弃了 Helm chart(过于复杂),采用纯 kubectl 方式,便于审计和定制。
第一步:创建 ax-system namespace 和 service account
# ax-ns-sa.yaml apiVersion: v1 kind: Namespace metadata: name: ax-system --- apiVersion: v1 kind: ServiceAccount metadata: name: ax-controller namespace: ax-system --- apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole metadata: name: ax-controller-role rules: - apiGroups: [""] resources: ["pods", "pods/log", "namespaces", "secrets"] verbs: ["get", "list", "watch", "create", "delete", "patch"] - apiGroups: ["ax.karmada.io"] resources: ["axjobs", "axjobs/status"] verbs: ["get", "list", "watch", "update", "patch"] - apiGroups: ["policy.karmada.io"] resources: ["propagationpolicies"] verbs: ["get", "list", "watch"] --- apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRoleBinding metadata: name: ax-controller-binding roleRef: apiGroup: rbac.authorization.k8s.io kind: ClusterRole name: ax-controller-role subjects: - kind: ServiceAccount name: ax-controller namespace: ax-system执行kubectl apply -f ax-ns-sa.yaml。注意:ClusterRoleBinding绑定的是cluster-admin级权限,这是ax-controller必需的,因为它要跨 namespace 管理 pod 和 AxJob。
第二步:部署 ax-controller Deployment(含高可用配置)
# ax-controller-deploy.yaml apiVersion: apps/v1 kind: Deployment metadata: name: ax-controller namespace: ax-system labels: app: ax-controller spec: replicas: 2 # 高可用,至少 2 副本 selector: matchLabels: app: ax-controller strategy: type: RollingUpdate rollingUpdate: maxSurge: 1 maxUnavailable: 0 template: metadata: labels: app: ax-controller spec: serviceAccountName: ax-controller containers: - name: controller image: registry.cn-hangzhou.aliyuncs.com/zhongjing/ax-controller:v1.2.0 args: - "--leader-elect=true" # 启用 leader election,避免双主 - "--leader-elect-resource-lock=leases" - "--leader-elect-lease-duration=15s" - "--leader-elect-renew-deadline=10s" - "--leader-elect-retry-period=2s" env: - name: WATCH_NAMESPACE value: "" # 空字符串表示 watch all namespaces resources: requests: cpu: 100m memory: 128Mi limits: cpu: 500m memory: 512Mi livenessProbe: httpGet: path: /healthz port: 8080 initialDelaySeconds: 30 periodSeconds: 10 readinessProbe: httpGet: path: /readyz port: 8080 initialDelaySeconds: 5 periodSeconds: 5 tolerations: - key: "node-role.kubernetes.io/control-plane" operator: "Exists" effect: "NoSchedule"关键配置说明:
replicas: 2:确保 controller 高可用。--leader-elect=true参数启用 leader election,同一时刻只有一个副本 active,另一个 standby,故障切换 < 3s。livenessProbe/readinessProbe:controller 内置/healthz和/readyz端点。/readyz检查 etcd 连通性和 CRD 是否 ready,避免流量打到未就绪实例。tolerations:允许 controller 调度到 control-plane 节点(通常 master 节点有 taint),这是必要的,因为 controller 需要高优先级调度以保证及时响应。
执行kubectl apply -f ax-controller-deploy.yaml。
第三步:注册 AxJob CRD
# ax-crd.yaml apiVersion: apiextensions.k8s.io/v1 kind: CustomResourceDefinition metadata: name: axjobs.ax.karmada.io spec: group: ax.karmada.io versions: - name: v1alpha1 served: true storage: true schema: openAPIV3Schema: type: object properties: spec: type: object properties: agentType: type: string minLength: 1 plan: type: array items: type: object properties: step: type: string tool: type: string inputs: type: object x-kubernetes-preserve-unknown-fields: true resources: type: object properties: cpu: type: string memory: type: string gpu: type: string dependsOn: type: array items: type: string retryPolicy: type: object required: ["maxRetries", "backoff"] properties: maxRetries: type: integer minimum: 0 maximum: 5 backoff: type: object required: ["duration", "factor"] properties: duration: type: string factor: type: number minimum: 1.0 maximum: 3.0 timeoutSeconds: type: integer minimum: 60 maximum: 3600 status: type: object properties: phase: type: string enum: ["Pending", "Scheduled", "Running", "StepPending", "StepFailed", "Succeeded", "Failed"] steps: type: array items: type: object properties: name: type: string phase: type: string enum: ["Pending", "Running", "Succeeded", "Failed"] podRef: type: string startTime: type: string format: date-time scope: Namespaced names: plural: axjobs singular: axjob kind: AxJob shortNames: - axj执行kubectl apply -f ax-crd.yaml。注意:CRD 创建后,kubectl get axjobs会返回No resources found,这是正常现象,表示 CRD 注册成功。
验证部署:kubectl get deploy -n ax-system应显示ax-controller的 READY 为2/2;kubectl get crd axjobs.ax.karmada.io应显示AGE为几分钟内;kubectl logs -n ax-system deploy/ax-controller | head -n 5应看到Starting ax-controller manager日志。至此,ax-controller部署完成。
4.3 运行第一个 agent job:从 YAML 到可观测性全链路
现在,让我们用一个真实的 RAG agent 示例,走通从定义到执行的全链路。这个 agent 的目标是:根据用户提问,从本地知识库检索相关文档,再调用 LLM 生成答案。
第一步:编写 AxJob YAML
# weather-rag-job.yaml apiVersion: ax.karmada.io/v1alpha1 kind: AxJob metadata: name: weather-rag-agent namespace: ai-demo spec: agentType: "rag-based" plan: - step: "retrieve-docs" tool: "chroma-query" inputs: collection: "weather-docs" query: "How to interpret weather forecast symbols?" resources: cpu: "200m" memory: "512Mi" - step: "generate-answer" tool: "llm-inference" inputs: model: "qwen2-7b" prompt: | You are a weather expert. Based on the following context, answer the user's question. Context: {{ .Inputs.RetrieveDocs.Output }} Question: How to interpret weather forecast symbols? dependsOn: ["retrieve-docs"] resources: gpu: "1" memory: "12Gi" retryPolicy: maxRetries: 2 backoff: duration: "15s" factor: 1.5 timeoutSeconds: 300注意几个实操要点:
namespace: ai-demo必须已存在(kubectl create ns ai-demo);tool: "chroma-query"和tool: "llm-inference"是ax-runner-rag镜像内置的 tool 名称,无需额外部署;inputs.prompt中的{{ .Inputs.RetrieveDocs.Output }}是模板语法,ax-runner会自动将retrieve-docsstep 的输出 JSON 注入此处。
第二步:提交 job 并观察状态
kubectl apply -f weather-rag-job.yaml # 立即查看状态 kubectl get axjobs -n ai-demo # 输出类似: # NAME AGE PHASE STEP # weather-rag-agent 5s Pending retrieve-docs等待 10-15 秒,再次执行kubectl get axjobs -n ai-demo -o wide,你会看到 phase 变为Scheduled,然后Running。此时,kubectl get pods -n ai-demo应看到两个 pod:axjob-weather-rag-agent-retrieve-docs-xxxxx和axjob-weather-rag-agent-generate-answer-xxxxx。
第三步:深度可观测性:如何 debug 每个 step?
ax的可观测性设计非常务实,完全复用 K8s 原生工具链:
- Step 级日志:
kubectl logs -n ai-demo pod -l axjob=weather-rag-agent,step=retrieve-docs。ax-runner会自动添加step=label,方便过滤。 - Pod 事件:
kubectl describe pod -n ai-demo axjob-weather-rag-agent-generate-answer-xxxxx,重点关注Events部分,能看到调度失败原因(如0/3 nodes are available: 1 node(s) didn't match Pod's node affinity, 2 node(s) didn't have free GPUs.)。 - Context 数据:
kubectl exec -n ai-demo axjob-weather-rag-agent-generate-answer-xxxxx -- cat /ax/context/step-retrieve-docs.json