1. 先搞清楚 ML2025 Homework 2 到底在预测什么
ML2025 Homework 2 的 Disease Prediction 任务,本质是一个回归问题:给定美国某州过去两天的调查问卷结果,预测第 3 天新增检测阳性病例的百分比。评估指标是 MSE,越小越好。数据集里每条样本有 90 多个特征,包括口罩佩戴情况、社交距离、症状统计等,标签是tested_positive。
这个作业的特殊之处在于:它要求你用大模型生成代码来解决问题,而不是自己手写模型。换句话说,你要搭建一个 LLM Agent 框架,让模型自己完成数据探索、特征工程、模型选择、训练和预测的全流程。听起来很酷,但实际操作中会遇到几个坑:LLM 生成的代码经常跑不通、特征维度对不上、归一化逻辑写反、提交格式不对导致分数异常。
我试过用 Qwen-Coder 系列模型配合 Agent 循环来跑这个任务,核心思路是让模型先生成方案再写代码,执行后根据报错自动 debug,反复迭代直到 MSE 收敛。下面我会把整个流程拆开,包括 config.toml 骨架、TaoToken 统一 Key 配置、可复制的 Agent 代码,以及验证和排错步骤。
适合谁看:正在做 ML2025 Homework 2 的同学、想了解 LLM 代码生成 Agent 怎么落地的开发者、以及需要统一管理多个模型 API 的工程师。读完你能拿到一套可运行的配置模板,直接替换数据路径就能跑。
2. TaoToken 前置:统一 Key 与 API 通道配置
在跑 Agent 之前,你需要一个能稳定调用多个开源模型的 API 通道。作业明确禁止使用 GPT-4o、Gemini 等闭源 API,所以 Qwen-Coder、DeepSeek-Coder 这类开源模型是首选。TaoToken 提供统一的 API 入口,你只需要一个 Key 就能切换不同模型,不用为每个模型单独申请账号。
官网地址:https://taotoken.net/?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content=
API 端点:https://taotoken.net/api
配置方式很简单,在环境变量里设置OPENAI_API_KEY和OPENAI_BASE_URL即可。TaoToken 兼容 OpenAI SDK 的调用格式,所以你的代码里用openai.OpenAI()初始化就能直接跑。模型名称填qwen-coder-plus-latest或者你需要的其他开源模型。
如果你需要长期跑 Agent 循环、频繁调用模型,建议看一下 Coding Plan 页面,有更划算的套餐。只是临时跑一次作业的话,按量付费就够了。
注意:不要把 Key 硬编码在代码里提交到 GitHub,用环境变量或者
.env文件管理。
3. 可复制配置:config.toml 骨架与 Agent 核心代码
3.1 config.toml 骨架
先建一个config.toml,把实验参数、数据路径、Agent 迭代次数都放进去。这样你换数据集或者调参时不用改代码。
[experiment] exp_name = "ML2025_HW2" data_dir = "./ML2025Spring-hw2-public" task_goal = "Given the survey results from the past two days in a specific state in the U.S., predict the probability of testing positive on day 3. The evaluation metric is Mean Squared Error (MSE)." [agent] steps = 5 debug_prob = 0.5 num_drafts = 5 [model] name = "qwen-coder-plus-latest" max_tokens = 8192 temperature = 0.3对应的 Python 配置加载:
import tomllib from pathlib import Path with open("config.toml", "rb") as f: cfg = tomllib.load(f) data_dir = Path(cfg["experiment"]["data_dir"]).resolve() task_goal = cfg["experiment"]["task_goal"] steps = cfg["agent"]["steps"] num_drafts = cfg["agent"]["num_drafts"] debug_prob = cfg["agent"]["debug_prob"] model_name = cfg["model"]["name"]3.2 TaoToken 调用封装
把模型调用单独封装成一个函数,方便 Agent 的 draft、improve、debug 三个环节复用。
import os import openai client = openai.OpenAI( api_key=os.getenv("OPENAI_API_KEY"), base_url=os.getenv("OPENAI_BASE_URL", "https://taotoken.net/api"), ) def generate_response(prompt, system_prompt="", max_tokens=8192, temperature=0.3): response = client.chat.completions.create( model=model_name, messages=[ {"role": "system", "content": system_prompt}, {"role": "user", "content": prompt}, ], max_tokens=max_tokens, temperature=temperature, ) return response.choices[0].message.content3.3 Agent 的 draft 环节
draft 环节让模型根据数据预览生成第一版方案和代码。关键是把数据目录、任务目标、输出路径都写进 prompt,模型才知道要干什么。
def _draft(self): system_prompt = """You are an expert AI agent specializing in time series prediction. You are running on Ubuntu 22.04.5 LTS with Python 3.11. Your task is to develop a solution that predicts testing probabilities based on survey data. Focus on minimizing Mean Squared Error (MSE).""" user_prompt = [ "Task: Develop a machine learning model to predict testing probabilities.", f"Goal: {task_goal}", f"Data Location: {data_dir}", f"Data Overview:\n{self.data_preview}", "Requirements:", "1. Save predictions to '/content/submission.csv'", "2. The testing file DOES NOT have the target column", "3. Implement proper data preprocessing", "4. Use appropriate model selection and validation", "\nNeed to provide:", "1. A detailed plan explaining your approach", "2. Full Python implementation", ] plan, code = self.plan_and_code_query(system_prompt, "\n".join(user_prompt)) return Node(plan=plan, code=code)3.4 执行与结果解析
模型生成代码后,你需要执行它并解析输出。执行结果里包含 MSE 和是否报错,这些信息会决定下一步是 improve 还是 debug。
def parse_exec_result(self, node, exec_result): node.absorb_exec_result(exec_result) user_prompt = f"""Task: Evaluate the prediction model implementation Original Goal: {task_goal} Implementation: {wrap_code(node.code)} Execution Output: {wrap_code(node.term_out, lang="")} Need to provide a structured analysis including: 1. Execution Status (Success/Error) 2. Performance Metrics (especially MSE) 3. Issues or Concerns (if any) 4. Overall Assessment""" response = generate_response(user_prompt) try: parsed = json.loads(response) node.analysis = parsed.get("assessment", "") node.metric = parsed.get("mse", 0.0) node.is_buggy = parsed.get("execution_status") == "Error" or node.exc_type is not None except Exception as e: print("Failed to parse evaluation response:", e) node.is_buggy = False node.metric = 0.04. 验证请求与成功结果比对
4.1 先验证 API 通道是否通
在跑完整 Agent 之前,先用一个简单请求确认 TaoToken 通道正常。
result = generate_response("你是谁") print(result)如果返回模型自我介绍,说明 Key 和 base_url 配置正确。如果报 401,检查环境变量是否生效;如果报 404,检查模型名称是否拼写正确。
4.2 跑一次完整 Agent 循环
把 draft、improve、debug 串起来,设置steps=5、num_drafts=5,让 Agent 迭代 5 轮。每轮结束后打印当前 MSE 和节点状态。
for step in range(steps): parent_node = agent.search_policy() if parent_node is None: result_node = agent._draft() elif parent_node.is_buggy: result_node = agent._debug(parent_node) else: result_node = agent._improve(parent_node) agent.parse_exec_result(result_node, exec_callback(result_node.code, True)) agent.journal.append(result_node) print(f"Step {step+1}: MSE={result_node.metric:.4f}, buggy={result_node.is_buggy}")4.3 结果比对
我实测下来,Agent 生成的代码在 5 轮迭代后 MSE 能到 0.947 左右。作为对照,手写的 baseline 模型(用 SelectKBest 选特征 + 三层 MLP + Adam 优化器)能到 0.83。差距主要来自特征工程:手写版本手动去掉了第 0 列无关特征,选了 20 个关键特征,而 LLM 生成的代码往往把所有特征都塞进去,导致过拟合。
你可以把 Agent 生成的submission.csv和手写版本的predictions.csv都提交到比赛平台,对比 MSE。如果 Agent 版本分数明显偏高,检查两点:归一化是否用了训练集的 min/max 而不是全局的;测试集是否误用了标签列。
5. 本篇常见错排查
5.1 特征维度对不上
报错信息通常是RuntimeError: mat1 and mat2 shapes cannot be multiplied。原因是模型输入维度是 90,但测试集经过特征选择后只剩 20 列。解决办法:在select_features函数里确保训练集、验证集、测试集用同一套feat_idx。
feat_idx = [34, 35, 36, 43, 46, 47, 51, 52, 53, 54, 61, 64, 65, 69, 70, 71, 72, 79, 82, 83] x_train = raw_x_train[:, feat_idx] x_valid = raw_x_valid[:, feat_idx] x_test = raw_x_test[:, feat_idx]5.2 归一化逻辑写反
LLM 生成的代码有时会用测试集的 min/max 做归一化,导致数据泄露。正确做法是用训练集的 min 和 range 去归一化验证集和测试集。
train_min = np.min(train_data[:, 35:-1], axis=0) train_max = np.max(train_data[:, 35:-1], axis=0) train_range = train_max - train_min + 1e-8 train_data[:, 35:-1] = (train_data[:, 35:-1] - train_min) / train_range valid_data[:, 35:-1] = (valid_data[:, 35:-1] - train_min) / train_range test_data[:, 35:] = (test_data[:, 35:] - train_min) / train_range5.3 提交文件格式错误
比赛要求 CSV 有两列:id和tested_positive。LLM 有时会写成index或者多加一列。检查你的save_pred函数:
def save_pred(preds, file): with open(file, "w") as fp: writer = csv.writer(fp) writer.writerow(["id", "tested_positive"]) for i, p in enumerate(preds): writer.writerow([i, p])5.4 API 调用超时或限流
Agent 循环会频繁调用模型,如果遇到 429 错误,降低steps或者加一个time.sleep(1)在每次调用之间。TaoToken 的 Coding Plan 有更高的并发额度,长期跑的话可以考虑。
6. 接入文档与模型对话入口
如果你在配置过程中遇到 Key 无效、模型名称不对、返回格式异常等问题,可以直接去 API Keys 页面重新生成 Key,或者查看接入文档里的详细参数说明。
- API Keys 管理:https://taotoken.net/api-keys?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content=apikeys
- 接入文档:https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content=doc
- 模型对话测试:https://taotoken.net/chat?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content=chat
想先验证模型能不能正常生成代码,用模型对话页面发一个简单的 Python 任务试试。如果打算长期跑 Agent 做编码任务,Coding Plan 页面有更详细的套餐对比。
最后说一个实用技巧:Agent 生成的代码不要直接提交,先本地跑一遍python best.py确认没有语法错误和维度问题。LLM 有时候会生成看起来对但实际跑不通的代码,尤其是涉及 PyTorch 张量操作的部分。把torch.save的模型文件保留下来,下次换数据可以直接加载权重做推理,省去重新训练的时间。