news 2026/9/29 11:39:25

Intern-S2-397B 使用指南:科学智能多模态模型的推理、工具调用与 Agent 集成

作者头像

张小明

前端开发工程师

1.2k 24
文章封面图
Intern-S2-397B 使用指南:科学智能多模态模型的推理、工具调用与 Agent 集成
  • 人工智能
  • 大模型
  • 基础模型
  • 多模态
  • 预训练
  • 强化学习

【免费下载链接】Intern-S2-397B

项目地址:https://ai.gitcode.com/InternLM/Intern-S2-397B
点击查看免费下载

本篇指南以开源模型仓库 InternLM/Intern-S2-397B 的官方 README 为核心,系统讲解该 397B 级多模态基础模型在科学智能与长程 Agent 场景下的完整使用链路:包括推荐的采样超参数、基于 LMDeploy / vLLM / SGLang 的服务化部署、OpenAI 兼容 API 的工具调用(Tool Calling)、思考/非思考模式切换、时间序列推理与预测,以及将模型接入自托管或官方 API 的 Agent 框架。读完本文,你将能独立完成 Intern-S2-397B 的部署配置、客户端调用和 Agent 集成,并理解其对话模板与底层分词实现的关键细节。

模型定位与核心特性

Intern-S2-397B 是面向**科学智能(Scientific Intelligence)与长程 Agent(Long-Horizon Agents)**的多模态基础模型,其能力构建沿三条关键维度扩展:预训练(pre-training)、强化学习任务覆盖(reinforcement-learning task coverage)与交互式 Agent 环境(interactive agent environments)。官方 README 将其核心特性概括为三点:

  1. 新的预训练范式(New Pre-training Paradigm):通过视觉预训练(visual pretraining),模型直接从科学文献的原始页面学习,在共享表示空间中联合建模符号语义(symbolic semantics)与视觉关系(visual relationships),无需中间的解析步骤。该范式保留文本-视觉对应关系,强化空间与视觉推理能力,并提升数据效率。
  2. 科学模态推理与生成(Scientific Modality Reasoning and Generation):通过在超过 20 个科学领域的强化学习任务上联合训练,模型在通用推理与专项科学任务(如生物分子相互作用设计、材料结构生成)上均取得较强表现。
  3. 通用与科学长程 Agent(General & Scientific Long-Horizon Agents):将多个 Agent 框架接入大规模沙箱环境进行黑盒 Agent 强化学习,提升泛化能力并抬升长程任务的能力上限。

从仓库的 config.json 可以看出该模型的底层实现骨架:架构类型为Qwen3_5MoeForConditionalGeneration,文本部分采用 60 层混合注意力结构(layer_types中每 4 层出现一个full_attention,其余为linear_attention),包含 512 个专家(num_experts)、每 token 激活 10 个专家(num_experts_per_tok),隐藏维度 4096,并配置了mtp_num_hidden_layers: 1的 MTP 投机解码模块;视觉编码器为 27 层(depth),patch size 16、时间 patch size 2(temporal_patch_size),输出隐层 4096 与文本侧对齐。权重索引 model.safetensors.index.json 显示模型权重总大小约 807 GB(BF16),分片为 188 个 safetensors 文件,并包含独立的mtp.*权重,与 MTP 加速推理相对应。

评测方法与基准

官方使用 OpenCompass。

关于该图的图例说明(原文档原意):下划线(Underline)表示开源模型中最佳表现,加粗(Bold)表示所有模型中的最佳表现。评测时的推理长度约束如下,复现评测时务必注意:

  • 文本推理基准:最大推理长度 256K tokens;
  • 多模态基准:最大推理长度 64K tokens。

需要说明的是,基准数字与具体榜单内容以官方发布为准,本文不再展开逐项数值;上述评测工具与长度配置已在 README.md 中明确说明,可作为复现时的基础约束。

快速开始:采样参数与服务化部署

推荐的采样参数

官方建议使用以下超参数以获得更好的生成结果:

top_p = 0.95 top_k = 50 min_p = 0.0 temperature = 0.8

需要留意的是,仓库自带的 generation_config.json 中的默认值(temperature: 0.6、top_k: 20、top_p: 0.95、do_sample: true)是模型发布时的生成配置,而 README 推荐的这组参数更适合作为在线服务请求中的显式采样参数(下文所有示例均显式传入temperature=0.8、top_p=0.95)。若追求确定性输出(如时序分析示例中采用temperature=0),则按具体任务覆盖即可。

支持的服务框架

Intern-S2-397B 可选用以下任一 LLM 推理框架部署:

  • LMDeploy(推荐,功能支持最完整,含时间序列推理)
  • vLLM
  • SGLang

各框架的详细部署示例(基础服务、MTP 投机解码、长上下文 YaRN 配置)见仓库中的 部署指南,本文后面的"部署配置速览"一节会提炼其要点。

部署后的服务同时暴露 OpenAI 兼容接口(/v1/chat/completions)与 Anthropic 兼容接口(/v1/messages),前者用于普通客户端与多数 Agent 框架,后者用于 Claude Code 等工具。

高级用法:Tool Calling(工具调用)

Tool Calling 让模型通过调用外部工具与 API 扩展能力。下面给出官方完整示例:基于 LMDeploy api server 提供的 OpenAI 兼容 API,让模型查询旧金山今明两天的天气。整个流程包含四个阶段:定义工具函数 → 声明 tools 描述 → 发起首轮请求 → 执行工具并把结果回填后发起第二轮请求。

from openai import OpenAI import json def get_current_temperature(location: str, unit: str = "celsius"): """Get current temperature at a location. Args: location: The location to get the temperature for, in the format "City, State, Country". unit: The unit to return the temperature in. Defaults to "celsius". (choices: ["celsius", "fahrenheit"]) Returns: the temperature, the location, and the unit in a dict """ return { "temperature": 26.1, "location": location, "unit": unit, } def get_temperature_date(location: str, date: str, unit: str = "celsius"): """Get temperature at a location and date. Args: location: The location to get the temperature for, in the format "City, State, Country". date: The date to get the temperature for, in the format "Year-Month-Day". unit: The unit to return the temperature in. Defaults to "celsius". (choices: ["celsius", "fahrenheit"]) Returns: the temperature, the location, the date and the unit in a dict """ return { "temperature": 25.9, "location": location, "date": date, "unit": unit, } def get_function_by_name(name): if name == "get_current_temperature": return get_current_temperature if name == "get_temperature_date": return get_temperature_date tools = [{ 'type': 'function', 'function': { 'name': 'get_current_temperature', 'description': 'Get current temperature at a location.', 'parameters': { 'type': 'object', 'properties': { 'location': { 'type': 'string', 'description': 'The location to get the temperature for, in the format \'City, State, Country\'.' }, 'unit': { 'type': 'string', 'enum': [ 'celsius', 'fahrenheit' ], 'description': 'The unit to return the temperature in. Defaults to \'celsius\'.' } }, 'required': [ 'location' ] } } }, { 'type': 'function', 'function': { 'name': 'get_temperature_date', 'description': 'Get temperature at a location and date.', 'parameters': { 'type': 'object', 'properties': { 'location': { 'type': 'string', 'description': 'The location to get the temperature for, in the format \'City, State, Country\'.' }, 'date': { 'type': 'string', 'description': 'The date to get the temperature for, in the format \'Year-Month-Day\'.' }, 'unit': { 'type': 'string', 'enum': [ 'celsius', 'fahrenheit' ], 'description': 'The unit to return the temperature in. Defaults to \'celsius\'.' } }, 'required': [ 'location', 'date' ] } } }] messages = [ {'role': 'user', 'content': 'Today is 2024-11-14, What\'s the temperature in San Francisco now? How about tomorrow?'} ] openai_api_key = "EMPTY" openai_api_base = "http://0.0.0.0:23333/v1" client = OpenAI( api_key=openai_api_key, base_url=openai_api_base, ) model_name = "internlm/Intern-S2-397B" # Must match the model ID served by your deployment. response = client.chat.completions.create( model=model_name, messages=messages, max_tokens=32768, temperature=0.8, top_p=0.95, extra_body=dict(spaces_between_special_tokens=False), tools=tools) print(response.choices[0].message) messages.append(response.choices[0].message) for tool_call in response.choices[0].message.tool_calls: tool_call_args = json.loads(tool_call.function.arguments) tool_call_result = get_function_by_name(tool_call.function.name)(**tool_call_args) tool_call_result = json.dumps(tool_call_result, ensure_ascii=False) messages.append({ 'role': 'tool', 'name': tool_call.function.name, 'content': tool_call_result, 'tool_call_id': tool_call.id }) response = client.chat.completions.create( model=model_name, messages=messages, temperature=0.8, top_p=0.95, extra_body=dict(spaces_between_special_tokens=False), tools=tools) print(response.choices[0].message)

使用要点:

  • model_name必须与部署时指定的模型 ID(internlm/Intern-S2-397B)完全一致;
  • extra_body=dict(spaces_between_special_tokens=False)用于保证特殊 token 两侧不加多余空格,确保工具调用格式被正确解析;
  • 工具结果需以role='tool'回填,并带上tool_call_id与首轮返回的 tool call 对应。

从底层对话模板 chat_template.jinja 可以印证工具调用的协议格式:当请求携带tools时,模板会在 system 消息中注入# Tools与<tools>...</tools>描述块,并约束模型"仅在需要调用函数时以固定 XML 格式回复,不得附带任何后缀",格式为嵌套在<tool_call>内的<function=name>与<parameter=key>value</parameter>;工具执行结果则包装在<tool_response>...</tool_response>中回填给模型,从而驱动多轮工具调用。这也解释了为什么必须为服务端配置对应的工具调用解析器(LMDeploy 下为--tool-call-parser interns2-preview,见部署章节)。

高级用法:思考(Thinking)与非思考模式切换

Intern-S2-397B默认开启思考模式(thinking mode),即在生成正式回答前先输出<think>推理过程,从而提升回答质量。思考模式可以通过在tokenizer.apply_chat_template中设置enable_thinking=False来关闭:

text = tokenizer.apply_chat_template( messages, tokenize=False, add_generation_prompt=True, enable_thinking=False # think mode indicator )

在线服务场景下,可以通过请求中的chat_template_kwargs.enable_thinking参数动态控制:

from openai import OpenAI import json messages = [ { 'role': 'user', 'content': 'who are you' }, { 'role': 'assistant', 'content': 'I am an AI' }, { 'role': 'user', 'content': 'AGI is?' }] openai_api_key = "EMPTY" openai_api_base = "http://0.0.0.0:23333/v1" client = OpenAI( api_key=openai_api_key, base_url=openai_api_base, ) model_name = "internlm/Intern-S2-397B" # Must match the model ID served by your deployment. response = client.chat.completions.create( model=model_name, messages=messages, temperature=0.8, top_p=0.95, max_tokens=2048, extra_body={ "chat_template_kwargs": {"enable_thinking": False} } ) print(json.dumps(response.model_dump(), indent=2, ensure_ascii=False))

官方同时给出明确建议:不建议在 Agent 任务中关闭思考模式(Note: We do not recommend disabling thinking mode for agentic tasks.),因为思考过程有助于 Agent 在长程任务中保持规划与决策质量。

结合 chat_template.jinja 的实现可以看到思考模式的机制:模板在拼接add_generation_prompt时会写入<|im_start|>assistant\n并追加<think>\n引导模型进入思考;而当enable_thinking is false(或请求中已包含时间序列数据)时,则改为追加<think>\n\n</think>\n\n的空思考块,直接进入正式回复。此外,模板会识别 assistant 消息中已有的<think>...</think>内容并将其剥离为reasoning_content,便于服务端按reasoning_content与content分离返回推理过程与最终答案。

高级用法:时间序列推理与预测(Time Series Demo)

时间序列推理目前仅在 LMDeploy 框架中支持。先按照 部署指南 用 LMDeploy 部署 Intern-S2-397B,再通过lmdeploy.vl.utils.encode_time_series_base64将时序文件编码为 base64 传入。下方示例从时间序列信号文件中检测地震事件(Earthquake event),并输出 P 波与 S 波的起始时间点索引。

请注意:在消息内容中,time_series_url与文本提示词(text prompt)的出现顺序可以任意排列。

from openai import OpenAI from lmdeploy.vl.utils import encode_time_series_base64 openai_api_key = "EMPTY" openai_api_base = "http://0.0.0.0:8000/v1" client = OpenAI( api_key=openai_api_key, base_url=openai_api_base, ) model_name = "internlm/Intern-S2-397B" # Must match the model ID served by your deployment. def send_base64(file_path: str, sampling_rate: int = 100): """base64-encoded time-series data.""" # encode_time_series_base64 accepts local file paths and http urls, # encoding time-series data (.npy, .csv, .wav, .mp3, .flac, etc.) into base64 strings. base64_ts = encode_time_series_base64(file_path) messages = [ { "role": "user", "content": [ { "type": "time_series_url", "time_series_url": { "url": f"data:time_series/npy;base64,{base64_ts}", "sampling_rate": sampling_rate }, }, { "type": "text", "text": "Please determine whether an Earthquake event has occurred in the provided time-series data. If so, please specify the starting time point indices of the P-wave and S-wave in the event." }, ], } ] return client.chat.completions.create( model=model_name, messages=messages, temperature=0, max_tokens=200, extra_body={ "chat_template_kwargs": {"enable_thinking": False} } ) def send_http_url(url: str, sampling_rate: int = 100): """http(s) url pointing to the time-series data.""" messages = [ { "role": "user", "content": [ { "type": "time_series_url", "time_series_url": { "url": url, "sampling_rate": sampling_rate }, }, { "type": "text", "text": "Please determine whether an Earthquake event has occurred in the provided time-series data. If so, please specify the starting time point indices of the P-wave and S-wave in the event." }, ], } ] return client.chat.completions.create( model=model_name, messages=messages, temperature=0, max_tokens=200, extra_body={ "chat_template_kwargs": {"enable_thinking": False} } ) def send_file_url(file_path: str, sampling_rate: int = 100): """file url pointing to the time-series data.""" messages = [ { "role": "user", "content": [ { "type": "time_series_url", "time_series_url": { "url": f"file://{file_path}", "sampling_rate": sampling_rate }, }, { "type": "text", "text": "Please determine whether an Earthquake event has occurred in the provided time-series data. If so, please specify the starting time point indices of the P-wave and S-wave in the event." }, ], } ] return client.chat.completions.create( model=model_name, messages=messages, temperature=0, max_tokens=200, extra_body={ "chat_template_kwargs": {"enable_thinking": False} } ) response = send_base64("./0092638_seism.npy") # response = send_http_url("https://huggingface.co/internlm/Intern-S1-Pro/raw/main/0092638_seism.npy") # response = send_file_url("./0092638_seism.npy") print(response.choices[0].message)

三个入口函数分别演示了三种数据来源:本地文件 base64 编码(send_base64)、HTTP(S) 直链(send_http_url)与file://本地路径(send_file_url)。其中encode_time_series_base64接受本地文件路径与 http url,可编码.npy、.csv、.wav、.mp3、.flac等格式的时序数据。sampling_rate默认 100,表示采样率信息。

从 chat_template.jinja 可以看出时间序列在模板层的表达:内容项中的time_series类型会被渲染为<|ts|><TS_CONTEXT><|/ts|>标记对(对应 tokenizer_config.json 中 id 248091/248092/248093 的特殊 token),同时模板会检测消息中是否含时间序列,若包含则生成"空思考块"直接进入回答,这与上述示例中显式传enable_thinking: False的目的是一致的——时序任务直接给出分析结果,不进入冗长思考。

时间序列预测(Forecasting)

对于时间序列预测任务,forecast_horizon是可选参数:设置为整数时,模型将输出恰好该长度的预测序列;设置为None时,模型根据文本提示词自行推断预测步长(horizon)。官方示例为电力负荷预测(electric load forecasting),输入历史每半小时的负荷数据与温度、湿度、气压等气象信息,预测未来 24 小时(48 个时间点)的负荷:

def forecast_base64(file_path: str, forecast_horizon: int | None = None): base64_ts = encode_time_series_base64(file_path) messages = [ { "role": "user", "content": [ { "type": "text", "text": ( "Please complete a electric load forecasting task. " "This dataset is based on historical electricity load data every half hour within 24 hours of the region, " "as well as data on minimum temperature, maximum temperature, humidity, air pressure, etc., " "to predict future load consumption every half hour within 24 hours. Here is the weather information for city TAS: " "Historical date weather: minimum temperature of 279.71K, maximum temperature of 285.83K, humidity of 85.0%, " "air pressure of 1003.0hPa. Forecast date weather: minimum temperature 280.54K, maximum temperature 286.47K, " "humidity 74.0%, air pressure 1007.0hPa. This data has no relevant effect information. " "Please predict the next 48 time points given information above." ), }, { "type": "time_series_url", "time_series_url": { "url": f"data:time_series/npy;base64,{base64_ts}", }, }, ], } ] return client.chat.completions.create( model=model_name, messages=messages, temperature=0, max_tokens=16, extra_body={ "chat_template_kwargs": {"enable_thinking": False}, "enable_forecasting": True, "forecast_horizon": forecast_horizon, }, ) response = forecast_base64("./load_20210803_0.npy", forecast_horizon=None) forecast = response.choices[0].message.ts_forecast print("Point forecast:", forecast.point_forecast) print("Quantile forecast:", forecast.quantile_forecast)

预测模式下需要额外传入两个请求级参数:enable_forecasting=True(开启预测输出)与forecast_horizon(预测步长)。返回结果位于message.ts_forecast,包含point_forecast(点预测)与quantile_forecast(分位数预测)两部分。注意预测任务中max_tokens只需 16,因为模型输出的是紧凑的预测 token 流而非长文本。

Agent 集成(Agent Integration)

Intern-S2-397B 可以通过两种方式接入 Agent 框架:连接自托管部署,或调用官方 Intern API。两种方式均同时覆盖主流 Agent 框架(OpenClaw、Hermes 等,它们都接受 OpenAI 兼容端点)与 Claude Code(使用 Anthropic 兼容端点)。

方式一:自托管部署(以 LMDeploy 为例)

先用 LMDeploy 按 部署指南 部署模型,假设服务运行在http://0.0.0.0:23333。

连接 Agent 框架:将 Agent 框架的 OpenAI 兼容端点指向 LMDeploy 服务的 base urlhttp://0.0.0.0:23333/v1。可用以下 curl 命令验证连通性:

curl http://0.0.0.0:23333/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer EMPTY" \ -d '{ "model": "internlm/Intern-S2-397B", "messages": [ {"role": "user", "content": "Hello"} ], "temperature": 0.8, "top_p": 0.95 }'

也可以在 Agent 框架中通过环境变量完成配置:

export OPENAI_API_KEY=EMPTY export OPENAI_BASE_URL=http://0.0.0.0:23333/v1 export OPENAI_MODEL=internlm/Intern-S2-397B

关键前提:启动 LMDeploy 时必须携带--tool-call-parser interns2-preview参数,工具调用才能被正确解析(否则 Agent 无法识别模型输出的<tool_call>结构)。

连接 Claude Code:LMDeploy 暴露了 Anthropic 兼容的/v1/messages端点,Claude Code 可直接对接。将以下配置写入~/.claude/settings.json:

{ "env": { "ANTHROPIC_BASE_URL": "http://127.0.0.1:23333", "ANTHROPIC_AUTH_TOKEN": "dummy", "ANTHROPIC_MODEL": "internlm/Intern-S2-397B", "ANTHROPIC_CUSTOM_MODEL_OPTION": "internlm/Intern-S2-397B" } }

方式二:官方 Intern API

若不想自托管,可注册官方 Intern API 并创建 API token(形如sk-xxxxxxxx)。

连接 Agent 框架:官方服务是 OpenAI 兼容的,任何 Agent 框架均可直接使用。在 CLI 或配置文件中设置 base url 为https://chat.intern-ai.org.cn/api/v1、模型名为intern-s2-397b。连通性验证:

curl https://chat.intern-ai.org.cn/api/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer sk-xxxxxxxx" \ -d '{ "model": "intern-s2-397b", "messages": [ {"role": "user", "content": "Hello"} ], "temperature": 0.8, "top_p": 0.95 }'

连接 Claude Code:将ANTHROPIC_BASE_URL指向 Intern 的 Anthropic 兼容网关即可:

{ "env": { "ANTHROPIC_BASE_URL": "https://chat.intern-ai.org.cn", "ANTHROPIC_AUTH_TOKEN": "your-api-token", "ANTHROPIC_MODEL": "intern-s2-397b", "ANTHROPIC_SMALL_FAST_MODEL": "intern-s2-397b" } }

随后用以下命令启动 Claude Code:

claude --model intern-s2-397b

注意两种方式下模型名不同:自托管时为部署时指定的internlm/Intern-S2-397B,官方 API 时为intern-s2-397b。

部署配置速览(摘自部署指南)

部署指南 给出了三种框架的完整启动命令,官方建议在H100 x8 或 H200 x8节点上部署。这里提炼核心配置,详细参数以原文档为准。

LMDeploy(≥ 0.14.0)

基础服务(无 MTP):启动lmdeploy serve proxy作为代理,再启动api_server:

# proxy server lmdeploy serve proxy --server-name ${proxy_server_ip} --server-port ${proxy_server_port} # api_server lmdeploy serve api_server \ internlm/Intern-S2-FP8 \ --model-name internlm/Intern-S2-397B \ --trust-remote-code \ --backend pytorch \ --dp 4 \ --ep 8 \ --enable-prefix-caching \ --proxy-url http://${proxy_server_ip}:${proxy_server_port} \ --reasoning-parser default \ --tool-call-parser interns2-preview

关键参数说明:--dp 4(数据并行度)、--ep 8(专家并行度,与 MoE 架构 512 专家相适配)、--reasoning-parser default(解析<think>推理块)、--tool-call-parser interns2-preview(解析工具调用,Agent 集成必需)。

MTP 投机解码(加速推理):在基础命令上追加--speculative-algorithm qwen3_5_mtp --speculative-num-draft-tokens 4 --max-batch-size 256。该能力与 config.json 中mtp_num_hidden_layers: 1及权重文件中的mtp.*权重对应。

长上下文推理(512k 示例):同时配置--session-len 512000与 YaRN RoPE 参数(通过--hf-overrides覆盖文本配置):

lmdeploy serve api_server \ internlm/Intern-S2-FP8 \ --model-name internlm/Intern-S2-397B \ --trust-remote-code \ --backend pytorch \ --dp 4 \ --ep 8 \ --enable-prefix-caching \ --reasoning-parser default \ --tool-call-parser interns2-preview \ --session-len 512000 \ --max-batch-size 64 \ --hf-overrides '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}'

这里rope_type: "yarn"与factor: 4.0将原生 262144(256K)的最大位置编码扩展为约 512K~1M 上下文;mrope_section: [11, 11, 10]为多模态旋转位置编码的分段配置,与 config.json 中的默认 RoPE 设置一致。

vLLM(≥ v0.22.1)

基础服务:

export VLLM_DEEP_GEMM_WARMUP=skip export VLLM_USE_DEEP_GEMM=0 export VLLM_FLASHINFER_MOE_BACKEND=latency vllm serve internlm/Intern-S2-FP8 \ --served-model-name internlm/Intern-S2-397B \ --trust-remote-code \ --tensor-parallel-size 8 \ --enable-auto-tool-choice \ --tool-call-parser qwen3_coder \ --reasoning-parser qwen3 \ --mm-encoder-tp-mode data

MTP 版本追加--speculative-config '{"method":"mtp","num_speculative_tokens":3}';长上下文版本需设置环境变量VLLM_ALLOW_LONG_MAX_MODEL_LEN=1并指定--max-model-len 1010000(约 1M),同时通过--hf-overrides传入与 LMDeploy 相同的 YaRN RoPE 覆盖。

SGLang(≥ v0.5.13)

基础服务:

python3 -m sglang.launch_server \ --model-path internlm/Intern-S2-FP8 \ --served-model-name internlm/Intern-S2-397B \ --trust-remote-code \ --tp-size 8 \ --mem-fraction-static 0.8 \ --enable-flashinfer-allreduce-fusion \ --reasoning-parser qwen3 \ --tool-call-parser qwen3_coder

MTP 版本需设置SGLANG_ENABLE_SPEC_V2=1并追加--speculative-algo 'NEXTN' --speculative-eagle-topk 1 --speculative-num-steps 3 --speculative-num-draft-tokens 4以及--mamba-scheduler-strategy extra_buffer。

三个框架统一以internlm/Intern-S2-FP8作为权重源、--served-model-name internlm/Intern-S2-397B作为对外模型 ID——这与 README 中所有客户端示例的model_name = "internlm/Intern-S2-397B"保持一致。

科学模态与分词器实现细节

作为科学智能模型,Intern-S2-397B 的分词器内建了对三类科学序列的专门处理。在 tokenizer_config.json 中可以看到相应特殊 token:<SMILES>/</SMILES>(化学分子式)、<protein>/</protein>、<dna>/</dna>、<rna>/</rna>,以及三个子词表的偏移量offset_SMILES: 248102、offset_PROT: 249126、offset_XNA: 250150,对应仓库内的 tokenizer_SMILES.model、tokenizer_PROT.model、tokenizer_XNA.model 三个 SentencePiece 模型文件。

对应的实现位于 tokenization_interns1.py:SmilesCheckModule、ProtCheckModule、XnaCheckModule三个自动检测模块会在文本中按正则与合法性规则识别候选序列并自动包裹标记。例如 SMILES 检测在安装了 RDKit 时调用Chem.MolFromSmiles做化学合法性校验(check_legitimacy_slow),未安装时退化为基于键类型、括号配平、环闭合等硬规则校验(check_legitimacy_fast);蛋白序列检测要求至少 27 个连续大写字母且排除纯核酸串;XNA 检测匹配连续 27 个以上的ATCGU字符。这种"自动检测 + 专用子词表"的机制,保证了模型在科学文本中对分子结构、蛋白序列等符号能以高保真方式分词与生成,是科学模态推理能力的基础设施。

结语

至此,你已经掌握 Intern-S2-397B 从部署到集成的完整使用链路:用推荐采样参数启动服务,通过 OpenAI 兼容接口完成工具调用与思考模式控制,用 LMDeploy 处理时间序列检测与预测,再把模型接入自托管 Agent 框架或官方 Intern API(含 Claude Code)。结合 config.json(混合注意力 MoE 架构)、chat_template.jinja(工具与思考协议)、tokenization_interns1.py(科学序列分词)与 deployment_guide.md(三框架部署)等仓库文件,即可在真实业务中复现并扩展官方示例。

  • 人工智能
  • 大模型
  • 基础模型
  • 多模态
  • 预训练
  • 强化学习

【免费下载链接】Intern-S2-397B

项目地址:https://ai.gitcode.com/InternLM/Intern-S2-397B
点击查看免费下载

相关推荐

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

版权声明: 本文来自互联网用户投稿,该文观点仅代表作者本人,不代表本站立场。本站仅提供信息存储空间服务,不拥有所有权,不承担相关法律责任。如若内容造成侵权/违法违规/事实不符,请联系邮箱:809451989@qq.com进行投诉反馈,一经查实,立即删除!
网站建设 2026/9/29 11:37:24

模拟IC设计全流程指南:从PDK到版图,告别玄学

做模拟IC设计这个行当&#xff0c;说起来有点尴尬——它明明是最古老的集成电路方向之一&#xff0c;从早期运放芯片诞生到现在已经好几十年&#xff0c;却老被当成“玄学”来看待。很多刚入行的同学拿着EDA工具打开PDK里的器件模型&#xff0c;对着原理图一调好几天&#xff0…

作者头像 李华
网站建设 2026/9/29 11:34:25

10 分钟跑通 UI-TARS Desktop:用一句自然语言操作你的电脑

10 分钟跑通 UI-TARS Desktop&#xff1a;用一句自然语言操作你的电脑 【免费下载链接】UI-TARS-desktop The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra 项目地址: https://gitcode.com/GitHub_Trending/ui/UI-TARS-deskto…

作者头像 李华
网站建设 2026/9/29 11:31:14

网络安全应急处置流程图:从静态PDF到可执行响应机制

简介&#xff1a;一份面向企业网络安全管理的应急处置流程图文档&#xff0c;适用于信息安全负责人、IT运维及应急响应人员&#xff0c;用于规范从预防、预警到事件分类、分级和处置的完整流程。文档依据国家相关标准编制&#xff0c;明确了由董事长任组长的信息安全领导小组与…

作者头像 李华
网站建设 2026/9/29 11:29:54

PDF合并免费工具推荐!电脑+小程序全覆盖,无水印超好用

日常办公、学习中&#xff0c;经常需要把多个PDF文件合并成一个&#xff0c;比如整理毕业论文、工作报表、合同资料、证件扫描件等。但很多工具要么收费开会员&#xff0c;要么合并后自带水印&#xff0c;要么限制文件大小和次数&#xff0c;非常影响使用体验。今天给大家整理一…

作者头像 李华
网站建设 2026/9/29 11:29:49

Ubuntu安装Edge浏览器完全指南:从apt源配置到卸载清理

1. 装 Edge 之前&#xff0c;先给 Ubuntu 做一次三分钟体检前几天把一台闲置笔记本装成了 Ubuntu 系统&#xff0c;第一个想装回的应用就是 Edge 浏览器。原因很简单&#xff1a;公司电脑和主力机都是 Windows&#xff0c;书签、密码、收藏夹全躺在微软账号里&#xff0c;换到 …

作者头像 李华
网站建设 2026/9/29 11:27:41

CNN与RNN选型、原理与实战:图像分类到序列预测

/* MD / 富文本中的 .toc(含博客园搬家等嵌套结构);.toc-box 在侧栏,不受影响 */#content_views .toc,/* 编辑器常在目录前后插入空 p(:empty 仍占 20px),一并去掉避免顶空隙 */#content_views.markdown_views > p:empty:has(+ .toc),#content_views.markdown_views …

作者头像 李华