- 人工智能
- 深度学习
- 推理引擎
- 本地部署
- 嵌入式
- 物联网
【免费下载链接】tflite-micro
Infrastructure to enable deployment of ML models to low-power resource-constrained embedded targets (including microcontrollers and digital signal processors).
导读
本文围绕 third_party/xtensa/examples/pytorch_to_tflite/pytorch_to_tflite_converter/README.md 展开,系统讲解如何把 PyTorch 生态中的 MobileNetV2 分类模型转换为 TFLite 格式的int8 全整数量化模型,以便最终部署到 tflite-micro 支持的低功耗嵌入式目标(如 Cadence Xtensa HiFi5 DSP)上。文章提供两条可复现的转换路径:TinyNN 一键直转(约 3 分钟)与PyTorch → ONNX → TFLite 标准链路(约 5 分钟),并完整覆盖环境搭建、量化校准、参数调优、推理验证,以及转换产物在仓库内的 Xtensa 平台测试验证,读完即可在 Google Colaboratory 上照做。
背景:为什么需要 PyTorch → TFLite(int8) 转换
tflite-micro 的目标是让机器学习模型在微控制器与数字信号处理器等低功耗、资源受限的嵌入式目标上运行。这类设备通常没有文件系统与动态内存管理,对模型体积、算子支持与数值格式都有严格约束。PyTorch 训练出的模型无法直接被 tflite-micro 加载,必须经过转换:
- 格式转换:PyTorch 模型 → TFLite FlatBuffer 格式(本仓库 tensorflow/lite/schema 目录中的
schema_generated.h即该格式的 C++ 绑定); - 量化:把 float32 权重与激活转换为 int8,从而显著缩小模型体积并利用嵌入式 DSP 的整数指令加速,这也是 MobileNetV2 这类图像模型能跑进嵌入式设备的前提。
仓库在 third_party/xtensa/examples/pytorch_to_tflite 目录中给出了完整示例:目录内已包含转换产物 mobilenet_v2_quantized_1x3x224x224.tflite、对应的 C 语言测试 pytorch_to_tflite_test.cc 以及两个转换 Notebook。其中 tinynn_pytorch_to_tflite_int8.ipynb 与 pytorch_to_onnx_to_tflite_int8.ipynb 分别对应 README 中介绍的两条转换路径。
两条路径的差异对比如下:
| 对比项 | 路径一:TinyNN 直接转换 | 路径二:PyTorch → ONNX → TFLite |
|---|---|---|
| 转换链路 | PyTorch → (TinyNN QAT) → TFLite | PyTorch → ONNX → TF SavedModel → TFLite |
| 依赖工具 | TinyNeuralNetwork、torch、tensorflow | onnx、onnxruntime、onnx-tf、tensorflow |
| 量化方式 | PostQuantizer 后训练量化 + QAT 校准 | 代表数据集(representative dataset)驱动全整数量化 |
| 预估耗时(README 标注) | ~3 分钟 | ~5 分钟 |
| 关键产物 | out/qat_model.tflite(int8) | mobilenet_v2_float32.tflite+mobilenet_v2_1.0_224_quant.tflite(int8) |
路径一:TinyNN 一键直转 PyTorch → TFLite(int8)
该路径在 Google Colaboratory 中完成,核心思路是:先用 TinyNN 的PostQuantizer对预训练 MobileNetV2 做后训练量化并得到带QuantStub/DeQuantStub的 QAT 模型,再借少量真实图片做校准前向,最后调用torch.quantization.convert与 TinyNN 的TFLiteConverter直接产出 int8 的.tflite文件。
1. 环境准备:安装 TinyNeuralNetwork
第一个 Notebook 通过pip从源码安装 TinyNeuralNetwork,安装日志显示它会一并拉取ruamel.yaml、python-igraph、tflite==2.3.0、PyYAML、flatbuffers、texttable等依赖:
!pip install git+https://github.com/alibaba/TinyNeuralNetwork.git随后导入转换所需的全部模块:
import random from glob import glob from PIL import Image import torch from torchvision import transforms import torchvision.models as models from tinynn.converter import TFLiteConverter from tinynn.graph.quantization.quantizer import PostQuantizer from tinynn.graph.tracer import model_tracer from tinynn.util.cifar10 import get_dataloader, train_one_epoch, validate from tinynn.util.train_util import DLContext, get_device, train2. 下载校准数据集
量化需要真实数据来统计激活值的动态范围。Notebook 下载的是 TensorFlow 官方教学使用的猫狗二分类数据集cats_and_dogs_filtered.zip(约 65 MB):
!wget --no-check-certificate \ https://storage.googleapis.com/mledu-datasets/cats_and_dogs_filtered.zip \ -O /content/cats_and_dogs_filtered.zipimport os import zipfile local_zip = '/content/cats_and_dogs_filtered.zip' zip_ref = zipfile.ZipFile(local_zip, 'r') zip_ref.extractall('/content') zip_ref.close()3. 模型追踪与 QAT 量化准备
TinyNN 通过model_tracer()上下文追踪模型结构,加载预训练的 MobileNetV2 并构造dummy_input(形状(1, 3, 224, 224),与仓库内转换产物文件名mobilenet_v2_quantized_1x3x224x224一致):
random.seed(0) with model_tracer(): model = models.mobilenet_v2(pretrained=True) model.eval() # Provide a viable input for the model dummy_input = torch.rand((1, 3, 224, 224)) quantizer = PostQuantizer(model, dummy_input, work_dir='out', config={'asymmetric': True, 'per_tensor': False}) qat_model = quantizer.quantize() print(qat_model)PostQuantizer的关键配置说明:
| 配置项 | 示例值 | 含义 |
|---|---|---|
work_dir | 'out' | 量化中间产物(onnx/qat 模型)的输出目录 |
config['asymmetric'] | True | 使用非对称量化(带 zero point),更贴合常见 int8 部署格式 |
config['per_tensor'] | False | 逐通道(per-channel)量化权重,精度损失更小 |
量化后的模型打印结果可以清楚看到 TinyNN 在每个卷积层后插入了HistogramObserver(用于校准期间统计激活值分布),并在输入/输出端插入了QuantStub/DeQuantStub,这正是后面 QAT 微调与转换的基础。
4. 校准前向:用真实图片驱动量化范围估计
将 QAT 模型搬到 GPU 上,从猫狗训练集中随机取 100 张图片,按 ImageNet 的标准预处理(Resize(256)→CenterCrop(224)→ToTensor→Normalize)后逐张前向,让各层的HistogramObserver收集激活分布:
if torch.cuda.device_count() > 1: qat_model = nn.DataParallel(qat_model) device = get_device() qat_model.to(device=device) dataset_list = glob('/content/cats_and_dogs_filtered/train/**/*', recursive=True) random.shuffle(dataset_list) for i in range(100): filename = dataset_list[i] input_image = Image.open(filename) preprocess = transforms.Compose([ transforms.Resize(256), transforms.CenterCrop(224), transforms.ToTensor(), transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]), ]) input_tensor = preprocess(input_image) input_tensor = torch.unsqueeze(input_tensor, 0) qat_model(input_tensor.to(device=device))5. 转换为真正量化的 TFLite 模型
校准完成后,在torch.no_grad()下执行torch.quantization.convert把带伪量化节点的 QAT 模型固化为使用量化内核的真实量化模型,再交给 TinyNN 的TFLiteConverter输出 int8 TFLite:
with torch.no_grad(): qat_model.eval() qat_model.cpu() # The step below converts the model to an actual quantized model, which uses the quantized kernels. qat_model = torch.quantization.convert(qat_model) # When converting quantized models, please ensure the quantization backend is set. torch.backends.quantized.engine = quantizer.backend # The code section below is used to convert the model to the TFLite format converter = TFLiteConverter(qat_model, dummy_input, tflite_path='out/qat_model.tflite', quantize_target_type='int8', input_transpose=False, fuse_quant_dequant=True) converter.convert()Notebook 注释中给出的TFLiteConverter参数说明:
quantize_target_type='int8':指定输出量化目标类型,如需其他数据类型可相应调整;strict_symmetric_check=True(可选):当需要对预定义 zero point 做严格对称量化检查时启用;input_transpose=False:控制输入张量是否需要转置(PyTorch 的NCHW与 TFLite 布局的差异处理);fuse_quant_dequant=True:融合相邻的 Quantize/Dequantize 节点,减少图冗余。
运行成功时控制台会打印INFO (tinynn.converter.base) Generated model saved to out/qat_model.tflite。
6. 在 TensorFlow Lite 中验证转换结果
最后一步用tf.lite.Interpreter加载生成的 int8 模型并跑真实图片。关键点在于:读取输入张量的scale与zero_point,用 PyTorch 的torch.quantize_per_tensor把预处理后的 float 图像量化成 qint8,再经int_repr取整数表示喂给解释器:
import urllib url, filename = ("https://github.com/pytorch/hub/raw/master/images/dog.jpg", "dog.jpg") try: urllib.URLopener().retrieve(url, filename) except: urllib.request.urlretrieve(url, filename) import tensorflow as tf import numpy as np tflite_model_path = '/content/out/qat_model.tflite' interpreter = tf.lite.Interpreter(model_path=tflite_model_path) interpreter.allocate_tensors() input_details = interpreter.get_input_details() output_details = interpreter.get_output_details() test_details = interpreter.get_input_details()[0] scale, zero_point = test_details['quantization'] print(scale) print(zero_point) from PIL import Image from torchvision import transforms input_image = Image.open(filename) preprocess = transforms.Compose([ transforms.Resize(256), transforms.CenterCrop(224), transforms.ToTensor(), transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]), ]) input_tensor = preprocess(input_image) input_tensor = torch.unsqueeze(input_tensor, 0) input_tensor = torch.quantize_per_tensor(input_tensor, torch.tensor(scale), torch.tensor(zero_point), torch.qint8) input_tensor = torch.int_repr(input_tensor).numpy() interpreter.set_tensor(input_details[0]['index'], input_tensor) interpreter.invoke() output_data = interpreter.get_tensor(output_details[0]['index']) print("Predicted value . Label index: {}, confidence: {:2.0f}%" .format(np.argmax(output_data), 100 * output_data[0][np.argmax(output_data)]))Notebook 记录的实测输出中,该模型的输入量化参数为scale=0.01871182955801487、zero_point=-15,推理得到Label index: 258(对应 ImageNet 中的金毛寻回犬类别),与路径二的量化模型结果一致。
路径二:PyTorch → ONNX → TFLite(int8)
当希望走业界通用的 ONNX 中间格式(便于后续对接 ONNX Runtime 等工具链)时,使用第二个 Notebook 的五步链路:导出 ONNX → 校验 → onnx-tf 转 SavedModel → 转 float32 TFLite → 代表数据集全整数量化。
1. 安装 ONNX 与 ONNX Runtime
!pip install onnx !pip install onnxruntime import numpy as np import torch import torch.onnx import torchvision.models as models import onnx import onnxruntime2. 从 torchvision 加载 MobileNetV2 并导出 ONNX
加载预训练模型并进入推理模式(日志显示权重文件mobilenet_v2-b0353104.pth约 13.6 MB,首次运行会自动下载):
model = models.mobilenet_v2(pretrained=True) model.eval()导出 ONNX 时,torch.onnx.export的参数直接决定了产物形态:
IMAGE_SIZE = 224 BATCH_SIZE = 1 x = torch.randn(BATCH_SIZE, 3, 224, 224, requires_grad=True) torch_out = model(x) torch.onnx.export(model, # model being run x, # model input (or a tuple for multiple inputs) "mobilenet_v2.onnx", # where to save the model export_params=True, # store the trained parameter weights inside the model file opset_version=10, # the ONNX version to export the model to do_constant_folding=True, # whether to execute constant folding for optimization input_names = ['input'], # the model's input names output_names = ['output'], # the model's output names dynamic_axes={'input' : {0 : 'BATCH_SIZE'}, 'output' : {0 : 'BATCH_SIZE'}})各参数作用:
| 参数 | 示例值 | 说明 |
|---|---|---|
export_params | True | 把训练好的权重内嵌进 ONNX 文件 |
opset_version | 10 | 目标 ONNX 算子集版本,需与下游转换工具兼容 |
do_constant_folding | True | 开启常量折叠,优化导出图 |
input_names/output_names | ['input']/['output'] | 指定输入/输出张量名,便于下游引用 |
dynamic_axes | {0: 'BATCH_SIZE'} | 声明 batch 维为动态轴 |
3. 校验 ONNX 并与 PyTorch 结果对齐
先用 ONNX 自带 checker 校验图结构合法性,再用 ONNX Runtime 跑一遍同样的输入,与 PyTorch 输出做数值对齐(rtol=1e-03, atol=1e-05):
onnx_model = onnx.load("mobilenet_v2.onnx") onnx.checker.check_model(onnx_model)ort_session = onnxruntime.InferenceSession("mobilenet_v2.onnx") def to_numpy(tensor): return tensor.detach().cpu().numpy() if tensor.requires_grad else tensor.cpu().numpy() ort_inputs = {ort_session.get_inputs()[0].name: to_numpy(x)} ort_outs = ort_session.run(None, ort_inputs) np.testing.assert_allclose(to_numpy(torch_out), ort_outs[0], rtol=1e-03, atol=1e-05) print("Exported model has been tested with ONNXRuntime, and the result looks good!")实测输出为Exported model has been tested with ONNXRuntime, and the result looks good!,确认导出无误。
4. ONNX → TF SavedModel → float32 TFLite
安装onnx-tf,用其prepare后端把 ONNX 图导出为 TensorFlow SavedModel 目录model_tf:
!pip install onnx-tf from onnx_tf.backend import prepare import onnx onnx_model_path = 'mobilenet_v2.onnx' tf_model_path = 'model_tf' onnx_model = onnx.load(onnx_model_path) tf_rep = prepare(onnx_model) tf_rep.export_graph(tf_model_path)随后用 TensorFlow 原生转换器得到 float32 版本的 TFLite:
import tensorflow as tf saved_model_dir = 'model_tf' tflite_model_path = 'mobilenet_v2_float32.tflite' converter = tf.lite.TFLiteConverter.from_saved_model(saved_model_dir) tflite_model = converter.convert() with open(tflite_model_path, 'wb') as f: f.write(tflite_model)5. float32 模型的推理验证
先用随机数据快速验证输入输出形状与基本推理通路,Notebook 输出显示输入形状为[1 3 224 224]:
import numpy as np import tensorflow as tf tflite_model_path = '/content/mobilenet_v2_float32.tflite' interpreter = tf.lite.Interpreter(model_path=tflite_model_path) interpreter.allocate_tensors() input_details = interpreter.get_input_details() output_details = interpreter.get_output_details() input_shape = input_details[0]['shape'] input_data = np.array(np.random.random_sample(input_shape), dtype=np.float32) interpreter.set_tensor(input_details[0]['index'], input_data) interpreter.invoke() output_data = interpreter.get_tensor(output_details[0]['index'])再用真实猫图验证(TensorFlow 侧读图、解码、resize 到 224 并 reshape 为[1, 3, 224, 224]),DUMP INPUT/DUMP OUTPUT表明输入名serving_default_input:0、输出名PartitionedCall:0,均为float32且quantization: (0.0, 0)(未量化):
image = tf.io.read_file('/content/cats_and_dogs_filtered/validation/cats/cat.2000.jpg') image = tf.io.decode_jpeg(image, channels=3) image = tf.image.resize(image, [IMAGE_SIZE, IMAGE_SIZE]) image = tf.reshape(image,[3,IMAGE_SIZE,IMAGE_SIZE]) image = tf.expand_dims(image, 0) interpreter.set_tensor(input_details[0]['index'], image) interpreter.invoke() output_data = interpreter.get_tensor(output_details[0]['index'])6. 代表数据集驱动 int8 全整数量化
float32 模型无法直接用于嵌入式部署,需要再走一次代表数据集量化。核心是提供representative_data_gen_1生成器:从猫狗训练集取 100 张图,做带数据增强的预处理(RandomCrop(224, padding=4)、Resize(224)、RandomHorizontalFlip、ToTensor、Normalize)后逐张yield:
converter = tf.lite.TFLiteConverter.from_saved_model(saved_model_dir) # This enables quantization converter.optimizations = [tf.lite.Optimize.DEFAULT] # This sets the representative dataset for quantization converter.representative_dataset = representative_data_gen_1 # This ensures that if any ops can't be quantized, the converter throws an error converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8] # For full integer quantization, though supported types defaults to int8 only, we explicitly declare it for clarity. converter.target_spec.supported_types = [tf.int8] # These set the input and output tensors to uint8 (added in r2.3) converter.inference_input_type = tf.int8 converter.inference_output_type = tf.int8 tflite_model = converter.convert() with open('mobilenet_v2_1.0_224_quant.tflite', 'wb') as f: f.write(tflite_model)这段代码的每个配置项都是 int8 全整数量化的关键:
| 配置项 | 值 | 作用 |
|---|---|---|
optimizations | [tf.lite.Optimize.DEFAULT] | 开启默认量化优化 |
representative_dataset | 生成器函数 | 提供校准数据,统计激活动态范围 |
target_spec.supported_ops | [TFLITE_BUILTINS_INT8] | 只允许 int8 内建算子,遇到无法量化的算子直接报错 |
target_spec.supported_types | [tf.int8] | 显式声明仅支持 int8 类型 |
inference_input_type/inference_output_type | tf.int8 | 输入输出张量也强制 int8,实现"全整数量化" |
7. int8 量化模型的推理验证
对量化模型跑 dog.jpg 时,读取到的输入量化参数为scale=0.020324693992733955、zero_point=-8,喂入的 int8 输入张量取值区间约在[-109, -27],推理输出同样为Label index: 258,说明量化前后预测类别保持一致:
import tensorflow as tf import numpy as np tflite_model_path = '/content/mobilenet_v2_1.0_224_quant.tflite' interpreter = tf.lite.Interpreter(model_path=tflite_model_path) interpreter.allocate_tensors() input_details = interpreter.get_input_details() output_details = interpreter.get_output_details() test_details = interpreter.get_input_details()[0] scale, zero_point = test_details['quantization'] print(scale) print(zero_point) # ... 与路径一相同的图像预处理与量化(quantize_per_tensor → int_repr)... interpreter.set_tensor(input_details[0]['index'], input_tensor) interpreter.invoke() output_data = interpreter.get_tensor(output_details[0]['index'])MobileNetV2(int8 量化)模型架构
README 指出示例模型即 MobileNetV2,并给出了量化模型的架构图。从 TinyNN 路径打印的量化模型结构可以看到,int8 量化后的 MobileNetV2 不再是一个普通的浮点骨干网,而是由 18 个带HistogramObserver的量化倒残差块(features_0~features_17)串成的主干、末尾Conv2d(320, 1280)+ReLU6、Dropout+Linear(1280, 1000)分类头,以及输入QuantStub、输出DeQuantStub和 10 个FloatFunctional加法节点(对应倒残差块的残差连接)共同组成的量化算子链:
(架构图位于 third_party/xtensa/examples/pytorch_to_tflite/images/qat_model.png,其中每个卷积层都插入了用于统计激活分布的HistogramObserver。)
在 tflite-micro 上验证转换产物(XTensa HiFi5)
转换只是第一步,最终要验证的是 TFLite 模型能被 tflite-micro 正确加载并推理。仓库已经把路径一/路径二产出的 int8 模型固化在 third_party/xtensa/examples/pytorch_to_tflite/mobilenet_v2_quantized_1x3x224x224.tflite,并由构建系统自动转换为 C 数组mobilenet_v2_quantized_1x3x224x224_model_data.h,配套的 pytorch_to_tflite_test.cc 演示了标准 TFLM 推理流程:
- 用
tflite::GetModel解析模型(HIFI5宏下编译,即仅针对 Xtensa HiFi5); - 使用 pytorch_op_resolver.h 中定义的
PytorchOpsResolver(即tflite::MicroMutableOpResolver<128>,最多注册 128 个算子)按需注册算子; - 分配 3 MB 的
tensor_arena,构建MicroInterpreter并AllocateTensors(); - 校验输入张量为
kTfLiteInt8、形状[1, 3, 224, 224],将内置的 pytorch_images_dog_jpg.h 中的 int8 狗图数据memcpy到输入缓冲; Invoke()后扫描 1000 维 int8 输出,打印Label index与Confidence。
构建接入由 Makefile.inc 完成:仅当OPTIMIZED_KERNEL_DIR=xtensa时启用该示例,TARGET_ARCH非 hifi5 时注册pytorch_to_tflite_test测试目标。同级 README.md 给出了完整的 Xtensa 工具链环境与构建命令:
# Setup Xtensa Tools $ set path = ( ~/xtensa/XtDevTools/install/tools/RI-2020.5-linux/XtensaTools/bin $path ) $ set path = ( ~/xtensa/XtDevTools/install/tools/RI-2020.5-linux/XtensaTools/Tools/bin $path ) $ setenv XTENSA_SYSTEM ~xtensa/XtDevTools/install/tools/RI-2020.5-linux/XtensaTools/config $ setenv XTENSA_CORE AE_HiFi5_LE5_AO_FP_XC $ setenv XTENSA_TOOLS_VERSION RI-2020.5-linux $ setenv XTENSA_BASE ~/xtensa/XtDevTools/install/ # Clean and build mobilenet_v2 model on TFLM $ make -f tensorflow/lite/micro/tools/make/Makefile clean $ make -f tensorflow/lite/micro/tools/make/Makefile TARGET=xtensa OPTIMIZED_KERNEL_DIR=xtensa TARGET=xtensa TARGET_ARCH=hifi5 test_pytorch_to_tflite_test -j其中XTENSA_CORE、XTENSA_TOOLS_VERSION等环境变量需要与本地安装的 Xtensa 工具链版本对应;该命令同时会使用 tensorflow/lite/micro/kernels/xtensa 下的 HiFi5 优化内核。从测试结构看,pytorch_to_tflite_test专用于验证"PyTorch 训练 → TFLite 量化转换 → TFLM 端到端推理"这条完整链路,标签258的输出也与两个转换 Notebook 的验证结果相互印证。
总结与注意事项
本文完整复现了 README 中介绍的两条 PyTorch → TFLite(int8) 转换路径:
- TinyNN 路径(~3 分钟):
model_tracer追踪 →PostQuantizer后训练量化 → 100 张图片校准 →torch.quantization.convert→ TinyNNTFLiteConverter直接产出 int8 模型,链路最短、依赖最少; - ONNX 路径(~5 分钟):
torch.onnx.export导出并校验 →onnx-tf转 SavedModel → float32 TFLite 验证 → 代表数据集 +TFLITE_BUILTINS_INT8全整数量化,链路标准、便于对接更多中间工具。
实操中需要注意以下几点:
- 量化必须配套校准数据:两条路径都依赖真实图片统计激活动态范围,校准集应与部署场景分布接近;
- 预处理与输入格式要对齐:ImageNet 风格预处理(
Resize(256)+CenterCrop(224)+Normalize)与输入量化(scale/zero_point→quantize_per_tensor→int_repr)必须与转换时一致,否则推理结果会明显退化; - 输出类型选择:嵌入式端(TFLM)通常偏好输入输出均为
int8的全整数量化模型,路径二通过inference_input_type/inference_output_type显式指定; - 算子兼容性:量化转换要求模型中所有算子都能被 TFLite int8 内建算子覆盖,
TFLITE_BUILTINS_INT8会在遇到不支持的算子时直接报错; - 部署验证闭环:转换完成后应像仓库 pytorch_to_tflite_test.cc 一样,在目标平台(如 Xtensa HiFi5)上跑真实输入并比对预测标签,确认端到端精度无损。
- 人工智能
- 深度学习
- 推理引擎
- 本地部署
- 嵌入式
- 物联网
【免费下载链接】tflite-micro
Infrastructure to enable deployment of ML models to low-power resource-constrained embedded targets (including microcontrollers and digital signal processors).
相关推荐
tflite-micro Xtensa 平台部署实战:PyTorch MobileNetV2 转 int8 TFLite 并在 HiFi5 DSP 上运行
tflite micro Xtensa 平台部署实战:PyTorch MobileNetV2 转 int8 TFLite 并在 HiFi5 DSP 上运行 本指
人工智能深度学习推理引擎本地部署嵌入式物联网OpenChatKit移动端部署:ONNX转换与TensorFlow Lite模型优化指南
OpenChatKit移动端部署:ONNX转换与TensorFlow Lite模型优化指南 引言:移动端大模型部署的挑战与解决方案 你是否还在为OpenChat
人工智能大模型NLP模型训练模型推理服务Qwen3-30B-A3B移动端部署:ONNX转换与TFLite模型优化指南
Qwen3 30B A3B移动端部署:ONNX转换与TFLite模型优化指南 引言:移动端大模型部署的痛点与解决方案 你是否还在为将300亿参数的Qwen3 3
大模型基础模型人工智能
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考