【免费下载链接】context-hub
Amazon Comprehend 是 AWS 的 NLP(自然语言处理)服务,提供语言检测、情感分析、实体识别、关键短语提取、语法分析和 PII(个人身份信息)检测等能力。本文基于 Context Hub 仓库中维护的 AWS 官方文档(content/aws/docs/comprehend/javascript/DOC.md),完整讲解如何通过 AWS SDK for JavaScript v3 的@aws-sdk/client-comprehend包,在 Node.js 中完成实时文本分析(同步 API)与基于 S3 的异步批量任务。读完本文,你将掌握客户端初始化、凭据配置、client.send(new Command(input))调用模式、六大常用工作流以及异步任务的状态轮询方法,并能在实际项目中直接复制运行。
本文档面向
@aws-sdk/client-comprehend版本3.1007.0,由 Context Hub 维护(frontmatter 中source: maintainer)。在 Context Hub 中,这篇文档的条目 ID 为aws/comprehend,语言变体为javascript,可通过chub get aws/comprehend --lang js拉取(详见 CLI 参考 与 SKILL.md)。
安装
使用 npm 安装 Comprehend 客户端包:
npm install @aws-sdk/client-comprehend如果你的代码需要显式加载 AWS 命名配置文件(named profile),例如通过fromIni读取~/.aws/credentials中的特定 profile,还需要安装凭据辅助包:
npm install @aws-sdk/credential-providers该包提供fromIni、fromEnv、fromSSO、fromTemporaryCredentials(assume-role 流程)等显式凭据加载工具,是@aws-sdk/client-comprehend之外最常用的配套依赖。
前置条件
基本凭据与区域
在创建客户端之前,先通过环境变量设置 AWS 凭据与区域:
export AWS_REGION="us-east-1" export AWS_ACCESS_KEY_ID="..." export AWS_SECRET_ACCESS_KEY="..." export AWS_SESSION_TOKEN="..." # 可选,仅在使用临时凭据时需要如果本地使用共享 AWS 配置文件(~/.aws/config与~/.aws/credentials),AWS_PROFILE也能被 AWS SDK for JavaScript v3 的标准凭据链识别:
export AWS_PROFILE="my-dev-profile" export AWS_REGION="us-east-1"在 Node.js 环境中,默认的凭据提供链(credential provider chain)通常已足够:它会依次检查环境变量、共享配置文件、ECS 任务凭据、EC2 实例元数据以及 IAM Identity Center(SSO)。也就是说,如果你的 AWS 访问权限来自上述任一渠道,通常无需手动注入凭据。
异步任务所需的 IAM 角色与 S3 位置
调用异步检测任务(如StartSentimentDetectionJobCommand)时,Amazon Comprehend 需要代为读取 S3 输入并写入 S3 输出,因此你必须提供一个 Comprehend 可以代入(assume)的 IAM 角色,以及输入、输出桶位置:
export COMPREHEND_DATA_ACCESS_ROLE_ARN="arn:aws:iam::123456789012:role/ComprehendDataAccessRole" export COMPREHEND_INPUT_S3_URI="s3://my-input-bucket/comprehend/input/" export COMPREHEND_OUTPUT_S3_URI="s3://my-output-bucket/comprehend/output/"注意:DataAccessRoleArn与你的应用用来调用 AWS 的凭据是两回事——前者是服务侧角色(授予 Comprehend 访问你的 S3 数据的权限),后者是你自己的 SDK 凭据。
初始化客户端
最小化 Node.js 客户端
import { ComprehendClient } from "@aws-sdk/client-comprehend"; const client = new ComprehendClient({ region: process.env.AWS_REGION ?? "us-east-1", });显式凭据
当你不希望依赖凭据链,而是直接在代码中传入凭据时:
import { ComprehendClient } from "@aws-sdk/client-comprehend"; const client = new ComprehendClient({ region: "us-east-1", credentials: { accessKeyId: process.env.AWS_ACCESS_KEY_ID, secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY, sessionToken: process.env.AWS_SESSION_TOKEN, }, });通过fromIni加载命名 profile
import { fromIni } from "@aws-sdk/credential-providers"; import { ComprehendClient } from "@aws-sdk/client-comprehend"; const client = new ComprehendClient({ region: "us-east-1", credentials: fromIni({ profile: "my-dev-profile" }), });核心调用模式
AWS SDK v3 的客户端统一采用client.send(new Command(input))模式:每个 API 对应一个*Command类,通过send方法发送,返回 Promise。例如情感检测:
import { ComprehendClient, DetectSentimentCommand, } from "@aws-sdk/client-comprehend"; const client = new ComprehendClient({ region: process.env.AWS_REGION ?? "us-east-1", }); const response = await client.send( new DetectSentimentCommand({ Text: "The delivery was fast and the packaging was excellent.", LanguageCode: "en", }), ); console.log(response.Sentiment, response.SentimentScore);Sentiment返回POSITIVE、NEGATIVE、NEUTRAL或MIXED;SentimentScore是包含Positive、Negative、Neutral、Mixed四个 0~1 概率值的对象,四个值之和为 1,可用于阈值判断。
常见工作流
先检测语言,再分析情感
大多数 Comprehend 同步文本分析 API 都要求提供LanguageCode。如果语言未知,先用DetectDominantLanguageCommand检测,再把返回的语言代码传入后续请求:
import { ComprehendClient, DetectDominantLanguageCommand, DetectSentimentCommand, } from "@aws-sdk/client-comprehend"; const client = new ComprehendClient({ region: process.env.AWS_REGION ?? "us-east-1", }); const text = "The delivery was fast and the packaging was excellent."; const languageResult = await client.send( new DetectDominantLanguageCommand({ Text: text, }), ); const languageCode = languageResult.Languages?.[0]?.LanguageCode; if (!languageCode) { throw new Error("Comprehend did not return a dominant language"); } const sentimentResult = await client.send( new DetectSentimentCommand({ Text: text, LanguageCode: languageCode, }), ); console.log(sentimentResult.Sentiment); console.log(sentimentResult.SentimentScore);DetectDominantLanguage返回Languages数组,每项包含LanguageCode与Score(置信度),按得分降序排列,取第一项即为最可能的语言。
对多篇短文本批量执行情感分析(单一语言码)
当你已经知道请求中的每篇文档都使用同一种语言时,使用批量 API 减少请求次数:
import { BatchDetectSentimentCommand, ComprehendClient, } from "@aws-sdk/client-comprehend"; const client = new ComprehendClient({ region: process.env.AWS_REGION ?? "us-east-1", }); const response = await client.send( new BatchDetectSentimentCommand({ LanguageCode: "en", TextList: [ "This product solved the problem quickly.", "Setup was confusing and took too long.", "Support answered within five minutes.", ], }), ); for (const result of response.ResultList ?? []) { console.log(result.Index, result.Sentiment, result.SentimentScore); } for (const error of response.ErrorList ?? []) { console.error(error.Index, error.ErrorCode, error.ErrorMessage); }注意:批量 API 对整批请求只接受一个LanguageCode,且TextList中每篇文档有长度上限(单文档约 5,000 字符)。ResultList中的Index对应TextList中的原始下标;失败的条目会出现在ErrorList中(含Index、ErrorCode、ErrorMessage),因此结果与错误要按Index对应处理。
从文档中提取实体
import { ComprehendClient, DetectEntitiesCommand, } from "@aws-sdk/client-comprehend"; const client = new ComprehendClient({ region: process.env.AWS_REGION ?? "us-east-1", }); const response = await client.send( new DetectEntitiesCommand({ Text: "Jane Doe from Example Corp met the AWS team in Seattle on Tuesday.", LanguageCode: "en", }), ); for (const entity of response.Entities ?? []) { console.log(entity.Text, entity.Type, entity.Score); }Entities中每项包含Text(实体原文)、Type(如PERSON、ORGANIZATION、LOCATION、DATE、QUANTITY等)、Score及起止偏移量。同样的调用模式适用于其他同步文本 API,例如DetectKeyPhrasesCommand(关键短语)与DetectSyntaxCommand(语法标注,返回SyntaxTokens及其PartOfSpeech),当它们更契合你的应用场景时直接替换即可。
先判断是否含 PII,再按需获取精确偏移
ContainsPiiEntitiesCommand只回答"文本中是否包含 PII、包含哪些类型的 PII 标签",适合作为前置门控(gate);DetectPiiEntitiesCommand才返回实体的精确起止偏移,适合脱敏场景:
import { ComprehendClient, ContainsPiiEntitiesCommand, DetectPiiEntitiesCommand, } from "@aws-sdk/client-comprehend"; const client = new ComprehendClient({ region: process.env.AWS_REGION ?? "us-east-1", }); const text = "Contact me at jane@example.com or 206-555-0100."; const contains = await client.send( new ContainsPiiEntitiesCommand({ Text: text, LanguageCode: "en", }), ); console.log(contains.Labels); const detailed = await client.send( new DetectPiiEntitiesCommand({ Text: text, LanguageCode: "en", }), ); for (const entity of detailed.Entities ?? []) { console.log(entity.Type, entity.BeginOffset, entity.EndOffset, entity.Score); }ContainsPiiEntities的Labels是形如{ Name: "EMAIL", Score: 0.99 }的标签列表;DetectPiiEntities的Entities中每项包含Type(如EMAIL、PHONE、NAME、CREDIT_DEBIT_NUMBER)、BeginOffset、EndOffset(字符偏移,可直接用于切片脱敏)与Score。注意:DetectPiiEntities本身不返回实体原文(Text为空),只提供偏移与类型,脱敏时需自行根据偏移截取原文本。
启动并轮询异步情感检测任务
当输入数据已经存放在 S3 中,或数据集规模超出实时文本 API 的限制(同步 API 单次Text约 5,000 字符)时,改用异步任务 API:
import { ComprehendClient, DescribeSentimentDetectionJobCommand, StartSentimentDetectionJobCommand, } from "@aws-sdk/client-comprehend"; const client = new ComprehendClient({ region: process.env.AWS_REGION ?? "us-east-1", }); const roleArn = process.env.COMPREHEND_DATA_ACCESS_ROLE_ARN; const inputS3Uri = process.env.COMPREHEND_INPUT_S3_URI; const outputS3Uri = process.env.COMPREHEND_OUTPUT_S3_URI; if (!roleArn || !inputS3Uri || !outputS3Uri) { throw new Error("Set COMPREHEND_DATA_ACCESS_ROLE_ARN, COMPREHEND_INPUT_S3_URI, and COMPREHEND_OUTPUT_S3_URI"); } const start = await client.send( new StartSentimentDetectionJobCommand({ JobName: "support-ticket-sentiment", LanguageCode: "en", DataAccessRoleArn: roleArn, InputDataConfig: { S3Uri: inputS3Uri, InputFormat: "ONE_DOC_PER_LINE", }, OutputDataConfig: { S3Uri: outputS3Uri, }, }), ); const jobId = start.JobId; if (!jobId) { throw new Error("Comprehend did not return a JobId"); } for (;;) { const detail = await client.send( new DescribeSentimentDetectionJobCommand({ JobId: jobId, }), ); const properties = detail.SentimentDetectionJobProperties; const status = properties?.JobStatus; console.log(status); if (status === "COMPLETED") { console.log(properties?.OutputDataConfig?.S3Uri); break; } if (status === "FAILED" || status === "STOPPED") { throw new Error(`Sentiment job ended with status ${status}`); } await new Promise((resolve) => setTimeout(resolve, 10000)); }要点说明:
InputDataConfig.InputFormat支持ONE_DOC_PER_LINE(每行一篇文档,推荐用于大批量文本)与ONE_DOC_PER_FILE(每个文件一篇文档)。- 输出结果会写入
OutputDataConfig.S3Uri指向的位置,DescribeSentimentDetectionJob返回的SentimentDetectionJobProperties.OutputDataConfig.S3Uri可能包含任务 ID 子路径(如s3://bucket/output/1234567890abcdef/)。 - 轮询间隔为 10 秒;实际生产代码可改用指数退避(如 10s → 20s → 40s…),并设置最大重试次数。
- 任务终态为
COMPLETED、FAILED、STOPPED(及STOP_REQUESTED过渡态),SUBMITTED、IN_PROGRESS属于进行中状态。
Comprehend 的其他异步任务 API 遵循完全相同的模式:Start*JobCommand启动 S3 任务,Describe*JobCommand轮询状态。可替换的对应关系包括StartEntitiesDetectionJobCommand/DescribeEntitiesDetectionJobCommand、StartKeyPhrasesDetectionJobCommand/DescribeKeyPhrasesDetectionJobCommand、StartDominantLanguageDetectionJobCommand/DescribeDominantLanguageDetectionJobCommand、StartPiiEntitiesDetectionJobCommand/DescribePiiEntitiesDetectionJobCommand、StartTopicsDetectionJobCommand/DescribeTopicsDetectionJobCommand等。
重要注意事项(Gotchas)
- 大多数同步文本 API 强制要求
LanguageCode:如果你的应用事先不知道语言,务必先用DetectDominantLanguageCommand检测,再将结果代码传给后续请求。 - 批量 API 只接受单个
LanguageCode:批量请求内所有文档必须同语言,发送前先按语言对文本分组。 - 同步 API 接收原始
Text,异步 API 读写 S3:前者请求体中直接携带文本;后者输入输出均通过 S3 URI 传递,且需要DataAccessRoleArn。 - 异步任务必须提供
DataAccessRoleArn:这是服务侧 IAM 角色,与你的 SDK 凭据相互独立,缺少时任务会以FAILED结束或启动即报权限错误。 ContainsPiiEntitiesCommand与DetectPiiEntitiesCommand职责不同:前者只返回存在哪些 PII 标签类型(Labels),后者才返回精确偏移(Entities中的BeginOffset/EndOffset),脱敏场景必须用后者。StartSentimentDetectionJobCommand只负责启动任务:它不等待任务完成,必须配合DescribeSentimentDetectionJobCommand轮询到终态并读取输出 S3 路径。- Comprehend 是区域化服务(regional):客户端
region与数据、IAM 角色必须同区域;区域不匹配时错误往往表现为权限问题或资源不存在,排查时先核对区域。
何时使用其他配套包
@aws-sdk/credential-providers:需要显式加载凭据时使用,包括fromIni、assume-role(fromTemporaryCredentials)以及其他基于 profile 的配置方式。- 其他 AWS SDK v3 服务客户端:当你的 Comprehend 工作流依赖 S3(上传/下载输入输出)、IAM(角色管理)或周边基础设施自动化时,组合使用对应的 v3 客户端(如
@aws-sdk/client-s3、@aws-sdk/client-iam)。
版本说明
- 本指南针对
@aws-sdk/client-comprehend版本3.1007.0(见文档 frontmatter 的versions字段)。 - 当前包面使用标准的 AWS SDK v3 命令模式:
client.send(new Command(input)),所有示例均基于该模式编写。 - 若你使用的是 Python 侧的 Comprehend 类型桩,可参考仓库中对应的 mypy-boto3-comprehend 指南;两篇文档由 Context Hub 分别维护为
javascript与python语言变体。
在 Context Hub 中获取与使用本文档
本文档是 Context Hub 仓库中面向 LLM/Agent 优化的文档条目之一。编码 Agent 在写 Comprehend 相关代码前,可通过chubCLI 拉取最新版本(而非依赖训练数据中可能过时的 API 记忆):
npm install -g @aisuite/chub chub search aws/comprehend --lang js # 检索条目 chub get aws/comprehend --lang js # 拉取本文档(JavaScript 变体)完整的命令说明见 CLI Reference 与 get-api-docs 技能;仓库总览见 README。该文档本身以 YAML frontmatter 标记元信息(条目名comprehend、语言javascript、版本3.1007.0、来源maintainer),与仓库的 内容规范 保持一致。
【免费下载链接】context-hub
相关推荐
使用 AWS SDK for Java 2.x 调用 Amazon Comprehend:文本分析实战指南
使用 AWS SDK for Java 2.x 调用 Amazon Comprehend:文本分析实战指南 Amazon Comprehend 是基于自然语言处
示例工程教程后端使用 AWS SDK for .NET 调用 Amazon Comprehend 自然语言处理 API 完整指南
使用 AWS SDK for .NET 调用 Amazon Comprehend 自然语言处理 API 完整指南 导读 本文以 dotnetv3/Compreh
示例工程教程后端使用 AWS SDK for Kotlin 调用 Amazon Comprehend:六个 NLP 检测与文档分类实战示例
使用 AWS SDK for Kotlin 调用 Amazon Comprehend:六个 NLP 检测与文档分类实战示例 导读 本文以 kotlin/serv
示例工程教程后端
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考