- 开发工具
【免费下载链接】Humanizer
Humanizer meets all your .NET needs for manipulating and displaying strings, enums, dates, times, timespans, numbers and quantities
导读
Humanizer 是 .NET 生态中处理字符串、枚举、日期、时间与数量的常用库,而Vocabularies类正是其全部英语单复数(Inflection)功能的"总入口"。本指南以版本 3.0.1 的 Humanizer.Vocabularies API 文档 为骨架,结合源码与测试,深入讲解Vocabularies.Default静态属性的定位、内置规则的构成,以及通过Vocabulary注册自定义复数、单数、不可数名词与缩写词(Acronym)的完整方法。读完本文,你将能精确理解Pluralize()/Singularize()/Humanize()背后的词库工作机制,并能在自己的项目中安全地扩展或覆盖默认规则。
一、Vocabularies 类:单一静态容器
在 Humanizer 中,所有屈折规则都集中在一个静态容器中。3.0.1 版文档将其定义为:
public static class Vocabularies从 src/Humanizer/Inflections/Vocabularies.cs 的源码可以看到,它通过Lazy<Vocabulary>延迟构建默认词库,并以只读静态属性暴露给整个进程:
static readonly Lazy<Vocabulary> Instance = new(BuildDefault, LazyThreadSafetyMode.PublicationOnly); public static Vocabulary Default => Instance.Value;关键设计点:
- 静态全局单例:
Default属性返回进程级的唯一Vocabulary实例,任何线程调用Pluralize()/Singularize()都会读取该实例中注册的规则。 - 延迟初始化:词库在第一次访问
Default时才构建,未访问过则不会产生构建开销;同时ApplyAcronyms/NormalizeAcronyms两个内部方法通过Instance.IsValueCreated判断,在词库未构建时直接返回原输入,避免无谓的锁竞争。 - 单一词库限制:文档明确说明"At this time, multiple vocabularies and removing existing rules are not supported"(当前不支持多词库,也不支持移除已有规则)。从源码可见
Vocabulary的构造函数是internal,外部只能通过Vocabularies.Default操作这一个实例;RemoveAcronym虽然是internal方法,仅用于测试中的清理场景。
二、Vocabularies.Default 属性:默认词库的定位
2.1 属性签名与职责
文档给出的属性签名为:
public static Humanizer.Vocabulary Default { get; }对应源码 src/Humanizer/Inflections/Vocabularies.cs。它的职责有三:
- 提供singular/plural 不规则形式(irregularities)的匹配规则;
- 支持自定义缩写词的大小写保护(custom acronym casing);
- 作为
Pluralize()、Singularize()与Humanize()三大扩展方法的底层规则来源。
例如 src/Humanizer/InflectorExtensions.cs 中,Pluralize()直接转发到默认词库:
public static string? Pluralize(this string? word, bool inputIsKnownToBeSingular = true) => Vocabularies.Default.Pluralize(word, inputIsKnownToBeSingular);这意味着你向Vocabularies.Default添加的任何规则,会立即对所有调用方生效——这正是它能"进程级自定义"的原因。
2.2 内置词库的组成
BuildDefault()方法(Vocabularies.cs)为默认词库预置了三类规则,构建完成后通过MarkRulesAsBuiltIn()标记:
① 复数化规则(AddPlural,约 20 条正则)覆盖常见英语复数形态,例如:
s$ → s (cat → cats) (ax|test)is$ → $1es (analysis → analyses) ([^aeiouy]|qu)y$ → $1ies (body → bodies) (x|ch|ss|sh)$ → $1es (box → boxes) (^[m|l])ouse$ → $1ice (mouse → mice) ^(ox)$ → $1en (ox → oxen) (quiz)$ → $1zes (quiz → quizzes)② 单数化规则(AddSingular,约 30 条正则)与复数规则互为逆映射,例如s$ → ""(cats → cat)、(n)ews$ → $1ews(保证 news 不被改写)、([dti])a$ → $1um(data → datum)、(^[m|l])ice$ → $1ouse(mice → mouse)等。
③ 不规则词(AddIrregular)如person/people、man/men、child/children、goose/geese、foot/feet、tooth/teeth、curriculum/curricula、zombie/zombies等。其中部分词以matchEnding: false注册(如is/are、was/were、this/these、bus/buses、die/dice),表示只在整词匹配时生效,不参与词尾匹配。
④ 不可数名词(AddUncountable,约 50 个)如staff、information、equipment、fish、sheep、deer、series、species、metadata(对应 issue 1132 的修复)、以及计量单位oz、tsp、tbsp、ml、l等。这些词在单复数转换时原样返回。
三、通过 Default 注册自定义规则
3.1 添加复数 / 单数正则规则
Vocabulary暴露了与内置规则同级的注册 API,见 src/Humanizer/Inflections/Vocabulary.cs:
public void AddPlural(string rule, string replacement) public void AddSingular(string rule, string replacement)规则为正则表达式(大小写不敏感),例如对 "bus → buses" 这类不规则词:
Vocabularies.Default.AddPlural("(bus)es$", "$1");自定义规则会追加到规则列表尾部。从ApplyRules的实现(Vocabulary.cs)看,匹配从列表末尾向前扫描,因此后注册的自定义规则优先于内置规则——这正是自定义覆盖默认行为的基础。测试 tests/Humanizer.Tests/InflectorTests.cs 验证了这一点:向独立Vocabulary注册AddPlural("meter per second", "meter rate units")后,Pluralize("meter per second")返回自定义结果而非默认的 "meters per second"。
3.2 添加不可数名词
public void AddUncountable(string word)见 Vocabulary.cs。不可数词被存放在HashSet<string>(大小写不敏感比较)中,在ApplyRules第一步即短路返回原词(Vocabulary.cs)。因此向Vocabularies.Default.AddUncountable("metadata")后,"metadata".Pluralize()与"metadata".Singularize()均返回 "metadata"。测试 InflectorTests.cs 还验证了整条短语(如 "meter per second")也能注册为不可数。
3.3 添加不规则词
public void AddIrregular(string singular, string plural, bool matchEnding = true)见 Vocabulary.cs。实现细节值得注意:
matchEnding: true(默认)时,取首字符后的子串拼成后缀正则,例如AddIrregular("person", "people")等价于注册(p)erson$ → $1eople与(p)eople$ → $1erson,即person 出现在长词末尾时也会被变换;matchEnding: false时使用^singular$与^plural$整词锚定,如内置的is → are、was → were,避免把 "this"/"his" 等词误伤。
3.4 添加缩写词(Acronym)
public void AddAcronym(string acronym)见 Vocabulary.cs。它注册一个"规范大小写"的缩写词,使Humanize()在拆分字符串时保留其官方写法。参数约束严格:
- 不能为
null(抛出ArgumentNullException); - 不能为空串或包含非字母字符(抛出
ArgumentException)。
测试 tests/Humanizer.Tests/StringHumanizeTests.cs 验证了null、空串与 "HS2"(含数字)均会被拒绝。
典型用法与效果(同文件 StringHumanizeTests.cs):
Vocabularies.Default.AddAcronym("iOS"); Vocabularies.Default.AddAcronym("API"); "IOS".Humanize() // => "iOS" "iOSSettings".Humanize() // => "iOS settings" "iOSAPISettings".Humanize() // => "iOS API settings" "iOS5".Humanize() // => "iOS 5"注意缩写词规则同样作用于 Humanize 之前的内部分词:"TheHTMLLanguage".Humanize()输出 "The HTML language","HTML5".Humanize()输出 "HTML 5"(见 StringHumanizeTests.cs)。注册的缩写词使用RegexOptions.IgnoreCase匹配,因此"hsAccess".Humanize()在注册 "HS" 后会输出 "HS access"。
四、扩展方法如何消费词库
4.1 Pluralize / Singularize
两者的完整签名(InflectorExtensions.cs):
public static string? Pluralize(this string? word, bool inputIsKnownToBeSingular = true) public static string Singularize(this string word, bool inputIsKnownToBePlural = true, bool skipSimpleWords = false)inputIsKnownToBeSingular/Plural:当不确定输入的单复数时传false,方法会先尝试反向变换再正向验证,避免双重变形的误判。测试 InflectorTests.cs 覆盖了该分支。skipSimpleWords:为true时跳过"仅以 s 结尾"的简单词,防止 "tires" 被错误单数化为 "tire",见 InflectorTests.cs。- 大小写保持:
MatchUpperCase(Vocabulary.cs)保证全大写输入得到全大写输出("PERSON" → "PEOPLE"),首字母大写输入保留首字母大写("Foot per Second" → "Feet per Second"),且使用 invariant culture(见 InflectorTests.cs 的 tr-TR 文化测试)。 - 复合比率词:形如 "X per Y" 的复合词只变形分子、保留分母("meter per second" → "meters per second"),实现位于
CompoundHeadLength(Vocabulary.cs),相关用例见 InflectorTests.cs。
4.2 Humanize 与缩写词
Humanize()内部通过Vocabularies.Default.ApplyAcronyms/NormalizeAcronyms(Vocabularies.cs)识别并保留注册的缩写词写法;Vocabulary.ApplyAcronyms的实现见 Vocabulary.cs,采用lock保证并发注册/匹配安全(对应测试 StringHumanizeTests.cs 的并行注册用例)。
五、使用边界与注意事项
- 进程级副作用:
Vocabularies.Default是全局单例,注册规则影响整个应用域的所有 Humanize/Pluralize/Singularize 调用。建议在应用启动阶段一次性注册,避免运行时反复增删。 - 只支持单一词库:当前版本不提供多词库隔离或删除规则的能力(
RemoveAcronym为 internal,仅供测试清理)。若需隔离实验,可如测试那样直接new Vocabulary()构造独立实例(构造函数为 internal,测试通过InternalsVisibleTo访问)。 - 正则规则是追加而非替换:自定义规则追加在列表尾部并在匹配时优先扫描,利用这一点可覆盖内置行为;但注意规则数量增长会线性影响每次转换的性能,不宜高频动态添加。
- 缩写词仅限字母:
AddAcronym拒绝含数字的输入(如 "HS2" 会抛ArgumentException),数字与字母混合场景由内置分词逻辑另行处理。
六、深入阅读
- 类定义与内置规则:src/Humanizer/Inflections/Vocabularies.cs
- 规则存储与匹配引擎:src/Humanizer/Inflections/Vocabulary.cs
- 扩展方法入口:src/Humanizer/InflectorExtensions.cs
- 测试覆盖:tests/Humanizer.Tests/InflectorTests.cs、tests/Humanizer.Tests/StringHumanizeTests.cs
- 配套 API 文档:Humanizer.Vocabulary
结语
Vocabularies是 Humanizer 屈折体系的单一入口:Vocabularies.Default承载了一套精心维护的美式英语规则库,并允许开发者以追加方式注册正则规则、不规则词、不可数名词与缩写词,从而精确控制Pluralize()、Singularize()与Humanize()的行为。理解它的延迟构建、逆序匹配与大小写保持机制,就能在真实项目中安全、高效地定制词库,而不至于陷入"规则不生效"或"全局污染"的坑。
- 开发工具
【免费下载链接】Humanizer
Humanizer meets all your .NET needs for manipulating and displaying strings, enums, dates, times, timespans, numbers and quantities
相关推荐
Humanizer 词汇表机制详解:Vocabularies 与 Vocabularies.Default 的规则注册与应用
Humanizer 词汇表机制详解:Vocabularies 与 Vocabularies.Default 的规则注册与应用 Vocabularies 是 Hu
开发工具Humanizer 屈折变换指南:深入解析 Humanizer.Inflections 的 Vocabularies 与 Vocabulary 词表体系
Humanizer 屈折变换指南:深入解析 Humanizer.Inflections 的 Vocabularies 与 Vocabulary 词表体系 导读
开发工具Humanizer Vocabularies 全面解析:掌握 .NET 单复数与缩写词规则的注册与定制
Humanizer Vocabularies 全面解析:掌握 .NET 单复数与缩写词规则的注册与定制 本文聚焦 Humanizer 的 Vocabularie
开发工具
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考