- 数据同步
- 前端
【免费下载链接】yjs
Shared data types for building collaborative software
本篇文章以 Yjs 仓库中的 attributing-content.md 为骨架,系统讲解 Yjs 用于实现"内容归属"(Attribution)的两大核心数据结构——IdSet与IdMap,以及它们如何支撑起类似 Google Docs 的变更标注能力。读完本文,你将掌握:如何用IdSet/IdMap高效表示 ID 区间、如何通过 diff/merge/intersect 等集合运算分离出"某个版本之后发生了哪些插入与删除",如何将变更归属(Attribution)映射到具体用户并借助渲染器(Renderer)输出带标注的 Delta,以及这些数据最终如何在网络传输中被压缩编码。
为什么需要内容归属:Google Docs 式版本追踪
在协同编辑场景中,仅仅知道"文档当前长什么样"往往不够。当用户在 Google Docs 中点击某个历史版本时,界面会展示类似这样的注解式变更——谁在何时插入了什么、删除了什么、修改了哪些格式:
# 例如:Bob 在上一版本 "hello " 之后追加了 "world" [{ insert: 'hello' }, { insert: 'world', color: 'blue', creator: 'Bob', when: 'yesterday' }] # 例如:Bob 从上一版本 "hello world" 中删除了 "world" [{ insert: 'hello' }, { insert: 'world', backgroundColor: 'red', creator: 'Bob', when: 'yesterday' }]Yjs 采用完全不同的思路来实现这一目标:所有变更在 Yjs 中都可以通过 ID 唯一标识(ID 由client与clock构成,见 src/utils/ID.js)。因此我们可以把"一段内容"抽象成一组 ID 区间,把"某个用户改了哪些内容"抽象成从 ID 区间到"属性(Attribution)"的映射。渲染时,toString()、toDelta()这类方法对无归属的内容原样输出,而对有归属的内容则附带额外信息(如creator、when),由编辑器决定如何用背景色等视觉效果呈现。
这套机制在仓库中的正式名称为Content Attribution(内容归属),其官方文档正是 attributing-content.md,底层实现分布在 src/utils/ids.js、src/utils/Renderer.js 与 src/utils/renderer-helpers.js 中。
两大基础数据结构:IdSet 与 IdMap
IdSet:ID 区间的紧凑表示
IdSet是 Yjs 中用于高效表示"一组 ID 区间"的数据结构(其前身名为DeleteSet,历史上用于表示已删除内容)。在 Yjs 中一切内容都由 ID 标识,因此任何内容集合——例如"文档里所有插入的内容"或"文档里所有被删除的内容"——都可以表达为若干条(client, clock, length)形式的 ID 区间。
从源码结构看,src/utils/ids.js 中的IdSet围绕客户端 ID(client)组织区间列表:内部为每个 client 维护一组排序合并后的IdRange,相邻区间会被合并,因此同样的内容集合占用空间极小。createIdSet()(ids.js)用于创建空集,insertIntoIdSet(ids.js)负责插入/合并区间。
IdMap:从 ID 到属性的映射
IdMap是相对较新的数据结构,用于将 ID 区间映射到属性列表(Attribution)。它与IdSet的结构类似,但每个区间额外携带一个Array<ContentAttribute>。一个ContentAttribute就是一条形如{ name: 'insert', val: 'Bob' }的记录,由createContentAttribute(name, val)(ids.js)创建,用于表达"这段区间被谁以什么方式修改"。
通用集合运算:diff、merge、intersect
IdSet与IdMap均支持全套标准集合运算,这正是归属功能的基础。文档明确指出:"We can perform all usual set operations onIdMaps andIdSets: diff, merge, intersect."
对应到 src/utils/ids.js 中的实现:
| 运算 | IdSet 版本 | IdMap 版本 | 语义 |
|---|---|---|---|
| 差集(diff) | diffIdSet(L721) | diffIdMap(L1666) | 从集合 A 中剔除集合 B 覆盖的区间 |
| 并集(merge) | mergeIdSets(L576) | mergeIdMaps(L1421) | 合并多个集合,属性按区间拼接去重 |
| 交集(intersect) | intersectSets(L773) | intersectMaps(L1673) | 取两集合共同覆盖的区间 |
此外还有createIdSetFromIdMap(L1485)可以把一个IdMap的覆盖范围提取为IdSet,供快速intersects/covers查询使用。
从文档中提取变更:三类关键工厂函数
要为一组变更打上归属,第一步是从文档结构中提取"变更集"。文档与源码提供了三个相互配合的工厂函数:
Y.createInsertSetFromStructStore(store, filterDeleted)(ids.js):遍历某个文档的 StructStore,返回其中所有插入内容的IdSet。第二个参数为false时不做删除过滤。Y.createDeleteSetFromStructStore(store)(ids.js):返回该文档中所有已删除内容的IdSet。Y.diffIdSet(a, b):计算a - b,用于"排除掉某个基准版本已有的内容",从而得到"从这个版本到那个版本之间新增的变更"。
例如,若要找出"ydoc相对于ydocVersion0的所有插入与删除",文档给出的核心模式是:
const insertionSet = Y.createInsertSetFromStructStore(ydoc.store, false) const deleteSet = Y.createDeleteSetFromStructStore(ydoc.store) // exclude the changes from `ydocVersion0` const insertionSetDiff = Y.diffIdSet(insertionSet, Y.createInsertSetFromStructStore(ydocVersion0.store, false)) const deleteSetDiff = Y.diffIdSet(deleteSet, Y.createDeleteSetFromStructStore(ydocVersion0.store))实战:为一段文本变更添加归属
文档给出了一个完整可运行的端到端示例,我们在此完整继承并逐步讲解。
第一步:构造基准文档与变更文档
// We create some initial content "Hello World!". Then we create another // document that will have a bunch of changes (make "Hell" italic, replace "World" // with "attributions"). const ydocVersion0 = new Y.Doc({ gc: false }) ydocVersion0.get().insert(0, 'Hello World!') const ydoc = new Y.Doc({ gc: false }) Y.applyUpdate(ydoc, Y.encodeStateAsUpdate(ydocVersion0)) const ytext = ydoc.get() ytext.applyDelta(delta.create().retain(4, { italic: true }).retain(2).delete(5).insert('attributions').done())要点说明:
gc: false至关重要。归属功能需要读取已被删除的内容(例如渲染被删除的 "World"),如果开启垃圾回收,这些内容会被 GC 掉,渲染器将无法还原。这一点在 Renderer.js 的AttributionsRenderer文档注释中也有明确要求:"Restoring deleted content requiresgc: falseon the doc"。ytext.applyDelta依次执行:retain(4, { italic: true })(前 4 个字符 "Hell" 加斜体)、retain(2)(跳过 "o ")、delete(5)(删除 "World")、insert('attributions')(插入新词)。- 这里
delta.create()是来自lib0/delta的 Delta 构建器(在 tests/attribution.tests.js 中同样通过import * as delta from 'lib0/delta'使用);Y.applyUpdate/Y.encodeStateAsUpdate是 Yjs 的标准更新同步 API。
第二步:计算基准版本与当前版本的变更差集
// this represents all insertions of ydoc const insertionSet = Y.createInsertSetFromStructStore(ydoc.store, false) const deleteSet = Y.createDeleteSetFromStructStore(ydoc.store) // exclude the changes from `ydocVersion0` const insertionSetDiff = Y.diffIdSet(insertionSet, Y.createInsertSetFromStructStore(ydocVersion0.store, false)) const deleteSetDiff = Y.diffIdSet(deleteSet, Y.createDeleteSetFromStructStore(ydocVersion0.store))此时insertionSetDiff中只包含 "attributions" 这个新插入的区间以及格式变更相关的区间,deleteSetDiff中只包含被删除的 "World" 区间——基准版本原有的 "Hello World!" 都被diffIdSet排除掉了。
第三步:将差集映射为归属(Attribution)
// assign attributes to the diff const attributedInsertions = Y.createIdMapFromIdSet(insertionSetDiff, [Y.createContentAttribute('insert', 'Bob')]) const attributedDeletions = Y.createIdMapFromIdSet(deleteSetDiff, [Y.createContentAttribute('delete', 'Bob')])Y.createIdMapFromIdSet(idset, attrs)(ids.js)把一份IdSet整体包装成IdMap,并给其中的每个区间都附上同一组属性。这里我们用createContentAttribute('insert', 'Bob')表示"这些内容是 Bob 插入的",用createContentAttribute('delete', 'Bob')表示"这些内容是 Bob 删除的"。
第四步:用渲染器输出带归属的 Delta
// now we can define a renderer that maps these changes to output. One of the // implementations is the TwosetRenderer const renderer = new Y.TwosetRenderer(attributedInsertions, attributedDeletions) // we render the attributed content with the renderer const attributedContent = ytext.toDelta({ renderer }) console.log(JSON.stringify(attributedContent.toJSON(), null, 2)) const expectedContent = delta.create().insert('Hell', { italic: true }, { format: { italic: ['Bob'] } }).insert('o ').insert('World', {}, { delete: ['Bob'] }).insert('attributions', {}, { insert: ['Bob'] }).insert('!') t.assert(attributedContent.equals(expectedContent))ytext.toDelta({ renderer })让渲染器介入 Delta 生成过程。从 renderer-helpers.js 的attributionJsonSchema(L7)可以确认归属信息的标准 JSON 形态:Delta 操作可以附带insert(字符串数组,表示插入者)、delete(字符串数组,表示删除者)、format({ 格式名: [归属者列表] })等字段,这正是文档示例输出中attribution对象的来源。
需要说明的是:文档中提到的TwosetRenderer属于示例性质的渲染器之一;在当前仓库中,AbstractRenderer的具体实现由 src/utils/Renderer.js 提供,包括AttributionsRenderer(L63,将{ inserts, deletes }两份归属图合并渲染)、DiffRenderer(L359,比较两个文档并归属两者之间的差异)、SnapshotRenderer(L613,面向快照场景)。它们输出的归属格式与文档示例保持一致——这一点由 tests/attribution.tests.js 中的testAttributionSession1(L143)等用例直接断言验证,例如期望insert('a', null, { insert: ['0'] })、insert('b', null, { delete: ['0'] })这样的输出形态。
第五步:解读渲染结果
上述代码会输出如下 JSON(文档原样给出):
{ "type": "delta", "children": [ { "type": "insert", "insert": "Hell", "format": { "italic": true }, "attribution": { // no "insert" attribution: the insertion "Hell" is not attributed to anyone "format": { "italic": [ // the formatting attribute "italic" was added by Bob "Bob" ] } } }, { "type": "insert", "insert": "o " // the insertion "o " has no attributions }, { "type": "insert", "insert": "World", "attribution": { // the insertion "World" was deleted by Bob "delete": [ "Bob" ] } }, { "type": "insert", "insert": "attributions", // the insertion "attributions" was inserted by Bob "attribution": { "insert": [ "Bob" ] } }, { "type": "insert", "insert": "!" // the insertion "!" has no attributions } ] }这段输出与 Google Docs 的版本注解语义高度一致:
"Hell"这个插入本身不归属于任何人,但它携带的italic格式是 Bob 添加的,因此attribution.format.italic = ["Bob"];"o "与"!"是基准版本就存在且未被修改的内容,没有任何归属;"World"是被 Bob 删除的内容,归属标记为attribution.delete = ["Bob"];"attributions"是 Bob 插入的新内容,归属标记为attribution.insert = ["Bob"]。
渲染器只负责把这些归属信息"翻译"进 Delta,至于最终用背景色(如红色高亮删除、蓝色高亮插入)等视觉效果呈现,则是编辑器层的职责——这正是文档所说的 "It will be the job of the editor to render those changes with background-color etc."。
多用户归属与按用户筛选
内容当然可以归属于多个用户。文档给出的示例是为同一段删除同时附加两条属性:
const attributedDeletions = Y.createIdMapFromIdSet(deleteSetDiff, [Y.createContentAttribute('insert', 'Bob'), Y.createContentAttribute('insert', 'OpenAI o3')])在真实协同系统中,更常见的做法是监听每个用户的 update 事件,动态累积全局归属表。仓库测试 testAttributionSession1 给出了这一模式的完整实现:
const globalAttributions = Y.createContentMap() // 见 src/utils/meta.js#L103 users.forEach(user => user.on('update', (update, _, ydoc, tr) => { if (!tr.local) return const userid = ydoc.clientID.toString() const contentIds = Y.createContentIdsFromUpdate(update) Y.insertIntoIdMap(globalAttributions.inserts, Y.createIdMapFromIdSet(contentIds.inserts, [Y.createContentAttribute('insert', userid)])) Y.insertIntoIdMap(globalAttributions.deletes, Y.createIdMapFromIdSet(contentIds.deletes, [Y.createContentAttribute('delete', userid)])) }))这里的Y.createContentMap()(meta.js)返回{ inserts, deletes }两份IdMap的包装结构,是归属渲染器的标准输入格式;createContentIdsFromUpdate从一次网络同步的 update 中提取插入/删除 ID 区间。测试随后通过Y.filterIdMap(...)筛选出仅属于某个用户的归属,构造出"只看 Bob 的改动"的专属渲染器,再用Y.undoContentIds撤销特定用户的改动——这些都是 src/utils/meta.js 提供的高层辅助能力。
测试 testYdocDiff 还验证了Y.diffDocsToDelta(ydocStart, ydocUpdated)这一高层 API:它直接比较两个文档并输出带{ insert: [] }/{ delete: [] }归属标记的 Delta(空数组表示"有变更但未命名归属者")。
AbstractRenderer 与自定义高亮渲染
AbstractRenderer是渲染器的抽象基类,负责把归属信息映射到内容上。它在 src/utils/renderer-helpers.js 中定义,核心契约有三个:
hasItem(item):判断某个 Item 是否命中归属集合。只有被hasItem认领的内容才会走渲染器路径,其余内容走通用快速路径原样渲染(见 Renderer.js 中AttributionsRenderer.hasItem的实现)。readContent(...):把内容按归属区间切分,逐个产出AttributedContent(renderer-helpers.js,携带content、clock、deleted、attrs与渲染行为标记)。contentLength(item):计算归属内容渲染后的长度,供遍历迭代器使用。
文档明确指出:通过这种抽象,"It is possible to highlight arbitrary content with this approach"——即任何内容片段都可以被任意高亮/标注,而不仅限于插入与删除。rendererContentLength(renderer-helpers.js)提供了通用长度规则(存活可计数内容全量渲染,其余为 0,除非渲染器认领),供各具体渲染器复用。渲染器还可以通过ObservableV2派发change事件,在归属发生变化后通知上层增量更新当前内容的归属标注。
从归属输出计算真实 Diff
由于归属输出本身就是结构化的 Delta,文档指出可以直接用它计算"纯 diff"(只包含删除与插入、去掉归属信息):
You could use the same output to calculate a real diff as well (consisting of deletions and insertions only, without Attributions).
换言之,同一份渲染结果既能驱动"版本注解"视图(展示谁改了哪里),也能驱动普通的差异视图(只展示改了什么)。这是toDelta({ renderer })输出模型带来的直接红利,也解释了为什么仓库测试(如 testAttributedEvents)会反复断言insert('world', null, { delete: [] })这类形态:归属存在与否、归属为空与否,都在同一份 Delta 结构中表达。
高效编码:RLE 与属性去重
归属数据最终需要与文档更新一起在网络上传输,因此编码效率至关重要。文档给出了两个硬性结论:
- ID 采用游程编码(run-length encoding):相邻的 ID 区间合并后以
(client, clock, length)三元组紧凑表达; - 属性去重、只编码一次:相同的属性(如
'insert'、'Bob')不会重复出现。 - 实测结果:上述示例整体仅编码为 27 字节。
在仓库实现中,这一高效编码由 src/utils/BlockSet.js 承担。writeBlockSet(L87)按 client 分组写出块引用(GC/Skip/Item),readBlockSet(L25)对称地读回;BlockSet.toIdSet()(L118)可把块集合还原为IdSet,exclude()(L144)则把被排除的区间替换为Skip占位块以保持编码紧凑。区间排序合并的逻辑位于 ids.js 的AttrRanges.getIds()(L1118)中:先按clock排序,再对重叠区间切分、对相邻同属性区间合并,这正是 RLE 压缩的算法基础。
测试验证与可靠性
归属功能并非纸面设计,仓库通过 tests/attribution.tests.js 提供了系统性的验证,可作为读者理解行为预期的活文档:
testAttributionSession1(L143):多用户协同下的全局归属累积、按用户筛选、撤销单用户改动;testAttributedEvents(L38)与testAttributionEvent(L180):在事件回调中通过event.getDelta({ renderer })获取带归属的变更 Delta,验证删除段落时子内容也会获得归属;testAttributionChange(L203):渲染器change事件在归属变化时触发,并提供三态(null 清除)归属更新;testInsertionsMindingAttributedContent(L62)与testInsertionsIntoAttributedContent(L80):在带归属渲染器下继续插入内容,文档实际状态依然正确,归属视图与真实状态互不干扰;testRdtDeltaAttributionSanity(L755):维护中的增量 Delta 缓存与全新深渲染结果保持一致。
小结
Yjs 的内容归属机制是一条完整的技术链路:以IdSet紧凑表示 ID 区间,以IdMap将区间映射到归属属性,通过diff/merge/intersect 集合运算分离版本间变更,借助AbstractRenderer家族把归属渲染进 Delta,最后以RLE + 属性去重的方式压缩编码。这套设计让 Yjs 在保持 CRDT 数据模型不变的前提下,原生支持"谁在何时改了什么"的版本注解能力,且与快照、撤销、多用户协同等既有能力自然衔接。无论是实现一个 Google Docs 风格的版本历史面板、一个审阅/建议模式,还是一个按用户着色的高亮层,上述数据结构与渲染器都是可以直接落地的核心构件。
- 数据同步
- 前端
【免费下载链接】yjs
Shared data types for building collaborative software
相关推荐
OneUptime AI 编程助手可观测性:基于 OpenTelemetry 追踪代码助手的用量、成本与员工归属
OneUptime AI 编程助手可观测性:基于 OpenTelemetry 追踪代码助手的用量、成本与员工归属 当团队中的工程师开始使用 Claude Cod
可观测性后端运维前端云原生微服务AI AgentYjs v14 Attribution 特性实战指南:用 Renderer 为 Delta 渲染作者归属、删除痕迹与变更高亮
Yjs v14 Attribution 特性实战指南:用 Renderer 为 Delta 渲染作者归属、删除痕迹与变更高亮 导读 Attribution(归属
数据同步前端Nuxt Content 内容编辑全指南:从可视化到代码编辑
Nuxt Content 内容编辑全指南:从可视化到代码编辑 前言 在现代内容管理系统中,内容编辑体验直接影响着开发者和内容创作者的效率。Nuxt Conten
前端CMS
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考