内核 LSM 追加写:吃下"写多读少、只追加"的消息流
- 一句话
- 宽列 + LSM-tree 存储引擎,为"海量追加写、极少改删、按时间范围读"的 workload 而生;写路径是 memtable + commit log 顺序追加,天然吞掉写入洪峰。Wide-column + LSM-tree storage engine, born for "massive append writes, rare updates/deletes, time-range reads" — the write path is memtable + commit log sequential appends, naturally absorbing write floods.
- 窄场景
- 聊天消息/事件流这类 workload——每天数十亿条写入、几乎不修改、极少删除、读取按 `(channel_id, 时间桶)` 范围扫描;团队处于从单机 MongoDB 向分布式演进的阶段。Chat-message/event-stream workloads — billions of writes per day, almost no modifications, rare deletes, reads as range scans over `(channel_id, time bucket)`; teams evolving from single-node MongoDB toward distributed.
- 机制
- MongoDB 的 B-tree 模型下,数据 + 索引必须能塞进内存才能保证延迟,这是随机写/随机读放大的结果,不是调参能解决的。Cassandra 的 LSM 写路径永远是顺序追加;宽列模型的 `(channel_id, bucket)` 分区键 + `message_id` 聚类键,让"同一频道按时间读"变成一次分区内的顺序扫描——这正是 LSM 最舒服的形状。Under MongoDB's B-tree model, data + indexes must fit in RAM to guarantee latency — that's random write/read amplification, not something tuning can fix. Cassandra's LSM write path is always sequential append; the wide-column model's `(channel_id, bucket)` partition key + `message_id` clustering key turns "read one channel by time" into one sequential scan inside a partition — exactly the shape LSM loves.
- 生产验证
- Discord 官方工程博客《How Discord Stores Billions of Messages》(2017,一线工程师撰写):2015 年底存到约 1 亿条消息时,MongoDB 数据和索引塞不进 RAM,延迟不可预测;迁到 Cassandra 后 12 节点、复制因子 3 时读 <5ms、写亚毫秒(https://discord.com/blog/how-discord-stores-billions-of-messages);2023 年续篇回顾确认了这段历史(https://discord.com/blog/how-discord-stores-trillions-of-messages)。Discord's engineering blog "How Discord Stores Billions of Messages" (2017, written by a frontline engineer): by late 2015, at ~100M messages, MongoDB data and indexes no longer fit in RAM and latency became unpredictable; after moving to Cassandra, 12 nodes at replication factor 3 delivered <5ms reads and sub-millisecond writes (https://discord.com/blog/how-discord-stores-billions-of-messages); the 2023 follow-up confirmed this history (https://discord.com/blog/how-discord-stores-trillions-of-messages).
- 竞品差距
- MongoDB——B-tree 模型,在该 workload 下先撞内存墙(Discord 亲历);ScyllaDB——同一宽列模型 + LSM,机制相同,是"换引擎"而非"换模型";DynamoDB——分区键模型类似,但 serverless 计费在该量级下成本不可比。MongoDB — B-tree model, hits the memory wall first on this workload (Discord lived it); ScyllaDB — same wide-column model + LSM, same mechanics, a "change of engine" not "change of model"; DynamoDB — similar partition-key model, but serverless billing is incomparable at that scale.
- 证据等级
- 社区实践(一线工程师生产分享,非厂商稿)。厂商 vs 用户无明显错位:官网也讲"write fast",但工程师记住的是"1 亿条消息时 MongoDB 延迟不可预测"这个具体断裂点。Community practice (frontline engineer production write-ups, not vendor copy). No vendor/user misalignment: the site also says "write fast," but what engineers remember is the concrete fracture point — "latency unpredictable at 100M messages on MongoDB."
- 最后核验
- 2026-10-012026-10-01