内核 Kubernetes 控制面的"唯一真相源":Raft 强一致 + MVCC watch
- 一句话
- 为"小数据、强一致、高可靠的协调"而生的分布式 KV:Raft 保证线性一致,MVCC revision + watch 让 k8s 的 controller 模式(list-watch)成为可能;它是"选主/配置/服务发现"的标准答案,不是通用 KV。A distributed KV born for "small data, strong consistency, high-reliability coordination": Raft election guarantees linearizability, MVCC revision + watch semantics make k8s's controller pattern (list-watch) possible; it's the standard answer for "leader election / config / service discovery" — not a general-purpose KV.
- 窄场景
- 集群元数据与协调状态——k8s 的所有资源对象、controller 选主 lease、服务发现;数据量小(GB 级)、读多写少、丢一条都不行。任何跑 k8s 的团队被动拥有,几乎不用选型。Cluster metadata and coordination state — all k8s resource objects, controller leader-election leases, service discovery; GB-scale data, read-heavy, every record must survive. Any team running k8s owns it passively — there's barely a selection decision.
- 机制
- 两层机制锁死这个生态位——① Raft:leader 处理所有写、多数派确认后才提交,天然线性一致;② MVCC:每次写递增全局 revision,watch 可以带 revision 断点续传——k8s apiserver 的 watch 缓存、controller 的 resync 全靠这个语义。官方 FAQ 的原话:leader 的存活本身就绑定在"能否及时 fsync"上,宁可丢 leader 也不接受一个"看起来活着但提交不了"的 leader。Two layers lock in this niche — ① Raft: the leader handles all writes, commits only after majority acknowledgment, linearizable by construction; ② MVCC: every write bumps a global revision, watches can resume from a revision breakpoint — the k8s apiserver's watch cache and controller resyncs depend entirely on this semantic. The official FAQ states it plainly: a leader's liveness is bound to "whether it can fsync in time" — etcd would rather lose a leader than accept one that "looks alive but can't commit."
- 生产验证
- 全球每一个生产 k8s 集群的控制面都跑在 etcd 上,这是最大规模的生产验证(社区共识级)。官方 FAQ 对"为什么磁盘抖动会丢 leader"的解释是该设计哲学的第一手陈述(https://github.com/etcd-io/website/blob/HEAD/content/en/docs/v3.5/faq.md)。Kubernetes itself — every production k8s cluster's control plane runs on etcd, the largest-scale production validation there is (community-consensus grade). The official FAQ's explanation of "why disk jitter loses leaders" is the first-hand statement of this design philosophy (https://github.com/etcd-io/website/blob/HEAD/content/en/docs/v3.5/faq.md).
- 竞品差距
- 31 款里没有同类(Consul/ZooKeeper 不在名单)。Redis 有 pub/sub 但无持久化、无 revision 语义,做选主是反模式;DynamoDB 有条件写可做分布式锁,但延迟与一致性语义不同且是云厂商绑定。No equivalent among the 31 (Consul/ZooKeeper aren't on the list). Redis has pub/sub but no persistence and no revision semantics — leader election on Redis is an anti-pattern; DynamoDB conditional writes can implement distributed locks, but with different latency/consistency semantics and cloud-vendor binding.
- 证据等级
- 社区共识(k8s 生态的既定事实)+ 官方文档。诚实备注:官网吹的是"reliable distributed key-value store",厂商和用户一致;错位在反方向——很多人误把它当通用 KV 用。Community consensus (an established fact of the k8s ecosystem) + official docs. Honest note: etcd's homepage never brags about performance — it brags about being a "reliable distributed key-value store." Vendor and users agree this time. The misalignment runs the other way: too many people misuse it as a general KV.
- 最后核验
- 2026-10-012026-10-01