- TrueTime + 外部一致性:把"全球线性化"做成产品
它是什么:每个数据中心 GPS + 原子钟,TrueTime API 返回的不是时间点而是区间 [earliest, latest];事务提交时取 commit_ts = TT.now().latest,然后 commit-wait(等到 TT.now().earliest > commit_ts 才释放锁、回 ACK),保证任何后开始的事务拿到更大的时间戳——全球任意两节点的因果序与实时序一致。 为什么是真本事:不靠单点授时中心、不靠中心化协调器,CockroachDB 的 HLC 方案在同样问题上只能做到可串行化而保不住跨键线性化(Jepsen 2017 确认其 causal-reverse 为 by-design)。Spanner 是极少数在论文之外、生产了 10 年以上的实现。 边缘真相:① 每次提交固定支付不确定性窗口(通常个位数 ms,且常被 Paxos quorum RTT 掩盖——官方口径);② TrueTime 是 Google 内建能力,你无法在别处复现,选了 Spanner 就等于把正确性赌在 Google 的时钟基建上;③ 若时钟不确定性超界,Spanner 选择阻塞而非冒险——这是保正确性的设计,但意味着极端时钟故障时写停摆。 证据:OSDI 2012 论文;Google Cloud 官方博客《Strict Serializability and External Consistency in Spanner》;2025 ACM SIGMOD Systems Award(厂商口径荣誉)。
- 无感水平扩展:自动分片、自动再平衡
它是什么:数据按主键 range 切分为 splits,负载/大小变化时自动分裂与迁移,无手动 reshard;加 node/PU 在线完成。 为什么是真本事:从单库 MySQL 长大到跨区多写,传统路线要经历"分库分表 + 应用层路由 + 一致性自理"的重构;Spanner 把这条路走完了且对外只暴露"加算力"。 边缘真相:① 扩展的是算力配额,热点键的单 split 吞吐天花板不会因加节点而消失——加 10 个节点也救不了一个热点主键(见深水区二);② 扩容按 node-hour 计费,缩容不及时就是烧钱;③ 自动再平衡的 split 迁移在后台吃 IO/网络,极端时可被观测到(Key Visualizer 可见)。
- 交错表(Interleaved Tables):物理共置的父子关系
它是什么:子表行与父表行按主键前缀物理存储在一起,INTERLEAVE IN PARENT... ON DELETE CASCADE;父+子的点查/join 是本地操作。 为什么是真本事:这是"分布式 join 很贵"问题的结构性解法——把访问模式编码进物理布局,90%+ 访问走父键的场景下,跨表读的延迟和成本都下来。 边缘真相:① 交错关系建表时确定,事后不可更改——建模错了要重建表迁数据;② 父子基数 1:N 且 N 极大(>10K 级)或子行独立高频更新时,交错反而制造热点/锁争用;③ ON DELETE CASCADE 对金融类数据是危险默认(删父带走一切),社区迁移指南明确建议金融表用 ON DELETE NO ACTION。 证据:官方 schema 设计文档;多份社区迁移技能指南一致建议。
- 双重 SQL 方言 + 多模型一库
它是什么:同一引擎同时讲 GoogleSQL 和 PostgreSQL 方言;之上叠加 Spanner Graph(GQL/openCypher)、ScaNN 向量检索(单索引 100 亿+向量,厂商口径)、全文检索、JSON/ARRAY/STRUCT 类型、Change Streams。 为什么是真本事:OLTP 主库顺手做图遍历、向量召回、CDC 下游,不用为每个模型再搭一套库——2026 年 Google 把它包装成"Agentic Data Cloud"底座,方向是对的。 边缘真相:① PG 方言是"子集 + 扩展",不是真 PG(见深水区五);② 图/向量/列存都是外挂能力,重度场景(大规模 OLAP、专业图分析)仍建议联邦到 BigQuery/专用引擎——官方自己也把重分析定位为"联邦",不是"原生"。
- Spanner Omni:GCP 之外的逃生舱
它是什么:2025 年发布的可下载容器化 Spanner,跑在 VM/Linux 容器/Kubernetes,可部署在 on-prem、AWS、Azure、边缘;支持与 GCP 托管版组成主备(热冷 failover)、多主权 jurisdiction 部署。 为什么是真本事:它部分回答了 Spanner 头号结构性批评——"只能跑在 Google Cloud"。对受监管行业(数据主权、灾备合规)是实质性解绑。 边缘真相:① 仍是专有许可,不是开源,商用条款/价格不明(需联系销售,待验证);② TrueTime 在 GCP 之外如何保证授时精度,官方细节未完全公开——全球强一致的魔法在自有机房里打几折,待验证;③ 自运维后,"零运维"这个最大卖点就没了。
- TrueTime + external consistency: making "global linearizability" a product
What it is: Every datacenter has GPS + atomic clocks; the TrueTime API returns not a point in time but an interval [earliest, latest]; on commit the transaction takes commit_ts = TT.now().latest, then commit-wait (locks are released and the ACK returned only after TT.now().earliest > commit_ts), guaranteeing that any transaction that starts later gets a larger timestamp — causal order and real-time order agree across any two nodes on the planet. Why it's a real capability: No single timekeeping center, no centralized coordinator; CockroachDB's HLC approach to the same problem achieves only serializability and cannot preserve cross-key linearizability (Jepsen 2017 confirmed its causal-reverse as by-design). Spanner is one of the very few implementations beyond a paper that has been in production for 10+ years. Edge truth: ① every commit pays a fixed uncertainty window (usually single-digit ms, often masked by Paxos quorum RTT — official docs); ② TrueTime is a built-in Google capability, you cannot reproduce it elsewhere — choosing Spanner means betting correctness on Google's clock infrastructure; ③ if clock uncertainty exceeds bounds, Spanner chooses to block rather than risk — the correctness-first design, but it means writes stall under extreme clock failure. Evidence: OSDI 2012 paper; Google Cloud official blog "Strict Serializability and External Consistency in Spanner"; 2025 ACM SIGMOD Systems Award (vendor-claimed honor).
- Transparent horizontal scaling: automatic sharding and rebalancing
What it is: Data is partitioned into splits by primary-key range, automatically split and migrated as load/size changes, with no manual resharding; adding nodes/PU completes online. Why it's a real capability: Growing from a single MySQL to multi-region multi-write traditionally means "application-level sharding + routing + DIY consistency" refactoring; Spanner walked that whole road and exposes only "add compute". Edge truth: ① what scales is the compute quota — the per-split throughput ceiling of a hot key won't disappear by adding nodes — adding 10 nodes won't save a single hot primary key (see deep-dive II); ② scaling is billed per node-hour, and scaling down late is burning money; ③ automatic rebalancing's split migration eats IO/network in the background and can be observed in extreme cases (visible in Key Visualizer).
- Interleaved tables: physically co-located parent-child relationships
What it is: Child rows are physically stored together with parent rows by primary-key prefix, INTERLEAVE IN PARENT... ON DELETE CASCADE; point lookups/joins of parent + child are local operations. Why it's a real capability: This is a structural answer to "distributed joins are expensive" — encode the access pattern into the physical layout, and for workloads where 90%+ of access goes through the parent key, cross-table read latency and cost both come down. Edge truth: ① the interleave relationship is fixed at table-creation time and cannot be changed afterwards — model it wrong and you're rebuilding tables and migrating data; ② when the parent-child fan-out is 1:N with very large N (>10K) or child rows are updated independently at high frequency, interleaving turns locality into a hotspot/lock-contention factory; ③ ON DELETE CASCADE is a dangerous default for financial data (deleting the parent takes everything), and community migration guides explicitly recommend ON DELETE NO ACTION for financial tables. Evidence: Official schema design docs; multiple community migration skill guides agree.
- Dual SQL dialects + multi-model in one engine
What it is: The same engine speaks both GoogleSQL and the PostgreSQL dialect; layered on top: Spanner Graph (GQL/openCypher), ScaNN vector search (10 billion+ vectors per single index, vendor claim), full-text search, JSON/ARRAY/STRUCT types, Change Streams. Why it's a real capability: The OLTP primary database can also do graph traversals, vector recall, and CDC downstream without standing up another database per model — in 2026 Google packaged it as the "Agentic Data Cloud" foundation, and the direction is sound. Edge truth: ① the PG dialect is a "subset + extensions", not real PG (see deep-dive V); ② graph/vector/columnar are bolt-on capabilities — heavy workloads (large-scale OLAP, professional graph analytics) are still better federated to BigQuery/specialized engines — the vendor itself positions heavy analytics as "federation", not "native".
- Spanner Omni: an escape hatch outside GCP
What it is: The 2025 release of downloadable containerized Spanner, running on VMs/Linux containers/Kubernetes, deployable on-prem, AWS, Azure, edge; supports forming primary/standby (hot-cold failover) pairs with the GCP managed edition and multi-sovereignty jurisdiction deployments. Why it's a real capability: It partially answers Spanner's #1 structural criticism — "it only runs on Google Cloud". For regulated industries (data sovereignty, disaster-recovery compliance) it is a substantive decoupling. Edge truth: ① still proprietary license, not open source, with unclear commercial terms/pricing (contact sales, to-be-verified); ② how TrueTime guarantees timekeeping precision outside GCP is not fully disclosed in official details — how much of the global-strong-consistency magic survives in your own datacenter is to-be-verified; ③ once you self-operate, the biggest selling point — "zero operations" — disappears.