6ix:MongoDB Atlas 迁往 PlanetScale Postgres 失败教训
成本敏感
读多写少
快速迭代/小团队
查询性能
- 决策
迁往 PlanetScale 的 Serverless Postgres(按 postgresql 产品计),把查询重写为 SQL。
Migrate to PlanetScale's Serverless Postgres (counted as the postgresql product) and rewrite queries in SQL.
- 结果
账单从 $5,749/月降到 $1,149/月(约 5 倍成本差);最热页面查询从 530ms 降到 9ms;据称一天内完成迁移(含 AI 辅助)。以上数字均为 6ix 公司博客单方口径,无第三方验证,"一天完成迁移"尤其只有单方说法。
The bill dropped from $5,749/month to $1,149/month (about a 5x cost difference); the hottest page query went from 530ms to 9ms; the company claims the migration finished in a day (with AI assistance). All figures above are the company's one-sided account on its own blog, with no third-party verification — the "one-day migration" claim especially so.
- 机制根因
在 6ix 的具体数据模型、查询、索引与 Atlas 计费配置下,迁移后 SQL 侧表现更好、账单更低。但这是单一样本的结论,不可推广:MongoDB 官方文档明确列出了聚合管道的重排、合并、索引利用、slot-based 执行等优化能力,不能从这一个 54GB / 28 TPS / 99.9% 读的应用推出"MongoDB 聚合管道优化器成熟度不如 SQL"的普遍规律。Atlas 按量计费对小团队是重税;多一套查询语言("Mongoose 税")是持续认知成本;Serverless Postgres 的连接池 + 分支模型对小团队更友好。
Under 6ix's specific data model, queries, indexes, and Atlas billing configuration, the SQL side performed better and billed lower after migration. But this is a single-sample conclusion and must not be generalized: MongoDB's official documentation lists aggregation-pipeline optimizations including reordering, coalescence, index utilization, and slot-based execution — one 54GB / 28 TPS / 99.9%-read app cannot support a general claim that "MongoDB's aggregation optimizer is less mature than SQL's". Atlas's usage-based billing is a heavy tax on small teams; maintaining a second query language (the "Mongoose tax") is an ongoing cognitive cost; Serverless Postgres's connection pooling plus branching model suits small teams better.
- 教训
读多写少的中小规模场景,文档库的灵活性收益可能被成本、查询性能与运维心智税吃掉;选型要把"账单结构"和"查询语言成熟度"计入硬指标,而不只看数据模型灵活。
In read-heavy, small-to-medium-scale scenarios, the flexibility gains of a document database can be eaten by cost, query performance, and operational mind-share tax; treat "billing structure" and "query language maturity" as hard criteria, not just data-model flexibility.
相关产品:PostgreSQL(社区版)、MongoDB 相关能力:MVCC + 查询优化器 + 事务性 DDL —— 复杂分析直接跑在 OLTP 主库上、Atlas:"set it and forget it" 的托管口碑 最后核验:2026-10-01
Airbnb:数百个 Aurora 实例合并为 TiDB,重建数据服务层(2023–2024) 成功经验
分片治理
数据库整合
成本优化
OLTP 扩展
- 场景
Airbnb's online data service layer long ran on Amazon Aurora (MySQL), growing with the business into hundreds of Aurora instances (instances, not clusters — PingCAP's original wording is "hundreds of Amazon Aurora instances") plus application-layer manual sharding to carry write traffic. The more shards, the more every MySQL version upgrade, incident investigation, and capacity-planning cycle multiplied operational work. Per PingCAP's official material, its data service layer was supported by "hundreds of Amazon Aurora instances and application-layer sharding" ("hundreds" is PingCAP's figure; no public data found for exact instance counts or QPS). (
https://static.pingcap.com/files/2024/11/18200414/TiDB_Amazon_Aurora_Datasheet.pdf)
- 决策
Airbnb 决定不再维持"分片 Aurora 舰队",把整个数据服务层重建在 TiDB(分布式 SQL、MySQL 协议兼容)上,目标是把分散的 KV 数据存储收敛成单一可信数据源。Airbnb 工程师在 2023 年 9 月 HTAP Summit 上公开分享了这一实践。选型逻辑:既然不能回到单机,就换一个自带自动分片、多写节点的分布式数据库,把分片逻辑从应用层下沉到数据库层。
Airbnb decided to stop maintaining its "sharded Aurora fleet" and rebuild the entire data service layer on TiDB (distributed SQL, MySQL-protocol compatible), converging scattered KV data stores into a single source of truth. An Airbnb engineer publicly shared the practice at HTAP Summit in September 2023. The selection logic: since there was no going back to a single node, switch to a distributed database with built-in auto-sharding and multi-writer nodes, pushing sharding logic down from the application layer into the database layer.
- 结果
PingCAP 官方 datasheet(2024 年 11 月版)宣称,通过数据库整合 Airbnb 降低了 60% 的数据库成本——该数字为 PingCAP 厂商口径,Airbnb 官方工程博客未披露对应数字,也未找到第三方独立复现,引用时须注明口径。迁移前后的 QPS、数据量、迁移耗时、TiDB 集群规模等数字未找到公开数据。
PingCAP's official datasheet (November 2024 edition) claims the consolidation cut Airbnb's database costs by 60% — that figure is PingCAP's vendor claim; Airbnb's own engineering blog has disclosed no corresponding number, and no independent third-party reproduction exists, so quote it with the source caveat. No public data found for pre/post-migration QPS, data volumes, migration duration, or TiDB cluster size.
- 机制根因
PingCAP 披露的迁移动因有四条:强制 MySQL 升级(Aurora 的升级窗口和节奏由云厂商控制)、缺乏数据库可观测性、写扩展能力有限、技术支持不佳。根因是架构模型差异:Aurora 本质是"单写 + 只读副本"的增强型单机模型,写入上限就是一台最大实例的上限;超过这个点只能靠应用层分片,而分片把一致性、重分片、DDL 等复杂性全部推给业务层,运维税随规模线性增长。TiDB 是多写分布式 SQL,自动分片 + Raft 复制把复杂性收回数据库内部——合并之后成本反而下降(实例数减少 + 释放业务层的分片代码维护人力)。代价是引入新数据库的迁移风险与团队学习成本,以及把鸡蛋放在一个新兴分布式系统的集中度风险。
PingCAP disclosed four migration drivers: forced MySQL upgrades (Aurora's upgrade windows and cadence controlled by the cloud vendor), lack of database observability, limited write scalability, and poor support. The root cause is an architecture-model mismatch: Aurora is essentially an enhanced single-node model ("single writer + read replicas") whose write ceiling is one maximum-size instance; past that point only application-layer sharding helps, and sharding pushes all consistency, re-sharding, and DDL complexity onto the business layer, with the ops tax growing linearly with scale. TiDB is multi-writer distributed SQL — auto-sharding + Raft replication pull that complexity back inside the database, so post-consolidation costs fell instead (fewer instances + freed application-layer sharding-code maintenance). The cost: migration risk and team learning curve of a new database, plus the concentration risk of betting on one emerging distributed system.
- 教训
应用层分片不是免费的,早期省下的架构成本会在版本升级、故障排查、扩容上以数倍奉还;当实例数以百计、应用层分片代码成为主要运维负担时,就是认真评估换架构的信号(经验法则,不是定律);把云厂商托管数据库的"强制升级节奏"计入 TCO,它不只是停机窗口,更是对所有分片逐一验证的隐性人力成本;数据库整合本身就是一种降本,分布式 SQL 的价值不只是写更快,而是把舰队式管理变成单一集群管理。
Application-layer sharding is not free — the architecture cost saved early gets repaid many-fold in version upgrades, incident response, and scaling; when instance counts reach the hundreds and application-layer sharding code becomes the dominant ops burden, that's the signal to seriously evaluate a new architecture (a rule of thumb, not a law). Count the cloud vendor's "forced upgrade cadence" in TCO: it's not just downtime windows, but the hidden labor of validating every shard one by one. Database consolidation is itself a cost cut: distributed SQL's value isn't just faster writes, it's turning fleet management into single-cluster management.
相关产品:Amazon Aurora、MySQL 相关能力:日志即数据库(log is the database)—— 存储计算分离的"真"实现 最后核验:2026-10-01
支付宝去 O:双十一做验证场的分布式替换(2010s) 成功经验
金融级事务
强一致
去 O
高并发写入
- 场景
The Double 11 payment peak demanded financial-grade transactions, strong consistency, and horizontal scaling all at once — a traditional single-node Oracle couldn't deliver all three. The following follows a September 2026 Sina Finance report (not Ant Group's official account): the claim that "Alipay's core traffic runs on OceanBase" is cited per that media report, and Ant's official side disclosed no equivalent detail in it. (
https://finance.sina.cn/stock/jdts/2026-09-07/detail-iniqymaa6344446.d.html)
- 决策
采用 OceanBase 分布式关系库替换 Oracle。
Replaced Oracle with the distributed relational database OceanBase.
- 结果
双十一连续十年以上作为验证场。以下量化数字均为 2026-09 媒体文章转述口径,非官方:TPC-C 7.07 亿 tpmC(媒体称打破 Oracle 九年纪录);Oracle 兼容 95%+;TNGD 案例同等硬件吞吐 +40%、4 万 TPS 压测零宕机;GCash 案例存储 −70%、成本 −40%。注意 TPC-C 是标准化基准,说明的是基准条件下的事务处理能力,不证明应用兼容性、真实查询/事务分布、迁移风险与故障恢复表现。
Double 11 served as the proving ground for more than ten consecutive years. The following figures are all 2026-09 media-report figures, not official: TPC-C 7.07 billion tpmC (media claim of breaking Oracle's nine-year record); 95%+ Oracle compatibility; TNGD case: +40% throughput on equal hardware, zero downtime under a 40k TPS stress test; GCash case: −70% storage, −40% cost. Note: TPC-C is a standardized benchmark — it speaks to transaction processing under benchmark conditions, not to application compatibility, real query/transaction mixes, migration risk, or failure-recovery behavior.
- 机制根因
分布式事务协议(2PC/Paxos 系)的正确实现是"去 O"的前提——用对协议不等于实现、运维与故障处理都正确,选型要看故障注入和恢复的工程证据,不能把"用了 Paxos"当成正确性证明;Oracle 高兼容大幅降低迁移成本;用真实业务峰值(双十一)做验证场而非实验室基准。
Correct implementation of distributed-transaction protocols (2PC/Paxos family) is the precondition for going off Oracle — using the right protocol does not equal correct implementation, operations, and failure handling; selection should rest on fault-injection and recovery engineering evidence, not on "it uses Paxos" as a proof of correctness; high Oracle compatibility drastically lowers migration cost; the real business peak (Double 11) is the proving ground, not lab benchmarks.
- 教训
金融级去 O 的核心公式 = 分布式事务正确性 + 高兼容降迁移成本 + 真实峰值验证;媒体与厂商的基准数字一律打折看,选型时盯住"压测零宕机"这类工程指标。
The formula for financial-grade off-Oracle migration = distributed transaction correctness + high compatibility to cut migration cost + real-peak validation; discount all media and vendor benchmark figures, and judge by engineering metrics like "zero downtime under stress."
来源
新浪财经转述(2026-09
Sina Finance report (2026-09
相关产品:Oracle Database(甲骨文)、OceanBase 相关能力:ELR(Early Lock Release)热点行更新、MySQL / Oracle 双模式兼容 最后核验:2026-10-01
拜耳:Field Answers 从自建 PostgreSQL 迁至 AlloyDB,收获季峰值平稳通过 成功经验
同构迁移
读扩展
季节性峰值
零应用改动
- 场景
拜耳作物科学(Bayer Crop Science)的 Field Answers 平台:收集并计算全球田间与温室运营中数十亿条观测数据(地图、表型观测、卫星影像、天气与土壤分层),支撑 R&D 管线中的选种、生产成本优化等决策。原架构是自建开源 PostgreSQL:主写节点既要承担写操作,又要负责向读节点复制变更。为新市场版块上线做压测时发现:写流量与读节点数双涨之下,主节点会成为瓶颈、复制延迟恶化,自建 PG 满足不了延迟与吞吐要求。
Bayer Crop Science's Field Answers platform collects and computes billions of observations across global field and greenhouse operations (maps, phenotypic observations, satellite imagery, weather and soil strata), supporting R&D pipeline decisions such as seed selection and production cost optimization. The original architecture was self-managed open-source PostgreSQL: the primary writer handled both writes and replication to reader nodes. Load testing ahead of onboarding a new market segment showed that rising write traffic plus more readers would overwhelm the primary, spiking replication lag - self-managed PG could not meet the latency and throughput bar.
- 决策
选择 AlloyDB for PostgreSQL。关键动因是 PG 兼容——零应用代码改动即可迁移,能赶上迫在眉睫的北美种植季;Google Cloud 团队在测试期提供支持兜底。数据战略延续 data mesh 思路:AlloyDB 负责在线数据,分析侧仍走 BigQuery。
Migrate to AlloyDB for PostgreSQL. The decisive factor was PG compatibility - zero application code changes, which made the aggressive timeline ahead of the North American planting season achievable; the Google Cloud team provided hands-on support during evaluation. The data strategy stayed data-mesh-flavored: AlloyDB for operational data, BigQuery for analytics.
- 结果
并行压测中,更小规格的 AlloyDB 实例相比原 PG 方案平均响应时间降低超 50%、吞吐提升 5 倍(拜耳 Global Data Assets 团队口述,经 Google Cloud 官方博客发布,属厂商渠道口径,引用须注明)。真正的考验是随后到来的首个收获季——农业的季节性决定了"晚几天就可能让产品上市推迟一整年",收获季平稳通过。AlloyDB 的"所有节点读同一份存储"架构让读流量扩展不再冲击主库,复制延迟保持低位。
In parallel load tests, a smaller AlloyDB instance cut average response times by over 50% and delivered 5x the throughput of the previous PostgreSQL setup (per Bayer's Global Data Assets team, published via the Google Cloud official blog - a vendor-channel account; cite with that caveat). The real test was the first peak harvest season that followed - in agriculture a delay of days can postpone a product launch by a full year - and it went smoothly. AlloyDB's "every node reads the same storage" design decoupled read scaling from the primary, keeping replication lag low.
- 机制根因
自建 PG 的读扩展走流复制,主库承担"写 + 复制源"双职责,读节点越多主库压力越大——这是 PG 读扩展的经典天花板。AlloyDB 存储计算分离后,计算节点(含读池)都从同一份分布式存储读数据,读扩展不再经过主库复制,主库只负责写。这是"PG 前端 + 自研存储层"架构(与 Aurora 同路线)带来的直接红利;PG 线协议兼容则让迁移零应用改动。
Self-managed PG read scaling rides on streaming replication, where the primary wears two hats - writer and replication source - so more readers mean more primary load: the classic PG read-scaling ceiling. With AlloyDB's disaggregated storage, compute nodes (including read pools) all read from the same distributed store, so read scale-out bypasses primary-side replication entirely and the primary only handles writes. That is the direct dividend of the "PG frontend + custom storage layer" architecture (the same route as Aurora); PG wire-protocol compatibility made the migration application-transparent.
- 教训
主库身兼"写 + 复制源"是 PG 架构读扩展的第一瓶颈,压测时要同时压"写涨 + 读节点涨"双维度;季节性业务的 deadline 是硬性的——把"收获季/播种季"写进迁移排期,PG 兼容带来的零改动迁移是赶 deadline 的关键;上云后第一件事仍是重新核定规格(拜耳用更小实例跑出更好成绩,超配可以直接砍掉)。
A primary doubling as "writer + replication source" is the first read-scaling bottleneck in PG architectures - load-test "rising writes AND rising readers" together. Seasonal businesses have hard deadlines: write "harvest/planting season" into the migration schedule, and PG-compatible zero-change migration is what makes such deadlines reachable. After landing, still re-size first (Bayer got better results on smaller instances - over-provisioning can be cut outright).
相关产品:Google AlloyDB(AlloyDB for PostgreSQL)、PostgreSQL(社区版)、Google BigQuery 相关能力:HN 祛魅 —— "它就是 Aurora 路线:PG 前端 + 自研存储层" 最后核验:2026-10-02
Character.AI:迁至 AlloyDB 后查询量 5 倍、延迟减半,Spanner 扛下海量摄入 成功经验
生成式 AI
读扩展
爆发式增长
- 场景
Character.AI 生成式 AI 平台用户与流量爆发式增长,平台的可扩展性与可靠性承压。AI 对话类应用的查询模式是"高并发、延迟敏感",传统单体 PG 架构下查询量与延迟难以兼得。
Character.AI's generative AI platform saw explosive user and traffic growth, straining platform scalability and reliability. Conversational AI workloads are high-concurrency and latency-sensitive - hard to scale on a classic single-node PG architecture without latency blowing up.
- 决策
把在线查询库迁移到 AlloyDB;同时用 Cloud Spanner 承担每天 TB 级的数据摄入,走"AlloyDB 服务查询、Spanner 扛摄入"的分工。配套使用 Cloud Run、BigQuery、GKE、TPU/GPU 等 Google Cloud 服务。
Migrate the online query database to AlloyDB, while Cloud Spanner takes on terabytes of daily data ingestion - a deliberate split of "AlloyDB serves queries, Spanner absorbs ingestion", alongside Cloud Run, BigQuery, GKE, and TPUs/GPUs across Google Cloud.
- 结果
迁移到 AlloyDB 后,查询量达到原来的 5 倍,查询延迟减半(Google Cloud 官方 YouTube 频道发布,属厂商渠道口径,引用须注明)。Spanner 每天可靠摄入 TB 级数据。
After migrating to AlloyDB, Character.AI serves five times the query volume at half the query latency (published on Google Cloud's official YouTube channel - a vendor-channel account; cite with that caveat). Spanner ingests terabytes of data every day with reliability and global scale.
- 机制根因
AlloyDB 的读扩展不走主库复制而走共享存储,读池可横向加节点扛查询洪峰;PG 兼容让迁移平滑。写入侧则交给为全球一致与海量写入设计的 Spanner——"查询与摄入解耦"是这套架构成立的关键,各自用最擅长的引擎。
AlloyDB read scale-out goes through shared storage rather than primary-side replication, so read pools can add nodes horizontally to absorb query floods; PG compatibility smoothed the migration. The write side was handed to Spanner, designed for global consistency and massive write throughput - "decoupling queries from ingestion" is what makes this architecture work, with each engine doing what it does best.
- 教训
AI 应用的流量是脉冲式的,选型先看"读扩展是否经过主库";单一数据库包打天下不如按读写特征分工(查询库 + 摄入库);本案例数字来自厂商渠道,生产选型时应以自己的压测为准。
AI application traffic is bursty - when evaluating, first ask whether read scale-out passes through the primary. One database for everything loses to splitting by read/write characteristics (a query store plus an ingestion store). The figures here come from a vendor channel - run your own load tests before betting production on them.
相关产品:Google AlloyDB(AlloyDB for PostgreSQL)、Google Spanner 相关能力:HN 祛魅 —— "它就是 Aurora 路线:PG 前端 + 自研存储层" 最后核验:2026-10-02
CME:全球最大衍生品交易所从 Oracle 转向 AlloyDB,Omni 做混合云垫脚石 成功经验
去 O 迁移
混合云
金融合规
- 场景
CME Group(芝加哥商品交易所集团)是全球领先的衍生品交易所,客户交易期货与期权,覆盖几乎所有可投资资产类别:交易吞吐极高,对正常运行时间与可用性的要求极为严苛。核心系统长期跑在传统专有数据库(Oracle)上,许可成本与技术债双重压力。
CME Group, the world's leading derivatives marketplace (futures and options across nearly every investable asset class), runs at extremely high transaction throughput with stringent uptime and availability demands. Core systems long ran on traditional proprietary databases (Oracle), under the dual pressure of licensing costs and technical debt.
- 决策
把"最严苛的企业级负载"交给 AlloyDB,从 Oracle 向 AlloyDB 迁移数个数据库(2023 年 10 月 Google Cloud Next 公布时迁移进行中);同时用可下载的 AlloyDB Omni 在本地先做现代化改造——不把客户承诺置于风险中,云之旅分步走。
Entrust its "most demanding enterprise workloads" to AlloyDB, migrating several databases from Oracle to AlloyDB (migration in progress when announced at Google Cloud Next, October 2023); use downloadable AlloyDB Omni to modernize on-premises first - never putting customer commitments at risk, taking the cloud journey in steps.
- 结果
CIO Sunil Cutinho 公开背书:"AlloyDB 给了我们需要的性能与扩展性;AlloyDB Omni 让我们在维持客户承诺、支撑最关键传统负载的同时,开始向 Google Cloud 现代化与迁移。"(Google Cloud 官方博客发布,属厂商渠道口径;迁移为进行中状态、非完成态,引用须注明。)
Public endorsement from CIO Sunil Cutinho: "AlloyDB gives us the performance and scalability we need, and AlloyDB Omni enables us to begin modernizing and migrating to Google Cloud, while maintaining our customer commitments and supporting our most mission critical traditional legacy workloads." (Published via the Google Cloud official blog - a vendor-channel account; the migration was in progress, not complete - cite with that caveat.)
- 机制根因
Omni 与云端 AlloyDB 是同一引擎的可下载版,本地现代化改造的应用"零改动"即可跑在云端 AlloyDB 上;PG 兼容大幅降低去 O 的 SQL/存储过程改写成本。强监管下"数据不能离境"不是不现代化的理由——先在本地换引擎,再迁云。
Omni is a downloadable build of the same engine as cloud AlloyDB, so applications modernized on-premises run unchanged on cloud AlloyDB later; PG compatibility sharply lowers the SQL/stored-procedure rewrite cost of an Oracle exit. Under strict regulation, "data cannot leave the country" is not an excuse to skip modernization - swap the engine locally first, move to the cloud second.
- 教训
金融级去 O 迁移的正确姿势是分步:Omni 做本地 in-place 现代化(合规边界内),云端 AlloyDB 做终态;选型时区分"引擎能力"与"部署形态",Omni 脱离 GCP 生态后差异化会减弱(见相关能力卡),终态仍应锚定云端;进行中的迁移案例引用时必须标注状态,不可写成"已完成"。
The right posture for financial-grade Oracle exits is stepwise: Omni for in-place on-premises modernization (inside the compliance boundary), cloud AlloyDB as the end state. Separate "engine capability" from "deployment shape" when evaluating - Omni's differentiation thins outside the GCP ecosystem (see the related capability card), so anchor the end state on the cloud service. Always mark in-progress migrations as such; never write them up as completed.
相关产品:Google AlloyDB(AlloyDB for PostgreSQL)、Oracle Database(甲骨文) 相关能力:HN 祛魅 —— "它就是 Aurora 路线:PG 前端 + 自研存储层" 最后核验:2026-10-02
Endear:Cloud SQL 长大后扛不住,迁 AlloyDB 一年实现连接数与 TPS 双 6 倍 成功经验
同构迁移
零停机切换
连接数扩展
- 场景
Endear 是全渠道零售 CRM 平台,聚合电商平台、POS、营销、客服、忠诚度等多源客户数据,做实时个性化与客户洞察。早期选 Cloud SQL(贴合当时的工作负载与预算),业务长大后"长大了穿不下":CPU、内存、连接数相继耗尽。选型评估了 CockroachDB(可扩展但价格结构长期更贵)、Spanner(强一致与水平扩展很强)、Bigtable(分析吞吐高),最终 AlloyDB 在性能、成本、灵活性、运维简单度上综合最优。
Endear is an omnichannel retail CRM platform aggregating customer data from ecommerce platforms, point-of-sale systems, marketing, customer service, and loyalty silos for real-time personalization and customer insights. It started on Cloud SQL (matching early workload and budget), then outgrew it: CPU, memory, and connection exhaustion arrived in sequence. The evaluation shortlist included CockroachDB (scalable but its pricing structure would cost more over time), Spanner (strong consistency and horizontal scale), and Bigtable (high analytics throughput); AlloyDB won on the blend of performance, cost, flexibility, and operational simplicity.
- 决策
Google Cloud Database Migration Service 做 Cloud SQL 到 AlloyDB 的持续复制(在 Cloud SQL 建复制槽、实时同步防丢数据);PgBouncer 做连接层,切换时只换密钥管理中的连接串,应用零停机;上线前做足性能压测验证新库扛得住。
Google Cloud Database Migration Service for continuous Cloud SQL-to-AlloyDB replication (a replication slot on Cloud SQL, real-time sync to prevent data loss); PgBouncer at the connection layer so cutover meant swapping connection strings in the secrets manager with zero application downtime; thorough performance testing before going live.
- 结果
一年后:连接数 6 倍(从单个 PgBouncer 池约 400–500 连接,涨到三个池、平均每池近 1000 连接)、TPS 6 倍(从约 1500 涨到主库峰值 5000、读集群 10000)、P99 聚合查询延迟低于 10 毫秒(Endear 团队口述,经 Google Cloud 官方博客发布,属厂商渠道口径,引用须注明)。
One year later: 6x connections (from one PgBouncer pool serving roughly 400-500 connections to three pools averaging just under 1,000 each), 6x transactions per second (from about 1,500 to a 5,000 TPS peak on the primary cluster and 10,000 TPS on the read cluster), and P99 aggregated query latency under 10ms (per the Endear team, published via the Google Cloud official blog - a vendor-channel account; cite with that caveat).
- 机制根因
DMS 基于 PG 原生复制的持续同步是"零停机"的底气——先实时跟上、再一次性切换;PgBouncer 把"换数据库"与"改应用配置"解耦,切换的是连接串而非应用;AlloyDB 读集群独立扩展后,读流量不再挤占主库。连接数先爆、TPS 后爆是 CRM 类聚合写入场景的典型顺序——连接池先于计算成为瓶颈。
DMS continuous replication on native PG replication is what makes "zero downtime" credible - sync in real time first, flip once. PgBouncer decouples "changing the database" from "changing application config" - the connection string moves, not the app. With AlloyDB's independently scalable read cluster, read traffic no longer crowds the primary. Connections blowing up before TPS is the typical order in CRM-style aggregated-write workloads - the pool becomes the bottleneck before compute does.
- 教训
先爆的往往是连接数不是 TPS,压测要压连接;PgBouncer + 密钥切换是 PG 系零停机迁移的标配,值得提前预埋;选型时把"三年后的价格曲线"算进去——CockroachDB 就是输在长期成本结构上;小团队别为"未来可能需要的全球一致"提前买 Spanner,为真实负载选型。
Connections usually exhaust before TPS does - load-test connections, not just throughput. PgBouncer plus secret-swapping is standard equipment for zero-downtime PG-family migrations; plant it early. Price the "year-three cost curve" into selection - CockroachDB lost on long-term cost structure. Small teams should not pre-buy Spanner for "global consistency they might need someday" - select for the real workload.
相关产品:Google AlloyDB(AlloyDB for PostgreSQL)、PostgreSQL(社区版)、CockroachDB、Google Spanner 相关能力:HN 祛魅 —— "它就是 Aurora 路线:PG 前端 + 自研存储层" 最后核验:2026-10-02
Lucius AI:一人公司用 AlloyDB 一库三用,ScaNN 索引让语义搜索快 47 倍 成功经验
向量检索
AI 运维
一人公司
- 场景
Lucius AI 是招标情报创业公司,平台覆盖英国、欧盟、印度、澳大利亚等五大洲 21 万+条招标信息,每晚从 13 个公共采购源摄入数据;两个生产区(欧洲、澳洲各一个 AlloyDB 集群,澳洲集群用客户管理加密密钥服务国防相关客户)。整个公司只有一个运营者,没有专职 DBA 与数据工程团队。关系型招标目录、审计日志、向量 embedding 原本需要关系库、向量库、日志库三套系统。
Lucius AI is a tender-intelligence startup covering 210,000+ tenders across the UK, EU, India, Australia and beyond, ingesting nightly from thirteen public procurement sources; two production regions (Europe and Australia, each on its own AlloyDB cluster, the Australian one using customer-managed encryption keys for defense-adjacent customers). The entire company has a single operator - no dedicated DBA or data engineering team. The relational tender catalog, audit logs, and vector embeddings would normally demand three systems: a relational store, a vector store, and a log store.
- 决策
AlloyDB for PostgreSQL 一库三用——关系型目录、文档元数据、审计日志、向量 embedding 全放同一引擎;语义搜索从"无索引暴力向量比对"迁到 ScaNN 索引;用开源 MCP Toolbox for Databases(alloydb-postgres 预置服务)把 AI agent 接进来做日常运维,权限严格最小化。
One AlloyDB for PostgreSQL for all three - relational catalog, document metadata, audit logs, and vector embeddings in the same engine; migrate semantic search from unindexed brute-force vector comparison to a ScaNN index; connect an AI agent via the open-source MCP Toolbox for Databases (prebuilt alloydb-postgres server) for day-to-day operations under strict least-privilege permissions.
- 结果
代表性生产查询从 1.14 秒降到 24 毫秒,语义搜索快 47 倍(Google Cloud 官方博客发布,属厂商渠道口径,引用须注明);语义索引重建:11.58 万条记录用 Gemini embedding 模型 10.6 分钟完成,API 花费约 3 美元,此后由 AlloyDB 自动 embedding 保持向量新鲜;库内 ai.rank 重排平均延迟 77 毫秒,无需独立重排微服务;AI agent 在最小权限下承担查询分析、数据新鲜度检查、事故取证(一次外部安全探测后,agent 几分钟内从审计日志还原出请求时间线)。
A representative production query dropped from 1.14 seconds to 24 milliseconds - semantic search 47x faster (published via the Google Cloud official blog - a vendor-channel account; cite with that caveat). Semantic index rebuild: 115,820 records embedded with the Gemini embedding model in 10.6 minutes for around three dollars in API spend, with AlloyDB auto embeddings keeping vectors fresh afterward. In-database retrieval reranking via the ai.rank function averages 77ms with no standalone reranking microservice. The AI agent, under least privilege, handles query analysis, data freshness checks, and incident forensics (after one external security probe, the agent reconstructed the request timeline from audit logs in minutes).
- 机制根因
ScaNN 是 Google 自研近似最近邻检索(Search/YouTube 同款技术),相对暴力比对的加速可达两个数量级(近似检索相对精确搜索的典型量级,取决于召回率要求);向量与业务数据同库带来统一备份计划与统一身份管理;MCP 侧 agent 使用专用 PG 角色(全库 SELECT + 单运维表 UPDATE,DROP/DELETE/TRUNCATE 直接禁用),把"AI 操作数据库"的爆炸半径关进笼子。另已启用列存引擎自动列存:40 个高频查询列一天内进入内存,报表查询加速而无需第二套分析库。
ScaNN is Google's in-house approximate nearest-neighbor retrieval (the same technology behind Search and YouTube), two orders of magnitude faster than brute-force comparison at comparable recall (the typical magnitude of approximate versus exact search, depending on recall requirements). Vectors living alongside business data unify backup schedules and identity management. On the MCP side, the agent uses a dedicated PG role (SELECT across the schema plus UPDATE on a single operational table; DROP, DELETE, TRUNCATE disabled outright), caging the blast radius of "AI operating the database". The columnar engine with auto-columnarization is also enabled - 40 frequently queried columns across four tables landed in memory within a day, accelerating reporting queries with no second analytical store.
- 教训
向量库不是必选项——向量与业务数据同库能省掉一套系统的运维、备份与身份管理;一人公司可以用 AI agent 当 DBA,但前提是最小权限 + 破坏性操作只留给人类;语义搜索"先跑通(暴力比对)、再建索引"是务实路径,索引建议本身都可以由 agent 在自动化性能审计里提出来。
A dedicated vector database is not mandatory - keeping vectors with business data saves operating, backing up, and securing a whole extra system. A one-person company can use an AI agent as its DBA, provided permissions are least-privilege and destructive operations stay human-only. "Get semantic search working first (brute force), then index" is the pragmatic path - the index recommendation itself can come out of the agent's automated performance audit.
相关产品:Google AlloyDB(AlloyDB for PostgreSQL) 相关能力:Vertex AI + ScaNN 向量 —— Google AI 生态红利 最后核验:2026-10-02
Regnology:合规报告 chatbot 用 AlloyDB 做动态向量库 成功经验
RAG
企业知识库
合规
- 场景
Regnology(监管科技公司)的合规报告 chatbot:需要理解复杂的监管术语、回答多样化的监管报告问题;grounding 数据是法规指南库、合规文档与历史报告数据。金融合规场景对"回答可溯源"要求高,向量检索的召回质量直接决定 chatbot 可用性。
Regnology (a regulatory technology company) built a regulatory-reporting chatbot that must understand complex regulatory terminology and answer diverse reporting questions; grounding data spans repositories of regulatory guidelines, compliance documents, and historical reporting data. Financial-compliance settings demand traceable answers, so vector retrieval recall quality directly determines whether the chatbot is usable.
- 决策
用 AlloyDB 做动态向量库,索引法规指南、合规文档与历史报告数据,把 RAG 的检索底座放在业务数据库里,而非另起专用向量库。
Use AlloyDB as a dynamic vector store indexing the regulatory guidelines, compliance documents, and historical reporting data - placing the RAG retrieval foundation inside the business database instead of standing up a dedicated vector store.
- 结果
合规分析师与报告专员以对话方式与 chatbot 交互,"节省时间、应对多样化的监管报告问题"(CIO Antoine Moreau 口述,经 Google Cloud 官方博客发布,属厂商渠道口径,引用须注明)。本案例公开细节较薄——仅一条 CIO 引言,无延迟/召回/规模数字,证据等级为"厂商渠道单引言",引用时不可脑补数字。
Compliance analysts and reporting specialists interact with the chatbot conversationally, "saving time and addressing diverse regulatory reporting questions" (per CIO Antoine Moreau, published via the Google Cloud official blog - a vendor-channel account; cite with that caveat). Public detail on this case is thin - a single CIO quote, no latency/recall/scale figures - so its evidence grade is "vendor-channel single quote"; do not invent numbers when citing it.
- 机制根因
AlloyDB AI 的向量能力(ScaNN 近似最近邻 + Vertex AI 闭环)让"embedding 生成—存储—检索"在 SQL 内闭环;向量与业务数据同库,合规文档的权限管控可以复用数据库已有的身份体系,这对金融合规是加分项。
AlloyDB AI's vector capabilities (ScaNN approximate nearest neighbor plus the Vertex AI loop) close "embedding generation, storage, and retrieval" inside SQL. Vectors living with business data lets compliance-document access control reuse the database's existing identity system - a plus in financial compliance.
- 教训
企业 RAG 不一定需要专用向量库——当召回延迟要求在几十毫秒量级、且文档权限需与业务数据一致时,同库向量是更省事的方案;但"一句话案例"的证据权重低,选型时应要求可复现的延迟/召回数字;金融合规场景下,向量库的权限模型与审计能力应与检索性能同等评估。
Enterprise RAG does not always need a dedicated vector store - when recall latency targets are in the tens of milliseconds and document permissions must match business data, in-database vectors are the lower-effort option. But a "one-quote case" carries low evidentiary weight; demand reproducible latency/recall figures before selecting. In financial compliance, a vector store's permission model and auditability deserve the same scrutiny as retrieval performance.
相关产品:Google AlloyDB(AlloyDB for PostgreSQL) 相关能力:Vertex AI + ScaNN 向量 —— Google AI 生态红利 最后核验:2026-10-02
SEEBURGER:集成平台 BIS 落子 AlloyDB,全球物流枢纽零代码改动上线 成功经验
SaaS 底座
零代码改动
全球部署
- 场景
SEEBURGER BIS Platform 是有 35 年历史的集成平台(iPaaS),提供 B2B/EDI、托管文件传输、应用集成、API 管理等能力,全球 14000+ 客户每天依赖它做集成,背后是数十亿美元的贸易额。其客户之一是服务 120 多个国家的全球物流与航运网络枢纽,对集成基础设施的可靠与可扩展要求极高。
SEEBURGER BIS Platform is a 35-year-old integration platform (iPaaS) offering B2B/EDI, managed file transfer, application integration, API management and more, with 14,000+ customers relying on it daily - billions of dollars in trade depend on its data infrastructure. One of its customers is a global logistics and shipping network hub serving over 120 countries, with extreme demands on integration reliability and scale.
- 决策
把 BIS 平台的云数据库选为 AlloyDB for PostgreSQL。五条选型标准:企业级特性(免运维杂活)、事务与分析混合负载能力、高可用、开发者效率(全托管 + 开放标准)、成本与部署灵活性(Omni 支持非 GCP 部署,避免锁定)。BIS 本就支持 PostgreSQL,AlloyDB 的 PG 兼容让采用零代码改动。
Select AlloyDB for PostgreSQL as the cloud database for the BIS Platform. Five selection criteria: enterprise-grade features (no operational toil), mixed transactional/analytical workload capability, high availability, developer productivity (fully managed + open standards), and cost/deployment flexibility (Omni supports non-GCP deployments, avoiding lock-in). BIS already supported PostgreSQL, so AlloyDB's PG compatibility meant zero code changes to adopt.
- 结果
"迁移到 AlloyDB 极其顺滑,我们一行代码都不用改,立刻看到了巨大的性能提升。它是我们最关键业务部署的完美平台。"——Ronald Heckendorff,SEEBURGER 集成架构总监(Google Cloud 官方博客发布,属厂商渠道口径,引用须注明)。物流枢纽客户的 iPaaS 方案上线后运营加速、全球服务提速,搭建与测试"没有任何问题或改动"。
"Migrating to AlloyDB was incredibly smooth. We didn't have to change any code, and immediately saw a huge performance improvement. It's the perfect platform for our most business-critical deployments." - Ronald Heckendorff, Director of Integration Architecture, SEEBURGER (published via the Google Cloud official blog - a vendor-channel account; cite with that caveat). The logistics hub's iPaaS rollout streamlined operations and accelerated global services; setup and testing completed "without any issues or changes".
- 机制根因
PG 线协议与 SQL 兼容是零代码改动的直接原因——应用层感知不到数据库换了;存储计算分离带来"更大部署的扩展性";Omni 的可下载形态给了 ISV"多云/本地都能交付"的灵活性,这是传统全托管云数据库给不了的。优化过的 vacuum、极低的复制延迟、自适应资源管理则砍掉了运维杂活。
PG wire-protocol and SQL compatibility is the direct reason for zero code changes - the application layer never noticed the database swap. Disaggregated compute and storage delivers "superior scalability for larger deployments"; Omni's downloadable form gives the ISV "deliver on any cloud or on-premises" flexibility that classic fully-managed cloud databases cannot. Optimized vacuuming, dramatically lower replication lag, and adaptive resource management removed the operational toil.
- 教训
ISV 把数据库作为产品底座时,"PG 兼容 = 零迁移成本"是最硬的选型论据之一;多云交付的 ISV 应把"能否脱离单一云部署"列为必选项(Omni 正是为此存在);平台厂商的证言比终端客户更稀缺——它证明的是"作为底座"的可靠性,而不只是一次迁移。
For ISVs building on a database, "PG compatibility = zero migration cost" is one of the hardest selection arguments. ISVs delivering multi-cloud should make "deployable off a single cloud" a mandatory criterion (Omni exists precisely for this). Testimony from a platform vendor is rarer than from an end customer - it proves reliability "as a foundation", not just a one-off migration.
相关产品:Google AlloyDB(AlloyDB for PostgreSQL)、PostgreSQL(社区版) 相关能力:HN 祛魅 —— "它就是 Aurora 路线:PG 前端 + 自研存储层" 最后核验:2026-10-02
Amazon 去 O:75PB、7500 个 Oracle 库的拆分 失败教训
成本敏感
去 O
高可用
多地域
- 决策
不做一对一替换,而是按访问模式拆分:DynamoDB 扛低延迟 KV、Aurora/RDS 扛事务、Redshift 扛分析,各团队按自身访问模式自选目标库。
Instead of a one-to-one replacement, the migration split by access pattern: DynamoDB for low-latency KV, Aurora/RDS for transactions, Redshift for analytics — each team choosing its own target database by its access pattern.
- 机制根因
单一巨型 RDBMS 承担所有访问模式时,license 成本与供应商锁定的运维代价被均摊到每笔交易上;purpose-built 拆分后,每种负载落到更匹配的架构(KV/文档/关系/列式各归其位),成本结构从根本上改变。
When a single giant RDBMS carries every access pattern, license costs and vendor-lock-in operational costs get amortized into every transaction; after a purpose-built split, each workload lands on a better-matched architecture (KV, document, relational, and columnar each in their rightful place), fundamentally changing the cost structure.
- 教训
去 O 的真正成本不是"换哪个兼容库",而是"一个数据库打天下"时 license + 锁定的隐性税;拆分后让一线团队按访问模式自选目标库,比顶层统一指定的成功率高。
The real cost of going off Oracle is not "which compatible database to swap in" but the hidden tax of license + lock-in from running "one database for everything"; after splitting, letting front-line teams pick their target database by access pattern succeeds more reliably than top-down mandates.
相关产品:Oracle Database(甲骨文)、Amazon DynamoDB、Amazon Aurora、Amazon Redshift 相关能力:授权 FUD —— "换到通用云要双倍 license"是人为商业壁垒、Serverless 零运维:流量不可预测时"先跑起来"的最短路径、日志即数据库(log is the database)—— 存储计算分离的"真"实现、RA3 存算分离 + 预留价:AWS 锁定团队的成本最优解 最后核验:2026-10-01
Apple 的 Cassandra 舰队(2021):16 万实例、100PB、1000+ 集群 成功经验
超大规模
写密集
多集群
去中心化
- 决策
把 Cassandra 作为大规模在线存储的默认选项长期投入;不追求少数巨型集群,而是拆成上千个中小集群,每个业务独立集群、独立爆炸半径。
Adopted Cassandra as the default choice for large-scale online storage with long-term investment; instead of a few giant clusters, split the fleet into thousands of small-to-medium clusters — one cluster per business, each with its own blast radius.
- 结果
按 Apache 官方公告口径(2021-07),160,000+ instances、100PB+ 数据、1000+ 集群,是已知全球最大 Cassandra 舰队。该数字适合说明部署规模,不适合推断单集群扩展能力。
Per the Apache project's official announcement (2021-07): 160,000+ instances, 100PB+ of data, 1,000+ clusters — the largest known Cassandra fleet in the world. The figure speaks to deployment scale; it should not be used to infer single-cluster scalability.
- 机制根因
无主对等架构 + LSM 追加写,让写吞吐随节点数接近线性增长(在该写密集小记录负载下;实际线性度取决于复制因子、一致性等级、网络拓扑与运维水平);tunable consistency 让每个业务按需在延迟与一致性之间选点;"多小集群"而非"少数大集群"的组织方式,把 gossip、repair、compaction 的运维复杂度锁在单个集群内部——规模化的真正敌人不是节点数,而是单个故障域的半径。
The masterless peer-to-peer architecture plus LSM append-only writes lets write throughput scale near-linearly with node count (under that write-heavy small-record workload; actual linearity depends on replication factor, consistency level, network topology, and operational maturity); tunable consistency lets each business pick its own point on the latency-consistency tradeoff; organizing as "many small clusters" instead of "a few large ones" confines gossip, repair, and compaction complexity inside each cluster — at this scale the real enemy is not node count but the radius of a single failure domain.
- 教训
写密集 + 可接受最终一致的场景,Cassandra 的扩展性在该场景成立,但"线性"二字要打折——取决于复制因子、一致性等级与运维;运维策略上要用"多小集群"控制爆炸半径,而不是把鸡蛋放在几个大集群里。选型时不仅要看单集群上限,还要看"舰队模式"下的运维成本曲线。
For write-heavy workloads that tolerate eventual consistency, Cassandra's scalability holds in this scenario — but discount the word "linear": it depends on replication factor, consistency level, and operations; operationally, prefer "many small clusters" to bound blast radius rather than putting everything in a few large clusters. When evaluating, look not only at single-cluster ceilings but at the ops-cost curve in "fleet mode."
相关产品:Apache Cassandra / ScyllaDB 相关能力:LSM 追加写:吃下"写多读少、只追加"的消息流 最后核验:2026-10-01
Capital One:PoC 验证 Aurora Global Database 故障切换后,仍为多活需求选了 DynamoDB(2020–2021) 失败教训
多活架构
选型决策
跨区容灾
反例
- 场景
Capital One 某团队的应用跑在美东/美西双区 active-active,服务层故障切换几乎瞬时,唯独 PostgreSQL 主库跨区切换要 10–15 分钟(把异区只读副本提升为主库再切流量)。现代应用期望零停机,数据库成了整条弹性链路上的短板。团队只考虑 AWS 托管方案(不想自己运维数据库),列出三个候选逐项对比。
A Capital One team's application ran active-active across AWS East and West regions: the service tier failed over near-instantly, but the PostgreSQL primary took 10-15 minutes to fail across regions (promote the remote read replica to primary, then reroute traffic). Modern applications expect zero downtime, and the database had become the weak link in an otherwise resilient chain. The team restricted itself to AWS managed options (no self-operated databases) and shortlisted three candidates for head-to-head comparison.
- 决策
① Aurora Global Database:优点是跨区故障切换 1 分钟内(AWS 口径 RPO 1 秒、RTO 小于 1 分钟),团队还专门做了 PoC,"切换效果极好",一度是首选;缺点是当时只支持 MySQL,与现有 PostgreSQL 应用有改造成本。② Aurora Multi-Master:多主 active-active,2019 年 8 月刚 GA,但仅限单区内多活,不满足跨区要求,出局。③ DynamoDB Global Tables:真正的跨区多活写,缺点是 NoSQL 数据模型要重写数据持久层、团队要重新学表设计。结论:只有 DynamoDB 满足"跨区多活写"这个硬需求,尽管它是迁移成本最高的选项。
(1) Aurora Global Database: one-minute cross-region failover (AWS's claim: 1-second RPO, sub-minute RTO); the team even ran a proof of concept and reported "excellent results" - it was briefly the favorite. Downside: at the time it was MySQL-only, while the application was PostgreSQL. (2) Aurora Multi-Master: multi-primary active-active, GA since August 2019, but confined to a single region - failed the cross-region requirement, eliminated. (3) DynamoDB Global Tables: genuine cross-region active-active writes; downside: the NoSQL data model required rewriting the persistence layer and relearning table design. Verdict: only DynamoDB satisfied the hard requirement of "writes accepted in every region," despite being the most expensive option to adopt.
- 结果
采用"双库并行"过渡:应用同时写 DynamoDB 和 PostgreSQL、读路径仍返回 PG 数据,生产跑数周零错误后彻底切掉 PG。据团队自述,应用层故障切换时间下降 99%,跨区故障切换的脚本和流程被彻底消除——"没有 failover,因为处处可写"。代价真实存在:表结构按访问模式重建、踩过"用 filter 当 query 使"导致全表扫描的坑、日期与空值类型全部手工处理。以上过程与数字引自 Capital One 技术博客作者 Kelly Jo Brown 的自述,属当事人一手记录。
A "dual-database" transition: the application wrote to both DynamoDB and PostgreSQL while reads still served PostgreSQL data; after several weeks in production with zero errors, PostgreSQL was cut off entirely. Per the team's own account, application failover time dropped 99% and the runbooks and scripts for regional database failovers were eliminated outright - "there are no failovers, because every region accepts writes." The costs were real: table structures rebuilt around access patterns, a full-table-scan incident from using a filter where a query belonged, and hand-rolled handling of dates and nulls. Process and figures come from Capital One tech blog author Kelly Jo Brown's first-party write-up.
- 机制根因
Aurora 的架构原语是"单写 + 存储层日志复制":Aurora Global Database 把跨区 RPO 做到秒级、RTO 压到 1 分钟内厂商口径,但写永远只有一个主区——它是"极快的故障切换",不是"多活"。Capital One 要的是任意区可写、故障时根本不需要切换动作,这超出了单写模型的表达能力。PoC 验证的是"切换够快",决策要的是"无需切换",两者差了一个架构维度——这就是 PoC 满分、选型仍被淘汰的原因。
Aurora's architectural primitive is "single writer plus storage-layer log replication": Aurora Global Database pushes cross-region RPO to seconds and RTO under a minute 厂商口径, but there is still exactly one writer region - it is "extremely fast failover," not "active-active." Capital One needed every region writable with no failover action at all, which is beyond what a single-writer model can express. The PoC validated "failover is fast enough"; the decision required "no failover needed" - a full architecture dimension apart. That is why a perfect PoC score still lost the selection.
- 教训
选型时把"failover 有多快"和"是否需要 failover"分成两个问题问,前者是优化项,后者是架构项,混在一起 PoC 会给出误导性高分;PoC 只能验证它测了的东西,验证不了它没测的语义(多活写语义);作者自己也承认 DynamoDB 不适合报表类负载——选多活的同时就要规划好分析链路的出路,不要等迁移完才发现。另注:文章发表后 Aurora Global Database 已补上 PostgreSQL 支持,但单写模型未变,本结论对"跨区多活写"需求依然成立。
In selection, ask "how fast is failover" and "is failover even needed" as two separate questions - the first is an optimization, the second is architecture; conflating them lets a PoC produce a misleadingly high score. A PoC can only validate what it measures (failover speed), never the semantics it never tested (active-active write semantics). The author herself concedes DynamoDB is a poor fit for reporting workloads - plan the analytics path's exit at selection time, not after migration. Note: Aurora Global Database has since added PostgreSQL support, but the single-writer model is unchanged, so this conclusion still holds for cross-region active-active write requirements.
相关产品:Amazon Aurora、Amazon DynamoDB 相关能力:日志即数据库(log is the database)—— 存储计算分离的"真"实现 最后核验:2026-10-02
道琼斯:行情数据平台从 SQL Server 迁至 Aurora,一个周六 8 小时完成切换 成功经验
异构迁移
零停机切换
全球多读副本
成本优化
- 场景
道琼斯行情数据平台(WSJ.com、Factiva、MarketWatch、Barron's 背后的行情源)已运行 20 年:本地 16 台数据库服务器(4 个主库跨机房镜像,其余做分发与订阅),整个平台约 200 台服务器分处两个机房。从 5 家数据商拉取行情,每秒处理数万条消息,大行情时流量暴涨 3–4 倍。迁移动因有二:砍掉昂贵的 SQL Server 许可与本地机房;合同变化只给团队约 1 年时间,而按历史经验这类迁移至少要 2 年外加团队翻倍。
Dow Jones's market data platform (the quote feed behind WSJ.com, Factiva, MarketWatch, and Barron's) had run for 20 years: 16 on-premise database servers (four primaries mirrored across datacenters, the rest split into distributors and subscribers), roughly 200 servers total across two datacenters. It pulled market data from five providers, processing tens of thousands of messages per second, with traffic spiking 3-4x during major market events. Two migration drivers: kill expensive SQL Server licensing and on-premise tin; contract changes gave the team about one year, while their experience said this class of migration needed at least two years plus a doubled team.
- 决策
用 AWS Schema Conversion Tool + Database Migration Service 把数据从本地 SQL Server 持续同步到 Aurora MySQL;迁移期间在客户端与 MarketData 服务之间架设 NGINX 代理层,请求先走旧链路,再整体切到云端 API。选定某一个周六:部署代理、更新 DNS、把全部客户端流量迁到 Aurora MySQL,全程 8 小时;两周后正式下线本地 SQL Server。目标架构是 Aurora Global 集群——弗吉尼亚 1 写 5 读、俄亥俄 6 读,写节点具备跨区 failover 能力,读副本尽量靠近数据商。
AWS Schema Conversion Tool plus Database Migration Service continuously replicated data from on-premise SQL Server to Aurora MySQL; during the migration an NGINX proxy tier sat between data clients and MarketData services, with requests initially flowing through the legacy path before flipping to the cloud APIs. On one chosen Saturday the team deployed the proxies, updated DNS, and moved all client traffic to Aurora MySQL within 8 hours; the on-premise SQL Server was shut down two weeks later. The target architecture was an Aurora Global cluster - one writer and five readers in Virginia, six readers in Ohio, close to the data providers, with writer failover to the opposite region.
- 结果
比计划提前 12 天交付,客户零感知、行情数据零差错(据工程副总裁 Mona Soni 口述)。综合硬件、许可、维护、机房空间与电力,成本下降超 50%;迁移后用 CloudWatch 复盘发现实例只用了 3% CPU、40% 内存——本地机房时代的超配习惯被原样搬上了云,随即缩容,"周支出又降了超 50%"(据工程经理 Luke Sawatsky 口述)。以上数字均引自 AWS 官方博客记述的客户口述,属厂商渠道口径,引用须注明。
Delivered 12 days ahead of schedule with zero customer disruption and zero data-accuracy issues (per VP Engineering Mona Soni). Counting hardware, licensing, maintenance, rack space, and power, costs fell by more than 50%; after the move, CloudWatch showed the instances using just 3% CPU and 40% of memory - the on-premise over-provisioning habit had been lifted onto the cloud unchanged - so the team downsized and "reduced weekly spend by more than 50%" again (per engineering manager Luke Sawatsky). All figures are customer statements reported via the AWS official blog - a vendor-channel account; cite with that caveat.
- 机制根因
Aurora Global Database 的跨区复制发生在存储层(redo 日志流)而非逻辑复制,因此"一写 + 十一读"横跨两个区的低延迟读扩展才成立,写 failover 也不需要重放日志追数据。8 小时切换可行的关键不是 DMS 同步有多快,而是 NGINX 代理层把"换数据库"与"客户端改 endpoint"解耦——切换的是流量路径而非数千个客户端配置。代价:依然是单写模型,写节点故障切换是分钟级而非秒级;且迁移后第一件事就该是重新核定实例规格。
Aurora Global Database replicates across regions at the storage layer (redo log stream), not via logical replication - which is what makes "one writer plus eleven readers" across two regions with low-latency reads viable, and writer failover needs no log replay to catch up. The 8-hour cutover was feasible not because DMS sync is fast, but because the NGINX proxy tier decoupled "changing the database" from "changing every client's endpoint" - the traffic path moved, not thousands of client configs. The price: still a single-writer model, so writer failover is measured in minutes, not seconds; and the first post-migration job should always be re-sizing instances.
- 教训
异构迁移最难的不是数据同步,而是"客户端不用动"的切换——代理/网关层是零停机迁移的标配,值得提前数月预埋;上云后第一件事是用监控重新核定规格,本地机房的超配习惯会直接变成云账单;把"许可到期日"写进迁移排期,它是比技术更硬的 deadline,道琼斯正是被它逼出了 8 小时切换方案。
The hardest part of a heterogeneous migration is not data sync but a cutover where "clients change nothing" - a proxy/gateway tier is standard equipment for zero-downtime migration and worth planting months ahead. The first thing to do after landing in the cloud is re-size with monitoring; on-premise over-provisioning habits convert directly into cloud bills. Write the "license expiry date" into the migration schedule - it is a harder deadline than any technical one, and at Dow Jones it is exactly what forced the 8-hour cutover design.
相关产品:Amazon Aurora、Microsoft SQL Server 相关能力:日志即数据库(log is the database)—— 存储计算分离的"真"实现 最后核验:2026-10-02
Instant:Aurora Postgres 大版本升级(13→16),4 人团队做出零停机切换(2024–2025) 成功经验
大版本升级
零停机
逻辑复制
踩坑复盘
- 场景
Instant(自称"现代 Firebase",开箱即用的实时后端)生产跑在单个 Aurora Postgres 实例(PG 13)上:数据不足 1TB,读约 180 万 tuples/秒、写约 500 tuples/秒;同步服务器监听 PG 的 WAL,把变更实时推送给浏览器客户端。2024 年 8 月开源后流量涨约 100 倍,12 月 Aurora CPU 开始打满,团队被迫一路把实例升到 db.r6g.16xlarge。复现慢查询时意外发现:同一查询在 PG 16 上比 PG 13 快 30% 以上——升级本身就是一轮性能优化。
Instant (self-described "modern Firebase," a realtime backend-as-a-service) ran production on a single Aurora Postgres instance (PG 13): under 1 TB of data, about 1.8 million tuples read per second and 500 tuples written per second; sync servers tailed the Postgres WAL to push changes to browser clients in realtime. After open-sourcing in August 2024, throughput grew roughly 100x, and by December the Aurora instance's CPU started spiking, forcing upgrades all the way to db.r6g.16xlarge. While reproducing slow queries the team stumbled on something: the same queries ran 30%+ faster on PG 16 than on PG 13 - the upgrade itself was a performance optimization.
- 决策
目标 PG 13→16、零停机,4 人团队列出试错清单逐个验证。① 原地升级:在克隆库实测约 15 分钟不可用(Lyft 称他们的 30TB 库原地升级要 30 分钟),出局。② Aurora 蓝绿部署:官方承诺约 1 分钟停机,实测却建失败——因为主库上有同步服务器监听 WAL 的活跃 replication slot,而 AWS 文档根本没提这个限制,出局。③ 照抄 Lyft 的"克隆→升级→复制":自定义函数在复制时因 search_path 找不到(修复:函数定义显式加 public. 前缀),更糟的是抽查发现丢了 13 条事务,出局。最终方案:全新起一个 PG 16 的 Aurora 库,用原生逻辑复制从零同步(publication + subscription,copy_data=true),7 步清单。
Target PG 13 to 16 with zero downtime; the 4-person team worked through a checklist, testing each option. (1) In-place upgrade: measured ~15 minutes of unavailability on a cloned database (Lyft reported 30 minutes for their 30 TB database) - rejected. (2) Aurora blue-green deployment: AWS promised about a minute of downtime, but creation failed in testing because the primary had active replication slots (used by the sync servers tailing the WAL) - a restriction the AWS docs never mentioned - rejected. (3) Lyft's "clone, upgrade, replicate" recipe: custom functions broke during replication because of search_path (fix: prefix function definitions explicitly with public.), and worse, a spot check found 13 missing transactions - rejected. Final approach: stand up a fresh PG 16 Aurora database and sync from scratch with native logical replication (publication + subscription, copy_data=true), in 7 steps.
- 结果
切换写流量用自研 failover 函数:暂停新事务→等 2.5 秒让存量事务完成→cancel 剩余长事务→用不可变 transactions 表(只插不改)对账确认目标库追平→setval 把序列拨到 max(id)+1000(逻辑复制不复制序列,不处理会主键冲突)→放行新连接指向新库。生产执行约 3.5 秒暂停,用户无感知。完整复盘于 2025-01-29 公开发布。以上数字与过程均引自 Instant 自家公开复盘,属当事人自述口径。
Write traffic was switched with a hand-rolled failover function: pause new transactions, wait 2.5 seconds for in-flight transactions to finish, cancel the stragglers, verify the target has caught up by reconciling against an immutable transactions table (insert-only), setval sequences to max(id)+1000 (logical replication does not replicate sequences - skip this and you get primary-key collisions), then release new connections onto the new database. Production execution paused for about 3.5 seconds; users noticed nothing. The full postmortem was published 2025-01-29. All figures and the process come from Instant's own public postmortem - a first-party account.
- 机制根因
Aurora 的托管式升级路径(原地/蓝绿)都假设"标准用法";Instant 用 WAL + replication slot 做实时同步,是合理但非标的架构,正好撞上蓝绿部署的隐性前置条件。逻辑复制跨大版本可行,但有三处暗坑:自定义函数 search_path、序列不复制、克隆升级路径可能静默丢数据——都必须用"不可变事务表 max(id) 对账"这类业务层校验兜底,而不能信任"复制状态正常"。零停机的本质是把连接控制权从控制台收回应用层:缩容到一台大同步服务器、用代码而非按钮做切换,暂停窗口才敢压到秒级。
Aurora's managed upgrade paths (in-place, blue-green) assume "standard usage"; Instant's WAL-tailing sync architecture is reasonable but non-standard, and it collided head-on with blue-green's undocumented precondition. Logical replication works across major versions but hides three traps: custom-function search_path, non-replicated sequences, and silent data loss in the clone-upgrade path - all of which demand application-level verification (reconciling max(id) against an immutable transaction table) rather than trusting "replication status: healthy." Zero downtime's essence was pulling connection control back from the console into the application: consolidate onto one big sync server and switch with code, not buttons, so the pause window can be squeezed to seconds.
- 教训
永远先做与生产等价的完整彩排(含活跃 replication slot 和真实客户端),只在控制台点按钮的 happy path 测试等于没测;大版本升级前先把慢查询在新版本上回放一遍,升级本身可能是性价比最高的优化;"零停机"有前提—— modest 规模 + 能收敛连接,小团队别照抄大厂(Lyft 30TB 用 Aurora 秒级克隆)的方案,大厂也别照抄小团队的手工 failover 算法,规模决定路径。
Always rehearse the full procedure against a production-equivalent environment - including active replication slots and real clients; testing only the console happy path is no test at all. Before a major version upgrade, replay slow queries on the new version first - the upgrade itself may be the highest-ROI optimization available. "Zero downtime" has preconditions: modest scale plus the ability to converge connections. Small teams should not copy big-company recipes (Lyft's 30 TB database used Aurora's near-instant cloning), and big companies should not copy small-team hand-rolled failover algorithms - scale determines the path.
相关产品:Amazon Aurora、PostgreSQL(社区版) 相关能力:— 最后核验:2026-10-02
三星电子:Samsung Account 11 亿用户从 Oracle 迁到 Aurora PostgreSQL(2018–2020) 成功经验
去 O 迁移
超大规模账户体系
多区域部署
成本优化
- 场景
Samsung Account 是三星设备与服务的统一身份入口(Bixby、SmartThings、Samsung Pay 都经它登录),用户总数 11 亿、约 4 亿月活,峰值约 8 万请求/秒。原系统跑在 IDC(自有机房)托管的 Oracle 上,已运行约 10 年,是典型的单体架构。首席架构师 Salva Jung 直言 Oracle"没为微服务架构做好准备,价格也不合理",且在老架构上做无停机扩容"既危险又昂贵",新设备与新服务带来的流量迟早撑不住。
Samsung Account is the unified identity gateway for Samsung devices and services (Bixby, SmartThings, Samsung Pay all sign in through it): 1.1 billion users, roughly 400 million monthly active, peaking around 80,000 requests per second. The legacy system ran on Oracle hosted in Samsung's own IDC (internet data center), in place for about 10 years as a classic monolith. Principal architect Salva Jung said Oracle was "not ready for microservices architecture, nor did it have reasonable pricing for it," and that scaling the old system without downtime had become "risky and costly" - incoming traffic from new devices and services would eventually overwhelm it.
- 决策
选定 Aurora PostgreSQL 为迁移目标,核心动因是兼容度:85–90% 的 Oracle 查询在 Aurora 的 PostgreSQL 兼容层里无需改写,近 3000 条查询的转换"几乎自动完成"。迁移用 AWS DMS 做异构在线迁移,源库全程保持在线服务用户。2018 年 10 月从欧盟区开始,分三个区(欧盟、中国、美国)逐区推进,每区数据量 2–4TB;DMS 3–4 天即可复制完 2–3TB 数据,随后逐一切换用户流量。仅欧盟区就用约 22 周完成 4TB 数据迁移;三区分别于 2019 年 4 月(欧盟)、2019 年 10 月(中国)、2020 年 3 月(美国)完成,全程"几乎没有停机"。
Samsung chose Aurora PostgreSQL as the migration target, driven above all by compatibility: 85-90% of existing Oracle queries ran unmodified on Aurora's PostgreSQL compatibility layer, making the conversion of nearly 3,000 queries "practically automatic." AWS DMS performed the heterogeneous migration online while the source database kept serving users. Starting October 2018 in the EU region, the rollout proceeded region by region (EU, China, US), each holding 2-4 TB of data; DMS replicated 2-3 TB of user data in 3-4 days, then user traffic was cut over one segment at a time. The EU leg alone took about 22 weeks for 4 TB of data; the three regions completed in April 2019 (EU), October 2019 (China), and March 2020 (US) - with "minimal downtime" throughout.
- 结果
迁移后每区可平滑扩展到 15 个 Aurora 只读副本,90% 请求延迟低于 60ms,云上自动化让功能交付更快。据三星 DBA Byungyul Ko 口述,Aurora PostgreSQL 的月度运维成本比 Oracle 低 44%,这还不算省掉的 IDC 许可费和 Oracle 那边另计的 22% 维护费。以上数字引自 AWS 官方案例页的三星受访人口述——属于经由厂商渠道发布的客户口径,引用时须注明。单体数据库架构也被拆解为适配微服务的数据库布局。
Post-migration each region scales seamlessly to 15 Aurora read replicas, 90% of request latency sits under 60 ms, and cloud automation ships features faster. Per Samsung DBA Byungyul Ko, monthly operational costs with Aurora PostgreSQL run 44% lower than Oracle - before counting the eliminated IDC license fees and the separate 22% Oracle maintenance fees. These figures come from named Samsung interviewees quoted on AWS's official case study page - i.e., customer claims published through a vendor channel; cite them with that caveat. The monolithic database architecture was also decomposed into a microservices-friendly database layout.
- 机制根因
Aurora"日志即数据库"的存储计算分离是这场迁移的底座:读副本共享同一份分布式存储日志,扩副本不需要复制数据,故障切换只是元数据层面的写角色变更——这正是三星敢于承诺"无缝扩展到 15 个副本"且延迟可控的原因。同时 Aurora 的 PostgreSQL 协议兼容让 85–90% 的 Oracle SQL 零改写落地,把去 O 迁移里最贵的人力环节(SQL 改写)压缩到最小。代价是三区仍是三个独立集群而非真正的全球多写,跨区用户行为分析还得另建数据湖(三星后续规划)。
Aurora's "log is the database" disaggregated storage/compute architecture underpins this migration: read replicas share one distributed storage log, so scaling reads needs no data copying, and failover is a metadata-level writer-role change - the structural reason Samsung could promise "seamless scaling to 15 replicas" with controlled latency. Meanwhile Aurora's PostgreSQL protocol compatibility let 85-90% of Oracle SQL land with zero rewrites, compressing the most expensive labor in any Oracle exit (SQL conversion). The price: three independent regional clusters rather than true global multi-writer, so cross-region user analytics still needs a separate data lake (on Samsung's roadmap).
- 教训
去 O 迁移真正的成本大头不是数据搬运而是 SQL 改写量,目标库的协议/方言兼容度直接决定迁移是否可行;超大规模账户体系选型,先看"读扩展要不要复制数据",存储计算分离的库在这类场景有结构性优势;厂商案例页的数字可以用,但必须标注口径与说话人身份(本例是三星 DBA 经 AWS 案例页口述),不要写成独立第三方实测。
The real cost driver of an Oracle exit is not data movement but SQL rewrite volume - the target database's protocol/dialect compatibility decides whether the migration is even feasible. For hyperscale account systems, ask first "does read scale-out require copying data" - disaggregated-storage databases hold a structural edge there. Vendor case-study numbers are usable but must carry their caveat and speaker identity (here: a Samsung DBA quoted via the AWS case study page), never presented as independent third-party measurement.
相关产品:Amazon Aurora、Oracle Database(甲骨文) 相关能力:日志即数据库(log is the database)—— 存储计算分离的"真"实现 最后核验:2026-10-02
Gojek:自研 Firehose 把 600 个 Kafka topic 实时灌进 BigQuery(2021–) 成功经验
流式写入
Kafka 入仓
Schema 演进
开源自研
- 场景
Gojek 有 19+ 条产品线、数百个微服务、多个 Kafka 集群,新 topic 几乎隔天就增加一个。数仓建在 BigQuery 上,既要跑历史分析,又要做近实时报表。最初的做法是每个 Kafka topic 一个独立代码库往 BQ 推数:topic 一多、字段一变,就要人工改代码又改表,还出过几次数据丢失,只能人工从 GCS 回补。随着业务扩张到多国,这种"脚本堆"彻底管不过来了。
Gojek runs 19+ product lines, hundreds of microservices, and multiple Kafka clusters, with new topics appearing almost every other day. Its BigQuery warehouse had to serve both historical analytics and near-real-time reporting. The original approach was one bespoke codebase per Kafka topic pushing into BigQuery: every new topic or field change meant manual code and table updates, and there were several data-loss incidents that had to be backfilled by hand from GCS. As the business expanded across countries, the script pile became unmanageable.
- 决策
Gojek 从零自研了 Beast(后演进为开源项目 Firehose,现属 ODPF):单代码库消费任意 topic,用 proto descriptor 驱动、无需为新 topic 写代码;Java blocking queue 把消费、转换、推送、提交 offset 做成四级独立流水线,只提交已确认写入 BQ 的 offset(at-least-once 语义);跑在 Kubernetes 上,按 topic 分区数水平扩展。Firehose 还对接 Stencil(schema registry),topic 结构一变就自动更新 BQ 表结构,无需人工干预。
Gojek built Beast from scratch (later evolved into the open-source Firehose, now under ODPF): a single codebase consuming any topic, driven by proto descriptors with no per-topic code; Java blocking queues forming a four-stage pipeline (consume, transform, push, commit offsets) that only commits offsets for batches confirmed written to BigQuery (at-least-once semantics); Kubernetes deployment scaling horizontally with topic partition counts. Firehose also integrates with Stencil (a schema registry) so topic schema changes automatically update BigQuery tables with no human intervention.
- 结果
Firehose 把 600 个 Kafka topic 流入 BigQuery,另有 700 个入 GCS;BigQuery 侧日均 60 亿事件、10+ TB 新增数据,分析师在数据产生 5 分钟内即可查询(以上规模数字由 Google Cloud 博客引 Gojek 披露,属厂商侧口径)。schema 变更零人工,官方称"为开发者省下数百小时"厂商口径。Beast/Firehose 已开源,多家公司在用。
Firehose streams 600 Kafka topics into BigQuery (plus 700 into Cloud Storage): an average of 6 billion events and 10+ terabytes ingested into BigQuery daily, queryable by analysts within five minutes of production (scale figures via the Google Cloud blog citing Gojek's disclosure — vendor-side claim). Schema changes need zero manual work, reportedly "saving developers hundreds of hours" 厂商口径. Beast/Firehose are open source and used by other companies.
- 机制根因
流式链路的核心矛盾是交付语义:Beast 选择 at-least-once + offset 对账,把"精确一次"的去重责任留给下游——这与 BigQuery streaming insert 的 insertId 仅做 best-effort 去重是同一道取舍,没有银弹。单代码库消灭了"N 个 topic × M 次 schema 变更"的人力乘数效应;K8s 无状态部署让吞吐随业务脉冲弹性伸缩;schema 自动化是关键——此前的数据丢失事故根子都在"人工改表"这一步。
The core tension in streaming ingest is delivery semantics: Beast chose at-least-once plus offset reconciliation, leaving exactly-once deduplication to downstream consumers — the same trade-off as BigQuery streaming inserts' best-effort insertId dedup; there is no free lunch. A single codebase killed the "N topics times M schema changes" labor multiplier. Stateless Kubernetes deployment lets throughput scale elastically with business pulses. Schema automation was the crux — the earlier data-loss incidents all traced back to the "human edits the table" step.
- 教训
流式入仓先定义交付语义(at-least-once 还是 exactly-once),再选工具,顺序反了必踩去重坑;schema 治理必须自动化,人工改表是数据丢失之源;高频小 schema 变更场景下,"通用消费器+注册中心" beats "一 topic 一脚本";把内部工具开源(ODPF)能换来社区的长期维护力,比养着内部孤儿项目划算。
Define delivery semantics (at-least-once vs exactly-once) before choosing streaming tooling — reversed order guarantees dedup pain. Schema governance must be automated; manual table edits are a data-loss factory. For high-frequency small schema changes, "generic consumer plus registry" beats "one script per topic." Open-sourcing internal tools (ODPF) buys long-term community maintenance, cheaper than nursing an internal orphan project.
相关产品:Google BigQuery 相关能力:— 最后核验:2026-10-02
HTTP Archive:研究员在"免费"公开数据集上跑出 $14,000 账单(反例,2024) 失败教训
账单爆表
公开数据集
SDK 盲区
默认无熔断
- 场景
HTTP Archive 是追踪"Web 如何被建成"的公益项目,把爬取的海量页面数据放在 BigQuery 公开数据集上。用户 Tim 经 Python 脚本(官方 GCP 库)跑查询,收到 Google $14,000 账单。他在论坛发帖抗议:"这个网站让人以为 public 数据集是给社区用的,结果它是 Google Cloud 的印钞机,一眨眼就能亏掉 $14k。"
HTTP Archive is a public-interest project tracking "how the web is built," hosting its massive crawl dataset on BigQuery as a public dataset. A user, Tim, ran queries via a Python script (the official GCP libraries) and received a $14,000 bill from Google. His forum post protested: the site makes the public dataset look like a community resource, "but it is instead a for-profit money maker for Google Cloud and you can lose tens of thousands of dollars."
- 决策
事件在论坛发酵。维护者回应:99% 的用户只看免费月报/年报,BigQuery 是给 1% 需要原始数据的 power user 的;$14,000 对应约 2.5 PB 扫描量(按 $6.25/TiB,数字吻合);Web UI 跑查询会显示预估扫描量,但 Python 库没有成本提示机制。维护者道歉并承诺在 FAQ 加显式收费警告。Tim 则主张两点:默认开启 cost controls(现在默认是关的)、$5k 熔断器——超限必须手动确认才继续跑。
The thread escalated. A maintainer responded: 99 percent of users only read the free monthly reports and the annual Web Almanac; BigQuery exists for the 1 percent of power users who need raw access. The $14,000 corresponded to about 2.5 petabytes processed (consistent with $6.25/TiB). The web UI shows an estimated scan size before running, but the Python library has no cost-preview mechanism. The maintainer apologized and promised an explicit billing warning in the FAQ. Tim argued for two defaults: cost controls enabled out of the box (currently off by default) and a $5k circuit breaker requiring manual confirmation to exceed.
- 结果
The Register 称已联系 Google 置评,文中未见回应;退款或补偿无公开记录(查证为无,不编造)。社区吵成两派:一派认为"不懂数据量就跑查询是用户自己蠢"(该回帖后被管理员隐藏),另一派认为厂商应默认设限,尤其对学生和学术用户。
The Register reports having contacted Google for comment; no response appears in the piece, and there is no public record of a refund or credit (verified absent — not invented). The community split: one camp called the complainant reckless for running queries without understanding data volumes (that reply was hidden by moderators), the other argued vendors should default to limits, especially for students and academics.
- 机制根因
on-demand 按扫描字节计费与"公开数据集免费"的心理模型直接冲突——数据集免费下载/查看 ≠ 查询免费。SDK 路径缺了 Web UI 的"预估扫描量"护栏,这是最隐蔽的坑。BigQuery 默认没有项目级花费上限,cost controls 要手动开——等于把熔断器的钥匙交给了最可能忘记的人。列存下一次 SELECT * 或宽表全扫,极易触达 PB 级,$6.25/TiB 的单价下 2.5 PB 就是 $14k。
On-demand per-byte billing collides head-on with the "public dataset is free" mental model — the dataset is free to view, not free to query. The SDK path lacks the web UI's estimated-scan guardrail, the stealthiest trap of all. BigQuery has no project-level spending cap by default; cost controls must be switched on manually — handing the circuit-breaker key to the person most likely to forget it. On a columnar engine, one SELECT * or full wide-table scan easily reaches petabyte scale, and at $6.25/TiB, 2.5 PB is $14k.
- 教训
任何 BigQuery 项目第一天就开 custom quota / cost controls,不要赌记性;SDK 跑查询前先 dry run 估算 bytes,再决定跑不跑;碰公开数据集先看表大小,"免费数据"不等于"免费查询"——把这句话写进团队手册;给学术/学生用户搭环境时,默认配额要从紧。
Enable custom quotas and cost controls on day one of any BigQuery project — don't bet on memory. Before running SDK queries, dry-run to estimate bytes, then decide. With public datasets, check table sizes first: "free data" does not mean "free queries" — put that sentence in the team handbook. When provisioning environments for academic or student users, start with tight default quotas.
相关产品:Google BigQuery 相关能力:On-demand 的"扫描税" 最后核验:2026-10-02
Monzo:150+ 微服务事件灌进"一张巨型 BigQuery 表",银行数据团队奠基(2016–) 成功经验
事件源架构
银行数仓
选型实话
全托管
- 场景
Monzo 是 2015 年创立的全云银行:150+ 微服务跑在 Kubernetes,事务库是 AWS 上的 Cassandra——不像关系库能打快照供分析用。2016 年数据团队奠基时决定用 BigQuery 做数仓。架构是事件源式的:所有微服务、App、网站发出的事件先被记录下来,独立的 analytics 服务做 enrich 和脱敏,然后灌进"一张巨型的 BQ 表"——事件 payload 存在 JSON blob 列里,新事件类型零 schema 变更,连 Stripe、Intercom 的 webhook 事件都能直接塞进来。
Monzo is a cloud-born bank founded in 2015: 150+ microservices on Kubernetes, with Cassandra on AWS as the transactional database — no relational snapshots available for analytics. When the data team was founded in 2016, BigQuery became the warehouse. The architecture is event-sourced: every event emitted by microservices, apps, and the website is recorded; a separate analytics service enriches and sanitizes them, then loads them into "one gigantic BQ table" — event payloads live in a JSON blob column, so new event types need zero schema changes, and even Stripe and Intercom webhook events slot right in.
- 决策
三原则:autonomy(人人可访问脱敏数据)、全托管分析栈、自动化。选 BQ 的理由很实在:存海量数据便宜、streaming insert 自带基础去重(对比 Redshift 不用先落 S3/GCS,少一跳)、全托管、数据实时可用。但博客罕见地把丑话说在前面:小查询有 2–3 秒固定开销(Looker 多数查询 3–10 秒,交互探索"quite sluggish");当时只支持按 ingestion date 分区,大表读一小部分也贵;legacy SQL 方言缺函数,脚本被迫写两倍长。
Three principles: autonomy (everyone gets access to sanitized data), fully managed analytics infrastructure, automation. The case for BigQuery was pragmatic: cheap storage of massive data, streaming inserts with basic deduplication (unlike Redshift, no staging on S3/GCS first — one hop fewer), fully managed, data available in real time. Rarely for a vendor-era blog post, the downsides were stated up front: 2-3 seconds of fixed overhead on small queries (most Looker queries took 3-10 seconds, making interactive exploration "quite sluggish"); at the time only ingestion-date partitioning existed, so reading a slice of a huge table was still expensive; the legacy SQL dialect lacked functions, forcing scripts twice as long as they should be.
- 结果
这套架构支撑了全行的数据决策与风控机器学习(Dataflow 提特征、TensorFlow 训练欺诈模型);Google Cloud 客户页称 Monzo 在 BQ 里存了近 19 PB、2000+ 数据模型厂商口径。数据团队从第一人起步,靠"托管+自动化"长期保持精干。
The architecture underpins bank-wide data decisions and fraud ML (Dataflow for feature extraction, TensorFlow for the fraud model). Google's customer page credits Monzo with nearly 19 petabytes in BigQuery and 2,000+ data models 厂商口径. The data team started as a one-person operation and stayed lean through "managed plus automation."
- 机制根因
事件源 + 列存数仓是天作之合:JSON blob 列用"读时 schema"换 schema 灵活性,代价是扫描成本——这正是 Monzo 当年预警的"读的时候不小心会很贵"。单巨表是 schema-on-read 的极端实践,适合事件种类爆炸但查询模式收敛的场景。小查询慢是 serverless 调度开销的固有税:BQ 的甜蜜点是大扫描,交互式小查询天然错位,所以 Monzo 早早规划了亚秒级 serving 层。
Event sourcing plus a columnar warehouse is a natural fit: the JSON blob column trades scan cost for schema flexibility — exactly the "reading it gets expensive if you are careless" warning Monzo issued back in 2016. The single giant table is schema-on-read taken to the extreme, suited to exploding event variety with converging query patterns. Slow small queries are the inherent tax of serverless scheduling overhead: BigQuery's sweet spot is big scans, so interactive small queries are structurally mismatched — which is why Monzo planned a sub-second serving layer early.
- 教训
选型博客敢写缺点是难得的诚实——"存便宜、读贵"这句 2016 年的预警就是今天的"扫描税";OLTP 快照不可得时,事件流是唯一的真相源;全托管的代价是"小查询税",serving 层要提前规划而不是事后补;团队精干的关键不是人少,而是把运维税交出去(managed)+ 把重复劳动自动化掉。
An engineering blog brave enough to list the downsides is rare honesty — the 2016 "cheap to store, expensive to read" warning is today's "scan tax." When OLTP snapshots are unavailable, the event stream is the only source of truth. Fully managed comes with a "small-query tax," so plan the serving layer up front rather than bolting it on later. A lean team isn't about headcount; it's about outsourcing the ops tax (managed) and automating the repetitive work.
相关产品:Google BigQuery、Apache Cassandra / ScyllaDB 相关能力:On-demand 的"扫描税" 最后核验:2026-10-02
Shopify:一个单查询差点每月烧掉 95 万美元(反例) 失败教训
账单爆表
On-demand 计费
聚簇剪枝
成本复盘
- 场景
Shopify 团队在给商家做营销工具的数据管线:Flink 管线(RocksDB 管内部状态)已 ingest 10 亿行,GA(全量发布)要扩量,Flink 扛不住持续增长。方案是找个外部 SQL 数仓:能原子加载 parquet、扛 60 requests/min、结果能回写 GCS。BigQuery 本来就在 Shopify 内部用着,顺手选中。团队先跑通再算账——把 10 亿行 parquet 灌进去跑查询,日志里跳出一行:`total bytes billed: 75462868992`,单次查询 75 GB。
A Shopify team was building a data pipeline for a merchant marketing tool: the Flink pipeline (RocksDB for internal state) had already ingested one billion rows, and the GA launch needed more headroom than Flink could sustainably provide. The plan: find an external SQL warehouse that could atomically load parquet, handle 60 requests per minute, and write results back to GCS. BigQuery was already in use inside Shopify, so it was the natural pick. The team got it working first and did the math later — after loading the billion-row parquet dataset and running the query, the log showed `total bytes billed: 75462868992`: 75 GB for a single query.
- 决策
工程师做了道小学算术:60 RPM × 60 分 × 24 小时 × 30 天 = 每月 259.2 万次查询;每次 75 GB 就是每月 1.944 亿 GB 扫描量;按 on-demand 单价,约 $949,218.75/月——"nearly $1 million"。决策立刻转向:按查询 WHERE 条件里的两个特征列建聚簇表,让引擎按谓词剪枝。同样的查询重跑:billed 508.1 MB(1/150),实际扫描 108.3 MB;月账单 → $1,370.67。
The engineer did quick arithmetic: 60 RPM x 60 minutes x 24 hours x 30 days = 2,592,000 queries per month; at 75 GB each that is roughly 194,400,000 GB scanned monthly — about $949,218.75/month under on-demand pricing, "nearly $1 million." The fix: cluster the table on two feature columns from the query's WHERE clause so the engine prunes by predicate. The identical query then billed 508.1 MB (150x less), with only 108.3 MB actually scanned — bringing the monthly projection down to about $1,370.67.
- 结果
一次聚簇把预估月账单从近百万美元打到一千多美元。博客还留下三条军规:别 SELECT *(只选需要的列)、分区表、用免费的表预览代替"跑个查询看看数据"。
One clustering change cut the projected monthly bill from nearly a million dollars to just over a thousand. The post leaves three standing rules: avoid SELECT * (select only needed columns), partition tables, and use the free table preview instead of "running a query to peek at the data."
- 机制根因
列存按"引用的列×命中的行"计费,LIMIT 救不了——引擎要先算完再截断。这是最毒的一点:查询跑得飞快,延迟一切正常,账单是唯一的报警器,而 bytes-scanned 恰恰是大多数人从不看的指标。聚簇的本质是把"谓词下推"变成物理布局:引擎按聚簇列找到命中区间就停止扫描,不再碰整表。on-demand 模型下没有熔断器,纪律就是熔断器。
Columnar billing charges on "columns referenced times rows touched," and LIMIT cannot save you — the engine computes everything before truncating. The insidious part: the query ran fast, latency looked perfect, and the bill was the only alarm — while bytes-scanned is exactly the metric most engineers never look at. Clustering turns predicate pushdown into physical layout: the engine stops scanning once it has the matching ranges instead of touching the whole table. Under on-demand pricing there is no circuit breaker; discipline is the circuit breaker.
- 教训
任何高频查询上线前,先看 bytes billed 而不是 latency——延迟正常不等于成本正常;聚簇/分区是成本设计,不是性能优化,要在建表时就定;把"预估月账单"写进管线发布 checklist(QPS × 单次扫描 × 单价);on-demand 下默认没有花费上限,别把"跑通了"当成"可以上线了"。
Before any high-frequency query ships, check bytes billed, not latency — fast does not mean cheap. Clustering and partitioning are cost design, not performance tuning, and belong in the table-creation decision. Put "projected monthly bill" (QPS x bytes per query x unit price) on the pipeline launch checklist. On-demand has no default spending cap; "it works" is not the same as "it is safe to launch."
相关产品:Google BigQuery 相关能力:On-demand 的"扫描税" 最后核验:2026-10-02
Spotify:欧洲最大 Hadoop 集群之一迁往 BigQuery,2 万日作业双跑近一年(2016–2018) 成功经验
数据仓库迁移
Hadoop 下云
双跑迁移
Serverless 化
- 场景
2016 年 Spotify 宣布 3 年投入 4.5 亿美元 all-in Google Cloud。数据侧是当时欧洲最大的 on-prem Hadoop 集群之一:每天 2 万个数据作业,依赖图极其复杂;服务侧另有近 1200 个微服务要搬。迁移动机按工程总监 Ramon van Alteren 的说法很直白:维护机房、做容量规划"对成为最好的音乐服务没有直接贡献",真正吸引他们的是 Google 的数据栈——BigQuery、Pub/Sub、Dataflow。
In 2016 Spotify announced a $450 million, three-year all-in bet on Google Cloud. On the data side sat one of Europe's largest on-premise Hadoop clusters: 20,000 data jobs a day with a fiendishly complex dependency graph, plus nearly 1,200 microservices to move on the services side. Per director of engineering Ramon van Alteren, the motivation was blunt: maintaining data centers and capacity planning "doesn't directly contribute to being the best music service in the world" — the real draw was Google's data stack: BigQuery, Pub/Sub, Dataflow.
- 决策
数据迁移否决了 big bang:即便有 160 Gbps 专线,全量拷贝也要两个月,"停机两个月我们就不是个生意了"。最终策略是"搬一个作业、拷一份它的依赖、下游没搬完就把输出拷回去"的双跑模式,主体迁移持续 6–12 个月。每个 sprint 团队二选一:forklift(原样搬运,赶时间用)或 rewrite(重写,理想路径但极耗时间);迁移中后期官方叫停 rewrite——"先搬过去再说,上了云还能改"。
A big-bang data migration was ruled out: even with a 160-gigabit-per-second link, copying everything would take two months, "we wouldn't be much of a business if we were down for two months." The strategy became dual-run: port one job, copy its dependencies over, and copy outputs back to the on-premise cluster for any downstream consumers not yet moved. The bulk migration lasted six to twelve months. Each sprint gave teams two options: "forklifting" (lift-and-shift, for the time-poor) or rewrite (the ideal path, but hugely time-consuming); the rewrite path was halted mid-migration — "just migrate it, you can rework it once it's on GCP."
- 结果
Spotify 数据栈完全跑在 BigQuery 上:每月 1000 万次查询与定时作业,处理 500 PB 数据(Computerworld 引 Google Cloud Next 2018 现场演讲)。服务质量"被严格度量,没有退化";事件投递管线(承载版权方版税结算)峰值从 80 万事件/秒涨到 300 万/秒。成本方面官方拒绝给数字:规模涨了太多,"没法同比"。
Spotify's data stack now runs entirely on BigQuery: 10 million queries and scheduled jobs per month, processing 500 petabytes of data (Computerworld, reporting the Google Cloud Next 2018 talk). Quality of service was "measured diligently" with no degradation; the event-delivery pipeline (which carries royalty payments to rights holders) grew from a peak of 800,000 events per second to 3,000,000. On cost, Spotify declined to give figures: the company had grown so much there was no like-for-like comparison.
- 机制根因
三层约束决定了策略。物理层:带宽×数据量算出的 big bang 停机窗口不可接受,这是数学不是偏好。依赖层:复杂的作业依赖图意味着单个作业无法独立搬迁,双跑+输出回拷是唯一不断下游的办法,代价是大量昂贵且复杂的 copy jobs(负责人 Baer 的忠告:"get out of hybrid as fast as you can")。人性层:工程师一旦动手重写就忍不住顺手重构架构,rewrite 路径会系统性拖慢迁移——迁移纪律本身是生产力。
Three layers of constraint dictated the strategy. Physics: bandwidth times data volume made the big-bang outage window unacceptable — arithmetic, not preference. Dependencies: the tangled job graph meant no job could move alone; dual-run plus output copy-back was the only way to keep downstream consumers alive, at the price of many expensive copy jobs ("get out of hybrid as fast as you can," warned migration lead Josh Baer). Human nature: once engineers start rewriting, they can't resist re-architecting too, and the rewrite path systematically slowed the migration — migration discipline is itself productivity.
- 教训
迁移前先算"物理账"(数据量÷带宽=最短停机窗口),它比技术选型更能决定策略;依赖图越复杂,双跑期越要压缩——copy jobs 是纯成本,每多一天都是浪费;给 rewrite 设熔断:先上云再重构,顺序反了就误了窗口期;迁移可视化(红/绿气泡看板)能替代大量状态汇报,士气工具也是生产工具。
Before migrating, do the "physics math" (data volume divided by bandwidth equals minimum outage window) — it constrains strategy more than technology preferences do. The more tangled the dependency graph, the shorter the dual-run window should be: copy jobs are pure cost, every extra day is waste. Put a circuit breaker on rewrites: move to the cloud first, refactor second — reversed order misses the window. A live migration visualization (red/green bubble board) replaced reams of status reporting; morale tooling is productivity tooling too.
相关产品:Google BigQuery 相关能力:Serverless 双轨计费 最后核验:2026-10-02
Wix:8900 万站长的分析仪表盘跑在 BigQuery 上(2016–) 成功经验
用户画像仪表盘
流式架构
Serving 分层
多租户
- 场景
Wix 是云建站平台(案例页口径:8900 万注册用户),想给站长提供多用途分析仪表盘:转化、clickstream、多维度筛选,还要能测"换高清图是否带来更多销售"。要求很硬:低延迟、亚 100ms 响应。原来的托管方案缺多租户支持、数据重导与恢复、按项目的安全隔离,撑不住。
Wix is a cloud website-building platform (89 million registered users, per the case page) that wanted to offer site owners multi-purpose analytics dashboards: conversions, clickstreams, multi-dimensional filters — even testing whether higher-resolution images drive more sales. The requirements were hard: low latency, sub-100ms responses. The incumbent managed-hosting solution lacked multi-tenant support, data re-import and recovery, and per-project security isolation.
- 决策
搭了一套分层流式架构:App Engine 收数 → Cloud Pub/Sub 排队;BigQuery 并行仓储原始数据(用于恢复与排障);Dataflow 流式处理 Pub/Sub 数据,输出进 Cloud Datastore;自研查询服务走 App Engine 按日期、过滤条件供仪表盘查询。注意 BigQuery 在这里的定位:它是"原始真相源+排障利器",而不是 serving 层——延迟敏感的查询走 Datastore+自研服务。
A layered streaming architecture: App Engine collects data into Cloud Pub/Sub; BigQuery warehouses the raw data in parallel (for recovery and troubleshooting); Dataflow processes the Pub/Sub stream into Cloud Datastore; a proprietary query service on App Engine serves dashboard queries by date range and filter. Note BigQuery's role here: it is the "raw source of truth plus troubleshooting tool," not the serving layer — latency-sensitive queries go to Datastore plus the bespoke service.
- 结果
Wix 数据服务高级总监 Gregory Bondar 称,用 GCP 搭仪表盘的成本不到自建的 20%(Wix 口径,经 Google Cloud 客户页发布);仪表盘成了 crowded 建站市场的差异化卖点,帮公司获客和留存;新垂直市场(音乐站看"过去 1/5/24 小时最热歌单"这类)的仪表盘能快速复制上线。
Per Gregory Bondar, Wix's Senior Director of Data Services, building dashboards on GCP cost less than 20 percent of building them in-house (Wix's figure, published via the Google Cloud customer page). Dashboards became a differentiator in a crowded site-builder market, helping win and retain customers; new vertical markets (e.g., music sites reporting the most popular playlists in the last 1/5/24 hours) could be spun up quickly.
- 机制根因
分层取舍是关键:BQ 存全量原始事件(便宜、SQL 可查,排障时直接翻原始记录),serving 用 Datastore 保亚 100ms——这与 Monzo"小查询慢"的观察一致,BigQuery 当时不适合高并发小查询。多租户隔离靠项目级安全隔离+灵活的资源分配,而不是在查询引擎里做行级隔离。成本对比的锚点是"自建"而非"友商",这是客户证言的常见口径,引用时需注明。
The layering trade-off is the point: BigQuery holds the full raw event firehose (cheap, SQL-queryable, invaluable when debugging), while Datastore guarantees sub-100ms serving — consistent with Monzo's observation that BigQuery was a poor fit for high-concurrency small queries at the time. Multi-tenant isolation came from per-project security isolation and flexible resource allocation, not row-level tricks inside the query engine. Note the testimonial's cost anchor is "vs building in-house," not "vs a competitor" — the standard framing of customer quotes, to be cited with that caveat.
- 教训
BigQuery 的正确位置常常是"真相源+排障",serving 另起一层,不要指望一个引擎包打天下;选型时把生态组合(Pub/Sub+Dataflow+BQ+Datastore)当整体评估,单产品视角会误判;客户证言里的成本数字锚点多是"vs 自建",横向对比时要打折听;仪表盘这类"高频小查询" workload,天生不适合按扫描计费的引擎。
BigQuery's rightful place is often "source of truth plus debugging," with serving handled by a separate layer — don't expect one engine to do everything. Evaluate the ecosystem bundle (Pub/Sub + Dataflow + BigQuery + Datastore) as a whole; a single-product lens misleads. Cost figures in customer testimonials are usually anchored to "vs in-house" — discount accordingly in head-to-head comparisons. Dashboard-style "high-frequency small query" workloads are structurally mismatched with per-scan billing.
相关产品:Google BigQuery 相关能力:— 最后核验:2026-10-02
字节跳动:18000 节点 ClickHouse/ByteHouse 撑起 700PB 实时分析(2022) 成功经验
实时分析
超大规模集群
存算架构
开源自研
- 场景
字节跳动 2017 年在实时分析版块试水 ClickHouse,2018—2019 年从单一业务扩展到 BI 分析、A/B 测试、模型预估等多个业务。截至 2022 年 3 月,其内部实时计算平台(自研 ByteHouse,基于 ClickHouse)总节点达 18000 个、单集群最大 2400 节点、支撑数据量最大 700PB、覆盖 80% 的字节业务。典型场景包括抖音线上活动的实时数据大屏、行为分析(事件/留存/漏斗)与精准营销的人群圈选。(
https://www.cnblogs.com/bytedata/p/17240123.html)
ByteDance piloted ClickHouse for realtime analytics in 2017, expanding in 2018–2019 from a single business to BI analytics, A/B testing, and model estimation. As of March 2022, its in-house realtime compute platform (ByteHouse, self-built on ClickHouse) ran 18,000 nodes in total, with the largest single cluster at 2,400 nodes, serving up to 700PB of data and covering 80% of ByteDance's businesses. Typical scenarios: realtime dashboards for Douyin live campaigns, behavioral analysis (events/retention/funnels), and audience targeting for precision marketing. (
https://www.cnblogs.com/bytedata/p/17240123.html)
- 决策
2017 年字节对实时数仓提出四项硬要求:支撑持续增长的海量数据(2019 年每天新增 100TB)、批流一体、明细+聚合双查、数千维度秒级交互响应,且成本可控。团队评估了 Redis、Apache 系等多种开源方案,但每种只能满足一到两点、维护多套系统成本过高;最终发现 ClickHouse 是唯一能同时满足全部要求的"All In One"选择。2020 年 ByteHouse 在字节内部正式立项(基于 ClickHouse 深度自研),2021 年经火山引擎对外服务。
In 2017 ByteDance set four hard requirements for its realtime warehouse: keep up with exploding data (100TB of new data daily by 2019), batch-streaming unification, both detail and aggregate queries, and second-level interactive response across thousands of dimensions — all cost-controlled. The team evaluated Redis, Apache-ecosystem, and other open-source options, but each satisfied only one or two requirements, and maintaining multiple systems was too costly; ClickHouse emerged as the only "All In One" choice satisfying all four. ByteHouse was formally launched internally in 2020 (deep self-build on ClickHouse) and offered externally via Volcengine in 2021.
- 结果
以下均为字节跳动数据平台官方文章自述口径(截至 2022 年 3 月):内部总节点 18000 个、单集群最大 2400 节点、支撑数据量最大 700PB、覆盖 80% 字节业务。实时监控场景写入 TPS 达 250 万/秒、端到端秒级可见且保障 Exactly Once;行为分析场景 90% 查询 5—7 秒内返回、整体 10 秒内响应;精准营销场景在自研优化器加持下 P95 响应 1 秒以内、部分达半秒,Bitmap 引擎使人群交并补计算获得 10—50 倍提升。全局查询 QPS 等指标未找到公开数据。
All figures below are ByteDance's data-platform team's own account (as of March 2022): 18,000 total internal nodes, 2,400-node largest cluster, up to 700PB served, 80% of businesses covered. Realtime monitoring: 2.5M writes/sec with second-level end-to-end visibility and Exactly Once semantics; behavioral analysis: 90% of queries returned in 5–7 seconds, all within 10 seconds; precision marketing: P95 under 1 second with their self-built optimizer (some at half a second), and a Bitmap engine delivering 10–50x speedups on audience set operations. No public data found for global query QPS.
- 机制根因
ClickHouse 胜出的机制在于列式存储 + 向量化执行 + 稀疏索引带来的海量数据扫描速度,且性能基准建立在磁盘而非内存之上——服务器成本随规模线性增长而非指数级,这是"省钱"得以成立的关键。但规模化也验证了开源 ClickHouse 的机制短板:ZooKeeper 依赖在超大规模下成为可用性瓶颈(硬件故障近乎每天发生、故障恢复常超 1 小时),字节自研 HaMergeTree 将 ZK 负载降到与数据量无关,并用 RocksDB 做元数据持久化,把恢复时间从 1—2 小时压缩到 3 分钟。代价是自研投入:走到 2400 节点规模,开源版本的运维工具、多表 Join、ZK 等短板都会被放大,必须有架构演进(存算分离)或自研兜底的准备。
ClickHouse won on columnar storage + vectorized execution + sparse indexing for massive scan speed, with the performance baseline built on disk rather than memory — server costs grow linearly, not exponentially, with scale, which is what makes the "savings" real. But scale also validated open-source ClickHouse's structural weaknesses: the ZooKeeper dependency became an availability bottleneck at hyperscale (hardware failures nearly daily, recovery often exceeding 1 hour), so ByteDance built HaMergeTree to decouple ZK load from data volume and persisted metadata in RocksDB, compressing recovery from 1–2 hours to 3 minutes. The cost is self-build investment: at 2,400-node scale, the open-source version's ops tooling, multi-table JOINs, and ZK weaknesses all get amplified — you need an architecture evolution plan (compute-storage separation) or in-house fallback ready.
- 教训
选型时不要为每个场景配一种引擎,"一款满足全部要求"的 All In One 引擎比拼凑多套开源系统更省长期运维成本;成本模型要按"磁盘线性 vs 内存指数"来算,大规模 OLAP 的成本优势往往来自存储层而非计算层;对开源引擎的规模上限做诚实评估,提前规划自研投入或架构演进,而不是等问题爆发。
Don't assign one engine per scenario — ByteDance's experience is that one engine satisfying all requirements costs less in long-term operations than stitching multiple open-source systems. Model cost as "linear on disk vs exponential in memory": hyperscale OLAP cost advantages usually come from the storage layer, not compute. Honestly assess an open-source engine's scale ceiling and plan self-build investment or architecture evolution in advance, instead of waiting for the blowup.
相关产品:ClickHouse 相关能力:海量事件的 ad-hoc 全表扫描、Kafka 表引擎准实时写入、没有优化器:JOIN 弱、更新慢、运维心智负担 最后核验:2026-10-01
WMG(华纳音乐):Tango 版税平台停用 Cassandra,迁回 PostgreSQL RDS——"得的不再是规模病,是复杂度病"(2026) 失败教训
回迁
NoSQL 回摆
复杂度税
双写一致性
关系模型
- 场景
华纳音乐集团创新实验室(WMG Innovation Lab)的版税处理平台 Tango,每年处理超 10 亿美元版税、250 万作品目录、年处理数据超 7TB。2012–2013 年 NoSQL 热潮中,团队为 Tango 选择了 Cassandra(自研 ORM 基于 2013 年的 1.2.x 客户端,耦合 Spring 3.1.x),指望水平扩展与高写入吞吐。十多年后,作者 Mark Lin 与 Patrick Cosmo 在官方工程博客发文复盘:Cassandra 已被完全停用(completely decommissioned),全部迁往 AWS PostgreSQL RDS——作者原话是"back to a traditional relational database"。
The Warner Music Group Innovation Lab's royalty-processing platform Tango handles over $1 billion in annual royalties across a 2.5-million-work catalog, processing 7+ TB of data a year. In the 2012–2013 NoSQL wave, the team built Tango on Cassandra (an in-house ORM on a 2013-era 1.2.x client, coupled to Spring 3.1.x), betting on horizontal scalability and high write throughput. A decade later, authors Mark Lin and Patrick Cosmo published a review on the official engineering blog: Cassandra has been completely decommissioned for Tango, everything migrated to AWS PostgreSQL RDS — in the authors' own words, "back to a traditional relational database."
- 决策
分阶段、按实体逐个迁移,拒绝"大爆炸"一次性切换(100 亿行数据,12 小时窗口若在第 8 小时失败,回滚是噩梦)。两个核心机制:一是"并行版税试算"——新旧两套实现并排跑 2 个季度,在强审计合规要求下逐实体比对结果的准确性与完整性,确认无误才下线旧实现;二是"历史回填"——实时数据同步后,用自研抽取管道在后台分批回填 100 亿历史行。
A phased, entity-by-entity migration — no big-bang cutover (with 10 billion rows, a failure eight hours into a twelve-hour window makes rollback a nightmare). Two core mechanisms: "parallel royalty runs" — old and new implementations ran side by side for 2 quarters, reconciling accuracy and completeness entity by entity under strict audit/compliance requirements before anything was decommissioned; and "historical backfilling" — once live data was syncing forward, custom extraction pipelines quietly backfilled the 10 billion historical rows in controlled batches.
- 结果
报表摄取与计算耗时下降约 20%(作者口径);Cassandra 时代平均每天至少 1 次节点故障、季度结算期高达每天 3 次,迁往全托管 RDS 后该问题消失;砍掉了外部 Elasticsearch 集群与自研 ORM,标准 SQL 与视图即可满足此前需要双写同步的查询;下游 BI 系统从"全表抽取"变为可做增量(delta)抽取。成本侧:迁移启动时已跑到 32 个节点,按数据增速预计 2 年内要扩到 64 节点,作者称成本呈指数级增长("cost grew exponentially",作者口径,未给绝对金额)。
Statement ingestion and calculation runtimes dropped roughly 20% (authors' claim); in the Cassandra era the team averaged at least 1 node outage a day, spiking to 3 a day during quarter-close royalty processing — gone after moving to fully managed RDS; the external Elasticsearch cluster and the in-house ORM were eliminated, with standard SQL and views covering queries that previously needed dual-write synchronization; downstream BI systems moved from full-table extracts to incremental (delta) extraction. On cost: the fleet had reached 32 nodes when the migration started, projected to hit 64 within ~2 years on data-volume growth, with cost growing "exponentially" in the authors' words (authors' claim; no absolute figures given).
- 机制根因
音乐版税是"深度互联的关系网"(用户→地区→分成比例→身份),而 Cassandra"hates relationships":为满足查询被迫建扁平化冗余表,外挂 Elasticsearch 做二级索引——相当于"在搜索引擎里重建了一个关系模型";每次更新用户属性都要双写 Cassandra 与 ES,同步延迟即不一致;团队把时间花在"照顾集群健康、保证 repair job 完成"而不是交付功能。作者的诊断句是全文转折点:"we weren't suffering from a scale problem anymore; we were suffering from a complexity problem."——2013 年关系型数据库确实扛不住他们设想的规模,但 2026 年的云托管 PG 早已不是当年的 PG,"规模优先"的技术栈变成了纯粹的复杂度税。
Music publishing is "a deeply interconnected web of relationships" (user → territory → royalty split → identity), while Cassandra "hates relationships": serving queries forced flattened, duplicated tables plus an external Elasticsearch cluster for secondary indexes — effectively "reconstructing a relational schema inside a search engine"; every user-attribute update meant dual-writing Cassandra and Elasticsearch, where sync lag equaled inconsistency; engineering time went to "babysitting cluster health and ensuring repair jobs completed" instead of shipping features. The authors' diagnosis is the line the whole post turns on: "we weren't suffering from a scale problem anymore; we were suffering from a complexity problem." In 2013 relational databases genuinely couldn't handle the scale they envisioned — but 2026's managed cloud Postgres is not 2013's Postgres, and the "scale-first" stack had become pure complexity tax.
- 教训
NoSQL 选型先回答"访问模式是关系型的吗":如果是,宽列存储的去关系化只是把 join 推迟到应用层和 ES 里付,账单会迟到但不会缺席;"双写 + 外部索引"是危险信号——当你开始在搜索引擎里重建关系模型时,就该回头算总账;十年前正确的架构决策会因基础设施进化而过期,作者建议"别让十年前的选择靠惯性决定你的架构未来";100 亿行级别的迁回必须做双跑对账,并行期按季度计,不可大爆炸切换。
For any NoSQL selection, first answer "is the access pattern relational?": if yes, a wide-column store's de-relationalization only defers the joins into the application layer and Elasticsearch — the bill arrives late but never waived. "Dual-write + external index" is a danger signal: when you start rebuilding a relational model inside a search engine, it's time to redo the math. An architecture decision that was right a decade ago expires as infrastructure evolves — the authors advise against letting decade-old choices dictate your architectural future out of sheer habit. A 10-billion-row move-back demands dual-run reconciliation; measure the parallel period in quarters, never big-bang it.
来源
WMG Innovation Lab 官方工程博客《Rethinking NoSQL: Why We Migrated From Cassandra to PostgreSQL RDS》(Mark Lin、Patrick Cosmo,2026-07-07
WMG Innovation Lab official engineering blog, "Rethinking NoSQL: Why We Migrated From Cassandra to PostgreSQL RDS" (Mark Lin, Patrick Cosmo, 2026-07-07
"back to a traditional relational database"为作者原话
"back to a traditional relational database" is the authors' wording
相关产品:Apache Cassandra / ScyllaDB、PostgreSQL(社区版) 相关能力:宽列存储的关系模型代价 最后核验:2026-10-02
Cloudflare:DEX 分析面弃选 ClickHouse——小批量写入场景下"五件套"接入成本太高(2025) 失败教训
PoC
选型评估
弃选 ClickHouse
时序分析
小团队
- 场景
Cloudflare Zero Trust 的 DEX(Digital Experience Monitoring)团队只有 3 名全栈工程师,要为 WARP 客户端的设备状态日志(每 2 分钟一条)建分析面。作者曾想用 ClickHouse,但按内部文档搭写入链路需要 Cap'n Proto/Protobuf → socket → logfwdr → Kafka → Concept:Inserter 五件套——系统图上凭空多出 5 个框。而 ClickHouse 的 MergeTree 为高吞吐批量写入优化,"每秒批量 < 1 次"的写入假设与"数百万设备每 2 分钟一条"的小批量写入正好相克,小写会引发写放大、资源争抢与限流。
Cloudflare Zero Trust's DEX (Digital Experience Monitoring) team — three full-stack engineers — needed an analytics plane for WARP client device-state logs (one per device every 2 minutes). The author wanted ClickHouse, but the internal docs' ingestion path required Cap'n Proto/Protobuf → socket → logfwdr → Kafka → Concept:Inserter — five new boxes on the architecture diagram. And ClickHouse's MergeTree, optimized for high-throughput batch inserts with a "fewer than one batch per second" assumption, is the exact opposite of millions of devices uploading one log each every 2 minutes — small writes trigger write amplification, resource contention, and throttling.
- 决策
先上原生 PostgreSQL(~200 inserts/sec 上线,查询几百毫秒),撑到 1000 inserts/sec、数十亿行后,7 天时间范围查询退化到数秒。随后在 canary PG 集群自建 TimescaleDB 实例做 apples-to-apples 对比:生产双写 + 两周回填,用真实 dashboard 查询做 side-by-side benchmark(3 个时间窗口 × 3 种 columnstore 模式,5 亿–10 亿行数据集)。
They shipped on vanilla PostgreSQL first (~200 inserts/sec at launch, queries in the hundreds of milliseconds), and when it scaled to 1,000 inserts/sec with billions of rows, 7-day queries degraded to seconds. They then ran an apples-to-apples comparison on a self-hosted TimescaleDB instance on their canary PG cluster: production dual-write plus a two-week backfill, with side-by-side benchmarks using real dashboard queries (3 time windows × 3 columnstore modes, 500M–1B row datasets).
- 结果
弃选 ClickHouse,选 TimescaleDB(PostgreSQL 扩展;TimescaleDB 不在本站 31 产品内,此处仅记录评估结论)。按 Cloudflare 自述口径:查询提升 5x–35x(随查询类型与时间窗口);压缩 1616GB→49GB(32.83x),"同等成本多存 33 倍数据"。入选理由:PG 生态完整保留、hypertable 自动分区、continuous aggregates 替代 cron 预聚合、3 人团队要简单。
ClickHouse was rejected in favor of TimescaleDB (a PostgreSQL extension; TimescaleDB is outside this site's 31 products and appears here only as the evaluation outcome). Per Cloudflare's own account: 5x–35x query improvements (varying by query type and time range); compression from 1616GB to 49GB (32.83x) — "33x more data retained for the same cost." Selection reasons: full Postgres ecosystem retained, hypertables auto-partitioned, continuous aggregates replacing cron pre-aggregation jobs, and simplicity for a 3-person team.
- 机制根因
这不是"ClickHouse 不够快",而是"接入成本与写入模型错配"。MergeTree 的每次 insert 写成独立 partition 靠后台合并——大批量写入时这是神器,小批量写入时就是写放大器。DEX 的 workload(低频、小批量、高设备数)恰好踩在反模式上。TimescaleDB 赢在"不用换数据库":PG 生态、SQL、工具链全保留,columnstore 和预聚合是加法而非换血。对 3 人团队,"少一个框"本身就是性能指标。
This wasn't "ClickHouse isn't fast enough" — it was ingestion-cost and write-model mismatch. MergeTree writes each insert as a separate partition and merges in the background — a superpower for large batch writes, a write-amplification machine for small ones. DEX's workload (low-frequency, small-batch, high device count) landed exactly on the anti-pattern. TimescaleDB won on "no database switch required": PG ecosystem, SQL, and tooling all retained; columnstore and pre-aggregation were additive, not a transplant. For a 3-person team, "one fewer box" is itself a performance metric.
- 教训
评估 OLAP 时先画写入链路图——"5 个框"的接入成本要和查询性能同权打分;小批量写入场景下,MergeTree 类引擎的默认假设就是反模式;小团队选型应把"运维复杂度"货币化计入 TCO。诚实注记:5x–35x、32.83x 均为 Cloudflare 自测口径(真实 dashboard 查询),未见第三方复现;这是 2025 年 7 月的结论,TimescaleDB 此后版本演进请自行核验。
When evaluating OLAP, draw the ingestion pipeline diagram first — the "five boxes" of onboarding cost deserve equal weight with query performance; under small-batch writes, MergeTree-class engines' default assumptions are the anti-pattern; small teams should monetize operational complexity into TCO. Honesty note: the 5x–35x and 32.83x figures are Cloudflare's self-tested claims (real dashboard queries) with no independent reproduction found; these conclusions date from July 2025 — re-verify against newer TimescaleDB releases.
相关产品:ClickHouse、PostgreSQL(社区版) 相关能力:MergeTree 批量写入假设与小写放大 最后核验:2026-10-02
PostHog:评估 6 款 OLAP 引擎后选 ClickHouse,自建 benchmark 跑出"一个数量级"领先(2021) 成功经验
PoC
选型评估
OLAP 选型
Postgres 替代
自建压测
- 场景
PostHog(开源产品分析平台)最初跑在 Heroku Postgres 上,随事件摄入量增长把 Postgres 推到极限:用户越多、用得越多,体验越差。2021 年团队立项为 Postgres 找替代者,候选包括 Pinot、Presto、Druid、TimescaleDB、CitusDB、ClickHouse。
PostHog (open-source product analytics platform) originally ran on Heroku Postgres, which hit its limits as event ingestion grew: the more users they had and the more those users used the product, the worse the experience got. In 2021 the team set out to find a Postgres replacement, with candidates including Pinot, Presto, Druid, TimescaleDB, CitusDB, and ClickHouse.
- 决策
按三个维度初筛——速度(实时结果)、复杂度(产品可自托管,不能要求用户装整个 Hadoop 栈)、查询接口(要标准 SQL;Druid 因"有 SQL 壳但不是 exactly SQL"被淘汰)。通过初筛后,团队研读公开 benchmark、参考 Cloudflare 用 ClickHouse 处理每秒 600 万请求的经验,然后自建测试集群跑自己的 benchmark。
An initial screen on three dimensions — speed (real-time results), complexity (the product is self-hostable, so users couldn't be asked to install an entire Hadoop stack), and query interface (standard SQL required; Druid was eliminated because, while it has a SQL wrapper, it's "not *exactly* SQL"). After the screen, the team studied public benchmarks, looked at Cloudflare's experience running 6M requests/sec on ClickHouse, then built a test cluster to run their own benchmarks.
- 结果
选定 ClickHouse。按 PostHog 自述口径,ClickHouse 在自家 benchmark 里"repeatedly performed an order of magnitude better"(反复比其他候选好一个数量级);压缩表现甚至超过 ORC/Parquet 序列化格式;从磁盘处理数据(不像 Presto 要求数据常驻内存);数据到达即处理,无需预聚合。迁移采用双写 + feature flag 逐查询切换的并行方案。
ClickHouse was selected. Per PostHog's own account, ClickHouse "repeatedly performed an order of magnitude better" than the other tools in their benchmarks; its compression even beat serialization formats like ORC and Parquet; it processes from disk (unlike Presto, which needs data in memory); and data is processed as it arrives with no pre-aggregation needed. Migration used a parallel-run approach with feature flags, switching queries over one by one.
- 机制根因
PostHog 选 ClickHouse 不是只看跑分,而是三个约束同时命中:列存 + C++ 带来的摄入/压缩比契合"事件分析" workload;类 PG/MySQL 的 SQL 方言让团队持续加功能不被查询层卡脖子;自托管简单契合开源产品的分发模式。反面教材同样真实:团队曾按每小时 1–2 次的文档建议反着用 mutations(跑到每分钟数百次),直接导致 outage——ClickHouse 的 MergeTree 假设"少量大文件",高频小 mutation 会让后台合并永远追不上。
PostHog didn't choose ClickHouse on benchmarks alone — three constraints aligned: columnar storage plus C++ fit the event-analytics workload's ingestion/compression profile; the Postgres/MySQL-like SQL dialect kept feature development from being bottlenecked by the query layer; self-hosting simplicity fit an open-source product's distribution model. The cautionary tale is equally real: the team ran mutations at hundreds per minute against documentation guidance of one or two per hour, causing an outage — ClickHouse's MergeTree assumes "few large files," and high-frequency small mutations let background merges fall permanently behind.
- 教训
选型三维度(速度/复杂度/接口)要先于 benchmark 定下来,否则跑分没有评判标准;"标准 SQL"在长期功能迭代中是生产力条款,不只是口味问题;任何引擎的"不要这么用"文档条款(如 mutation 频率)都要当硬约束读,PostHog 的 outage 就是交过学费的注脚。诚实注记:"好一个数量级"为 PostHog 自测口径,未见第三方复现;TimescaleDB/CitusDB 不在本站 31 产品内,此处仅作评估名单记录。
Fix the evaluation dimensions (speed/complexity/interface) before benchmarking, or the numbers have no rubric; "standard SQL" is a productivity clause for long-term feature velocity, not just a taste preference; treat every engine's "don't do this" documentation clause (e.g., mutation frequency) as a hard constraint — PostHog's outage is the paid-tuition footnote. Honesty note: "an order of magnitude better" is PostHog's self-tested claim with no independent reproduction found; TimescaleDB/CitusDB are outside this site's 31 products and appear here only as evaluation-list records.
相关产品:ClickHouse、PostgreSQL(社区版) 相关能力:列存压缩与 MergeTree 的写入假设 最后核验:2026-10-02
Cloudflare:ClickHouse 做每秒 600 万请求的 HTTP 实时分析 成功经验
实时分析/OLAP
高吞吐摄入
物化视图
高可用
- 决策
用 ClickHouse 替换旧管线最脆弱的环节,Kafka 解耦摄入与查询。
Replace the most fragile link of the old pipeline with ClickHouse, with Kafka decoupling ingestion from querying.
- 结果
36 节点 × 3 副本扛下全网流量;月 1.5 万亿 page views 可查;物化视图预聚合常用维度,查询延迟可控。
36 nodes × 3 replicas carry all global traffic; 1.5 trillion monthly page views queryable; materialized views pre-aggregate common dimensions, keeping query latency under control.
- 机制根因
append-mostly 日志流 + 列式存储 + 物化视图预聚合,是该负载特征的标准答案;Kafka 让摄入与查询解耦,背压可控,单点故障被隔离在管线之外。
Append-mostly log streams plus columnar storage plus materialized-view pre-aggregation is the textbook answer for this workload shape; Kafka decouples ingestion from queries, keeps backpressure manageable, and isolates single-point failures outside the pipeline.
- 教训
"不自研"的机制条件:负载特征(只追加、按时间与维度聚合)与成熟工具的抽象精确匹配时,不要自研聚合层——列式 OLAP + MQ 缓冲是成熟范式。演进策略同样重要:先替换最脆弱的单点、验证范式,再整体演进,而不是一次性重写整条管线。
The mechanism condition for "not building it yourself": when the workload shape (append-only, aggregate by time and dimension) precisely matches a mature tool's abstraction, don't build a custom aggregation layer — columnar OLAP plus an MQ buffer is the proven pattern. Evolution strategy matters equally: replace the most fragile single point first, validate the pattern, then evolve the whole pipeline — never rewrite the entire pipeline in one shot.
相关产品:ClickHouse 相关能力:海量事件的 ad-hoc 全表扫描、物化视图 + 表引擎的 DDL 管道哲学 最后核验:2026-10-01
DoorDash:特征存储引入 CockroachDB 做 Redis 补充,单位存储云支出降 75%(2023) 成功经验
特征存储
Redis 降本
混合存储
在线推理
- 场景
DoorDash 机器学习平台 2021–2022 年特征数量增长超过 10 倍,在线特征存储跑在 AWS ElastiCache Redis 上(超 100 节点的大集群)。特征暴增迫使团队每周扩容一次 Redis 集群,而大集群扩容是 2–3 天的蓝绿流程——从日备份起新集群、重放一天写入、切流量、删旧集群,极易出错且耗时不定,有时因缺 AWS 机型还得找 AWS 支持重试。
From 2021 to 2022, the number of ML features created by practitioners at DoorDash grew more than 10x. The online feature store ran on AWS ElastiCache Redis (clusters of 100+ nodes). The feature explosion forced the team to upscale a Redis cluster about once a week, and upscaling a large cluster was a 2–3 day blue-green process — spin up a new cluster from the daily backup, replay a day of writes, switch traffic, delete the old cluster — error-prone with unpredictable duration, sometimes requiring AWS support when instance types were unavailable.
- 决策
ML 平台团队(Brian Seo、Kunal Shah,2023 年 3 月 21 日发表工程博客)决定引入 CockroachDB 作为在线特征存储的补充后端,与 Redis 混合使用:超低延迟场景留给 Redis,其余搬到磁盘型 CockroachDB。选型理由是零停机升级/扩容、按负载自动伸缩,以及磁盘存储让高基数特征的单位成本大幅下降。
The ML platform team (Brian Seo and Kunal Shah, in an engineering blog post dated March 21, 2023) decided to add CockroachDB as a supplementary backend for the online feature store, used alongside Redis: ultra-low-latency use cases stayed on Redis, everything else moved to disk-based CockroachDB. The rationale was zero-downtime upgrades and scaling, load-based auto-scaling, and far cheaper per-byte cost for high-cardinality features on disk.
- 结果
单位特征值存储的云支出平均下降 75%,延迟仅小幅上升(DoorDash 工程博客口径,未找到第三方独立复现)。过程有波折:最初一行一值的 KV 模型下,63 台 m6i.8xlarge 峰值约 200 万行/秒写入、CPU 均值约 30%,但 CPU 会冲到 50–70% 且吞吐腰斩到不足 100 万行/秒;按此测算仅为 Redis 成本的 30% 左右,未达预期。团队把同一 entity 的特征收敛为 JSON map(一行多值)后,写效率提升最高 300%,读延迟降 50%,部分场景读性能接近 Redis,才最终拿到 75% 的降本数字。以上数字均为博客自述口径。
Cloud spend per feature value stored fell 75% on average, with only a minimal latency increase (DoorDash engineering blog figure; no independent third-party reproduction found). The path was bumpy: with the initial one-row-per-value KV model, 63 m6i.8xlarge instances peaked at roughly 2 million rows/sec with ~30% average CPU, but CPU would spike to 50–70% while throughput collapsed by half to under 1 million rows/sec; back-of-the-envelope math put costs at only ~30% of Redis — not the win hoped for. After condensing each entity's features into a JSON map (multiple values per row), write efficiency rose up to 300%, read latency dropped 50%, and some workloads read nearly as fast as Redis — only then did the 75% figure land. All figures are as self-reported in the blog.
- 机制根因
Redis 是内存数据库,特征这种高基数、持续膨胀的数据集用内存装,成本随特征数大致成比例上涨(取决于内存价格与压缩效果),且 ElastiCache 大集群扩容本质是"重建集群",运维税极高。CockroachDB 的 range 有序分片加 LSM 磁盘存储把单位字节成本降下来;JSON 收敛减少了单个特征占用的 range 数量和写放大,一次写入携带多个特征值,正好命中 LSM 批量写入的甜点。代价有三:批量 INSERT 过大(每查询超 1000 个值)会让最慢节点拖住整个集群;新表从单个 range 起步有预热期,需预切分或限流;可串行化隔离下写冲突多时 CPU 开销显著。选型的本质是"用延迟换成本",且必须按 CockroachDB 的存储模型重塑 schema 才能拿到收益。
Redis is an in-memory database: for a high-cardinality, ever-growing feature dataset, cost grows roughly in proportion to feature count (depending on memory pricing and compression), and upscaling a large ElastiCache cluster is effectively "rebuilding the cluster" — an enormous ops tax. CockroachDB's ordered range sharding plus LSM disk storage brought per-byte cost down; JSON condensing reduced the number of ranges each feature occupied and cut write amplification, with each write carrying multiple feature values — right in the LSM batch-write sweet spot. Three costs: oversized batch INSERTs (>1,000 values per query) let the slowest node stall the whole cluster; new tables start from a single range and need a warm-up (pre-splitting or throttled writes); under serializable isolation, write contention burns significant CPU. The trade is fundamentally "latency for cost," and the schema must be reshaped for CockroachDB's storage model to capture the gains.
- 教训
混合存储不是"多装一个数据库",而是按延迟与成本把工作负载分级;直接把 Redis 的 KV 模型平移到 SQL 表会踩坑(大 batch、单 range 预热、写冲突),schema 必须为 LSM 与 range 分片重新设计;降本数字要看分母——DoorDash 的 75% 是"单位存储价值"口径,且建立在 JSON 收敛重构之后,照搬原始模型只能拿到约 30%。
Hybrid storage is not "install one more database" — it is tiering workloads by latency and cost. Porting Redis's KV model straight onto SQL tables hits every pitfall (giant batches, single-range warm-up, write contention); the schema must be redesigned for LSM and range sharding. Read cost-reduction numbers by their denominator — DoorDash's 75% is "per value stored" and only after the JSON-condensing rework; copying the original model would have yielded only ~30%.
相关产品:CockroachDB、Redis / Valkey 相关能力:— 最后核验:2026-10-02
DoorDash:新零售履约后端从 PostgreSQL 迁到 CockroachDB,支撑 10 倍增长(2023) 成功经验
PostgreSQL 迁移
履约后端
影子读
特性开关
- 场景
DoorDash 从外卖扩张到便利店、杂货等新零售,商户与 SKU 数量指数级增长。履约后端的 store_items 物化视图(商品目录、库存、价格)跑在 PostgreSQL 上,表迅速涨到 500GB——这正是 DoorDash 内部单表上限,超限后表变得不可靠;高峰期大批量非批处理 upsert 让整体服务延迟翻倍、数据库 CPU 超过 80%;且单写节点位于单个可用区,整个新零售业务的可用性系于一个 AZ。
As DoorDash expanded from food delivery into new verticals like convenience and grocery, merchant and SKU counts grew exponentially. The fulfillment backend's store_items materialized view (catalog, inventory, pricing) ran on PostgreSQL and quickly hit 500GB — DoorDash's internal single-table limit, beyond which tables become unreliable. During peak hours, large volumes of non-batched, non-partitioned upserts doubled overall service latency and pushed database CPU past 80%; the single writer sat in one availability zone, tying the entire new-verticals business's availability to a single AZ.
- 决策
New Verticals Fulfillment 团队(Yin Zhang、Nikhil Pujari、Kevin Chen、ThulasiRam Peddineni,2023 年 2 月 7 日发表工程博客)决定把存储引擎换成 CockroachDB,分四个里程碑迁移:先建 gRPC 门面服务 RFDS(Retail Fulfillment Data Service)收敛所有数据访问;再做 schema 改造与数据回填;然后双写加影子读比对(用 Guava MapDifference 抓数据 skew);最后特性开关灰度切流。
The New Verticals Fulfillment team (Yin Zhang, Nikhil Pujari, Kevin Chen, ThulasiRam Peddineni, in an engineering blog post dated February 7, 2023) decided to switch the storage engine to CockroachDB, in four milestones: first build a gRPC facade service, RFDS (Retail Fulfillment Data Service), consolidating all data access; then schema changes and backfill; then dual writes plus shadow-read comparison (using Guava's MapDifference to catch data skew); finally a feature-flagged gradual traffic cutover.
- 结果
迁移消除了扩展瓶颈,多组查询反而更快:按 store_id 的批量查询延迟下降约 38%(store_id 在 CockroachDB 是复合主键首列,在 PostgreSQL 只是普通二级索引);store_id + merchant_supplied_id 的实时查询快 10 倍(复合主键对复合二级索引);按 dd_menu_item_ids 的查询持平。团队称现在可支撑 10 倍负载增长(DoorDash 自述口径,迁移后的绝对 QPS、数据量未公开)。以上数字均为博客自述,未经独立验证。
The migration eliminated the scaling bottleneck and several query patterns got faster: bulk queries by store_id dropped ~38% in latency (store_id is the first column of the composite primary key in CockroachDB but only an ordinary secondary index in PostgreSQL); real-time queries on store_id + merchant_supplied_id ran 10x faster (composite primary key vs. composite secondary index); queries on dd_menu_item_ids were on par. The team reports being ready for 10x load growth (DoorDash's own claim; absolute post-migration QPS and data volumes were not disclosed). All figures are self-reported in the blog and independently unverified.
- 机制根因
PostgreSQL 单写模型下所有 upsert 走一个 writer,500GB 单表加大批量写入等于写入天花板;CockroachDB 的 shared-nothing 多写把写入分散到各 range leader。但性能提升的关键不是"换数据库"本身,而是借机做的三处 schema 手术:高频更新列(价格、库存、标记位)拆成独立 column family,避免整行重写,更新性能提升 5 倍以上;二级索引从 8 个砍到 2 个;废弃大表 join,改服务层内存 join。影子读阶段是关键:双读比对把遗漏的更新路径全部暴露,feature flag 让回滚随时可做。代价是应用必须接受分布式 SQL 约束(join 变贵、索引要精简),且迁移期维护了双写链路。
Under PostgreSQL's single-writer model every upsert funnels through one writer — a 500GB table plus peak bulk writes is a hard write ceiling; CockroachDB's shared-nothing multi-writer spreads writes across range leaders. But the performance wins came less from "switching databases" than from three schema surgeries done along the way: frequently updated columns (price, availability, flags) were split into a separate column family to avoid full-row rewrites, improving update performance more than 5x; secondary indexes were cut from eight to two; large-table joins were deprecated in favor of in-memory joins in the service layer. The shadow-read phase was decisive: dual-read comparison exposed every missed update path, and the feature flag kept rollback always available. The price: the application had to accept distributed-SQL constraints (joins get expensive, indexes must be lean), and a dual-write pipeline had to be maintained during migration.
- 教训
分布式 SQL 迁移最值钱的一步往往是"被迫还技术债"——门面服务、去 join、索引审计这些在 PostgreSQL 上也能做,但没人愿意动;影子读加双写加特性开关是生产迁移的标准三件套,不要信一次性 cutover;单表 500GB 这类内部红线是架构换代的信号灯,不是靠"优化一下"能绕过去的。
The most valuable step in a distributed-SQL migration is often "paying down tech debt under cover" — facade services, de-joining, index audits are all doable on PostgreSQL too, but nobody wants to touch them; shadow reads plus dual writes plus feature flags are the standard three-piece kit for production migrations — never trust a one-shot cutover; an internal red line like a 500GB table limit is the signal to change architecture, not something to route around with "a bit of tuning."
相关产品:CockroachDB、PostgreSQL(社区版) 相关能力:PG 线路兼容 —— 从 PG 迁来的摩擦力最低的分布式 SQL 最后核验:2026-10-02
Form3:英国支付平台把 CockroachDB 横跨 AWS/GCP/Azure 三云,单云故障照常结算(2020–) 成功经验
多云
支付清算
跨云仲裁
监管合规
- 场景
英国云原生支付公司 Form3 提供 Faster Payments 等清算通道接入,最初跑在 AWS RDS for PostgreSQL 上。监管与业务连续性要求不能把命系在一家云上;团队曾考虑在 GCP 按 AWS 技术栈复制一套(SQS 换 PubSub 等),但两套行为不一致的平台维护成本极高,且"大爆炸"式迁移风险大。
Form3, a UK cloud-native payments company providing Faster Payments scheme access, originally ran on AWS RDS for PostgreSQL. Regulatory and business-continuity requirements meant it could not stake the business on a single cloud; the team considered replicating the AWS stack on GCP (SQS→PubSub and the like) but rejected it — two behaviorally inconsistent platforms would be very expensive to maintain, and a big-bang migration was too risky.
- 决策
平台工程团队(负责人 Kevin Holditch,Form3 官方播客 Ep 38 主讲)重构为 V2 多云架构:AWS、GCP、Azure 三云各跑一个 Kubernetes 集群,用私有网络打通;NATS JetStream 与 CockroachDB 集群横跨三云;CockroachDB 副本因子为 3,每个 range 的三个副本分落三云,写 quorum 取三取二——任意一朵云整体故障,写入仍可继续。
The platform engineering team (led by Kevin Holditch, Head of Platform Engineering, the featured guest on Form3's official .tech podcast Ep 38) rebuilt as a V2 multi-cloud architecture: one Kubernetes cluster per cloud on AWS, GCP, and Azure, joined over a private network; NATS JetStream and a CockroachDB cluster spanned across all three clouds; CockroachDB ran with replication factor 3, each range's three replicas landing on different clouds, and write quorum set at two of three — any single cloud failing outright, writes continue.
- 结果
建成"一个逻辑系统横跨三朵云":一笔支付可以在 GCP 发起、在另一朵云继续,全程不绑定单云;产品团队只需对接一套云无关架构。运维细节来自 Form3 工程师 Rogger Fabri 在 RoachFest 2023 的分享:备份用 v23.1 起的 locality-restricted backup 分别落在 GCP 与 AWS,全量每 12 小时、增量每 5 分钟,RPO 为 5 分钟;可观测性用 Grafana、Prometheus、logz.io,告警进 PagerDuty。支付 TPS、延迟等业务数字未公开。
The result is "one logical system across three clouds": a payment can start on GCP and continue on another provider mid-journey, never pinned to one cloud, and product teams integrate against a single cloud-agnostic architecture. Operational detail comes from Form3 engineer Rogger Fabri's RoachFest 2023 talk: backups use locality-restricted backups (since v23.1) landing on both GCP and AWS, full every 12 hours and incremental every 5 minutes, for a 5-minute RPO; observability runs on Grafana, Prometheus, and logz.io with alerts paging via PagerDuty. Payment TPS and latency figures were not disclosed.
- 机制根因
支付的核心约束是"钱不能丢、不能重"——Raft quorum 写天然满足:三取二确认才算提交,单云故障不丢已确认交易。从 PostgreSQL 迁过来最大的摩擦是访问模式:CockroachDB 的分布式执行让"看似简单的查询"实际跨节点取数,需要重做应用设计与性能优化。跨云延迟加在每次 quorum 写上,这是多活每天交的税;好在支付场景对单次写入延迟的容忍度高于对可用性的要求。备份跨云双写则是"备份的备份",防的是云厂商级故障。
Payments have one core constraint — money must neither be lost nor duplicated — and Raft quorum writes satisfy it natively: a write counts only after two-of-three clouds acknowledge, so a single-cloud failure loses no committed transaction. The biggest friction migrating from PostgreSQL was access patterns: CockroachDB's distributed execution means "a seemingly simple query" actually fetches across nodes, requiring application redesign and performance tuning. Cross-cloud latency is added to every quorum write — the daily tax of multi-active operation; fortunately payments tolerate per-write latency far better than they tolerate downtime. Cross-cloud dual backups are "a backup of the backup," guarding against cloud-vendor-level failure.
- 教训
真正的多云不是"每个云一套",而是"一套系统横跨多云",而这要求数据层本身支持跨云 quorum——主从复制模型天然做不到;跨云 CockroachDB 的代价是写延迟永久性垫高,选型前先确认业务是否付得起这笔税;RPO 5 分钟这个数字来自增量备份频率,不是数据库"零丢失"的魔法,两件事别混为一谈。
Real multi-cloud is not "one stack per cloud" but "one system across clouds" — and that requires the data layer itself to support cross-cloud quorum, which primary/replica models fundamentally cannot do; the price of cross-cloud CockroachDB is a permanently elevated write latency — confirm the business can afford that tax before selecting; a 5-minute RPO comes from incremental backup frequency, not from database magic about "zero loss" — don't conflate the two.
相关产品:CockroachDB、PostgreSQL(社区版) 相关能力:多活生存性 —— 丢一个 region 业务不中断 最后核验:2026-10-02
Jepsen 独立测试在 CockroachDB beta 版揪出两处可串行化漏洞,均在 1.0 前修复(2016–2017) 失败教训
一致性测试
可串行化
第三方验证
beta 缺陷
- 场景
CockroachDB 在 1.0 GA(2017 年 5 月)前宣称提供可串行化隔离(serializable,ANSI 最高隔离级别)。团队自己用 Jepsen 框架做过测试,但为求独立验证,2016 年秋聘请 Jepsen 作者 Kyle Kingsbury 评审并扩展测试集,费用由 Cockroach Labs 承担。
Before its 1.0 GA (May 2017), CockroachDB claimed serializable isolation (the highest ANSI SQL isolation level). The team had tested with the Jepsen framework themselves, but to get independent verification, in fall 2016 they hired Jepsen author Kyle Kingsbury to review and extend the test suite, with Cockroach Labs footing the bill.
- 决策
Kingsbury 在 beta-20160829 至 beta-20160908 版本上运行扩展测试:register(单键线性一致性)、bank(转账总额守恒)、sequential、G2(反依赖环)等。测试发现两个新 bug:一是 timestamp cache 缺陷——两个事务被分配相同时间戳时可能产生不一致;二是内部重试导致某些事务被应用两次。两处都是货真价实的可串行化违规。
Kingsbury ran the extended suite against beta-20160829 through beta-20160908: register (single-key linearizability), bank (conservation of total balances), sequential, G2 (anti-dependency cycles), and others. Testing found two new bugs: a timestamp-cache defect that could produce inconsistencies whenever two transactions were assigned the same timestamp, and internal retries that could apply certain transactions twice. Both were genuine serializability violations.
- 结果
两个 bug 分别在 beta-20160915 与 beta-20161013 中修复,扩展后的 Jepsen 套件并入每夜回归,Cockroach Labs 发博客《CockroachDB beta passes Jepsen testing》公示。但测试同时钉死了两个"按设计"的弱保证:CockroachDB 只保证可串行化,不保证严格可串行化(strict serializability)——跨多 key 的事务可能不按实时顺序被观察到,读在有限场景下可能是 stale 的;且一切保证的前提是节点时钟偏移在阈值内(默认 250ms),超限后"一切保证作废",部分节点会自我关闭。
The two bugs were fixed in beta-20160915 and beta-20161013 respectively, the expanded Jepsen suite was folded into nightly regression runs, and Cockroach Labs published the blog post "CockroachDB beta passes Jepsen testing." But the testing also pinned down two by-design weaker guarantees: CockroachDB promises serializability, not strict serializability — multi-key transactions may be observed out of real-time order, and reads can be stale in limited cases; and every guarantee is conditional on node clock offset staying within a threshold (250ms by default) — beyond it, "all bets are off" and some nodes shut themselves down.
- 机制根因
CockroachDB 用混合逻辑时钟(HLC)代替 Spanner 的 TrueTime 原子钟,跑在普通 NTP 同步的硬件上;为换性能放弃了 Spanner 的 commit-wait(写后等待),只做可串行化而不做外部一致性。timestamp cache bug 的本质是相同时戳的事务在缓存判定上撞车,破坏了"读到已提交写的单调性";double-apply 则是两阶段提交内部重试路径缺了幂等。Jepsen 的价值在于用形式化测试把"文档声称"和"实际语义"的缝隙量化出来——Cockroach Labs 当时的文档写着"no possibility of reading stale data",测试证明这句话只在单 key 且时钟良好时成立。
CockroachDB uses hybrid logical clocks (HLC) instead of Spanner's TrueTime atomic clocks, running on commodity NTP-synchronized hardware; for performance it dropped Spanner's commit-wait and settled for serializability without external consistency. The timestamp-cache bug was, at its core, same-timestamp transactions colliding in cache decisions and breaking read-your-committed-write monotonicity; double-apply was a missing idempotency in the two-phase-commit internal retry path. Jepsen's value was quantifying the gap between "documented claims" and "actual semantics" with formal testing — Cockroach Labs' docs then said "no possibility of reading stale data," and the tests showed that sentence only held for single keys with healthy clocks.
- 教训
厂商的一致性宣称要看限定词:serializable 不等于 strict serializable,差的两个字是"跨 key 实时序";分布式 SQL 的正确性高度依赖时钟同步,NTP 偏移监控不是可选项,而是正确性基础设施;第三方独立测试(尤其是厂商付费请人来挑刺)在 beta 期做,比 GA 后出事故便宜得多——Cockroach Labs 这次是正面示范。
Read the qualifiers on a vendor's consistency claims: serializable ≠ strictly serializable, and those two missing words are "real-time order across keys"; distributed SQL correctness leans heavily on clock synchronization — NTP offset monitoring is correctness infrastructure, not optional; paying a third party to poke holes (especially the vendor paying someone to find fault) during beta is far cheaper than a post-GA incident — Cockroach Labs set a positive example here.
相关产品:CockroachDB、Google Spanner 相关能力:— 最后核验:2026-10-02
CockroachDB 2024 年砍掉免费 Core 版:年入超 1000 万美元的自建用户按 CPU 核付费 失败教训
开源许可
BSL
选型风险
供应商锁定
- 场景
CockroachDB 早期 Apache 2.0 开源,2019 年转 BSL 1.1(Core 版可免费自建、限制商用 DBaaS),另有功能更多的 Enterprise 版。2024 年 8 月 CEO Spencer Kimball 宣布:随 24.3 版本(11 月 18 日生效)砍掉 Core,合并为单一 Enterprise 许可(CockroachDB Software License,源码可看不可自由用);年入 1000 万美元以下的企业、个人、学生、研究者继续免费(Enterprise Free),超线企业按部署机器的 CPU 核数付费;云托管产品不受影响。
CockroachDB started Apache 2.0 open source, moved to BSL 1.1 in 2019 (Core edition free to self-host, restricted from commercial DBaaS), alongside a fuller-featured Enterprise edition. In August 2024 CEO Spencer Kimball announced that with version 24.3 (effective November 18) Core would be retired and folded into a single Enterprise license (the CockroachDB Software License — source available to read, not to freely use); businesses under $10 million in annual revenue, plus individuals, students, and researchers, stay free (Enterprise Free), while above-the-line businesses pay per CPU core on the machines hosting the database; the cloud product was unaffected.
- 决策
这是厂商的单方面商业决策。动因 Kimball 自己说得很直白:越来越多有规模的公司"将就着用免费 Core、绕开 Enterprise 付费","我们的免费 Core 成了我们最精明的竞争对手"(TechCrunch 专访原话)。社区反应激烈:Percona 联合创始人 Peter Zaitsev 称"CockroachDB 完成了去开源化,正在变成又一个 Oracle";OpenUK CEO Amanda Brock 称对 Cockroach 的转向"并不意外"。
A unilateral commercial decision by the vendor. Kimball stated the motive plainly: a growing number of scaled businesses were "compromising on using the full capabilities of CockroachDB, eschewing the Enterprise license for free usage of Core" — "our 'core' [free] offering has become one of our savviest competitors" (his words in a TechCrunch interview). Community reaction was sharp: Percona co-founder Peter Zaitsev said "CockroachDB has completed its transition away from open source... becoming yet another Oracle"; OpenUK CEO Amanda Brock said the move was "not really a surprise."
- 结果
政策已生效(同样适用于 2024 年 11 月 18 日之后发布的 23.1 及更新版本的补丁)。实际影响分层:小公司反而赚到——免费版现在包含全部 Enterprise 功能;年入超 1000 万美元的自建用户则面临"按核付费"的新账单,且免费资格需每年重新验证营收、遥测不可关闭。证据等级说明:未找到公开的"某大公司因此迁出 CockroachDB"的实证——搜到的多为个人或小项目的选型规避(如 GitHub 上的分布式 SQL 对比文档以此为由排除 CockroachDB),"迁移出走潮"属于传闻级别,此处如实标注为未找到证据。
The policy took effect (it also applies to patches of 23.1 and later issued after November 18, 2024). Impact is tiered: small companies actually gained — the free edition now includes all Enterprise features; self-hosted users above $10M annual revenue face a new per-core bill, and free-tier eligibility must be re-verified against revenue every year with telemetry that cannot be opted out. Evidence grading: no public evidence was found of any specific large company migrating off CockroachDB over this — what surfaced were selection-stage rejections by individuals and small projects (e.g., distributed-SQL comparison docs on GitHub excluding CockroachDB on licensing grounds); talk of a "migration exodus" is rumor-grade and is marked here as no evidence found.
- 机制根因
这是 VC 驱动型"伪开源"数据库的经典剧本:Apache 2.0 获客、 BSL 防云厂商、专有许可变现。CockroachDB 的特殊之处在于 Kimball 承认产品太稳定了——"跑生产几乎不需要支持",于是免费版成了付费版的替代品,厂商只能靠改许可收费。对用户的实质风险不是"今年多付钱",而是许可变更的单方面性:今天 1000 万美元线是恩赐,明天线划在哪、遥测收什么,都是厂商说了算;2025 年底官方进一步把新版本开发转入私有仓库,社区连代码可见性都在收缩。
The classic VC-backed "fauxpen source" playbook: Apache 2.0 for adoption, BSL against cloud vendors, proprietary licensing for monetization. CockroachDB's twist is Kimball's own admission that the product got too stable — "you can go a long time without having any kind of support needs" — turning the free edition into a substitute for the paid one, leaving license changes as the only lever. The substantive risk to users is not "paying more this year" but the unilateral nature of license changes: today's $10M line is a gift, tomorrow's line and telemetry scope are the vendor's call; in late 2025 the company further moved new-version development into private repositories, shrinking even source visibility for the community.
- 教训
选型评估必须把"许可变更风险"单列一条:看贡献者是否集中于一家公司、看历史改许可次数(CockroachDB:Apache 2.0 到 BSL 再到专有,三次);Zaitsev 的忠告值得引用——警惕绝大多数贡献来自单一公司的"开源"项目;社区治理的项目(如 PostgreSQL)是这类风险的对冲。年入接近 1000 万美元线的公司尤其要算清:免费的前提是每年自证没超线,这本身就是合规成本。
Selection evaluations must score "license-change risk" as its own line item: check whether contributions concentrate in one company and count past license changes (CockroachDB: Apache 2.0 → BSL → proprietary, three times); Zaitsev's warning is worth quoting — be cautious about open-source projects where the vast majority of contributions come from a single company that doesn't welcome contributions fairly; community-governed projects (e.g., PostgreSQL) hedge this risk. Companies near the $10M line should price in the compliance cost itself: staying free means proving every year you haven't crossed it.
相关产品:CockroachDB、PostgreSQL(社区版) 相关能力:— 最后核验:2026-10-02
Nubank:信用卡授权系统从内存存储迁到 CockroachDB,on-prem 三地部署(2021) 成功经验
信用卡授权
金融核心
私有化部署
强一致
- 场景
巴西数字银行 Nubank 2019–2021 年客户从 1200 万涨到 4000 万,业务扩张到墨西哥、哥伦比亚。信用卡授权服务(决定每笔交易批不批准)最初跑在 Java 堆内存存储上,多实例之间靠 Kafka 消息同步变更;内存方案运维复杂、扩展性差,撑不住跨国增长。
Brazilian digital bank Nubank grew from 12 million customers in 2019 to 40 million in 2021, expanding into Mexico and Colombia. Its credit card authorization service (which approves or declines every transaction) originally ran on Java-heap in-memory storage, with Kafka messages syncing changes between instances — operationally complex and poorly scalable, unable to carry cross-country growth.
- 决策
Authorizer 团队评估了 Cassandra、CouchDB 等多个方案,最终选择 CockroachDB:授权场景要求强一致(批错一笔就是资金损失)、水平扩展、运维简单。团队在私有化环境三个地点先行部署,把授权应用迁了过去;看中的还有 PostgreSQL 线路兼容带来的低接入摩擦,以及自动化备份、垃圾回收、滚动升级。
The Authorizer team evaluated several options including Cassandra and CouchDB, and chose CockroachDB: authorization demands strong consistency (one wrong approval is a direct financial loss), horizontal scalability, and operational simplicity. The team deployed CockroachDB on-prem in three locations first and migrated the authorization application; PostgreSQL wire compatibility kept developer onboarding friction low, and automated backups, garbage collection, and rolling upgrades reduced ops burden.
- 结果
CockroachDB 成为授权系统的关键任务数据库(Cockroach Labs 厂商博客口径)。公开的量化结果很少:客户数 1200 万到 4000 万是公司增长数字,不是数据库性能数字;授权 QPS、延迟、集群规模均未公开。证据等级说明:选型事实与部署形态来自厂商博客,Nubank 官方未在自家工程博客披露对应细节,引用时须注明口径。
CockroachDB became the mission-critical database behind the authorization system (per Cockroach Labs' vendor blog). Public quantified results are thin: the 12M→40M customer figure is company growth, not a database performance metric; authorization QPS, latency, and cluster size were not disclosed. Evidence grading: the selection facts and deployment shape come from the vendor blog; Nubank itself has published no corresponding detail on its own engineering blog — quote with the source caveat.
- 机制根因
授权是典型的"一致性优先、可用性也不能丢"负载:内存存储加 Kafka 对账本质是最终一致,脑裂时可能出现重复授权或错误拒绝;Cassandra 的最终一致模型需要应用层自己处理冲突,对金融授权是错配。CockroachDB 的 Raft 多副本加可串行化事务把一致性收回数据库内部,range 自动分片让扩容变成"加节点指向集群"。代价是 on-prem 自运维:团队自己扛时钟同步(NTP 偏移超阈值则一致性保证失效,参见 Jepsen 案例)、备份与升级——比内存方案简单,但比托管云数据库重。
Authorization is the classic "consistency first, availability non-negotiable" workload: in-memory storage plus Kafka reconciliation is effectively eventual consistency, and a split brain can produce duplicate authorizations or wrongful declines; Cassandra's eventual-consistency model pushes conflict handling onto the application — a mismatch for financial authorization. CockroachDB's Raft replication plus serializable transactions pull consistency back inside the database, and automatic range sharding turns scaling into "add a node and point it at the cluster." The cost is self-operated on-prem: the team owns clock synchronization (if NTP offset exceeds the threshold, consistency guarantees lapse — see the Jepsen case), backups, and upgrades — simpler than the in-memory setup, but heavier than a managed cloud database.
- 教训
金融授权类场景选型时,"一致性模型"是第一性原理——Cassandra 们不是不好,是错配;厂商博客客户故事的可信度取决于细节颗粒度(这篇有团队名、部署形态、评估过谁,可用但数字要降级引用);on-prem 跑分布式 SQL 省了云账单,但把时钟与运维还给了自己。
For financial authorization workloads, the consistency model is the first principle of selection — Cassandra and friends aren't bad, they're mismatched; a vendor customer story's credibility tracks its granularity of detail (this one names the team, the deployment shape, and the evaluated alternatives — usable, but downgrade the numbers); running distributed SQL on-prem saves the cloud bill but hands clocks and operations back to you.
相关产品:CockroachDB、Apache Cassandra / ScyllaDB 相关能力:多活生存性 —— 丢一个 region 业务不中断 最后核验:2026-10-02
Coinbase:MongoDB 连接风暴后,把高频 KV 查询迁往 DynamoDB(约 2025) 失败教训
突发流量
连接风暴
多库并存
金融交易
- 场景
Users Service is the core of Coinbase's identity platform — a hard dependency of every critical user journey, read-heavy with strict demands on read latency and write correctness. Crypto market traffic is driven by price volatility: bursty and unpredictable, with a single market event amplifying into millions of RPS of reads at the bottom of deep call stacks. The service originally relied entirely on MongoDB, but some requests bypassed the cache and hit the database directly, escalating at traffic peaks into connection storms and MongoDB "death spirals" that took the whole service down once triggered. (
https://www.coinbase.com/blog/scaling-identity-how-coinbase-serves-1-5M-reads-second)
- 决策
Coinbase 转向联邦持久化:先分析访问模式,把访问最频繁的 K-V 查询数据集从 MongoDB 迁到 DynamoDB——DynamoDB 向客户端提供免连接池管理的托管请求接口,不需要维持活跃连接,为不可预测流量下的稳定性能而设计(准确说不是"数据库本身没有状态",而是"连接管理"这个失效维度被托管层拿走了);MongoDB 保留服务文档型查询和灵活模型。代价是 DynamoDB 没有原生唯一约束,Coinbase 自研了一套框架,用辅助表 + 多项事务保证需要唯一性的列上的原子写入;其余文档负载继续留在 MongoDB。
Coinbase moved to federated persistence: after analyzing access patterns, it migrated the most frequently accessed K-V query dataset from MongoDB to DynamoDB — DynamoDB offers clients a managed request interface with no connection-pool management, maintains no active connections, and is designed for stable performance under unpredictable traffic (to be precise, it's not that "the database has no state", but that the "connection management" failure dimension is taken over by the managed layer); MongoDB kept serving document queries and the flexible data model. The cost: DynamoDB has no native uniqueness constraints, so Coinbase built its own framework using auxiliary tables + multi-item transactions to guarantee atomic writes on columns requiring uniqueness; remaining document workloads stayed on MongoDB.
- 结果
架构改造(Fragment API + 联邦存储 + 乐观并发控制 + 负载削减)后,Users Service 在市场行情高峰期可持续承受 150 万+ 读/秒——这是整体架构改造的结果,非 DynamoDB 迁移单项的贡献,迁移时间点未单独披露(以上均为 Coinbase 官方博客《Scaling Identity》口径)。连接风暴量化:高流量期蓝绿部署时每分钟近 60K 新增连接;2020 年 4 月一次长时间事故中,单主机连接尝试数超过 MongoDB 128K 单机上限;自研 Go 连接复用代理 mongobetween 上线后,MongoDB 外连总数降低约 20 倍(以上出自 Coinbase 官方博客《Scaling connections with Ruby and MongoDB》)。DynamoDB 迁移前后的延迟/可用性对比、迁移耗时、成本变化未找到公开数据。
After the architecture overhaul (Fragment API + federated storage + optimistic concurrency control + load shedding), Users Service can sustain 1.5M+ reads/sec through market peaks — the result of the overall overhaul, not of the DynamoDB migration alone, and the migration's timing was not separately disclosed (all per Coinbase's official "Scaling Identity" blog). Connection-storm figures: nearly 60K new connections per minute during blue-green deploys at high traffic; in one prolonged April 2020 incident, connection attempts on a single host exceeded MongoDB's 128K per-machine limit; the homegrown Go connection-multiplexing proxy "mongobetween" cut total external MongoDB connections ~20x (per Coinbase's "Scaling connections with Ruby and MongoDB"). No public data found for pre/post-migration latency/availability comparisons, migration duration, or cost changes.
- 机制根因
MongoDB 是连接型有状态数据库,每个应用进程维持连接池;Coinbase 的 CRuby(GVL)应用每台机器就有 10–20K 外连,蓝绿部署瞬间实例数翻倍、连接数翻倍,形成"突发→绕过缓存打库→连接暴涨→失败→重试再增连接"的正反馈死亡循环,在业务最需要可用性的峰值时刻恰恰最可能触发。DynamoDB 的托管请求模型切断了这条反馈链:客户端没有连接池概念、按需弹性,不存在"连接数"这个失效维度。根本教训不是 MongoDB 弱,而是把"高频 K-V 点查"这种访问模式放在客户端需自管连接池的数据库上是模式错配;DynamoDB 缺唯一约束的代价则通过辅助表 + 事务在应用层补回,是选型时就要算进总拥有成本的隐性成本。
MongoDB is a connection-oriented stateful database with a connection pool per application process; Coinbase's CRuby (GVL) apps held 10–20K external connections per machine, and blue-green deploys instantly doubled instances — and connections — forming a positive-feedback death spiral ("spike → cache bypass → connection explosion → failures → retries add more connections") most likely to trigger exactly when the business needed availability most. DynamoDB's managed request model cuts that feedback loop: no client-side connection-pool concept, elastic on demand, no "connection count" failure dimension. The real lesson isn't that MongoDB is weak, but that placing a "hot K-V point-lookup" access pattern on a database whose connection pools the client must manage itself is a pattern mismatch; the cost of DynamoDB's missing uniqueness constraints was paid back at the application layer via auxiliary tables + transactions — a hidden cost that belongs in TCO from day one.
- 教训
突发流量场景下优先评估"连接管理模型":客户端需自管连接池的数据库,在峰值会把连接数放大成雪崩;把"连接管理"列为和延迟/吞吐同级的评估项;不要让一个数据库包揽一切,先做访问模式分析再按模式拆分负载;算清隐性成本,DynamoDB 的唯一约束要自研框架,没有免费的弹性。
For bursty traffic, evaluate the "connection management model" first: databases whose connection pools the client must manage itself can amplify connection counts into avalanches at peaks — rank "connection management" alongside latency/throughput as a selection criterion. Don't let one database do everything: analyze access patterns first, then split load by pattern. Account for hidden costs: DynamoDB's uniqueness constraints need a homegrown framework — there is no free elasticity.
相关产品:Amazon DynamoDB、MongoDB 相关能力:Serverless 零运维:流量不可预测时"先跑起来"的最短路径、单表设计的心智税:"access pattern 先行"既是超能力也是枷锁 最后核验:2026-10-01
Biogen:本地 HPC 一周宕机后迁云,70 万变异注释从 2 周缩到 15 分钟 成功经验
基因组学
本地迁云
Delta Lake
药物靶点发现
- 场景
Biogen 用人类遗传证据给在研管线排序、发现新基因靶点,核心数据源是英国生物银行(UK Biobank)50 万志愿者的健康与基因组数据,PB 级。但本地数据中心撑不住了:存储不够、网络带宽传不动这么大数据,2018 年其高性能计算集群直接宕机了一周。基因组技术与信息学高级总监 David Sexton 的原话是"我们真的需要一种新的数据范式"。旧范式下,一条注释 70 万个变异的管道要跑 2 周。
Biogen ranks its drug portfolio and discovers new gene targets using human genetic evidence, anchored on petabytes of health and genomic data from 500,000 UK Biobank volunteers. Its on-prem data center could not keep up: insufficient storage, network bandwidth that could not move that much data — and in 2018 its high-performance compute cluster went down for a full week. David Sexton, senior director of genome technology and informatics, put it bluntly: "We really needed a new data paradigm." Under the old paradigm, one pipeline annotating 700,000 variants took 2 weeks.
- 决策
Biogen 与 DNAnexus 和 Databricks 合作,把本地基础设施整体迁到 AWS,用 Databricks for Genomics 运行时覆盖从原始数据处理到大规模统计分析的全谱需求;用 Delta Lake 重做变异注释管道;按基因组位置做重度垂直分区(数千列的元数据下这是关键),并把 Spark Hive Metastore 接入其平台访问控制模型做数据安全。
Biogen partnered with DNAnexus and Databricks to move its on-prem infrastructure to AWS, using the Databricks for Genomics runtime to cover the full spectrum from raw data processing to large-scale statistical analysis; rebuilt the variant-annotation pipeline on Delta Lake; applied heavy vertical partitioning by genomic location (critical with thousands of metadata columns); and wired the Spark Hive Metastore into its platform access-control model for data security.
- 结果
Databricks 官方客户页宣称:原来 2 周处理 70 万变异的管道,优化后约 15 分钟注释 200 万变异;基于 UK Biobank 数据找到 6 个基因中影响人类寿命的变异,识别出 2 个新药靶点,并产出阿尔茨海默病、帕金森病相关洞见。数字为 Databricks 厂商口径(与 DNAnexus 联合实施,效果不可单独归因于 Databricks),引用须注明。2018 年宕机一周与 Sexton 的引言有公开出处。
Databricks' official customer story claims the pipeline that took 2 weeks for 700,000 variants now annotates 2 million variants in about 15 minutes; UK Biobank analysis surfaced lifespan-impacting variants across six genes, two new drug targets, and insights into Alzheimer's and Parkinson's disease. Figures are Databricks' vendor claims from a joint implementation with DNAnexus — the effect cannot be attributed to Databricks alone, so quote with the caveat. The 2018 week-long outage and Sexton's quote are publicly sourced.
- 机制根因
本地 HPC 的病根是"存储、带宽、算力"三者刚性绑定:数据进不来,算力再强也白搭。迁云把三者解耦——S3 存 PB 级数据、带宽弹性、算力按需;Delta Lake 在对象存储上提供 ACID 与版本化,让基因组这种"写一次、反复重注释"的负载敢做大规模重算;按基因组位置分区把"按区间查变异"变成局部扫描。这是典型的"云不是更便宜,而是让不可能变为可能"。
On-prem HPC's disease was rigid coupling of storage, bandwidth, and compute: if data cannot get in, the strongest compute is useless. The cloud move decoupled all three — S3 holds petabytes, bandwidth is elastic, compute is on-demand; Delta Lake brings ACID and versioning to object storage, so a "write once, re-annotate repeatedly" genomics workload can afford massive recomputation; partitioning by genomic location turns "query variants by interval" into local scans. A textbook case of "the cloud is not cheaper, it makes the impossible possible".
- 教训
突发性科研负载(一个队列空了、全员同时重跑)和本地 HPC 的固定容量是天然冲突,弹性比峰值性能更重要;基因组数据的分区键必须按生物学访问模式(基因组位置)设计,按时间/随意分区的表在关联分析时会付出数倍代价;迁云的真实收益常常不在账单上,而在"宕机一周"这种风险的消除——TCO 里要给可靠性定价。
Bursty research workloads (an empty queue, everyone re-running at once) inherently conflict with fixed on-prem capacity — elasticity beats peak performance. Genomic partition keys must follow biological access patterns (genomic location); tables partitioned by time or arbitrarily pay multiples on association analysis. And the real payoff of a cloud move is often not on the bill but in eliminating risks like "a week of downtime" — price reliability into TCO.
来源
Databricks 官方客户故事《Customer Story: Biogen》(厂商口径,含 2 周→15 分钟、2 个新靶点数字
Databricks official customer story "Customer Story: Biogen" (vendor claim
source of the 2-weeks-to-15-minutes and two-new-targets figures
—
相关产品:Databricks 相关能力:SQL+ML 同一数据底座 最后核验:2026-10-02
Comcast:语音遥控背后的 Lakehouse——PB 级遥测数据与数百个模型的统一平台 成功经验
实时遥测
语音交互
模型生命周期管理
流批一体
- 场景
Comcast 是连接数千万家庭的传媒科技公司,其娱乐系统每天产生数十亿事件,2000 多万支语音遥控器带来 PB 级的语音与视频遥测数据,需要做 sessionization 才能分析。旧架构有三重病:数据量撑爆 IT 基础设施;管道脆弱、频繁失败且难以恢复,大量小文件拖慢下游机器学习的数据摄入;全球分散的数据科学家用不同语言写脚本,代码难以共享复用;数百个模型从训练到部署全靠手工,慢且不可复制;模型还要部署到云、本地甚至终端设备等互不相干的环境。
Comcast, a media and technology company connecting tens of millions of households, generates billions of events daily from its entertainment system, with 20M+ voice remotes producing petabytes of voice and video telemetry that must be sessionized before analysis. The old architecture had three compounding problems: data volumes strained IT infrastructure; fragile pipelines failed frequently and were hard to recover, with masses of small files slowing data ingestion for downstream ML; globally dispersed data scientists wrote scripts in different languages with little code sharing or reuse; hundreds of models went from training to deployment entirely by hand — slow and unrepeatable; and models had to be deployed to disjoint environments spanning cloud, on-prem, and even end devices.
- 决策
Comcast 把整条链路搬上 Databricks Data + AI 平台:用 Delta Lake 做视频/语音原始遥测的摄入、 enrichment 和初加工;用托管 MLflow(经 Kubeflow 做模型服务)管理数百个模型的完整生命周期;分析师侧继续用 Tableau 消费数据。选型逻辑是"一个平台同时解决数据工程和 ML 工程",而不是再拼一套拼凑方案。
Comcast moved the whole chain onto the Databricks Data + AI Platform: Delta Lake for ingestion, enrichment, and initial processing of raw video/voice telemetry; managed MLflow (with model serving via Kubeflow) for the full lifecycle of hundreds of models; analysts kept consuming data through Tableau. The selection logic was "one platform solving data engineering and ML engineering together" rather than assembling another patchwork.
- 结果
Databricks 官方客户页宣称:Delta Lake 优化数据摄入后,计算资源从 640 台机器降到 64 台,计算成本降为 1/10 且性能更好;为 200 名用户做平台 onboarding 所需的运维人力从 5 人降到 0.5 人;模型部署从"数周"缩短到"数分钟"。以上数字均为 Databricks 厂商口径,Comcast 官方工程博客未披露对应数字,也未找到第三方独立复现,引用须注明口径。公开可查的定性结果是支撑其获奖的语音交互体验的模型迭代速度明显加快。
Databricks' official customer story claims: Delta Lake ingestion optimization cut compute from 640 machines to 64 — a 10X compute cost reduction with better performance; DevOps headcount for onboarding 200 users fell from 5 to 0.5 FTE; model deployment shrank from "weeks" to "minutes". All figures are Databricks' vendor claims; Comcast's own engineering blog discloses no corresponding numbers, and no independent third-party reproduction exists — quote them with the source caveat. The publicly verifiable qualitative outcome is materially faster iteration on the models behind its award-winning voice experience.
- 机制根因
根因有三层。其一是小文件税:流式写入在对象存储上产生海量小文件,查询规划与 listing 开销吞掉大部分算力,Delta Lake 的文件优化(compaction)直接对冲了这一点。其二是流批一体:同一份 Delta 表同时服务实时摄入与历史回填,避免了 Lambda 架构的两套管道。其三是托管红利:autoscaling + spot 实例把"按峰值买机器"变成"按需用算力",640→64 本质是把闲置容量还给了云。代价是把核心数据链路绑定在 Databricks 的托管运行时与 DBU 计价上,迁移出去的成本不低。
Three layers. First, the small-file tax: streaming writes produce vast numbers of tiny files on object storage, and query planning plus listing overhead eats most of the compute — Delta Lake's file optimization (compaction) directly offsets this. Second, unified batch and streaming: one Delta table serves both real-time ingestion and historical backfill, eliminating Lambda architecture's dual pipelines. Third, the managed dividend: autoscaling plus spot instances turn "buy for peak" into "pay for what you use" — 640→64 is essentially returning idle capacity to the cloud. The cost is binding the core data path to Databricks' managed runtime and DBU pricing, making a future exit expensive.
- 教训
Lakehouse 的真正价值不是"快",而是把数据工程、分析和 ML 收敛到同一份数据上——Comcast 的数百个模型如果还在多套系统间搬运数据,部署周期不可能从周降到分钟;小文件治理是数据湖的隐性税,不做 compaction,算力有一半花在"找文件"上;托管平台的降本主要来自弹性(autoscaling/spot),而不是引擎更快,选型时要把"运维人力释放"计入 TCO。
The lakehouse's real value is not "speed" but converging data engineering, analytics, and ML onto the same data — had Comcast's hundreds of models still been shuttled between systems, deployment cycles could never have gone from weeks to minutes. Small-file governance is the hidden tax of data lakes: without compaction, half your compute goes to "finding files". Managed-platform savings come mostly from elasticity (autoscaling/spot), not a faster engine — count "freed ops headcount" in TCO during selection.
来源
Databricks official customer story "Customer Story: Comcast" (vendor claim
—
相关产品:Databricks 相关能力:SQL+ML 同一数据底座 最后核验:2026-10-02
Condé Nast:37 个品牌的数据孤岛收敛,基础设施年省约 600 万美元 成功经验
数据孤岛治理
成本优化
统一治理
媒体个性化
- 场景
Condé Nast 旗下有 Vogue、The New Yorker、GQ 等 37 个品牌,每个品牌各自为战形成数据孤岛。数据工程高级总监 Nana Yaw Essuman 的原话是"我们在全公司范围内都挣扎着打破数据孤岛"。碎片化的结果是:分析效率低、决策慢、跨品牌做个性化内容推荐时连"同一个用户"都对不齐——在媒体业拼个性化体验的竞争里,这是致命伤。
Conde Nast houses 37 distinct brands — Vogue, The New Yorker, GQ among them — each operating its own data fiefdom. Nana Yaw Essuman, senior director of data engineering, said it plainly: "We struggled to find ways to break out of our data silos across the organization." Fragmentation meant slow analysis, slow decisions, and an inability to resolve "the same user" across brands for personalized content — a fatal weakness in a media industry competing on personalized experience.
- 决策
Condé Nast 选择 Databricks Data + AI 平台做"统一消费者视图":lakehouse 架构承载全品牌数据管道;用 Unity Catalog 做统一的数据访问治理;用 Databricks SQL 让分析师直接在治理后的数据上做分析与决策。选型逻辑不是"换个更快的查询引擎",而是"先有统一可信的数据底座,个性化才有地基"。
Conde Nast chose the Databricks Data + AI Platform for a "unified view of the consumer": the lakehouse architecture carries all brands' data pipelines; Unity Catalog provides unified data-access governance; Databricks SQL lets analysts analyze and decide directly on governed data. The selection logic was not "a faster query engine" but "a unified, trusted data foundation first — personalization needs ground to stand on".
- 结果
Databricks 官方客户页宣称,采用后年度基础设施成本降低约 600 万美元,同时实现全球一致的报表口径与更高的数据准确性。该数字为 Databricks 厂商口径,未披露统计口径(是否含人力、是否对比迁云前本地成本),也未找到第三方独立复现,引用须注明口径。定性层面可确认的是团队响应市场趋势的速度与个性化内容交付能力提升。
Databricks' official customer story claims annual infrastructure costs fell by approximately $6 million, alongside globally consistent reporting and higher data accuracy. The figure is Databricks' vendor claim — the accounting basis is undisclosed (whether it includes headcount, whether it is measured against pre-migration on-prem costs), and no independent third-party reproduction exists, so quote it with the source caveat. What is qualitatively verifiable is faster response to market trends and stronger personalized content delivery.
- 机制根因
孤岛的本质是"每个品牌一套烟囱",成本不在存储而在重复建设:37 套管道、37 套口径、37 次治理。Lakehouse 把 37 个烟囱收敛成一份 Delta 数据 + 一套 Unity Catalog 治理,边际成本从"每新增一个品牌加一套"变成"加一张表"。600 万美元的节省大概率主要来自关停冗余系统与合并运维,而非查询更快——这是"整合降本"而非"性能降本",两者的决策含义完全不同。
Silos are essentially "one chimney per brand", and the cost is not storage but duplicated construction: 37 pipelines, 37 metric definitions, 37 governance efforts. The lakehouse converges 37 chimneys into one Delta dataset plus one Unity Catalog governance layer; marginal cost goes from "each new brand adds a stack" to "adds a table". The $6M saving most likely comes from decommissioning redundant systems and consolidating operations, not from faster queries — this is consolidation-driven savings, not performance-driven savings, and the two imply very different decisions.
- 教训
数据孤岛的税是隐性的:不只体现在多花的机器钱,更体现在"同一个用户对不齐"导致做不出的业务;治理(Unity Catalog 这类)必须和平台迁移同步做,先迁数据后补治理等于给孤岛换了个新地址;评估 lakehouse 收益时,把"关停了多少旧系统"放在"查询快了多少"前面算,后者是锦上添花,前者才是大头。
The data-silo tax is invisible: it shows up not only as extra machine spend but as the business you cannot build because "the same user" cannot be resolved. Governance (Unity Catalog-class tooling) must move together with platform migration — migrating data first and bolting on governance later just gives the silos a new address. When evaluating lakehouse returns, count "how many legacy systems were decommissioned" before "how much faster queries got" — the latter is a bonus, the former is the main prize.
来源
Databricks official customer story "Customer Story: Conde Nast" (vendor claim
—
source of the ~$6M figure
—
相关产品:Databricks 相关能力:Unity Catalog 统一治理 最后核验:2026-10-02
Regeneron:40 万例外显子组 + 电子病历,基因组查询从 30 分钟降到 3 秒 成功经验
基因组学
药物靶点发现
ETL 加速
PB 级分析
- 场景
Regeneron 遗传学中心(RGC)建了全球最全面的遗传数据库之一,把 40 多万人的外显子测序数据与其电子健康病历配对,用于发现新药靶点。但数据高度分散在约 10TB 的基因组 + 临床数据集中,800 多亿个数据点让 legacy 架构扩不动:团队光是做 ETL 就要花几天(最长三周),全量查询一次跑 30 分钟,数据科学家根本不敢做"全表扫描式"的探索——而这恰恰是基因关联发现最需要的工作方式。
The Regeneron Genetics Center (RGC) built one of the world's most comprehensive genetics databases, pairing sequenced exomes with electronic health records from more than 400,000 people to discover new drug targets. But the data sat decentralized across roughly 10TB of genomic and clinical datasets, and 80B+ data points overwhelmed the legacy architecture: teams spent days — up to three weeks — just on ETL, and a full-dataset query took 30 minutes. Data scientists simply avoided "full-table-scan style" exploration, which is exactly the working mode gene-association discovery needs most.
- 决策
RGC 把分析栈搬到跑在 AWS 上的 Databricks:用 Spark 驱动的管道重做 ETL,用交互式 workspace 让生物信息学家、数据科学家和计算生物学家在同一份数据上协作。选型关键是"弹性算力 + 同一份数据的交互式分析",而不是再买一台更大的本地 HPC。
RGC moved its analytics stack to Databricks on AWS: Spark-powered pipelines rebuilt ETL, and interactive workspaces let bioinformaticians, data scientists, and computational biologists collaborate on the same data. The selection hinged on "elastic compute plus interactive analysis on one copy of the data", not on buying a bigger on-prem HPC.
- 结果
Databricks 官方客户页宣称:全数据集查询从 30 分钟降到 3 秒(600 倍);ETL 从 3 周缩短到 2 天。该数字为 Databricks 厂商口径,Regeneron 官方未披露可交叉验证的基准细节,引用须注明口径。定性层面可确认的是团队得以支撑"以前不可能"的新分析方式,把精力从搭集群转向找靶点。
Databricks' official customer story claims full-dataset queries dropped from 30 minutes to 3 seconds (600x), and ETL from 3 weeks to 2 days. These figures are Databricks' vendor claims; Regeneron has disclosed no independently verifiable benchmark details, so quote them with the source caveat. What is qualitatively verifiable is that the team can now run analyses that were previously impossible, shifting effort from cluster-tending to target-finding.
- 机制根因
基因组数据的访问模式是"宽表 + 全基因组扫描",legacy 架构的瓶颈不在 CPU 而在数据布局与调度:分散的小批量 ETL 导致数据科学家排队等数仓窗口。Spark 的分布式扫描把 800 亿数据点的全表查询变成可交互操作;托管集群管理把 DevOps 工作自动化,相当于把原来花在"等集群、调集群"上的时间还给了科研。代价是这类扫描型负载对算力极敏感,DBU 账单随探索频率大致成比例增长(取决于 workload 形状与自动终止配置),成本 governance 必须跟上。
Genomic access patterns are "wide tables plus genome-wide scans" — the legacy bottleneck was data layout and scheduling, not CPU: fragmented batch ETL forced scientists to queue for warehouse windows. Spark's distributed scan turned full-table queries over 80 billion data points into interactive operations; managed cluster automation returned the time previously spent "waiting on and tuning clusters" to research. The cost: scan-heavy workloads are extremely compute-sensitive, so the DBU bill grows roughly in proportion to exploration frequency (depending on workload shape and auto-termination settings) — cost governance has to keep up.
- 教训
科研型数据团队的瓶颈往往不是算法,而是"问一个问题要等多久"——30 分钟到 3 秒的差别决定了科学家敢不敢做探索性分析;ETL 从 3 周到 2 天说明很多"数据准备慢"不是数据本身的问题,是架构把简单任务复杂化了;把 HPC 思维(买大机器)换成云原生思维(弹性 + 共享数据),是基因组学规模化的前提,但要为"探索越便宜、探索越频繁"的账单做好预算机制。
For research data teams the bottleneck is rarely the algorithm, it is "how long one question takes" — the difference between 30 minutes and 3 seconds decides whether scientists dare to explore. ETL going from 3 weeks to 2 days shows that much "slow data prep" is not a data problem but an architecture making simple tasks complicated. Swapping the HPC mindset (buy bigger machines) for a cloud-native one (elasticity plus shared data) is the precondition for genomics at scale — but budget for the dynamic where cheaper exploration means more frequent exploration.
来源
Databricks official customer story "How Regeneron is discovering new treatments with AI" (vendor claim
—
相关产品:Databricks 相关能力:SQL+ML 同一数据底座 最后核验:2026-10-02
Riot Games:1 亿玩家的行为数据,14 名数据科学家"人手一个集群" 成功经验
游戏遥测
玩家行为分析
自助式集群
数据科学协作
- 场景
2017 年的 Riot Games,《英雄联盟》拥有约 1 亿玩家。高级数据科学家 Wesley Kerr 在当年 Spark Summit 主题演讲中披露:约 2% 的对局存在严重不良行为(仇恨言论、种族/性别歧视),团队要用数据科学改善玩家体验、识别并治理这些行为。数据规模是"玩家级"的全量行为数据,全部流入 S3 上的 Hive 数仓;旧的交互方式是直接查 Hive,又慢又不适合数据科学家的探索式工作。
In 2017, Riot Games' League of Legends had roughly 100 million players. Senior data scientist Wesley Kerr disclosed in his Spark Summit keynote that year that about 2% of games contained serious abuse — hate speech, racism, sexism — and the team used data science to improve player experience and police that behavior. The data was player-level behavioral data at full scale, all flowing into a Hive warehouse on S3; the old interaction pattern was querying Hive directly — slow and unfit for data scientists' exploratory work.
- 决策
Kerr 的原话是"We rely on DataBricks for all of our deployments"。约 14 名数据科学家每人管理自己的集群,随用随开、随手可关,在 Databricks + Spark 上找数据、做分析;Hive 数仓继续做 S3 上的存储层,查询层切到 Databricks/Spark,因为"对我们的数据科学场景快得多"。这是一个 2017 年的早期 lakehouse 形态:对象存储做存储、弹性 Spark 做算力,中间没有重型数仓。
Kerr's own words: "We rely on DataBricks for all of our deployments." About 14 data scientists each managed their own cluster — spun up on demand, torn down by hand — finding data and analyzing through Databricks and Spark; the Hive warehouse stayed as the S3 storage layer while the query layer moved to Databricks/Spark because it was "much quicker for our data science use cases". This was an early-2017 lakehouse shape: object storage for storage, elastic Spark for compute, no heavy warehouse in between.
- 结果
数据科学家获得自助式集群管理能力,不再排队等数仓资源;玩家行为分析(包括聊天文本的不良行为识别)得以规模化。需要诚实说明的是:该演讲未披露任何性能数字(查询提速倍数、集群规模、成本),"查证为无"——只有定性结论可引用。另外该环节由 Databricks 赞助(SiliconANGLE 页面底部有披露声明,称赞助方不干预编辑内容),引用时应一并注明。
Data scientists gained self-service cluster management and stopped queueing for warehouse resources; player behavior analysis (including chat-text abuse detection) scaled up. The honest caveat: the keynote disclosed no performance numbers (no query speedup multiples, cluster sizes, or costs) — verified absent, so only qualitative conclusions may be cited. Also note the segment was sponsored by Databricks (disclosed at the bottom of the SiliconANGLE page, which states sponsors have no editorial control) — cite with that disclosure attached.
- 机制根因
核心是"存算分离 + 自助算力":S3 存全量行为数据保证便宜可扩展,Databricks 集群只为活跃的分析会话付费,14 个人互不阻塞。对比当时主流的"Hadoop 集群排队"模式,瓶颈从"等资源"变成"想问题"。不良行为识别这类"先探索、再建模"的负载,特别吃交互式分析的响应速度——这正是 notebook + 弹性 Spark 的甜点区。代价是 2017 年还没有 Delta Lake/Unity Catalog 这类治理层,数据质量与权限全靠团队自律。
The core is "disaggregated storage and compute plus self-service power": S3 holds the full behavioral dataset cheaply and scalably, Databricks clusters bill only for active analysis sessions, and 14 people never block each other. Compared with the era's mainstream "queue for the Hadoop cluster" pattern, the bottleneck moved from "waiting for resources" to "thinking about problems". Abuse detection — an "explore first, model later" workload — lives on interactive analysis responsiveness, exactly the notebook-plus-elastic-Spark sweet spot. The cost: in 2017 there was no Delta Lake or Unity Catalog governance layer, so data quality and access control rested entirely on team discipline.
- 教训
数据团队的规模化不靠"更大的中央集群",靠"每人都能自助开算力"——14 个数据科学家 14 个集群,协作效率来自不互相等待;存储与查询解耦(S3 + Spark)是后来 lakehouse 的雏形,2017 年就有人在生产里跑通了;引用厂商赞助环节的演讲要做来源分级:事实部分(用了 Databricks、人手集群)可信度高,效果数字缺失就老实写"查证为无",不脑补。
Scaling a data team does not come from "a bigger central cluster" but from "everyone can spin up compute themselves" — 14 data scientists with 14 clusters collaborate well because nobody waits on anyone. Decoupling storage from querying (S3 + Spark) was the prototype of the later lakehouse, already running in production in 2017. And when citing a vendor-sponsored talk, grade the source: factual claims (they used Databricks, one cluster per person) carry high confidence; where effect numbers are missing, write "verified absent" honestly instead of filling the gap.
来源
SiliconANGLE《A game of data science: the analytics architecture behind Riot Games》(2017-07-20,Spark Summit 2017 主题演讲 + theCUBE 专访
SiliconANGLE "A game of data science: the analytics architecture behind Riot Games" (2017-07-20
Spark Summit 2017 keynote plus theCUBE interview
—
相关产品:Databricks 相关能力:SQL+ML 同一数据底座 最后核验:2026-10-02
Seemplicity:默认配置跑生产,日账单 2000 美元,几天砍到 500 美元 失败教训
成本失控
DBU 计价
实例选型
FinOps
- 场景
Seemplicity 是一家做 AI 驱动漏洞暴露面管理的网络安全公司,核心是一条 7×24 连续运行的 Databricks 作业(每小时跑一轮,主数据入流)加若干日调度管道。随着功能不断上线,DBU 账单持续走高,团队起初把这当成"业务增长的自然代价",直到某个月突破了与 Databricks 签的月度 commitment(包量),才停下来看钱到底花在了哪。作者 Igal Drayerman 在公司工程博客自述:几天集中优化,日计算账单从 2000 美元降到 500 美元,降幅 75%。
Seemplicity is a cybersecurity company offering an AI-driven exposure management platform. Its core is a single continuously running Databricks job (one run per hour, ingesting the main data inflow) plus a handful of daily scheduled pipelines. As features shipped, the DBU bill kept climbing, and the team initially treated it as "the natural cost of growth" — until they breached their monthly Databricks commit, which forced a look at where the money was actually going. Igal Drayerman writes on the company engineering blog: a few days of focused work cut the daily compute bill from $2,000 to $500, a 75% reduction.
- 决策
复盘发现账单里全是"默认配置税"。第一,Photon 一直开着:收 2 倍 DBU 费率、提速超 2 倍才回本,而他们以 I/O 密集(读写 Delta 表、流式摄入)和内存密集(GraphFrame)为主,Photon 达不到 2 倍加速——关掉等于白捡钱。第二,实例选反了:Databricks 的 DBU 费率按实例家族区分,通用型单价反而高于内存优化型;换成内存优化型后内存翻倍(64GB)、单价减半,主管道耗时几乎不变。第三,Autoloader 默认目录 listing 模式每次扫描 S3 要 15 分钟,切事件通知模式(SNS+SQS)后降到 3 分钟。第四,Delivery Guard 的 resync 逻辑从独立管道改成主管道内联,独立管道成本归零。第五,删掉没人看的指标查询,有用的挪到数据已物化的阶段。
The review found the bill was full of "default-setting taxes". First, Photon had been left on: Photon carries a 2x DBU multiplier and only breaks even past 2x speedup, but their workloads are I/O-bound (reading/writing Delta tables, streaming ingestion) and memory-intensive (GraphFrame) — Photon never reached the 2x bar, so turning it off was free money. Second, the wrong instance family: they had always used general-purpose instances, yet Databricks DBU rates differ by instance family, and general-purpose carries a higher DBU rate than memory-optimized; switching to memory-optimized doubled RAM (64GB) at roughly half the Databricks hourly rate, with the main pipeline's runtime barely moving. Third, Autoloader ran in default directory-listing mode, spending 15 minutes per scan on S3 LIST calls; switching to event-notification mode (SNS+SQS push) cut it to 3 minutes. Fourth, the Delivery Guard resync logic lived in a separate pipeline — inlining it into the main pipeline drove the standalone pipeline's ongoing cost to near zero. Fifth, metric queries nobody watched were deleted, and useful ones were relocated to pipeline stages where data was already materialized.
- 结果
日账单 2000→500 美元(公司自述口径,非第三方审计)。还发现反直觉规律:autoscaling 下把实例从 64GB/8 核砍半到 32GB/4 核,成本只降约 35% 而非 50%——集群会自动多拉节点补齐,降本是非线性的。账单回到 commitment 之内,团队才保住"继续把 workload 迁进 Databricks"的战略选项——作者原话:按原成本轨迹,他们会被迫重估整个平台选型。
Daily bill $2,000 → $500 (company engineering blog's self-reported figures, not third-party audited). Along the way they hit a counter-intuitive law: under autoscaling, halving instances again from 64GB/8 cores to 32GB/4 cores cut cost only ~35%, not 50% — the cluster simply provisions more nodes to compensate, so savings are nonlinear. Back under commit, the team preserved the strategic option of "keep migrating workloads into Databricks" — in the author's words, the old cost trajectory would have forced a full platform re-evaluation.
- 机制根因
DBU 计价的坑在于"单价不透明且与实例家族挂钩":同样的 vCPU,通用型实例的 DBU 费率高于内存优化型,直觉("通用型最划算")在这里是错的。Photon 的 2 倍乘数是"按开启收费"而非"按加速效果收费",I/O bound 的负载开 Photon 等于纯交税。Autoloader 目录 listing 模式的 LIST 请求在深目录树下既慢又贵,是流式摄入的经典暗坑。根因总结:团队把"能跑"当成了"跑得对",没有任何人对 DBU 账单负责,直到 commitment 被突破。
The DBU pricing trap is that "rates are opaque and tied to instance family": for the same vCPU, general-purpose instances carry a higher DBU rate than memory-optimized ones — intuition ("general-purpose is the safe default") is wrong here. Photon's 2x multiplier is charged "for being on", not "for speedup delivered", so I/O-bound workloads running Photon pay pure tax. Autoloader's directory-listing mode turns S3 LIST requests over deep trees into a slow, expensive classic streaming-ingestion pitfall. In short: the team confused "it runs" with "it runs right", and nobody owned the DBU bill until the commit broke.
- 教训
Databricks 的自由度(实例家族、Photon、摄入模式、集群类型)就是"旋钮地狱":每个默认选项都在默默收税,不 benchmark 就等于默认交税;FinOps 不是财务的事,是数据工程的事——system.billing.usage 这类系统表要有人定期看,tag 要打全,预算告警要在超支前响;成本优化最大的杠杆往往不是重构架构,而是"把选错的旋钮拨回去",几天工作换 75% 降幅,ROI 远高于任何新功能。
Databricks' freedom (instance families, Photon, ingestion modes, cluster types) is exactly "knob hell": every default quietly taxes you, and skipping benchmarks means paying by default. FinOps is not finance's job, it is data engineering's — system tables like system.billing.usage need regular readers, tags must be complete, and budget alerts must fire before overruns. The biggest cost lever is usually not re-architecting but "turning the wrong knobs back" — days of work for a 75% cut beats the ROI of any new feature.
来源
Seemplicity company engineering blog "How We Reduced Our Databricks Jobs Compute Costs by 75% - And You Can Too" by Igal Drayerman (2026-05-07
—
相关产品:Databricks 相关能力:Photon 向量化引擎、"旋钮地狱":自由度的代价是工程师小时 最后核验:2026-10-02
Datadog 的数据库选型实践(2024–2025):负载分离与三代时序存储 成功经验
负载分离
CDC
多租户
异步复制
- 场景
On Datadog's shared Postgres, once a single org's metrics tables crossed the 50,000 threshold, page loads slowed and facet filters became unreliable — real-time search and faceted aggregation are fundamentally different workloads from OLTP. Meanwhile, real-time timeseries storage went through two generations: gen-one Cassandra had strong write scalability but couldn't support the breadth and complexity of real-time queries needed for alerting and analysis, and struggled returning large datasets; gen-two Redis was fast and flexible but single-threaded, forced tradeoffs between snapshotting and durability, and suffered rare but severe memory-management failure modes. (
https://www.datadoghq.com/blog/engineering/rust-timeseries-engine/)
- 决策
两条线。搜索线:把搜索查询从共享 Postgres 搬走,复制过程中做反范式化,落到专用搜索平台;基于 Debezium + Kafka + Elasticsearch,用 Temporal 自动化整条 pipeline 的开通,以 Schema Registry(Avro)管多租户 schema 演进。时序线:自研第三代 Rust 实时时序存储引擎(RTDB),对 I/O、内存布局、CPU 利用率全栈可控。
Two tracks. Search track: move search queries off the shared Postgres, denormalizing during replication into a dedicated search platform; built on Debezium + Kafka + Elasticsearch, with Temporal automating end-to-end pipeline provisioning and a Schema Registry (Avro) managing multi-tenant schema evolution. Timeseries track: built a third-generation, purpose-built Rust real-time timeseries storage engine (RTDB) with full-stack control over I/O, memory layout, and CPU usage.
- 结果
Per the official blog, search page loads dropped from ~30s to ~1s (up to 97% reduction) with replication lag held around 500ms; the platform then grew into a company-wide capability: Postgres-to-Postgres (unwinding the shared monolith), Postgres-to-Iceberg (event-driven analytics), Cassandra source onboarding, and cross-region Kafka replication. (
https://www.datadoghq.com/blog/engineering/cdc-replication-search/)
- 机制根因
核心洞察是"负载分离"——OLTP 与搜索/聚合对存储的要求正交,硬塞进一个库里两边都受罪;CDC(PG logical replication + Debezium + Kafka)是解耦手段,复制时反范式化一次,下游各取所需;异步复制被坦然接受(搜索场景 500ms 延迟完全可接受),不为不需要的地方付同步代价;时序三代演进则说明:当 workload 足够特殊(超高写 + 实时查询 + 大结果集),通用存储的"木桶短板"会逐个暴露,自研是算过账的选择,不是情怀。
The core insight is "workload separation" — OLTP and search/aggregation demand orthogonal things from storage, and forcing both into one store punishes both sides. CDC (Postgres logical replication + Debezium + Kafka) is the decoupling mechanism: denormalize once during replication, let each downstream consumer take what it needs. Async replication is accepted without apology (500ms lag is perfectly fine for search) — don't pay synchronization costs where they aren't needed. The three-generation timeseries evolution shows the other side: when a workload is special enough (extreme writes + real-time queries + large result sets), a general-purpose store's shortest staves reveal themselves one by one, and building your own becomes the accounted-for choice, not an indulgence.
- 教训
先按 workload 拆分问题,再为每个 workload 选存储;CDC + 反范式化是"保 OLTP、解放查询"的标准解法,500ms 级延迟对搜索类场景通常可接受;自研存储的正当性来自 workload 的特殊性证明(三代试错),而不是第一天就造轮子。
Split the problem by workload first, then pick storage per workload; CDC plus denormalization is the standard answer for "protect the OLTP, liberate the queries," and ~500ms lag is usually acceptable for search-class scenarios. The legitimacy of self-built storage comes from proving the workload's specialness (three generations of trial), not from building from day one.
相关产品:Apache Cassandra / ScyllaDB、PostgreSQL(社区版) 相关能力:LSM 追加写:吃下"写多读少、只追加"的消息流 最后核验:2026-10-01
Discord 换掉 Cassandra:JVM GC 与 p99 延迟 失败教训
长尾延迟 p99
JVM GC
热点分区
运维负担
- 决策
Cassandra 跑到 2022 年,迁往 API 兼容的 ScyllaDB;自研 Rust 迁移器与请求合并的数据服务层。
The Cassandra chosen in 2017 was replaced in 2022 with API-compatible ScyllaDB; Discord built a Rust-based migrator plus a request-coalescing data-service layer.
- 结果
177 节点压缩到 72;读 p99 从 40–125ms 降到 15ms,写 p99 从 5–70ms 波动收敛为稳定 5ms;自研迁移器 320 万 msg/s、9 天完成全量迁移。数字为 Discord 自报,属单方口径。
177 nodes shrank to 72; read p99 fell from 40–125ms to 15ms, write p99 from fluctuating 5–70ms to a stable 5ms; the custom migrator moved 3.2M msg/s and finished the full migration in 9 days (the Spark-based plan was estimated at 3 months). Figures are from Discord's own talks and blog — self-reported.
- 机制根因
JVM GC 是运行时层面的尾延迟源,调参只能缓解、无法根除;ScyllaDB 的 C++ shard-per-core 架构(无 GC、工作负载感知调度)正好对症;API 兼容让"换引擎不换模型"成为可能,迁移风险集中在数据搬运。
JVM GC is a runtime-level source of tail latency — tuning can only mitigate it, never eliminate it. ScyllaDB's C++ shard-per-core architecture (no GC, workload-aware scheduling) addresses exactly that; API compatibility made "swap the engine, keep the model" possible, concentrating migration risk on data movement rather than model rewrites.
- 教训
p99 被 GC 和热点分区绑架时,换运行时架构(C++/shard-per-core)比调参有效一个数量级;超大迁移自研专用迁移器比通用方案快一个数量级;先保协议与模型兼容,再换引擎。
When p99 is held hostage by GC and hot partitions, changing the runtime architecture (C++/shard-per-core) beats tuning by an order of magnitude. For mega-migrations, a purpose-built migrator can be an order of magnitude faster than generic tooling. Secure protocol/model compatibility first, then swap the engine.
相关产品:Apache Cassandra / ScyllaDB 相关能力:ScyllaDB:shard-per-core 消灭 JVM GC 的运维税 最后核验:2026-10-01
Discord 抛弃 MongoDB:单副本的内存墙 失败教训
高并发写入
内存墙
分片策略
长尾延迟 p99
- 场景
In early 2015 Discord built its first version in two months, storing all messages in a single MongoDB replica set (a deliberate choice), but kept a data-access abstraction layer from day one. Around November 2015 stored messages reached about 100 million; once data plus indexes no longer fit in RAM, random-read latency became unpredictable from page faults; the team had decided at selection time not to adopt MongoDB sharding, which they considered complicated and not known for stability. By the January 2017 blog post, volume had grown from 100 million total messages to 120 million new messages per day. (
https://discord.com/blog/how-discord-stores-billions-of-messages)
- 决策
2017 年起把消息存储迁往 12 节点 Cassandra 集群。
From 2017, message storage moved to a 12-node Cassandra cluster.
- 结果
Cassandra 承接了消息量从 1 亿条总量到每天 1.2 亿条新增的增长(2017-01 博客口径);但批量删除数百万消息留下大量 tombstone,读放大引发 GC 停顿,gc_grace_seconds 从 10 天降到 2 天。
The Cassandra cluster carried message growth from 100 million total messages to 120 million new messages per day (per the January 2017 post). The migration also surfaced an LSM pitfall: bulk-deleting millions of messages left mountains of tombstones, and read amplification triggered GC pauses — later mitigated by dropping gc_grace_seconds from 10 days to 2.
- 机制根因
Discord 官方原文只说"数据加索引装不进内存后延迟不可预测",没有归因于某个存储引擎——不要把这口锅扣给 WiredTiger。这是 Discord 当时"单个 MongoDB 副本集"架构在该负载下的上限:读模式极度随机(50/50 读写比、大量冷数据随机点查),工作集一旦超出内存,缺页抖动直接打在 p99 上。结论必须限定:它证明的是单副本集方案在该负载下到顶,不证明"MongoDB 的硬上限就是内存",更不能推广为"文档库过内存就该迁 Cassandra"。Cassandra 侧:LSM-tree 追加写使磁盘 I/O 永远顺序,分区键 ((channel_id, bucket), message_id) 让单频道消息天然局部。但删除在 LSM 系不是免费的——墓碑迫使读路径合并所有 SSTable,读放大直接打在延迟上。
Discord's official post only says "the data and the index could no longer fit in RAM and latencies started to become unpredictable" — it blames no storage engine, so don't pin this on WiredTiger. This was the ceiling of Discord's then-architecture ("a single MongoDB replica set") under that workload: extremely random reads (about a 50/50 read/write ratio, lots of cold-data point lookups), where page-fault jitter hits p99 directly once the working set exceeds memory. The conclusion must stay scoped: it proves a single-replica-set design topped out under that load — not that "MongoDB's hard ceiling is memory", and certainly not a general rule that document stores should move to Cassandra past RAM. On the Cassandra side: LSM-tree append writes keep disk I/O always sequential, and the partition key ((channel_id, bucket), message_id) keeps a single channel's messages naturally local. But deletion is not free on LSM: tombstones force reads to merge all SSTables before a live row can be found, and read amplification lands directly on latency.
- 教训
为迭代速度选"简单方案"可以,但必须同时预留迁移抽象层(Discord 从第一天就这么干的);对当时的单副本集方案,数据+索引超出内存即撞上扩展上限——但这是单样本结论,不要外推为所有文档库的通用规律;删除在 LSM 系是写操作,按墓碑成本设计 TTL 与清理策略。
Picking the "simple option" for iteration speed is fine, but only with a migration abstraction layer in place from day one (Discord did exactly that). For their single-replica-set design, data-plus-indexes exceeding memory was the scaling ceiling — but that's a single-sample conclusion, don't extrapolate it into a universal law for all document stores. And deletion on LSM is a write operation: design TTLs and cleanup around tombstone cost.
相关产品:MongoDB、Apache Cassandra / ScyllaDB 相关能力:LSM 追加写:吃下"写多读少、只追加"的消息流、tombstone 读放大:"删除"是最贵的写操作 最后核验:2026-10-01
DoorDash:全集群调优 Cassandra,成本降 35%(2024) 成功经验
调优降本
压缩策略
布隆过滤器
读写隔离
- 决策
基础设施存储团队联合各产品团队开展数月的集中调优:按读写特征把压缩策略从 STCS 切换为 LCS,并联动把 bloom_filter_fp_chance 从 0.01 调到 0.1(否则切换后 OOM);把每日全表快照从"只读 DC 扛批量扫描"改为 token range 随机游走分散负载,最终整个只读 DC 下线;规避跨分区的批量读写反模式;针对性 JVM GC 调优(偏好 young gen、压制 old gen 回收,G1)。
The infrastructure storage team ran a months-long joint tuning effort with product teams: switched compaction strategy from STCS to LCS by read/write profile, and raised bloom_filter_fp_chance from 0.01 to 0.1 in the same move (otherwise OOM after the switch); replaced "read-only DC absorbing batch scans" with token-range random walks to spread the daily full-table snapshot load, eventually decommissioning the entire read-only DC; eliminated cross-partition batch read/write anti-patterns; and did targeted JVM GC tuning (favoring young-gen, suppressing old-gen collection, G1).
- 结果
全 Cassandra 集群成本下降约 35%;单位经济性从每 1 美元投入处理 23 KB/秒吞吐量提升到 59 KB/秒,提升约 154%。两数字均出自 DoorDash 官方工程博客正文与 Figure 1,由 DoorDash 存储团队统计的其全集群成本与吞吐量口径。
Total Cassandra fleet cost fell ~35%; unit economics rose from 23 KB/sec of throughput per dollar to 59 KB/sec — about a 154% improvement. Both figures come from the DoorDash engineering blog's text and Figure 1, measured by DoorDash's storage team across their full fleet.
- 机制根因
STCS 让大小渐增的 SSTable 长期共存,读放大高、需要预留大量闲置磁盘;LCS 分层后读路径更可预测。切换到 LCS 后 SSTable 更碎更分散,若不放大布隆过滤器误判率,读路径的布隆过滤器会吃光内存导致 OOM——"换压缩策略必须联动布隆过滤器"是根本原因。Cassandra 所有 DC 互联并参与复制,只读 DC 过载产生的延迟会以背压形式回流主 DC,所以"用 DC 做读写隔离"在 Cassandra 里是假隔离,这是下掉只读 DC 后主 DC 反而受益的原因。代价是数月的跨团队调优投入,以及 LCS 写放大高于 STCS、对写密集型表的取舍判断。
STCS lets ever-larger SSTables coexist, with high read amplification and large idle-disk reservations; LCS's leveled layout makes the read path more predictable. After switching to LCS, SSTables are smaller and more scattered — without raising the bloom-filter false-positive rate, the read path's bloom filters eat all memory and OOM — "changing compaction strategy must be paired with bloom-filter tuning" is the root cause. All Cassandra DCs are interconnected and participate in replication, so latency from an overloaded read-only DC flows back to the primary DC as backpressure — "using DCs for read-write isolation" is fake isolation in Cassandra, which is why the primary DC benefited after the read-only DC was removed. The cost: months of cross-team tuning effort, plus the judgment call that LCS's higher write amplification trades off against write-heavy tables.
- 教训
默认配置在规模下全是坑,压缩策略、布隆过滤器、一致性级别必须按 workload 成对调优,调优前先对齐访问模式做审计;先调透再扩容,35% 降本来自调优而非加机器,"规模问题=买更多节点"是贵且懒的答案;复制拓扑决定隔离边界,在 Cassandra 里用独立 DC 做读写隔离是架构幻觉,真正的隔离要回到查询本身。
Default configurations are all landmines at scale — compaction strategy, bloom filters, and consistency levels must be tuned as pairs against the workload, with an access-pattern audit first. Tune before scaling: the 35% savings came from tuning, not from adding machines — "scale problems = buy more nodes" is the expensive, lazy answer. Replication topology decides isolation boundaries: in Cassandra, a separate DC for read-write isolation is an architectural illusion; real isolation goes back to the queries themselves.
相关产品:Apache Cassandra / ScyllaDB 相关能力:可调一致性的运维税:gossip dance 与 compaction 死亡螺旋 最后核验:2026-10-01
浩瀚深度:三物理机受控对比后 ClickHouse 换 Doris,单表 13PB/534 万亿行稳定运行半年(2025) 成功经验
PoC
选型评估
OLAP 选型
受控对比测试
超大规模
- 场景
浩瀚深度(SHA: 688292,科创板上市公司)旗下顺水云(StreamCloud)大数据平台,需满足客户每日万亿级增量数据的写入与查询。团队对 MPP 数据库做过多轮选型测试,并在生产环境试过 Greenplum、ClickHouse 等多个方案。ClickHouse 体系的痛点:ZSTD 压缩因性能开销频繁报"too many parts"致入库积压、被迫退回 LZ4(存储成本高);节点数据不均衡、坏盘无法自动迁移,人工持续干预;并发查询多了性能明显下降;多表/大表 JOIN 能力不足。
Hohandeep (SHA: 688292, STAR Market listed) and its StreamCloud big-data platform needed to serve customers' daily trillion-row incremental writes and queries. The team ran multiple rounds of MPP database evaluations and had tried Greenplum, ClickHouse, and others in production. Pain points with the ClickHouse stack: ZSTD compression's overhead triggered frequent "too many parts" errors and ingestion backlogs, forcing a retreat to LZ4 (higher storage cost); skewed data across nodes with no automatic migration on disk failure, requiring constant manual intervention; visible query degradation under concurrent load; and inadequate multi-table/large-table JOIN capability.
- 决策
用三台物理机模拟生产环境数据和业务,对 Doris 与 ClickHouse 做受控对比(同数据量、同字段、不同排序键与索引类型,各 3 次冷查询):前缀索引 Doris 为 ClickHouse 2 倍以上;BloomFilter 2 倍;倒排索引 5 倍以上;全表扫描接近(ClickHouse 在 IS_IP_ADDRESS_IN_RANGE 函数上略胜)。据此评估迁移后查询响应提升超 2 倍。实施初期采用 ClickHouse 与 Doris 双跑并行验证,而非直接切换。
Three physical machines simulated production data and workloads for a controlled Doris-vs-ClickHouse comparison (same data volume, same columns, varying sort keys and index types, 3 cold queries each): prefix index — Doris 2x+ faster than ClickHouse; BloomFilter — 2x; inverted index — 5x+; full-table scans were close (ClickHouse slightly ahead on the IS_IP_ADDRESS_IN_RANGE function). Estimated post-migration query response improvement: over 2x. The rollout started with ClickHouse and Doris running in parallel for verification, not a flag-day cutover.
- 结果
选定 Apache Doris 并替换 ClickHouse(MySQL 语法使迁移便捷)。按浩瀚深度自述口径:117 节点集群、单表 13PB/534 万亿行、日均导入 145TB(峰值 158TB)、稳定运行半年以上;导入接口机 32 台→23 台(省超 28%);ZSTD 存储较 LZ4 降 6%;单副本压缩率约 4 倍(6.5PB,双副本+倒排+ZSTD);单 SQL 响应提升近 2 倍、批量查询提升近 30%。压测期的问题(写锁超时、Compaction 假死、事务积压)均已解决并分享至社区。
Apache Doris was selected to replace ClickHouse (MySQL syntax made migration easy). Per Hohandeep's own account: 117-node cluster, single table of 13PB / 534 trillion rows, ~145TB ingested daily (158TB peaks), stable for 6+ months; importer machines cut from 32 to 23 (28%+ saved); ZSTD storage 6% lower than LZ4; ~4x single-replica compression ratio (6.5PB total with dual replicas + inverted index + ZSTD); single-SQL response nearly 2x faster, batch queries ~30% faster. Load-test issues (write-lock timeouts, compaction stalls, transaction backlogs) were all resolved and shared back with the community.
- 机制根因
ClickHouse 的痛点集中在"存算一体 + 无自动均衡"的架构特性:坏盘/倾斜要人肉干预,并发一高就退化——这在百 TB 规模可忍,在 13PB/日增百 TB 规模是不可接受的运维税。Doris 的胜出点不是单项跑分,而是"索引丰富度 + 自动均衡 + MySQL 协议"的组合:倒排索引补上了日志场景的检索短板,tablet 自动均衡消掉了坏盘人工干预,MySQL 协议让迁移只是改配置和 SQL。值得注意的是 PoC 本身也不顺利(bucket 设 480 致 Compaction 假死、磁盘写满致事务积压)——"选型测试也测出了新引擎的坑",这恰恰是受控 PoC 的价值:坑暴露在压测期而非生产期。
ClickHouse's pain points centered on the architectural traits of "coupled storage-compute with no automatic rebalancing": disk failures and skew needed manual care, and concurrency degraded — tolerable at hundreds of TB, an unacceptable ops tax at 13PB with hundreds of TB ingested daily. Doris won not on a single benchmark but on the combination of "rich indexes + automatic rebalancing + MySQL protocol": inverted indexes filled the log-search gap, tablet auto-balancing eliminated manual disk-failure handling, and the MySQL protocol reduced migration to config and SQL changes. Notably, the PoC itself wasn't smooth (480 buckets causing compaction stalls, a full disk causing transaction backlogs) — "the evaluation also surfaced the new engine's pitfalls," which is precisely the value of a controlled PoC: pitfalls surface during load testing, not in production.
- 教训
超大规模选型的第一性指标是"坏盘/倾斜时要不要人肉干预",而不是查询快几倍;PoC 要包含故障与压测(浩瀚深度的 bucket 数、磁盘满盘教训都是压测逼出来的);双跑并行验证是替换核心引擎时的必要成本。诚实注记:所有倍数与规模数字均为浩瀚深度自述口径(阿里云开发者社区第一人称文章),未见第三方独立复现;"国内最大单表"为作者自述。
At hyper-scale, the first-order selection criterion is "does a disk failure or skew require manual intervention," not "how many times faster queries are"; PoCs must include failure and load testing (Hohandeep's bucket-count and full-disk lessons were forced out by load tests); parallel-run verification is a necessary cost when replacing a core engine. Honesty note: all multiples and scale figures are Hohandeep's self-reported claims (first-person article on Alibaba Cloud Developer Community) with no independent reproduction found; "largest known single table in China" is the author's own claim.
来源
Hohandeep StreamCloud team, "Hohandeep: From ClickHouse to Doris, Powering Hyper-Scale Analytics on a 13PB, 534-Trillion-Row Single Table" (2025-08, Alibaba Cloud Developer Community, first-person)
https://developer.aliyun.com/article/1678348
相关产品:Apache Doris、ClickHouse 相关能力:倒排索引与 tablet 自动均衡 最后核验:2026-10-02
快手:ClickHouse+Elasticsearch 双栈收敛到 Doris,广告分析延迟降 64%-90%(2026) 成功经验
广告分析
系统收敛
主键更新
ES 替代
- 场景
Kwai (developer of Kling AI, 400 million DAU) runs an ad platform where ad material data lived in MySQL + Elasticsearch while ad performance data (impressions/clicks/spend) lived in ClickHouse, forcing report queries to JOIN across two systems inside ClickHouse. Scale: hundreds of billions of ad material rows heading toward trillions, daily new rows up 3.5x year-over-year in Q1 2025 (300 million new rows/day), 700+ core fields, 4,000+ query templates. Pain: 35% slow-query rate on Elasticsearch with 1.4s average latency; a single ES shard could not hold datasets above 1 billion rows, and scaling out required costly data redistribution; no end-to-end observability between the two systems, so troubleshooting meant cross-system investigation. (
https://medium.com/@VeloDB_poweredby_ApacheDoris/from-clickhouse-elasticsearch-to-apache-doris-how-kwai-unified-trillion-scale-ad-analytics-31528f41513d)
- 决策
团队把 Doris、ClickHouse、Elasticsearch 拉到同一压测基准上对比写入吞吐、查询延迟、存储压缩和全文检索。ClickHouse 最先被淘汰——它不支持 Unique Key 更新,而广告素材表需要高频主键更新,这是硬约束;剩下 ES 和 Doris 全面对比,Doris 在写入吞吐、查询性能、存储效率、运维简单度上全部胜出。迁移分三阶段:试点验证(关键词推广场景双管线校验)→ 核心迁移(广告素材数据迁入、ES 集群下线)→ 全量收官,自建统一分析引擎 Bleem(Doris 核心 + Alluxio 缓存层 + OneSQL 统一网关)。
The team benchmarked Doris, ClickHouse, and Elasticsearch head-to-head on write throughput, query latency, storage compression, and full-text search. ClickHouse was eliminated first: it does not support unique key updates, while the ad material tables required high-frequency primary-key updates -- a hard constraint. Between ES and Doris, Doris won across the board on write throughput, query performance, storage efficiency, and operational simplicity. Migration ran in three phases: pilot validation (keyword promotion scenario with dual-pipeline verification), core migration (ad material data moved in, ES cluster decommissioned), and full completion, building a unified analytics engine called Bleem (Doris core + Alluxio cache layer + OneSQL unified gateway).
- 结果
以下均为 VeloDB 官方渠道口径(2026-03 文章),未经独立第三方复现:关键词推广页平均延迟降 64%、创意推广页延迟降 90%;写入吞吐 3 倍,单表实时写入峰值 300 万行/秒/节点;存储效率比 ES 高 60%(分区策略+ZSTD 压缩),支撑万亿行单表;慢查询率从 35% 降到 5% 以下;统一可观测性让平均排障时间降 80%。
All figures below are VeloDB's official-channel account (March 2026 article), not independently reproduced: average latency down 64% on the keyword promotion page and 90% on the creative promotion page; write throughput up 3x, with single-table realtime ingestion peaking at 3 million rows/sec per node; 60% better storage efficiency than Elasticsearch (partitioning strategy + ZSTD compression), supporting trillion-row single tables; slow-query rate from 35% down to below 5%; unified observability cut average troubleshooting time by 80%.
- 机制根因
决定性机制是"Unique Key 模型 + 倒排索引"的组合:广告素材表用 UNIQUE KEY(account_id, id) + 倒排索引,一张表同时承担 ES 的全文检索和高并发主键更新——这正是 ClickHouse(无行级 upsert 语义)和 ES(更新模型弱、单分片上限)的各自短板面。Doris 存算分离/耦合双模式支撑大表弹性扩展;万分区裁剪、数据倾斜处理、Stream Load 参数调优是落地时的工程关键。代价:迁移中 SeaTunnel 两阶段提交在 Spark 推测执行下出现数据重复,团队引入 ZooKeeper 分布式锁兜底——统一引擎不等于零运维,大规模写入一致性仍要自己解决。
The decisive mechanism is the combination of the Unique Key model and the inverted index: the ad material table uses UNIQUE KEY(account_id, id) plus an inverted index, so one table carries both Elasticsearch's full-text search and high-concurrency primary-key updates -- exactly the two weaknesses of ClickHouse (no row-level upsert semantics) and ES (weak update model, per-shard limits). Doris's coupled and decoupled storage-compute modes support elastic scaling of large tables; partition pruning at 10,000+ partitions, data-skew handling, and Stream Load parameter tuning were the engineering keys to landing it. The cost: during migration, SeaTunnel two-phase commits produced duplicate writes under Spark speculative execution, and the team added a ZooKeeper distributed lock as a guard -- a unified engine does not mean zero ops; large-scale write consistency still has to be solved yourself.
- 教训
先淘汰不满足硬约束的选项(主键更新),再比性能,避免在错误候选上浪费压测;统一引擎的收益不只是延迟数字,更是端到端可观测性带来的排障效率;大规模迁移三阶段走、双管线并行校验数据一致性,不要一次切完。
Eliminate candidates that fail hard constraints (primary-key updates) before benchmarking, so no stress-test effort is wasted on the wrong option; the payoff of a unified engine is not just latency numbers but troubleshooting efficiency from end-to-end observability; run large migrations in phases with dual-pipeline data consistency checks instead of cutting over in one shot.
来源
VeloDB official blog "From ClickHouse + Elasticsearch to Apache Doris: How Kwai Unified Trillion-Scale Ad Analytics" (2026-03-26, vendor ecosystem channel
—
相关产品:Apache Doris、ClickHouse 相关能力:一个 Doris 干掉 ES+ClickHouse+HBase、Unique Key 模型的实时更新 最后核验:2026-10-02
美团:300+ 集群、数十 PB 的 Doris 统一服务层与版本治理 成功经验
多引擎收敛
服务层统一
超大规模集群
版本治理
- 场景
美团数据平台长期多引擎并存(Hadoop、Kylin、Druid 等,各管一摊)。四个倒逼选型的挑战:分析报表要秒级返回、部分面向业务的场景要亚秒;核心业务表数千亿行;交易/运营/用户/流量/商户各业务线查询模式、并发、时效要求差异大;BI 系统合并后底层要能扛统一分析。Doris 被定位在"数据服务层":实时链路 Kafka→Flink→Doris 支撑实时大盘与在线分析,离线链路 Hive/HDFS→Spark 清洗建模→同步到 Doris 对外服务。两个定义需求的典型负载:到店餐饮 BD 效能分析(数百维度、几十个聚合指标、要求精确去重,近似计数不可接受)、商户经营报表(面向商户的高并发点查,查询挡在商户等 dashboard 刷新的路径上)。(
http://velodb.io/blog/how-meituan-consolidated-its-analytics-stack-on-apache-doris)
Meituan's data platform long ran multiple engines side by side (Hadoop, Kylin, Druid, each owning its workload). Four challenges forced the search: analytical reports must return in seconds, some business-facing scenarios need sub-second; core business tables passed hundreds of billions of rows; transactions, operations, users, traffic, and merchants each demand different query patterns, concurrency, and freshness; and BI consolidation meant whatever sat underneath had to carry unified analytics. Doris was placed in the data serving layer: the realtime path runs Kafka -> Flink -> Doris for live dashboards and online analysis, while the offline path lands in Hive/HDFS where Spark cleans and models before syncing results into Doris for serving. Two workloads defined the requirements: in-store dining BD productivity analysis (hundreds of dimensions, dozens of aggregate metrics, exact deduplication required -- approximate counts unacceptable) and merchant operating reports (high-concurrency point lookups sitting in the path of merchants waiting on dashboard refreshes). (
http://velodb.io/blog/how-meituan-consolidated-its-analytics-stack-on-apache-doris)
- 决策
选 Doris 做统一服务层收敛多引擎;随后把 Doris 没扛过的新负载不断搬上来,遇到问题自研解决并回馈社区:为外卖报表多表 JOIN 超时自研 Colocate Join;为流量/用户分析的精确去重做 Bitmap 系列优化;为"生产与查询抢资源"建 Spark on Doris 架构;为 300+ 集群的版本治理(0.1 到 2.1 的版本跨度)用跨集群复制 CCR 做不停服升级。
Doris was chosen as the unified serving layer to consolidate the engines; new workloads were then moved onto Doris one after another, with problems solved in-house and contributed back: Colocate Join for food-delivery report multi-table JOIN timeouts; a Bitmap optimization series for exact deduplication in traffic/user analysis; a Spark-on-Doris architecture to isolate data production from query serving; and cross-cluster replication (CCR) to upgrade 300+ clusters across versions 0.1 to 2.1 without downtime.
- 结果
以下为 VeloDB 文章口径(改编自美团技术专家在 Apache Doris 线上活动的分享),未经独立第三方复现:平台规模 300+ 集群、数十万 CPU 核、数十 PB 数据,最大单集群 10PB;Colocate Join 内测 join 查询平均 3 倍提升,外卖报表回到超时窗口内;Bitmap 优化后亿级基数下指标计算平均快 4-5 倍、单表查询 10 秒内返回;升级+治理完成后整体查询写入性能提升 20%-40%、极端场景 50%,全集群稳定性明显改善。
Figures below are the VeloDB article's account (adapted from a talk by a Meituan technical expert at an Apache Doris online event), not independently reproduced: fleet scale of 300+ clusters, hundreds of thousands of CPU cores, tens of PB of data, with the largest single cluster at 10PB; Colocate Join improved join query performance roughly 3x on average in internal testing, bringing food-delivery reports back inside their timeout windows; after Bitmap optimization, metric computation at hundred-million cardinalities ran 4x-5x faster on average with single-table queries returning within 10 seconds; after the upgrade and governance work, overall query and ingest performance improved 20%-40% (50% in extreme cases) with noticeably better fleet stability.
- 机制根因
Colocate Join 把参与 JOIN 的数据预先分布到同一节点,计算找数据而非数据找计算,消除 shuffle 开销;Bitmap 正交分桶控制数据分布基数,解决高基数精确去重的内存与速度问题;CCR 用 binlog 增量追赶把一次性迁移变成持续追赶,新老集群并行、DNS 一切换即完成 PB 级不停服升级。代价在治理侧:版本跨度带来语义差异、元数据膨胀(个别老集群副本数达千万级、影响 FE 恢复)、升级前必须先做元数据治理——规模上去后,难题从性能变成治理。
Colocate Join pre-distributes join-participating data onto the same nodes so computation moves to where the data already is, eliminating shuffle cost; Bitmap orthogonal bucketing controls data-distribution cardinality, solving the memory and speed problem of high-cardinality exact deduplication; CCR's binlog incremental catch-up turns a one-time migration into continuous sync, so old and new clusters run side by side and a DNS switch completes a PB-scale zero-downtime upgrade. The cost sits in governance: the version span brought semantic differences and metadata bloat (some legacy clusters had replica counts in the tens of millions, affecting FE recovery), and metadata cleanup had to precede the upgrade -- at that scale, the hard problems stop being about performance and become about governance.
- 教训
统一服务层之后真正的挑战是治理(版本、元数据、集群健康),不是查询性能;PB 级集群不要停机重载升级,用增量追赶+双跑业务语义校验;把自研能力(Colocate Join、Bitmap 优化、CCR 回补)回馈社区,能把一次性投入变成长期公共资产。
After unifying the serving layer, the real challenge is governance (versions, metadata, cluster health), not query performance; do not reload PB-scale clusters offline for upgrades -- use incremental catch-up plus dual-run business-semantics validation; contributing in-house work (Colocate Join, Bitmap optimizations, CCR backports) back to the community turns one-time investment into a lasting shared asset.
来源
VeloDB official blog "How Meituan consolidated its analytics stack on Apache Doris" (adapted from a talk by a Meituan technical expert at an Apache Doris online event
—
相关产品:Apache Doris 相关能力:一个 Doris 干掉 ES+ClickHouse+HBase 最后核验:2026-10-02
网易游戏:六引擎收敛到 Doris 统一湖仓,每天 1500 万查询(2026) 成功经验
六引擎收敛
湖仓一体
Bitmap 去重
游戏行业
- 场景
NetEase Games' data platform handles 15 million queries per day across 200+ internal projects on petabyte-scale storage. The original stack ran six specialized systems: Hive + Spark for batch processing, Trino for ad-hoc queries, Elasticsearch + HBase for realtime lookups, and ClickHouse for analytics. Pain: data passed through multiple systems before reaching analysts, so realtime analytics could not keep up; Hive/Spark/Trino interactive query performance was insufficient, while HBase/ES/ClickHouse had limited support for complex JOINs; six systems meant six sets of dedicated ops expertise; every new requirement meant separate Spark/Trino/HBase jobs plus Elasticsearch DSL, forcing analysts to learn multiple tools. (
https://medium.com/@VeloDB_poweredby_ApacheDoris/netease-games-from-elasticsearch-hbase-and-clickhouse-to-a-unified-apache-doris-lakehouse-686362fa1bc1)
- 决策
两阶段收敛。阶段一:Doris 替换 ES、HBase、ClickHouse 做实时层(Doris 当时已是网易内部使用最广的 OLAP 引擎之一)。阶段二(从 Doris 2.1 起):再替换 Hive、Spark、Trino,建成以 Doris 为核心的统一湖仓,自研 SmartSQL 做内部查询路由(支持"Doris 查湖"和"Doris 统一湖仓"两种模式)。
Two-phase consolidation. Phase 1: replace Elasticsearch, HBase, and ClickHouse with Doris as the realtime layer (Doris was already one of the most widely used OLAP engines inside NetEase). Phase 2 (starting on Doris 2.1): replace Hive, Spark, and Trino as well, building a unified lakehouse with Doris at its core, plus an in-house SmartSQL query router supporting two patterns -- "Doris as query engine on the lake" and "Doris as unified lakehouse".
- 结果
以下均为 VeloDB 官方渠道口径(2026-08 文章),未经独立第三方复现:现运行 20+ Doris 集群、数百节点,服务 200+ 项目、每天 1500 万查询、PB 级存储。宽表分析场景:Doris 与 ClickHouse 查询性能相当,但显著更易运维(降低运维人员技术门槛)。用户行为分析:14 亿记录数据集上做 Bitmap 优化后,峰值内存从 54GB 降到 4.2GB、查询时间从 20 秒降到 2 秒以内。联邦查询场景:Doris 跨源查询性能是 Presto 的 2-3 倍。
All figures below are VeloDB's official-channel account (August 2026 article), not independently reproduced: 20+ Doris clusters across hundreds of nodes now serve 200+ projects with 15 million daily queries on petabyte-scale storage. Wide-table analytics: Doris query performance is comparable to ClickHouse but significantly easier to operate (lower technical bar for ops staff). User behavior analysis: after Bitmap optimization on a 1.4-billion-record dataset, peak memory dropped from 54GB to 4.2GB and query time from 20 seconds to under 2 seconds. Federated queries: Doris runs 2-3x faster than Presto.
- 机制根因
Doris 的倒排索引+主键点查覆盖 ES/HBase 的实时检索需求,列存 MPP 覆盖 CK 的分析需求——"一个引擎同时扛检索和 OLAP"是收敛成立的前提;Bitmap 索引与物化视图加速用户行为分析(14 亿记录精确去重);Workload Group 资源隔离替代 Trino 的一长串资源限制参数,大查询不再拖垮集群。诚实备注:超大 ETL 作业仍回退到原 Hive/Spark/Trino 引擎池——不是所有负载都硬搬,这是务实做法。
Doris's inverted index plus primary-key point lookups cover the realtime retrieval needs of ES/HBase, while its columnar MPP covers ClickHouse's analytics -- "one engine carrying both search and OLAP" is what makes the consolidation viable; Bitmap indexes and materialized views accelerate user behavior analysis (exact deduplication over 1.4 billion records); Workload Group resource isolation replaces Trino's long list of resource-limit parameters so large queries no longer drag the cluster down. Honest note: extremely large ETL jobs still fall back to the original Hive/Spark/Trino engine pool -- not every workload was force-migrated, which is the pragmatic choice.
- 教训
收敛分阶段走,先实时层后批处理层,不要一次掀桌子;SQL 方言兼容(生产 99%+)和 Hive UDF 兼容是迁移成本的决定项;保留回退路径(超大 ETL 留原引擎)比"全部迁移"的口号更可持续。
Consolidate in phases -- realtime layer first, batch layer second -- instead of ripping everything out at once; SQL dialect compatibility (99%+ in production) and Hive UDF compatibility decide migration cost; keeping a fallback path (large ETL stays on the old engines) is more sustainable than a "migrate everything" slogan.
来源
VeloDB official blog "NetEase Games: From Elasticsearch, HBase, and ClickHouse to a Unified Apache Doris Lakehouse" (2026-08-26, vendor ecosystem channel
—
相关产品:Apache Doris、ClickHouse 相关能力:一个 Doris 干掉 ES+ClickHouse+HBase 最后核验:2026-10-02
腾讯音乐数据平台:ClickHouse 迁 Doris,部分列更新省存储 42%(2023) 成功经验
ClickHouse 迁移
部分列更新
语义层
工程师一手
- 场景
A first-hand account by Tencent Music data platform engineers Jun Zhang and Kai Dai: data assets over a music library serving 800 million MAU (songs, lyrics, melodies, albums, artists), with the offline warehouse TDW producing 800+ tags, 1,300+ metrics, and 80+ source tables. The original architecture imported flat tables into ClickHouse for analysis and used Elasticsearch for search and audience targeting. Pain: ClickHouse did not support partial column updates -- latency from any one data source delayed the whole flat-table pipeline and hurt timeliness; pouring everything into flat tables wasted storage; ClickHouse's storage-compute coupling and heavily interdependent components raised cluster instability risk; federated queries across ClickHouse and Elasticsearch came with tedious connection issues. (
https://github.com/apache/doris-website/blob/HEAD/blog/Tencent-Data-Engineers-Why-We-Went-from-ClickHouse-to-Apache-Doris.md)
- 决策
Doris 的 Aggregate 模型支持实时部分列更新:Spark→Kafka→Flink 预聚合→Doris/ES;大宽表按更新频率拆成小表、用 Doris 多表 JOIN 与 ES 外表联邦查询替代"一个大宽表灌到底";引入语义层统一标签指标定义;列名用 ID 映射(如 song_name 存为 a4)规避标签频繁上下线带来的加列/删列开销。
Doris's Aggregate model supports realtime partial column updates: Spark -> Kafka -> Flink pre-aggregation -> Doris/ES; large flat tables were split by update frequency into smaller tables, replaced by Doris multi-table JOINs and ES external-table federated queries instead of "one giant flat table"; a semantic layer unified tag/metric definitions; column names were mapped to IDs (e.g., song_name stored as a4) to avoid the add/drop column cost of tags going online and offline frequently.
- 结果
以下均为两位工程师自述口径(2023-03-07 文章),未经独立第三方复现:存储成本降 42%、开发成本降 40%;每日离线导入时间降 75%、CUMU compaction 评分从 600+ 降到 100;新标签上线 10 分钟后可查;人群圈选场景要求秒级响应。
All figures below are the two engineers' own account (March 7, 2023 article), not independently reproduced: storage costs down 42%, development costs down 40%; daily offline ingestion time down 75%, CUMU compaction score from 600+ down to 100; newly added tags queryable 10 minutes after onboarding; audience-targeting scenarios required second-level response.
- 机制根因
Aggregate 模型的部分列更新是 ClickHouse 机制性缺失的能力——各源表 ETL 节奏不一、只涉及部分标签指标时,不必等全量宽表;Doris 联邦查询(ES 外表自动映射 schema)减少数据搬运;FE/BE 双进程、无外部依赖的架构降低运维复杂度。诚实备注:Doris 1.1.3 不支持修改列名,团队用 MySQL 元数据表做"名→ID"映射绕行,直到 1.2.0 的 Light Schema Change 才解决——版本能力边界要写进方案。
Partial column update in the Aggregate model is a capability ClickHouse structurally lacks -- when source tables run ETL at different paces touching only some tags and metrics, there is no need to wait for the full flat table; Doris federated queries (ES external tables with automatic schema mapping) reduce data movement; the two-process FE/BE architecture with no external dependencies lowers ops complexity. Honest note: Doris 1.1.3 did not support renaming columns, so the team worked around it with a MySQL metadata table mapping names to IDs, fixed only by Light Schema Change in 1.2.0 -- version capability boundaries belong in the plan.
- 教训
把"数据会变"(部分列更新)列为硬约束再选型,而不是事后补救;宽表按更新频率拆分比一股脑灌宽表更省存储、吞吐更高;追踪版本能力边界(列名修改、schema change 代价),避免方案建立在"下个版本会修"的假设上。
Make "data changes" (partial column updates) a hard constraint before selecting, not something patched afterward; splitting flat tables by update frequency saves more storage and raises throughput than pouring everything into one wide table; track version capability boundaries (column renames, schema-change costs) instead of building plans on "the next release will fix it".
来源
Apache Doris official blog "Tencent data engineer: why we went from ClickHouse to Apache Doris?" (2023-03-07, by Tencent Music data platform engineers Jun Zhang & Kai Dai, first-hand production account
—
相关产品:Apache Doris、ClickHouse 相关能力:Unique Key 模型的实时更新、MySQL 协议零学习成本接入 最后核验:2026-10-02
腾讯音乐:Elasticsearch 迁 Doris 做统一检索引擎,成本降 80%(2025) 成功经验
ES 替代
统一检索
倒排索引
成本优化
- 场景
Tencent Music Entertainment (TME, 800 million MAU) content library platform: full-text search (finding artists/songs by flexible conditions) + tag-based audience segmentation over billions of records (sub-second) + aggregation analysis. The original hybrid architecture used Elasticsearch for full-text search and segmentation and Doris for OLAP analytics, with Doris's ES Catalog providing a unified query interface. Pain: enormous ES storage overhead; full data writes exceeding 10 hours as volumes grew, nearing the business's operational limits; a multi-component architecture with complex maintenance, redundant storage, and data inconsistency risk. (
https://github.com/apache/doris-website/blob/HEAD/blog/tencent-music-migrate-elasticsearch-to-doris.md)
- 决策
Doris 2.0 引入倒排索引后,TME 把全文检索、标签圈选、聚合分析全部交给 Doris:维度表用 Unique Key 模型+部分列更新(标签频繁变更),事实表用 Aggregate 模型按天分区;迁移经自研 SuperSonic(Headless BI)把 ES DSL 翻译成 SQL、切换预定义指标的数据源,业务无感。
After Doris 2.0 introduced the inverted index, TME handed full-text search, tag-based segmentation, and aggregation analysis entirely to Doris: dimension tables use the Unique Key model with partial column updates (tags change frequently), fact tables use the Aggregate model partitioned by day; the migration went through their in-house SuperSonic (Headless BI), which translates ES DSL into SQL and switches pre-defined metrics' data sources, invisible to the business.
- 结果
以下均为 VeloDB 官方渠道口径(2025-04-17 文章),未经独立第三方复现:整体运维成本降 80%(单业务日全量数据 ES 697.7GB → Doris 195.4GB);写入 4 倍(全量导入从 10 小时以上降到 3 小时内);告警频率从每天 20+ 次降到每月个位数;复杂标签查询从分钟级降到秒级。
All figures below are VeloDB's official-channel account (April 17, 2025 article), not independently reproduced: overall operational costs down 80% (one business's daily full data: 697.7GB in Elasticsearch vs 195.4GB in Doris); 4x write performance (full ingestion from 10+ hours down to under 3 hours); alert frequency from 20+ per day down to single digits per month; complex tag queries from minutes down to seconds.
- 机制根因
Doris 内核级倒排索引(MATCH_ANY/MATCH_ALL/MATCH_PHRASE、中英文分词、任意 AND/OR/NOT 组合)是替代 ES 的决定性能力——2.0 之前 TME 只能维持混合架构,版本与能力发布时间线对得上是选型关键;Unique Key 模型支撑维度表高频部分列更新;Resource Group 物理隔离(核心/普通)+ Workload Group 逻辑隔离保障多业务互不干扰。代价:迁移依赖自研的 SuperSonic 语义层做 DSL 翻译,没有这个中间层,ES 查询习惯无法平滑搬迁。
Doris's kernel-level inverted index (MATCH_ANY / MATCH_ALL / MATCH_PHRASE, English and Chinese tokenization, arbitrary AND/OR/NOT combinations) is the decisive capability for replacing ES -- before 2.0, TME could only sustain the hybrid architecture, so aligning version choice with the capability's release timeline was the key selection call; the Unique Key model supports high-frequency partial column updates on dimension tables; Resource Group physical isolation (core vs normal) plus Workload Group logical isolation keep multiple businesses from interfering with each other. The cost: the migration depended on the in-house SuperSonic semantic layer for DSL translation -- without that middle layer, ES query habits could not move over smoothly.
- 教训
ES 替代的门槛是倒排索引成熟度,版本选型必须对齐能力发布时间线;用语义层(Headless BI)解耦数据源,迁移对业务透明;成本账要算"存储+写入+运维"三项,单看查询性能会低估统一架构的收益。
The bar for ES replacement is inverted-index maturity -- version choice must align with the capability's release timeline; decouple data sources with a semantic layer (Headless BI) so migration is transparent to the business; count cost as storage + writes + operations -- judging by query performance alone understates the unified architecture's payoff.
来源
Apache Doris official blog "How Tencent Music saved 80% in costs by migrating from Elasticsearch to Apache Doris" (2025-04-17, by VeloDB Engineering Team, vendor ecosystem channel
—
相关产品:Apache Doris 相关能力:一个 Doris 干掉 ES+ClickHouse+HBase 最后核验:2026-10-02
小米:Doris+Paimon 统一湖仓,查询延迟 60 秒到 10 秒(2026) 成功经验
湖仓一体
Paimon
物化视图
高并发
- 场景
Xiaomi's OLAP platform originally ran multiple engines (Presto, Druid, Doris) on multiple storage formats (Iceberg, Paimon, Doris, Druid), with minute/hourly/daily data scattered across systems -- redundant and inconsistent -- and each engine maintaining its own data modeling and access control, driving governance costs up. The team first built a Doris (query engine) + Paimon (open table format) lakehouse, then upgraded Doris from "querying the lake only" to a realtime data warehouse -- hot data lands in Doris internal tables, using materialized views, JOIN capability, and write-back for warehouse-grade performance. (
https://www.velodb.io/blog/unified-lakehouse-apache-doris-apache-paimon-xiaomi)
- 决策
统一计算引擎为 Doris+Spark(Doris 管实时交互分析、Spark 管离线批处理)、统一存储为 Paimon;冷热分层:冷/历史数据放 Paimon(低成本、开放格式),热数据与高频聚合进 Doris 内表。针对性优化三处:Paimon 聚合下推到 Doris C++ 引擎(绕过单线程 Java SDK 的 merge 瓶颈)、快照级增量物化视图、HDFS 读超时 60 秒→100 毫秒+数据缓存。
Unified compute on Doris + Spark (Doris for realtime interactive analytics, Spark for offline batch) and unified storage on Paimon; hot-cold tiering with cold/historical data in Paimon (low cost, open format) and hot data plus high-frequency aggregations in Doris internal tables. Three targeted optimizations: push Paimon aggregation down into the Doris C++ engine (bypassing the single-threaded Java SDK merge bottleneck), snapshot-level incremental materialized views, and HDFS read timeout cut from 60 seconds to 100 milliseconds plus data caching.
- 结果
以下均为 VeloDB 官方渠道口径(2026-03 文章),未经独立第三方复现:平均查询延迟 60 秒→10 秒(6 倍);聚合查询 40 秒→8 秒(5 倍);高并发场景查询延迟降到原来的 25%-75%,并发从 5 提到 80,并发吞吐是 Presto 的 5 倍;HDFS 长尾 P99 翻倍、整体性能提升 10%。上述优化已全部回馈 Apache Doris 社区。
All figures below are VeloDB's official-channel account (March 2026 article), not independently reproduced: average query latency from 60s down to 10s (6x); aggregation queries from 40s down to 8s (5x); in high-concurrency scenarios query latency reduced to 25%-75% of the original, concurrency scaled from 5 to 80 queries, and concurrency throughput 5x that of Presto; HDFS long-tail P99 doubled with 10% overall improvement. All of these optimizations were contributed back to the Apache Doris community.
- 机制根因
Doris 原生 Parquet Reader 直读 Paimon 数据文件 + 分布式 Hash 聚合,替代 Paimon Java SDK 单线程排序合并,这是 5 倍聚合加速的来源;快照级增量 MV 只读指定 snapshot 范围、避免全量重算;HDFS 快速失败重试(100ms 超时)+ 本地缓存解决长尾抖动。代价/边界:Doris 做纯查询引擎时,Paimon Merge-on-Read 表读取是已知短板——湖仓一体不是"只查湖",热数据必须进内表或缓存才能拿到数仓级性能。
Doris's native Parquet reader reads Paimon data files directly plus distributed hash aggregation, replacing the Paimon Java SDK's single-threaded sort-merge -- the source of the 5x aggregation speedup; snapshot-level incremental MVs read only the specified snapshot range, avoiding full recomputation; fast-fail HDFS retries (100ms timeout) plus local caching tame long-tail jitter. The cost/boundary: when Doris acts purely as a query engine, reading Paimon Merge-on-Read tables is a known weakness -- the right lakehouse posture is "cold data on the lake, hot data in the warehouse", not querying the lake directly for everything.
- 教训
湖仓一体的正确姿势是"冷数据放湖、热数据进仓",而不是所有查询都直查湖;长尾延迟靠"超时+重试+缓存"三件套系统性解决;把优化回馈社区,小米的快照级增量 MV 已成为 Doris 公共能力。
The correct lakehouse posture is "cold data on the lake, hot data in the warehouse", not pointing every query at the lake; fix long-tail latency systematically with the timeout + retry + cache trio; contributing optimizations back to the community turns Xiaomi's snapshot-level incremental MV work into a shared Doris capability.
来源
VeloDB official blog "Unified Lakehouse with Apache Doris + Paimon: Xiaomi Achieves 6x Faster Performance" (2026-03, vendor ecosystem channel
—
相关产品:Apache Doris 相关能力:一个 Doris 干掉 ES+ClickHouse+HBase 最后核验:2026-10-02
Dropbox 的 MySQL 元数据帝国(2013–2016):Edgestore 与 Magic Pocket 成功经验
元数据与Blob分离
强一致
EB级存储
无聊技术
- 场景
Dropbox's core hard problem is metadata: who owns which file, with what permissions, at which version — demanding strong consistency, high availability, and global distribution. Heavily sharded MySQL carried all metadata in the early days; as the business grew, the marginal cost of "bolting more features onto MySQL" kept rising. (
https://dropbox.tech/infrastructure/reintroducing-edgestore)
- 决策
Two tracks in parallel. Metadata track: built Edgestore in-house — a strongly consistent, read-optimized, horizontally scalable, geo-distributed metadata store; in the official blog's words, "rather than continue bolting more features and tooling to our MySQL infrastructure, we built Edgestore to abstract away the database altogether." Dozens of internal and external products share one Edgestore deployment. Blob track: built the Magic Pocket immutable block store, keeping the block index layer on sharded MySQL (hash → cell/bucket/checksum). (
https://dropbox.tech/infrastructure/inside-the-magic-pocket)
- 结果
Magic Pocket 达到 EB 级规模(据 InfoQ 报道);Edgestore 成为数十个产品的统一元数据底座,一致性缓存、事件流、跨数据中心复制等硬骨头只啃一次。
Magic Pocket reached exabyte scale (per InfoQ's reporting); Edgestore became the unified metadata substrate for dozens of products — the hard problems (consistent caches, event streams, cross-datacenter replication) solved once and reused everywhere.
- 机制根因
元数据与 blob 分离是关键一刀——blob 用不可变块 + SHA-256 内容寻址,天然去重、消灭一致性复杂度;索引层(block index)故意用"无聊"的分片 MySQL,因为团队对它的运维能力最有把握,"足够好且运维熟"胜过"理论上更优";Edgestore 则把真正难的东西(多租户隔离、一致性缓存、事件流、跨 DC 复制)做成平台能力,一次投入、处处复用。
Separating metadata from blobs is the decisive cut — blobs use immutable blocks plus SHA-256 content addressing, which gives deduplication for free and eliminates consistency complexity. The index layer (block index) deliberately uses "boring" sharded MySQL, because that is the system the team knows how to operate best — "good enough and operationally familiar" beats "theoretically superior." Edgestore then turns the genuinely hard parts (multi-tenancy isolation, consistent caches, event streams, cross-DC replication) into platform capabilities: invested once, reused everywhere.
- 教训
创新要放在刀刃上:索引/元数据层用你最熟的无聊技术,把自研火力集中在存储引擎和平台抽象层;当"在现有库上打补丁"的成本超过"抽象掉它"时,就是自研平台的时机——但前提是你已经用旧方案撑到了那个规模。
Put innovation where the blade is: use the boring technology you know best for the index/metadata layer, and concentrate in-house firepower on the storage engine and platform abstraction. When the cost of "patching the old database" exceeds the cost of "abstracting it away," that is the moment for a self-built platform — but only after the old stack has carried you to that scale.
相关产品:MySQL 相关能力:简单 OLTP 下"最不折腾周末"的运维体感 最后核验:2026-10-01
FinQore:财务报告管线从 8 小时压到 8 分钟,Postgres 数据集逐步被替换(2025) 成功经验
财务报表自动化
ETL 管线
Postgres 替换
AI Agent 实时 RAG
60 倍加速
- 场景
FinQore 是面向 CFO 及其财务团队的收入智能(revenue intelligence)平台,要从客户的账单系统、ERP、CRM、产品系统等多源抽取财务、客户、产品数据,拼成按客户业务逻辑分段的"收入立方体"(revenue cube),实现月度/季度报告周期的自动化。原来跑在 Postgres 上的数据管线一次要跑 8 小时,财务团队想即时分析、探索、落地数据时只能干等。(
https://motherduck.com/case-studies/finqore/)
FinQore is a revenue intelligence platform serving CFOs and finance teams. It pulls financial, customer, and product data from billing platforms, ERPs, CRMs, and product systems, stitching them into a "revenue cube" segmented by each customer's business logic to automate monthly and quarterly reporting cycles. Its data pipeline on Postgres took 8 hours per run, leaving finance teams waiting whenever they wanted to analyze, explore, or operationalize data instantly. (
https://motherduck.com/case-studies/finqore/)
- 决策
FinQore 把整条管线押注在 DuckDB + MotherDuck 上:用 MotherDuck 处理和统一多源数据,跑一条代码生成的专有管线,每天产出一份全公司唯一的财务可信数据源;前端全部用 MotherDuck,并"系统性地把原来从 Postgres 拉的数据集替换掉"。选型逻辑是开发体验一致:本地 DuckDB 开发调试,云端 MotherDuck 跑生产,同一引擎。
FinQore bet the whole pipeline on DuckDB + MotherDuck: MotherDuck processes and unifies multi-source data through a bespoke code-generated pipeline that produces a single, daily-updated, reliable source of financial data for the entire organization. "We use MotherDuck for everything on the front end," the team says, and is "systematically replacing some of the datasets we're pulling from Postgres." The selection logic was developer-experience consistency: develop and debug on local DuckDB, run production on MotherDuck in the cloud — the same engine everywhere.
- 结果
联合创始人兼 CTO Jim O'Neill 原话:"Our data pipelines used to take eight hours. Now they're taking eight minutes"(管线从 8 小时降到 8 分钟,约 60 倍),并称"and I see a world where they take eight seconds"。在此基础上 FinQore 推出了指标浏览器和 AI 财务分析师 Qori——基于 MotherDuck + DuckDB 做实时 RAG(检索增强生成),让 agent 的回答引用真实、最新的财务数字而非模型幻觉。"8 小时→8 分钟"为 MotherDuck 官方案例口径,管线覆盖的数据量、并发规模未披露,引用时须注明口径。
Co-founder and CTO Jim O'Neill's own words: "Our data pipelines used to take eight hours. Now they're taking eight minutes" (about a 60x improvement), adding "and I see a world where they take eight seconds." On top of that, FinQore launched a metrics explorer and Qori, an AI financial analyst that does real-time RAG (retrieval-augmented generation) over MotherDuck + DuckDB, so the agent's answers cite real, current financial numbers instead of hallucinations. The "8 hours to 8 minutes" figure is MotherDuck's official case-study claim; pipeline data volumes and concurrency levels were not disclosed, so quote it with the source caveat.
- 机制根因
财务报告是典型的"多源抽取→统一变换→高频聚合"负载,瓶颈在引擎模型而非硬件:Postgres 是行式 OLTP 引擎,大范围扫描聚合时逐行开销大;DuckDB 列式向量化执行一次处理一批列,在这类大范围扫描聚合的场景下比行式逐行快约一个数量级(该场景的实测对比,非通用加速比)。进程内模型还省掉了"ETL 进数仓"的数据搬运链路。MotherDuck 则解决了纯 DuckDB 的短板:单机文件无法多用户协作、云端无 serverless 弹性——把同一个 DuckDB 引擎搬到云上,管线才能从笔记本走进生产调度。
Financial reporting is a classic "multi-source extract, unify, transform, aggregate heavily" workload, and the bottleneck is the engine model, not the hardware. Postgres is a row-oriented OLTP engine with heavy per-row overhead on wide scans and aggregations; DuckDB's columnar vectorized execution processes batches of columns at a time, roughly an order of magnitude faster than row-by-row scanning on this scenario's wide-scan aggregations (a scenario-specific comparison, not a universal speedup ratio). The in-process model also eliminates the data-movement chain of "ETL into a warehouse." MotherDuck fixes pure DuckDB's weaknesses — a single-machine file can't serve multiple collaborators and has no serverless elasticity in the cloud — by moving the same DuckDB engine to the cloud so the pipeline can graduate from notebooks into production scheduling.
- 教训
OLTP 库兼职 OLAP 的天花板往往比预期来得早,8 小时管线里的大头通常是"用错引擎"而非"数据太大";进程内 OLAP + 云端同引擎的组合,可以替代传统的重 ETL 链路,开发、测试、生产用同一套 SQL 与语义;AI agent 要可信地回答业务问题,底座必须是"快到能实时查"的分析引擎,RAG 的延迟预算直接决定了架构选型。
The ceiling of running OLAP on an OLTP database arrives earlier than expected — most of an 8-hour pipeline is usually "wrong engine," not "too much data." In-process OLAP plus the same engine in the cloud can replace a heavyweight ETL chain, with one SQL dialect and one semantic layer across dev, test, and production. For an AI agent to answer business questions credibly, the foundation must be an analytical engine fast enough for real-time lookups — RAG's latency budget directly dictates the architecture choice.
相关产品:DuckDB、PostgreSQL(社区版) 相关能力:进程内 OLAP:pip install 即拥有的分析引擎 最后核验:2026-10-02
GoodShip:Postgres 频繁超时 → 亚秒级 AI 货运分析师,每租户独立 Duckling(2026) 成功经验
AI Agent
物流分析
多租户隔离
Postgres 超时
serverless 计费
- 场景
GoodShip is an all-in-one platform for freight orchestration, procurement, and pricing, with its analytics backend running on Postgres. As the customer base grew, complex queries — joining late-delivery records to financials, surfacing carrier penalties, aggregating years of operational data — slowed down until they timed out, with large historical queries returning no results at all. The larger datasets held about 4 million records spanning three to four years — not enormous by data warehouse standards, but enough to expose Postgres's limits on analytical workloads and degrade the end-user experience. (
https://motherduck.com/case-studies/goodship-ai-transportation-analyst/)
- 决策
评估备选:ClickHouse 是有力竞争者,但查询方言偏离 Postgres SQL,长期维护负担重;Starburst、Apache Pinot 要自己管集群,"moving parts 太多"。最终选 MotherDuck:SQL 方言友好、serverless 基建、计费模型贴合"早上集中看数"的间歇型负载。团队基于 MotherDuck 的 remote MCP 不到一周搭出 Laney v1——业界第一个 AI 货运分析师,2026 年 1 月上线生产,随后把整个报表后端都切到 MotherDuck。
The team evaluated alternatives: ClickHouse was a contender, but its query dialect diverged from Postgres SQL in ways that felt unsustainable to maintain; Starburst and Apache Pinot required managing clusters — too many moving parts. They chose MotherDuck: a friendly SQL dialect, serverless infrastructure, and a pricing model matching their bursty "review data in the morning" usage pattern. The team built v1 of Laney — the industry's first AI Transportation Analyst — on MotherDuck's remote MCP in under a week; it went to production in January 2026, and the team has since flipped their entire reporting backend to MotherDuck.
- 结果
Postgres 上 60 秒超时无结果的查询,经 MotherDuck 亚秒返回;hypertenancy 架构下每客户一个 service account、一份独立数据库切片,agent 生成的 SQL 在结构上无法访问其他客户的数据(对比 Postgres 里"只是一个 tenant_id 字段"的软隔离);serverless 按用计费,"早上看数、然后去解决问题"的间歇负载不再为空转集群付费。VP Engineering Eric Dillon 称客户反响"rapturous"(热烈)。以上均为 MotherDuck 官方案例口径。
Queries Postgres could not return in sixty seconds before timeout now return sub-second through MotherDuck. With the hypertenancy architecture, each customer gets one service account and one isolated database slice, so agent-generated SQL structurally cannot access another customer's data — versus Postgres, where isolation was "just a tenant ID column on the tables." Serverless billing means the bursty workload ("review data in the morning, then go solve problems") pays for compute when used, not for an idle cluster. VP of Engineering Eric Dillon called the customer reaction "rapturous." All figures are MotherDuck's official case-study claims.
- 机制根因
Agent 式分析是串行推理链:查一次、看结果、决定下一步,每步 30 秒延迟产出的不是"慢仪表盘"而是"不可用的 agent"——延迟在这里是可用性问题。DuckDB 列式引擎解决扫描聚合速度;hypertenancy 用"每租户独立 DuckDB 进程(Duckling)"的物理隔离替代 WHERE 子句的逻辑隔离,这对"SQL 由 AI 生成"的场景至关重要:隔离必须是结构上不可能越界,而非约定不越界。实时性需求(如标记争议、静音告警)则用 DuckDB 的 Postgres 扩展把 Postgres 实时数据与 MotherDuck 快照 union 起来。
Agentic analytics is a sequential reasoning chain — query once, read the result, decide the next step — so 30-second latency per step produces not a slow dashboard but an unusable agent: latency here is an availability problem. The DuckDB columnar engine fixes scan-and-aggregate speed; hypertenancy replaces WHERE-clause logical isolation with physical isolation ("one DuckDB process per tenant," a Duckling), which is critical when SQL is AI-generated — isolation must be structurally impossible to cross, not merely agreed upon. Real-time needs (flagging a dispute, muting an alert) are handled by unioning live Postgres data with the MotherDuck snapshot via DuckDB's Postgres extension.
- 教训
给 AI agent 选数据底座时,先算推理链的延迟预算,再谈功能;多租户 + agent 生成 SQL 的组合下,租户隔离要做到物理/结构级别,逻辑隔离挡不住不可预测的生成 SQL;间歇型负载别为闲时集群付费,serverless 计费与" bursts"型使用模式天然契合;迁移可以分两步走——先解决最痛的 agent 查询,再翻转整个报表后端。
When choosing a data foundation for an AI agent, budget the reasoning chain's latency first and talk features second. The multi-tenant plus agent-generated-SQL combination demands physical/structural tenant isolation — logical isolation cannot contain unpredictable generated SQL. Don't pay for idle clusters under bursty workloads; serverless billing naturally fits burst-shaped usage. Migration can go in two steps: fix the most painful agent queries first, then flip the entire reporting backend.
相关产品:DuckDB、PostgreSQL(社区版)、ClickHouse 相关能力:单写者 + 无 Server:并发与多租户的天花板 最后核验:2026-10-02
Hugging Face:38 万数据集的 SQL 控制台由 DuckDB WASM 驱动,hf:// 直查免下载(2024–2026) 成功经验
AI 数据集
浏览器内分析
WASM
hf:// 协议
零下载查询
- 决策
DuckDB 与 Hugging Face 合作(2024 年 5 月随 DuckDB v0.10.3 发布),在 httpfs 扩展上增加 hf:// 协议,把数据集仓库映射成可读路径,查询直接读远端文件原地分析;HF Data Studio 的 SQL 控制台由 DuckDB WASM 驱动,完整 OLAP 引擎跑在用户浏览器里;DuckDB v1.2.1 起 CLI 自带本地 UI(duckdb --ui,由 MotherDuck 与 DuckDB Labs 合作开发),大查询切到本地跑,绕开浏览器资源限制。
DuckDB and Hugging Face worked together (announced May 2024 with DuckDB v0.10.3) to add the hf:// path scheme on top of the httpfs extension, mapping a dataset repository onto readable paths so queries analyze files exactly where they are. Hugging Face Data Studio's SQL console is powered by DuckDB WASM, running a full OLAP engine inside the user's browser. Starting with DuckDB v1.2.1, the CLI ships a local UI (duckdb --ui, built by MotherDuck and DuckDB Labs together) so heavy queries run locally, bypassing browser resource limits.
- 结果
一行 SQL 直查云端数据集:select * from 'hf://datasets/glaiveai/reasoning-v1-20m@~parquet/default/train/*.parquet' limit 500,其中 @~parquet 走 HF 的 Parquet 转换,列式读取只扫需要的列;Data Studio 里点"Copy for DuckDB CLI"一键把查询搬到本地 UI。数据集越来越多以 Parquet 发布,而这正是 DuckDB 原生能读、无需全量物化进内存的格式。
One line of SQL queries a cloud dataset directly: select * from 'hf://datasets/glaiveai/reasoning-v1-20m@~parquet/default/train/*.parquet' limit 500, where @~parquet uses Hugging Face's Parquet conversions so columnar reads scan only the needed columns. Data Studio's "Copy for DuckDB CLI" button moves a query to the local UI in one click. Datasets are increasingly published as Parquet — exactly the format DuckDB reads natively without materializing everything into memory.
- 机制根因
httpfs 的 range read 只拉取查询需要的行组,网络传输量与"查询的选择性"成正比而非与"文件大小"成正比;WASM 把完整 OLAP 引擎塞进浏览器,计算下推到用户侧,服务端只做文件托管——查询并发由用户设备分担,这是中心化数仓做不到的成本结构;Parquet 作为开放列式格式,是"免下载查询"得以成立的地基。
httpfs range reads fetch only the row groups a query needs, so network transfer scales with query selectivity, not file size. WASM packs a complete OLAP engine into the browser, pushing compute to the user side while the server only hosts files — query concurrency is absorbed by user devices, a cost structure no centralized warehouse can match. Parquet as an open columnar format is the foundation that makes "query without downloading" possible.
- 教训
开放列式格式(Parquet)+ 轻量可嵌入引擎,足以让数据平台把"查询"这件事外包给用户侧;AI 数据集这类读多写少、探索式访问的场景,server 模式是多余成本;协议层创新(hf://)有时比引擎优化更能改变用户体验——它消灭的是"下载"这个步骤本身。
An open columnar format (Parquet) plus a lightweight embeddable engine is enough for a data platform to outsource "querying" itself to the user side. For read-heavy, exploratory workloads like AI datasets, server mode is unnecessary cost. Protocol-level innovation (hf://) sometimes changes user experience more than engine optimization — it eliminates the "download" step itself.
相关产品:DuckDB 相关能力:被嵌入的分析引擎:BI 产品的"标配内核" 最后核验:2026-10-02
Layers:躲开上游调价暴涨,每租户一个 DuckDB 做多租户分析(2025) 成功经验
多租户 SaaS
成本优化
客户画像分析
110ms SLA
Tinybird 迁移
零拷贝分析
- 场景
Layers 给零售品牌做站内搜索、推荐和用量分析,流量又大又抖:典型一天 1 亿请求,大促时逼近 10 亿。客户-facing 仪表盘要求查询落在 110ms 以内,否则拖慢店面体验;不同租户行为差异极大,小品牌涓涓细流、大零售商 firehose 级数据量,留存窗口需求也不同。团队最初用 Postgres + 流式管线 + serverless 数据 API 拼凑,Postgres 跑分析先到瓶颈,又搬到 Tinybird(serverless ClickHouse)。随后 Tinybird 调价——按 MotherDuck 官方案例 TL;DR 表的口径,"每租户预计成本涨 100 倍";同一页面的 Results 总结则写成"避免了预计 1000 倍的成本上涨",两处数字不一致(均为 MotherDuck 厂商口径),引用时须注明。
Layers powers on-site search, recommendations, and usage analytics for retail brands, with traffic that is large and spiky: 100M requests on a typical day, pushing toward the billion mark during promotions. Customer-facing dashboards must answer within 110ms or they drag down storefront performance, and tenants behave very differently — boutique brands trickle data while big-box retailers send firehose volumes, with different retention needs. The team first stitched together Postgres plus a streaming pipeline plus a serverless data API; Postgres hit the analytics bottleneck first, so they moved to Tinybird (serverless ClickHouse). Then Tinybird changed its pricing — per MotherDuck's official case study TL;DR table, a "projected 100x increase in cost for each tenant"; the same page's Results section says the team "avoided a projected 1,000x cost increase." The two figures disagree (both are MotherDuck vendor claims), so quote them with the caveat. (
https://motherduck.com/case-studies/layers-multi-tenant-data-warehouse/)
- 决策
Layers 暂停路线图重做架构,核心两条:第一,每租户独立 DuckDB 引擎(MotherDuck 的 hypertenancy),一个 SaaS 客户就是一个 mini 数仓,进程级隔离、按租户计量 CPU 秒数;第二,放弃"永远在线的流式摄入"心态,改走批式对象存储:Cloudflare Pipelines 每 5–15 分钟把压缩 Parquet 微批落到 R2,MotherDuck 对文件零拷贝就地查询,不经过 ETL 进专有存储。
Layers paused roadmap work and rebuilt the architecture around two ideas. First, one DuckDB engine per tenant (MotherDuck hypertenancy): each SaaS customer gets its own mini data warehouse with process-level isolation and per-tenant metering of CPU-seconds. Second, abandon the "always-on streaming ingestion" mindset for batched object storage: Cloudflare Pipelines land compressed Parquet micro-batches to R2 every 5 to 15 minutes, and MotherDuck queries the files in place with zero-copy analytics — no ETL into proprietary storage.
- 结果
躲开了上游调价带来的成本暴涨;小租户边际成本降到"几分钱的零头"(fractions of a penny),得以新开免费档;仪表盘回到 110ms SLA(少了一次经过外部查询 API 的网络跳转);大零售商的长扫描不再跟小租户抢资源;留存窗口变成销售谈判筹码——免费档 5 天、企业版 180+ 天,按租户改生命周期规则即可,无需数据迁移。以上均为 MotherDuck 官方案例口径。
The team dodged the cost explosion from the vendor's pricing change. Marginal cost for small tenants fell to "fractions of a penny," unlocking a freemium tier. Dashboards are back inside the 110ms SLA (one fewer network hop through an external query API). Long scans from large retailers no longer contend with small tenants' workloads. Retention windows became a sales lever — 5 days for free tier, 180+ days for enterprise — adjusted per tenant by editing lifecycle rules, with no data migration. All figures are MotherDuck's official case-study claims.
- 机制根因
DuckDB 轻到可以按需实例化,这是"每租户一个引擎"的前提——共享集群架构下大租户的重查询必然产生 noisy neighbor,而独立 DuckDB 进程天然隔离。对象存储范式解决成本:Parquet 落对象存储每 GB 每月几分钱,MotherDuck 直接读文件,历史数据"留着也无妨",留存从成本项变成可配置项。批式替代流式的前提是诚实面对 SLA:客户要的"实时"是 5–15 分钟新鲜度,不是毫秒级,架构不必为伪需求付流式税。
DuckDB is light enough to be instantiated on demand, which is what makes "one engine per tenant" possible — under a shared-cluster architecture, a large tenant's heavy queries inevitably create noisy neighbors, while separate DuckDB processes are isolated by construction. The object-storage paradigm fixes the cost side: Parquet in object storage costs pennies per GB-month, and MotherDuck reads files directly, so keeping history "just in case" is nearly free and retention becomes a configuration item instead of a cost item. Replacing streaming with batching only works when the SLA is honestly defined: what customers meant by "real-time" was 5-to-15-minute freshness, not millisecond latency, so the architecture stopped paying the streaming tax for a phantom requirement.
- 教训
多租户 SaaS 的分析成本必须能按租户归因计量,否则大租户在悄悄补贴小租户、定价无从谈起;vendor 的定价模型变更是真实的架构风险,开放格式(Parquet)+ 可移植引擎是最便宜的保险;先问清"实时"的真实定义再选架构,分钟级批和毫秒级流的成本可差数个数量级(架构量级差异,取决于数据量与实时性 SLA);AI 推荐/ML 训练可以直接复用同一套租户隔离的分析查询,不必另起炉灶。
A multi-tenant SaaS must be able to attribute and meter analytics cost per tenant, or large tenants silently subsidize small ones and pricing has no basis. A vendor's pricing-model change is a real architectural risk; open formats (Parquet) plus a portable engine are the cheapest insurance. Define what "real-time" actually means before choosing an architecture — minute-level batch and millisecond streaming can differ in cost by orders of magnitude (an architectural magnitude difference, depending on data volume and latency SLA). ML recommendation and model-training workloads can reuse the same tenant-isolated analytical queries instead of building a separate stack.
来源
MotherDuck official case study "Multi-Tenant Data Warehouse: Per-Customer Analytics" (vendor claim
—
note the 100x vs 1,000x discrepancy between the TL
—
相关产品:DuckDB、ClickHouse、PostgreSQL(社区版) 相关能力:进程内 OLAP:pip install 即拥有的分析引擎、单写者 + 无 Server:并发与多租户的天花板 最后核验:2026-10-02
Rill:DuckDB 作为默认内嵌 OLAP 引擎,BI 项目开箱即有分析库(2024–2026) 成功经验
BI 内嵌引擎
默认 OLAP
运营智能仪表盘
零运维
开箱即用
- 场景
Rill is a code-first BI / operational-intelligence platform built for low-latency, high-performance, drill-down dashboards with alerts and scheduled reports. Every project needs an analytical foundation that is fast and interactive out of the box — users cannot be asked to stand up a warehouse and load data first. The default OLAP engine therefore had to be zero-ops, fast to start, and fast to query. (
https://github.com/rilldata/agent-skills/blob/HEAD/skills/rill-development/SKILL.md)
- 决策
把 DuckDB 定为默认内嵌 OLAP 引擎:rill.yaml 里不配置 olap_connector 时,Rill 自动初始化一个托管的 duckdb OLAP 数据库并作为默认连接器;模型(model)默认把表物化进默认 OLAP 连接器;managed: true 时 Rill 自动置备,用户不碰基建与凭证。重负载用户可显式换 ClickHouse、Druid、Pinot、StarRocks,或直连 Snowflake/BigQuery/Databricks 等外部 OLAP。
DuckDB became the default embedded OLAP engine: when rill.yaml does not configure olap_connector, Rill automatically initializes a managed duckdb OLAP database and uses it as the default connector. Models materialize tables into the default OLAP connector, and with managed: true Rill provisions everything — users never touch infrastructure or credentials. Heavy workloads can explicitly switch to ClickHouse, Druid, Pinot, StarRocks, or live-connect to Snowflake, BigQuery, and Databricks.
- 结果
Rill 官方 agent-skill 文档确认上述默认行为;Rill Cloud 对托管 duckdb 按实际 CPU/内存/磁盘用量计费。Rill 的"operational intelligence"定位(低延迟下钻、告警、定时报告)得以开箱即用:从 YAML + SQL 代码到可交互仪表盘,中间不需要用户做任何数据库运维。
Rill's official agent-skill documentation confirms this default behavior, and Rill Cloud bills managed DuckDB by actual CPU, memory, and disk usage. Rill's "operational intelligence" promise — low-latency drill-downs, alerts, scheduled reports — works out of the box: from YAML plus SQL code to an interactive dashboard with zero database operations in between.
- 机制根因
DuckDB 进程内、无依赖、启动快,适合"每个项目一个库"的托管模式——这正是传统 server 式 OLAP 做不到的粒度;列式向量化保证仪表盘下钻交互的延迟;Rill 的驱动针对各引擎特性做专门优化,而 DuckDB 是默认路径,意味着它的"默认体验"决定了产品的第一印象。managed duckdb 也是 Rill 仅有的两种可自动置备 OLAP 之一(另一种是 ClickHouse)。
DuckDB is in-process, dependency-free, and fast to start, which suits the "one database per project" managed model — a granularity traditional server-style OLAP cannot reach. Columnar vectorized execution keeps dashboard drill-down latency low. Rill's drivers are heavily optimized per engine, and DuckDB is the default path, meaning its default experience defines the product's first impression. Managed DuckDB is also one of only two OLAP engines Rill can auto-provision (the other being ClickHouse).
- 教训
BI/数据产品的引擎选型逻辑里,"默认路径必须零运维"优先于峰值性能;DuckDB 的价值不只是快,而是"把 OLAP 从服务变成库",让上层产品敢把它当默认项——这是 SQLite 当年走过的路;可替换架构(默认 DuckDB、重负载换 ClickHouse/数仓)比"一个引擎打天下"更诚实,也更能留住从小到大的用户。
In engine selection for BI/data products, "the default path must be zero-ops" outranks peak performance. DuckDB's value is not just speed — it turns OLAP from a service into a library, letting products above it dare to make it the default, the same road SQLite walked. A replaceable architecture (DuckDB by default, ClickHouse or a warehouse for heavy loads) is more honest than "one engine for everything," and retains users as they grow from small to large.
相关产品:DuckDB、ClickHouse、StarRocks、Snowflake 相关能力:被嵌入的分析引擎:BI 产品的"标配内核" 最后核验:2026-10-02
UDisc:dbt 任务从 6 小时压到 30 分钟,飞盘 App 的全量球场统计上线(2024) 成功经验
dbt 迁移
用户统计仪表盘
MongoDB 分析短板
中小团队选型
Hex 生态
- 场景
UDisc is the leading disc golf app (16,000+ courses worldwide, roughly 90% market share) with MongoDB as its transactional database. Annual usage reports were stitched together with manual scripts, and the course-stats dashboard for volunteer ambassadors initially dared to show only 30 days of data — any more would slow down the production database. As disc golf exploded during the pandemic and data volumes surged, "ad hoc analytics queries on MongoDB alone" became nearly impossible to sustain. (
https://motherduck.com/case-studies/udisc-motherduck-sports-management/)
- 决策
团队把主流方案试了个遍:ClickHouse、Snowflake、Databricks、BigQuery、Postgres,"对我们这种体量来说都太贵、太复杂"。2024 年春选定 MotherDuck,两个触因:一是看到仪表盘前端 Hex 自己在用 DuckDB,二是 MotherDuck 底座就是开源 DuckDB 引擎。作为员工持股的精益创业团队,他们要的是"简单且便宜"。
The team tried everything: ClickHouse, Snowflake, Databricks, BigQuery, and Postgres — "too expensive and too complex for our use case." In spring 2024 they chose MotherDuck for two reasons: they learned that Hex (part of their dashboard front-end) used DuckDB themselves, and MotherDuck is built on the open-source DuckDB engine. As an employee-owned lean startup, they needed something simple and cost-effective.
- 结果
POC 对比——同样一个典型查询,MotherDuck 5 秒出结果,Postgres 跑 2 分钟还没跑完;同一份充分优化过的 dbt 任务,Postgres 上 6 小时(还不算稍改查询就要重做的索引工作),MotherDuck 上 30 分钟(12 倍)。上线后大使们几秒钟就能打开某球场的全生命周期历史+全部统计图表——"这在以前根本不可能"。年度使用报告从"好几个 1 小时脚本"变成"几秒钟出全年统计"。以上均为 MotherDuck 官方案例口径。
POC comparison — a typical query took 5 seconds in MotherDuck versus still incomplete after 2 minutes in Postgres; the same fully-optimized dbt job took 6 hours in Postgres (not counting index rebuilds often required after small query changes) versus 30 minutes in MotherDuck (12x). After launch, ambassadors could load a course's lifetime history with all stats and charts in seconds — "just not even possible before." The annual usage report went from "several 1-hour scripts" to "all the year's stats in a couple of seconds." All figures are MotherDuck's official case-study claims.
- 机制根因
文档库/行式库做大聚合是模型错配:MongoDB 的聚合管线为事务文档设计,扫全量做统计时没有列式裁剪与向量化;Postgres 同理,行式存储在大范围聚合下逐行开销占主导。DuckDB 列式存储只读需要的列、向量化批量算,这类聚合查询通常快得多(取决于扫描占比,列式裁剪与向量化是主因)。关键洞察是"数据量不大≠不需要列式引擎"——UDisc 自认"数据库整体不算大",但查询模式是 OLAP,引擎模型对了,小数据也能差出数量级(该迁移的实测对比,非通用加速比)。
Running large aggregations on a document/row store is a model mismatch. MongoDB's aggregation pipeline is designed for transactional documents, with no columnar pruning or vectorization when scanning everything for stats; Postgres has the same row-store problem, where per-row overhead dominates wide aggregations. DuckDB's columnar storage reads only the needed columns and computes in vectorized batches, so this class of aggregation queries is usually much faster (depending on scan share; columnar pruning and vectorization are the main drivers). The key insight: "small data does not mean no columnar engine needed" — UDisc considered its database "not all that large," but the query pattern was OLAP, and with the right engine model even small data shows order-of-magnitude gaps (a measured comparison from this migration, not a universal speedup ratio).
- 教训
OLTP/文档库兼职 OLAP 的天花板来得比数据量天花板早得多,"只放 30 天数据保生产库"是典型的预警信号;中小团队选型第一性原理是"简单+便宜",重型数仓的复杂度税对小团队是净负担;dbt + DuckDB 让本地开发与云端生产跑同一引擎,dev/prod 一致性是 MotherDuck 这类方案的隐藏价值;生态信号(Hex 自己用 DuckDB)有时比基准测试更能说明问题。
The ceiling of running OLAP on an OLTP/document database arrives far earlier than the data-volume ceiling — "only 30 days of data to protect the production DB" is the classic warning sign. For small teams, the first principle of selection is "simple and cheap"; a heavyweight warehouse's complexity tax is a net burden on a small team. dbt + DuckDB runs the same engine for local development and cloud production, and that dev/prod consistency is the hidden value of this kind of solution. Ecosystem signals (Hex using DuckDB themselves) sometimes say more than benchmarks.
相关产品:DuckDB、MongoDB、PostgreSQL(社区版) 相关能力:进程内 OLAP:pip install 即拥有的分析引擎 最后核验:2026-10-02
Amazon:2004 年圣诞宕机催生 Dynamo,购物车成为 DynamoDB 的起点(2004–2012) 成功经验
起源故事
键值存储
高可用设计
电商大促
- 场景
On December 12, 2004, all of Amazon was running on relational databases when a massive rack-cluster failure took the entire site down on the busiest day of the year. The postmortem revealed that roughly 70% of Amazon's storage access was fundamentally key-value: "Give me my shopping cart, give me this, give me that - one attribute and give me the result of that." Multi-datacenter operations were worse: Vogels recalled that the team once deliberately pulled the plug on a data center to see what would happen, and "all the failures were with relational databases" - plugging the data center back into live operations was "a nightmare." For shopping carts, lost data means lost sales, so availability is the hard requirement, while traditional replication technology "chooses consistency over availability" - a model mismatch. (
https://siliconangle.com/2021/06/07/digging-covers-aws-timestream-database-amazon-cto-werner-vogels/)
- 决策
Amazon 内部立项 Dynamo:去中心化、无主节点的键值存储,用一致性哈希做分区复制、向量时钟做版本管理、quorum 读写、gossip 做成员感知与故障检测,把"一致性/可用性/成本/性能"的权衡旋钮交还给应用层(购物车选高可用+最终一致,冲突由应用合并)。2007 年 10 月以 SOSP 论文形式公开发表,发表前已在生产环境支撑多个核心服务。
Amazon built Dynamo internally: a decentralized, leaderless key-value store using consistent hashing for partitioned replication, vector clocks for versioning, quorum reads/writes, and gossip-based membership and failure detection - handing the availability/consistency/cost/performance tradeoff knobs back to the application layer (the cart chose high availability plus eventual consistency, with conflicts merged by the application). It was published as a SOSP paper in October 2007, after already supporting several core services in production.
- 结果
SOSP 2007 论文披露的生产数据:购物车服务单日处理数千万请求、带来超 300 万次结账,会话状态服务同时承载数十万活跃会话,扛过假日购物季极端峰值"没有任何宕机"。2012 年 1 月 18 日,DynamoDB 作为全托管服务正式发布(GeekWire 报道),发布时内部已有 Cloud Drive、IMDb、Kindle、广告平台在用,首批外部客户包括 SmugMug 和 Elsevier。以上单日请求/结账数字出自论文一手生产数据;"70% 键值""首个服务是购物车"为 Vogels 2021 年口述回忆。
Production figures disclosed in the SOSP 2007 paper: the Shopping Cart Service served tens of millions of requests resulting in well over 3 million checkouts in a single day, the session-state service carried hundreds of thousands of concurrently active sessions, and Dynamo "scaled to extreme peak loads efficiently without any downtime during the busy holiday shopping season." On January 18, 2012, DynamoDB launched as a fully managed service (GeekWire), already running Amazon's internal Cloud Drive, IMDb, Kindle, and advertising platform, with early external customers including SmugMug and Elsevier. The daily request/checkout figures come from first-hand production data in the paper; the "70% key-value" and "shopping cart was the first service" claims are Vogels's 2021 recollections.
- 机制根因
第一层是访问模式与存储模型的匹配:购物车、会话、畅销榜都是主键点查,不需要关系模型的复杂查询能力,为用不上的 JOIN 和 ACID 付出"昂贵硬件+资深 DBA"的代价是浪费;且关系型复制在分区故障时优先保一致性,直接违反"永远可写"的业务要求。第二层是产品化教训:Dynamo 虽好,但每个团队自己运维 Dynamo 集群的复杂度成了内部推广障碍——Vogels 在 2012 年发布时承认,内部服务方抱怨"要自己成为 Dynamo 专家、自己做一致性/性能/可靠性的权衡";于是 DynamoDB 的核心产品决策是"服务化":全托管、按表声明吞吐量,运维复杂度由 AWS 吞掉。技术验证在前(2007),服务化在后(2012),顺序不能反。
The first layer is access-pattern/storage-model fit: carts, sessions, and bestseller lists are primary-key point lookups that need none of the relational model's complex query power - paying for unused JOINs and ACID with "expensive hardware and highly skilled personnel" is waste, and relational replication prioritizes consistency during partitions, directly violating the "always writable" business requirement. The second layer is a productization lesson: Dynamo worked, but the operational complexity of each team running its own Dynamo cluster blocked internal adoption - at the 2012 launch Vogels admitted internal service owners complained about having to "become experts on Dynamo" and make the consistency/performance/reliability tradeoffs themselves. So DynamoDB's core product decision was "servicification": fully managed, declare throughput per table, with AWS absorbing the operational complexity. Technical validation came first (2007), the managed service second (2012) - the order cannot be reversed.
- 教训
从访问模式反推存储模型:70% 是 KV 就别硬上关系型,为用不上的功能交税是架构原罪;可用性需求决定一致性取舍,而不是反过来——先问"丢了这条数据损失多少钱",再定一致性级别;内部自研技术的最大采用障碍往往不是性能而是运维复杂度,"托管化"是技术走向产品的关键一跃;一篇生产系统论文的价值在于"先在真实流量里跑过",Dynamo 论文发表前已支撑核心服务,这是它区别于纯学术工作的根本。
Derive the storage model from the access pattern: if 70% is key-value, don't force it into a relational database - paying for features you never use is an architectural original sin. Availability requirements decide the consistency tradeoff, not the other way around: first ask "how much does losing this record cost," then set the consistency level. The biggest adoption barrier for internally built technology is usually operational complexity, not performance - "managed-ification" is the key leap from technology to product. A production-systems paper is valuable precisely because it "ran in real traffic first" - Dynamo was already supporting core services before publication, which is what fundamentally separates it from purely academic work.
相关产品:Amazon DynamoDB 相关能力:Serverless 零运维:流量不可预测时"先跑起来"的最短路径 最后核验:2026-10-02
Canva:感知哈希反向图片搜索上线即不可用——分区键的数据分布教训(2022) 失败教训
分区键设计
数据倾斜
容量规划
上线事故
- 场景
Canva's ever-growing media library needed deduplication and content moderation, so the team built an in-house reverse image search on perceptual hashes (pHash): images are hashed and stored in DynamoDB, and queries find similar images within a Hamming-distance threshold. They implemented multi-index hashing (Norouzi 2012) on DynamoDB: split the perceptual hash into N segments, use each segment prefixed with its slot index as the partition key and the image ID as the sort key; at query time split the query hash the same way, run parallel Queries against each slot, then merge, deduplicate, and filter out results above the Hamming-distance threshold in the application. They chose 4 segments for recall guarantees (the pigeonhole principle guarantees all matches within Hamming distance 3 are found). (
https://www.canva.dev/blog/engineering/simple-fast-and-scalable-reverse-image-search-using-perceptual-hashes-and-dynamodb/)
- 决策
按论文算法实现并直接上线,假设感知哈希在全库上近似均匀分布——每段槽位下的候选集都很小,合并过滤的代价可忽略。
Implement per the paper's algorithm and ship it, assuming perceptual hashes are approximately uniformly distributed across the corpus - so each slot's candidate set would be tiny and the merge-filter cost negligible.
- 结果
上线即不可用:查询跑数分钟、读容量消耗是预期的 20 倍、返回的有效结果却很少。DynamoDB 每次查询返回海量候选项,几乎全在合并步骤被过滤掉,等于花 20 倍的读容量读了一堆垃圾。这是 Canva 工程博客亲笔记录的生产事故复盘,非第三方转述。
Unusable immediately after rollout: queries took minutes to run, consumed 20x more read capacity than expected, and returned very few valid results. Each DynamoDB query returned a massive candidate set that was almost entirely discarded in the consolidation step - effectively paying 20x read capacity to read garbage. This is a production-incident retrospective written in Canva's own engineering blog, not a third-party retelling.
- 机制根因
两层根因都是"数据分布"问题,而 DynamoDB 恰恰把数据分布的责任交给了应用层的键设计。第一,用户经常上传重复或轻微修改的图片,哈希的唯一性分布远不如预期——每个槽位下的候选比模型假设多得多。第二更致命:感知哈希对低复杂度图片(如单色矢量图)会产生高度均匀的哈希(如 `AAAAAAAA`),数十万张这类图片挤在同一个分区键下;查询一旦命中这些槽位,DynamoDB 一次返回数十万行。DynamoDB 按分区键哈希做物理分区,单个分区键的所有 item 落在同一分区——低基数分区键就是热点分区,论文算法的"均匀分布"假设被真实用户数据击碎。修复:窗口数降到原来的 1/4,并跳过低复杂度图片不建哈希;问题查询的返回行数从数十万降到数十。修好后系统规模:100 亿+图片哈希,平均 40ms、p95 60ms,峰值 2000+ qps(约 10k RCU)——证明 DynamoDB 扛得住这个量级,锅在建模不在产品。
Both root causes are data-distribution problems, and DynamoDB is exactly the database that hands data-distribution responsibility to the application's key design. First, users commonly upload duplicates or slightly modified images, so the hash uniqueness distribution was far worse than assumed - far more candidates per slot than the model predicted. Second and more fatal: perceptual hashes produce highly uniform hashes for low-complexity images (e.g., single-colored vectors like `AAAAAAAA`), packing hundreds of thousands of such images under a single partition key; whenever a query hit those slots, DynamoDB returned hundreds of thousands of rows at once. DynamoDB hashes the partition key for physical placement, so every item under one partition key lands on the same partition - a low-cardinality partition key is a hot partition, and real user data shattered the paper algorithm's uniform-distribution assumption. The fix: reduce the window count to one quarter and skip hashing low-complexity images; problematic queries went from hundreds of thousands of returned rows to tens. After the fix the system holds 10B+ image hashes at 40ms average, 60ms p95, and 2000+ queries/sec peak (about 10k RCU) - proving DynamoDB handles the scale fine; the fault was in the modeling, not the product.
- 教训
DynamoDB 不拯救糟糕的键设计——分区键的基数和分布是上线前必须用真实数据验证的第一性原理,"哈希应该很均匀"是原罪级假设;低基数或值分布极不均匀的属性绝不能裸做分区键,要么加盐打散,要么在写入侧过滤;上线前用生产量级+生产分布的数据做压测,而不是用均匀随机数据;反例的价值在于区分"产品不行"和"用法不对":Canva 修好建模后同一套 DynamoDB 跑出 40ms 平均延迟,说明这是 access pattern 建模事故,不是选型事故。
DynamoDB does not rescue bad key design - partition-key cardinality and distribution are first principles that must be validated against real data before launch; "the hashes should be fairly uniform" is an original-sin assumption. Never use a low-cardinality or heavily skewed attribute as a bare partition key: either salt-and-scatter it or filter at write time. Load-test with production-scale, production-distribution data, not uniformly random data. The value of an anti-pattern case is distinguishing "the product is bad" from "the usage was wrong": after fixing the modeling, the same DynamoDB delivered 40ms average latency - this was an access-pattern modeling incident, not a selection incident.
相关产品:Amazon DynamoDB 相关能力:单表设计的心智税:"access pattern 先行"既是超能力也是枷锁 最后核验:2026-10-02
Cognito Forms:双平台实测后弃选 DynamoDB——按预测付费不适合可变负载(约 2014–2015) 失败教训
PoC
选型评估
弃选 DynamoDB
KV 选型
计费模型
- 场景
Cognito Forms(在线表单 SaaS)创立之初做数据库选型。先定大方向:SQL vs NoSQL——按其调研口径,NoSQL 比企业级 SQL 便宜约 $5,000/月/TB,且 NoSQL 可按负载分区。定下 NoSQL 后,把候选收敛到 Amazon DynamoDB 与 Microsoft Azure Table Storage。
Cognito Forms (online form SaaS) ran a database selection at founding. First the big direction — SQL vs NoSQL: per their research, NoSQL ran about $5,000/month/TB cheaper than enterprise SQL, and NoSQL could be partitioned by load. With NoSQL chosen, candidates narrowed to Amazon DynamoDB vs Microsoft Azure Table Storage.
- 决策
列出两家的 pros/cons 对照表后,团队把应用同时在双平台搭建并跑负载测试。DynamoDB 的问题:单次 scan/query 结果集限 1MB,查大结果要手动分页重查;predictive pricing——按预测的吞吐量付费而非实际用量,可变负载下要么超付要么被限流;高负载下需持续监控、手动调 throughput。Azure Table Storage:按实际用量付费,超额数据"无缝接受";PaaS 形态省运维;.NET 技术栈无缝集成。
After a pros/cons table for both, the team built the application on both platforms and ran load tests. DynamoDB's problems: single scan/query result sets capped at 1MB, forcing manual paginated re-queries; predictive pricing — paying for predicted rather than actual throughput, a bad fit for variable usage; sustained high load required constant monitoring and manual throughput adjustments. Azure Table Storage: pay for actual usage, excess data "seamlessly accepted"; PaaS shape saved ops effort; seamless .NET stack integration.
- 结果
弃选 DynamoDB,选 Azure Table Storage。按 Cognito Forms 自述口径,负载测试结论是"DynamoDB 高负载下需持续监控调 throughput,Azure 无缝接受超额数据"。代价是接受 Azure 当时的短板:美国仅 4 个数据中心、自动扩缩容备份等功能仍在 beta。
DynamoDB was rejected; Azure Table Storage won. Per Cognito Forms' own account, load testing concluded that "DynamoDB required constant monitoring to adjust throughput under high volume, while Azure seamlessly accepted the excess data." The price was accepting Azure's then-shortcomings: only 4 US data centers, auto-scaling backups and similar features still in beta.
- 机制根因
这是一个"计费模型即架构约束"的案例。DynamoDB 的 provisioned throughput 把容量规划前置——对负载可预测的大厂这是成本优化器,对负载不可预测的初创公司这是税:要么为峰值超付,要么在峰值时被限流后再手动调参。Azure 的按量付费把规划后置,匹配了初创公司的不确定性。选型时"谁便宜"取决于负载曲线的形状,而不取决于单价表。
A case of "pricing model as architecture constraint." DynamoDB's provisioned throughput front-loads capacity planning — a cost optimizer for large companies with predictable load, a tax for startups with unpredictable load: either overpay for peaks or get throttled at peaks and re-tune manually. Azure's usage-based pricing deferred planning, matching startup uncertainty. In selection, "which is cheaper" depends on the shape of the load curve, not the price list.
- 教训
KV 选型先看计费模型再看性能——predictive pricing 与可变负载是天然互斥;"1MB scan 上限"这类 API 配额要在 PoC 里实测对业务查询模式的影响;小团队选 PaaS 形态,本质是买"不用雇 DBA"的期权。诚实注记:该文未标注发布日期,据"美国仅 4 个数据中心""autoscaling backups 仍在 beta"等内容推断约 2014–2015 年;DynamoDB 此后推出 on-demand 按量计费模式,本案例的计费结论已过期,仅保留"计费模型要进评估项"的方法论价值。
In KV selection, evaluate the pricing model before performance — predictive pricing and variable workloads are inherently at odds; API quotas like the "1MB scan limit" must be tested against real query patterns in PoC; a small team choosing PaaS is essentially buying the "no DBA hire needed" option. Honesty note: the article carries no publication date; ~2014–2015 is inferred from "only 4 US data centers" and "autoscaling backups still in beta." DynamoDB has since launched on-demand pricing, so this case's pricing conclusion has expired — its lasting value is methodological: put the pricing model on the evaluation scorecard.
来源
Cognito Forms official blog, "Why We Chose Azure Table Storage over Amazon DynamoDB" (undated
—
相关产品:Amazon DynamoDB 相关能力:预置吞吐计费与可变负载的错配 最后核验:2026-10-02
SmugMug:从 MySQL 迁到 DynamoDB,一年内搬完 Flickr 的数百 PB(2018–2019) 成功经验
照片平台
MySQL 迁移
规模扩展
收购整合
- 场景
SmugMug's online photo-metadata layer long ran on traditional MySQL, and as data volumes and traffic grew, query response times became unpredictable at scale. The team first tried reshaping MySQL into key-value access patterns, but still could not get predictable latency at scale. After acquiring Flickr in 2018, SmugMug faced a hard task: moving Flickr out of Yahoo's data centers and onto the AWS cloud - hundreds of petabytes of data, tens of billions of photos, over 100 million users, with a migration window of about one year. (
https://aws.amazon.com/blogs/database/motivations-for-migration-to-amazon-dynamodb/)
- 决策
SmugMug 把照片平台的元数据从 MySQL 迁到 DynamoDB;随后借助 DynamoDB 的弹性,把 Flickr 的工作负载从 Yahoo 数据中心迁入 AWS。选型逻辑很直接:照片元数据的访问模式是按 ID 点查,与 DynamoDB 的分区键模型契合(点查模式是该契合的前提,扫描/复杂查询为主的 workload 不适用此结论);收购整合的时间窗口容不下"自建分片集群再调优"的长周期。
SmugMug moved its photo platform metadata from MySQL to DynamoDB, then used DynamoDB's elasticity to migrate Flickr's workloads from Yahoo data centers into AWS. The selection logic was direct: photo metadata access is point lookup by ID, a good fit for DynamoDB's partition-key model (point lookups are the premise of this fit; workloads dominated by scans or complex queries do not share this conclusion), and the acquisition-integration timeline left no room for a long "build your own sharded cluster and tune it" cycle.
- 结果
AWS 官方数据库博客(2023)口径:迁移后 SmugMug 获得"与存储量无关的可预测查询响应时间";Flickr 迁移在 1 年内完成,数百 PB、数百亿张照片、1 亿+用户搬上 DynamoDB 及其他 AWS 服务——"如果你在 Flickr 上看一张照片,你就是在和 DynamoDB 交互"。以上数字均为 AWS 厂商口径,未找到 SmugMug 工程博客披露的对应数字,也无第三方独立复现,引用须注明口径。迁移前后的延迟分布、成本变化未找到公开数据。
Per AWS's official Database Blog (2023): after migrating, SmugMug got predictable query response times "regardless of their storage volume"; the Flickr migration completed within one year - hundreds of petabytes, tens of billions of photos, and 100M+ users moved onto DynamoDB and other AWS services: "If you are viewing a photo on Flickr, you are interacting with DynamoDB." All figures above are AWS's vendor claim; no corresponding numbers were found in SmugMug's own engineering blog, and no independent third-party reproduction exists - quote them with the source caveat. No public data found for latency distributions or cost changes before and after the migration.
- 机制根因
MySQL 的 B-tree + 缓存体系在数据量增长时延迟分布变宽(p99 长尾恶化),而 DynamoDB 的"按分区键哈希做数据放置+请求路由"让延迟与数据量解耦:表多大,点查都是一次分区定位。SmugMug 先"把 MySQL 当 KV 用"的中间态很有说服力——问题不在 SQL 语言本身,而在存储引擎的扩展模型:单机引擎靠纵向扩展和缓存续命,分布式分区引擎靠加分区接近线性扩展(取决于分区键均匀度与一致性配置)。收购整合场景放大了 DynamoDB 的价值:迁移窗口按月算,没有时间留给分片键设计评审和压测调优,"开箱即用的弹性"是买时间。
MySQL's B-tree plus cache architecture widens the latency distribution as data grows (p99 tail degrades), while DynamoDB's "hash the partition key for data placement plus request routing" decouples latency from data volume: however large the table, a point lookup is a single partition lookup. SmugMug's intermediate step of "using MySQL as a KV store" is telling - the problem was never the SQL language, but the storage engine's scaling model: single-node engines survive on vertical scaling and caching, while partitioned engines scale near-linearly by adding partitions (depending on partition-key uniformity and consistency configuration). The acquisition scenario amplified DynamoDB's value: with the migration window measured in months, there was no time for sharding-key design reviews and load-test tuning - "elasticity out of the box" was buying time.
- 教训
当延迟的可预测性(p99/p999 分布)比功能丰富度更重要时,就该换存储引擎,而不是在原引擎上继续打补丁;"把关系型数据库当 KV 用"的中间态是危险信号,说明访问模式和引擎模型已经分叉,越晚承认迁移成本越高;收购后的技术栈整合是 DynamoDB 这类免运维弹性服务的典型高价值场景——时间窗口比单价更重要;厂商口径的迁移规模数字必须标注,不把"数百 PB"当成可复现的工程指标。
When latency predictability (the p99/p999 distribution) matters more than feature richness, change the storage engine instead of patching the old one. "Using a relational database as a KV store" is a danger signal: the access pattern and the engine model have diverged, and the later you admit it, the higher the migration cost. Post-acquisition stack consolidation is a classic high-value scenario for zero-ops elastic services like DynamoDB - the time window matters more than the unit price. Vendor-claimed migration scale figures must be labeled; never treat "hundreds of petabytes" as a reproducible engineering metric.
相关产品:Amazon DynamoDB、MySQL 相关能力:Serverless 零运维:流量不可预测时"先跑起来"的最短路径 最后核验:2026-10-02
Snapchat:故事收件箱从 GCP 迁到 DynamoDB,除夕峰值零人工干预(约 2022–2023) 成功经验
社交应用
突发流量
跨云迁移
故事流
- 场景
A set of Snapchat's critical storage use cases ran on Google Cloud Platform, among them "story inboxes" (collections of story posts) - a core but difficult feature: it had to sustain millions of writes per second while storage and throughput costs on GCP ran high. The team had an additional strategic motive: expanding cloud providers beyond GCP alone. Social-app traffic is inherently spiky - New Year's Eve is the year's largest spike, and capacity provisioned year-round for that peak is a year-round tax. (
https://aws.amazon.com/blogs/database/motivations-for-migration-to-amazon-dynamodb/)
- 决策
把故事收件箱等关键存储从 GCP 迁到 DynamoDB,目标是在扛住数百万写/秒的同时显著降本,并实现多云供应商布局。选型逻辑:故事流是写密集、突发、按用户/故事 ID 点查的 KV 负载,不需要复杂查询,DynamoDB 的自动分区与弹性伸缩正好对准"峰值不可预测"这个痛点。
Move story inboxes and other critical storage from GCP to DynamoDB, aiming to sustain millions of writes per second while cutting costs significantly, and to establish a multi-cloud provider footprint. The selection logic: the stories feed is a write-heavy, spiky, point-lookup-by-user/story-ID key-value workload with no need for complex queries - DynamoDB's automatic partitioning and elastic scaling aim exactly at the "unpredictable peaks" pain point.
- 结果
AWS 官方数据库博客(2023)口径:Snapchat 把 100% 用户迁到 DynamoDB;除夕夜峰值流量"无需任何人工干预"扛过;"省掉了 GCP 上绝大部分存储与吞吐成本,每年节省数百万美元"。另据 AWS Data Roadshow 2023 宣讲材料引 Snap 基础设施工程负责人 Dave Killian:"这已是 Snapchat 最稳定的系统之一",相比老系统每年省掉数百万美元的过度预置成本。以上均为 AWS 方面口径,未找到 Snap 工程博客披露的对应数字与迁移时间线,确切迁移年份未找到公开数据,标题年份为按 AWS 材料发布时间的约数。
Per AWS's official Database Blog (2023): Snapchat migrated 100% of its users to DynamoDB; peak New Year's Eve traffic was absorbed "with no operator intervention"; it "eliminated most of the storage and throughput costs incurred in GCP, saving millions of dollars per year." Separately, AWS Data Roadshow 2023 deck materials quote Snap Infrastructure Engineering Lead Dave Killian: this is now "one of the most stable systems at Snapchat," eliminating "millions of dollars of over-provisioning cost per year" versus the legacy system. All of the above is AWS-side sourcing; no corresponding figures or migration timeline were found in Snap's own engineering blog, and no public data pins down the exact migration year - the year in the title is an approximation based on when the AWS materials were published.
- 机制根因
故事收件箱的写入是"追加帖子到用户收件箱"的 KV 写,读是按 ID 点查——访问模式简单,但写入速率在节假日呈脉冲式。传统方案要为脉冲峰值预置全年容量,闲置部分就是纯成本;DynamoDB 的分区自动分裂+按需/自动伸缩把"为除夕夜买单"变成"为实际流量买单"。零人工干预过峰值的技术底座是:分区是 DynamoDB 内部管理的物理单元,热点时自动分裂、请求路由自动跟随,不需要 DBA 半夜起床加节点或改分片键。这是"弹性"作为成本项的典型案例:省掉的不是机器单价,而是为峰值预置的闲置容量。
Story-inbox writes are "append a post to a user's inbox" KV writes, and reads are point lookups by ID - a simple access pattern, but write rates arrive in holiday pulses. Conventional setups must provision year-round capacity for the pulse peak, and the idle portion is pure cost; DynamoDB's automatic partition splitting plus on-demand/auto scaling turns "paying for New Year's Eve" into "paying for actual traffic." The technical foundation of zero-intervention peak riding: partitions are physical units managed inside DynamoDB - they split automatically under heat and request routing follows, with no DBA waking up at midnight to add nodes or redesign shard keys. This is the classic case of "elasticity as a cost line item": what gets saved is not the per-machine price but the idle capacity provisioned for peaks.
- 教训
突发流量场景下,把"为峰值预置的闲置容量"计入 TCO 再做选型,弹性本身就是钱;跨云迁移的真实动因往往是成本+多云战略,不只是技术优劣,写案例时不要把商业动因洗成纯技术叙事;社交/内容型产品的故事流、收件箱、时间线是 DynamoDB 的经典甜点区——写密集、KV 点查、峰值脉冲,三者占其二就值得评估;厂商宣讲材料里引用的客户原话(如"最稳定的系统之一")可以引用,但必须标注这是 AWS 转述的口径。
For spiky traffic, put "idle capacity provisioned for peaks" into the TCO before choosing - elasticity itself is money. The real drivers of cross-cloud migration are usually cost plus multi-cloud strategy, not pure technical merit; don't launder business motives into a purely technical narrative when writing the case. Stories feeds, inboxes, and timelines in social/content products are DynamoDB's classic sweet spot - write-heavy, KV point lookups, pulsed peaks; any two of the three merit an evaluation. Customer quotes relayed in vendor decks (like "one of the most stable systems") may be cited, but must be labeled as AWS-relayed sourcing.
相关产品:Amazon DynamoDB 相关能力:Serverless 零运维:流量不可预测时"先跑起来"的最短路径 最后核验:2026-10-02
CARE Risk Solutions:177 小时的 Oracle 报表在 EDB 上跑进 90 小时 成功经验
去 O
金融
报表性能
ISV
- 场景
CARE Risk Solutions, a multinational risk-management ISV for banking, financial services and insurance, hit performance and scale walls with customers on Oracle Exadata: one large public-sector bank's regulatory profitability report - hundreds of millions of records, hundreds of branches, hundreds of GL headers - took Oracle Exadata nearly 177 hours per run and sometimes could not be generated at all, so the bank repeatedly missed regulatory filing deadlines; 2TB of temp tablespace was not enough, and months of Oracle support went nowhere. (
https://enterprisedb.com/resources/customer-story/care-risk-solutions-accelerates-performance-growth-edb)
- 决策
全产品线从 Oracle 迁到 EDB(Postgres Advanced Server 12 + Postgres Enterprise Manager + 备份恢复工具 + Failover Manager),由 EDB 合作伙伴 Chemtrols Infotech 负责迁移方法论与实施;迁移前先用"最痛的那个 workload"做同等负载内部验证。
Migrated the entire product portfolio off Oracle onto EDB (Postgres Advanced Server 12, Postgres Enterprise Manager, Backup and Recovery Tool, Failover Manager), with EDB partner Chemtrols Infotech owning the migration methodology and execution; before committing, the team validated performance in-house against the same painful workload.
- 结果
EDB 官方口径(具名客户技术与创新总监 Amarjeet Tiwari 自述):事务处理提升 30%、报表生成时间降低 49%(177 小时→约 90 小时,仅做迁移与总账科目合并、无其他改动)、全产品组合迁移仅用 90 天、迁移后签下 3 家此前"大到不敢接"的新客户。
Per EDB's account, in the words of named customer Amarjeet Tiwari, Director of Technology and Innovations: 30% improvement in transaction processing, 49% reduction in report generation time (177 hours down to about 90, with no changes besides the migration and merging ledger heads), the full portfolio migrated in just 90 days, and 3 new customers signed that were previously "too large to perform effectively."
- 机制根因
瓶颈在数据库层(临时表空间撑爆),换库即换执行引擎;EDB 自动化迁移工具链承担了大部分迁移工作量,90 天完成全产品组合迁移并通过多客户数据集测试。
The bottleneck was at the database layer (temp tablespace exhaustion), so changing the database changed the execution engine; EDB's automated migration toolkit carried most of the migration workload, and the 90-day portfolio migration passed testing against multiple customer datasets.
- 教训
先拿最痛的 workload 做同等负载验证再全量迁移;性能问题会直接转化为商业天花板(接不了大客户),换库的 ROI 不只看 license,还要看它打开的市场空间。
Validate against your most painful workload at equal scale before the full migration; performance problems convert directly into a commercial ceiling (customers you cannot take on), so a database switch's ROI is not only about licenses but also about the market it unlocks.
相关产品:EDB Postgres / EDB Postgres Advanced Server(EPAS)、Oracle Database(甲骨文)、PostgreSQL(社区版) 相关能力:EPAS 的 Oracle PL/SQL 兼容 —— 去 O 路上改动最小的 PG 最后核验:2026-10-02
FBI:用 EDB Postgres Advanced Server 做 Oracle 迁移 成功经验
去 O
政务
许可成本
零重写
- 场景
The US Federal Bureau of Investigation needed to move mission-critical applications - case management and background checks - off its "cumbersome Oracle infrastructure" into the AWS cloud while protecting some of the government's most sensitive data; Oracle's exorbitant and rigid licensing fees promised to make the migration prohibitively expensive, with database costs staying sky-high afterward. (
https://mktgsite.enterprisedb.com/sites/default/files/2025-06/FBI_case_study_v6.pdf)
- 决策
选用 EDB Postgres Advanced Server(EPAS),利用其原生 Oracle 兼容性实施迁移——"Postgres 看起来、用起来就像 Oracle",应用无需重写代码。
Chose EDB Postgres Advanced Server (EPAS) for its native Oracle compatibility - "Postgres looks and feels just like Oracle" - so applications migrated with no need to recode.
- 结果
EDB 官方口径:迁移过程"简单且具成本效益"、最小化业务中断、零数据丢失;降低供应商锁定,获得跟随新技术演进的灵活性。以上均为 EDB 单方口径(客户成功故事 PDF),无具名技术细节与独立第三方验证。
Per EDB's own account: the migration was "easy and cost-effective" with minimal disruption and no data loss; vendor lock-in was reduced and the agency gained flexibility to evolve with new technology. All figures and claims are EDB's one-sided account (customer success story PDF) with no named technical detail or independent third-party verification.
- 机制根因
EPAS 的 Oracle 兼容层让应用代码零重写成为可能,消除了去 O 迁移中最大的人力成本项;订阅制替代 Oracle 按 CPU/用户数的刚性许可,从根本上改变成本结构。
EPAS's Oracle compatibility layer made zero-rewrite migration possible, eliminating the largest labor cost in a de-Oracle project; subscription pricing replacing Oracle's rigid per-CPU/per-user licensing fundamentally changed the cost structure.
- 教训
许可费用不仅是"贵",还能直接卡死现代化项目——FBI 案例里 Oracle 许可差点让迁移本身无法立项;兼容性是"不重写"承诺的兑现关键,选型时要对着自己的代码做 PoC 验证,而不是采信宣传数字。
License fees are not just "expensive" - they can kill a modernization project outright, as Oracle's nearly did the FBI's migration; compatibility is what cashes the "no rewrite" check, so validate it against your own code in a PoC rather than trusting marketing numbers.
相关产品:EDB Postgres / EDB Postgres Advanced Server(EPAS)、Oracle Database(甲骨文)、PostgreSQL(社区版) 相关能力:EPAS 的 Oracle PL/SQL 兼容 —— 去 O 路上改动最小的 PG 最后核验:2026-10-02
Lone Wolf:Oracle 许可从 100 万美元砍到 10 万,EPAS 救了公司一命 成功经验
去 O
成本敏感
ISV
许可
- 决策
迁到 EDB Postgres Advanced Server(EPAS);过渡期团队同时向 Oracle 和 EDB 双写摄入数据;EDB 称这是"最顺滑的 Oracle 逃生通道",接口与程序可原样运行、只省掉许可费。
Migrated to EDB Postgres Advanced Server (EPAS); during transition the team ingested data in parallel into both Oracle and EDB. EDB describes this as "the most fluent path out of Oracle" - interfaces and programs keep running as on Oracle, minus the licensing expense.
- 结果
具名客户工程总监 Vladimir Sanchez 自述(EDB 官方博客刊登):数据库许可成本从约 100 万美元降到约 10 万美元(−90%);一次 Oracle 侧灾难性故障中,靠着并行摄入到 EDB 的那份数据逃过一劫——"要不是 EPAS,我的团队同时向 Oracle 和 EDB 摄入数据,我们就倒闭了。EDB 把公司从灾难里救了出来。"
In the words of named customer Vladimir Sanchez, Director of Engineering (published on EDB's official blog): database licensing costs fell from approximately $1 million to $100,000 (about 90%); during a catastrophic failure of the legacy Oracle database, the parallel EPAS ingestion saved the business - "Had it not been for EPAS and my team ingesting data in parallel into Oracle and EDB, we would have been out of business. EDB saved our company from catastrophe."
- 机制根因
EPAS 让 Oracle 接口/程序"原样运行",去 O 不用付重写税,这是许可成本能砍 90% 的前提;双写摄入本质上是迁移期的廉价保险,故障时成了救命稻草。
EPAS lets Oracle interfaces and programs run as-is, so leaving Oracle costs no rewrite tax - the precondition for a 90% license cut; dual ingestion during transition is cheap insurance against the source's single point of failure, and in this case it became the lifeboat.
- 教训
许可成本会从"贵"变成"生死线";迁移过渡期的双写看似浪费,实则是对抗源端单点故障的期权;"最顺滑的逃生通道"这类评价来自真实逃生经历才有分量,选型时多问"谁真的跑过"。
License costs can move from "expensive" to "existential"; dual-write during a migration transition looks wasteful but is an option against source-side failure; "smoothest escape route" claims only carry weight when they come from someone who actually escaped - ask who has really done it.
相关产品:EDB Postgres / EDB Postgres Advanced Server(EPAS)、Oracle Database(甲骨文)、PostgreSQL(社区版) 相关能力:EPAS 的 Oracle PL/SQL 兼容 —— 去 O 路上改动最小的 PG、Isabel Group 的 TCO 实证 —— 去 O 的"省钱"有人算过账 最后核验:2026-10-02
Mastercard 在 EDB Postgres 上实现支付零停机 成功经验
高可用
零停机
支付
多数据中心
- 场景
Mastercard's payment gateway (acquired in 2014, Postgres came with it); a 99.999% uptime floor is a company-wide minimum requirement. Two big pain points: the time to switch processing to the disaster recovery site, and how to run server maintenance tasks such as vacuum full and reindex. Shared by Mastercard's Lead BizOps Engineer at EDB's Postgres Vision 2020. (
https://www.enterprisedb.com/blog/why-mastercards-secret-zero-downtime-postgres?lang=ja)
- 决策
采用 EDB 的 xDB 多主复制(xDB MMR,基于 Postgres 9.4 逻辑复制)替代二进制流复制做灾备;利用 xDB 的数据过滤能力做跨洲数据迁移以满足个人信息数据属地合规要求,并在目标端用触发器对写入数据叠加业务逻辑。
Adopted EDB's xDB multi-master replication (xDB MMR, built on Postgres 9.4 logical replication) to replace binary streaming replication for DR; used xDB's data filtering for cross-continent data migration to comply with mandates on where personally identifiable data may reside, with triggers on the target side applying additional business logic as data is written.
- 结果
切换到灾备站点"不仅更快,而且可靠地更快";真正做到"零停机";xDB 被用作实时数据迁移工具,跨洲迁移策略非常成功。以上为具名客户自述、经 EDB 官方博客刊登,无独立第三方验证。
Switching to the DR site got "not just quicker, but reliably quicker"; Mastercard could "truly offer zero downtime"; xDB proved an invaluable real-time data migration tool and the cross-continent migration strategy was very successful. All from a named customer's own account, published on EDB's official blog, with no independent third-party verification.
- 机制根因
逻辑复制相对二进制流复制的核心差异在于可过滤数据子集、目标端可配置函数与触发器在写入时再加工数据;xDB MMR 让多数据中心并发处理支付并保持站点同步成为可能(需少量应用改造,以及围绕数据冲突的大量测试)。
Logical replication's core edge over binary streaming replication is filtering data subsets and letting the target server run functions/triggers that further manipulate data on write; xDB MMR made concurrent payment processing across multiple data centers with sites kept in sync feasible (requiring minor app changes and heavy testing around data conflicts).
- 教训
对支付网关这种"每秒无数笔交易"的系统,缩短 failover 时间是经济必需而非锦上添花;逻辑复制"可过滤、可加工"的特性让同一套机制同时解决灾备与数据主权合规两件事;多主复制的数据冲突处理必须前置大量测试,不能只看功能有无。
For a payment gateway processing countless card transactions per second, cutting failover time is an economic necessity, not a nice-to-have; logical replication's "filterable and transformable" nature lets one mechanism solve DR and data-sovereignty compliance at once; multi-master conflict handling needs extensive upfront testing, not just a feature checkbox.
相关产品:EDB Postgres / EDB Postgres Advanced Server(EPAS)、PostgreSQL(社区版)、Oracle Database(甲骨文) 相关能力:— 最后核验:2026-10-02
新韩 EZ 财险:Oracle 迁 EDB,半年回本、年运维成本降超 50% 成功经验
去 O
保险
零停机
成本敏感
- 决策
采用 EDB Postgres Advanced Server 实施迁移,EDB 提供持续的全球支持服务保障数据库平稳运行。
Migrated on EDB Postgres Advanced Server, with EDB providing ongoing comprehensive global support to keep database operations running smoothly.
- 结果
EDB 官方口径:零停机完成稳定迁移;6 个月收回投资成本;年度运维成本降低超 50%。均为 EDB 单方口径,无独立第三方验证,技术细节未公开。
Per EDB's own account: stable migration completed with zero downtime; investment costs recouped within six months; annual operating costs reduced by more than 50%. All EDB's one-sided claims, no independent third-party verification, and no technical detail disclosed.
- 机制根因
(公开信息有限)EPAS 的 Oracle 兼容降低迁移改写量,是"零停机+快速回本"的前提;订阅制替代 Oracle 许可是运维成本下降的主因。
(Public detail is thin.) EPAS's Oracle compatibility reduces migration rewrite volume - the precondition for "zero downtime plus fast payback"; subscription pricing replacing Oracle licensing is the main driver of the opex drop.
- 教训
保险这类强监管行业也敢把核心 DBMS 从 Oracle 换到 PG,关键看两点:迁移期零停机、回本周期可量化(6 个月);但本案例公开信息过少,选型时应要求厂商提供同等规模保险客户的可验证细节,不宜只看 headline 数字。
Even heavily regulated insurers will move a core DBMS off Oracle - the bar is zero-downtime migration and a quantifiable payback window (six months); but with so little public detail, a serious evaluation should demand verifiable references from comparable insurance customers rather than headline numbers alone.
相关产品:EDB Postgres / EDB Postgres Advanced Server(EPAS)、Oracle Database(甲骨文)、PostgreSQL(社区版) 相关能力:EPAS 的 Oracle PL/SQL 兼容 —— 去 O 路上改动最小的 PG、Isabel Group 的 TCO 实证 —— 去 O 的"省钱"有人算过账 最后核验:2026-10-02
欧洲流媒体商:AWS Postgres 两次故障后迁到 EDB 全托管云服务 成功经验
云迁移
高可用
流媒体
成本敏感
- 场景
An unnamed European streaming technology provider (ultra-low-latency cloud gaming and betting streams, bound by SLAs) had placed knowledge-base workloads on AWS Postgres; two AWS environment outages hit in a short window. The first was resolved fairly quickly by AWS premium support; the second coincided with a Postgres upgrade "supposed to take 3 minutes" that instead caused 3+ hours of downtime. AWS's troubleshooting prescription was quadrupling the existing cloud infrastructure. SLA payouts ran into the tens of thousands of Euros, plus subscriber and brand losses. (
https://www.Enterprisedb.com/resources/customer-story/streaming-provider-gets-back-in-the-game-with-continuous-uptime-from-edb-biganimal-on-aws)
- 决策
重新选型后迁到全托管的 EDB Postgres AI Cloud Service(原 BigAnimal,跑在 AWS 上),用 EDB Migration Toolkit + Migration Portal 自动化迁移;看中的是 EDB 作为"懂 Postgres 的运营方"的排障能力与白手套支持。
After re-evaluating vendors, migrated to the fully managed EDB Postgres AI Cloud Service (formerly BigAnimal) on AWS, automated with EDB Migration Toolkit and Migration Portal; the draw was EDB as a Postgres-savvy operator with white-glove support.
- 结果
EDB 官方口径:高延迟与反复 outage 被消除;Postgres 成本降低 30%(且无需翻四倍基础设施);DBA 获得底层访问权限与深度洞察。客户匿名,均为 EDB 单方口径,无独立第三方验证。
Per EDB's own account: high latency and repeated outages eliminated; Postgres costs cut 30% without quadrupling infrastructure; DBAs gained underlying infrastructure access and deeper insight into their Postgres environment. Customer unnamed, all EDB's one-sided claims, no independent third-party verification.
- 机制根因
托管 PG 的问题不在 PG 内核而在运营能力——升级窗口失控、排障只能靠"加机器"是云厂商通用支持的典型短板;EDB 这类"懂 PG 的运营方"卖的是故障处理能力,不是实例规格。
The managed-Postgres problem was operational, not in the Postgres kernel - runaway upgrade windows and "just add machines" troubleshooting are the classic shortfalls of generic cloud-vendor support; a Postgres-savvy operator like EDB sells incident-handling capability, not instance sizes.
- 教训
选托管数据库是在选"出事时谁来修";把"升级本应 3 分钟、实则 3 小时"写进选型 checklist,比看功能矩阵更能避坑;匿名案例的数字(−30%)引用时要打折听。
Choosing a managed database is choosing "who fixes it when it breaks"; writing "upgrade meant to take 3 minutes, actually took 3 hours" into the evaluation checklist avoids more pain than any feature matrix; treat numbers from unnamed cases (the 30%) with a discount.
相关产品:EDB Postgres / EDB Postgres Advanced Server(EPAS)、PostgreSQL(社区版) 相关能力:— 最后核验:2026-10-02
美国无线运营商:100TB Oracle Exadata 两周迁到 EDB 成功经验
去 O
电信
百 TB 迁移
成本敏感
- 决策
迁到 EDB Postgres Platform(作为记录系统数据库)+ Cloudera(Hadoop)承接归档与分析;利用 EDB 为 Postgres 开发的分区、异构数据库链路(heterogeneous database links)等 Oracle 迁移增强特性;用 EDB Data Adaptors 让 PG 与 Cloudera 数据双向互通、对 DBA 呈现为 Postgres 表。
Moved to the EDB Postgres Platform as the database of record plus Cloudera (Hadoop) for archival and analytics; used EDB's Oracle-migration enhancements for Postgres such as partitioning and heterogeneous database links; EDB Data Adaptors let data flow both ways between Postgres and Cloudera, appearing to DBAs as Postgres tables.
- 结果
EDB 官方口径:应用运行成本降低"数百万美元";迁移两周完成(由 EDB 主导项目管理以确保快速成功);分析能力相比原方案增强。均为 EDB 单方口径,客户匿名,无独立第三方验证。
Per EDB's own account: running the application got "several million dollars" cheaper; the migration was completed in two weeks with EnterpriseDB managing the project; analytics capabilities improved over the prior solution. All EDB's one-sided claims, customer unnamed, no independent third-party verification.
- 机制根因
分区与异构 DB Link 是 Oracle 重度用户迁移时的关键兼容抓手;EDB Data Adaptors 让结构化(PG)与非结构化(Hadoop)数据对 DBA 呈现为"一张表",避免应用大改;"记录系统"与"分析归档"拆分到 PG+Hadoop 各归其位,比单一 RDBMS 扛所有负载更便宜。
Partitioning and heterogeneous DB links are the key compatibility handholds for heavy Oracle users migrating out; EDB Data Adaptors present structured (Postgres) and unstructured (Hadoop) data to DBAs as "one table," avoiding application rewrites; splitting "system of record" and "analytics archive" onto Postgres + Hadoop respectively is cheaper than one RDBMS carrying every workload.
- 教训
100TB 级 Oracle Exadata 迁出可以是"两周级"工程,但前提是厂商级迁移方法论与工具链全程护航;迁移不只是换库,更是把不同访问模式拆到各自合适的架构上。
A 100TB Oracle Exadata exit can be a two-week engineering effort - but only with vendor-grade migration methodology and tooling shepherding it end to end; migration is not just swapping the database but placing each access pattern on the architecture that fits it.
相关产品:EDB Postgres / EDB Postgres Advanced Server(EPAS)、Oracle Database(甲骨文)、PostgreSQL(社区版) 相关能力:EPAS 的 Oracle PL/SQL 兼容 —— 去 O 路上改动最小的 PG 最后核验:2026-10-02
Jepsen 独立测试 etcd 3.4.3:KV 严格可串行化全过,分布式锁却"从根本上不安全"(2020) 失败教训
选型评估
独立测试
分布式锁
线性一致性
CNCF
- 场景
etcd 是 Kubernetes 等系统的元数据存储与协调原语,常被同时用作 KV 存储和分布式锁。2020 年 1 月 Jepsen 测试 etcd 3.4.3(5 节点 Debian 集群),验证 KV 操作、分布式锁与 watch 的安全属性。本次由 CNCF(Linux 基金会下属中立基金会,etcd 托管方)资助,非厂商付费。
etcd is the metadata store and coordination primitive behind Kubernetes and similar systems, often used simultaneously as a KV store and a distributed lock service. In January 2020, Jepsen tested etcd 3.4.3 (5-node Debian cluster), verifying the safety properties of KV operations, distributed locks, and watches. This work was funded by the CNCF (the neutral Linux Foundation subsidiary that hosts etcd) — not by a vendor.
- 决策
用 Knossos 线性一致性检查器验证 register/set/append 多键事务;故障注入包括网络分区(单节点隔离、多数/少数分裂、非传递分区)、进程暂停/崩溃、针对性杀主、数百秒时钟偏移、动态成员变更;另设 lock 与 watch 专项工作负载。
Used the Knossos linearizability checker to verify register/set/append multi-key transactions; fault injection included network partitions (single-node isolation, majority/minority splits, non-transitive partitions), process pauses/crashes, targeted leader kills, clock skew up to hundreds of seconds, and dynamic membership changes; plus dedicated lock and watch workloads.
- 结果
KV 操作默认严格可串行化(strict serializable),在全部故障下均成立——Jepsen 罕见的正面结论;watch 按序交付每个变更。但 etcd lock 从根本上不安全:健康集群、无故障时多客户端可同时持有同一把锁;"等待锁后未重查租约有效性"的实现 bug 加剧了风险;用 etcd 锁保护内存集合更新时,2 秒租约下可稳定复现约 18% 的确认更新丢失(Jepsen 报告原文口径)。etcd 官方发布 companion blog 回应,锁租约检查 bug 在 master 修复,并修订了 API guarantees 文档(删去"sequential 是分布式系统最强一致性"等错误表述)。
KV operations are strict serializable by default and held up under every fault — a rare unqualified positive from Jepsen; watches delivered every change in order. But etcd locks are fundamentally unsafe: multiple clients could hold the same lock simultaneously even in healthy clusters with no faults; an implementation bug (failing to re-check lease validity after waiting for a lock) worsened the risk; using etcd locks to guard updates to an in-memory set reliably reproduced ~18% loss of acknowledged updates with 2-second lease TTLs (per the Jepsen report). The etcd team published a companion blog post in response, fixed the lease-check bug on master, and revised the API guarantees documentation (removing incorrect claims like "sequential is the strongest consistency guarantee in distributed systems").
- 机制根因
KV 与锁的安全属性来自完全不同的机制:KV 走 Raft 全序状态机,天然严格可串行化;锁则依赖租约(lease)+ 客户端心跳,在异步系统里"故障检测"本质不可靠——按 Kleppmann 的论述,锁服务为保活必须在持有者"疑似"死亡时强制释放锁,而疑似死亡不等于真死,互斥性就破了。这是理论天花板,不是修个 bug 能解决的;实现 bug 只是让破窗来得更快。
The KV store's and the lock service's safety properties come from entirely different mechanisms: KV rides a Raft totally-ordered state machine, hence strict serializability for free; locks depend on leases plus client heartbeats, and failure detection is fundamentally unreliable in asynchronous systems — per Kleppmann's argument, a lock service must forcibly release a lock when its holder is *suspected* dead to preserve liveness, but suspected-dead ≠ actually-dead, and mutual exclusion breaks. That is a theoretical ceiling, not something a bug fix can resolve; the implementation bug only broke the window sooner.
- 教训
选型时把"KV 存储"和"分布式锁"当两个产品评估——etcd 前者满分、后者不及格;需要互斥必须配合 fencing token 之类机制;任何分布式锁宣称都要先问"持有者假死时怎么办"。诚实注记:3.4.3 是 2020 年版本,etcd 已演进到 3.5/3.6,但"异步系统里分布式锁的根本不安全性"是理论结论,不随版本过期;租约检查 bug 的修复状态请按当前版本核验。
Evaluate "KV store" and "distributed lock" as two separate products during selection — etcd scores full marks on the former and fails the latter; pair mutual exclusion needs with fencing tokens or similar mechanisms; greet every distributed-lock claim by asking "what happens when the holder is falsely suspected dead." Honesty note: 3.4.3 is a 2020 release and etcd has since moved to 3.5/3.6, but "distributed locks are fundamentally unsafe in asynchronous systems" is a theoretical conclusion that does not expire with versions; re-verify the lease-check fix against the current release.
来源
Jepsen《Jepsen: etcd 3.4.3》(2020-01-30,CNCF 资助
Jepsen, "Jepsen: etcd 3.4.3" (2020-01-30, CNCF-funded
etcd 官方 companion blog 回应(见该报告内链)
etcd's official companion blog post in response (linked from the report)
相关产品:etcd 相关能力:Raft KV 的严格可串行化与分布式锁的能力边界 最后核验:2026-10-02
K3s v1.19.1:放弃实验性 dqlite,嵌入式 etcd 成为官方 HA 数据存储(2020) 成功经验
高可用
边缘计算
架构选型
云原生
- 场景
K3s 是 Rancher 面向边缘与资源受限环境的轻量 Kubernetes 发行版,单二进制、默认内嵌 SQLite。SQLite 只支持单 server;多 server 高可用此前只能外接 MySQL/PostgreSQL/外部 etcd,或使用实验性的 dqlite(分布式 SQLite)。边缘场景(零售门店、工厂车间)往往没有专职运维,外接数据库的运维成本不可接受,而 dqlite 生态几乎为零。
K3s is Rancher's lightweight Kubernetes distribution for edge and resource-constrained environments: a single binary with embedded SQLite by default. SQLite only supports a single server; multi-server HA previously required an external MySQL/PostgreSQL/etcd, or the experimental dqlite (distributed SQLite). Edge locations (retail stores, factory floors) rarely have dedicated DBAs or ops staff, so the operational cost of an external database was unacceptable, while dqlite had essentially zero ecosystem.
- 决策
K3s v1.19.1(2020 年 9 月)把实验性 dqlite 替换为嵌入式 etcd 作为集群化数据存储方案:首个 server 用 `--cluster-init` 初始化,后续 server 加入形成 Raft 多数派(内嵌组件版本为 etcd v3.4.13-k3s1)。官方 release notes 原话:"we've made this change in order to leverage the existing effort and knowledge that has gone into operating Kubernetes with etcd"——直接复用社区已有的 etcd 运维知识,而非另起一套;附带收益是顺手获得 etcd 快照/恢复能力。对 dqlite 存量集群明确为 breaking change,不提供升级路径。
K3s v1.19.1 (September 2020) replaced experimental dqlite with embedded etcd as the clustered datastore: the first server bootstraps with `--cluster-init`, additional servers join to form a Raft majority (embedded component version etcd v3.4.13-k3s1). The official release notes put it plainly: "we've made this change in order to leverage the existing effort and knowledge that has gone into operating Kubernetes with etcd" - reusing the community's existing etcd operational knowledge instead of building a new one, with etcd snapshot/restore support as a side benefit. For existing dqlite clusters it was explicitly a breaking change with no upgrade path.
- 结果
嵌入式 etcd 在 v1.19.5+k3s1 转为正式支持,成为 K3s 官方文档推荐的生产级 HA 路径("Embedded etcd (Raft consensus): Suitable for production. Requires 3 or 5 server nodes");dqlite 实验终止。此后 K3s 的 HA 数据面与上游 Kubernetes 同构,etcd 排障经验(调优、备份、恢复)在两个生态间可直接复用。
Embedded etcd became fully supported in v1.19.5+k3s1 and is now the production-grade HA path in the official K3s docs ("Embedded etcd (Raft consensus): Suitable for production. Requires 3 or 5 server nodes"); the dqlite experiment was terminated. K3s's HA data plane is now isomorphic with upstream Kubernetes, so etcd operational knowledge (tuning, backup, recovery) transfers directly between the two ecosystems.
- 机制根因
选型逻辑是"生态复用 > 技术尝鲜"。etcd 已是 Kubernetes 控制面的事实标准存储,运维知识、工具链(etcdctl、快照)、人才池都是现成的;dqlite 虽轻但生态为零,K3s 团队要自己承担全部运维知识的生产成本。嵌入式 etcd 让 K3s 的 HA 数据面与上游同构,问题可互相借鉴——这是用"标准件"替代"自研件"的胜利。
The selection logic was "ecosystem reuse beats technical novelty." etcd was already the de facto standard store for the Kubernetes control plane, with ready-made operational knowledge, tooling (etcdctl, snapshots), and talent pool; dqlite was lighter but had no ecosystem, leaving the K3s team to bear the full cost of producing operational knowledge themselves. Embedded etcd made K3s's HA data plane isomorphic with upstream - a win for "standard parts" over "homegrown parts."
- 教训
基础设施选型时,"社区已验证的运维知识存量"是硬资产;为省体积引入生态孤岛反而抬高长期 TCO;与上游同构的数据面让排障经验可复用,边缘场景更应如此——现场没人能帮你调一个没人用过的分布式存储。
In infrastructure selection, "the stock of community-validated operational knowledge" is a hard asset; adopting an ecosystem island to save binary size raises long-term TCO instead. A data plane isomorphic with upstream makes troubleshooting experience reusable - especially at the edge, where nobody on site can debug a distributed store nobody else runs.
相关产品:etcd 相关能力:Kubernetes 控制面的"唯一真相源":Raft 强一致 + MVCC watch 最后核验:2026-10-02
kOps 生产集群 etcd 配额耗尽:60 节点集群控制面冻结实录(2025) 失败教训
运维事故
容灾
配额
资源泄漏
- 场景
某 kOps 管理的生产 Kubernetes 集群(60+ 节点、2000+ pod)。值班期间 Slack 先报 cert-manager 宕机,随后 kubectl 无响应、ArgoCD 与内部工具相继变黑;业务流量尚正常,但控制面已瘫痪。master 节点在 LB 健康检查中全部 unhealthy,重启 master 无效。
A kOps-managed production Kubernetes cluster (60+ nodes, 2,000+ pods). During on-call, Slack first reported cert-manager down, then kubectl went unresponsive, followed by ArgoCD and internal tooling going dark; user traffic was still being served, but the control plane was down. All master nodes showed unhealthy in load-balancer health checks, and rebooting masters did not help.
- 决策
排查发现 kube-apiserver 被 OOMKilled,先把 master 机型从 c5.4xlarge 升到 c5.9xlarge(kOps 命令行),但集群仍不健康——事前没有任何内存压力信号,内存并非真因。继续深挖控制面,在 etcd 日志里看到 "etcdserver: mvcc: database space exceeded":etcd 逻辑配额(2.2GB)被打满,触发 NOSPACE alarm 后全集群拒绝一切写。
Investigation found kube-apiserver being OOMKilled, so masters were first upgraded from c5.4xlarge to c5.9xlarge via the kOps CLI - but the cluster stayed unhealthy, and there had been no memory-pressure signals beforehand, so memory was not the true cause. Digging deeper into the control plane revealed "etcdserver: mvcc: database space exceeded" in etcd logs: the etcd logical quota (2.2GB) was full, and the NOSPACE alarm made the whole cluster reject all writes.
- 结果
按 etcd 官方维护指南做 compact + defrag,从 2.2GB 配额中释放约 800MB,集群自行恢复;次日重启了 coredns、Prometheus 等与 apiserver 失联的 pod。根因是 logging-operator 命名空间里数千个无用的 Secret(`default-logging-fluentd-configcheck-app-*`)把 revision 历史撑爆——膨胀的是 revision 历史,不是有效数据。
Following the official etcd maintenance guide, a compact + defrag freed about 800MB out of the 2.2GB quota and the cluster healed itself; the next day, pods that had lost contact with the API server (CoreDNS, Prometheus) were restarted. The root cause was thousands of useless Secrets (`default-logging-fluentd-configcheck-app-*`) in the logging-operator namespace bloating the revision history - the bloat was revision history, not live data.
- 机制根因
etcd 默认永久保留 MVCC revision 历史;高 churn(此处是 operator 失控批量创建 Secret)让 bbolt 文件膨胀到 quota 上限;NOSPACE 是集群级写冻结——kubectl apply/create 全拒,但读不受影响。compact 只标记旧 revision 可回收、defrag 才真正归还空间,两步缺一不可;且这是逻辑配额耗尽,不是物理磁盘满,加磁盘没用。
etcd retains MVCC revision history forever by default; high churn (here, an operator uncontrollably creating Secrets in bulk) inflated the bbolt file to the quota limit; NOSPACE is a cluster-wide write freeze - kubectl apply/create all rejected, while reads kept working. Compact only marks old revisions reclaimable, defrag actually returns the space - both steps are required. And this was logical quota exhaustion, not a full physical disk, so adding disk space would not have helped.
- 教训
给 Secret/ConfigMap 等临时资源设 retention/清理策略,operator 失控创建资源是真实风险;监控 `etcd_mvcc_db_total_size_in_bytes` 相对 quota 的水位,而非只看物理磁盘;OOM 式的表面症状会误导排查方向,控制面失联时优先查 etcd 配额与 leader 状态;kOps 这类"半托管"方案同样需要 etcd 运维手册,不能假设有人替你管。
Set retention/cleanup policies for ephemeral resources like Secrets and ConfigMaps - a runaway operator creating resources is a real risk. Monitor `etcd_mvcc_db_total_size_in_bytes` against the quota watermark, not just physical disk usage. OOM-like surface symptoms mislead triage; when the control plane goes dark, check etcd quota and leader status first. Semi-managed solutions like kOps still need an etcd operations handbook - don't assume someone else is managing it for you.
相关产品:etcd 相关能力:MVCC 修订历史无上限:"database space exceeded"让全集群写冻结 最后核验:2026-10-02
Kubernetes 1.6 宣布支持 5000 节点:"由 CoreOS 的新版 etcd v3 驱动"(2017) 成功经验
超大规模
扩展
架构选型
云原生
- 场景
2017 年初,Kubernetes 社区要把可扩展性 SLO 从 2000 节点推到 5000 节点(15 万 pod),支撑搜索、游戏等超大规模负载。1.5 时代 apiserver 后端还是 etcd v2(HTTP+JSON API、无 MVCC、无高效 watch),成为扩展瓶颈。
In early 2017, the Kubernetes community wanted to push its scalability SLO from 2,000 to 5,000 nodes (150,000 pods) for hyperscale workloads like search and gaming. In the 1.5 era the API server backend was still etcd v2 (HTTP+JSON API, no MVCC, no efficient watch) - the scaling bottleneck.
- 决策
Kubernetes 1.6 把 apiserver 默认存储后端升级到 etcd v3(CoreOS 新版),官方发布公告原话:"This 150% increase in total cluster size, powered by a new version of etcd v3 by CoreOS"。从 1.5 升级的集群需规划数据迁移窗口——存储后端的代际切换被当作 release headline 而非脚注。
Kubernetes 1.6 upgraded the API server's default storage backend to etcd v3 (the new CoreOS release), with the official announcement stating verbatim: "This 150% increase in total cluster size, powered by a new version of etcd v3 by CoreOS." Clusters upgrading from 1.5 had to plan a data-migration window - a storage-backend generational switch treated as a release headline, not a footnote.
- 结果
1.6 官宣支持 5000 节点/15 万 pod 集群,成为当年扩展性里程碑的公开背书;此后 etcd 一直是 Kubernetes 控制面的唯一真相源,v3 的 MVCC/watch 语义支撑了 controller list-watch 模式与后续全部扩展性工作。
1.6 officially supported 5,000-node / 150,000-pod clusters, the public commitment behind that year's scalability milestone; etcd has remained the single source of truth for the Kubernetes control plane ever since, with v3's MVCC/watch semantics underpinning the controller list-watch pattern and all subsequent scalability work.
- 机制根因
etcd v3 的 MVCC 让 watch 可以带 revision 断点续传、range 查询可高效批量,apiserver 的 watch 缓存与 controller resync 模式才得以成立;gRPC 二进制多路复用替代 v2 的 HTTP/JSON,大幅降低 apiserver↔etcd 的协议开销。这是"存储语义升级解锁上层架构"的典型——瓶颈不在 etcd 本身的吞吐,而在 v2 的 API 语义表达力不够。
etcd v3's MVCC made watches resumable from a revision and range queries efficiently batchable, which is what made the API server's watch cache and the controller resync pattern possible; gRPC's binary multiplexing replaced v2's HTTP/JSON and sharply cut API-server-to-etcd protocol overhead. A classic case of "a storage-semantics upgrade unlocking the upper architecture" - the bottleneck was never etcd's raw throughput, but v2's API semantics not being expressive enough.
- 教训
控制面存储的语义(MVCC/watch)决定整个编排系统的扩展天花板;基础设施的代际升级值得为"语义"而非只为"性能数字"买单;官方把第三方组件写进发布公告头条,说明 etcd 已是 Kubernetes 扩展性故事的一等公民——选型时看它被谁依赖,比看 benchmark 更说明问题。
The semantics of the control-plane store (MVCC/watch) set the scaling ceiling of the entire orchestration system; a generational infrastructure upgrade is worth buying for "semantics," not just performance numbers. When official release notes headline a third-party component, that component is a first-class citizen of the scalability story - who depends on it says more than any benchmark.
相关产品:etcd 相关能力:Kubernetes 控制面的"唯一真相源":Raft 强一致 + MVCC watch 最后核验:2026-10-02
CoreOS etcd-operator 归档:在 Kubernetes 里再套一层 operator 管 etcd 走不通(2020) 失败教训
运维
云原生
架构选型
控制器模式
- 场景
CoreOS 2016 年推出 etcd-operator——最早的主流 operator 范式项目之一,用 CRD(EtcdCluster)在 Kubernetes 内自动管理 etcd 集群的创建/扩缩容/故障转移/滚动升级/备份恢复,一度被视为"operator 模式"的样板工程,star 达 1.8k。
In 2016 CoreOS launched etcd-operator - one of the earliest mainstream operator-pattern projects - using a CRD (EtcdCluster) to automatically manage etcd cluster creation, resizing, failover, rolling upgrades, and backup/restore inside Kubernetes. It was long held up as the model example of the operator pattern, reaching 1.8k stars.
- 决策
2020 年 3 月,coreos/etcd-operator 仓库被官方归档为只读,README 声明:"This project is no longer actively developed or maintained. The project exists here for historical reference."(本项目不再积极开发与维护,仅作历史参考。)其 helm chart 随后废弃,依赖它的下游部署模式(如 CoreDNS 外部 DNS 方案中依赖 etcd-operator chart 的那一种)被迫改道。
In March 2020, the coreos/etcd-operator repository was archived read-only, with the README stating: "This project is no longer actively developed or maintained. The project exists here for historical reference." Its Helm chart was deprecated afterward, forcing downstream deployment patterns that depended on it (such as the CoreDNS external-DNS setup that required the etcd-operator chart pre-installed) to change course.
- 结果
社区共识收敛为:etcd 跑静态 Pod(kubeadm/kOps 部署)或直接用托管控制面,而非"在 Kubernetes 里用 operator 套娃管 etcd";etcd 社区另起 etcd-operator Working Group 重新思考可用性问题,而非复活该项目。
Community consensus converged on running etcd as static Pods (kubeadm/kOps deployments) or simply using a managed control plane - not "nesting an operator inside Kubernetes to manage etcd." The etcd community started a separate etcd-operator Working Group to rethink usability instead of reviving the project.
- 机制根因
etcd 是"先有鸡"的组件——operator 自身跑在 Kubernetes 上,而 Kubernetes 的控制面又依赖 etcd;用 k8s 原语去运维 k8s 的根存储,形成循环依赖,故障域纠缠(控制面抖动时 operator 自身也可能失联)。Raft 成员变更的正确性对时机与顺序极敏感,operator 的"声明式调谐"与 Raft 的"过程式成员协议"语义错配,边界情况极易写错——而这里的"写错"直接等价于丢数据。
etcd is a "chicken first" component - the operator itself runs on Kubernetes, while Kubernetes's control plane depends on etcd; managing Kubernetes's root store with Kubernetes primitives creates a circular dependency with entangled failure domains (when the control plane jitters, the operator itself can lose connectivity). Raft membership changes are extremely sensitive to timing and ordering, and the operator's "declarative reconciliation" semantically mismatches Raft's "procedural membership protocol" - edge cases are easy to get wrong, and getting them wrong here means losing data directly.
- 教训
不要在被管理系统的上层再套一层同构的自动化去管它的根依赖;operator 模式适合无状态或可重建的工作负载,对"一致性状态机"要极度谨慎;选型时看项目的维护状态——归档本身就是信号,1.8k star 也救不了一个语义错配的设计。
Don't stack another layer of isomorphic automation on top of a managed system to manage its own root dependency; the operator pattern fits stateless or rebuildable workloads, but demands extreme caution for "consistency state machines." When selecting, watch a project's maintenance status - archival is itself a signal, and 1.8k stars can't save a semantically mismatched design.
相关产品:etcd 相关能力:Kubernetes 控制面的"唯一真相源":Raft 强一致 + MVCC watch 最后核验:2026-10-02
etcd 引入 Antithesis 自治测试:830 小时模拟 4.5 年,挖出影响全稳定版的关键 watch bug(2025) 成功经验
数据一致性
正确性测试
运维
云原生
- 场景
etcd 是 Kubernetes 的主数据存储,一致性是第一优先级。v3.5 发布后暴露了若干正确性问题,维护者为此开发了基于属性的 robustness 测试框架(多类型流量回放 + 随机故障注入 + 线性一致性检查 + watch 语义校验)。但原框架"像蒙眼扔飞镖"——bug 靠运气发现且难以复现。
etcd is Kubernetes's primary datastore, so consistency is the top priority. After the v3.5 release surfaced several correctness issues, the maintainers built a property-based robustness testing framework (varied traffic replay + random fault injection + linearizability checks + watch-semantics verification). But the original framework was "like throwing darts while blindfolded" - bugs were found by luck and were hard to reproduce.
- 决策
etcd 团队把既有 robustness 测试搬上 Antithesis 确定性仿真平台做"自治测试":整个 etcd 集群跑在确定性 hypervisor 里,平台完全控制网络行为、线程调度、系统时钟等一切非确定性来源;改用声明式属性断言("数据一致性永不被违反""watch 事件永不丢失")作为要主动打破的目标,自动搜索能违反属性的故障序列。测试覆盖 3 节点与单节点集群,故障类型包括网络延迟/拥塞/分区、线程暂停、进程 kill、时钟抖动、CPU 限流;同时用含已知 bug 的旧版本验证方法有效性,再测 3.4/3.5/3.6 稳定版与 main 分支。
The etcd team ported its robustness tests onto the Antithesis deterministic simulation platform for "autonomous testing": the whole etcd cluster runs inside a deterministic hypervisor where the platform controls every source of non-determinism (network behavior, thread scheduling, system clocks). Declarative property assertions ("data consistency is never violated", "a watch event is never dropped") became targets to actively break, with automated search for fault sequences that violate them. Tests covered 3-node and single-node clusters with network latency/congestion/partitions, thread pauses, process kills, clock jitter, and CPU throttling; older releases with known bugs validated the methodology before testing the 3.4/3.5/3.6 stable releases and the main branch.
- 结果
830 个 wall-clock 小时模拟出约 4.5 年的使用量;找回全部已知 bug("棕色 M&M"回归集),并在 main 分支发现若干新问题——其中一个"存在于所有稳定版本"的关键 watch bug 是此前测试从未发现的;还暴露了团队自研线性一致性检查器模型自身的缺陷。v3.6 发布公告同步披露了该流程中发现并修复的三个数据不一致 bug:负载下崩溃的数据不一致(v3.5.0 引入,v3.5.3 修复,issue/13766)、单节点 durability API 保证被破坏(历史遗留,v3.4.21/v3.5.5 修复,issue/14370)、defrag 期间崩溃的 revision 不一致(v3.5.0 引入,v3.5.6 修复)。以上均为官方口径。
830 wall-clock hours simulated roughly 4.5 years of usage; all known bugs were re-found (the "Brown M&M" regression set), and several new issues were discovered on the main branch - including a critical watch bug present in all stable releases that previous testing had never caught; the effort also exposed a flaw in the team's own linearizability checker model. The v3.6 announcement disclosed three data-inconsistency bugs found and fixed through this process: data inconsistency when crashing under load (introduced in v3.5.0, fixed in v3.5.3, issue/13766), broken durability API guarantee in single-node clusters (a legacy issue, fixed in v3.4.21/v3.5.5, issue/14370), and revision inconsistency when crashing during defragmentation (introduced in v3.5.0, fixed in v3.5.6). All per official sources.
- 机制根因
分布式一致性 bug 往往需要"故障的精确组合序列"才能触发,传统随机故障注入命中概率极低;确定性仿真把"发现 bug"变成可复现、可系统搜索的优化问题。属性式断言(不变式)比场景式断言更能覆盖未知的故障交互——你要找的不是"某个场景",而是"任何违反不变式的路径"。
Distributed consistency bugs usually need a precise combination sequence of faults to trigger, which random fault injection hits with near-zero probability; deterministic simulation turns "finding bugs" into a reproducible, systematically searchable optimization problem. Property-based assertions (invariants) cover unknown fault interactions far better than scenario-based ones - you search for "any path violating the invariant," not "a specific scenario."
- 教训
对一致性攸关的基础设施,测试预算应投向"可复现的系统性搜索"而非更多的随机用例;确定性仿真让每一次 bug 发现都可精确重放,这是传统混沌工程给不了的;把已知 bug 做成回归集持续验证测试方法本身有效——先证明你的测试能抓到已知的 bug,再相信它能抓到未知的。
For consistency-critical infrastructure, spend the testing budget on "reproducible systematic search" rather than more random cases; deterministic simulation makes every bug discovery exactly replayable, which traditional chaos engineering cannot offer. Keep known bugs as a regression set that continuously validates the testing methodology itself - first prove your tests catch the bugs you know, then trust them to catch the ones you don't.
相关产品:etcd 相关能力:Kubernetes 控制面的"唯一真相源":Raft 强一致 + MVCC watch 最后核验:2026-10-02
一块跑了 7 年的磁盘拖垮控制面:etcd WAL fsync p99 达 500ms 的复盘(2025) 失败教训
磁盘 IO
延迟
运维
硬件故障
- 场景
某裸金属自建 Kubernetes 集群触发 kube-prometheus-stack 的 KubeAPIErrorBudgetBurn 告警,可用性掉到 90.9%,错误预算快速耗尽。初步排查 CPU/内存/PID 压力/网络/kubelet 全部正常;etcd `endpoint health` 返回 healthy(15.7ms)——极具迷惑性。
A bare-metal, self-managed Kubernetes cluster fired the kube-prometheus-stack `KubeAPIErrorBudgetBurn` alert; availability dropped to 90.9% and the error budget was burning fast. Initial checks of CPU, memory, PID pressure, network, and kubelet were all clean; etcd `endpoint health` returned healthy (15.7ms) - deeply misleading.
- 决策
作者绕开 health 检查,直看 etcd 日志与 Prometheus 指标:etcd 大量 "apply request took too long"(took 409ms,期望 100ms),连只读 range 都要等 Raft 达成一致(400ms);WAL fsync p99 在 300–500ms 之间(健康线应 <10ms);CoreDNS 本地健康检查 >1s、metrics-server 报 Handler timeout——均为 apiserver 等 etcd 的继发症状。`df` 确认 etcd 与 Longhorn 等共享 /dev/sda;`smartctl -t long` 发现 Command_Timeout=85(>0 即硬件故障信号)、Power_On_Hours=62247(7.1 年连续运行,寿命剩余 9%)。
The author bypassed the health check and went straight to etcd logs and Prometheus metrics: a flood of "apply request took too long" (took 409ms vs. 100ms expected), with even read-only range requests waiting on Raft agreement (400ms); WAL fsync p99 sat between 300-500ms (healthy is under 10ms); CoreDNS local health checks exceeded 1s and metrics-server threw handler timeouts - all secondary symptoms of the API server waiting on etcd. `df` showed etcd sharing /dev/sda with Longhorn and others; a `smartctl -t long` test found Command_Timeout=85 (any value above zero signals hardware failure) and Power_On_Hours=62247 (7.1 years of continuous operation, 9% lifetime remaining).
- 结果
更换故障磁盘后 fsync p99 回到 10ms 以下,错误预算燃烧停止。事后补了三条基线:对 `etcd_disk_wal_fsync_duration_seconds` p99 告警(50ms 告警、100ms paging,300ms 即事故);监控 etcd 节点磁盘 IO 饱和度(与 Longhorn 等 IO 大户混部时 IO 利用率持续 >50% 即告警);用 smartctl_exporter 把 Command_Timeout 做成指标,>0 立刻告警。作者同时指出:把 etcd 迁到独立磁盘可滚动完成,无需停机。
Replacing the failing drive brought fsync p99 back under 10ms and stopped the budget burn. Three new baselines followed: alert on `etcd_disk_wal_fsync_duration_seconds` p99 (warn at 50ms, page at 100ms, 300ms means incident); monitor disk IO saturation on etcd nodes (alert when sustained above 50% with IO-heavy neighbors like Longhorn); export Command_Timeout via smartctl_exporter and alert the moment it goes non-zero. The author also noted etcd can be moved to a dedicated disk as a rolling change with no downtime.
- 机制根因
etcd 写路径 fsync-bound,leader 的存活绑定在"能否及时 fsync"上;磁盘物理命令超时 → fsync 停滞 → Raft 心跳 missed → 短暂丢 leader → apiserver 请求超时 → 全控制面抖动。`endpoint health` 只反映"此刻能否响应",对渐进式磁盘退化完全失明——health 为绿而 p99 已 500ms 是本案最贵的教训。
etcd's write path is fsync-bound, and a leader's survival is tied to "can it fsync in time"; physical disk command timeouts stalled fsyncs, Raft heartbeats were missed, the leader was briefly lost, API server requests timed out, and the whole control plane jittered. `endpoint health` only reflects "can it respond right now" and is completely blind to gradual disk degradation - green health with 500ms p99 was the most expensive lesson of this incident.
- 教训
别把 etcd endpoint health 当磁盘健康信号;WAL fsync p99 才是 etcd 磁盘的真实体温计(<10ms 健康、50ms 告警、100ms paging);etcd 必须独占磁盘,勿与 Longhorn 等 IO 大户混部;裸金属集群应把 SMART 指标纳入监控——"Kubernetes 把硬件抽象得太好,好到让人忘了硬件存在"。
Never use etcd endpoint health as a disk-health signal; WAL fsync p99 is etcd's true disk thermometer (under 10ms healthy, 50ms warn, 100ms page). Give etcd a dedicated disk - never co-locate it with IO-heavy workloads like Longhorn. Bare-metal clusters should bring SMART metrics into monitoring: "Kubernetes abstracts hardware so completely that it's easy to forget hardware exists."
相关产品:etcd 相关能力:写吞吐天花板 + fsync 偏执:"慢盘 = 丢 leader = 控制面抖动" 最后核验:2026-10-02
Figma 单 Postgres → 垂直分区 → 逻辑分片(2020–2024):扩展是阶梯不是跳跃 成功经验
渐进式扩展
垂直分区
零停机
小团队
- 决策
不换库,分三阶段推进:①垂直分区,按领域(files 库、users 库……)拆成十几个库,快速买到 runway;②量化每种瓶颈(CPU/IO/表大小/写入行数),预测每个分片的 runway;③自研 colos:Postgres 视图做逻辑分片 + Go 查询代理,9 个月、零停机、每步可回滚。
Stayed on Postgres in three phases: (1) vertical partitioning — split by domain (files DB, users DB, …) into a dozen-plus databases to buy runway fast; (2) quantified every bottleneck type (CPU/IO/table size/write rows) and forecast each shard's runway; (3) built colos in-house: logical sharding via Postgres views plus a Go query proxy, nine months, zero downtime, every step rollback-safe.
- 结果
2020 年的 AWS 最大单机 PG → 2022 年底十几个垂直分区 → 逻辑分片;100 倍增长下未发生全局事故。VACUUM 曾引发可靠性事件,倒逼出分片。
From the largest single-node PG on AWS in 2020 to a dozen-plus vertical partitions by end of 2022, then logical sharding; no global incidents across 100x growth. A VACUUM-induced reliability incident forced the sharding move.
- 机制根因
垂直分区是"按领域拆",复杂度远低于水平分片,先拿 2–3 年 runway;视图 + 代理实现"逻辑先行、物理随后",应用层无感;先量化瓶颈再动手,避免过早分片。
Vertical partitioning is "split by domain," far simpler than horizontal sharding, buying 2–3 years of runway first; views plus proxy achieve "logical first, physical later," invisible to the application; quantifying bottlenecks before acting avoids premature sharding.
- 教训
PG 规模化的真实天花板常常不是吞吐,而是 VACUUM/长事务这类运维债——要提前监控、提前建模;扩展路线每一步都应先有可测量的瓶颈模型(runway),按需升级阶梯,而不是一步跳到终局架构。
The real ceiling on scaling Postgres is often not throughput but operational debt like VACUUM and long transactions — monitor and model them early; each step of the scaling ladder should be driven by a measurable bottleneck model (runway), upgrading the ladder as needed rather than jumping straight to the end-state architecture.
相关产品:PostgreSQL(社区版) 相关能力:进程模型连接天花板 + VACUUM 运维税 最后核验:2026-10-01
FriendFeed 在 MySQL 上实现无模式存储(2009) 成功经验
无模式
索引演进
在线变更
- 场景
2009 年 FriendFeed 高速迭代,最大的敌人是 schema 变更:2.5 亿行的大表加索引、改列,ALTER 锁表动辄数小时到数天,新功能上线节奏被数据库拖死。(Bret Taylor 2009 年博客,经存档;AWS James Hamilton 评论)
In 2009 FriendFeed was iterating fast, and its biggest enemy was schema change: on 250-million-row tables, adding an index or altering a column meant ALTER TABLE locking the table for hours to days — the database was throttling feature velocity. (Bret Taylor's 2009 blog, via archive; commentary by AWS's James Hamilton)
- 决策
不把 schema 暴露给 MySQL:实体存成一张分片表的 MEDIUMBLOB(zlib 压缩的 Python dict);索引拆成独立的"两列小表"(属性值→实体 ID),由离线索引构建进程持续维护;查询时先查索引表拿 ID,再回原表取实体,应用层做残余过滤。
Don't show the schema to MySQL at all: entities went into a sharded table as MEDIUMBLOBs (zlib-compressed Python dicts); indexes were split into standalone two-column tables (attribute value → entity ID), maintained continuously by an offline index-building process; queries first hit the index table for IDs, then fetched entities from the primary table, with residual filtering in the application layer.
- 结果
据 Bret Taylor 原文,加新属性、建新索引从"以周计"缩短到"以天计",且不再需要主从切换等高危运维操作;该模式后来被 Uber Schemaless 等发扬光大。
Per Bret Taylor's post, adding new attributes and building new indexes went from "weeks" to "days," with no more scary operational work like master/slave swaps; the pattern was later carried forward by systems like Uber's Schemaless.
- 机制根因
关系 schema 的真正成本不在"建模"而在"演进"——在 2009 年的 MySQL(5.0/5.1 时代),大表 ALTER 往往意味着锁表重写物理存储,表越大越贵。注意这是历史条件:MySQL 8.x 的 online DDL 已支持 INSTANT/INPLACE 等低代价操作,部分变更只改元数据并允许并发 DML;本案例的"ALTER=重写"论断不可直接套用于现代 MySQL。FriendFeed 把"存储"和"索引"解耦:存储层只保证按主键取实体(MySQL 最稳的能力),索引变成可随意创建/删除的派生数据(最终一致,由后台 Cleaner 进程修复)。代价是放弃数据库端 JOIN 与二级索引一致性(索引与实体非原子更新,应用层必须容忍短暂不一致并做残余过滤),以及把查询优化器的活儿搬到应用层自己干。
The real cost of a relational schema isn't "modeling" — it's "evolution": ALTER fundamentally rewrites physical storage, and the bigger the table, the more expensive — under the MySQL 5.0/5.1 of 2009, where big-table ALTERs typically meant locking table rebuilds. Note this is a historical condition: MySQL 8.x online DDL supports low-cost INSTANT/INPLACE operations, with some changes touching only metadata while concurrent DML continues; the "ALTER = rewrite" claim must not be applied to modern MySQL as-is. FriendFeed decoupled "storage" from "indexing": the storage layer only guaranteed primary-key entity fetch (MySQL's most solid capability), while indexes became disposable derived data (eventually consistent, repaired by a background Cleaner process). The price: giving up database-side JOINs and secondary-index consistency (index and entity updates aren't atomic, so the app must tolerate brief inconsistency and do residual filtering), plus moving the query optimizer's job into application code.
- 教训
迭代速度被 ALTER 卡住时有两条路:搞在线 DDL 工具链(gh-ost),或干脆不让数据库知道 schema。后者适合"读多写少、查询模式简单、迭代极快"的场景;查询一复杂就别硬套。
When ALTER TABLE is throttling iteration speed, there are two ways out: an online-DDL toolchain (gh-ost), or simply not letting the database know the schema. The latter fits "read-heavy, simple query patterns, extremely fast iteration"; don't force it once queries get complex.
相关产品:MySQL 相关能力:— 最后核验:2026-10-01
GitHub:不换库,自研 gh-ost 把在线 DDL 做成工具链 成功经验
零停机 schema 变更
运维工具链
MySQL 生态
- 决策
不换库,而是把"在线 DDL"做成工具链:自研 gh-ost。
Don't switch databases — turn "online DDL" into a toolchain instead: build gh-ost in-house.
- 结果
gh-ost 开源后成为 MySQL 生态标准的在线改表工具之一,被大量公司采用。
After open-sourcing, gh-ost became one of the standard online schema-migration tools in the MySQL ecosystem, adopted by many companies.
- 机制根因
无触发器设计——通过读 binlog 追增量、影子表 + 原子切换完成迁移,避免了触发器在主库上的写放大与风险;把"改表不停机"从数据库内核能力变成了可版本化、可审计、可回滚的外部工具。
A triggerless design — it tails the binlog for deltas and migrates via a ghost table plus atomic cutover, avoiding the write amplification and risk that triggers impose on the primary; "zero-downtime ALTER" moved from being a database-kernel capability to a versionable, auditable, rollback-capable external tool.
- 教训
"不换"的机制条件:痛点集中在运维操作面(schema 变更停机),可以用工具链补齐,而数据模型本身没有错。换库是重构数据模型,工具链是补齐操作能力——先判断痛点在哪一层,再决定动哪一层。把运维痛点做成产品,收益会外溢到整个生态。
The mechanism condition for "not switching": the pain is concentrated in the operational surface (schema-change downtime) and can be fixed with tooling, while the data model itself is sound. Switching databases rebuilds the data model; tooling completes the operational surface — first diagnose which layer the pain lives in, then decide which layer to touch. Productizing an operational pain point pays dividends across the whole ecosystem.
相关产品:MySQL 相关能力:在线 DDL / 零停机 schema 变更工具链(gh-ost + pt-osc) 最后核验:2026-10-01
Infisical:MongoDB 迁往 PostgreSQL 失败教训
文档模型
递归/关系结构
迁移工程
零停机
- 决策
迁往 PostgreSQL:数十种数据结构重写、数百个查询重写,用 LevelDB 做旧→新 ID 映射,逐表迁移实现零停机。
Migrate to PostgreSQL: dozens of data structures and hundreds of queries rewritten, old→new ID mapping done with LevelDB, table-by-table migration achieving zero downtime.
- 结果
迁移完成,核心树形查询改用递归 CTE;代价是数十种数据结构、数百个查询全部重写。
The migration completed, with core tree queries rewritten as recursive CTEs; the price was rewriting dozens of data structures and hundreds of queries.
- 机制根因
递归树在文档模型里难表达、难高效查询,"灵活"反噬为查询税;关系模型 + 递归 CTE 才是树形结构的自然归宿;迁移工程本身(ID 映射、逐表双跑)是巨大成本,当初选型时未计入。
Recursive trees are hard to express and hard to query efficiently in a document model — "flexibility" backfires into a query tax; the relational model plus recursive CTEs is the natural home for tree structures; the migration engineering itself (ID mapping, table-by-table dual runs) is a huge cost that was never factored into the original selection.
- 教训
当关系/递归结构成为核心而非点缀时,文档模型的灵活就是查询税;选型权重应给"数据模型与核心结构的匹配度",并把未来迁移工程的成本提前计入决策。
When relational/recursive structures are the core rather than decoration, a document database's flexibility is a query tax; give "fit between data model and core structure" real weight in selection, and price future migration engineering into the decision up front.
相关产品:PostgreSQL(社区版)、MongoDB 相关能力:MVCC + 查询优化器 + 事务性 DDL —— 复杂分析直接跑在 OLTP 主库上 最后核验:2026-10-01
Instagram:不换库,应用层分片撑过爆发增长 成功经验
高并发写入
水平扩展
应用层分片
ID 设计
- 决策
不换数据模型,保留 PostgreSQL 与 SQL 语义,在应用层实现分片;把分片号嵌入 64 位自增 ID(41 位时间戳 + 13 位分片号 + 10 位序列),ID 天然可排序且生成无需跨节点协调。
Keep the data model, PostgreSQL, and SQL semantics; shard at the application layer instead. The shard ID was embedded in a 64-bit auto-generated ID (41-bit timestamp + 13-bit shard + 10-bit sequence), making IDs naturally sortable and generatable with zero cross-node coordination.
- 结果
支撑了被 Facebook 收购后的爆发式增长;这套 ID 方案后来成为分布式 ID 设计的经典范式。
It carried Instagram through the explosive growth after the Facebook acquisition; the ID scheme became a canonical pattern for distributed ID design.
- 机制根因
瓶颈只是单机容量,不是数据模型——查询模式与关系模型依然契合。分片号内嵌在 ID 里,路由在应用层一次算出,彻底避免了分布式协调;事务与 SQL 语义完整保留,应用层几乎无感。
The bottleneck was single-machine capacity, not the data model — the query patterns still fit the relational model. With the shard ID embedded in the ID, routing is computed once at the application layer, eliminating distributed coordination entirely; transactions and SQL semantics stayed intact, nearly invisible to application code.
- 教训
"不换"的机制条件有三:瓶颈只是容量而非模型错配、查询模式与关系模型契合、跨分片查询可以避免。满足时,"分片现有库"比"换模型"便宜一个数量级。ID 设计是分片架构的第一公民——分片键必须在第一天就想好,事后补救的成本远高于事前设计。
The mechanism conditions for "not switching" are three: the bottleneck is capacity rather than model mismatch, query patterns fit the relational model, and cross-shard queries can be avoided. When they hold, "shard the existing database" is an order of magnitude cheaper than "switch the model." ID design is the first-class citizen of sharding architecture — the shard key must be decided on day one; retrofitting it costs far more than designing it upfront.
相关产品:PostgreSQL(社区版) 相关能力:— 最后核验:2026-10-01
LinkedIn 自研 Espresso 替换 Oracle(2012–2018):特性组合倒逼自研 成功经验
多地域/全球化
文档模型
强一致
二级索引
- 决策
自研分布式文档库 Espresso,逐步替换 Oracle,而非在现有产品上做妥协。
Built the distributed document store Espresso in-house and migrated off Oracle gradually, instead of compromising on existing products.
- 结果
截至 2017 年 8 月:约 100 个集群、约 420TB 的 SoT 数据、峰值约 200 万 qps;系统设计发表于 SIGMOD 论文。(迁移规模数字来自 2018 年迁移幻灯片,为 B 级转述口径)
As of August 2017: ~100 clusters, ~420TB of system-of-record data, peak ~2M qps; the design was published as a SIGMOD paper. (Migration scale figures come from the 2018 migration slide deck — second-hand, grade B.)
- 机制根因
层级文档模型(灵活与结构兼得)、timeline 一致性(多地可读写的折中方案)、实时局部二级索引、层级内事务、on-the-fly schema 演进——五个特性组合起来,当时任何现成产品都给不出。
Hierarchical document model (flexibility plus structure), timeline consistency (a multi-region read/write tradeoff), real-time local secondary indexes, intra-hierarchy transactions, on-the-fly schema evolution — no existing product offered all five together.
- 教训
当选型需求的本质是"特性组合"而非单一指标时,硬套现成产品的适配成本会超过自研成本;但自研的门槛是一个 LinkedIn 级别的 infra 团队,且必须先有论文级的设计再动手,否则只是把选型债换成自研债。
When the selection requirement is fundamentally a "feature combination" rather than a single metric, forcing an off-the-shelf product to fit costs more than building; but the bar is a LinkedIn-grade infra team, and a paper-grade design must come before code, or you just swap selection debt for build-it-yourself debt.
相关产品:Oracle Database(甲骨文) 相关能力:出走潮 —— Amazon 下线 Oracle 与"每两周停机 10 分钟做 schema 变更" 最后核验:2026-10-01
DBS 银行:东南亚最大银行的"去闭源"之路 成功经验
去闭源
成本优化
金融合规
- 场景
DBS (Development Bank of Singapore, Southeast Asia's largest bank; per MariaDB's site: roughly $333 billion in total assets, nearly three million financial transactions monthly) had long relied on closed-source proprietary databases, and its technology team wanted an open-source agenda: no CPU license fees, elastic scaling without cost escalation, and compliance with financial regulations (data-at-rest encryption, user auditing). (
https://mariadb.com/resources/customer-stories/dbs/)
- 决策
A conservative start: pilot MariaDB on one non-critical application first; on success, expand to three applications (including the payment service module, a key component); within about two years, 30+ applications went live on MariaDB in production, including the bank's most complex corporate banking systems; the commercial banking and integrated payment engines went live in August 2017 as the first mission-critical apps; MaxScale's CDC modules streamed OLTP data to Hadoop in real time for live insights. (
https://mariadb.com/resources/customer-stories/dbs/)
- 机制根因
专有数据库的成本结构是"横向扩展=成本线性膨胀"(CPU license 计费);MariaDB 开源且无 CPU 费用,扩展只付硬件和运维;MaxScale 代理把读写分离、CDC 从应用层解耦,Hadoop 实时流不侵入 OLTP 主链路;静态加密+审计插件直接满足银行合规。
Proprietary databases have a cost structure where "scaling out = linear cost growth" (CPU-based licensing); MariaDB is open source with no CPU fees, so scaling only pays for hardware and operations; the MaxScale proxy decouples read-write splitting and CDC from the application layer, keeping the Hadoop real-time stream off the OLTP critical path; data-at-rest encryption plus the audit plugin directly satisfy banking compliance.
- 教训
高监管行业做开源替换,渐进式验证(非关键试点→支付模块→公司银行核心)比"大爆炸"替换风险低得多;DBS 还内部办了两场 DBS/MariaDB 技术大会做布道——组织认同比技术验证更难拿下。另注意:本案例的成本数字全部来自厂商侧,严谨选型时应要求客户侧可验证的 TCO 对账。
For open-source replacement in heavily regulated industries, gradual validation (non-critical pilot, then payment module, then corporate banking core) is far less risky than a big-bang swap; DBS also ran two internal DBS/MariaDB tech conferences to evangelize - winning organizational buy-in is harder than winning the technical validation. Note: every cost figure in this case comes from the vendor side; a rigorous evaluation should demand customer-verifiable TCO reconciliation.
相关产品:MariaDB、Oracle Database(甲骨文) 相关能力:治理红利 —— "真正开放的 MySQL"、授权 FUD —— "换到通用云要双倍 license"是人为商业壁垒 最后核验:2026-10-02
LeadDesk:400 万通电话/周的多租户分片 成功经验
多租户
分片扩展
MaxScale
- 决策
采用 MariaDB + MaxScale 代理做分片(sharding):MaxScale 把数据库操作与应用层解耦,应用仍走 MySQL 协议、分片路由由代理层承担,数据库变更无需停机、下线应用。
Adopt MariaDB plus the MaxScale proxy for sharding: MaxScale decouples database operations from the application layer, applications keep speaking the MySQL protocol while the proxy handles shard routing, so database changes happen without taking applications offline.
- 结果
合作宣布时口径(非事后复盘):LeadDesk CEO Olli Nokso-Koivisto 称"用 MariaDB MaxScale 做分片,扩展在技术上没有上限",且快速上线、无需改造应用。以上为厂商合作宣布阶段的表述,无第三方事后验证。
Per the partnership announcement (not a post-hoc review): LeadDesk CEO Olli Nokso-Koivisto said that with MariaDB MaxScale sharding "there is no technical limit for scalability," with fast deployment and no application changes required. These are vendor-announcement-stage statements with no third-party follow-up verification.
- 机制根因
MaxScale 作为数据库代理承担分片路由,存量 LAMP 应用零改造是关键——迁移成本主要在数据重分布而非应用重写;多租户场景下按租户/时间分片,20TB 数据分散到多节点。
MaxScale as a database proxy takes on shard routing, and zero changes to the legacy LAMP applications is the key win - migration cost concentrates on data redistribution rather than application rewrites; in a multi-tenant setup, sharding by tenant or time spreads 20TB of data across nodes.
- 教训
MySQL 系做分片不一定要上 Vitess 或自研中间件:MaxScale 这类协议层代理对"存量 LAMP 应用 + 多租户"是改动最小的横向扩展路径;但要注意这是厂商方案宣布口径,生产口碑样本少,选型时应要求同规模的真实生产参照。
Sharding a MySQL-family database does not always require Vitess or homegrown middleware: a protocol-layer proxy like MaxScale is the least-invasive horizontal scaling path for "legacy LAMP apps plus multi-tenancy"; but note this is a vendor solution announcement, production references at this scale are scarce, and an evaluation should demand real production references of comparable size.
相关产品:MariaDB、MySQL 相关能力:— 最后核验:2026-10-02
MariaDB plc 的过山车:6.72 亿美元估值到 3730 万美元被私有化 失败教训
厂商风险
开源商业化
云战略反复
- 场景
MariaDB the company (MariaDB plc - distinct from the MariaDB Foundation: the GPLv2 core is foundation-stewarded while the company builds BSL/proprietary products around it): listed on the NYSE via SPAC (Angel Pond) in December 2022 at a $672 million valuation; then the stock collapsed - layoffs plus a "going concern" warning in April 2023, another 28% workforce cut (about 84 people) in October alongside a $26.5 million loan while axing strategic product lines including the SkySQL cloud DBaaS and the Xpand distributed engine, and spinning SkySQL out as an independent company in December; in February 2024 K1 Investment Management offered $0.55 per share (about $37.3 million total), completed the acquisition and delisted the company in September, installing Rohit de Souza as CEO. (
https://www.theregister.com/software/2024/09/10/private-equity-firm-completes-mariadb-buyout/273959?td=keepreading)
- 决策
公司层面收缩回核心企业版 Server,砍掉云 DBaaS 与分布式产品线;接受私募股权收购、私有化求生。
At company level, retreat to the core Enterprise Server, kill the cloud DBaaS and distributed product lines; accept a private-equity buyout and go private to survive.
- 结果
Roughly 94% of public-market value evaporated ($672M to $37.3M); SkySQL customers faced product-line termination and forced migration; epilogue: in March 2025 the K1-owned MariaDB announced it would rebuild a DBaaS (the CEO conceding that competing head-on with hyperscalers back then was "crazy" and the economics never worked), while the spun-out SkySQL Inc. raised a $6.6 million seed round in December 2024 to keep operating independently. All figures from public reporting. (
https://www.theregister.com/software/2025/03/12/mariadb-reboots-dbaas-plans-with-open-source-at-the-core/1042458)
- 机制根因
开源流行度不等于商业化能力——SPAC 上市过早、云 DBaaS 与 hyperscaler 正面竞争导致 SkySQL 经济模型不成立(CEO de Souza 原话:产品没真正 ready、靠向云厂商包量转售,企业大客户自己谈的折扣永远更好);核心矛盾是社区版免费 + 云厂商托管分流,企业版订阅撑不起上市公司的增长故事。
Open-source popularity does not equal monetization ability - a premature SPAC listing, and a cloud DBaaS competing head-on with hyperscalers, broke SkySQL's economics (CEO de Souza's words: the product was never really ready, and the model relied on reselling capacity committed from cloud vendors, while large enterprises always negotiate better discounts themselves); the core contradiction: a free community edition plus cloud-vendor hosted alternatives left enterprise subscriptions unable to support a public company's growth story.
- 教训
选开源数据库要看"公司"和"项目"两层风险:内核有基金会托管(GPLv2 不会消失)是安全垫,但商业支持、云托管、roadmap 投入会随公司财务剧烈波动;SkySQL 用户的教训是把生产跑在厂商专有云服务上,等于把命运押在厂商的资产负债表上。这也正是"治理红利"卡片的现实注脚:基金会托管的 GPLv2 内核熬过了 plc 的崩盘。
Evaluating an open-source database means assessing two layers of risk - the company and the project: a foundation-stewarded core (GPLv2 cannot disappear) is a safety net, but commercial support, cloud hosting, and roadmap investment swing violently with the company's finances; SkySQL customers learned that running production on a vendor's proprietary cloud service ties your fate to the vendor's balance sheet. It is also a real-world footnote to the "governance dividend" card: the foundation-hosted GPLv2 core outlived the plc's collapse.
相关产品:MariaDB 相关能力:治理红利 —— "真正开放的 MySQL" 最后核验:2026-10-02
ServiceNow 出走:Xanadu 版本换上 Postgres 系 RaptorDB 失败教训
厂商替换
性能天花板
SaaS
- 决策
不做 MariaDB 原地升级,而是引入 Postgres 系新库;同时首次推出数据库分档:Pro 档(付费,更高性能)与标准档;迁移走 guided 流程,官方口径称对谨慎的管理员来说"无缝嵌入、不耗时"。
No in-place MariaDB upgrade; instead a new Postgres-family database, plus the company's first-ever database tiering: a paid Pro tier (higher performance) alongside a standard tier; migration follows a guided process that the company says is a "seamless drop-in" for a careful admin.
- 结果
ServiceNow 官方口径:Pro 档下整体事务时间 −53%、报表/分析/列表拉取快 27 倍、跨工作流事务吞吐 3 倍;顺带把数据库变成新的增收层(此前 SaaS 从不按数据库性能分档收费)。以上数字均为厂商单方口径,无独立验证。
Per ServiceNow's figures: on the Pro tier, overall transaction times down 53%, reports/analytics/list views 27x faster, 3x transactional throughput across workflows; the database quietly becomes a new revenue layer (the SaaS had never charged by database performance tier before). All figures are vendor one-sided claims with no independent verification.
- 机制根因
客户数据量与性能诉求持续增长后,MariaDB/MySQL 系在复杂查询与分析加速上的天花板显现;Postgres 系的优化器与可扩展性给了 ServiceNow 做"专有增强 + 分档收费"的技术支点;MySQL 协议生态的护城河,挡不住"27 倍"量级的性能差距叙事。
As customer data volumes and performance demands kept growing, the ceiling of the MariaDB/MySQL family on complex queries and analytics acceleration showed; the Postgres family's optimizer and extensibility gave ServiceNow the technical footing for "proprietary enhancement plus tiered pricing"; the MySQL protocol ecosystem's moat could not hold against a "27x" performance-gap narrative.
- 教训
旗舰客户不是永久的——2018 年"database of choice"的 8.5 万库规模,6 年后成了出走案例。对 MySQL 系选型者的启示:把未来全部押在单一协议生态上,要有"头部 SaaS 会为性能切换内核"的预案;评估厂商时除了看功能,更要看"离开成本"和"可替代性"。
Flagship customers are not forever - the 85,000-database "database of choice" of 2018 became an exit case six years later. For MySQL-family evaluators: betting the entire future on a single protocol ecosystem demands a contingency plan for "a top SaaS switching engines for performance"; when assessing a vendor, look beyond features at exit cost and substitutability.
相关产品:MariaDB、PostgreSQL(社区版) 相关能力:MVCC + 查询优化器 + 事务性 DDL —— 复杂分析直接跑在 OLTP 主库上 最后核验:2026-10-02
ServiceNow 巅峰:8.5 万个 MariaDB 库、每小时 250 亿查询 成功经验
SaaS 多租户
超大规模
高可用
- 决策
多实例架构 + MariaDB TX 承载高可用(每层冗余、追求近乎完美的可用性);与 MariaDB 共建产品 roadmap(实时 DDL 等),ServiceNow Ventures 参与 MariaDB C 轮融资、其开发运维 SVP Pat Casey 进入 MariaDB 董事会。
A multi-instance architecture with MariaDB TX delivering high availability (redundancy at every layer, near-perfect availability as the goal); co-building the product roadmap with MariaDB (real-time DDL and more), with ServiceNow Ventures joining MariaDB's Series C round and its SVP of Development and Operations Pat Casey joining MariaDB's board.
- 结果
厂商口径下达到 85k 库/250 亿 qph 的规模,支撑全球 2000 强中 40% 企业的工单流;合作期内联合交付实时 DDL(MariaDB TX 3.0)。以上数字全部来自 MariaDB 公司 2018 年通稿,无独立第三方验证。
Per vendor figures, the deployment reached 85k databases and 25 billion queries per hour, powering workflows for 40% of the world's 2,000 largest companies; the partnership jointly shipped real-time DDL (MariaDB TX 3.0). All figures come from MariaDB Corporation's 2018 press releases, with no independent third-party verification.
- 机制根因
SaaS 多实例模型天然适配"海量中小库":MariaDB 单实例轻量 + MySQL 协议生态,让每个租户库独立、数据隔离、可按客户节奏升级;集中式大库(Oracle RAC 这类)在"实例海"场景反而是成本与运维负担。
The SaaS multi-instance model naturally fits "a sea of small-to-medium databases": MariaDB's lightweight single instance plus the MySQL protocol ecosystem let every tenant database stay independent, data-isolated, and upgradeable on each customer's schedule; a centralized big database (like Oracle RAC) would instead be a cost and operations burden in this "instance ocean" scenario.
- 教训
MariaDB 在"海量小库"场景的 scale-out 靠的是实例数量而非单库规模;这类旗舰客户的成功高度依赖厂商共建(roadmap 输入、联合开发)——而这恰恰是后来故事反转的伏笔(参见关联反例 mariadb-servicenow-leaves:2024 年 ServiceNow 切换到 Postgres 系 RaptorDB)。
MariaDB's scale-out in the "massive small-database" scenario comes from instance count, not single-database size; a flagship customer's success like this depends heavily on vendor co-development (roadmap input, joint engineering) - which is precisely the foreshadowing of the story's reversal (see the companion anti-pattern case mariadb-servicenow-leaves: ServiceNow moved to the Postgres-based RaptorDB in 2024).
相关产品:MariaDB 相关能力:— 最后核验:2026-10-02
Wikipedia 迁入 MariaDB:世界最大百科的 2013 年"信任投票" 成功经验
开源治理
大规模迁移
读密集
- 场景
Wikipedia in 2013: all production databases for the English and German Wikipedias plus Wikidata migrated from the Facebook fork of MySQL 5.1 (r3753) to MariaDB 5.5. The English Wikipedia's databases alone peaked at roughly 50k queries per second daily, with 80% of that load (about 40k qps) carried by just two replica servers; the most common query type had a median execution time of about 0.2ms and a p95 of about 50ms. (
https://diff.wikimedia.org/2013/04/22/wikipedia-adopts-mariadb/)
- 决策
不选 Oracle MySQL 5.5 而选 MariaDB 5.5:MariaDB 的优化器增强、Percona XtraDB 的特性集(如 buffer pool LRU 列表持久化,避免新服务器昂贵的预热),叠加 MySQL 5.5 的功能;同等重要的是治理——Wikimedia 作为自由文化运动支持者,偏好无"双重许可代码库"的自由软件项目,并公开支持 MariaDB 基金会作为非营利托管方。
MariaDB 5.5 over Oracle MySQL 5.5: MariaDB's optimizer enhancements, the Percona XtraDB feature set (e.g., persisting the buffer pool LRU list, avoiding costly warmups on new servers), layered on top of MySQL 5.5's features. Equally important was governance - as supporters of the free culture movement, Wikimedia prefers free software projects without bifurcated code bases under different licenses, and publicly backed the MariaDB Foundation as a not-for-profit steward.
- 结果
Production-grade A/B validation first: one English Wikipedia production replica was taken out of rotation, upgraded to MariaDB 5.5.30, and gradually reweighted against a machine still running the Facebook fork. For the most common query type, p95 dropped from 56ms to 43ms and the average from 15.4ms to 12.7ms; most query types ran 4-15% faster, a few 5% slower, nothing aberrant. Remaining replicas were then upgraded one by one and masters rotated in - the switch was seamless. All figures come from Wikimedia's official blog post. (
https://diff.wikimedia.org/2013/04/22/wikipedia-adopts-mariadb/)
- 机制根因
MariaDB 5.5 = MySQL 5.5 功能超集 + XtraDB + 优化器改进,相当于把 Facebook fork 已验证的 InnoDB 增强路线装进可长期维护的社区发行版;Wikimedia 的决策权重里"谁拥有未来"(治理)与"跑得更快"(性能)并列,而 Oracle MySQL 在治理项上直接出局。
MariaDB 5.5 is a superset of MySQL 5.5 features plus XtraDB plus optimizer improvements - essentially the Facebook fork's proven InnoDB enhancement line packaged into a community release maintainable long-term. In Wikimedia's decision, "who owns the future" (governance) weighed as heavily as "runs faster" (performance), and Oracle MySQL was disqualified on governance alone.
- 教训
大规模 MySQL 系迁移成功的关键不是基准测试分数而是迁移方法:生产环境外跑 MariaDB 从库做兼容性测试、用 pt-upgrade 回放生产读查询对比两边响应、用 tcpdump + pt-query-digest 做 A/B 性能验证、逐台灰度。另外升级暴露了 MediaWiki 代码里依赖"unsigned 整数溢出回绕"的隐性行为——新版严格语义下会报错,升级前必须做应用兼容性测试。
What makes a large-scale MySQL-family migration succeed is not benchmark scores but migration method: compatibility testing on MariaDB replicas outside production, replaying production read queries against both servers with pt-upgrade, A/B performance validation via tcpdump plus pt-query-digest, and rolling upgrades server by server. The upgrade also exposed implicit behavior in MediaWiki code that relied on unsigned-integer overflow wraparound - which errors under the new version's stricter semantics - so application compatibility testing is mandatory before upgrading.
相关产品:MariaDB、MySQL 相关能力:治理红利 —— "真正开放的 MySQL" 最后核验:2026-10-02
贝壳找房:房源智能搜索——300 万向量的毫秒级相似房源推荐 (2021) 成功经验
房产搜索
个性化推荐
多向量融合
余弦距离
A/B 表切换
- 场景
贝壳找房(被称作"中国版 Zillow")的房源库需要一套 AI 搜索引擎:根据用户偏好、搜索历史与房源特征自动推荐相似房源,帮购房者更快找到房子、帮经纪人更快成交。房源的"相似"是多维的:户型、面积、朝向、装修、漆色等特征大多是非结构化的,传统结构化查询表达不了"和这套房子感觉像"的需求。
Beike (often called "China's Zillow") needed an AI search engine over its property listings: automatically recommend similar homes from user preferences, search history, and listing attributes, helping buyers find homes faster and agents close deals sooner. "Similarity" between homes is multi-dimensional — floor plan, size, orientation, interior finishings, paint colors — mostly unstructured signals that structured queries cannot express ("find me homes that feel like this one").
- 决策
用深度学习模型(CNN/RNN/BERT)把房源特征转成特征向量,导入 Milvus 建索引存储;查询时按输入房源、搜索条件或用户画像做相似度检索。系统按输入房源的户型、面积、朝向等特征抽取 4 组特征向量集合,分别在 Milvus 中做相似检索,再把 4 组检索结果对比融合,推荐相似房源;相似度用余弦距离计算。数据更新采用 A/B 表切换机制:前 T 天数据存 A 表,第 T+1 天起写入 B 表,第 2T+1 天起重写 A 表,依此类推。
Deep-learning models (CNN/RNN/BERT) converted listing attributes into feature vectors, which were indexed and stored in Milvus; queries ran similarity search against an input listing, search criteria, or user profile. For each input listing, the system extracted 4 collections of feature vectors (floor plan, size, orientation, etc.), ran similarity search per collection in Milvus, then compared and fused the four result sets to recommend similar homes. Similarity was computed with cosine distance. Data refreshes used an A/B table-switching scheme: the first T days of data lived in table A, day T+1 started writing to table B, day 2T+1 rewrote table A, and so on.
- 结果
在 300 万以上向量的数据集上平均查询耗时 113 毫秒厂商口径。原文称 Milvus 在万亿级数据集上仍能保持高效,本案的 300 万向量属于"相对较小"的规模。
Average query latency of 113 milliseconds on a dataset of over 3 million vectors (vendor figure). The source notes Milvus stays efficient at trillion-vector scale, making this 3M-vector database "relatively small" by comparison.
- 机制根因
非结构化房源特征→数值向量→相似度计算,是向量检索的标准链路;本案的特别之处是"多组向量分别检索再融合"——一套房源用 4 组不同特征向量表达,而不是把所有特征压成一个向量,召回时各组独立检索、结果层再对齐。这对应了推荐系统里"多路召回"的思想:不同特征视角互为补充。
Unstructured listing features to numeric vectors to similarity computation is the standard vector-search pipeline; this case's twist is "retrieve per vector group, then fuse" — one listing expressed as 4 distinct feature-vector collections rather than a single mashed-together vector, with independent retrieval per group and alignment at the result layer. That mirrors the multi-channel recall idea in recommender systems: different feature perspectives complement each other.
- 教训
"一个向量一把梭"不是唯一解法。当实体的相似度是多维的(户型像、装修像、朝向像可能是三回事),拆成多组向量分别检索再融合,往往比硬塞进一个向量空间更可控。代价是查询放大 4 倍,要在延迟预算里留好余量。
"One vector to rule them all" is not the only answer. When an entity's similarity is genuinely multi-dimensional (similar floor plan vs. similar interior vs. similar orientation can be three different things), splitting into multiple vector groups with separate retrieval and late fusion is more controllable. The price is 4x query fan-out — budget the latency accordingly.
相关产品:Milvus 相关能力:— 最后核验:2026-10-02
Shopee:短视频向量检索引擎——从 Milvus 1.x+Mishards 到 2.x 的架构升级 (2023) 成功经验
短视频
视频召回
视频去重
版权匹配
版本迁移
存算分离
- 场景
Shopee 为对抗 TikTok 等短视频平台、上线了多媒体理解(MMU)业务:Shopee Video 短视频功能及独立短视频 App。视频、图像、音频、文本等海量非结构化数据涌入,传统数据库难以处理;内部还并存着视频召回、视频去重、视频推荐等多套用不同技术栈打造的系统,都重度依赖向量检索能力。Shopee 需要一个能无缝嵌入这些异构系统、随数据量增长快速扩容的统一向量检索引擎。
To compete with TikTok-style short-video platforms, Shopee launched a Multimedia Understanding (MMU) business: the Shopee Video feature and a standalone short-video app. Massive volumes of unstructured data — video, images, audio, text — flooded in, overwhelming traditional databases. Internally, several systems built on different tech stacks (video recall, video deduplication, video recommendations) all depended heavily on vector search. Shopee needed a unified vector search engine that could slot into these heterogeneous systems and scale out as data grew.
- 决策
MMU 团队调研了多种开源向量搜索引擎后选定 Milvus 作为向量检索引擎的地基,从零搭建检索系统。第一阶段用 Milvus 1.x(1.1 + Mishards 分布式方案);随着业务与数据量增长,再升级到 Milvus 2.x。
After researching various open-source vector search engines, the MMU team chose Milvus as the foundation and built their vector search systems from scratch. Phase one used Milvus 1.x (1.1 plus Mishards in a distributed setup); as the business and data volumes grew, they upgraded to Milvus 2.x.
- 结果
支撑了 1 亿以上 embedding 向量的存储与检索厂商口径。1.x 阶段 Mishards 的默认分片策略偶发 segment 在只读节点间分布不均导致延迟,团队用"多套 Mishards 集群共享数据库与 S3 存储桶"的方式缓解。升级到 2.x 后,稳定性、可扩展性与多副本能力带来变革:实时检索延迟降低、可用性提高,日志与监控成本下降,系统架构与运维得以简化。视频召回系统以 Milvus 为基石做 Top-K 候选召回再经排序算法精排;版权匹配系统把已发布视频特征全部向量化入库,新上传视频做相似度匹配(预处理、特征提取、结果排序、复检四个模块)识别盗版;去重系统用 Top-K 相似检索加批量检索、聚类与指纹分配消除重复内容。
It now supports storage and search over 100M+ embedding vectors (vendor figure). In the 1.x phase, Mishards' default sharding strategy occasionally produced uneven segment distribution across read-only nodes, causing latency; the team worked around it by running multiple sets of Mishards clusters sharing databases and S3 buckets. The 2.x upgrade was transformative: better stability, scalability, and multi-replica capability delivered lower real-time retrieval latency, higher availability, and cheaper logging and monitoring — simplifying both architecture and operations. The video recall system uses Milvus as its cornerstone for Top-K candidate recall followed by post-ranking; the copyright-match system vectorizes every released video's features and matches each new upload by similarity (four modules: pre-processing, feature extraction, results sorting, rescan) to catch pirated content; the deduplication system combines Top-K similarity search with batch searching, clustering, and fingerprint assignment to remove duplicates.
- 机制根因
Mishards 是 Milvus 1.x 时代的分片中间件:把上游请求拆成子模块分发到各服务再聚合结果,本身不是原生的分布式协调层,分片策略不均就会出现"部分只读节点过载、部分空闲"的倾斜。2.x 的云原生存算分离架构加多副本,让 Top-K 召回可以直接走分布式接口,不再需要外挂缓存与"多套集群共享存储"的拼装方案——架构升级消灭的是一整类运维补丁。
Mishards was the sharding middleware of the Milvus 1.x era: it split upstream requests into sub-modules, fanned them out to sub-services, and aggregated the results — not a natively distributed coordination layer. Skewed sharding meant some read-only nodes overloaded while others idled. Milvus 2.x's cloud-native disaggregated architecture with multi-replica support lets Top-K recall run directly over distributed interfaces, eliminating the need for bolt-on caches and "multiple clusters sharing one store" workarounds — the architecture upgrade deleted an entire category of operational patches.
- 教训
1.x+Mishards 是"能用、但部署维护成本高"的过渡方案(原文明确承认"incurred significant deployment and maintenance costs"),版本演进可以带来架构级红利,但是否值得承担迁移成本,要看业务是否真的撞到了旧架构的天花板。Shopee 的判断标准很务实:延迟与可用性成为瓶颈时才升级,而不是追新。
Milvus 1.x plus Mishards was a "workable but expensive to deploy and maintain" transitional setup (the source openly admits "significant deployment and maintenance costs"). Version upgrades can deliver architecture-level dividends, but the migration cost is only worth paying when the business genuinely hits the old architecture's ceiling. Shopee's bar was pragmatic: upgrade when latency and availability become the bottleneck, not to chase novelty.
来源
Zilliz 客户故事《Shopee Elevates Its Multimedia Business with Milvus》(由 Shopee MMU 团队撰写、授权转载
Zilliz customer story "Shopee Elevates Its Multimedia Business with Milvus" (written by the Shopee MMU team, reposted with permission
相关产品:Milvus 相关能力:十亿级分布式向量检索(存算分离 + 分片的云原生架构)、自托管运维复杂度——"defusing a bomb",全场最高 最后核验:2026-10-02
Tokopedia:语义搜索让商品搜索"聪明 10 倍"——FAISS/Vearch 选型后的 Mishards 高可用 (2022) 成功经验
语义搜索
电商搜索
选型对比
高可用
Mishards
关键词广告
- 场景
Tokopedia 是印尼最大电商平台:9000 万月活用户、860 万商户、覆盖印尼 98% 的行政区厂商口径。商品搜索长期用 Elasticsearch 做关键词检索与排序:ES 把关键词存成 ASCII/UTF 数值序列、建倒排索引,用词频、词距等统计信号打分——完全不理解语义。用户搜"跑鞋"和"运动鞋"在 ES 眼里是两组无关字符。
Tokopedia is Indonesia's largest e-commerce platform: 90 million monthly active users, 8.6 million merchants, reaching 98% of Indonesia's administrative regions (vendor figures). Product search ran on Elasticsearch keyword search and ranking: ES stores keywords as sequences of ASCII/UTF numeric codes, builds inverted indexes, and scores by term frequency, proximity, and other statistical signals — with zero understanding of semantics. To ES, "running shoes" and "sneakers" are unrelated character strings.
- 决策
团队把关键词编码为携带语义的特征向量,对 GitHub 上的多个向量检索方案(FAISS、Vearch、Milvus)做了 POC 与负载测试,最终选 Milvus,理由有三:开箱即用(拉 Docker 镜像调参即可);索引选择面广(除 FAISS 外还有 HNSW、DISK_ANN、ScaNN 等共 11 种);文档齐全、社区支持可靠。先在广告业务落地:用 Milvus 做特征向量检索引擎,把低填充率关键词匹配到高填充率关键词;开发环境先跑单机 standalone 节点验证价值,再因"单节点宕机=整个服务不可用"的风险切到高可用架构:1 个写节点 + 2 个只读节点 + 1 个 Mishards 中间件,用 Ansible 编排部署在 GCP 上。
The team encoded keywords as meaning-carrying feature vectors and ran POCs plus load tests on several GitHub vector search stacks — FAISS, Vearch, and Milvus — picking Milvus for three reasons: it was user-friendly (pull the Docker image, tune parameters); it supported a broader index menu (11 indexes including HNSW, DISK_ANN, and ScaNN alongside FAISS); and it had solid documentation with reliable community support. First landing was the ads business: Milvus as the feature-vector search engine matching low-fill-rate keywords to high-fill-rate keywords. A standalone node in the DEV environment proved the value first; then, because "a crashed standalone node would take the whole service down," the team moved to HA: one writable node, two read-only nodes, and one Mishards middleware instance, orchestrated with Ansible playbooks on GCP.
- 结果
上线后广告关键词匹配带来 10 倍的点击率(CTR)与转化率(CVR)提升(厂商口径/客户自述)。Tokopedia 软件工程师 Rahul Yadav 原话:"Our search system has been much more intelligent, stable, and reliable using Milvus."
Keyword matching in ads delivered a 10x lift in click-through rate (CTR) and conversion rate (CVR) (vendor figure / customer-reported). Rahul Yadav, Software Engineer at Tokopedia: "Our search system has been much more intelligent, stable, and reliable using Milvus."
- 机制根因
FAISS 本质是算法库——快,但没有数据管理、高可用与监控,也不是分布式系统,生产环境要自己造"壳";Vearch 等方案在负载测试中落选。Mishards 作为 1.x 时代的分片中间件,负责把上游请求拆成子模块、分发到各子服务再聚合返回,是"拼装式高可用"的典型:能用,但拓扑(1 写 2 读 1 中间件)要自己用 Ansible 编排维护。
FAISS is fundamentally an algorithm library — fast, but with no data management, high availability, monitoring, or distributed design; production use means building the "shell" yourself. Vearch lost in load testing. Mishards, the 1.x-era sharding middleware, split upstream requests into sub-modules, fanned them out to sub-services, and aggregated results back — the classic "bolted-on HA": workable, but the topology (1 writer, 2 readers, 1 middleware) had to be orchestrated and maintained by hand with Ansible.
- 教训
务实的落地路径是"单机验证价值 → 集群保障可用",而不是一上来就搭分布式。但也要看清:Mishards 时代的高可用是拼装出来的,中间件本身就是运维负担;如果团队没有专人维护这套拓扑,不如直接上托管服务或等 2.x 的原生分布式。
The pragmatic rollout path is "standalone proves value, cluster guarantees availability" — not distributed from day one. But see it clearly: Mishards-era HA was bolted on, and the middleware itself is an operational burden. Without a dedicated owner for that topology, a managed service or the natively distributed 2.x is the saner choice.
来源
Zilliz 客户故事《Tokopedia Achieved a 10x Smarter Search with Milvus》(由 Tokopedia 软件工程师 Rahul Yadav 撰写、授权转载
Zilliz customer story "Tokopedia Achieved a 10x Smarter Search with Milvus" (written by Tokopedia software engineer Rahul Yadav, reposted with permission
相关产品:Milvus 相关能力:索引动物园——按数据冷热与硬件选索引 最后核验:2026-10-02
Trend Micro:APK 病毒检测——从 MySQL 到 Faiss 再到 Milvus 的选型三级跳 (2023) 成功经验
病毒检测
APK 安全
选型对比
MySQL 迁移
实时威胁检测
监控
- 场景
Trend Micro 移动安全团队负责从 Google Play 等外部渠道爬取 APK(Android 安装包),用自研算法检测携带病毒的 APK。样本库膨胀到千万级、日增几十万后,原有方案撑不住了:既要毫秒级相似检索做实时威胁判定,又要跟上每天几十万新增样本的入库速度。
Trend Micro's mobile-security team crawls external APKs (Android application packages) from sources like Google Play and applies proprietary algorithms to detect virus-carrying APKs. As the sample library grew into the tens of millions with hundreds of thousands of new samples daily, the old stack broke down: the system needed millisecond similarity search for real-time threat verdicts and ingestion fast enough to keep up with hundreds of thousands of daily additions.
- 决策
三级跳。第一级 MySQL:项目早期用关系型数据库做 APK 相似搜索,数据量小的时候 SQL 查询够用;数据量到千万级、日增几十万后查询延迟飙升、高并发下出现瓶颈。第二级 Faiss:Facebook 2017 年开源的相似搜索算法库,速度快、索引选项多(IndexFlatL2、IndexFlatIP、HNSW、IVF),但本质是"裸算法库"——无数据管理、无高可用、无监控、非分布式,生产环境要自己造存储层与运维体系。第三级 Milvus:用 C++ 实现的完整向量检索引擎,集成了 Faiss/NMSLIB/Annoy 等主流索引库,兼有直观 API(可按场景选索引类型)、高可用与分布式架构、Prometheus+Grafana 原生监控。团队还做了横向基准对比厂商口径:ES 600ms@100 万向量/128 维,ES+阿里云 900ms@2000 万,Milvus 27ms@10 亿以上,SPTAG 记为"Not good",ES+nmslib+faiss 插件方案 90ms@1.5 亿——Milvus 在延迟与规模两项同时胜出。
Three stages. Stage one, MySQL: the relational database handled APK similarity search early on, when SQL queries sufficed for a small dataset; at tens of millions of samples with hundreds of thousands added daily, query latency soared and high-concurrency bottlenecks appeared. Stage two, Faiss: Facebook's 2017 similarity-search library was fast with many index options (IndexFlatL2, IndexFlatIP, HNSW, IVF), but it was essentially a "bare algorithm library" — no data management, no high availability, no monitoring, not distributed — so production meant hand-building the storage and ops layers. Stage three, Milvus: a complete C++ vector search engine integrating mainstream index libraries (Faiss, NMSLIB, Annoy), with an intuitive API for per-scenario index selection, HA and distributed architecture, and native Prometheus/Grafana monitoring. The team also published a head-to-head benchmark (vendor figures): ES 600ms at 1M vectors/128 dims, ES on Alibaba Cloud 900ms at 20M, Milvus 27ms at 1B+, SPTAG rated "Not good", and an ES+nmslib+faiss plugin combo 90ms at 150M — Milvus won on latency and scale simultaneously.
- 结果
ThashSearch 服务上线数月,端到端查询平均延迟稳定在 95 毫秒以内厂商口径;300 万 192 维向量约 10 秒入库厂商口径,跟上了每日几十万新增样本的节奏。高级研究工程师 Wei Huang 评价:"Milvus delivers unparalleled performance and flexibility... make it an indispensable tool in our APK security efforts." 团队还在规划用 Milvus 的字符串类型 ID 干掉现有的 Redis 缓存层以简化架构。
The ThashSearch service ran live for months with average end-to-end query latency under 95 milliseconds (vendor figure); 3 million 192-dimensional vectors ingested in about 10 seconds (vendor figure), keeping pace with hundreds of thousands of new daily samples. Wei Huang, Senior Research Engineer: "Milvus delivers unparalleled performance and flexibility... make it an indispensable tool in our APK security efforts." The team is also planning to use Milvus string-type IDs to eliminate its Redis caching layer and simplify the architecture.
- 机制根因
MySQL 做相似搜索是"用关系模型干向量活":B-Tree 索引对高维最近邻毫无帮助,数据量上去后只能靠全表暴力比对,延迟与并发双崩。Faiss 解决了"算得快",但生产系统还需要"管得住"(数据管理、副本、监控、水平扩展)——Milvus 的价值正在于把 Faiss 级别的索引库装进了一个分布式数据库的壳里。这是"库 vs 数据库"的经典分野:算法库不管数据生命周期,数据库才管。
MySQL doing similarity search is "a relational model doing a vector's job": B-Tree indexes do nothing for high-dimensional nearest neighbors, so beyond a certain scale every query degrades toward a brute-force full scan — latency and concurrency collapse together. Faiss solved "compute fast," but production systems also need "manage well" (data lifecycle, replicas, monitoring, horizontal scaling) — Milvus's value is packaging Faiss-class index libraries inside a distributed database shell. The classic library-vs-database divide: algorithm libraries don't own the data lifecycle; databases do.
- 教训
向量选型至少要过四问:数据管理(谁管向量的增删与持久化)、高可用与监控(挂了谁知道、谁接管)、分布式扩展(数据量翻 10 倍怎么办)、索引灵活性(不同特征向量能否选不同索引)。"检索快"只是必要条件,Trend Micro 在 Faiss 上栽的跟头就是只看了速度。
Vector-database selection needs at least four questions: data management (who owns vector CRUD and durability), HA and monitoring (who notices and who takes over on failure), distributed scaling (what happens at 10x data), and index flexibility (can different feature vectors use different indexes). "Fast retrieval" is necessary but not sufficient — Trend Micro's Faiss detour is what happens when you only look at speed.
来源
Zilliz customer story "Trend Micro: Leveraging Milvus for Advanced APK Security" (MySQL-to-Faiss-to-Milvus journey, benchmark table, sub-95ms latency, 3M vectors in ~10s ingestion)
https://zilliz.com/customers/trend-micro
相关产品:Milvus 相关能力:索引动物园——按数据冷热与硬件选索引、十亿级分布式向量检索(存算分离 + 分片的云原生架构) 最后核验:2026-10-02
唯品会:推荐系统从 Elasticsearch 迁到 Milvus,10 倍加速与一线运维教训 (2024) 成功经验
电商推荐
Elasticsearch 迁移
读写分离
Java 客户端
索引预热
压测调参
- 场景
唯品会用户超 5200 万、年订单 2.7 亿厂商口径,推荐系统是电商命脉。旧系统用 Elasticsearch 做向量检索:Top-K 检索平均 300ms,叠加后续处理阶段后最终响应长达秒级;共享索引织成"复杂迷宫",构建与维护成本飙升;即便尝试用 hashing 插件抢救 ES 也未达预期。团队需要一套更快、更便宜维护的向量检索栈。
VIPSHOP, with 52M+ customers and 270M annual orders (vendor figures), lives or dies by its recommender system. The legacy Elasticsearch-based vector search was the bottleneck: Top-K retrieval averaged 300ms, and with downstream processing the final response stretched into seconds; shared indexes became a "complex labyrinth" with soaring build and maintenance costs; even a hashing-plugin rescue attempt on ES fell short. The team needed a faster, cheaper-to-run vector search stack.
- 决策
评估后选 Milvus 重建推荐系统,看中的是分布式部署、多语言 SDK、读写分离——"outshone Elasticsearch and other vector search technologies"(原文)。架构:深度学习模型把商品特征转成向量,经 MySQL + ETL 工具写入 Milvus;用户的查询与购买偏好同样向量化,在 Milvus 中做 ANN 相似检索取 Top-K;采用读写分离部署策略。
After evaluation, VIPSHOP rebuilt its recommender on Milvus, drawn by distributed deployment, multi-language SDKs, and read-write separation — a combination that "outshone Elasticsearch and other vector search technologies" (source wording). Architecture: a deep-learning model converts product features into embeddings, ingested into Milvus via MySQL and an ETL tool; user queries and purchase preferences are vectorized the same way, and Milvus runs ANN similarity search for Top-K; the deployment uses a read-write separation strategy.
- 结果
向量查询耗时降到 30ms 以内,相对 ES 方案整体 10 倍加速厂商口径;数据更新与召回链路的三个核心服务(用户向量获取、Milvus 检索、调度合并)响应延迟均为 10ms;推荐系统维护成本下降。
Vector query latency dropped below 30ms — a 10x overall acceleration over the ES solution (vendor figure); the three core services on the update-and-recall path (user-vector acquisition, Milvus search, scheduling merge) each responded in 10ms; recommender maintenance costs fell.
- 机制根因
ES 的 Top-K 高延迟源于其倒排索引基因——为文本检索设计的架构做向量 ANN 是"跨界打工";共享索引则让不同业务的向量数据互相干扰,调参与维护成本指数上升。Milvus 的读写分离让读多写少的推荐场景下读节点可独立扩展,分布式架构让计算与存储随业务量各自伸缩。
ES's Top-K latency comes from its inverted-index DNA — an architecture built for text retrieval doing vector ANN is working outside its design center; shared indexes let different business lines' vector data interfere with each other, driving tuning and maintenance costs up exponentially. Milvus's read-write separation lets read nodes scale independently in this read-heavy recommendation workload, and the disaggregated architecture scales compute and storage separately with the business.
- 教训
本案最有价值的是原文坦承的四个运维坑(这也是 Milvus 自托管复杂度的真实注脚):1)Milvus Java 客户端没有原生重连机制(连接常驻召回服务内存中),唯品会自建了连接池并在客户端与服务端之间加心跳检测保活;2)新 collection 预热不足会导致偶发慢查询,对策是对新 collection 做模拟查询预热;3)检索性能与精度是一对 trade-off,必须按业务场景做压测、设合理的阈值;4)静态数据场景下,先全量导入数据、后建索引更高效。选型时只看 benchmark 数字的人,往往会在这四个坑里挨个摔一遍。
The most valuable part of this story is the four operational pitfalls the source openly documents (a genuine footnote to Milvus's self-hosted complexity): 1) the Milvus Java client has no built-in reconnection mechanism (connections live in the recall service's memory), so VIPSHOP built its own connection pool with heartbeat checks between client and server; 2) insufficient warm-up of new collections causes occasional slow queries — countered by simulating queries against new collections; 3) retrieval performance vs. accuracy is a trade-off that must be settled by load testing against the specific business scenario, with a sensible threshold; 4) for static data, importing everything first and building the index later is more efficient. Anyone who picks a vector DB on benchmark numbers alone will trip over each of these four in turn.
来源
Zilliz 博客《Milvus Made VIPSHOP's Ecommerce Recommender 10x Faster》(52M 用户/270M 订单、ES 300ms→Milvus <30ms、四个运维教训
Zilliz blog "Milvus Made VIPSHOP's Ecommerce Recommender 10x Faster" (52M users / 270M orders, ES 300ms to Milvus sub-30ms, four operational lessons
相关产品:Milvus、MySQL 相关能力:十亿级分布式向量检索(存算分离 + 分片的云原生架构)、自托管运维复杂度——"defusing a bomb",全场最高 最后核验:2026-10-02
小米:手机浏览器 AI 新闻推荐的向量召回层 (2021) 成功经验
新闻推荐
召回排序
BERT
实时增量更新
移动端
- 场景
小米为预装在手机上的自带浏览器做差异化:在每季度 4000 万台以上的出货量厂商口径面前,浏览器需要一套 AI 新闻推荐引擎,根据用户搜索历史与兴趣推荐相似内容。新闻推荐的核心矛盾是"相关性 vs 时效性":文章库每时每刻都在新增,推荐必须基于用户行为与最新内容实时产出。
Xiaomi wanted to differentiate the mobile browser preinstalled on its phones — over 40 million smartphones shipped per quarter (vendor figure) — with an AI news recommendation engine that suggests similar content from user search history and interests. News recommendation is a tug-of-war between relevance and freshness: the article pool grows every minute, and recommendations must reflect both user behavior and the latest content in real time.
- 决策
推荐系统拆成三段:向量化、ID 映射、近似最近邻(ANN)服务。向量化用基于 BERT 的 SimBert 模型(12 层、hidden size 768;中文 L-12_H-768_A-12 持续训练,训练任务为 metric learning + UniLM,在单张 TITAN RTX 上训练了 117 万步,Adam 优化器、学习率 2e-6、batch size 128);ANN 检索层选用 Milvus 做核心数据管理平台;ID 映射负责取回点击量、浏览量等业务指标。
The system was split into three stages: vectorization, ID mapping, and an approximate nearest neighbor (ANN) service. Vectorization used SimBert, a BERT-based model (12 layers, hidden size 768; continued training from Chinese L-12_H-768_A-12 on "metric learning + UniLM" for 1.17 million steps on a single TITAN RTX with the Adam optimizer, learning rate 2e-6, batch size 128). Milvus was chosen as the core data-management platform for the ANN layer; ID mapping retrieved business metrics like page views and clicks.
- 结果
系统采用"T-1 天全量更新 + T 天增量更新"的数据策略:定期删除旧数据、插入 T-1 天的处理后数据,新产生的数据实时入库;入库后即做相似度检索,召回结果再按点击率等指标排序后推送给用户。小米选择 Milvus 的理由原文是"fast, reliable, and requires minimal configuration and maintenance"厂商口径;其分布式版本"极大降低了自建检索层的工作量"(原文)。
The data strategy was "full refresh for the first T-1 days, incremental updates for the following T days": old data deleted on a schedule, processed T-1-day data inserted, and newly generated data ingested in real time; similarity search ran immediately after insertion, and retrieved articles were re-sorted by click-through rate and other factors before being pushed to users. Xiaomi picked Milvus because it is "fast, reliable, and requires minimal configuration and maintenance" (vendor wording); its distributed version "greatly reduces the workload of building a retrieval layer" (source wording).
- 机制根因
新闻推荐是"高频更新 + 实时检索"的双重要求:向量库必须同时擅长快速写入新向量和毫秒级 ANN 检索。Milvus 的动态向量数据管理让"边更新边查"成为可能,而召回与排序的分离(向量库只管 Top-K 召回,点击率等业务指标管精排)是推荐系统的标准分工——不要指望向量库替你做业务排序。
News recommendation demands high-frequency updates plus real-time retrieval at once: the vector store must ingest new vectors fast and serve millisecond ANN search simultaneously. Milvus's dynamic vector-data management makes "update while serving" feasible, and the recall/rank split (the vector DB only handles Top-K recall; business metrics handle fine ranking) is the standard division of labor in recommender systems — don't expect the vector DB to do your business ranking for you.
- 教训
推荐系统的召回层与排序层要用不同的技术选型:ANN 检索解决"从百万级文章库里快速捞出几千篇候选",点击率、时效性等解决"哪篇放最前"。把两层混在一起(比如在向量库里硬算业务权重)是常见的反模式。
The recall layer and the ranking layer deserve different technology choices: ANN retrieval narrows millions of articles to thousands of candidates fast, while click-through rate and freshness decide what goes on top. Mixing the two — e.g., hard-coding business weights inside the vector DB — is a common anti-pattern.
相关产品:Milvus 相关能力:— 最后核验:2026-10-02
Momentic:PostgreSQL+Redis 缓存层迁往 ClickHouse 失败教训
快速迭代/小团队
实时分析/OLAP
架构简化
物化视图
- 场景
The cache tier of Momentic, an AI testing SaaS, serving 12 billion cache operations per day; a small team maintaining a two-layer architecture of PostgreSQL (storing cached data) plus Redis (hot tier). (
https://momentic.ai/blog/postgres-to-clickhouse-migration, single-source company blog)
- 决策
下线 Redis 整层,把缓存搬到 ClickHouse:UPDATE 改写为 INSERT + ReplacingMergeTree 后台去重,用物化视图维护 commit 时间戳;迁移按"双写 → 双读 + 一致性校验 → 切读 → 下线旧库"推进。
Retire the entire Redis layer and move the cache to ClickHouse: rewrites turned UPDATEs into INSERTs with ReplacingMergeTree background deduplication, with materialized views maintaining commit timestamps; the migration followed "dual-write → dual-read + consistency checks → switch reads → decommission the old stack".
- 结果
Redis 整层被删除,架构少一层;UPDATE 语义由"追加插入、后台去重"实现。数字与细节来自公司博客单一信源,无第三方验证。
The whole Redis layer was deleted, one fewer layer in the architecture; UPDATE semantics are now implemented as "append inserts, deduplicate in the background". Figures and details come from a single-source company blog with no third-party verification.
- 机制根因
缓存 workload 是高频 UPDATE,落在 PG MVCC 上就是写放大 + 表膨胀;ClickHouse 的 ReplacingMergeTree 把"更新"变成"追加插入",范式正好匹配高频覆写缓存,物化视图则承担派生状态的维护。但代价要说全:ReplacingMergeTree 的去重要等后台 merge 触发,非常规查询可能读到同一键的多版本(需 FINAL 或版本列过滤,FINAL 有查询代价);并发覆盖、删除与读后写一致性都要单独设计——它不等于传统 OLTP 的 UPDATE 语义。
Cache workloads are high-frequency UPDATEs, which on PG's MVCC mean write amplification plus table bloat; ClickHouse's ReplacingMergeTree turns "updates" into "append inserts", a paradigm that fits high-frequency overwrite caches exactly, while materialized views take over maintaining derived state. But the costs must be stated in full: ReplacingMergeTree dedup only happens when background merges trigger; ad-hoc queries can read multiple versions of the same key (needing FINAL or a version-column filter, and FINAL has query cost); concurrent overwrites, deletes, and read-after-write consistency all need separate design — it is not traditional OLTP UPDATE semantics.
- 教训
用分析型 DB 承担缓存/事件存储时,关键范式转换是把 UPDATE 改写成 INSERT,而不是硬扛原语义;双写双读校验是低风险迁移的标准模板;对 Momentic 这种容忍短暂多版本、读多写少的缓存场景,少一层确实少一份运维税——但需要强一致读后写的场景别硬套,省掉的层会以一致性设计成本的形式回来。
When using an analytical database for cache/event storage, the key paradigm shift is rewriting UPDATEs as INSERTs rather than fighting the native semantics; dual-write/dual-read verification is the standard low-risk migration template; for a cache scenario like Momentic's — tolerant of brief multi-version reads, read-heavy — one fewer layer genuinely means one fewer operational tax. But don't force it onto scenarios needing strongly consistent read-after-write: the removed layer comes back as consistency-design cost.
相关产品:PostgreSQL(社区版)、Redis / Valkey、ClickHouse 相关能力:MVCC + 查询优化器 + 事务性 DDL —— 复杂分析直接跑在 OLTP 主库上、物化视图 + 表引擎的 DDL 管道哲学 最后核验:2026-10-01
Etsy:在 Treasury 功能上试水 MongoDB 失败,迁回 MySQL 分片——"再运维一种生产数据库是巨大的时间浪费"(2012) 失败教训
回迁
NoSQL 回摆
多栈运维税
功能级回迁
- 场景
2010 年左右,Etsy 为新功能 Treasury(精选橱窗)试水 MongoDB,工程师 Dan McKinley 当时还在 Etsy 官方工程博客 Code As Craft 上写过选型时的思考。结果他自己的总结是"an abject failure"(彻头彻尾的失败):同事 Ryan Young 最终把整个功能 port 回了 MySQL 分片——而在此期间,Etsy 的 MySQL 分片体系已经成熟。对 Treasury 功能本身而言这是 MongoDB→MySQL 单向;放在 Etsy 公司层面,是 MySQL→试水 MongoDB→迁回 MySQL 的完整回摆。
Around 2010, Etsy piloted MongoDB for a new feature, Treasury (curated showcases); engineer Dan McKinley wrote about his thinking at the time on Etsy's official engineering blog, Code As Craft. His own verdict: "an abject failure." Colleague Ryan Young ended up porting the entire feature back to the MySQL shards — which had come to maturity in the meantime. For the Treasury feature itself this is a one-way MongoDB→MySQL move; at the Etsy company level it is a complete MySQL→MongoDB-pilot→back-to-MySQL swing.
- 决策
没有"调优再战",直接迁回。Dan 的原话是复盘全文的核心:"adding another kind of production database was a huge waste of time"(再引入一种生产数据库是巨大的时间浪费)。他列的账单很具体:日志、监控、慢查询优化、init 脚本、绘图、复制、分片策略、rebalance 策略、备份、恢复——"probably like 50 other things",每一种数据库都要做一遍;实践中团队只会认真做其中一套,另一套就成了"贫民窟"(ghetto)。而 2010 年的他们"几乎是全世界第一批"在生产环境啃下 MongoDB 这套运维清单的人。
No "tune and try again" — straight back. Dan's line is the core of the whole retrospective: "adding another kind of production database was a huge waste of time." His bill is concrete: logging, monitoring, slow query optimization, init scripts, graphing, replication, sharding strategy, rebalancing strategy, backups, restoration — "probably like 50 other things" — every item once per database; in practice a team only does one of the two stacks properly, and the other becomes a ghetto. And in 2010 they were "the first people in the world to attempt several of these bulletpoints" in production with MongoDB.
- 结果
Treasury 功能迁回 MySQL 分片后稳定运行(作者口径)。值得注意的是 Dan 的诚实注记:失败原因"大概不是你想象的那些"——关于 MongoDB 默认不安全、单机 durability、数据丢失传言,"none of that ever affected us"(没有一条真正影响到我们)。他也不反对别人从零开始用 MongoDB。真正的杀手是"第二种数据库"本身。
The Treasury feature ran stably after moving back to the MySQL shards (author's account). Note Dan's honesty footnote: the failure was "probably not any of the ones you're imagining" — the rumors about MongoDB's unsafe defaults, single-server durability, and data loss: "none of that ever affected us." Nor does he discourage anyone from starting fresh with MongoDB. The real killer was "the second database" itself.
- 机制根因
"多栈运维税"是隐性、线性、不可外包的成本。MongoDB 的无 schema 红利(开发快)是显性的、前置的;运维清单(备份恢复、分片、监控、慢查询)是隐性的、后置的,且与第一种数据库"完全不复用"。Dan 的判断句是全文转折点:"There is no panacea for your scaling problem. You still have to think about how to store your data so that you can get it out of the database."(扩展性问题没有灵丹妙药,你终究要想清楚数据怎么存才取得出来。)当 MySQL 分片在这期间成熟到"足够好",MongoDB 相对 MySQL 的差异化就不值得再付一份运维税——"for most purposes, it's pretty hard to make the case that MySQL and Mongo are really sufficiently different."
The polyglot ops tax is a hidden, linear, non-outsourceable cost. MongoDB's schemaless dividend (fast development) is visible and upfront; the ops checklist (backup/restore, sharding, monitoring, slow queries) is hidden and deferred, and shares nothing with the first database. Dan's pivot line: "There is no panacea for your scaling problem. You still have to think about how to store your data so that you can get it out of the database." Once MySQL sharding matured to "good enough" during the pilot, MongoDB's differentiation over MySQL no longer justified a second ops tax — "for most purposes, it's pretty hard to make the case that MySQL and Mongo are really sufficiently different."
- 教训
引入第二种数据库前先算"双份运维清单":不是 license 或性能对比,而是日志/监控/备份/分片/on-call 全部乘以二;新功能试水新数据库时设"回迁止损线"——Etsy 的止损动作是等人(Ryan Young)而不是等技术成熟;"足够好"的存量栈是新技术的隐形天花板:MySQL 分片成熟之日,就是 MongoDB 试水价值归零之时;复盘时区分"技术本身的问题"和"引入第二种技术的问题",两者对策完全不同。
Before introducing a second database, price the "doubled ops checklist": not license or benchmark comparisons, but logging/monitoring/backup/sharding/on-call times two; set a move-back stop-loss when piloting a new database on a new feature — Etsy's stop-loss was a person (Ryan Young), not waiting for the technology to mature; a "good enough" incumbent stack is the invisible ceiling for new tech: the day MySQL sharding matured was the day the MongoDB pilot's value went to zero; in retrospectives, separate "problems with the technology itself" from "problems with introducing a second technology" — the remedies are completely different.
来源
Dan McKinley(Etsy 工程师)个人博客《Why MongoDB Never Worked Out at Etsy》(2012
Dan McKinley (Etsy engineer) personal blog, "Why MongoDB Never Worked Out at Etsy" (2012
实名复盘,含"abject failure"、Ryan Young 迁回、"huge waste of time"清单等细节
named retrospective with details including "abject failure", Ryan Young's port-back, and the "huge waste of time" checklist
the adoption side is corroborated by his contemporaneous post on Etsy's official engineering blog Code As Craft, but the abandonment retrospective is this single source — no second independent source found
—
相关产品:MongoDB、MySQL 相关能力:简单 OLTP 下"最不折腾周末"的运维体感 最后核验:2026-10-02
卫报:240 万篇文章 10 个月用户无感知——从 MongoDB 迁回关系型 PostgreSQL(2018) 失败教训
回迁
NoSQL 回摆
自建运维税
全托管
JSONB 迁移
- 场景
英国《卫报》的自研 CMS 系统 Composer(记者写稿工具),是其线上全部内容(文章、直播帖、图集、视频)的"唯一事实源",约 230 万条内容。2010 年代初,卫报为摆脱 Oracle 系旧架构(2000 年代 Vignette CMS + Oracle;2005–2009 年 J2EE 单体 + Oracle,300 多张表、1 万行 Hibernate XML 配置),由首席软件架构师 Mat Wall 主导选型 MongoDB(见其 QCon London 2011 演讲《Why I chose MongoDB for guardian.co.uk》,内有"Can we migrate from Oracle to NoSQL?"一页)。2015 年 7 月一场热浪导致自有机房故障后,卫报加速上 AWS,为 MongoDB 购买了 OpsManager 与官方支持合同。但 AWS 上两次严重停机(每次至少 1 小时全站无法发稿)、OpsManager 1→2 升级极其耗时、每年至少 2 个月工程时间耗在数据库管理上,加上昂贵的年度支持费,团队决定弃用 MongoDB。严格说是 A→B→C(Oracle→MongoDB→PostgreSQL),"回迁"发生在范式层面:从 NoSQL 回到关系型。
The Guardian's in-house CMS, Composer (the tool journalists write in), is the "source of truth" for everything published online — articles, live blogs, galleries, video — roughly 2.3 million content items. In the early 2010s the Guardian moved off its Oracle-era stack (2000s: Vignette CMS + Oracle; 2005–2009: J2EE monolith + Oracle, 300+ tables, 10,000 lines of Hibernate XML config) onto MongoDB, led by Lead Software Architect Mat Wall (see his QCon London 2011 talk "Why I chose MongoDB for guardian.co.uk", which includes a "Can we migrate from Oracle to NoSQL?" slide). After a July 2015 heatwave knocked out their own data centre, the Guardian accelerated onto AWS and bought MongoDB OpsManager plus an official support contract. But two severe AWS outages (each blocking all publishing for at least an hour), a painfully time-consuming OpsManager 1-to-2 upgrade, at least two months of engineering time per year burned on database management, and a hefty annual support fee pushed the team to abandon MongoDB. Strictly this is A→B→C (Oracle→MongoDB→PostgreSQL); the "moving back" happened at the paradigm level: from NoSQL back to relational.
- 决策
先等 DynamoDB(当时 DynamoDB 不支持静态加密,卫报等了约 9 个月仍未果,放弃),最终选 AWS RDS 上的 PostgreSQL。关键决策点是 JSONB 列类型:可以用最小的数据模型改动把文档结构搬过去,未来还保留转向更关系化模型的选择权;加上"PostgreSQL 足够成熟,几乎每个问题在 Stack Overflow 上都有答案"。迁移策略是"双跑 + 代理 diff":2017 年 7 月底开始写全新的 APIV2(Scala + doobie,PostgreSQL 后端);2018 年 1 月在预生产环境 CODE 做完整迁移演练;写 Scala 代理把线上流量同时发给新旧两个 API 并记录响应差异;用 GoReplay 做流量回放验证;最后通过 DNS CNAME 一次切换,客户端无感。
They first waited on DynamoDB (which lacked encryption at rest at the time; after about nine months of waiting they gave up), and finally chose PostgreSQL on AWS RDS. The key decision point was the JSONB column type: the document model could be moved over with minimal data-model changes, with the option of going more relational later; plus "Postgres is mature — almost every question we had was already answered on Stack Overflow." The migration strategy was "dual-run + proxy diff": work on a brand-new APIV2 (Scala + doobie, PostgreSQL backend) began at the end of July 2017; a full migration rehearsal ran in the CODE pre-production environment in January 2018; a Scala proxy sent live traffic to both the old and new APIs and logged response differences; GoReplay replayed traffic for validation; the final cutover was a single DNS CNAME change, invisible to clients.
- 结果
10 个月、240 万篇文章(作者口径;文首另记为约 230 万条内容条目)迁移完成,所有 Mongo 相关基础设施关闭:Mongo 实例从 OpsManager 解绑并终止。切换瞬间一次点击、没有出故障。但诚实记录:迁移期间代理本身导致过两次各约 2 分钟的生产抖动(作者原话,团队评估后决定不修代理、把时间花在迁移本身)。此前困扰他们的两类停机(Mongo 侧)在 RDS 上不再出现——代价是把数据库管理交给了 AWS 全托管。
Ten months and 2.4 million migrated articles later (author's figures; the post also describes the corpus as approximately 2.3m content items), all Mongo-related infrastructure was switched off: Mongo instances were detached from OpsManager and terminated. The cutover itself was one click with nothing breaking. Honestly recorded: the proxy caused two production wobbles of about two minutes each during the migration (author's words; the team decided not to fix the proxy and spent the time on the migration instead). The two classes of outage that had plagued them on MongoDB no longer occurred — at the price of handing database management to fully-managed AWS.
- 机制根因
压垮 MongoDB 的不是性能,是"自建运维税"。Composer 是写重型应用(记者每停一下键盘就写一次库),但并发只有几百用户——"not exactly high performance computing",MongoDB 的扩展性优势在这里用不上;反而每次停机都发生在最要命的时刻(无法发稿),而 OpsManager 承诺的"无忧数据库管理"没有兑现:光是升级 OpsManager 本身就需要专家知识,"one-click upgrade"因 MongoDB 版本间认证 schema 变更而落空。团队的原话是全文题眼:"Database management is important and hard – and we'd rather not be doing it ourselves."(数据库管理重要且艰难——而我们宁愿不自己干。)当数据库的差异化价值为零、运维成本全是税时,全托管的关系型数据库就是理性终点;JSONB 则让"文档模型"不再是留在文档库里的理由。
What broke MongoDB was not performance but the self-hosting ops tax. Composer is write-heavy (it writes on every journalist keystroke pause) yet serves only a few hundred concurrent users — "not exactly high performance computing" — so MongoDB's scalability advantage never mattered; instead every outage struck at the worst possible moment (unable to publish), and OpsManager's promised hassle-free database management never materialized: upgrading OpsManager itself required specialist knowledge, and its "one-click upgrade" promise fell through due to authentication-schema changes between MongoDB versions. The team's line is the thesis of the whole post: "Database management is important and hard – and we'd rather not be doing it ourselves." When a database's differentiated value is zero and its ops cost is pure tax, a fully managed relational database is the rational endpoint; JSONB removed the last reason ("document model") to stay on a document store.
- 教训
NoSQL 的"免运维"承诺要按"出事时谁来修"来验算:支持合同在两次停机中都没帮上忙,最后是靠团队自己(一次靠一位在阿布扎比沙漠边缘接起电话的同事);评估替代方案时把"等待上游功能"设止损线(卫报给 DynamoDB 加密功能设了约 9 个月,超时即换);迁移期间为临时组件(代理)投入的修复时间要封顶——它是注定要删的代码;JSONB 这类"关系库里的文档能力"是回迁的润滑剂:数据模型不用重写,回迁阻力小很多。
Verify a NoSQL "no-ops" promise against "who fixes it when it breaks at 3am": the support contract helped in neither outage — the team fixed both themselves (once thanks to a colleague who picked up the phone from a desert on the outskirts of Abu Dhabi); when evaluating alternatives, put a stop-loss on "waiting for an upstream feature" (the Guardian gave DynamoDB's encryption feature about nine months, then switched); cap the repair time invested in throwaway migration scaffolding like the proxy — it is code destined for deletion; document capabilities inside relational databases, like JSONB, are the lubricant of moving back: no data-model rewrite needed, far less resistance.
来源
Guardian Engineering 官方工程博客《Bye bye Mongo, Hello Postgres》(2018-11-30
Guardian Engineering official blog "Bye bye Mongo, Hello Postgres" (2018-11-30
相关产品:Oracle Database(甲骨文)、MongoDB、PostgreSQL(社区版) 相关能力:JSONB —— 文档能力"反杀"专用文档库 最后核验:2026-10-02
Jepsen 独立测试 MongoDB 4.2.6:最强读写关注下仍违反快照隔离(2020) 失败教训
选型评估
独立测试
多文档事务
快照隔离
读写关注
- 场景
MongoDB 4.2 引入多文档事务后宣称"full ACID transactions"、"among the strongest data consistency guarantees"。2020 年 5 月 Jepsen 独立测试 MongoDB 4.2.6(2 分片×3 节点复制集 + configsvr,共 9 节点),用 Elle 事务一致性检查器验证多文档事务是否真能提供快照隔离。本次为 Jepsen 独立发起、无偿测试。
After MongoDB 4.2 introduced multi-document transactions, MongoDB claimed "full ACID transactions" and "among the strongest data consistency guarantees". In May 2020, Jepsen independently tested MongoDB 4.2.6 (2 shards × 3-node replica sets + configsvr, 9 nodes total), using the Elle transaction consistency checker to verify whether multi-document transactions truly delivered snapshot isolation. This was independently initiated, uncompensated work by Jepsen.
- 决策
测试设计对照多档读写关注——从默认到最强(read concern `snapshot` + write concern `majority`),并注入隔离主节点的网络分区,观察事务在故障与正常运行两种状态下的行为。
The test design compared multiple read/write concern levels — from defaults to the strongest (read concern `snapshot` + write concern `majority`) — while injecting network partitions isolating primary nodes, observing transaction behavior both under faults and during normal operation.
- 结果
即使在最强读写关注下仍违反快照隔离:read skew(读偏斜)、循环信息流、重复写、读到"自己未来的写"(retrocausal 异常)。更关键的是事务会把库/集合级设置的读写关注**降级**为 `local`/`w:1`,导致丢确认写、脏读;`snapshot` 读关注必须配 `majority` 写关注才有效,连只读事务也不例外。无故障正常运行时约 10% 事务出现异常(1461/13914,Jepsen 报告原文口径)。MongoDB 在报告发布 11 天后确认事务重试机制存在 bug,补丁排入 4.2.8。
Snapshot isolation was violated even at the strongest read/write concerns: read skew, cyclic information flow, duplicate writes, and reads observing "their own future writes" (retrocausal anomalies). More critically, transactions **downgraded** database/collection-level read/write concerns to `local`/`w:1`, causing loss of acknowledged writes and dirty reads; the `snapshot` read concern only worked when paired with `majority` write concern — even for read-only transactions. Roughly 10% of transactions exhibited anomalies during normal, fault-free operation (1461/13914, per the Jepsen report). Eleven days after publication, MongoDB confirmed a bug in the transaction retry mechanism, with a patch scheduled for 4.2.8.
- 机制根因
问题不止于单个实现 bug,更在于默认配置哲学。MongoDB 长期为保性能采用弱默认值(写关注默认 w:1),而事务作为新功能继承了这套"降级"逻辑——事务内直接忽略库/集合级的读写关注设置。用户以为"我开了 snapshot 就是快照隔离",实际还要逐事务再配 write concern `majority`。这是对"默认安全"预期的系统性违背:安全不是默认项,而是需要逐事务手动开启的选项。
The problem went beyond a single implementation bug to a default-configuration philosophy. MongoDB long favored weak defaults for performance (default write concern w:1), and transactions — as a new feature — inherited this "downgrade" logic, silently ignoring database/collection-level read/write concern settings inside transactions. Users who believed "I set snapshot, so I get snapshot isolation" still had to configure write concern `majority` per transaction. This is a systematic violation of the "safe by default" expectation: safety was not the default but a per-transaction manual opt-in.
- 教训
评估文档库事务能力时,不要看营销页的"ACID"字样,要看默认配置下的实际行为;关键事务逐条显式设置读写关注;任何"full ACID"宣称都值得用独立测试报告交叉验证。诚实注记:这是 4.2.6(2020 年 5 月)的结论,事务重试 bug 已在 4.2.8 修复;但"事务降级读写关注"是文档化的设计选择,选型时需按当前版本重新核验,不能默认已改。
When evaluating a document database's transaction capabilities, don't read the marketing page's "ACID" claims — read its actual behavior under default configuration; set read/write concerns explicitly on every critical transaction; cross-check any "full ACID" claim against independent test reports. Honesty note: these conclusions apply to 4.2.6 (May 2020); the retry bug was fixed in 4.2.8. But "transactions downgrading read/write concerns" was documented design behavior — re-verify against the current version during selection rather than assuming it changed.
来源
Jepsen《Jepsen: MongoDB 4.2.6》(2020-05-15,独立测试、无偿
Jepsen, "Jepsen: MongoDB 4.2.6" (2020-05-15, independent, uncompensated
MongoDB 确认事务重试 bug、补丁排入 4.2.8(该报告 2020-05-26 更新注记)同上
MongoDB confirmed the transaction retry bug with a patch scheduled for 4.2.8 (update note on the same report, 2020-05-26)
相关产品:MongoDB 相关能力:多文档事务的快照隔离与读写关注语义 最后核验:2026-10-02
Jepsen 独立测试 MySQL 8.0.34:默认 REPEATABLE READ 名不副实,RDS 版连串行化都保不住(2023) 失败教训
选型评估
独立测试
隔离级别
InnoDB
RDS
- 场景
MySQL InnoDB 的默认隔离级别是 REPEATABLE READ。2023 年 12 月 Jepsen 独立复核 Kleppmann 2014 年 Hermitage 项目的结论,用 Elle 事务一致性检查器验证 MySQL 8.0.34 四档隔离级别的真实语义,并顺带测试了 AWS RDS Multi-AZ DB Cluster。本次为 Jepsen 独立发起、无偿测试(Peter Alvaro 与 Kyle Kingsbury 合著)。
MySQL InnoDB's default isolation level is REPEATABLE READ. In December 2023, Jepsen independently re-examined the conclusions of Kleppmann's 2014 Hermitage project, using the Elle transaction consistency checker to verify the true semantics of MySQL 8.0.34's four isolation levels, and additionally tested AWS RDS Multi-AZ DB Cluster. This was independently initiated, uncompensated work by Jepsen (co-authored by Peter Alvaro and Kyle Kingsbury).
- 决策
对照 Adya 形式化定义(PL-2.99)与 ANSI SQL 标准,测试对象包括单节点 InnoDB、binlog 一主多从,以及 AWS RDS MySQL 集群。
Measured against Adya's formal definitions (PL-2.99) and the ANSI SQL standard, the test targets included single-node InnoDB, binlog primary-replica replication, and an AWS RDS MySQL cluster.
- 结果
默认的 REPEATABLE READ 不满足 PL-2.99(ANSI RR),甚至违反快照隔离与单调原子视图(MAV):G2-item、G-single(read skew)、lost update、内部一致性违反——Jepsen 原文口径为"略强于 Read Committed 而已"。附带发现:AWS RDS MySQL 集群频繁违反可串行化(主库执行成功的 CREATE DATABASE,从库一小时不恢复,复制脆弱)。截至报告页面,未见 Oracle 官方回应(查证为无);Jepsen 建议 AWS 修改 RDS 默认配置或在文档中明确说明限制。
The default REPEATABLE READ does not satisfy PL-2.99 (ANSI RR), and even violates snapshot isolation and monotonic atomic view (MAV): G2-item, G-single (read skew), lost update, and internal consistency violations — "somewhat stronger than Read Committed," in the report's own words. As a lagniappe: AWS RDS MySQL clusters routinely violated serializability (a CREATE DATABASE that succeeded on the primary never recovered on replicas within an hour — fragile replication). No official response from Oracle appears on the report page (verified absent); Jepsen recommended AWS change RDS defaults or document the limitations explicitly.
- 机制根因
隔离级别同名不同义是行业通病——Jepsen 指出,在其测过的库里只有 SQL Server 的 RR 对应 PL-2.99,PostgreSQL 的 RR 实际是快照隔离。MySQL 的 RR 是"基于首次读的快照 + 写时可见性例外"的混合语义,文档里那句"DML 语句不一定读快照"就是裂缝所在。RDS 的问题则出在托管层的复制实现,而非 InnoDB 本身——选型时"MySQL"和"RDS 上的 MySQL"是两个不同的评估对象,不能默认等同。
Same-named isolation levels with different semantics are an industry-wide affliction — Jepsen notes that among the databases it has tested, only SQL Server's RR corresponds to PL-2.99, while PostgreSQL's RR is actually snapshot isolation. MySQL's RR is a hybrid semantics of "snapshot from first read, plus visibility exceptions on writes" — the documentation's caveat that "the snapshot does not necessarily apply to DML statements" is exactly where the crack is. The RDS problem lies in the managed layer's replication implementation, not InnoDB itself — "MySQL" and "MySQL on RDS" are two different evaluation targets and must not be assumed equivalent.
- 教训
不要依赖隔离级别名称做正确性假设,关键路径用显式 `SELECT ... FOR UPDATE` 加锁;评估云托管版本时要把"托管层的复制语义"单独列为评估项;同名隔离级别跨库对比前先对齐语义定义。
Never rely on isolation level names for correctness assumptions; use explicit `SELECT ... FOR UPDATE` locking on critical paths; when evaluating a cloud-managed edition, list "the managed layer's replication semantics" as its own evaluation item; align semantic definitions before comparing same-named isolation levels across databases.
来源
Jepsen《Jepsen: MySQL 8.0.34》(2023-12-19,独立测试、无偿
Jepsen, "Jepsen: MySQL 8.0.34" (2023-12-19, independent, uncompensated
相关产品:MySQL 相关能力:InnoDB 隔离级别的真实语义与文档裂缝 最后核验:2026-10-02
Netflix 逃离 Oracle:每两周停一次机的 schema 变更 失败教训
多地域/全球化
云原生
零停机 schema 变更
高可用
- 决策
早期用 Oracle 单库承载核心数据;2011–2013 年间经 SimpleDB 过渡,最终迁往 Cassandra。
Oracle as the single core database early on; between 2011 and 2013 Netflix moved via SimpleDB to Cassandra.
- 结果
到 2013 年约 95% 数据在 Cassandra 上,Oracle 被彻底移除(Netflix 自述,单方口径)。
By 2013 roughly 95% of its data ran on Cassandra and Oracle was fully removed (figures from Netflix's own account — self-reported).
- 机制根因
传统 RDBMS 把 DDL 当作需要锁表窗口的操作,而全球流媒体没有"维护窗口"——这是商业模型与运维模型的双重错配;Cassandra 无主架构为廉价节点设计,节点随意增减,与云弹性同构;集中式主库与"一切皆可失败"的云原生可用性模型根本冲突。
Traditional RDBMS treats DDL as an operation needing a lock window, while a global streaming service has no maintenance window — a mismatch of both business and operations models. Cassandra's masterless architecture was designed for cheap nodes that can be added or removed freely, isomorphic to cloud elasticity; a centralized primary fundamentally conflicts with a cloud-native "everything can fail" availability model.
- 教训
当数据库成为可用性单点、且无法在弹性硬件上水平扩展时,集中式 RDBMS 不适合云原生全球服务;schema 演进的停机成本必须计入选型总账,不能只看功能矩阵。
When the database becomes an availability single point and can't scale horizontally on elastic hardware, a centralized RDBMS doesn't fit a cloud-native global service. The downtime cost of schema evolution must be booked in the selection ledger — not just the feature matrix.
相关产品:Oracle Database(甲骨文)、Apache Cassandra / ScyllaDB 相关能力:出走潮 —— Amazon 下线 Oracle 与"每两周停机 10 分钟做 schema 变更" 最后核验:2026-10-01
Notion 的 Postgres 分片之路(2021–2023):480 逻辑分片扛住 2000 亿 blocks 成功经验
水平分片
应用层路由
零停机
租户隔离
- 场景
In Notion, everything is a block, and a block is a row in Postgres. By early 2021 more than 20 billion block rows were squeezed into a single Postgres instance, with data doubling every 6–12 months; the official blog says it has since grown past 200 billion blocks — hundreds of terabytes compressed. (
https://www.notion.com/blog/building-and-scaling-notions-data-lake)
- 决策
No microservices split, no database swap — in 2021 Notion horizontally sharded Postgres: 32 physical instances × 15 logical shards = 480 logical shards; in 2023 they grew to 96 physical × 5 logical, keeping the logical total at 480. The shard key is workspace_id (one workspace's data always lives on one shard); logical shards are implemented as Postgres schemas, routing lives in the application, and PgBouncer manages connections. (
https://www.notion.com/blog/sharding-postgres-at-notion)
- 结果
据官方博客,480 逻辑分片承载 2000 亿+ blocks、数百 TB 数据;2021→2023 两次物理扩容均未重分片、未停机;单 workspace 查询天然落在单个分片,跨分片 join 基本消失。
Per the official blog, 480 logical shards carry 200B+ blocks and hundreds of terabytes; both physical expansions (2021→2023) completed without resharding or downtime; single-workspace queries naturally land on one shard, so cross-shard joins effectively disappeared.
- 机制根因
分片键 = 主导查询模式(用户一次只查一个 workspace),同租户数据共置,消灭了分片架构最痛的跨分片 join;逻辑分片与物理实例解耦,480 是高度可合成数(可被 2、3、4、5、6、8、10…整除),物理扩容只是搬运逻辑分片,无需重算路由;"单体应用 + 分片数据"的组合保住了开发效率,只把复杂度放在真正需要的地方。
The shard key equals the dominant query pattern (users query one workspace at a time); co-locating one tenant's data eliminates the most painful part of sharded architectures — cross-shard joins. Decoupling logical shards from physical instances, and picking 480, a highly composite number (divisible by 2, 3, 4, 5, 6, 8, 10…), means physical growth is just moving logical shards around with no re-routing math. The "monolith app + sharded data" combo preserved development velocity, putting complexity only where it was truly needed.
- 教训
分片键必须等于你的主导访问模式,而不是数据的某种"自然"属性;逻辑分片数选高度可合成数,为未来物理扩容留好算术空间;分片是最后手段——先吃尽垂直扩展和读副本,Notion 也是单实例撑到 200 亿行才动手。
Your shard key must equal your dominant access pattern, not some "natural" attribute of the data; pick a highly composite number of logical shards to leave arithmetic room for future physical growth; sharding is the last resort — exhaust vertical scaling and read replicas first. Notion itself rode a single instance to 20 billion rows before acting.
相关产品:PostgreSQL(社区版) 相关能力:— 最后核验:2026-10-01
Nubank 选择 Datomic 不可变数据库(2016–):把审计合规变成免费属性 成功经验
金融级审计
时间一等公民
HTAP
强一致
- 决策
核心账务直接采用 Datomic(不可变事实数据库),而非传统 RDBMS。
Adopted Datomic (an immutable fact database) for the core ledger, instead of a traditional RDBMS.
- 结果
支撑 Nubank 从初创成长为拉美最大数字银行;历史回放能力多次用于微服务拆分与从 bug 中恢复。(技术分享 PDF 为第三方托管副本,证据降级至 B+)
Carried Nubank from startup to Latin America's largest digital bank; historical replay was used repeatedly for service splits and bug recovery. (The tech-talk PDF is a third-party-hosted copy; evidence downgraded to B+.)
- 机制根因
时间一等公民——事实不可变、只追加,任意时点查询免费获得,审计合规从成本变成副产品;读写分离——读在应用端(peer library 缓存),写经 transactor 串行,保证可序列化 ACID,OLAP 不打扰 OLTP,无需 ETL 做实时分析;历史回放——拆分服务时按时间重放事务,保住账目完整;存储后端可换——PII 数据用 RDS Postgres + EBS 加密,非 PII 用 DynamoDB。
Time as a first-class citizen — facts are immutable and append-only, so querying any point in time is free, turning audit compliance from a cost into a byproduct; read/write separation — reads at the application side (peer library cache), writes serialized through the transactor for serializable ACID, while OLAP queries never disturb OLTP, enabling real-time analytics with no ETL; historical replay — services are split by replaying transactions in time order, keeping ledgers intact; swappable storage backend — PII data on RDS Postgres with EBS encryption, non-PII on DynamoDB.
- 教训
审计要求本质上是"历史状态可查",不可变 + 时间一等公民把它变成数据模型的免费属性;但代价是小众生态与 Clojure 绑定——选型时要把"生态存活期"和"架构收益"放在同一张天平上称,不要只看模型优雅。
Audit requirements are fundamentally "historical state must be queryable," and immutability plus time-as-first-class-citizen makes that a free property of the data model; the price is a niche ecosystem and Clojure lock-in — weigh "ecosystem longevity" against "architecture payoff" on the same scale, not model elegance alone.
相关产品:无 相关能力:— 最后核验:2026-10-01
百丽时尚迁移OceanBase踩坑复盘:唯一键语义、OMS大字段与执行计划三连Bug(2025) 失败教训
迁移踩坑
数据一致性
执行计划
版本缺陷
- 场景
同一项目(百丽财务系统MyCat→OceanBase)的另一面:迁移跑通了,但在数据校验与SQL治理阶段连踩四个坑。源端是约束宽松的MyCat分片:历史脏数据(ID重复)、唯一键只在分区内生效;目标端是全局约束的OceanBase;中间还隔着OMS→Kafka的异构链路。任何一处语义理解偏差都会变成丢数,而丢数在财务系统是不可接受的。
The other side of the same project (Belle's finance system, MyCat → OceanBase): the migration went through, but data verification and SQL governance hit four pitfalls in a row. The source was loosely-constrained MyCat shards — historical dirty data (duplicate IDs), unique keys enforced only within a partition; the target was globally-constrained OceanBase; and an OMS→Kafka heterogeneous link sat in between. Any misunderstood semantic difference becomes lost rows, which is unacceptable in a finance system.
- 决策
团队没有"一遍过"的幻想,建了三道防线:自研crc32分块校验工具做全量比对;流量回放(全量日志→CSV→双端对比)抓兼容与性能问题;上线前用Kafka把数据反向同步回MyCat/Oracle,保留回切预案。
The team had no "one-shot" illusions and built three lines of defense: a self-built crc32 chunk-comparison tool for full verification; traffic replay (full logs → CSV → dual-side comparison) to catch compatibility and performance issues; and pre-cutover Kafka reverse-sync back to MyCat/Oracle as a rollback plan.
- 结果
数据校验抓出4个问题:MyCat历史遗留ID重复;唯一键"分区级vs全局"约束范围差异导致丢数;OMS主键Hash分区+下游唯一键在并发消费时丢数;OMS 4.2.5.2之前版本>4K的LOB字段前镜像缺失,下游被置空(4.2.5.2已修复)。SQL治理抓出执行计划三连Bug:local rescan规则误裁剪走高代价hash join(V4.2.5.3之前版本,4.2.5.3修复)、tablegroup空分区导致手动均衡任务卡住(V4.2.5.3,4.2.5.4修复)、含复制表的SQL无法复用执行计划(V4.2.5.4,4.2.5.5修复)。另发现OceanBase离线DDL会直接锁表,团队自研"提交工单前先在测试环境跑一遍,看Table ID是否变化"来识别Online/Offline。
Data verification caught 4 issues: historical duplicate IDs inherited from MyCat; row loss from the partition-level vs global unique-key constraint scope difference; row loss from OMS primary-key hash partitioning combined with downstream unique keys under concurrent consumption; and missing LOB before-images for fields over 4K in OMS versions before 4.2.5.2, nulling downstream columns (fixed in 4.2.5.2). SQL governance caught three execution-plan bugs: local-rescan rules wrongly pruning the cheap nested-loop plan into a costly hash join (before V4.2.5.3, fixed in 4.2.5.3), manual rebalance jobs hanging on empty partitions in tablegroups (V4.2.5.3, fixed in 4.2.5.4), and unreusable execution plans for SQL touching duplicate tables (V4.2.5.4, fixed in 4.2.5.5). They also found OceanBase offline DDL locks tables outright, so they built a pre-submit check — run the DDL in a test environment and watch whether the Table ID changes — to distinguish online from offline DDL.
- 机制根因
坑的共性是"语义搬家":MyCat的唯一键、分区、Binlog都是分片局部的,OceanBase是全局的——约束从分区级变全局级,旧脏数据立刻现形;Kafka按主键Hash分区后,不同分区的并发写入叠加下游唯一键,消费顺序一乱就丢数,这是把"有序"假设从单库搬到分布式消息队列必然要付的代价。执行计划三连Bug是另一类:优化器在新版本快速迭代中引入的回归,说明"追最新版"本身就是一种风险敞口。DDL锁表暴露的是运维心智差异:MySQL系习惯pt-osc式的Online DDL,而OceanBase离线DDL直接锁表,必须前置识别。
The common thread is "semantic relocation": MyCat's unique keys, partitions, and Binlog are all shard-local, while OceanBase's are global — constraints move from partition scope to global scope and old dirty data surfaces immediately; after Kafka hash-partitions by primary key, concurrent writes from different partitions interact with downstream unique keys and lose rows the moment consumption order skews — the inevitable price of moving an "ordered" assumption from a single database to a distributed message queue. The three plan bugs are a different species: optimizer regressions introduced during fast version iteration, which means "chasing the latest release" is itself a risk exposure. The DDL locking issue exposes an ops-mindset gap: MySQL practitioners expect pt-osc-style online DDL, while OceanBase offline DDL locks the table and must be identified up front.
- 教训
分布式迁移最大的风险不在数据库内核,在数据链路的语义差异——画链路图时把"唯一键范围、分区语义、消费顺序"三项显式标出来;校验工具必须自研或深度定制,通用工具覆盖不了异构语义差;生产环境优先选经过多个补丁验证的版本而非最新版;DDL上线前先在测试环境验证Online/Offline,把Table ID变化检查做进发布工单。
The biggest risk in a distributed migration is not the database kernel but semantic differences across the data links — when drawing the link diagram, explicitly mark "unique-key scope, partition semantics, consumption order"; verification tooling must be self-built or deeply customized because generic tools don't cover heterogeneous semantic gaps; prefer production versions validated through multiple patch releases over the newest release; verify online/offline DDL in a test environment before rollout and bake the Table-ID-change check into the release ticket.
来源
Lu Wenhao (Database Lead, Belle Fashion), "Migrating from Sharded MyCat to OceanBase: Experience Summary and Issue Catalog of Belle's Core Finance System Migration" (CSDN blog
—
no publication date shown on page
—
相关产品:OceanBase、MySQL 相关能力:— 最后核验:2026-10-02
百丽时尚:MyCat分库分表迁OceanBase,成本核算从10小时压到20分钟(2025) 成功经验
分库分表下线
成本优化
性能提升
零售财务
- 场景
百丽时尚是中国头部鞋服集团(20+品牌、8000+线下门店、覆盖300+城市),科技中心统筹零售、库存、财务等核心系统。其财务系统长期跑在MyCat分库分表架构上(按大区分片、一主两从,主机房北京),演进中暴露三类问题:大区合并/调整要跨库搬数据,脚本复杂、回滚窗口小;MyCat只做基础路由,复杂SQL与分布式事务支持弱,财务模块被迫把分片表改成全局表,数据冗余、维护成本上升;扩容要重划大区再全量迁移,性能瓶颈只能靠垂直升配。(百丽时尚数据库负责人卢文豪署名复盘,CSDN)
Belle Fashion is one of China's largest footwear/apparel groups (20+ brands, 8,000+ offline stores, 300+ cities), with its tech center running core retail, inventory, and finance systems. Its finance system long ran on a MyCat sharded architecture (sharded by sales region, one primary plus two replicas, primary IDC in Beijing). Three pain classes emerged: merging or resizing regions required cross-shard data moves with complex scripts and tiny rollback windows; MyCat only does basic routing with weak support for complex SQL and distributed transactions, forcing finance modules to convert sharded tables into global tables — data redundancy and maintenance cost grew; scaling out meant re-sharding regions plus full data migration, so performance bottlenecks could only be addressed by vertical hardware upgrades. (First-hand retrospective by Lu Wenhao, Belle Fashion's database lead, on CSDN)
- 决策
科技中心把选型锁定在原生分布式数据库,提两项硬诉求:高可用、零数据丢失;在线弹性扩展,不停服不搬数据。多轮评估后选定OceanBase,迁移分三步走:先梳理全链路数据流转(上游Otter同步、下游Binlog入仓与Oracle报表),再用自研crc32分块比对工具做数据校验,最后用流量回放(全量日志→CSV→双端对比)做SQL兼容与性能治理。
The tech center narrowed selection to native distributed databases with two hard requirements: high availability with zero data loss, and online elastic scaling without downtime or data relocation. After multiple evaluation rounds they chose OceanBase, migrating in three steps: first map the full data flow (upstream Otter sync, downstream Binlog-to-warehouse and Oracle reporting links); then data verification with a self-built crc32 chunk-comparison tool; finally SQL compatibility and performance governance via traffic replay (full query logs → CSV → dual-side comparison).
- 结果
成本核算跑批从MyCat下的10小时压到20分钟,约30倍;存储从20.3TB压到1.3TB,降96.7%(作者称一半功劳是OceanBase高压缩,一半是MyCat冗余数据被整合释放);服务器从37台减到10台,硬件成本从207万元降到84万元,降59.4%。约4.5万个SQL ID回放后只发现1个不兼容(SQL末尾"--"注释在OceanBase报错)。以上数字均为百丽DBA负责人单方披露,无第三方独立复现。
The costing batch job dropped from 10 hours on MyCat to 20 minutes, roughly 30x. Storage shrank from 20.3 TB to 1.3 TB, down 96.7% (the author attributes half to OceanBase's high compression and half to redundant MyCat data being consolidated away). Servers went from 37 to 10, hardware cost from RMB 2.07M to 0.84M, down 59.4%. Replaying ~45,000 SQL IDs surfaced only one incompatibility (a trailing "--" comment errors on OceanBase). All figures are single-party disclosures by Belle's DBA lead; no independent third-party reproduction found.
- 机制根因
MyCat是"中间件+MySQL分片"的拼接架构,分布式能力(事务、复杂查询、扩缩容)都不在数据库内核里,瓶颈只能靠业务层妥协绕行(改全局表、分区内收敛),绕一次留一笔技术债。OceanBase把分片、事务、压缩收进同一个分布式内核:Paxos多副本让RPO=0成为架构免费项;单机分布式一体化使扩容不再触发数据重分布;编码+两层压缩把存储压到1/20以下。代价是下游链路重构:MyCat时代下游吃Binlog,OceanBase无原生Binlog,只能走租户级单线程的Binlog Service(高峰有瓶颈)或OMS→Kafka(需自研Kafka到Oracle的同步工具)。
MyCat is a stitched architecture — middleware plus MySQL shards — where distributed capabilities (transactions, complex queries, scaling) live outside the database kernel, so every bottleneck is worked around with application-layer compromises (global tables, partition-local access), each leaving tech debt. OceanBase pulls sharding, transactions, and compression into one distributed kernel: Paxos multi-replica makes RPO=0 an architectural freebie; the standalone/distributed unified design means scaling no longer triggers data redistribution; encoding plus two-tier compression pushes storage below 1/20th. The price is downstream link re-engineering: downstream consumers fed on Binlog under MyCat, and OceanBase has no native Binlog — the options are a tenant-level single-threaded Binlog Service (a bottleneck at peak) or OMS→Kafka (requiring a self-built Kafka-to-Oracle sync tool).
- 教训
分库分表中间件的账要在"合区、扩容、复杂SQL"三件事上一次性结算,规模越大越痛;迁移前先画全数据流转图,下游消费链路才是真实工作量;流量回放比抽样测试便宜且覆盖全;MySQL协议兼容在OLTP场景基本可用(4.5万SQL仅1坑),但执行计划、分区裁剪这类语义层差异仍需逐条治理。
Sharded-middleware debt gets settled all at once on three items — region merges, scaling, and complex SQL — and hurts more at larger scale. Map the complete data-flow graph before migrating; the downstream consumption links are the real workload. Traffic replay is cheaper and more complete than sampled testing. MySQL-protocol compatibility is broadly usable for OLTP (1 pitfall in 45,000 SQLs), but semantic-layer differences — execution plans, partition pruning — still need per-statement governance.
来源
Lu Wenhao (Database Lead, Belle Fashion), "Migrating from Sharded MyCat to OceanBase: Experience Summary and Issue Catalog of Belle's Core Finance System Migration" (CSDN blog
—
no publication date shown on page
—
相关产品:OceanBase、MySQL 相关能力:数据编码 + 两层压缩带来的迁移降本、MySQL / Oracle 双模式兼容 最后核验:2026-10-02
常熟农商行:从2018年试点到两地三中心,交易处理能力提升46倍(2018–2022) 成功经验
两地三中心
容灾
农商行
核心改造
- 场景
江苏常熟农商行2018年开始引入OceanBase,是较早试水的农商行。当时诉求很实际:传统集中式数据库在数据容量(装不下)和高并发(放不下)上同时见顶,而农商行"开门红"这类营销大促又有明显的脉冲流量。路线是先边缘后核心:业务中台、手机银行、大零售营销等近30个应用系统逐步上线,先跑起来再说。(《中国经营报》2022-08-15报道)
Changshu Rural Commercial Bank in Jiangsu began adopting OceanBase in 2018, an early mover among rural commercial banks. The motivation was practical: traditional centralized databases were hitting the ceiling on both data volume (couldn't hold it) and high concurrency (couldn't serve it), while the bank's "strong start of year" marketing campaigns generate obvious pulse traffic. The route was edge-first, core-later: nearly 30 application systems — business middle-platform, mobile banking, mass-retail marketing — went live incrementally. (China Business Journal, 2022-08-15)
- 决策
2020年起,该行借助OceanBase做两地三中心改造,把容灾能力从"机房级"提升到"数据中心级"。这是关键一跃:从"用分布式数据库跑业务"到"用分布式数据库做容灾底座",意味着把最核心的可用性命脉交了出去。
Starting in 2020, the bank rebuilt its disaster-recovery topology on OceanBase, moving from a two-site/three-center setup and lifting DR capability from "computer-room level" to "data-center level." This was the key leap: from "running business on a distributed database" to "building disaster recovery on a distributed database" — handing over the most critical availability lifeline.
- 结果
据该行金融科技总部科技运营中心负责人唐明向《中国经营报》透露:每秒交易处理能力提升46倍,批量代发处理量每分钟超20万笔,日终批处理缩短到10分钟以内。数字为银行方负责人受访口径、经《中国经营报》报道,无第三方独立复现;46倍的对比基线(相对何种架构)原文未明确,引用时需注意。
Per Tang Ming, head of the technology operations center at the bank's fintech headquarters, speaking to China Business Journal: transaction processing capacity per second up 46x, batch payroll disbursement over 200,000 transactions per minute, and end-of-day batch processing shortened to under 10 minutes. Figures are the bank executive's on-the-record statements as reported by China Business Journal; no independent third-party reproduction found. The baseline for the 46x comparison (relative to which architecture) is not specified in the original report — cite with care.
- 机制根因
两地三中心在传统架构下靠存储层复制+应用层切换,RPO/RTO都依赖人工编排,演练一次伤筋动骨。OceanBase的Paxos多副本把"三中心"做成了数据库原生语义:多数派提交即RPO=0,副本自动接管把RTO压到秒级,容灾从"演练项目"变成"日常属性"。46倍TPS提升的另一半来自架构简化:原来"集中式DB+外挂缓存/分片"的链路被单集群替代,省掉了跨系统协调。农商行先边缘后核心的渐进路线,本质是用时间换信任——近30个系统先跑两年,容灾改造才敢动手。
In traditional architectures, a three-data-center topology relies on storage-layer replication plus application-layer failover, with RPO/RTO dependent on manual orchestration — each DR drill is painful. OceanBase's Paxos multi-replica design makes "three centers" a native database semantic: majority-quorum commits mean RPO=0, automatic replica takeover pushes RTO to seconds, and disaster recovery turns from a "drill project" into a daily property. The other half of the 46x TPS gain comes from architectural simplification: the old "centralized DB plus bolted-on cache/sharding" chain was replaced by a single cluster, eliminating cross-system coordination. The bank's edge-first, core-later progression is essentially trading time for trust — nearly 30 systems ran in production for two years before the DR rebuild dared to proceed.
- 教训
容灾选型先看"容灾是不是数据库的原生语义",外挂式容灾的RTO永远受制于最慢的人工环节;中小银行的稳妥路线是"边缘先行、核心跟进",用两三年生产运行换决策信心;看"提升N倍"类数字先问基线是什么——46倍的含金量取决于分母,选型时盯住"批量代发20万笔/分钟、日终10分钟"这类绝对值工程指标。
In DR selection, first ask whether DR is a native semantic of the database — bolt-on DR's RTO is always hostage to the slowest manual step. The prudent route for small and mid-size banks is "edge first, core later," trading two to three years of production runtime for decision confidence. When reading "Nx improvement" figures, always ask what the baseline is — the 46x figure's meaning depends on its denominator; during selection, anchor on absolute engineering metrics like "200,000 payroll transactions per minute" and "10-minute end-of-day batch."
来源
《中国经营报》记者李晖《新技术与金融的双向奔赴》(2022-08-15
China Business Journal, reporter Li Hui, "The Two-Way Journey of New Technology and Finance" (2022-08-15
相关产品:OceanBase 相关能力:— 最后核验:2026-10-02
GCash:菲律宾最大电子钱包数百库零宕机迁移,存储降70%(2026) 成功经验
出海
零宕机迁移
降本
高并发
- 场景
GCash是菲律宾最大电子钱包,用户覆盖菲律宾约一半人口。原架构是MySQL集群:连接数有限,高并发时频繁不稳定——电子钱包的脉冲式流量(全民转账/发红包)下,连接池先被打满,接着雪崩。数百个MySQL库还要逐个运维,版本、备份、扩容都是N倍工作量。(新浪财经转述"星海情报局"2026-09报道)
GCash is the Philippines' largest e-wallet, covering roughly half the country's population. The original architecture was a MySQL cluster: limited connection counts, frequently unstable under high concurrency — under the pulse-style traffic of e-wallets (nationwide transfers and red-packet gifting), the connection pool fills up first and then avalanches. Hundreds of MySQL databases each needed individual operations — versions, backups, and scaling all multiplied N-fold. (Sina Finance republishing "Xinghai Intelligence" report, Sep 2026)
- 决策
换OceanBase。迁移策略是零宕机无缝迁移:数百个数据库逐个切换,业务无感知。技术负责人John证实,迁移后系统在每秒处理数百万笔交易的同时保持稳定。
Switch to OceanBase. The migration strategy was seamless zero-downtime cutover: hundreds of databases switched one by one, invisible to the business. Tech lead John confirmed the system stays stable while processing millions of transactions per second after migration.
- 结果
存储需求降低约70%,资源成本降低超40%。数百个库完成零宕机迁移。以上数字为媒体转述GCash技术负责人John的口径,无独立第三方复现;迁移耗时、切换窗口细节未披露。
Storage demand down ~70%, resource cost down over 40%. Hundreds of databases migrated with zero downtime. Figures are media-reported from GCash tech lead John; no independent third-party reproduction found. Migration duration and cutover-window details were not disclosed.
- 机制根因
MySQL集群的天花板是"连接数×单机写入"的乘积,钱包类脉冲流量的特征是瞬间并发远超均值数十倍,连接池模型在这种方差下最先崩。OceanBase的多租户+分布式架构同时解了两个结:计算上多节点分摊连接与写入,脉冲被摊薄;存储上行列混存编码+两层压缩把约70%存储直接省掉——省存储不只是省钱,更是把"数百个库"的备份、扩容窗口同步缩小。数百库零宕机迁移可行的前提是MySQL协议兼容:应用层连接串几乎不用改,才能一个库一个库地灰度切换。
A MySQL cluster's ceiling is the product of connection count and single-machine writes; wallet pulse traffic spikes instantaneous concurrency tens of times above the mean, and the connection-pool model collapses first under that variance. OceanBase's multi-tenant distributed architecture unties both knots: on compute, multiple nodes share connections and writes so pulses get diluted; on storage, hybrid row-column encoding plus two-tier compression directly eliminates ~70% of storage — saving storage is not just saving money, it also shrinks the backup and scaling windows for "hundreds of databases" in sync. The precondition for a zero-downtime migration of hundreds of databases is MySQL-protocol compatibility: application connection strings barely change, so cutover can proceed database by database.
- 教训
脉冲型业务选型的第一指标是"方差承载力"而非均值TPS;存储压缩在钱包场景是双重收益——降本+缩短运维窗口;数百库迁移不要追求"大爆炸"切换,协议兼容带来的灰度能力比迁移工具更重要;"零宕机"是结果,手段是"应用不用改+逐个切",选型时就要验证这两条。
For pulse-shaped workloads, the first selection metric is "variance capacity," not average TPS. Storage compression is a double win for wallets — lower cost plus shorter operations windows. Don't chase a "big bang" cutover for hundreds of databases; the canary capability from protocol compatibility matters more than migration tooling. "Zero downtime" is the outcome; the means are "no application changes plus one-at-a-time cutover" — verify both during selection.
来源
新浪财经转述"星海情报局"《千万东南亚人的日常交易:谁在改写亚太数据库的竞争坐标系?》(2026-09-07
Sina Finance republishing "Xinghai Intelligence", "Millions of Southeast Asians' Daily Transactions: Who Is Redrawing APAC's Database Competitive Map?" (2026-09-07
相关产品:OceanBase、MySQL 相关能力:数据编码 + 两层压缩带来的迁移降本 最后核验:2026-10-02
TNGD:马来西亚最大电子钱包换掉底层数据库,同等硬件吞吐提升40%(2026) 成功经验
出海
高并发
强一致
数字支付
- 场景
TNGD(Touch 'n Go eWallet)是马来西亚最大电子钱包,2600万用户,承载着马来西亚85%成年人口的日常支付(媒体口径)。导火索是一次全国性刺激计划:数百万用户几乎同时涌入,流量激增、响应变慢,"日常运行良好的架构开始吃力"。技术负责人事后说,那是团队第一次意识到"当前的设计、架构和技术平台远远低于目标预期"。需求很明确:扛住极端峰值的高并发、每笔交易强一致、零宕机。(新浪财经转述"星海情报局"2026-09报道)
TNGD (Touch 'n Go eWallet) is Malaysia's largest e-wallet with 26 million users, carrying the daily payments of 85% of Malaysia's adult population (media figure). The trigger was a nationwide stimulus program: millions of users flooded in nearly simultaneously, traffic surged and response times degraded — "an architecture that ran fine day-to-day started to strain." The tech lead later said it was the first time the team realized "the current design, architecture, and technology platform fell far short of target expectations." The requirements were clear: survive extreme peak concurrency, strong consistency per transaction, zero downtime. (Sina Finance republishing "Xinghai Intelligence" report, Sep 2026)
- 决策
摆在桌上的候选包括亚马逊Aurora、谷歌Spanner等"更熟悉"的全球厂商,TNGD最终选了OceanBase。选型逻辑是场景同构:TNGD的问题(极高并发+强一致+零宕机)正是OceanBase在支付宝双十一里打磨了十年的拿手场景。TNGD技术团队事后用"零额外学习成本"形容迁移的顺滑。
Candidates on the table included more familiar global vendors such as Amazon Aurora and Google Spanner; TNGD ultimately chose OceanBase. The selection logic was scenario isomorphism: TNGD's problem (extreme concurrency + strong consistency + zero downtime) is exactly the scenario OceanBase spent a decade honing through Alipay's Double-11 festivals. TNGD's tech team later described the migration as having "zero additional learning cost."
- 结果
同等硬件规格下吞吐提升40%,系统在4万笔/秒压力测试下零宕机升级。TNGD高管Leslie Lip称:"OceanBase不仅仅是一个数据库,而是一个能够随着公司野心一起成长的技术平台。"以上数字与引言均为媒体转述口径,无独立第三方复现;迁移耗时、迁移前后延迟对比等数字未披露。
Throughput up 40% on identical hardware specs, with a zero-downtime upgrade demonstrated under a 40,000 TPS stress test. TNGD executive Leslie Lip: "OceanBase is not just a database, but a technology platform that can grow with the company's ambitions." Figures and quotes are media-reported; no independent third-party reproduction found. Migration duration and before/after latency comparisons were not disclosed.
- 机制根因
全国性刺激计划是典型的"热点+洪峰"叠加:连接数、热点账户行锁、写放大三重挤压。传统路线要么靠中间件分片(把一致性推给业务层),要么靠单机升配(天花板明显)。OceanBase的分布式多写+Paxos强一致把"峰值吸收"做进内核:多节点分摊写入、多数派提交保证强一致、在线扩缩容应对脉冲。40%提升的另一半来自少了一层中间件——分片逻辑下沉进数据库后,省掉了路由层的网络与协调开销。
A nationwide stimulus program is a classic hotspot-plus-flood stack: connection counts, hotspot account row locks, and write amplification squeezing together. The traditional routes are middleware sharding (pushing consistency onto the application layer) or vertical single-machine upgrades (a low ceiling). OceanBase's distributed multi-writer design plus Paxos strong consistency bakes "peak absorption" into the kernel: multiple nodes share the write load, majority-quorum commits guarantee strong consistency, and online scaling absorbs pulses. Half of the 40% gain comes from removing a middleware layer — once sharding logic sinks into the database, the routing layer's network and coordination overhead disappears.
- 教训
选型先找"场景同构"的验证场:TNGD选OceanBase不是看功能列表,而是看双十一这种同构极端场景的十年实战记录;"零额外学习成本"的另一面是团队已有MySQL心智,兼容性把迁移成本压到最低;出海选型里"全球大厂更熟悉"不等于"更合适",峰值形态匹配度才是第一优先级。
In selection, first find a "scenario-isomorphic" proving ground: TNGD chose OceanBase not from a feature list but from a decade of real combat records in the isomorphic extreme scenario of Double-11. The flip side of "zero additional learning cost" is a team already fluent in MySQL — compatibility pushes migration cost to the floor. In overseas selection, "more familiar global vendor" does not equal "better fit"; peak-shape match is the first priority.
来源
新浪财经转述"星海情报局"《千万东南亚人的日常交易:谁在改写亚太数据库的竞争坐标系?》(2026-09-07
Sina Finance republishing "Xinghai Intelligence", "Millions of Southeast Asians' Daily Transactions: Who Is Redrawing APAC's Database Competitive Map?" (2026-09-07
相关产品:OceanBase、Amazon Aurora、Google Spanner 相关能力:— 最后核验:2026-10-02
Pinterest 的 MySQL 分片:从 NoSQL 废墟退回成熟技术(2012) 成功经验
水平分片
ID 设计
成熟技术
去 NoSQL
- 场景
2011 年 Pinterest 爆发式增长,基础设施全线过载。团队试了 Cassandra、Membase、MongoDB 等多个 NoSQL,"全部以灾难性方式挂掉"(Marty Weiner 原文);MySQL 读备库则带来成堆的缓存与复制延迟 bug。(Pinterest 工程博客,2015)
In 2011 Pinterest's explosive growth overloaded every piece of infrastructure. The team tried several NoSQL systems — Cassandra, Membase, MongoDB — "all of which eventually broke catastrophically" (Marty Weiner); MySQL read replicas, meanwhile, produced piles of caching and replication-lag bugs. (Pinterest Engineering Blog, 2015)
- 决策
退回 MySQL,做应用层分片:8 台 EC2 起步,每台主主复制做热备,生产只读写主库("永远不要在生产环境读写备库");4096 个虚拟分片,配置放 ZooKeeper;64 位 ID 编码分片信息(16 位分片 ID + 10 位类型 + 36 位本地 ID),任何服务解析 ID 即可路由,无需查路由表;对象存 JSON blob,关系用单向映射表,JOIN 搬到应用层。
Retreat to MySQL with application-layer sharding: start with 8 EC2 instances, each master-master replicated for hot standby, production reading and writing only the master ("You never want to read/write to a slave in production"); 4,096 virtual shards with config in ZooKeeper; 64-bit IDs encoding shard info (16-bit shard ID + 10-bit type + 36-bit local ID) so any service can route by parsing the ID — no routing-table lookup; objects stored as JSON blobs, relationships in unidirectional mapping tables, JOINs moved to the application layer.
- 结果
2012 年初上线,至博客发表时(2015)稳定运行 3.5 年,作者称"可能会永远跑下去";支撑当时 500 亿+ Pin;三年里只做过一次 ALTER。
Launched in early 2012 and running stably for 3.5 years by the time of the blog post (2015); the author wrote it would "likely be in there forever"; it supported 50B+ pins at the time, with exactly one ALTER TABLE in three years.
- 机制根因
Pinterest 的选择逻辑是"故障模式优先":NoSQL 的故障是"丢数据、脑裂"(不可接受),MySQL 单机故障是"可预测、可修复"(主主切换+备库)。分片把 MySQL 的故障域切小,而 ID 编码分片信息消灭了路由这一单点依赖——这是整个设计最精妙处:路由信息内嵌在数据标识里,分片配置只在扩容时变化。代价是放弃跨分片 JOIN、外键与二级索引一致性,以及分片键一旦选定几乎不可更改(靠"整片搬迁"而非"逐行重分布"来扩容)。
Pinterest's decision logic was "failure mode first": NoSQL failures meant "data loss, split brain" (unacceptable), while single-MySQL failures were "predictable and repairable" (master-master failover + standby). Sharding shrank MySQL's failure domains, and encoding shard info into IDs eliminated routing as a single point of dependency — the most elegant part of the design: routing information embedded in the data identifier itself, with shard config changing only on scale-out. The price: giving up cross-shard JOINs, foreign keys, and secondary-index consistency — and the shard key is effectively immutable once chosen (scale by moving whole virtual shards, not row-by-row redistribution).
- 教训
选型时先问"它坏的时候是什么死法",再问"它快不快"。2011 年 NoSQL 的运维成熟度撑不起 Pinterest 的增长,而"无聊但可靠"的 MySQL + 精心设计的 ID,分片后稳定跑了十年。过早追新技术的税,最终都是用数据丢失来交。
When choosing a database, ask "how does it die" before "how fast is it." In 2011, NoSQL operational maturity couldn't carry Pinterest's growth, while "boring but reliable" MySQL plus a well-designed ID scheme ran stably for a decade after sharding. The tax on chasing new technology too early is ultimately paid in lost data.
相关产品:MySQL 相关能力:— 最后核验:2026-10-01
富友支付:高并发交易上 PolarDB,数据库整体成本大降、性能明显提升 成功经验
支付交易
高并发
成本优化
- 场景
富友支付是持牌第三方支付机构,高并发交易与海量数据给数据库带来三重压力:性能瓶颈、扩展性差、运维复杂。
Fuyou Payment is a licensed third-party payment institution. High-concurrency transactions and massive data put threefold pressure on its databases: performance bottlenecks, poor scalability, and complex operations.
- 决策
做云原生架构升级,引入云原生数据库 PolarDB(MySQL 版)。
Upgrade to a cloud-native architecture with the cloud-native database PolarDB (MySQL-compatible edition).
- 结果
官方客户案例口径:数据库整体成本大幅下降,性能明显提升。(阿里云官网产品页客户案例栏,属厂商渠道口径;未披露具体数字与对比基线,引用时须注明。)
Per the official customer story: overall database cost dropped substantially and performance improved markedly. (Vendor-channel claim from the Alibaba Cloud product page's customer stories section; no specific figures or comparison baseline disclosed - cite with that caveat.)
- 机制根因
支付交易是典型的"高并发写入叠加读写混合"负载:PolarDB 存算分离下只读节点可接近线性扩展(取决于复制延迟与读一致性要求),分担查询压力,主节点专注写入;分钟级弹性应对营销活动、节假日的交易洪峰,避免按峰值常年持有资源。
Payment traffic is a classic "high-concurrency writes plus mixed read/write" workload: under PolarDB's storage-compute disaggregation, read replicas scale out near-linearly to absorb query pressure (depending on replication lag and read-consistency requirements) while the primary focuses on writes; minute-level elasticity absorbs transaction floods from marketing campaigns and holidays, avoiding year-round peak provisioning.
- 教训
支付机构选型云原生数据库,TCO 核算要把"弹性节省下来的峰值预留"算进去,而不只对比单价;官方案例未给数字时,诚实标注"未披露"比转述模糊表述更有价值。
When a payment company evaluates cloud-native databases, TCO math must include the peak headroom that elasticity eliminates, not just unit pricing. When the official story gives no numbers, honestly marking "undisclosed" is more valuable than paraphrasing vague claims.
相关产品:PolarDB 相关能力:共享存储一写多读 —— 阿里云版的"日志即数据库" 最后核验:2026-10-02
MiniMax:AI 独角兽用 PolarDB Limitless 扛住 2.36 亿用户的潮汐流量 成功经验
AI 数据底座
弹性扩缩容
成本优化
- 场景
MiniMax(稀宇科技)是全球领先的 AI 公司,自研全模态大模型,旗下有海螺 AI、星野等产品。星野等 C 端陪伴应用带来海量多模态数据与明显的潮汐流量:高峰期对话写入暴涨、低谷期迅速回落,按峰值常年持有数据库资源的成本极高。
MiniMax is a global AI company building its own full-modality models, with consumer products including Hailuo AI and Talkie. Companion-style AI apps generate massive multimodal data and sharp tidal traffic: conversation writes spike at peak hours and collapse off-peak, making peak-provisioned database capacity prohibitively expensive.
- 决策
基于阿里云 PolarDB Limitless 构建智能数据底座,承载核心业务数据。
Build the intelligent data substrate on Alibaba Cloud PolarDB Limitless for core business data.
- 结果
官方客户案例口径:千亿级对话表性能提升 3 倍、秒级弹性扩缩容、存储成本下降 75%,支撑 2.36 亿用户高效服务。(以上数字来自阿里云官网客户案例栏,属厂商渠道口径,未披露测试方法;2.36 亿用户数与 MiniMax 港股公告"截至 2025 年 12 月 31 日累计服务逾 2.36 亿名用户"一致。)
Per the official customer story on aliyun.com: 3x performance on hundred-billion-row conversation tables, second-level elastic scaling, 75% lower storage cost, serving 236 million users efficiently. (All figures are vendor-channel claims from the Alibaba Cloud product page; no test methodology disclosed. The 236M user figure is consistent with MiniMax's HKEX announcement of "over 236 million users served as of Dec 31, 2025".)
- 机制根因
Limitless 是 PolarDB 存算分离架构的弹性扩展形态:存储层统一、计算节点可按需秒级扩缩,潮汐流量来时加节点、退潮时释放,成本只为实际使用的算力买单;共享存储也避免了传统分库分表每次扩容都要做的数据重分布。
Limitless is the elastic-scaling form of PolarDB's disaggregated storage-compute architecture: a unified storage layer with compute nodes that scale out and in within seconds. Add nodes when the tide comes in, release them when it recedes - you pay only for the compute actually used. Shared storage also avoids the data redistribution that classic sharding forces on every scale-out.
- 教训
AI 陪伴/对话类业务的数据库选型,第一优先级不是峰值 TPS,而是"弹性速度 × 成本曲线":写放大严重的对话表叠加潮汐流量,存算分离加秒级弹性的收益远大于单纯堆实例规格;但"性能提升 3 倍"类数字是厂商口径,引用时必须标注来源并说明测试条件缺失。
For AI companion and conversation workloads, the first selection criterion is not peak TPS but "elasticity speed x cost curve": write-amplified conversation tables plus tidal traffic reward storage-compute disaggregation with second-level elasticity far more than simply upsizing instances. But "3x performance" style figures are vendor claims - always cite the source and note the missing test conditions.
相关产品:PolarDB 相关能力:共享存储一写多读 —— 阿里云版的"日志即数据库" 最后核验:2026-10-02
欧派家居:从 Oracle 迁到 PolarDB,部分 SQL 比 Oracle 快 3 到 5 倍 成功经验
去 O 迁移
家居零售
成本优化
- 场景
欧派家居的系统架构对 Oracle 生态高度依赖:存储过程、PL/SQL 写法深入业务代码,Oracle 许可与运维成本随业务扩张不断走高。
Oppein Home's system architecture was deeply dependent on the Oracle ecosystem: stored procedures and PL/SQL idioms ran through its business code, while Oracle licensing and operations costs kept climbing with business growth.
- 决策
上云并迁移到 PolarDB(PostgreSQL 版,高度兼容 Oracle 语法),用其多读架构与云计算算力替换 Oracle。
Move to the cloud on PolarDB (PostgreSQL edition, highly Oracle-compatible), replacing Oracle with its multi-reader architecture and cloud compute.
- 结果
官方客户案例口径:享受云计算的高效算力,部分 SQL 执行速度比 Oracle 快 3 至 5 倍,整体业务效率大幅提升。(阿里云官网产品页客户案例栏,属厂商渠道口径;"部分 SQL"未说明占比与类型,引用时须注明。)
Per the official customer story: enjoying the cloud's efficient compute, some SQL runs 3 to 5 times faster than on Oracle, with overall business efficiency greatly improved. (Vendor-channel claim from the Alibaba Cloud product page's customer stories section; the share and type of "some SQL" are unspecified - cite with that caveat.)
- 机制根因
PolarDB-PG 的 Oracle 兼容模式降低了 PL/SQL 改写成本;一写多读架构把原来压在 Oracle 单机上的报表、查询类负载分流到只读节点,主库只跑交易——很多"比 Oracle 快"的案例本质是架构红利,而非单机引擎碾压。
PolarDB-PG's Oracle compatibility mode lowers PL/SQL rewrite cost; the single-writer/multi-reader architecture diverts the reporting and query workloads that used to sit on one Oracle box onto read replicas, leaving the primary for transactions only. Much of the "faster than Oracle" story is architectural dividend, not single-engine domination.
- 教训
去 O 选型的真实账本有三栏:许可成本、改造成本、架构红利;"部分 SQL 快 3 到 5 倍"这类表述必须追问是哪部分、占比多少,否则无法作为选型依据。
A de-Oracle business case has three columns: license cost, migration cost, and architectural dividend. Claims like "some SQL 3-5x faster" must be followed up with "which SQL, and what share" - otherwise they cannot serve as selection evidence.
相关产品:PolarDB、Oracle Database(甲骨文) 相关能力:— 最后核验:2026-10-02
数云:天猫 CRM 服务商用 PolarDB,把双十一升配从 8 小时压到 20 分钟 成功经验
大促弹性
多租户 SaaS
高并发
- 场景
杭州数云信息技术有限公司(数云)是面向电商的 CRM SaaS 服务商,多租户部署在公有云上,上百个 PolarDB 实例,单实例数据量 2 到 3TB。双十一期间单集群订单量超 4 亿,数据库要扛住交易洪峰又不能提前数月锁定昂贵规格。
Hangzhou Shuyun Information Technology (Shuyun) is a CRM SaaS provider for e-commerce, multi-tenant on the public cloud, with over a hundred PolarDB instances at 2 to 3TB per instance. During Double 11 a single cluster handles over 400 million orders; the database must survive the transaction flood without locking in expensive peak specs months ahead.
- 决策
核心业务从传统 MySQL 迁移到 PolarDB,利用其分钟级弹性升配应对大促。
Move core business from traditional MySQL to PolarDB, using its minute-level elastic scale-up for the shopping festival.
- 结果
据阿里云官方案例库(2022 年 5 月版,数云自述):传统 MySQL 升配最大实例需要 6 到 8 小时,而 PolarDB 节点升配只需 10 到 20 分钟、增加节点只需 5 到 8 分钟,20 分钟内即可完成 10TB 级数据集群的升配;32 并发以上的 OLTP 写能力达到普通 MySQL 的 2 到 3 倍;单实例最大支持 100TB 存储。(官方渠道口径及客户证言;压测条件未披露。)
Per Alibaba Cloud's official case library (May 2022 edition, in Shuyun's own words): scaling up the largest traditional MySQL instance took 6 to 8 hours, while PolarDB upgrades a node in 10 to 20 minutes and adds a node in 5 to 8 minutes - a 10TB data cluster finishes scale-up within 20 minutes; OLTP write throughput above 32 concurrency reaches 2 to 3 times that of ordinary MySQL; a single instance supports up to 100TB of storage. (Official-channel claims plus customer testimony; benchmark conditions undisclosed.)
- 机制根因
PolarDB 的计算与存储分离:升配只换计算节点规格,数据不动(不像传统 MySQL 升配要迁移数据文件);只读节点走存储层物理复制,分钟级拉起;大促前一两天做弹性升级、大促后降配,成本只为洪峰那几天买单。
PolarDB separates compute from storage: scaling up only swaps compute-node specs while data stays put (unlike traditional MySQL scale-up, which migrates data files); read replicas use physical replication at the storage layer and spin up in minutes. Scale up a day or two before Double 11, scale down after - you pay for peak capacity only during the flood.
- 教训
大促型业务选数据库,"升配速度"和"加节点速度"是跟 TPS 同等重要的硬指标:传统方案 6 到 8 小时的升配窗口逼着你提前数月买峰值规格,分钟级弹性把容量规划从"猜峰值"变成"看板调参";客户证言里"双十一前一两天做弹性升级,双十一期间 IOPS 很稳定,连接数只用到当前规格的一半"正是这种工作流的写照。
For festival-driven businesses, "scale-up speed" and "add-node speed" are hard metrics on par with TPS: a 6-to-8-hour scale-up window forces you to buy peak specs months early, while minute-level elasticity turns capacity planning from "guess the peak" into "tune the dashboard". Shuyun's testimony - "we did the elastic upgrade a day or two before Double 11; IOPS stayed stable through the festival and connections only reached half of the provisioned spec" - is exactly that workflow in action.
相关产品:PolarDB、MySQL 相关能力:共享存储一写多读 —— 阿里云版的"日志即数据库" 最后核验:2026-10-02
心动网络:爆款手游用 PolarDB,支撑百万级玩家同时在线 成功经验
游戏出海
高并发
高可用
- 场景
心动网络(中国互联网百强企业,前身为 VeryCD)业务覆盖游戏研发运营与 TapTap 游戏社区全球化运营。游戏出海需要在国内、东南亚、欧美统一部署;活动峰值时需支撑 100 万级玩家同时在线;游戏版本发布、服务端软硬件故障重启时,需要数据库快速恢复读取能力。
XD Inc. (a Top-100 Chinese internet company, formerly VeryCD) runs game development and operations plus the global TapTap gaming community. Going global requires unified deployment across China, Southeast Asia, Europe and the US; event peaks must support a million concurrent players; and game releases or server hardware/software restarts demand fast recovery of database read capacity.
- 决策
采用 PolarDB 云原生数据库方案(一主一读集群)构建全部业务系统。
Build all business systems on the PolarDB cloud-native database (one-primary-one-replica clusters).
- 结果
据阿里云官方案例库(2022 年 5 月版):PolarDB 为千万级用户在线手游保驾护航;所有实例一主一读,性能为 MySQL 的 3 倍;100% 兼容 MySQL 5.6/8.0 及生态工具,业务无缝迁移;主实例故障时 30 到 60 秒内完成切换,数据三副本一致性存储。(官方渠道口径;"3 倍性能"未披露测试条件。)
Per Alibaba Cloud's official case library (May 2022 edition): PolarDB safeguards online mobile games with tens of millions of users; every instance runs one primary plus one read replica with 3x MySQL performance; 100% compatibility with MySQL 5.6/8.0 and ecosystem tools enabled seamless migration; primary failover completes in 30 to 60 seconds with three-replica consistent storage. (Official-channel claims; the "3x performance" test conditions are undisclosed.)
- 机制根因
存算分离加一写多读:读流量由只读节点分担,服务端重启后"惊群式"的数据加载被只读节点吸收,主库不被打垮;三副本物理复制保证 RPO 为 0,故障切换只做主备角色切换而非数据重放,所以能压到 30 到 60 秒。
Disaggregated storage-compute plus single-writer/multi-reader: read replicas absorb query traffic and soak up the "thundering herd" data loading after server restarts, so the primary never gets crushed; three-replica physical replication keeps RPO at zero, and failover is just a primary/standby role switch rather than data replay - which is why it fits in 30 to 60 seconds.
- 教训
游戏开服、合服、版本发布这类"脉冲式"负载,数据库选型的关键指标是"故障与重启后的恢复速度",而不只是稳态 TPS;MySQL 兼容度决定迁移成本,心动的"无缝迁移"建立在 100% 协议兼容之上,换引擎前先做 SQL 兼容性扫描。
For "pulse-shaped" loads like game launches, server merges and version releases, the key selection metric is recovery speed after failure and restart, not just steady-state TPS. MySQL compatibility determines migration cost - XD's "seamless migration" rested on 100% protocol compatibility, so run a SQL compatibility scan before switching engines.
相关产品:PolarDB、MySQL 相关能力:共享存储一写多读 —— 阿里云版的"日志即数据库" 最后核验:2026-10-02
小鹏汽车:PolarDB for PostgreSQL 扛起智驾业务每日 TB 级大表更新 成功经验
智能驾驶
大表优化
HTAP
- 场景
小鹏汽车智能辅助驾驶业务每天产生 TB 级数据:车辆回传的感知、轨迹数据汇入大表,既要支撑每天 7000 万行级别的数据更新,又要做秒级分析查询。社区版 PostgreSQL 在大表上的查询与并发更新慢,成为瓶颈。
XPeng's intelligent assisted-driving business generates terabytes of data daily: perception and trajectory data uploaded by vehicles flows into huge tables that must absorb around 70 million row updates per day while also serving second-level analytical queries. Community PostgreSQL was too slow on large-table queries and concurrent updates, becoming the bottleneck.
- 决策
采用 PolarDB PostgreSQL 版,用其大表优化与弹性跨机并行查询(ePQ)能力承载智驾数据。
Adopt PolarDB for PostgreSQL, using its large-table optimization and elastic parallel query across nodes (ePQ) for the driving data.
- 结果
官方客户案例口径:在小鹏智能辅助驾驶业务上实现每日 TB 级大数据表的 7000 万行更新和大数据表秒级分析查询。(阿里云官网产品页客户案例栏,属厂商渠道口径;未披露硬件规格与对比基线。)
Per the official customer story: 70 million row updates per day on terabyte-scale tables plus second-level analytical queries on those tables in XPeng's assisted-driving business. (Vendor-channel claim from the Alibaba Cloud product page's customer stories section; hardware specs and comparison baseline undisclosed.)
- 机制根因
ePQ 把单机 PG 难以并行的大表扫描与聚合拆到多计算节点并行执行;行列混存与大表优化降低 TB 级宽表的更新与查询代价——同一份数据既做事务更新又做实时分析,是 HTAP 能力的直接体现。
ePQ splits large-table scans and aggregations that single-node PostgreSQL cannot parallelize across multiple compute nodes; hybrid row-column storage and large-table optimization cut the update and query cost on terabyte-scale wide tables - transactional updates and real-time analytics on the same data, a direct expression of HTAP capability.
- 教训
智驾、IoT 这类"每天 TB 级写入叠加即席分析"的负载,选型时要把"大表更新吞吐"和"分析查询延迟"放在同一张表里评估,纯 OLTP 或纯 OLAP 产品都会瘸腿;厂商案例数字缺基线时,只转述、不演绎。
For workloads like assisted driving and IoT - "terabytes written daily plus ad-hoc analytics" - evaluate "large-table update throughput" and "analytical query latency" on the same scorecard; pure-OLTP or pure-OLAP products will both limp. When vendor case numbers lack a baseline, paraphrase, don't extrapolate.
相关产品:PolarDB 相关能力:— 最后核验:2026-10-02
长沙营智:易撰资讯搜索 DRDS+PolarDB 双引擎,5TB 级数据秒级检索 成功经验
资讯搜索
混合架构
HTAP
- 场景
长沙营智信息科技有限公司旗下"易撰"网是资讯搜索平台:海量文章、视频资讯入库,每天产生大量业务数据;业务特点是写入并发高,同时又要支持大范围时间维度的多维复杂查询与统计——单一引擎两头都吃力。
Changsha Yingzhi Information Technology's "Yizhuan" is a news search platform: massive volumes of articles and video news are ingested daily, generating huge business data. The workload combines high-concurrency writes with wide time-range, multi-dimensional complex queries and statistics - one engine struggles at both ends.
- 决策
写链路用 DRDS 加 RDS 承载高并发写入,查询链路用 DRDS 加 PolarDB 承载大范围多维复杂查询;PolarDB 的分布式存储与高性能成为大范围时间查询的关键。
Use DRDS plus RDS for the high-concurrency write path, and DRDS plus PolarDB for the wide-range multi-dimensional complex query path; PolarDB's distributed storage and high performance became the key to wide time-range queries.
- 结果
据阿里云官方案例库(2022 年 5 月版,长沙营智技术总监刘涛证言):PolarDB 的海量存储与高性能满足 5TB 到 10TB 级数据存储,完全满足其业务的大数据量存储需求;PolarDB 满足了复杂大范围查询的需求,同时还能保证事务支持。(官方渠道口径;未披露具体查询延迟数字。)
Per Alibaba Cloud's official case library (May 2022 edition, testimony by Yingzhi CTO Liu Tao): PolarDB's massive storage and high performance cover 5TB to 10TB of data storage, fully meeting the business's big-data storage needs; PolarDB satisfied the complex wide-range query requirements while still guaranteeing transactional support. (Official-channel claims; specific query latency figures undisclosed.)
- 机制根因
典型的"写走分库分表、查走云原生数仓化"分工:DRDS 做分片路由扛写入,PolarDB 的共享存储层天然就是一份完整的海量数据,计算节点做大范围扫描与聚合;事务支持让"查询结果可写回、可对账",这是纯 OLAP 引擎给不了的。
A classic division of labor - "writes go to sharding, queries go to cloud-native storage": DRDS does shard routing for writes, while PolarDB's shared storage layer is inherently one complete copy of the massive dataset, with compute nodes doing wide-range scans and aggregations. Transactional support means "query results can be written back and reconciled" - something pure OLAP engines cannot offer.
- 教训
资讯、内容类"高并发写入叠加大范围时间查询"的业务,一个引擎打天下往往两头不讨好;按读写链路拆分引擎时,查询侧选型要同时考核"海量存储成本"与"复杂查询是否带事务",缺一不可。
For news and content businesses with "high-concurrency writes plus wide time-range queries", one engine for everything usually fails at both ends. When splitting engines by read/write path, the query side must be evaluated on both "massive-storage cost" and "whether complex queries come with transactions" - both are non-negotiable.
相关产品:PolarDB、MySQL 相关能力:共享存储一写多读 —— 阿里云版的"日志即数据库" 最后核验:2026-10-02
韵达快递:客户管家跑在 PolarDB 分布式版上,QPS 峰值近 10 万 成功经验
物流
分布式数据库
高并发
- 场景
韵达快递的"客户管家"是其核心业务场景之一:快递网点、客户服务的在线业务对数据库的并发与延迟都很敏感,大促期间流量洪峰明显。
Yunda Express's "Customer Butler" is one of its core business scenarios: the online business of courier outlets and customer service is sensitive to both database concurrency and latency, with clear traffic floods during shopping festivals.
- 决策
把客户管家作为首个核心业务场景,运行在 PolarDB 分布式版上。
Run Customer Butler - the first core business scenario - on PolarDB Distributed Edition.
- 结果
官方客户案例口径:自上线以来数据库运行平稳,整体 QPS 峰值接近 10 万,SQL 响应时间稳定在 5 毫秒以内,很好地支撑了整个管家平台的平稳运行。(阿里云官网产品页客户案例栏,属厂商渠道口径;未披露数据规模与节点数。)
Per the official customer story: the database has run steadily since launch, with overall QPS peaking near 100,000 and SQL response time held within 5 milliseconds, supporting the Butler platform's smooth operation. (Vendor-channel claim from the Alibaba Cloud product page's customer stories section; data volume and node count undisclosed.)
- 机制根因
分布式版通过分片把写入压力打散到多节点,QPS 随节点数近似线性增长(取决于分片键均匀度与跨分片查询比例);SQL 层做分布式优化与路由下推,压住跨分片查询的延迟——这是"接近 10 万 QPS"与"5 毫秒响应"同时成立的前提。
The distributed edition spreads write pressure across nodes via sharding, so QPS grows roughly linearly with node count (depending on shard-key uniformity and cross-shard query ratio); the SQL layer does distributed optimization and route pushdown to keep cross-shard query latency down - the precondition for "near-100K QPS" and "5ms response" to hold at the same time.
- 教训
物流、电商这类核心交易链路上分布式数据库,验收指标要同时看吞吐(QPS 峰值)与尾延迟(响应时间),只谈其一都是耍流氓;首个核心业务上分布式版之前,先做分片键设计评审,避免上线后跨分片查询拖垮延迟。
When putting a distributed database on a core transaction path like logistics or e-commerce, acceptance criteria must cover both throughput (peak QPS) and tail latency (response time); quoting only one is misleading. Before the first core business goes on the distributed edition, review the sharding-key design, or cross-shard queries will drag latency down after launch.
相关产品:PolarDB 相关能力:— 最后核验:2026-10-02
Flipkart:信任与安全团队用 Qdrant 把欺诈检出从 9 小时压到 1 分钟 成功经验
实时风控
多模态相似检索
HBase 替代
高维向量
- 场景
Flipkart 信任与安全团队负责平台反欺诈,需要在客户与商家提交的数据(尤其是图片)上做大规模相似检索,识别重复退货、虚假商家索赔等模式。旧方案是 HBase + 局部敏感哈希(LSH):能跑批处理,但跟不上实时反欺诈节奏——在历史数据里找相似图片最慢要 9 小时;且生产模型的向量维度高达 2048,高维索引压力大。
Flipkart's Trust & Safety team fights platform fraud and abuse, running large-scale similarity searches over customer- and seller-submitted data — especially images — to spot patterns like repeat returns or duplicate seller claims. The old approach, HBase with locality-sensitive hashing (LSH), worked for batch analysis but could not keep up with real-time fraud prevention: finding similar images in historical data could take up to nine hours; and production embedding models produce 2048-dimensional vectors, adding indexing pressure.
- 决策
多个开源向量数据库做 POC 后选 Qdrant:官方 Debian 包契合 Flipkart 内部基础设施(部署灵活性);高效的 HNSW 索引能同时处理读写;支持高维向量。随后建成多租户相似检索服务,覆盖实时图片相似度欺诈检出、非结构化地址聚类(改善最后一公里配送路由)、内部 GenAI 的 RAG 检索层。
After proof-of-concept evaluations of multiple open-source vector databases, the team chose Qdrant: official Debian packaging fit Flipkart's internal infrastructure (deployment flexibility); efficient HNSW indexing handling simultaneous reads and writes; support for high-dimensional embeddings. They then built a multi-tenant similarity service covering real-time image-similarity fraud detection, unstructured address clustering (improving last-mile delivery routing), and a RAG retrieval layer for internal GenAI initiatives.
- 结果
从批处理转向实时检索后,欺诈检出时间从 9 小时降到 1 分钟以内(Flipkart 工程师 Sourabh Sarkar 口述,经 Qdrant 官方博客发布,厂商渠道口径);"过去批处理要几小时的事,现在一分钟内搞定——这对在欺诈影响客户之前拦截至关重要"(同上)。团队正把 K8s 部署的 Qdrant 标准化为全公司各团队的 embedding 存储。
Moving from batch to real-time search cut fraud detection time from nine hours to under one minute (stated by Flipkart engineer Sourabh Sarkar, via Qdrant official blog — vendor-channel source); "What used to take hours in our old batch workflows can now be done in under a minute. That change has been crucial in stopping fraud before it impacts customers." (same source). The team is standardizing Kubernetes-based Qdrant as the embedding store across groups company-wide.
- 机制根因
LSH + HBase 的近似检索是"离线批"的基因,无法满足"写入即查"的实时风控;Qdrant 的 HNSW 支持读写并发,图片上传后立即可被相似检索命中;2048 维高维向量下 HNSW 的内存/索引效率是 POC 胜出的硬指标。多租户服务把一次选型复用到欺诈、地址聚类、RAG 三个场景,摊薄了引入新组件的成本。
LSH-on-HBase approximate search is batch by DNA — it cannot serve write-then-query real-time fraud workflows; Qdrant's HNSW supports concurrent reads and writes, so an uploaded image is immediately matchable by similarity search; HNSW memory/index efficiency at 2048 dimensions was the hard metric that won the POC. The multi-tenant service amortizes one selection decision across fraud, address clustering, and RAG — three scenarios sharing the cost of a new component.
- 教训
风控类场景的选型第一指标是"从事件发生到可检出的时间",不是向量基准的 QPS;高维向量(2000+ 维)一定要在 POC 里用真实模型维度压测,很多库在 768 维的漂亮数字到 2048 维会变脸;能复用到多团队的组件才值得引入——Flipkart 把一次选型做成平台服务,是中小团队也该学的"选型摊销"思路。
For fraud-prevention scenarios, the primary selection metric is "time from event to detectability," not vector-benchmark QPS; always POC with real model dimensions for high-dimensional vectors (2000+ dims) — many engines' pretty 768-dim numbers change face at 2048; only components reusable across teams are worth adopting — Flipkart turning one selection into a platform service is the amortization mindset smaller teams should copy.
来源
Qdrant official blog, "Building real-time multimodal similarity search in Flipkart Trust & Safety with Qdrant" (vendor-channel source, named engineer Sourabh Sarkar, SDE-III, Trust & Safety at Flipkart)
https://qdrant.tech/blog/case-study-flipkart/
相关产品:Qdrant 相关能力:1M–100M 区间的性价比之王——Rust 单机的高 QPS 与低延迟 最后核验:2026-10-02
Garden:专利情报创业公司的 filterable HNSW 选型——语料从 2000 万扩到 2 亿+ 成功经验
可过滤 HNSW
标量量化
两次迁移
成本下降 10 倍
- 场景
Garden 是纽约专利情报创业公司,用大规模 AI 分析全球专利语料(2 亿+ 专利)加 TB 级真实世界数据。单件专利可达上百页、携带约 2000 个元数据字段(管辖区、授权日、专利族 ID、权利要求依赖等);每件专利切成语义块后产生"数亿"向量。工程需求:向量检索必须支持外科手术级过滤(国家×日期范围×技术标签的任意组合)。
Garden is a New York patent-intelligence startup applying large-scale AI to the global patent corpus (200M+ patents) plus terabytes of real-world data. A single patent can run 100+ pages and carries roughly 2,000 metadata fields (jurisdiction, grant date, family ID, claim dependencies); each patent is chunked into semantically meaningful pieces, producing "many hundreds of millions" of vectors. The engineering requirement: vector search with surgical-grade filtering across arbitrary combinations of country, date range, and technology tags.
- 决策
经历了两次迁移。第一版用全托管向量服务:几十 GB 数据每月约 $5000,且没有原生 filterable HNSW,被迫为每个"国家×日期×技术标签"组合建独立索引,还看不到基础设施内部,排障又慢又贵。第二版迁到自托管开源替代:省了钱,但两人团队要自己 on-call、工作时间升级,且过滤能力的老问题还在。看到 Qdrant 的 filterable HNSW 博客后第三次迁移:可过滤 HNSW 是决定因素,Qdrant Cloud 的托管 Rust 底座解决了 24×7 运维;8-bit 标量量化让热向量驻内存、冷向量落盘;周末一次脚本化 ETL 就把 GCS 里的向量灌进 Qdrant Cloud。
Two migrations preceded Qdrant. First, a fully-managed vector service: tens of gigabytes already cost about $5,000/month, with no native filterable HNSW forcing a separate index per country/date/tag combination, and no infrastructure visibility making troubleshooting slow and expensive. Second, a self-hosted open-source alternative: cheaper, but a two-person team on on-call duty, upgrades during business hours, and the same filtering limitations. After reading Qdrant's filterable-HNSW blog post, the third migration: filterable HNSW was the deal-maker, Qdrant Cloud's managed Rust backbone offloaded 24x7 ops; 8-bit scalar quantization keeps hot vectors in RAM with colder full-precision embeddings on disk; a weekend of scripted ETL pushed GCS-held embeddings into Qdrant Cloud.
- 结果
可检索专利语料从约 2000 万扩到 2 亿+;管理向量量从千万级到数亿级;典型查询延迟从 250–400ms 降到 p95 <100ms;每 GB 存储成本降约 10 倍;还孵化出全新收入线——高置信度侵权分析产品(客户点一次按钮,几分钟内拿到权利要求图表级分析)。以上数字来自 Qdrant 官方博客整理的 KPI 对照表,厂商渠道口径,引用须注明。
Addressable patent corpus grew from ~20M to 200M+; vectors under management from tens of millions to hundreds of millions; typical query latency from 250-400ms to <100ms p95; cost per stored GB down ~10x; plus a brand-new revenue line — a high-confidence infringement-analysis product (clients click a button and get claim-chart-quality analysis in minutes). Figures from the KPI table in Qdrant's official blog — vendor-channel source, cite as such.
- 机制根因
没有 filterable HNSW 时,"过滤"只能靠建 N 个预过滤索引实现——组合爆炸是成本与复杂度的根源;in-graph 过滤下推把过滤做进 HNSW 遍历,一份索引服务任意过滤组合,索引数从"组合数"降到 1;标量量化(8-bit)+ 冷热分层让读多、突发的工作负载在有限内存下跑出 <100ms p95。两次失败的迁移恰好证明:托管省心和过滤能力缺一不可。
Without filterable HNSW, "filtering" meant pre-building N indexes — combinatorial explosion was the root of cost and complexity; in-graph filter pushdown folds filtering into HNSW traversal, so one index serves arbitrary filter combinations — index count drops from "number of combinations" to 1; scalar quantization (8-bit) plus hot/cold tiering delivers <100ms p95 on limited memory for a read-heavy, bursty workload. The two failed migrations prove the point: managed peace of mind and filtering capability are both non-negotiable.
- 教训
过滤密集型场景选型,先问"过滤条件有多少种组合"——答案是"很多"时,没有原生 filterable ANN 的库直接出局;小团队别为省托管费自己 on-call,两次迁移的学费证明"源代码透明 + 托管运维"是可以兼得的;周末能完成的迁移(ETL 脚本只改几行)说明 ingestion API 贴近开源惯例是迁移摩擦力的决定因素。
For filter-heavy scenarios, first ask "how many filter combinations exist" — if the answer is "many," any engine without native filterable ANN is out; small teams should not trade managed hosting for self on-call — two migrations' tuition proves source transparency plus managed ops can coexist; a weekend migration (ETL script changed by a few lines) shows an ingestion API close to open-source conventions is the deciding factor in migration friction.
相关产品:Qdrant 相关能力:过滤 + 混合检索的工程完成度——payload 索引、in-graph 过滤下推、稀疏/稠密 RRF 服务端融合 最后核验:2026-10-02
HubSpot:20 亿+向量的 VaaS 平台,Breeze AI 选型 Qdrant 成功经验
超大规模
向量即服务
多团队平台
自研运维
- 场景
HubSpot 把 Qdrant 做成内部 Vector-as-a-Service 平台,服务 38+ 团队、200+ 索引、140+ 集群、5 个地域,向量总量超 200 亿,写入峰值 10 万 QPS。驱动场景是旗舰智能助手 Breeze AI:个性化、上下文感知的实时推荐与 RAG,检索速度和准确率直接决定用户参与度;同时数据量与交互量持续快速增长,系统不能随规模退化。
HubSpot turned Qdrant into an internal Vector-as-a-Service platform serving 38+ teams, 200+ indexes, 140+ clusters across 5 regions, with over 20 billion vectors and write peaks of 100k QPS. The driving workload is Breeze AI, its flagship intelligent assistant: personalized, context-aware real-time recommendations and RAG, where retrieval speed and accuracy directly shape user engagement — with data volume and interaction counts growing fast, and the system must not degrade with scale.
- 决策
多款向量数据库对比评估后选 Qdrant:检索与排序的性能显著胜出;开发者友好的集成加速了开发节奏;named vectors、多向量检索、稀疏向量、混合检索、多阶段查询与加权 rerank 等能力对长期路线图友好;on-prem 部署满足数据管控要求。运维上从手动 Helm 部署迁移到自研 K8s Operator(滚动升级、自动扩缩、自愈)。
After evaluating multiple vector databases, HubSpot chose Qdrant: retrieval and ranking performance won decisively; developer-friendly integration accelerated development cycles; named vectors, multi-vector search, sparse vectors, hybrid search, multi-stage queries, and weighted rerank fit the long-term roadmap; on-prem deployment met data-control requirements. Operationally, the team migrated from manual Helm deployments to a self-built Kubernetes Operator (rolling upgrades, autoscaling, self-healing).
- 结果
Breeze AI 检索延迟下降、推荐与 RAG 应用满足实时要求;向量检索集成复杂度降低,工程资源释放回模型与体验迭代;平台支撑 AI 交互量持续增长而未出现基础设施瓶颈。HubSpot 技术负责人 Srubin Sethu Madhavan 评价:"我们选择它是因为易于部署和大规模下的高性能,结果一直令人印象深刻"(经 Qdrant 官方博客发布,厂商渠道口径)。
Breeze AI retrieval times dropped, and recommendation/RAG workloads meet real-time requirements; vector-search integration complexity fell, freeing engineering resources back to models and UX iteration; the platform sustains growing AI interaction volumes without infrastructure bottlenecks. Srubin Sethu Madhavan, Technical Lead at HubSpot: "We chose it for its ease of deployment and high performance at scale, and we have been consistently impressed with its results" (via Qdrant official blog, vendor-channel source).
- 机制根因
单机 Rust 引擎的高 QPS/低延迟是能扛住 10 万写 QPS 的底盘;量化与 on-disk 索引把 200 亿量级的存储成本压到可承受;named vectors + 稀疏向量 + RRF 服务端融合让"推荐+关键词+语义"多路检索一次查询完成,避免多系统拼凑。代价是分布式运维回到自己手里——HubSpot 自研 Operator 本身就是"官方运维工具链没到无脑程度"的证据。
The Rust single-node engine's high QPS/low latency is the floor that absorbs 100k write QPS; quantization and on-disk indexing keep 20-billion-vector storage costs bearable; named vectors + sparse vectors + server-side RRF fusion complete multi-path retrieval (recommendation + keyword + semantic) in a single query, avoiding multi-system patchwork. The price: distributed operations land back on your own team — HubSpot building its own Operator is itself evidence the official ops tooling is not yet hands-off.
- 教训
超大规模选型要把"写峰值"和"团队数"写进合同式验收标准,而不仅是读延迟;Qdrant 的分布式是"功能完整、运维自理"——选之前先诚实评估自己有没有平台团队能写 Operator;多租户/多团队平台场景下,named vectors 与多阶段查询这类"一次查询做多件事"的能力,比单纯的向量延迟数字更能决定架构复杂度。
Hyperscale selection must put "write peaks" and "team count" into contract-style acceptance criteria, not just read latency; Qdrant's distributed story is "feature-complete, ops are yours" — honestly assess whether you have a platform team able to write an Operator before committing; in multi-tenant/multi-team platform scenarios, capabilities like named vectors and multi-stage queries that do many things in one query matter more for architectural complexity than raw vector-latency numbers.
相关产品:Qdrant 相关能力:1M–100M 区间的性价比之王——Rust 单机的高 QPS 与低延迟、分布式 HA 与索引单一性——"单机之王"的天花板,以及 hybrid 的三个生产坑 最后核验:2026-10-02
Lyzr:Agent 平台从 Weaviate/Pinecone 迁 Qdrant,查询延迟降 90%+ 成功经验
Agent 基础设施
Weaviate 替代
高并发检索
成本下降
- 场景
Lyzr Agent Studio 是 AI Agent 平台,部署了 100+ 跨行业 Agent。早期栈用 Weaviate(另对 Pinecone 做基准):1500 条向量、10–20 个知识检索 Agent、每 Agent 每分钟 5–10 查询时一切正常(延迟 80–150ms)。知识库超过 2500 条、Agent 并发超过 100 后系统开始吃力:查询延迟涨近 4 倍到 300–500ms,高峰期 Agent 等待向量结果超时、下游决策逻辑受影响;索引操作变慢,吃 CPU 和内存,数据更新出现瓶颈。
Lyzr Agent Studio is an AI agent platform with 100+ agents deployed across industries. The early stack used Weaviate (with Pinecone benchmarked): fine at 1,500 vector entries, 10-20 knowledge-search agents, 5-10 queries per agent per minute (80-150ms latency). Past 2,500 entries and 100 concurrent agents, the system strained: query latency rose nearly 4x to 300-500ms; at peak, agents timed out waiting for vector results, breaking downstream decision logic; indexing slowed, consumed excess CPU and memory, and bottlenecked data updates.
- 决策
按可扩展性、索引性能、查询延迟/吞吐、一致性、资源效率、真实负载基准六个维度评估替代方案,选 Qdrant:HNSW 索引支持在线更新无需停机重建;高并发下延迟稳定。
Evaluated alternatives on six dimensions — scalability, indexing performance, query latency/throughput, consistency, resource efficiency, real-load benchmarks — and chose Qdrant: HNSW indexing supports live updates without downtime or reindexing; latency stays stable under high concurrency.
- 结果
查询延迟降到 20–50ms(P99),比 Weaviate/Pinecone 好 90% 以上;每分钟 1000+ 查询、100+ 并发 Agent 下性能稳定,峰值吞吐超 250 QPS;大数据集 ingestion 快 2 倍;基础设施成本降约 30%。以上数字来自 Qdrant 官方博客整理的三方对比表(Weaviate 300–500ms / Pinecone 250–450ms / Qdrant 20–50ms P99;索引 3h / 2.5h / 1.5h;吞吐 ~80 / ~100 / >250 QPS),厂商渠道口径,引用须注明。同文还收录两个下游部署:NTT Data 把 IT 变更请求的 Agent 从 Azure Cosmos DB 迁到 Qdrant 后长尾查询准确率明显提升;NPD 在 6 个网站的客服 Agent 消除了之前方案的延迟尖峰。
Query latency dropped to 20-50ms (P99), >90% better than Weaviate/Pinecone; stable at 1,000+ queries per minute with 100+ concurrent agents, peak throughput above 250 QPS; large-dataset ingestion 2x faster; infrastructure costs down ~30%. Figures from the three-way comparison table in Qdrant's official blog (Weaviate 300-500ms / Pinecone 250-450ms / Qdrant 20-50ms P99; indexing 3h / 2.5h / 1.5h; throughput ~80 / ~100 / >250 QPS) — vendor-channel source, cite as such. The same piece covers two downstream deployments: NTT Data moved an IT-change-request agent from Azure Cosmos DB to Qdrant with substantially better long-tail retrieval accuracy; NPD's customer-facing agents across six websites eliminated the latency spikes of the previous solution.
- 机制根因
Agent 是"检索循环" workload——一次任务触发多次检索,单次检索慢 300ms 会被循环放大成秒级;Weaviate/Pinecone 在高并发下延迟退化,本质是索引与查询争资源,而 Qdrant 的 Rust 单二进制 + 可调 HNSW 参数在同样硬件下留出更多余量。选型教训的普适点:demo 规模(1500 条向量)下"都够用"的结论,到生产并发下会系统性失效——基准必须按生产并发打。
Agents are a retrieval-loop workload — one task triggers many retrievals, so a 300ms slower single retrieval multiplies into seconds; Weaviate/Pinecone degraded under concurrency because indexing and querying contend for resources, while Qdrant's Rust single binary plus tunable HNSW parameters leaves more headroom on the same hardware. The generalizable lesson: "everything is fine" at demo scale (1,500 vectors) fails systematically at production concurrency — benchmarks must be run at production concurrency.
- 教训
Agent 场景的向量选型,第一指标是"高并发下的 P99 延迟"而不是"单查询延迟";demo 数据量下的基准没有决策价值,压测要按生产并发(Lyzr 是 100+ 并发 Agent、1000+ QPM)来;检索是 Agent 的内循环——"AI 太慢"的用户投诉,根因经常是向量库不在 LLM;从 Cosmos DB 这类通用库迁过来的案例说明:向量检索作为独立组件选型的时代已经到了。
For agent vector selection, the primary metric is P99 latency under high concurrency, not single-query latency; demo-scale benchmarks have no decision value — load-test at production concurrency (Lyzr: 100+ concurrent agents, 1,000+ QPM); retrieval is the agent's inner loop — "AI is too slow" complaints are often rooted in the vector store, not the LLM; migrations from general-purpose stores like Cosmos DB show the era of selecting vector search as a standalone component has arrived.
来源
Qdrant official blog, "How Lyzr Supercharged AI Agent Performance with Qdrant" (vendor-channel source, with three-way benchmark table and NTT Data / NPD downstream cases)
https://qdrant.tech/blog/case-study-lyzr/
相关产品:Qdrant、Weaviate 相关能力:1M–100M 区间的性价比之王——Rust 单机的高 QPS 与低延迟 最后核验:2026-10-02
Mixpeek:MongoDB kNN 做多模态特征库扛不住,迁 Qdrant 检索快 40% 成功经验
MongoDB 替代
多模态特征库
混合检索
RRF 原生
多向量
- 场景
Mixpeek 是多模态数据处理与检索平台(视频、图像、音频、文本),创始人 Ethan Steininger 是前 MongoDB 搜索专家。特征库(feature store)需要支撑越来越复杂的检索模式:稠密+稀疏向量混合检索带元数据预过滤。MongoDB Atlas 的向量搜索在两处掉链子:做视频 embedding 的 ColBERT 式 late interaction 需要多向量索引,MongoDB kNN 支持不了;客户要做程序化广告投放的反向视频搜索,在海量对象集合里找高转化视频片段,MongoDB 的通用特征库效率低下。
Mixpeek is a multimodal data processing and retrieval platform (video, image, audio, text), founded by Ethan Steininger, a former MongoDB search specialist. Its feature stores needed to support increasingly complex retrieval patterns: hybrid retrievers combining dense and sparse vectors with metadata pre-filtering. MongoDB Atlas vector search failed in two places: ColBERT-style late interaction over video embeddings requires multi-vector indexing, which MongoDB kNN cannot support; and a customer needed reverse video search for programmatic ad serving — finding high-converting video segments across massive object collections — inefficient with MongoDB's general-purpose feature stores.
- 决策
评估了 Postgres + pgvector、MongoDB kNN 等选项后选 Qdrant 做特征库:向量检索专精 + 与检索管线集成顺滑;原生多向量索引是实现 ColBERT 式 late interaction 的前提;原生 RRF(Reciprocal Rank Fusion)把混合检索的合并逻辑收进服务端。
After evaluating Postgres with pgvector, MongoDB kNN, and others, Mixpeek chose Qdrant for its feature stores: vector-search specialization plus smooth integration with retrieval pipelines; native multi-vector indexing as the prerequisite for ColBERT-style late interaction; native RRF (Reciprocal Rank Fusion) pulling hybrid-merge logic server-side.
- 结果
混合检索代码量减少 80%("原来维护复杂自定义逻辑合并多个特征库的结果,Qdrant 的 RRF 让混合检索器的实现简化了 80%"——创始人原话,经 Qdrant 官方博客发布,厂商渠道口径);十亿级特征集合上 prefetch 并行检索让查询时间从约 2.5s 降到 1.3–1.6s(快 40%);SageMaker 做特征提取时,数据库查询曾是显著瓶颈,迁后查询开销降 50%,ingestion 管线被捋顺。
Hybrid retriever code reduced by 80% ("We eliminated hundreds of lines of code with what was previously a MongoDB kNN Hybrid search when we replaced it with Qdrant as our feature store." — founder, via Qdrant official blog, vendor-channel source); on billion-feature collections, prefetch parallel retrieval cut query times from ~2.5s to 1.3-1.6s (40% faster); database queries had been a significant bottleneck while running SageMaker feature extraction — post-migration query overhead dropped 50%, streamlining ingestion pipelines.
- 机制根因
通用数据库的向量搜索是"附加功能",多向量、稀疏向量、服务端融合这类检索原语不在其设计中心——MongoDB kNN 撑不起 late interaction 是架构性差距;Qdrant 把 RRF、prefetch、named vectors 做成一等公民,混合检索从"应用层拼胶水"变成"一次查询声明意图";前 MongoDB 搜索专家亲手把自家老东家的方案换掉,是"专精打败通用"在这个场景最有分量的背书。
A general-purpose database's vector search is an add-on — retrieval primitives like multi-vector, sparse vectors, and server-side fusion are not at its design center; MongoDB kNN's inability to support late interaction is an architectural gap; Qdrant makes RRF, prefetch, and named vectors first-class, turning hybrid retrieval from application-layer glue into a single declarative query; a former MongoDB search specialist replacing his old employer's solution is the weightiest possible endorsement of specialization over generality in this scenario.
- 教训
多模态/混合检索场景选型,先列出"必须原生的检索原语清单"(多向量?稀疏?服务端融合?预过滤?),通用库的向量功能是"能跑",不是"好用";代码量减少 80% 这类指标比 QPS 更能说明"专精"的价值——它直接等于工程师时间;特征提取(SageMaker)与特征检索的瓶颈要分开看,数据库只能解决后者,别指望换库解决 embedding 推理慢。
For multimodal/hybrid retrieval selection, first list the "must-be-native retrieval primitives" (multi-vector? sparse? server-side fusion? pre-filtering?) — a generalist's vector feature "runs" but is not "good"; an 80% code reduction says more about specialization than QPS does — it directly equals engineer hours; separate the feature-extraction (SageMaker) bottleneck from the feature-retrieval bottleneck — a database only fixes the latter, don't expect a swap to fix slow embedding inference.
相关产品:Qdrant、MongoDB、PostgreSQL(社区版) 相关能力:过滤 + 混合检索的工程完成度——payload 索引、in-graph 过滤下推、稀疏/稠密 RRF 服务端融合 最后核验:2026-10-02
OpenTable:Concierge AI 点餐助手的稀疏向量 + 重度过滤选型 成功经验
稀疏向量
元数据过滤
RAG 准确性
托管省心
- 场景
OpenTable 构建 AI 餐饮助手 Concierge,用自然语言回答餐厅相关问题。业务优先级明确:第一是"可答率"(绝大多数问题都能答),第二是准确性(错的菜单信息会同时伤害用户与餐厅信任)。检索层需求苛刻:需要稀疏向量做关键词扩展 + 细粒度过滤;查询经常要从 6 万多家餐厅中精确定位到"一家",集合实际是稀疏的,对过滤性能要求极高。
OpenTable built Concierge, an AI dining assistant that answers restaurant questions in natural language. Business priorities were explicit: answerability first (most questions must be answerable), accuracy second (wrong menu items erode both diner and restaurant trust). The retrieval layer faced demanding requirements: sparse embeddings for keyword expansion plus fine-grained filtering; queries often narrowed to a single restaurant out of 60,000+, making collections effectively sparse and putting heavy demands on filtering performance.
- 决策
选型 Qdrant,三个理由直接对应业务优先级:稀疏向量处理是关键差异点(许多向量库在这种"集合极度稀疏"条件下 HNSW 图质量退化,Qdrant 的优化避免了性能跌落);高精度过滤性能可预测,满足延迟预算;Qdrant Cloud 让部署比自托管简单("创建 Qdrant Cloud 集群是这个项目里最容易的部分之一,它就是能用"——OpenTable 团队原话,经官方博客发布,厂商渠道口径)。
Chose Qdrant for three reasons mapped directly to business priorities: sparse-embedding handling as the key differentiator (many vector databases see HNSW graph quality degrade under such sparse conditions, while Qdrant's optimizations avoided the drop); reliable, predictable high-precision filtering to hit the latency budget; and Qdrant Cloud offering a simpler deployment path than self-hosting ("Creating a Qdrant Cloud cluster was one of the easiest parts of the project. It just worked." — OpenTable team, via official blog, vendor-channel source).
- 结果
Concierge 按期上线并全球发布,达成延迟目标且保持高可答率,上线后几乎无需调优;运营上 Qdrant 成为栈里最稳定的组件之一("自从投产以来,它是栈里无摩擦的一部分"——团队原话,同上厂商渠道口径)。
Concierge launched globally from day one, met its latency target with high answerability, and needed almost no post-launch tuning; operationally, Qdrant became one of the most stable components in the stack ("Since running it in production, it is a frictionless part of the stack." — team quote, same vendor-channel source).
- 机制根因
重过滤场景下 HNSW 遍历若图质量差,过滤下推的剪枝效率会崩——Qdrant 在稀疏集合下保持图质量是选型成立的工程前提;稀疏向量(关键词扩展)与稠密向量同库混合,避免了"关键词一路、语义一路"的双系统拼凑;托管化把向量层的运维成本压到零,团队把精力留给了模型与体验迭代。
In heavy-filter scenarios, if HNSW graph quality degrades, filter-pushdown pruning efficiency collapses — Qdrant's graph quality under sparse collections was the engineering premise that made the choice work; keeping sparse (keyword) and dense vectors in one store avoided the dual-system patchwork of keyword path plus semantic path; managed hosting drove vector-layer ops cost to zero, leaving the team free to iterate on models and UX.
- 教训
选型时先把"过滤选择性"摆到台面上——能过滤到单餐厅的场景里,过滤性能才是第一指标,单纯的向量延迟基准会误导;稀疏向量不是"锦上添花",在关键词确定性要求高的场景(菜单、菜名)它是准确性底线;小团队做 AI 功能时,把向量层交给托管服务换取迭代速度是划算的买卖。
Put "filter selectivity" on the table during evaluation — in scenarios that filter down to a single restaurant, filtering performance is the primary metric, and pure vector-latency benchmarks will mislead; sparse vectors are not a nice-to-have — where keyword determinism matters (menus, dish names), they are the accuracy floor; for small teams shipping AI features, trading vector-layer ops for managed hosting is a worthwhile deal for iteration speed.
相关产品:Qdrant 相关能力:过滤 + 混合检索的工程完成度——payload 索引、in-graph 过滤下推、稀疏/稠密 RRF 服务端融合 最后核验:2026-10-02
TripAdvisor:十亿级评论多模态数据激活,AI 行程规划带来 2–3 倍收入提升 成功经验
生成式 AI
行程规划
用户图谱
多模态检索
收入挂钩
- 场景
TripAdvisor 是全球最大旅行指南平台,每月数亿活跃用户、1100 万商户、超 10 亿条用户评论与贡献(其中包含数亿张图片),另有酒店、餐厅、体验等多年的行为数据。数据资产长期处于"沉睡"状态,尤其非结构化内容没有被激活。数据与 AI 负责人 Rahul Todkar(曾负责 LinkedIn 的数据与向量系统)上任后推动转型:把散落的评论、图片、行为数据统一成向量可检索的用户图谱(user graph),让行程规划、推荐、搜索都跑在向量检索之上。
Tripadvisor is the world's largest travel guidance platform, with hundreds of millions of monthly users, 11 million businesses, over a billion reviews and contributions (including hundreds of millions of images), and years of behavioral data across hotels, restaurants, and experiences. That data asset sat largely dormant — especially its unstructured content. Rahul Todkar, Head of Data and AI (previously built LinkedIn's data and vector systems), drove a transformation: unifying scattered reviews, images, and behavior into a vector-retrievable user graph so trip planning, recommendations, and search all run on vector search.
- 决策
选择 Qdrant 作为向量数据库底座;用向量构建多维用户图谱(酒店偏好、餐饮选择、旅行风格、用户行为统一表征);上线生成式 AI 产品 Trip Planner(对话式行程规划),并把向量检索扩展到搜索重构(从"过滤器+标签页"转向对话式双向搜索)。
Chose Qdrant as the vector database foundation; built a multidimensional user graph on vectors (unified representations of hotel preferences, dining choices, travel styles, and user behavior); shipped Trip Planner, a generative-AI itinerary builder, and extended vector search to reinvent search itself (from filter-and-tab interfaces toward conversational, bidirectional search).
- 结果
使用 Trip Planner 的旅行者带来的收入是传统界面的 2–3 倍(TripAdvisor 公司口径,经 Qdrant 官方博客发布,属厂商渠道口径,引用须注明;收入对比口径为"使用生成式 AI 体验的用户 vs 未使用的用户",非严格 A/B 实验,解读时需留有余地)。
Travelers engaging with Trip Planner generate 2 to 3 times more revenue than those using traditional interfaces (Tripadvisor company figure, published via Qdrant's official blog — vendor-channel source, cite as such; the comparison is users of the GenAI experience vs non-users, not a randomized A/B test, so interpret with margin).
- 机制根因
超 10 亿多模态条目的统一向量表征,把原来分散在各业务线的偏好信号变成可查询的用户图谱;向量检索天然适合"模糊偏好 + 上下文"的对话式查询,而传统倒排/过滤器系统做不到语义泛化。收入提升的本质不是检索变快了,而是原来沉睡的数据资产第一次被用于影响购买决策。
A unified vector representation of 1B+ multimodal items turned preference signals scattered across business lines into a queryable user graph; vector retrieval is a natural fit for fuzzy, context-rich conversational queries that inverted-index/filter systems cannot generalize over. The revenue lift came not from faster retrieval but from a dormant data asset finally influencing purchase decisions.
- 教训
向量数据库的 ROI 故事不止于"延迟降低",能把数据资产变成收入杠杆的选型论证更容易通过;先有一个能产生业务数字的旗舰场景(Trip Planner),再扩展到全站搜索改造,是大平台落地的稳妥顺序;收入倍数类指标一定要注明统计口径(使用人群 vs 未使用人群存在自选择偏差)。
The ROI story for a vector database goes beyond latency reduction — selection arguments framed as a revenue lever on data assets clear approval more easily; land one flagship scenario with measurable business numbers (Trip Planner) before re-platforming site-wide search — the safe sequence for large platforms; revenue-multiplier metrics must always carry their statistical definition (engaged users vs non-users have self-selection bias).
相关产品:Qdrant 相关能力:— 最后核验:2026-10-02
leboncoin:100+ ElastiCache 实例平迁 Valkey——许可证倒逼的零代码改动迁移 (2026) 成功经验
Valkey 迁移
许可证合规
零停机
基础设施即代码
- 场景
分类广告平台 leboncoin 的存储团队(Staff 工程师 Flavio Gurgel、Jonathan Lubin、Naeva Mallet)管理着 100 多个 Amazon ElastiCache 实例,内存里近 10 亿个 key,个别集群达 500GiB。2024 年 3 月 Redis 改许可证(RSALv2/SSPL)后,继续守旧版本意味着两头风险:AWS 对旧引擎版本的 Extended Support 额外费用,以及"跑在没人维护的分支上"。
The storage team at classified-ads platform leboncoin (staff engineer Flavio Gurgel with Jonathan Lubin and Naeva Mallet) manages more than 100 Amazon ElastiCache instances holding nearly a billion keys in memory, with individual clusters reaching around 500GiB. After Redis changed its license in March 2024 (RSALv2/SSPL), staying on old versions meant two-sided risk: extra AWS Extended Support fees for legacy engine versions, and running on a branch nobody maintains.
- 决策
迁到 Valkey:Linux 基金会项目、BSD 3-Clause、RESP 协议兼容。迁移前全量实例已经跑在 Redis 7.2.4——正好是 Valkey 的分叉基线,于是策略定为"基础设施层的事,不碰业务代码":用 AWS 官方提供的 Redis→Valkey 迁移机制,先在预生产环境验证流程(含 TLS 证书链检查),再按每批 5 个实例分波 rollout;每个实例"迁移 + Terraform 状态更新"约 20 分钟,可并行。
Move to Valkey: a Linux Foundation project, BSD 3-Clause, RESP-protocol compatible. Every instance was already on Redis 7.2.4 - exactly Valkey's fork baseline - so the strategy was "an infrastructure-layer concern, no application code touched": use AWS's official Redis-to-Valkey migration mechanism, validate the runbook (including TLS certificate-chain checks) in pre-production first, then roll out in waves of five instances; each instance's "migrate + Terraform state update" took about 20 minutes and parallelized.
- 结果
历时数月完成 100+ 实例迁移,业务代码零改动、客户端库零更换、应用团队几乎无感。唯一的事故是一次漏掉的 Terraform apply,导致单个非关键微服务短暂配置丢失——作者明确指出这是迁移流程失误,与 Valkey 本身无关;迁移后未观察到兼容性、稳定性或性能问题。
All 100+ instances migrated over a couple of months with zero changes to backend code, zero client-library swaps, and application teams barely noticing. The single incident was a missed Terraform apply that briefly left one non-critical microservice misconfigured - the author states plainly this was a migration-procedure mistake, unrelated to Valkey itself; no compatibility, stability, or performance issues were observed after migration.
- 机制根因
能"零改动"的前提是三层对齐:协议层(RESP 没变)、数据格式层(RDB 兼容)、版本层(7.2.4 正好是分叉点,无需跨版本升级)。AWS 的迁移机制本质是"换引擎不断连接":过程中可能有短暂连接抖动,但 Redis 客户端自带重连,所以无感。真正的风险不在 Valkey 而在周边:TLS 证书链、IaC 状态漂移——而唯一的事故也恰恰出在 IaC 这边,印证了风险判断。
"Zero changes" was possible because three layers aligned: protocol (RESP unchanged), data format (RDB compatible), and version (7.2.4 is exactly the fork point, so no cross-version upgrade). AWS's migration mechanism is essentially "swap the engine without dropping connections": brief connection blips may occur, but Redis clients reconnect on their own, so nobody notices. The real risk was never Valkey but the surroundings - TLS chains, IaC state drift - and the one incident landed exactly on the IaC side, confirming the risk assessment.
- 教训
许可证是架构决策的一部分:选型时把"协议开源且可分叉"计入 TCO,SSPL 这类变更的成本最终由用户承担;大舰队迁移的关键是"分叉基线对齐版本"——先把全量实例升到 7.2.4 再迁,比直接跨版本跳省掉一整类问题;IaC 状态与真实 infra 的双写必须原子化,漏一次 apply 就够喝一壶,checklist 要写在流程里而不是记在脑子里。
Licensing is part of architecture decisions: count "protocol is open and forkable" in TCO, because the cost of changes like SSPL ultimately lands on users. The key to fleet-scale migration is "align versions to the fork baseline" - upgrading everything to 7.2.4 first eliminates a whole class of problems versus jumping across versions. And dual-writes between IaC state and real infrastructure must be atomic: one missed apply is enough to ruin your day, so put the checklist in the runbook, not in someone's head.
来源
leboncoin tech blog "Redis to Valkey migration on AWS: How we migrated more than 100 ElastiCache instances without changing our microservices" (June 2026, by Flavio Gurgel
—
相关产品:Redis / Valkey 相关能力:Redis 换协议 → Valkey 分支:许可证事件重塑选型 最后核验:2026-10-02
Redis 许可证事件:一次"防云厂商"的决策,8 天逼出 Valkey 分叉 (2024) 失败教训
开源许可证
社区分叉
SSPL
厂商治理
- 场景
2024 年 3 月 20 日,Redis 公司宣布:从 7.4 起,Redis 采用 RSALv2 + SSPLv1 双许可证,不再使用 BSD 3-Clause。官方理由:云厂商把 Redis 的开源投入"商品化"、只赚钱不回馈。但 SSPL 要求"提供托管服务就必须开源全部管理代码",OSI 明确不承认它是开源许可证——"source available"不等于 open source。15 年来 Redis 生态恰恰建立在"BSD + 云厂商深度集成"之上。
On March 20, 2024, Redis Inc. announced that starting with 7.4, Redis would be dual-licensed under RSALv2 + SSPLv1, abandoning BSD 3-Clause. The stated reason: cloud vendors were "commoditizing" Redis's open-source investment without giving back. But SSPL demands that anyone offering a managed service open-source all management code, and the OSI explicitly does not recognize it as an open-source license - "source available" is not open source. For 15 years, the Redis ecosystem had been built precisely on "BSD plus deep cloud-vendor integration."
- 决策
Redis 公司的算盘是"逼云厂商签商业协议"。但它误判了筹码:AWS、Google Cloud 既是最大的分发渠道,也是核心贡献者——当时的 Redis 核心维护者 Madelyn Olson 就是 AWS 工程师。许可证变更 8 天后(2024 年 3 月 28 日),Linux 基金会宣布成立 Valkey:从 Redis 7.2.4(最后一个 BSD 版本)分叉,BSD 3-Clause 不变;创始支持者包括 AWS、Google Cloud、Oracle、Ericsson、Snap,技术领导委员会由多位前 Redis 贡献者组成。
Redis Inc.'s bet was "force the cloud vendors into commercial agreements." It misread its leverage: AWS and Google Cloud were both the largest distribution channels and core contributors - Redis's core maintainer at the time, Madelyn Olson, was an AWS engineer. Eight days after the license change (March 28, 2024), the Linux Foundation announced Valkey: forked from Redis 7.2.4 (the last BSD release), keeping BSD 3-Clause; founding supporters included AWS, Google Cloud, Oracle, Ericsson, and Snap, with a technical leadership committee of former Redis contributors.
- 结果
社区分裂几乎是即时的:核心贡献者出走 Valkey,AWS、Google Cloud 等把托管服务转向 Valkey(Linux 基金会 2024 年 3 月声明的支持者名单;leboncoin 迁移时 AWS 已提供官方 Redis→Valkey 迁移机制)。Redis 公司被迫回头:Redis 8 发布时改用 AGPL 许可证回归开源,CEO Rowan Trollope 公开承认"这次变更伤害了我们和社区的关系"("SSPL is not truly open source")。
The community split was nearly instant: core contributors walked to Valkey, and AWS, Google Cloud, and others steered managed services toward Valkey (per the Linux Foundation's March 2024 supporter list; by the time of leboncoin's migration, AWS already offered an official Redis-to-Valkey migration path). Redis Inc. was forced to reverse course: Redis 8 launched under the AGPL open-source license, with CEO Rowan Trollope publicly admitting "the change hurt our relationship with the Redis community" ("SSPL is not truly open source").
- 机制根因
这是"平台型开源"的经典困境:Redis 的价值一半在代码,一半在"所有人都默认它在"的生态位——客户端、教程、云集成、运维知识。改许可证打击的正是后者,而后者恰恰是云厂商和贡献者共建的。当"默认选项"可以被 8 天分叉带走,说明护城河从来不是代码,而是社区共识。Redis 赌的是"云厂商离不开 Redis",现实是云厂商离不开的是"RESP 协议的生态",不是"Redis 这家公司"。
The classic "platform open source" dilemma: half of Redis's value was the code, half was the ecosystem position of "everyone just assumes it's there" - clients, tutorials, cloud integrations, operational knowledge. The license change struck at the latter, which was co-built by cloud vendors and contributors. When the "default option" can be forked away in 8 days, the moat was never the code - it was community consensus. Redis bet that "cloud vendors can't live without Redis"; the reality was they couldn't live without "the RESP-protocol ecosystem," not "the Redis company."
- 教训
选型时把许可证当作架构属性:BSD/Apache 项目的"可分叉性"本身就是一种保险,SSPL/RSAL 类项目要预先评估"如果社区分叉,你跟哪边";厂商的"防云"叙事和用户的"防锁定"叙事是零和的,站队前先看贡献者名单——当核心维护者都在云厂商一边时,许可证变更就是在和自己的生态开战。许可证不是法务部的事,是架构师的事。
Treat licensing as an architecture attribute: the "forkability" of BSD/Apache projects is itself insurance, and SSPL/RSAL-style projects demand an up-front answer to "if the community forks, which side are you on?" A vendor's "anti-cloud" narrative and a user's "anti-lock-in" narrative are zero-sum - check the contributor roster before picking sides. When the core maintainers sit on the cloud vendors' side, a license change is a war against your own ecosystem. Licensing is not legal's problem; it is the architect's problem.
来源
InfoWorld "Redis moves to source-available licenses" (March 2024
Linux Foundation "Linux Foundation Launches Open Source Valkey Community" (2024-03-28
—
—
Forkable "Redis returns to open source" (Redis 8 moved to AGPL
—
相关产品:Redis / Valkey 相关能力:Redis 换协议 → Valkey 分支:许可证事件重塑选型 最后核验:2026-10-02
Netflix:Dynomite——给 Redis 套上 Dynamo 的多数据中心外壳 (2013–2014) 成功经验
多数据中心复制
高可用
Dynamo 架构
协议兼容
- 场景
2013 年的 Netflix 全站跑在 AWS 上:Redis/Memcached 都是单机架构,常规主从复制扛不住它的流量规模、也跨不了地域,而 Netflix 要的是"任意节点可写、多数据中心互备"。官方集群方案同样缺位(2015 年才来),在 Netflix 的微服务规模下让每个客户端自己分片不可维护。Dynomite 仓库 2013-10-10 建仓,2014 年 11 月 4 日正式开源(Apache 2.0)。
Netflix in 2013 ran entirely on AWS: Redis and Memcached were single-server architectures, conventional master-slave replication could neither handle its traffic scale nor span regions, yet Netflix needed "writable from any node, replicated across datacenters." The official clustering story was equally absent (it would arrive in 2015), and client-side sharding was unmaintainable at Netflix's microservice scale. The Dynomite repo was created 2013-10-10 and formally open-sourced on November 4, 2014 (Apache 2.0).
- 决策
给每个 Redis/Memcached 实例旁挂一个"dynamo 层"协进程,组成 datacenter/rack/token 拓扑;借 Amazon Dynamo 论文的思路(token 环、quorum、gossip 式节点通信)给单机存储补上跨 DC 复制和高可用。README 原话:"The ultimate goal with Dynomite is to be able to implement high availability and cross-datacenter replication on storage engines that do not inherently provide that functionality." 它讲各存储引擎的原生协议,现有工具链不用换。
Attach a "dynamo layer" sidecar process to every Redis/Memcached instance, forming a datacenter/rack/token topology; borrow the Amazon Dynamo paper's ideas (token ring, quorums, gossip-style node communication) to graft cross-DC replication and high availability onto single-server stores. The README's own words: "The ultimate goal with Dynomite is to be able to implement high availability and cross-datacenter replication on storage engines that do not inherently provide that functionality." It speaks each engine's native protocol, so existing tooling keeps working.
- 结果
支撑了 Netflix 当年跨地域的 KV 需求;配套的 Dyno 客户端处理故障转移(可切到远端 rack/DC),dynomite-manager 做集群管理。代价是运维复杂度:部分 Redis 命令不支持、单节点只持有一个 token(无 vnode),本质是"用一套分布式系统的复杂度换 Redis 的简单"。SiliconANGLE 当时的报道点出了动因:流量太大用不了常规主从,而大规模分片"极其复杂"。
It carried Netflix's cross-region KV needs of the era; the companion Dyno client handled failover (to a remote rack or DC) and dynomite-manager handled cluster management. The price was operational complexity: some Redis commands unsupported, one token per node (no vnodes) - essentially trading a distributed system's complexity for Redis's simplicity. SiliconANGLE's coverage at the time nailed the motive: traffic too large for conventional master-slave, while sharding at scale was "immensely complex."
- 机制根因
Redis 的复制是"主从星型",写只能打到主;Dynomite 把每个节点变成对等节点,写可以打到任意节点,由持有该 token 的节点落本地 Redis、再异步复制到其他 rack/DC。故障时客户端切到远端副本。这是对 CAP 的显式选择:跨 DC 异步复制等于接受最终一致性,换来"地域级灾难也不停写"——这正是云上多活业务要的语义。
Redis replication is a "master-slave star" - writes go to the master only. Dynomite turns every node into a peer: writes can land on any node, the token owner persists to local Redis and asynchronously replicates to other racks and DCs; on failure the client shifts to a remote replica. This is an explicit CAP choice: cross-DC async replication means accepting eventual consistency in exchange for "writes survive region-level disasters" - exactly the semantic multi-active cloud businesses want.
- 教训
给单机存储"套壳"做分布式是可行的,但壳的复杂度会反噬:命令兼容性、拓扑管理、客户端配合缺一不可,三个里面任何一个跟不上,壳就比裸 Redis 更难运维。今天回看,Redis Cluster 和云托管吃掉了这类中间层的生存空间——选型时要问:"这个壳解决的痛,官方方案或云服务是不是已经解决了?"如果答案是 yes,就不要再造一层。
Wrapping a single-server store in a distributed shell works, but the shell's complexity bites back: command compatibility, topology management, and client cooperation are all mandatory - if any one lags, the shell is harder to run than bare Redis. In hindsight, Redis Cluster and managed cloud services ate this middle layer's reason to exist. The selection question to ask: "Is the pain this shell solves already solved by the official solution or a cloud service?" If yes, don't build another layer.
来源
Netflix Dynomite official README (github.com/Netflix/dynomite
—
SiliconANGLE "Netflix open-sources database cloudification engine" (2014-11-04
—
相关产品:Redis / Valkey 相关能力:— 最后核验:2026-10-02
Stack Overflow:两台 Redis 扛起全网问答缓存——L1/L2 分层与"简单到不用操心" (2016–2019) 成功经验
缓存分层
共享缓存
发布订阅
简单架构
- 场景
Stack Overflow 全网问答跑在极小的自建机房里。2019 年 Nick Craver 公开的数字:全站每天 3.08 亿次 HTTP 命中;Redis 层每天处理 15.9 亿条命令、峰值 8.7 万/秒、1.24 亿个活跃 key,但服务器平均 CPU 占用只有 2.01%(最忙的实例不到 1%),256GB 内存用了不到 96GB。2016 年的架构帖是更早的口径:每月约 1600 亿次操作,每个实例 CPU 不到 2%。"远没到 Redis 的极限"——这是 Nick 的原话。
Stack Overflow's entire Q&A network runs out of a remarkably small self-built datacenter. Numbers published by Nick Craver in 2019: 308 million HTTP hits per day site-wide; the Redis layer processes 1.59 billion commands per day, peaks at 86,982 commands/sec, holds 124 million active keys - yet average server CPU sits at just 2.01% (under 1% on the busiest instance), with less than 96GB of 256GB RAM in use. His 2016 architecture post gave the earlier figures: about 160 billion ops per month, every instance under 2% CPU. "Nowhere near Redis's limits" - Craver's own words.
- 决策
架构坚持"本机内存 L1 + Redis 共享 L2"两层:miss 逐层下钻,回填时两层都写。多租户按站点拆 key 空间(全局缓存用 Redis db0,站点缓存按 Sites 表 ID 分 db),值用 protobuf-net 二进制序列化。Redis 的 pub/sub 另作两用:一个 Web 节点删缓存时广播清掉其他节点的 L1;以及 websocket 实时推送(通知、票数、新回答)。
The architecture holds to two tiers - per-web-server in-memory L1 plus shared Redis L2: misses drill down tier by tier, and fills write back to both. The multi-tenant keyspace is split per site (global cache on Redis db 0, per-site caches on separate DBs keyed by the Sites table ID), values serialized with protobuf-net. Redis pub/sub pulls double duty: broadcasting L1 invalidations when one web node deletes a cache entry, and powering websocket realtime pushes (notifications, vote counts, new answers).
- 结果
这套架构从 2016 沿用到 2019 几乎没变,Nick 的原话是"它是那种我们根本不用操心的 infra"(It's a piece of infrastructure we just don't worry about),同时保持主从 HA。2019 年他还做了"反缓存"实验:把问题页侧边栏缓存去掉、直接查 SQL,性能几乎无差别——用来论证"只在需要时才缓存"的纪律。
The setup barely changed from 2016 to 2019. Craver's verdict: "It's a piece of infrastructure we just don't worry about" - while keeping master/slave HA. In 2019 he even ran an "anti-caching" experiment: dropping the question-page sidebar cache and querying SQL directly showed almost no performance difference, used to argue the discipline of "only cache when you need to."
- 机制根因
Redis 在这里赢的不是功能多,而是延迟数量级:同机房网络往返 0.17ms,小对象一次取回 0.2–0.5ms——比本地 RAM 慢数千倍,但比回源 SQL 快约一个数量级(该配置下的实测对比:回源 SQL 为毫秒量级,非 Redis 的通用加速比),恰好卡在 L1 和源之间的甜点位。L1 解决"同一台机器重复命中",L2 解决"请求被 LB 打到不同机器";pub/sub 复用同一套连接做失效广播,省掉了一套消息系统。
Redis wins here not on features but on latency magnitude: 0.17ms network roundtrip in the same datacenter, 0.2-0.5ms for a small-object fetch - thousands of times slower than local RAM, yet about an order of magnitude faster than hitting source SQL (a comparison measured in this configuration: source SQL sits at the millisecond level; not a universal Redis speedup ratio), landing exactly in the sweet spot between L1 and origin. L1 solves "repeat hits on the same box," L2 solves "requests landing on different boxes behind the load balancer"; pub/sub reuses the same connections for invalidation broadcasts, eliminating a separate messaging system.
- 教训
缓存选型先算延迟账再看功能:先确认 Redis 处在"L1 与源之间"的中间层定位,再决定缓存什么;多租户 key 空间隔离(db/前缀)上线第一天就要做对,事后改 key 命名等于迁库;"不用操心"本身是架构指标——Nick 把"我们维护着最流行的 .NET 客户端 StackExchange.Redis,有问题自己能修"列为选型理由,可观测性和可修复性比 benchmark 数字重要。
Do the latency math before comparing features: first confirm Redis sits in the middle tier between L1 and origin, then decide what to cache. Multi-tenant keyspace isolation (DBs/prefixes) must be right on day one - renaming keys later is a migration. "Nothing to worry about" is itself an architecture metric: Craver lists "we maintain the most popular .NET client, StackExchange.Redis, so we can fix it ourselves" as a selection reason - observability and fixability matter more than benchmark numbers.
来源
Nick Craver "Stack Overflow: How We Do App Caching - 2019 Edition" (Aug 6, 2019
—
Nick Craver "Stack Overflow: The Architecture - 2016 Edition" (Feb 17, 2016
—
相关产品:Redis / Valkey、Microsoft SQL Server 相关能力:— 最后核验:2026-10-02
Bonnier News:嫌 Redshift 又慢又难管,连 Elasticsearch 一起关掉 all-in BigQuery(2019–2020) 失败教训
新闻媒体
Redshift 迁出
实时看板
WLM 运维
BigQuery
- 场景
Bonnier News(瑞典媒体集团,Per Näslund,Bonnier News Tech,2020-01-07 复盘)。当时用 Redshift 做分析,另有一个"过时的"Elasticsearch 集群扛实时看板和监控。想换的理由很直接:Redshift sluggish(慢),WLM 的管理工作是想甩掉的负担;编辑部依赖的实时看板 Redshift 做不了,被迫另维护一套 ES,整套架构"不必要地复杂"。
Bonnier News (Swedish media group; retrospective by Per Naslund, Bonnier News Tech, Jan 7, 2020). The team used Redshift for analytics plus an "outdated" Elasticsearch cluster for real-time dashboards and monitoring. The reasons for switching were blunt: Redshift was sluggish, and administering workload management was a burden they wanted to shed; the real-time dashboards their editorial staff depended on were something Redshift could not deliver, forcing a separate ES cluster that made the whole setup "unnecessarily complex."
- 决策
先定需求:实时读写、查询快、不用管索引、一个查询不拖慢其他查询、闲置时便宜、成本可预测、SQL 接口。候选:Athena(很快排除——更像 S3 上的即席查询工具,不算完整数仓)、BigQuery、Snowflake。实测:约 400 行/秒流式写入,BigQuery 4 秒可查,Snowflake 约 1 分钟;流式成本(每天约 10GB):BigQuery $0.5/天,Snowflake $3.5/天(Snowpipe)+$15/天(每 30 秒触发一次 ALTER PIPE REFRESH 的 XSMALL 集群);存储:Snowflake $23/TB/月,BigQuery 估算 $15/TB/月。2019 年 9 月初拍板 BigQuery——性能和成本对比"并不一边倒",IAM 和"分析师本来就在用 BigQuery"才是决定性因素;2020 年初目标下线 Redshift 和 Elasticsearch。
Requirements first: real-time load and queries, fast queries, no index management, one query must not slow down others, cheaper when idle, predictable costs, SQL interface. Candidates: Athena (dismissed quickly - more an ad-hoc query tool over S3 than a full warehouse), BigQuery, Snowflake. Hands-on tests: streaming about 400 rows/second, BigQuery had data queryable within 4 seconds, Snowflake took about 1 minute; streaming cost (about 10GB/day): BigQuery $0.5/day vs Snowflake $3.5/day (Snowpipe) + $15/day (an X-SMALL cluster firing ALTER PIPE REFRESH every 30 seconds); storage: Snowflake $23/TB/month vs an estimated $15/TB/month for BigQuery. BigQuery won in early September 2019 - the performance and cost comparison "wasn't too obvious," and IAM plus "analysts were already using BigQuery" tipped the scales; the goal was to take down Redshift and Elasticsearch in early 2020.
- 结果
关掉一批数据库和系统后"省了大笔成本"(具体数字未披露);分析师"这些天看起来开心多了",全公司数据终于可以去一个地方查。
Shutting down a batch of databases and systems "cut a lot of costs" (no figure disclosed); analysts "seem a lot happier these days," with one place to find and combine data from the whole organization.
- 机制根因
Redshift 的批处理基因(WLM 队列、微批加载)与"新闻编辑部实时看板"的秒级需求天然错位;为补实时能力叠一个 ES,是"用第二个系统的运维复杂度为第一个系统的架构短板买单"。BigQuery 赢在 serverless 的"闲置零成本 + 流式原生",而 Snowflake 当时的流式要靠 Snowpipe + 定时刷新的拼凑方案——这次选型的本质是"为实时需求选流式原生架构",不是"谁的查询跑得更快"。作者也诚实标注了测算缺陷:成本对比只用了 Snowflake medium 集群,small 集群"可能价格接近,但无法确定"。
Redshift's batch DNA (WLM queues, micro-batch loading) is structurally misaligned with a newsroom's second-level dashboard needs; bolting on Elasticsearch to cover real-time is "paying for the first system's architectural shortfall with a second system's operational complexity." BigQuery won on serverless "zero idle cost plus streaming-native," while Snowflake's streaming at the time was a Snowpipe-plus-scheduled-refresh patchwork - the real contest was "pick streaming-native architecture for real-time needs," not "whose queries run faster." The author honestly flags the flaw in his own math: the cost comparison only used a Snowflake medium cluster, and the small cluster "might end up pretty close, but we can't know for sure."
- 教训
当数仓满足不了实时需求时,叠第二个系统(ES 等)是常见但昂贵的补丁,先算清双系统运维总账再动手;流式场景下"数据多久可查"比"查询跑多快"更决定选型;WLM 这种"手动调优旋钮"在团队小、需求杂时是净负担——"不用管"本身就是一种功能。与 Faire、Robin 的 Snowflake 路线不同,Bonnier 证明了迁出潮里还有第三条路:需求里"实时"权重足够高时,serverless + 流式原生会压倒一切。
When the warehouse can't serve real-time needs, stacking a second system (ES and friends) is a common but expensive patch - total up the dual-system operations bill before doing it. For streaming workloads, "how soon data is queryable" decides the pick more than "how fast queries run." Manual tuning knobs like WLM are a net burden for small teams with messy requirements - "nothing to manage" is itself a feature. Unlike the Snowflake routes of Faire and Robin, Bonnier proves a third road in the exodus: when "real-time" weighs enough in requirements, serverless plus streaming-native beats everything else.
来源
Per Naslund (Bonnier News Tech), "Why We Picked Google BigQuery over Snowflake as Our New Data Warehouse Solution" (Jan 7, 2020
—
相关产品:Amazon Redshift、Google BigQuery 相关能力:迁移潮:"反例价值"本身 最后核验:2026-10-02
FanDuel:DC2→RA3→数据共享→Serverless,三次迭代喂饱三倍增长(2018–2023) 成功经验
体育博彩
数据平台
RA3 存算分离
数据共享
Redshift Serverless
- 场景
FanDuel(Flutter Entertainment 旗下,做体育博彩、每日梦幻体育、赛马和在线赌场)各产品线自建的本地数仓逐渐过时,数据团队建了新的 Global Data Platform,以 Amazon Redshift 为唯一可信数据源,支撑风控、盈利、交叉销售等全局分析。首个 Redshift 集群 2018 年上线,用的是 DC2 计算型节点。到 2021 年,负载几乎是 2018 年的三倍,团队靠"不停加节点 + 反复调 WLM"续命,但用户争用越来越严重,加节点已不是办法。
FanDuel (part of Flutter Entertainment; sportsbooks, daily fantasy sports, horse racing, online casinos) saw its product-specific, often on-premises data warehouses go obsolete, so the data team built a new Global Data Platform with Amazon Redshift as the single trusted source for risk, profitability, and cross-sell analysis. Its first Redshift cluster launched in 2018 on DC2 compute nodes. By 2021 workloads had almost tripled since 2018, and the team kept up by "continuously adding nodes and experimenting with WLM" - but user contention kept getting worse, and adding nodes was no longer the answer.
- 决策
分三步演进。2021 年评估 RA3 对 DC2,确认性能相当后上线 RA3 集群,核心目标是存算分离、存储计算独立扩展;2022 年转向数据共享架构:生产者集群专做 ELT,消费者集群隔离分析负载——先给做 dbt 迁移的数据工程师开 dev/test 消费者(可读生产数据做模型验证而不影响生产),2022 年春开出第一个生产消费者、同年夏开出第二个,把重分析负载搬离主集群;2022–2023 年上线体育博彩事件流微批(micro-batch)入仓,并把最难预测的风控与交易(risk and trading)负载迁到 Redshift Serverless(自动 WLM、只按查询计费),高峰期 Serverless 端点可直接读写 provisioned 集群。
A three-step evolution. In 2021 FanDuel evaluated RA3 vs DC2 and, once satisfied performance matched, launched an RA3 cluster with the explicit goal of separating storage and compute so each could scale independently. In 2022 it moved to a data-sharing architecture: a producer cluster dedicated to ELT, consumer clusters isolating analytics - first dev/test consumers for data engineers migrating legacy code to dbt (real production data for model validation without touching production), then the first production consumer in spring 2022 and a second that summer, moving heavy analytics off the main cluster. In 2022-2023 it launched event-based streaming micro-batches for sportsbook data and moved the hardest-to-predict risk-and-trading workloads to Redshift Serverless (automatic WLM, pay-per-query), whose endpoints can read and write provisioned clusters during peaks.
- 结果
关键业务 SLA 快了 3 倍;全负载平均查询效率提升 55%(内部 KPI"Query Efficiency"衡量用户等待查询的时间);博客原文称业务成本总体节省达十倍("tenfold",具体基线与计算口径未披露)。拆分负载后查询并发翻倍、排队减少;C-suite 营收报表的 SLA 大幅提前,在超级碗之前就已达成——此前从未做到。
The most critical business SLAs finished three times faster; average query efficiency rose 55% (an internal "Query Efficiency" KPI measuring time users spent waiting on queries); the post claims an overall tenfold saving in business cost ("tenfold" is the post's wording; baseline and methodology undisclosed). Splitting workloads doubled query concurrency and cut queueing; C-suite revenue reports hit a much earlier SLA - achieved before the Super Bowl, which had never happened before.
- 机制根因
三次迭代解决的是三个不同的问题,顺序不能错:RA3 解决"为存数据买计算"(存储增长问题);数据共享解决"多类负载抢同一集群"(争用域切分),让 WLM 配置可以按负载定制而不是一锅烩;Serverless 解决"风控交易负载波动不可预测"(弹性问题)。单集群模型下,一个坏查询的代价由全公司分摊;把争用域逐层切小后,"查询为什么慢"从玄学变成可归因。事件流微批则把"日报 T+1"变成了"准实时",这是批处理架构天然给不了的。
The three iterations solved three different problems, in an order that mattered. RA3 solved "buying compute to store data" (the storage-growth problem). Data sharing solved "many workload classes fighting over one cluster" (contention-domain splitting), letting WLM be tuned per workload instead of one-size-fits-all. Serverless solved "risk-and-trading load is unpredictable" (the elasticity problem). Under the single-cluster model one bad query's cost was socialized across the company; slicing contention domains turned "why is this query slow" from folklore into something attributable. Event micro-batching moved "T+1 daily reports" toward near-real-time, which batch architecture could never deliver.
- 教训
单集群扛多类负载时,先切分争用域再谈调优;RA3、数据共享、Serverless 各自回答一个问题,不要指望一次架构升级全解决;给每一类负载找到"最便宜且够用"的计算形态(稳定 ELT 用 provisioned、波动分析用 Serverless),比"一个大集群打天下"更省;用业务指标(SLA、Query Efficiency)而不是集群 CPU 来验收数仓改造。
When one cluster carries many workload classes, split contention domains before tuning. RA3, data sharing, and Serverless each answer one question - don't expect one architecture upgrade to fix everything. Matching each workload class to its cheapest sufficient compute shape (provisioned for steady ELT, Serverless for spiky analytics) beats "one big cluster for everything." Judge warehouse rebuilds by business metrics (SLAs, Query Efficiency), not cluster CPU.
来源
AWS Big Data Blog "How FanDuel adopted a modern Amazon Redshift architecture to serve critical business workloads" (Nov 22, 2023, co-written with FanDuel principal data architects Sreenivasa Mungala and Matt Grimm
—
相关产品:Amazon Redshift 相关能力:— 最后核验:2026-10-02
McDonald's:从 Teradata 迁到 Redshift,十大市场日报从 3 小时压到 9 秒 成功经验
餐饮零售
Teradata 迁移
每日经营报表
云数仓
- 场景
McDonald's 在 65 个市场经营数千家餐厅,需要跑每日高管报表、周趋势报告,还要回答菜单与促销、厨房设计、供应链等复杂业务问题。数据量与数据类型持续增长,遗留 Teradata 数仓的扩展与成本跟不上节奏,公司决定迁往云数仓,并要求迁移后仍是"唯一可信数据源"。
McDonald's operates thousands of restaurants across 65 world markets and needs daily executive reports, weekly trend reporting, plus answers to hard business questions about menus and promotions, kitchen design, and supply chains. Data volumes and variety kept growing while the legacy Teradata warehouse's scaling economics fell behind, so the company decided to move to a cloud data warehouse - with the requirement that it remain the single source of truth afterward.
- 决策
在数据咨询商 Wavicle 主导下,把数仓从 Teradata 迁到 Amazon Redshift,用 Talend 做数据集成、Tableau 做可视化。Redshift 承载每日/每周趋势报告与更复杂的分析,覆盖数千家餐厅的经营数据。
Led by data consultancy Wavicle, the warehouse moved from Teradata to Amazon Redshift, with Talend for data integration and Tableau for visualization. Redshift carries daily and weekly trend reporting plus more sophisticated analytics across thousands of restaurants.
- 结果
十大市场的每日报表从 3–4 小时压到 9 秒;test-to-market(新品上市测试)项目的数据分析效率提升 87%(Wavicle 口径);落地用例包括产品组合分类、供应商质量与成本洞察、季节性销售预测、第三方外卖平台 KPI 分析等。
Daily reports for the top 10 markets went from 3-4 hours to 9 seconds; test-to-market initiatives saw an 87% improvement in data-analytics efficiency (Wavicle's figure). Live use cases include product-mix classification, supplier quality and cost insight, seasonal sales forecasting, and third-party delivery KPI analysis.
- 机制根因
Teradata 时代的瓶颈不在"算不动",而在批处理窗口与"为峰值买单"的成本结构:日报跑 3–4 小时,本质是行存架构 + 固定容量下全量扫描的代价。Redshift 的列存 MPP 把"每日全量经营报表"变成可并行裁剪的扫描,9 秒是列存 + 排序键裁剪共同作用的结果。迁移的真正价值不只是换引擎,而是借机把工具链一起现代化——旧 ETL 与新数仓的接口标准化后,"加一个分析用例"不再需要数仓团队排期。
The Teradata-era bottleneck was not "can't compute" but the batch window and the "pay for peak" cost structure: a 3-4-hour daily report is the price of row-store architecture plus full scans on fixed capacity. Redshift's columnar MPP turned "daily full operations reports" into parallelizable, prunable scans - 9 seconds is columnar storage plus sort-key pruning working together. The migration's real value went beyond swapping engines: modernizing the toolchain at the same time meant adding a new analytics use case no longer required warehouse-team scheduling once legacy ETL and the new warehouse spoke through standardized interfaces.
- 教训
遗留数仓迁移的 ROI 往往不在"查询快了 X 倍",而在"报表 SLA 从小时级到秒级"带来的决策节奏变化;选云数仓时把 ETL/BI 工具链一起规划,避免"新仓配旧管道"把老问题搬上云;用"十大市场日报 9 秒"这种业务方可感知的指标验收迁移,而不是只看 TPC 式基准。
A legacy-warehouse migration's ROI is usually not "queries got X times faster" but the decision-cadence change from hour-level to second-level report SLAs. Plan the ETL/BI toolchain together with the cloud warehouse choice, or you will lift old problems onto the cloud with "new warehouse, old pipelines." Accept a migration by business-felt metrics like "top-10-market daily reports in 9 seconds," not by TPC-style benchmarks alone.
来源
Wavicle Data Solutions case study "McDonald's: Data Driven Solutions" (Teradata to Redshift migration, Talend + Tableau
—
相关产品:Amazon Redshift 相关能力:— 最后核验:2026-10-02
Nasdaq:从 70 节点 Redshift 集群到湖仓架构,账单流程 40 分钟压到 4 分钟(2014–2019) 成功经验
证券交易所
市场数据
湖仓一体
Redshift Spectrum
夜间批处理
- 场景
Nasdaq 运营 27 个市场、近 4000 家上市公司,每晚要把当天的订单、报价、交易、撤单等受保护交易数据全部入库,在次日开市前跑完账单、报表和监管报送。2014 年,为扩大规模、提升性能、降低运维成本,Nasdaq 把遗留本地数仓迁到 Amazon Redshift。到 2018 年集群扩至 70 个节点,每晚从数千个数据源摄取 300–550 亿条记录、超过 4TB;2018 年初市场波动加剧时单日峰值约 550 亿条。硬约束是时间窗口:收盘到次日开市之间既要写完数百亿条,又要同时读出来做报表——"数据加载拖慢了我们报表的交付"(软件工程副总裁 Robert Hunt)。
Nasdaq operates 27 markets with nearly 4,000 listed companies. Every evening it must ingest the day's protected trading data - orders, quotes, trades, cancellations - and finish billing, reporting, and surveillance before the next market open. In 2014, to gain scale and performance and lower operational costs, Nasdaq moved from a legacy on-premises data warehouse to Amazon Redshift. By 2018 the cluster had grown to 70 nodes, ingesting 30-55 billion records and over 4TB nightly from thousands of sources, peaking at about 55 billion records in a single day during the early-2018 volatility spike. The hard constraint was the time window: hundreds of billions of records had to be both written and read between market close and the next open - "Data loading delayed the delivery of our reports" (Robert Hunt, vice president of software engineering).
- 决策
2019 年 1 月 Nasdaq 参加 AWS Data Lab,用 4 天时间把数仓实现推倒重来:Redshift 退为纯计算层,所有交易所数据先以消息形式写入 S3 归档,再用 Redshift Spectrum 直接在 S3 上查询——即湖仓架构,账单、报表、监控等下游流程都由 S3 上的消息驱动。
In January 2019 Nasdaq joined an AWS Data Lab and rebuilt its warehousing implementation in four days: Redshift stepped back to a pure compute layer, all exchange data was first written to S3 as archived messages, and Amazon Redshift Spectrum queried S3 directly - a lake house architecture with billing, reporting, and surveillance driven by the S3 messages downstream.
- 结果
据 AWS 案例页引述 Robert Hunt,一个账单流程从约 40 分钟降到 4 分钟(降幅 90%);全部流程平均提升 60–70%。摄取侧不再受数仓写入吞吐卡脖子,查询侧可以直接对 S3 做并行查询。
Per the AWS case study quoting Robert Hunt, one billing process went from about 40 minutes to 4 minutes (a 90% improvement); processes overall improved 60-70% on average. Ingest was no longer choked by warehouse write throughput, and queries could run in parallel directly against S3.
- 机制根因
改造的本质是把"存"和"算"解耦。70 节点集群的膨胀是存算耦合时代的典型症状:数据量涨只能加节点,而节点自带计算,于是"为存数据被迫买计算"。S3 接管持久化与并行摄取后,Redshift 只保留弹性计算,Spectrum 让冷热数据共享同一套 SQL;夜批窗口里"加载与查询抢同一批资源"的死锁被物理隔离打破。这正是后来 RA3(本地 SSD 缓存 + S3 托管存储)要产品化的思路,Nasdaq 在 2019 年用手工架构先走了一遍。
The rebuild decoupled storage from compute. The 70-node sprawl was the classic symptom of the coupled era: data growth forced node additions, and every node carried compute, so the team was "buying compute to store data." With S3 owning durable, parallel ingest and Redshift keeping only elastic compute, Spectrum gave hot and cold data one SQL surface, and the overnight deadlock of "loads and queries fighting over the same resources" was broken by physical isolation. This is exactly the idea RA3 (local-SSD cache + S3 managed storage) would later productize - Nasdaq walked it first with a hand-built architecture in 2019.
- 教训
数仓扩到几十个节点时,先问"节点里多少算力是为存储买单的",而不是继续加节点;"写多读少、窗口固定"的夜批场景,S3 + 计算层的组合优于"全量 COPY 进仓";架构改造的验收标准应该是单个关键流程的量化提速(如 40 分钟→4 分钟),而不是感觉"变快了"。
When a warehouse grows to dozens of nodes, ask how much of each node you bought for storage before adding more nodes. For "write-heavy, fixed-window" overnight batches, S3 plus a compute layer beats "COPY everything into the warehouse." Judge an architecture rebuild by a quantified speedup on one critical process (40 minutes to 4), not by a feeling that things got faster.
来源
AWS official case study "Nasdaq Case Study" (quoting Robert Hunt
—
AWS "Nasdaq Migrates to a More Modern Data Lake Architecture" (2019 Data Lab + Spectrum lake house rebuild
—
相关产品:Amazon Redshift 相关能力:— 最后核验:2026-10-02
Robin:为嵌入式分析迁出 Redshift,SQL 语义暗坑比数据搬运更费工(2023) 失败教训
Redshift 迁出
嵌入式分析
SQL 方言差异
迁移校验
维护窗口
- 场景
Robin(工程师 Eric Wurtzbacher 2023 年 9 月第一手复盘)。数仓跑在 Redshift 上,管道是 RDS 导出(EventBridge 编排)→ Python 脚本加载 → dbt 转换。2023 年 1 月团队决定迁往 Snowflake,三个动因:① 即将接入第三方嵌入式分析工具,负载将大增,Redshift 单集群扛不住;② 降本——餐巾纸测算 Snowflake 更便宜,迁移四个月后验证"确实便宜了";③ 不再为扩缩容的维护窗口和停机操心——产品面向付费客户,有 uptime 协议要守。此外还看中 zero-copy-cloning、time travel、内置监控等"quality of life"改进。
Robin (first-hand retrospective by engineer Eric Wurtzbacher, September 2023). The warehouse ran on Redshift with a pipeline of RDS exports (orchestrated by EventBridge), a Python loading script, and dbt transformations. In January 2023 the team decided to move to Snowflake for three reasons: (1) an upcoming third-party embedded analytics tool would sharply increase load on the single Redshift cluster; (2) cost - back-of-the-napkin math said Snowflake would be cheaper, and four months after migration "this has held true"; (3) no more maintenance windows or downtime when scaling up or down - the product faces paying customers with uptime agreements to honor. They also looked forward to quality-of-life improvements like zero-copy cloning, time travel, and integrated monitoring.
- 决策
双仓并行、逐服务切换。Snowflake 账号用 Terraform 起;Python 加载脚本把 boto3 Redshift driver 换成 Snowflake connector,COPY 换成 COPY INTO(ON_ERROR=CONTINUE 救了不少脏数据);dbt-redshift 包换成 dbt-snowflake,全量 SQL 翻译——作者直言"这部分是最多的工作"。QA 不追求 100% 对等:所有表行数一致、每列数值加总一致、字符串列 distinct 计数一致,差异能解释即可。
Run both warehouses in parallel and switch service by service. The Snowflake account was stood up with Terraform; the Python loader swapped the boto3 Redshift driver for the Snowflake connector, COPY became COPY INTO (ON_ERROR=CONTINUE saved plenty of dirty rows); the dbt-redshift package became dbt-snowflake with full SQL translation - "this section was most of the work," the author writes. QA did not chase 100% parity: row counts match, every column aggregates to the same value, string columns count-distinct the same - differences just had to be explainable.
- 结果
5 个月完成迁移,作者认为在性能、成本、开发者体验三方面都值得。迁移中最大的时间黑洞不是数据搬运,而是 SQL 边缘语义差异:greatest()/least() 遇到 null 时 Redshift 直接忽略、Snowflake 返回 null;Redshift 时间戳基准精度是 1 位小数、Snowflake 是 3 位,cast 成字符串再 md5 哈希会对不上;墨西哥 2022 年取消夏令时——Redshift(尽管每周维护)没跟进时区库变更,Snowflake 跟进了;'infinity' 时间戳 Snowflake 不认;type、start 这类保留关键字作列名必须加引号。
The migration took about 5 months and was judged worth it on performance, cost, and developer experience. The biggest time sink was not moving data but SQL edge-semantic differences: Redshift's greatest()/least() ignores nulls while Snowflake's returns null; Redshift timestamps have a base scale of 1 decimal place vs Snowflake's 3, so casting to strings for md5 hashing mismatches; when Mexico abolished daylight saving time in 2022, Redshift (despite weekly maintenance) never picked up the timezone change while Snowflake did; 'infinity' timestamps are rejected by Snowflake; reserved keywords like type and start used as column names need quoting.
- 机制根因
Redshift"单集群共享大脑"模型下,嵌入式分析这种不可控的外部负载没有隔离手段,只能靠扩缩容硬扛,而扩缩容=维护窗口=可能的停机——触碰了面向付费客户产品的红线,这是三动因里最结构性的一条。SQL 暗坑的本质是两家对 SQL 标准边缘语义(null 传播、时间戳精度、时区库版本)的实现分歧:dbt 包替换只是开始,真正的成本在逐函数的语义核对与"可解释差异"的 QA 哲学。
Under Redshift's "one shared brain per cluster" model there is no isolation lever for uncontrollable external load like embedded analytics - only scaling through it, and scaling means maintenance windows and possible downtime, which crossed the red line for a paying-customer-facing product. That is the most structural of the three drivers. The SQL pitfalls are implementation divergences on the edges of the SQL standard (null propagation, timestamp precision, timezone-library versions): swapping the dbt package is only the start; the real cost is per-function semantic verification and a QA philosophy of "explainable differences."
- 教训
评估迁仓成本时,SQL 翻译 + QA 的人力往往超过数据搬运本身,预算要按"逐函数核对"来估;迁移前先列边缘语义清单(null 行为、时间戳精度、时区、保留字、JSON null 类型)逐项验证;"可解释的差异"比"100% 一致"更务实,100% 对等"需要不可思议的努力"(作者原话);当负载增长来自"不可控的第三方"(嵌入式分析)而非自身业务时,弹性隔离比单集群调优更重要。与 Faire 的九个月零停机双仓迁移(重工程)相比,Robin 这个样本重的是"为什么离开"与"翻译税"——同一迁出潮的另一面。
When costing a warehouse migration, SQL translation plus QA labor usually exceeds the data-moving work - budget for per-function verification. List the edge-semantic checklist (null behavior, timestamp precision, timezones, reserved words, JSON null types) and verify each before cutover. "Explainable differences" beats "100% identical" as a QA target - full parity "would take an incredible amount of effort" (author's words). When load growth comes from an uncontrollable third party (embedded analytics) rather than your own business, elastic isolation matters more than single-cluster tuning. Compared with Faire's nine-month zero-downtime dual-warehouse migration (an engineering story), Robin is the other side of the same exodus: why teams leave, and the translation tax.
来源
Eric Wurtzbacher (Robin), "The Process for our Redshift to Snowflake Migration" (Sep 21, 2023
—
相关产品:Amazon Redshift、Snowflake 相关能力:迁移潮:"反例价值"本身 最后核验:2026-10-02
Salesforce "Sayonara":十年去 O 战略(2013–2023,目标未证实完成) 成功经验
去 O
商业数据库锁定
长期主义
PostgreSQL
- 决策
Salesforce launched an internal database project codenamed "Sayonara" (Japanese for "goodbye"), targeting full independence from Oracle by 2023 (per The Information, citing a former employee); in parallel it did two things: starting in 2013 it hired core PostgreSQL developers (including Tom Lane), bringing world-class PG engineering in-house; and new business lines (Heroku, MuleSoft, etc.) were built off Oracle from day one. (
https://www.infoworld.com/article/2265383/how-postgresql-just-might-replace-your-oracle-database.html)
- 结果
2019 年披露时目标 2023 年完成;但截至 2026-10,公开渠道未见"核心 CRM 已下线 Oracle"的宣告,多方分析认为主平台仍在 Oracle 上(Cloud Wars;cirra.ai 2023 年分析)。这是一次"未完成"的去 O——但战略动作本身是教科书级的,结论以公开报道口径为准。本案例标为"成功经验",指的是其方法论(能力建设路径)值得学习,不可引用为"Salesforce 已用 PostgreSQL 替换 Oracle 核心"的证据——后者截至 2026-10 仍无公开证实。
When disclosed in 2019 the target was completion by 2023; but as of 2026-10, no public announcement of "core CRM off Oracle" has surfaced, and multiple analyses conclude the main platform still runs on Oracle (Cloud Wars; cirra.ai 2023 analysis). This is an "unfinished" Oracle exit — but the strategy itself is textbook. Conclusions follow public reporting. This case is filed as a "success story" for its methodology (the capability-building path), which is worth studying — it must not be cited as evidence that "Salesforce has replaced Oracle with PostgreSQL at its core", which remains publicly unconfirmed as of 2026-10.
- 机制根因
Salesforce 的多租户架构本就把 Oracle 当"存 EAV 大宽表的 KV"用——关系能力早被自己的元数据抽象层架空,Oracle 的价值只剩"稳定+售后"。去 O 的真正资产因此不是一次迁移,而是把数据库从"供应商产品"变成"自有工程能力":招 PG 内核级人才、自研抽象层、新业务先行、核心系统最后动。代价是时间:2013 年续约 Oracle 到 2023 年目标,整整十年窗口——"换系统记录数据库"是公司史上最大变更(Wikibon),只能按代际推进,不能按季度考核。
Salesforce's multi-tenant architecture already used Oracle as "a KV store for giant EAV wide tables" — the relational capabilities had long been hollowed out by its own metadata abstraction layer, leaving Oracle's value as "stability + support." So the real asset of going off-Oracle was never a single migration but turning the database from a "vendor product" into an "in-house engineering capability": hire PG kernel-level talent, build your own abstraction layer, new businesses first, core system last. The price is time: from the 2013 Oracle contract renewal to the 2023 target is a full decade — swapping the system-of-record database is "the largest change in Salesforce's history" (Wikibon); it moves at generational pace, not quarterly pace.
- 教训
去 O 不要立项成"迁移项目",要立项成"能力建设":先有人(内核级人才)、再有抽象层(让上层不感知底层库)、新业务先行、核心系统最后动。连 Salesforce 都需要十年窗口,CTO 对"一年去 O"的承诺要打骨折听。
Don't charter off-Oracle as a "migration project" — charter it as "capability building": people first (kernel-level talent), then an abstraction layer (so upper layers never feel the underlying database), new businesses first, core system last. If even Salesforce needs a ten-year window, a CTO should heavily discount any "off-Oracle in one year" promise.
相关产品:Oracle Database(甲骨文)、PostgreSQL(社区版) 相关能力:出走潮 —— Amazon 下线 Oracle 与"每两周停机 10 分钟做 schema 变更" 最后核验:2026-10-01
Sentry 迁向 ClickHouse:事件流的 MVCC 写放大 失败教训
实时分析/OLAP
写入放大
高吞吐
运维简单
- 场景
Sentry, an error-tracking SaaS, ingests a continuous high-write event stream and searches it by multi-dimensional tags, with a team that wanted simple operations. Early on it stored events in PostgreSQL and used Redis for tag indexes — but the high-write stream generated massive numbers of dead tuples under PG's MVCC, wasting I/O on scanning dead rows. The team evaluated Impala, Druid, Pinot, Presto, Drill, BigQuery, Spanner, and Spark Streaming, then chose ClickHouse with its own Snuba query layer. (
https://blog.sentry.io/introducing-snuba-sentrys-new-search-infrastructure/)
- 决策
2019 年把事件与 tag 搜索迁往 ClickHouse / Snuba,下线 PostgreSQL + Redis 方案。
In 2019 event and tag search moved to ClickHouse / Snuba, retiring the PostgreSQL + Redis design.
- 结果
Tagstore 数据从 TB 级压缩到 GB 级;约 40% 查询(告警规则)要求写后立即可读,迁移后满足。数字为 Sentry 自报,属单方口径。
The tagstore shrank from terabytes to gigabytes; roughly 40% of queries (alerting rules) required read-your-writes, which the migration satisfied. Figures are from Sentry's own engineering blog — self-reported.
- 机制根因
Sentry 的死元组问题是复合的,不是"每次 INSERT 都制造死行":tagstore 对聚合计数的高频 UPDATE、保留期到期时的批量删除、每新增一种查询维度就要加索引/反范式结构并回填——三者叠加制造死元组潮汐,autovacuum 跟不上,I/O 被浪费在扫描死行上。注意纯 INSERT 本身不直接产生死行。ClickHouse 侧:主键排序加列式压缩带来数量级压缩比;"不耍小聪明"的查询规划器加 PREWHERE 让延迟可预测;复制只需 ZooKeeper;写后立即可读满足告警实时性。
Sentry's dead-tuple problem was composite — not "every INSERT creates a dead row": frequent UPDATEs to aggregate counters in the tagstore, bulk deletes when retention periods expired, and new indexes/denormalized structures plus backfills for every new query dimension combined into a dead-tuple tide that autovacuum couldn't keep up with, wasting I/O on scanning dead rows. Note that a plain INSERT by itself does not directly create dead rows. On the ClickHouse side: primary-key sorting plus columnar compression delivers order-of-magnitude compression; a query planner that "doesn't try to be clever" plus PREWHERE makes latency predictable; replication needs only ZooKeeper, keeping ops simple; and read-your-writes satisfies alerting's real-time demand.
- 教训
append-mostly 事件流的 OLAP,列式加排序键的压缩与扫描红利远超行存 MVCC;选型时把"运维简单"和"写后立即可读"列为硬指标;存储引擎的写入范式必须与 workload 的写入范式同构。
For append-mostly event OLAP, the compression and scan dividends of columnar + sort keys far outweigh row-store MVCC. Put "operational simplicity" and "read-your-writes" on the hard-requirements list during selection — not just query features. The storage engine's write paradigm must be isomorphic to the workload's write paradigm.
相关产品:PostgreSQL(社区版)、ClickHouse 相关能力:海量事件的 ad-hoc 全表扫描、MVCC + 查询优化器 + 事务性 DDL —— 复杂分析直接跑在 OLTP 主库上 最后核验:2026-10-01
Shopify 的 MySQL 分片与 Pod 隔离架构(2015–2018) 成功经验
水平分片
故障隔离
多租户
在线迁移
- 场景
By 2015 Shopify had hit the end of buying bigger database servers and was forced to shard MySQL by merchant. Sharding introduced a new problem: cross-shard fan-out scattered through the codebase — when one shard went down, the affected operations became unavailable platform-wide. The more shards, the bigger the blast radius of a single failure. (
https://shopify.engineering/a-pods-architecture-to-allow-shopify-to-scale)
- 决策
2016 年重组运行时架构,引入 Pod:每个 Pod 是一组商户 + 完全隔离的数据存储集合(MySQL 分片、Redis 等);负载均衡层的 Sorting Hat 按规则把每个请求路由到唯一 Pod,禁止任何跨 Pod 调用;配套自研 Ghostferry 做在线数据迁移、Pod Mover 做分钟级容灾切换。
In 2016 Shopify reorganized its runtime around pods: each pod is a set of merchants plus a fully isolated set of datastores (MySQL shards, Redis, etc.); a Sorting Hat in the load balancers routes every request to exactly one pod, and no cross-pod calls are allowed; the in-house Ghostferry handles online data migration and Pod Mover handles disaster-recovery cutover within about a minute.
- 结果
据官方博客,单个 Pod 故障不再扩散为平台事故;Pod Mover 可在一分钟内把 Pod 切到恢复数据中心且不丢请求;Ghostferry 已在生产环境迁移数十 TiB 级分片数据(据 Shopify 在 Percona Live 的分享口径,70+ TiB)。
Per the official blog, a single pod failure no longer cascades into a platform outage; Pod Mover can shift a pod to its recovery datacenter in about a minute without dropping requests or jobs; Ghostferry has moved tens of TiBs of sharded data in production (70+ TiB per Shopify's Percona Live talk).
- 机制根因
分片解决的是"容量",Pod 解决的是"爆炸半径"——这是两个正交问题。Shopify 的洞察是:多租户 SaaS 里真正的故障域单位不是"数据库实例"而是"租户集合";把计算、存储、任务队列按租户集合做硬隔离后,扩容=加 Pod(线性),故障=丢 Pod(有界)。代价是彻底放弃跨分片事务与跨 Pod 查询,所有跨商户分析走 CDC 进 Kafka 再入数仓,应用层必须接受最终一致。
Sharding solves "capacity"; pods solve "blast radius" — two orthogonal problems. Shopify's insight: in multi-tenant SaaS the real failure-domain unit isn't the "database instance" but the "tenant set"; hard-isolating compute, storage, and job queues by tenant set makes scaling = add pods (linear) and failures = lose pods (bounded). The price is giving up cross-shard transactions and cross-pod queries entirely — all cross-merchant analytics flow through CDC into Kafka and then the warehouse, and the application must accept eventual consistency.
- 教训
分片之后必须回答"分片挂了影响谁";如果答案是"全平台",分片只做了一半。租户天然可分的业务(电商、SaaS),按租户集合做硬隔离是性价比最高的扩展模型。
After sharding you must answer "who is affected when a shard dies"; if the answer is "the whole platform", the sharding job is only half done. For businesses with naturally separable tenants (e-commerce, SaaS), hard isolation by tenant set is the most cost-effective scaling model.
相关产品:MySQL 相关能力:— 最后核验:2026-10-01
Slack:不换 MySQL,加 Vitess 代理层解决分片键僵化 成功经验
水平扩展
分片键演进
大客户热点
迁移工程
- 场景
A team-chat app with MySQL sharded by workspace at the application layer; oversized customer workspaces became shard hotspots, and application-layer sharding had welded the shard key into the code, unable to evolve with the business (URL cited from a PlanetScale article; open it in a real browser to confirm the original is reachable:
https://slack.engineering/scaling-datastores-at-slack-with-vitess)
- 决策
引入 Vitess 做代理层,而不是重写应用分片逻辑;历时约 3 年灰度完成迁移。
Introduce Vitess as a proxy layer instead of rewriting the application's sharding logic; the migration took about 3 years of gradual rollout.
- 结果
99% 的 MySQL 流量经 Vitess;峰值从 0 做到 230 万 QPS;分片键可按 user、channel、workspace 灵活重构,并向上游社区贡献代码。
99% of MySQL traffic goes through Vitess; peak grew from zero to 2.3M QPS; the shard key can be flexibly restructured by user, channel, or workspace, and the team contributed code back upstream.
- 机制根因
应用层分片的最大代价不是性能而是僵化——分片键焊死在代码里,大客户热点无解。Vitess 代理层把路由与业务逻辑解耦,支持在线 DDL 与在线 reshard;MySQL 本体不动,事务语义、生态工具与运维知识全部复用。
The biggest cost of application-layer sharding isn't performance but rigidity — a shard key welded into code leaves large-customer hotspots unsolvable. The Vitess proxy layer decouples routing from business logic, enabling online DDL and online resharding; MySQL itself stayed untouched, so transactional semantics, ecosystem tooling, and operational knowledge were all reused.
- 教训
"不换"的机制条件:MySQL 本体的事务与生态仍是资产,只是路由层僵化——此时加代理层而非换库。分片键需要随业务演进时,代理层比应用层分片灵活一个数量级。大迁移用 3 年灰度是正常节奏,不是失败;能灰度、可回滚的迁移设计本身就是选型的一部分。
The mechanism condition for "not switching": MySQL's transactions and ecosystem are still assets — only the routing layer is rigid, so add a proxy layer rather than switching databases. When shard keys must evolve with the business, a proxy layer is an order of magnitude more flexible than application-layer sharding. A 3-year gradual migration is a normal cadence, not a failure; a migration design that can roll out gradually and roll back is itself part of the selection decision.
相关产品:MySQL 相关能力:Vitess 水平分片 —— "第二增长曲线" 最后核验:2026-10-01
Adobe:在 Snowflake 上组装组合式 CDP,数据不动、受众联邦 成功经验
营销科技
组合式 CDP
零拷贝
实时受众
- 场景
营销数据散在 CDP、数仓和业务系统里,传统做法是把数仓数据再拷一份进 CDP:双份存储、双份治理,还有延迟。隐私与数据安全的要求又让"拷来拷去"越来越贵,营销部门要实时个性化,数据团队却被困在搬运里。
Marketing data was scattered across the CDP, the warehouse, and business systems; the traditional fix copied warehouse data into the CDP a second time — double storage, double governance, plus latency. Privacy and security demands made "copying things around" ever more expensive, while marketing wanted real-time personalization and data teams stayed stuck on plumbing.
- 决策
Adobe 与 Snowflake 共建组合式 CDP。Federated Audience Composition 让营销人员直接在 Snowflake 的企业全量数据上创建、丰富受众,数据不进 Adobe;profile enrichment 把 Snowflake 的属性与聚合喂给实时画像,做"当下"个性化;双向零拷贝同步让 Real-Time CDP 的画像数据在 Snowflake 侧直接可用;Customer Journey Analytics 与 Snowflake 间的数据镜像让增删改自动反射。
Adobe and Snowflake co-built a composable CDP. Federated Audience Composition lets marketers create and enrich audiences directly on the enterprise's full Snowflake data without it entering Adobe; profile enrichment feeds Snowflake attributes and aggregates into real-time profiles for in-the-moment personalization; bi-directional zero-copy sync makes Real-Time CDP profile data directly usable inside Snowflake; data mirroring between Customer Journey Analytics and Snowflake reflects inserts, updates, and deletes automatically.
- 结果
受众创建与模型训练不再需要搬运数据;同一份企业数据既服务营销激活,也服务数仓内的分析与 ML;治理做一次,两边生效。Adobe 引述调研称 61% 的营销与技术高管认为更个性化的体验是 2025 年增长主因——该数字为 Adobe 引述的第三方调研口径。
Audience building and model training no longer require moving data; the same enterprise dataset serves marketing activation and in-warehouse analytics/ML alike; governance done once applies on both sides. Adobe cites survey data that 61% of marketing and technology executives see more personalized experiences as a primary 2025 growth driver — a third-party survey figure quoted by Adobe.
- 机制根因
"把计算推到数据处"替代"把数据搬到计算处"。联邦查询下,权限、血缘不断裂,没有第二份拷贝可泄露、可过期。组合式 CDP 的本质是承认"数据重力":企业全量数据在 Snowflake,CDP 应该长在数据上,而不是把数据吸进 CDP。
"Push compute to the data" replaces "drag data to the compute." Under federated queries, permissions and lineage never break — there is no second copy to leak or go stale. The composable CDP's essence is accepting data gravity: the enterprise's full data lives in Snowflake, so the CDP should grow on top of the data rather than sucking the data into the CDP.
- 教训
CDP 选型先看数据重力,谁离全量数据近谁赢;复制即负债——每份拷贝都是未来的对账与治理成本;零拷贝架构下,安全与合规反而更容易,因为永远只管一个源头。
Evaluate CDPs by data gravity first — whoever sits closest to the full dataset wins. Copying is liability: every duplicate is future reconciliation and governance cost. Under zero-copy architecture, security and compliance get easier, because there is only ever one source to govern.
来源
Adobe 官方商业博客《Adobe and Snowflake partner to power a composable CDP》(Adobe 第一方
Adobe official business blog, "Adobe and Snowflake partner to power a composable CDP" (first-party Adobe
相关产品:Snowflake 相关能力:Secure Data Sharing + Marketplace 最后核验:2026-10-02
Capital One:从 Teradata 迁往 Snowflake,先给"无限弹性"装上刹车(2017–2022) 成功经验
银行核心数仓
云迁移
成本治理
联邦治理
- 场景
Capital One 的数据仓库多年前跑在 Teradata 一体机上,存储与并发处处受限;一次版本升级是长达六个月的流程——厂商把硬件发到机房安装,伴随一到两周停机。2017 年起,该行启动向 Snowflake 的整体迁移,目标是把整个数据工作负载搬上公有云。
Capital One's data warehouse long ran on Teradata appliances, constrained on storage and concurrency; a single version upgrade was a six-month ordeal — the vendor shipped hardware to its data centers and installed it, with one to two weeks of outage. Starting in 2017, the bank began a full migration to Snowflake, aiming to move its entire data workload onto the public cloud.
- 决策
Capital One 预判 Snowflake"存储计算无限"的架构是双刃剑:过去按年买 license,用多用少一个价;上云后用多少付多少,不建治理必失控。于是迁移同时自建联邦式治理工具:业务线自助开通、自主管理,成本策略与最佳实践由中央工具强制执行。这套工具后来产品化为 Slingshot。
Capital One judged Snowflake's "unlimited storage and compute" architecture a double-edged sword: the old world bought annual licenses where usage barely mattered; in the cloud you pay for what you use, and without governance, cost and data sprawl would run wild. So alongside the migration it built federated governance tooling: lines of business self-provision and manage their own environments, while cost policies and best practices are enforced by central tooling. That tooling was later productized as Slingshot.
- 结果
双仓并行一年多完成切换,Teradata 1.8 万张表只迁了 6000 张——迁移顺带做了资产盘点。据工程副总裁 Salim Syed 口径:消除超 5.5 万小时手动变更(另一次专访口径为近 5 万小时),查询成本降 43%,整体节省近 27%;数仓从 200TB 扩到 50PB(250 倍),查询量涨 5–6 倍,6000–7000 用户每天跑数百万查询。2022 年这套能力被装进新成立的 Capital One Software 对外售卖。
The cutover completed after more than a year of running both warehouses in parallel — and only 6,000 of Teradata's 18,000 tables were migrated, since the move doubled as a data-asset rationalization. Per VP of Engineering Salim Syed: over 55,000 hours of manual changeover eliminated (a separate interview cites nearly 50,000 hours), query costs down 43%, and nearly 27% overall cost savings; the warehouse grew 250-fold from 200TB to 50PB, query volume rose 5–6x, and 6,000–7,000 users run millions of queries a day. In 2022 the capability was packaged into the newly formed Capital One Software and sold externally.
- 机制根因
弹性计费把"用量纪律"写进了成本公式。分析师习惯把数据子集拷进私人沙箱(内部戏称 "data drunk"),存储与计算在无人察觉中膨胀。根因是旧世界买断制养成的用量习惯,与新世界按量付费的错配。联邦治理的本质是把成本意识编译进工具链:配额、告警、低效查询提醒默认开启,而不是靠事后审计。
Elastic billing writes "usage discipline" into the cost formula. Analysts habitually copy data subsets into private sandboxes ("data drunk," in internal parlance), and storage plus compute inflate unnoticed. The root cause is a mismatch between usage habits formed under perpetual licenses and pay-as-you-go pricing. Federated governance compiles cost-awareness into the toolchain itself: quotas, alerts, and inefficient-query nudges on by default, instead of after-the-fact audits.
- 教训
迁移第一天就要设计治理,事后"把精灵放回瓶子"极难;数据民主化的前提是护栏而非审批——让业务自助,把最佳实践做成默认项;小团队支撑数千用户的唯一出路是工具化,不是加人。
Design governance on day one of a cloud warehouse migration — putting the genie back in the bottle afterwards is brutally hard. Data democratization needs guardrails, not approvals: let the business self-serve, but make best practices the default. The only way a small team supports thousands of users is tooling, not headcount.
来源
SiliconANGLE, "Capital One sees win-win in selling software built on Snowflake cloud" (Jul 9, 2022 — Teradata upgrade outages, 6,000/18,000 tables, 55,000 hours, −43% query cost, 200TB→50PB, Slingshot productization)
https://siliconangle.com/2022/07/09/capital-one-sees-win-win-selling-software-built-snowflake-cloud/
相关产品:Snowflake 相关能力:弹性计费的"隐形闲置税" 最后核验:2026-10-02
Faire:从 Redshift 迁往 Snowflake,九个月零停机大迁移(2022–2023) 成功经验
批发电商
Redshift 迁移
零停机迁移
数据治理
- 场景
Faire 2017 年成立,数年做到 1000 多人、60 万零售商、8.5 万品牌,数据需求涨得比人快。数仓跑在 Redshift 上,支撑 100 多个 Airflow ETL、Mode 报表和 SageMaker 机器学习。Redshift 存算耦合,所有负载抢同一集群;且当时只支持串行隔离级别——每天 9 点 Mode 报表的长 SELECT 会阻塞 ETL 的 ALTER TABLE RENAME,一个慢查询能拖住整条队列,WLM 调优成了玄学。
Founded in 2017, Faire grew within years to 1,000+ people serving 600K retailers and 85K brands, with data demand outpacing headcount. Its warehouse ran on Redshift, supporting 100+ Airflow ETLs, Mode dashboards, and SageMaker machine learning. Redshift coupled compute and storage, so every workload fought over one cluster; worse, it then offered only serializable isolation — the daily 9 AM Mode report's long SELECT would block an ETL's ALTER TABLE RENAME, letting one slow query stall the whole queue, and WLM tuning became black magic.
- 决策
迁往 Snowflake,立下硬指标:100% 表迁移、数据对等度 ≥95%、零停机。策略是双仓并行九个月:Kinesis 事件流用 Snowpipe 双写、历史数据从 S3 回填;Stitch 另起一套部署同步 MySQL 等关系源;引入 transformation layer 做 RAW 与 ANALYTICS 逻辑分离;自研 Redshift→Snowflake SQL 解析器,Airflow 双分支自动同步;用 Datafold 做跨仓数据对等校验,并众包给全公司在 Mode 看板上认领。
Migrate to Snowflake with hard targets: 100% of tables moved, ≥95% data parity, zero downtime. The playbook was nine months of dual-running both warehouses: Kinesis event streams dual-written via Snowpipe with historical backfill from S3; a second Stitch deployment syncing MySQL and other relational sources; a new transformation layer separating RAW from ANALYTICS; a homegrown Redshift→Snowflake SQL parser with dual-branch Airflow auto-sync; Datafold for cross-warehouse parity checks, crowdsourced company-wide through a Mode dashboard.
- 结果
九个月完成全量切换(ETL、BI、ML、特征服务),350 多张 Mode 核心报表通过逆向 Mode API 批量迁移,迁移工具开源为 github.com/Faire/snowflake-migration。据博客称,查询排队时间与 Airflow 作业运行时显著改善;如今月均 400 万+ ETL 查询,特征库每天支撑 1 亿次在线预测;全仓 PII 通过 object tagging 加 transformation layer 实现列级脱敏,治理一次做对。
Full cutover in nine months (ETLs, BI, ML, feature serving), with 350+ core Mode reports bulk-migrated via a reverse-engineered Mode API; the migration tooling was open-sourced as github.com/Faire/snowflake-migration. Per the blog, query queue times and Airflow job runtimes improved significantly; today the platform serves 4M+ ETL queries per month, the feature store powers 100M online predictions per day, and warehouse-wide PII masking is enforced at column level through object tagging plus the transformation layer.
- 机制根因
Redshift 的瓶颈不在单查询快慢,而在"所有负载共享一个大脑":锁、WLM 队列、存储空间是全局资源,一个坏查询的代价由全公司分摊。Snowflake 存算分离加多仓隔离把竞争域切小,"查询为什么慢"从玄学变成可观测、可归因——query tag 直接查 information_schema.query_history 就能定位最贵的 DAG。transformation layer 则把"上游结构变更"与"下游稳定"解耦,这是 Redshift 时代缺失的一层。
Redshift's bottleneck was never single-query speed but "every workload sharing one brain": locks, WLM queues, and storage were global resources, so one bad query taxed the whole company. Snowflake's disaggregated storage/compute plus multi-warehouse isolation shrank the contention domain, turning "why is my query slow" from mysticism into something observable and attributable — query tags let anyone find the priciest DAG via information_schema.query_history. The transformation layer decoupled "upstream schema change" from "downstream stability," a layer the Redshift era lacked.
- 教训
迁移首先是组织工程,先拿全员 buy-in 并写进 OKR;能自助就不手工——校验众包、报表迁移工具化决定了九个月的可行性;用 task-specific Airflow operator 约束 SQL,把一致性做进工具而不是文档里。
Migration is an org-engineering project first — secure company-wide buy-in and write it into OKRs. Automate instead of hand-carrying: crowdsourced validation and tooled report migration are what made nine months feasible. Constrain SQL with task-specific Airflow operators; bake consistency into tools, not documents.
相关产品:Snowflake、Amazon Redshift 相关能力:— 最后核验:2026-10-02
JetBlue:Snowflake + Fivetran 现代数据栈,把数据工程师 40% 的时间还给洞察 成功经验
航空
现代数据栈
ELT
客户 360
- 场景
JetBlue 的数据散落在 ServiceNow、Qualtrics、JIRA、Salesforce、第三方 SQL Server 和本地 Oracle 里,是典型的"意大利面式"架构。数据工程总经理 Ashley Van Name 统计:同行的数据工程师约 40% 时间花在搭建和测试管道上,真正产生洞察的时间被严重挤压。
JetBlue's data was scattered across ServiceNow, Qualtrics, JIRA, Salesforce, third-party SQL Servers, and an on-premises Oracle database — classic spaghetti architecture. Ashley Van Name, general manager of data engineering, noted that at peer firms data engineers spend roughly 40% of their time building and testing ingestion pipelines, severely squeezing the time left for actual insight.
- 决策
以 Snowflake 为中心搭建现代数据栈,用 Fivetran 的"管道即服务"(200+ 连接器)替代自建 ETL;坚持"原始数据先进仓"——Qualtrics 飞后问卷、SaaS 数据以原始形态入仓再建模;本地 Oracle 运维库用 Fivetran HVR 实时复制进仓。团队理念是"把数据当产品"。
Build a modern data stack centered on Snowflake, replacing hand-built ETL with Fivetran's pipeline-as-a-service (200+ connectors); insist on "raw data lands first" — post-flight Qualtrics surveys and SaaS data enter the warehouse in raw form before modeling; replicate the on-prem Oracle operations database in real time via Fivetran HVR. The team's mantra: treat data as a product.
- 结果
造管道的时间从 40% 降到 10%,工程师 90% 时间转向数据消费与产品化。飞后问卷直入 Snowflake 原始层,分析师自助取数,推动了机上娱乐与数字交互的改进;运维库实时进仓后可做及时干预;全公司报表逐渐收敛到同一口径,离"单一可信源"更近。
Pipeline-building time fell from 40% to 10%, with engineers spending 90% of their time on data consumption and productization. Post-flight surveys land raw in Snowflake for analyst self-service, driving improvements to in-flight entertainment and digital interactions; the real-time operations feed enables timely maintenance interventions; company-wide reporting is converging on one set of definitions, closer to a single source of truth.
- 机制根因
ELT 把"搬运"标准化外包,把"建模"留给最懂业务的人。自建管道的隐性成本不在开发那几周,而在常年 40% 人力的维护税;Fivetran 把连接器变成商品,Snowflake 的弹性计算吸收峰值波动。原始层进仓保留了"重算"的可能——这是 ETL 先清洗后入库所丢失的东西,也是分析师敢自助的底气。
ELT outsources the commoditized "moving" and keeps the "modeling" with the people who know the business. The hidden cost of hand-built pipelines isn't the weeks of development but the permanent 40%-headcount maintenance tax; Fivetran commoditizes connectors while Snowflake's elastic compute absorbs peaks. Landing raw data preserves the option to recompute — the thing ETL's cleanse-then-load destroys, and the reason analysts dare to self-serve.
- 教训
数据工程师的时间是最贵的成本,省管道钱不如省人力;管道自建是隐性税,标准化环节越早外包越好;"单仓加原始层"是口径统一的前提,没有它,Customer 360 只是口号。这正是 40% 与 10% 之差的来源。
Data engineers' time is the most expensive cost — saving headcount beats saving pipeline dollars. Hand-built pipelines are a hidden tax; outsource standardized links as early as possible. "One warehouse plus a raw layer" is the precondition for consistent metrics; without it, Customer 360 is just a slogan.
相关产品:Snowflake、Oracle Database(甲骨文) 相关能力:— 最后核验:2026-10-02
JLL:Snowflake 遇性能瓶颈,转向 Databricks Lakehouse 统一 80+ 国家分析(2024–2025) 失败教训
性能瓶颈
Lakehouse 迁移
全球分析
变更管理
- 场景
JLL(仲量联行)是全球商业地产服务巨头,旗下 JLL Technologies 部门要把 80 多个国家的数据统一起来,做实时 BI 和 AI 驱动洞察。但原有数仓跟不上:数据量持续增长、查询出现性能瓶颈、传统 BI 工具链受限,全球 120 多名分析师各自为战。全球 BI 技术总监 Kristopher Curtis 的原话是,旧系统在"数据处理的速度和可靠性"上已成主要挑战,迁移不是一次技术升级,而是"维持竞争力的战略必需"。
JLL (Jones Lang LaSalle), a global commercial real estate services giant, needed to unify data across 80+ countries through its JLL Technologies division to deliver real-time BI and AI-driven insights. The incumbent warehouses couldn't keep up: growing data volumes, query performance bottlenecks, and legacy BI limitations, with 120+ analysts working in silos. Kristopher Curtis, JLL's Global BI Technology Director, put it bluntly — the speed and reliability of data processing on legacy systems had become the core challenge, and migrating was "a strategic necessity to maintain our competitive edge," not just a tech upgrade.
- 决策
JLL 与 Databricks 建立战略合作,把分析统一到 Lakehouse:Databricks SQL 承载报表与分析管道,Unity Catalog 做跨 80 多国数据的统一治理,Notebook + SQL dashboard 让分析师和工程师在同一份数据上协作建模全球物业数据、市场趋势与客户 KPI。同时配套做了"Databricks Odyssey"游戏化培训体系——从 Rookie 到 Rockstar 的闯关认证、徽章、积分,把 100 多名员工的换平台阵痛变成组织能力升级。Credly 上由 JLL 官方签发的 JLL Databricks Rockstar/Pro 徽章可佐证该培训体系真实存在。
JLL entered a strategic partnership with Databricks and consolidated analytics on the lakehouse: Databricks SQL for reporting and analytics pipelines, Unity Catalog for unified governance of data spanning 80+ countries, and notebooks plus SQL dashboards so analysts and engineers model global property data, market trends, and client KPIs on the same data. Alongside the platform move, JLL built "Databricks Odyssey," a gamified training program — Rookie-to-Rockstar progression, badges, points — turning the pain of switching platforms into an organization-wide capability upgrade. JLL-issued Databricks Rockstar/Pro badges on Credly corroborate that the program genuinely exists.
- 结果
按 Databricks 官方客户故事口径,切换后同数据源、同数据量下处理时间比原本地数仓降低 60%,数据摄取从小时级降到分钟级;为每个客户配独立集群后,"不再受他人资源消耗影响",跨租户的资源争抢依赖消失。Odyssey 上线几个月内 175 人注册、完成 1250 多次挑战提交、75 人拿证(2024 年 8 月启动)。基于 Databricks 建成的 JLL Azara 成为其商业地产数据洞察平台。以上数字均为 Databricks 厂商口径,未找到第三方独立复现,引用须注明。
Per Databricks' official customer story, processing time dropped 60% versus the previous on-premises warehouse on the same data sources and volumes, ingestion went from hours to minutes, and per-client clusters eliminated cross-tenant resource contention ("you're not impacted by other people's resource consumption"). Within months of its August 2024 launch, Odyssey had 175 enrolled users, 1,250+ challenge submissions, and 75 certifications awarded. JLL Azara, the purpose-built commercial real estate data-and-insights platform, was built on Databricks. All figures are vendor-channel claims with no independent third-party reproduction found — cite with the source noted.
- 机制根因
这是一个"数仓→湖仓"的典型迁移逻辑。Snowflake 这类云数仓在单一市场、单一团队时表现很好,但 JLL 的痛点是三维叠加:80 多国的数据主权与治理复杂度、分析师/工程师/数据科学家三类人群的协作割裂、以及对 AI/ML 的前瞻需求。Snowflake 的封闭存储格式让"一份数据同时服务 BI 和 ML"变得昂贵(要么导数据、要么双份治理),而 Lakehouse 的开放格式(Delta Lake)+ 统一治理(Unity Catalog)把边际成本结构改了:新增一个国家/一条业务线是"加一张表",而不是"再搭一套烟囱"。60% 的提速大概率不只是引擎更快,而是"关停旧链路 + 消除跨租户争抢 + 数据不动"的综合效应——评估这类迁移收益时,要把架构简化放在引擎 benchmark 前面算。
This is the canonical warehouse-to-lakehouse migration logic. A cloud warehouse like Snowflake performs well for a single market or team, but JLL's pain was three-dimensional: governance complexity across 80+ countries, collaboration silos between analysts, engineers, and data scientists, and forward-looking AI/ML demand. Snowflake's closed storage format makes "one copy of data serving both BI and ML" expensive — either move the data or govern it twice — while the lakehouse's open format (Delta Lake) plus unified governance (Unity Catalog) changes the marginal cost structure: adding a country or business line becomes "add a table," not "build another silo." The 60% speedup is most likely a composite of decommissioned legacy pipelines, eliminated noisy-neighbor contention, and no-copy data access — not engine benchmarks alone. When evaluating such migrations, count architecture simplification before engine speed.
- 教训
从 Snowflake 这类成熟数仓迁出的触发条件通常不是"查询慢了 10%",而是"治理、协作、AI 三个维度同时撞墙";Databricks 自己的《Big Book of Data Warehousing and BI》(2025)第 2.21 节明确写 JLL"pivoted from Snowflake",但 JLL 官方客户页只说"legacy client data warehouses/on-premises solutions",Snowflake 这一具体信源来自厂商渠道,引用时必须标注口径;换平台最大的成本不是数据搬运,而是人的迁移——JLL 把培训做成游戏化认证体系(Odyssey)是本案例最值得抄的作业,技术迁移计划里要给变更管理单独列预算和里程碑;"每个客户独立集群"的做法消除了 noisy neighbor,但代价是集群 sprawl,需要配套的成本治理,否则省下的性能会变成浪费的 DBU;行业交叉印证(以下非 JLL 官方披露,而是迁移咨询机构 playbook 的共识,带商业立场):LatentView 称这类迁移中代码转换通常占总工作量的 30–50%,真正的长尾是存储过程、JS UDF 与 TASK/STREAM 管道,成熟 Snowflake 环境里 20–40% 是无人消费的死负载、迁前应先清理,Unity Catalog 必须在迁移前完成设计、事后补等于重迁,回本周期约 12–18 个月;Dateonic 指出 Snowflake 的 CLUSTER BY 在 Databricks 没有一对一替代、需换成 Liquid Clustering,Photon 会被 Python UDF 绕过、迁移前须审计 UDF,TASK/STREAM 管道无法平移、必须重写为 Delta Live Tables。这些共识解释了为什么 JLL 需要 Odyssey 这种量级的培训投入——人的迁移和代码重写才是大头;一线工程师实录交叉印证(匿名 Medium 实录,2025-01,单一样本、客户匿名,不可独立验证):一次 30TB/2000+ 表/66 库的真实迁移里,900+ 个 Snowflake SQL 脚本的转换是主体工作量;迁移期必须建回 Snowflake 的 egress 管道、把 Databricks 金层同步回 Snowflake 以保住 Power BI 不中断("并行运行"原则的实战形态);具体坑包括 Array 类型先按 STRING 落临时表再重灌、JSON 解析在 from_json 与 Variant 预览版之间做性能选型、S3 与 Databricks 分属不同 VPC 要做 peering、迁移中被改动的脚本用数据字典跟踪漂移。
The trigger for leaving a mature warehouse like Snowflake is rarely "queries got 10% slower" — it's hitting the wall on governance, collaboration, and AI readiness at the same time; Databricks' own *Big Book of Data Warehousing and BI* (2025), §2.21, explicitly says JLL "pivoted from Snowflake," while JLL's customer page only says "legacy client data warehouses / on-premises solutions" — the Snowflake-specific attribution comes from a vendor channel and must be labeled as such; the biggest cost of switching platforms isn't moving data, it's moving people — JLL's gamified certification program (Odyssey) is the most copy-worthy part of this case, and migration plans should budget change management as its own line item with milestones; per-client clusters kill the noisy-neighbor problem but create cluster sprawl, which needs cost governance of its own, or the saved performance turns into wasted DBUs; cross-industry corroboration (not disclosed by JLL — consensus from migration consultancies' playbooks, which carry commercial interest): LatentView puts code conversion at 30–50% of total migration effort, with the long tail in stored procedures, JavaScript UDFs, and TASK/STREAM pipelines; 20–40% of a mature Snowflake estate is typically dead workloads nobody consumes and should be retired before migrating; Unity Catalog must be designed before migration starts — retrofitting it afterwards amounts to a re-migration; payback runs 12–18 months. Dateonic notes Snowflake's CLUSTER BY has no one-to-one equivalent and must be replaced with Liquid Clustering, Photon is bypassed by Python UDFs (audit UDFs before migrating), and TASK/STREAM pipelines cannot be lifted and shifted — they must be rewritten as Delta Live Tables. A first-hand practitioner account (anonymous Medium post, Jan 2025, single unverifiable sample, client unnamed) corroborates the mechanics from the trenches: in one 30TB / 2,000+ table / 66-database migration, converting 900+ Snowflake SQL scripts was the bulk of the work; egress pipelines syncing Databricks gold tables back into Snowflake kept Power BI running during the transition (the parallel-run principle in practice); concrete gotchas included staging Array columns as STRING before re-ingestion, evaluating JSON parsing across from_json vs. the Variant preview, VPC peering between S3 and Databricks, and a data dictionary tracking schema drift from scripts modified mid-migration. Together these explain why JLL needed an Odyssey-scale training investment: moving people and rewriting code is the bulk of the work.
来源
Databricks official customer story *Shaping the future of real estate for a better world* (JLL, with on-the-record quotes from Kristopher Curtis, Srihari Kumar, and Paul Chapman
Databricks《Big Book of Data Warehousing and BI》(2025)第 2.21 节"JLL — Upskilling Program for Data Warehouse Migration"(p.62)原文"Facing increasing data volumes, performance bottlenecks and legacy BI limitations, JLL pivoted from Snowflake to a strategic partnership with Databricks"(厂商渠道口径
Databricks *Big Book of Data Warehousing and BI* (2025), §2.21 "JLL — Upskilling Program for Data Warehouse Migration" (p.62), verbatim: "Facing increasing data volumes, performance bottlenecks and legacy BI limitations, JLL pivoted from Snowflake to a strategic partnership with Databricks" (vendor-channel attribution
Dateonic《Snowflake to Databricks Migration Partner: A Step-by-Step Guide》(迁移咨询机构 playbook,带商业立场
Dateonic *Snowflake to Databricks Migration Partner: A Step-by-Step Guide* (migration consultancy playbook, commercial interest
its named "Finanzwelt" case study fails verification — no such company exists, likely a pseudonym — and must not be cited
—
—
相关产品:Snowflake、Databricks 相关能力:Unity Catalog 统一治理 最后核验:2026-10-02
卡夫亨氏:用 Snowflake 安全数据共享,与零售商"同坐一张桌子"(2020–2021) 成功经验
消费品
供应链
安全数据共享
零售协同
- 场景
卡夫亨氏的数据基建以本地 Hadoop 为主,2020 年数字化转型要求整体上云。与零售商的协同长期靠报表、门户和 API:零售商要为每家制造商维护应用、重复拷贝数据;CPG 这边拼文件、验数,费力且有时滞,双方看的从来不是同一份实时数据。
Kraft Heinz's data infrastructure was largely on-premises Hadoop when its 2020 digital transformation demanded a move to the cloud. Collaboration with retailers long ran on reports, portals, and APIs: retailers maintained apps and duplicated data extracts per manufacturer, while the CPG side stitched files and validated them — laborious, laggy, and never a shared real-time view.
- 决策
按可扩展性、敏捷/速度/性能、云中立三条标准选中 Snowflake(跑在 Azure 上),并把"安全数据共享"列为零售行业的关键差异点。同时用 Marketplace 直接启用第三方数据(如 Johns Hopkins 大学的 COVID 数据),几分钟即开即用,省掉数据准备与测试。
Snowflake (on Azure) was selected against three criteria — scalability, agility/speed/performance, and cloud agnosticism — with Secure Data Sharing named a key differentiator for retail. Third-party data, such as Johns Hopkins University's COVID-19 dataset, was switched on directly from the Marketplace in minutes, skipping data preparation and testing.
- 结果
疫情期间 9 个月下掉本地数仓,5 千亿条记录进入 Snowflake,形成全球统一数据枢纽;安全库存模型 8–10 周上线验证。与零售商的 Joint Value Planning 改为:零售商直接更新共享表,卡夫亨氏零延迟看到同一份数据,"相当于同坐一张桌子",联合看供应链、库存、销售与分销。
On-premises warehouses were decommissioned in nine months despite the pandemic, with half a trillion records landing in Snowflake as a single global data hub; safety-stock models were built and validated in 8–10 weeks. Joint Value Planning with retailers changed shape: retailers update shared tables directly, Kraft Heinz sees the same data with zero delay — "equivalently sitting at the same table" — jointly reading supply chain, inventory, sales, and distribution.
- 机制根因
共享的是数据的"引用"而非拷贝——零复制、权限可管到表/列,双方读的是同一份物理数据,天然无时滞、无需对账。Marketplace 的数据产品同样免搬运,数据科学团队把精力从管道 plumbing 转到建模。网络效应:接入的零售商越多,协同价值越大。
What is shared is a reference to the data, not a copy — zero-copy sharing with table/column-level permissions means both sides read the same physical data, so there is inherently no lag and no reconciliation. Marketplace data products are equally movement-free, freeing data scientists from pipeline plumbing for modeling. Network effects apply: the more retailers join, the greater the collaboration value.
- 教训
协同速度取决于数据共享的摩擦力;"少搬运等于少出错",省掉的 ETL 不只是机器成本;第三方数据即开即用是云数仓被低估的红利。注意:本案例主要来源为 Snowflake 赞助/官方渠道,数字为受访高管口径,未见第三方独立复现。
Collaboration speed is set by the friction of data sharing; "less movement means fewer things break" — the ETL eliminated saves more than machine cost. Instant-on third-party data is an underrated dividend of cloud warehouses. Caveat: this case's sources are Snowflake-sponsored/official channels, figures are as stated by interviewed executives, and no independent third-party reproduction was found.
来源
HBR 赞助内容《How Data-Sharing Is Helping to Power a Global CPG Company》(Snowflake 赞助,厂商相关口径
HBR sponsor content, "How Data-Sharing Is Helping to Power a Global CPG Company" (sponsored by Snowflake — vendor-adjacent
Snowflake 官方博客《Kraft Heinz Fosters Collaboration with Retail Partners》(厂商口径
Snowflake official blog, "Kraft Heinz Fosters Collaboration with Retail Partners" (vendor claim
相关产品:Snowflake 相关能力:Secure Data Sharing + Marketplace 最后核验:2026-10-02
2024 年 UNC5537 凭证窃取事件:160 多家 Snowflake 租户因未开 MFA 被批量拖库 失败教训
安全事件
凭证失窃
身份治理
供应链风险
- 场景
2024 年 4 月起,财务动机攻击组织 UNC5537 用信息窃取木马盗来的凭证,批量登录 Snowflake 客户租户。Mandiant 确认超 160 家客户被波及、约 165 家收到通知,具名受害者包括 Ticketmaster、Santander、Advance Auto Parts。攻击者把窃得数据挂上暗网论坛售卖并勒索受害者。
Starting April 2024, the financially motivated threat actor UNC5537 used credentials stolen by infostealer malware to log into Snowflake customer tenants at scale. Mandiant confirmed more than 160 customers hit, with roughly 165 notified; named victims include Ticketmaster, Santander, and Advance Auto Parts. The attackers listed stolen data for sale on cybercrime forums and extorted victims.
- 决策
涉事租户事前的共同"决策"是:不强制 MFA、服务账号凭证多年不轮换、不配置网络 allowlist。事发后 Mandiant 介入调查通报,Snowflake 声明自身企业环境未被入侵,涉事客户承认事件并启动补救。
The victims' shared pre-incident "decision" was: no enforced MFA, service-account credentials unrotated for years, and no network allow lists. Afterward Mandiant investigated and disclosed, Snowflake stated its own enterprise environment was never breached, and affected customers acknowledged the incidents and began remediation.
- 结果
攻击者自称窃取 Ticketmaster 5.6 亿条记录、挂牌 50 万美元,Santander 3000 万条、索价 200 万美元——均为攻击者单方面宣称;两家公司承认发生了数据安全事件,但未证实条数。Mandiant 定性:手法"并不新颖也不复杂",规模完全来自受害者的普遍疏忽。
The attackers claimed 560M Ticketmaster records listed at $500,000 and 30M Santander records at $2M — both one-sided claims by the attackers; the two companies confirmed security incidents but never verified the figures. Mandiant's verdict: the technique was "not particularly novel or sophisticated" — the scale came entirely from victims' widespread negligence.
- 机制根因
三重缺失叠加——租户未开 MFA、凭证多年未轮换(有的来自 2020 年的木马日志)、未配置网络 allowlist;Snowflake 当时默认不强制 MFA,拿到用户名密码即获完全访问。SaaS 数仓是"数据引力中心",单个服务账号沦陷等于整仓沦陷;重灾区是第三方承包商的个人电脑(游戏、盗版软件带毒)。共享责任模型下,责任全部落在客户侧,平台默认配置不会替你兜底。
Three missing controls stacked — no MFA on tenants, credentials unrotated for years (some from 2020 stealer logs), no network allow lists; Snowflake did not require MFA by default, so a valid username and password meant full access. A SaaS warehouse is a data-gravity center: one compromised service account equals the whole warehouse. Third-party contractors' personal PCs (games, pirated software carrying stealers) were the worst-hit vector. Under the shared-responsibility model, all of it landed on the customer — platform defaults would not cover for them.
- 教训
服务账号同样要 MFA、轮换、allowlist,"系统账号不用管"是错觉;默认配置不等于安全配置,上云第一天就要按最高标准收紧;你的数据可能在供应商的 Snowflake 里——第三方的安全水位就是你的水位;攻击者永远先吃低 hanging fruit,基础卫生比高级威胁模型更重要。
Service accounts need MFA, rotation, and allow lists too — "system accounts don't need attention" is an illusion. Default configuration is not secure configuration; harden to the highest standard on day one. Your data may live in your vendor's Snowflake — their security posture is your blast radius. Attackers always eat the lowest-hanging fruit first; basic hygiene matters more than advanced threat models.
相关产品:Snowflake 相关能力:— 最后核验:2026-10-02
Google F1:AdWords 计费后端从分片 MySQL 迁到 Spanner,支撑 100TB+ 广告数据(2012–2013) 成功经验
Google 内部
广告计费
分片 MySQL 替换
层次 schema
- 场景
Google AdWords 是公司核心收入系统,旧后端是分片 MySQL:扩容难、rebalance 更难,复杂 join 逼着业务层精心设计分片键,resharding 经常要改应用。2012 年初 F1 在 Spanner 上投产,接管全部 AdWords 广告活动数据:100+ TB,数十万 QPS,SQL 查询每天扫描数十万亿行。
Google AdWords is the company's core revenue system; its old backend was sharded MySQL: hard to scale, harder to rebalance, complex joins forced the business layer to design shard keys with great care, and resharding often meant changing the application. In early 2012, F1 went into production on Spanner, taking over all AdWords advertising campaign data: 100+ TB, hundreds of thousands of requests per second, with SQL queries scanning tens of trillions of rows per day.
- 决策
与 Spanner 联合研发 F1——分布式 SQL 查询引擎 + Spanner 做存储层。核心设计:层次 schema(Customer→Campaign→AdGroup 主键前缀),子表行与父表行物理共置;异步 schema 变更(在线、零停机);乐观事务 + 变更历史自动记录发布。
Co-develop F1 with Spanner — a distributed SQL query engine with Spanner as the storage layer. Core design: hierarchical schema (Customer→Campaign→AdGroup key prefixes) with child-table rows physically colocated with parent rows; asynchronous online schema changes with zero downtime; optimistic transactions plus automatic change-history publishing.
- 结果
论文口径:五个九可用性(含计划外故障),Web 应用可观测延迟相对旧 MySQL 系统未增加。代价诚实披露:跨洲同步复制导致提交延迟 50–150ms,靠批量、并行、异步读与 ORM 显式化等应用层模式消化。
Per the paper: five-nines availability (including unplanned failures), with no increase in web-application observable latency versus the old MySQL system. The cost was honestly disclosed: cross-continent synchronous replication causes 50–150ms commit latency, absorbed through application-level patterns — batching, parallelism, asynchronous reads, and ORM explicitness.
- 机制根因
分片 MySQL 的根本矛盾是"分片键 = 业务模型的紧箍咒":join 必须沿分片键走,跨分片查询要么禁止要么在应用层拼。F1 的层次 schema 把"物理共置"声明在 schema 里:同一客户的所有广告数据落在同一个 Spanner directory,单客户事务走单分片、免 2PC;读同一客户的 Campaign+AdGroup 是一次 range 读 + 有序归并。这是 Spanner 交错表思想的源头——schema 设计即物理布局设计。
Sharded MySQL's fundamental contradiction is "shard key = straitjacket on the business model": joins must follow the shard key, and cross-shard queries are either forbidden or stitched together in the application. F1's hierarchical schema declares physical colocation in the schema: one customer's entire ad data lands in a single Spanner directory, single-customer transactions stay on one shard with no 2PC, and reading a customer's Campaigns+AdGroups is one range read plus an ordered merge. This is the origin of Spanner's interleaved-table idea — schema design is physical-layout design.
- 教训
分布式 SQL 的延迟税(50–150ms 提交延迟)是真实存在的,F1 用"应用层模式重构"而非"调参"来消化——选型时要评估团队是否愿意为新延迟模型重写数据访问层;schema 变更频率应作为选型输入:变更越频,越需要 F1 式的在线异步 DDL;"五个九 + 延迟不增"的结论来自 Google 第一方论文口径,外部复现时网络拓扑(datacenter 间约 100ms)差异会直接改变延迟数字。
Distributed SQL's latency tax (50–150ms commit latency) is real; F1 absorbed it through "application-pattern refactoring," not tuning — selection must assess whether the team is willing to rewrite its data-access layer for a new latency model; schema-change frequency should be a selection input: the more frequent the changes, the more you need F1-style online asynchronous DDL; the "five nines + no latency increase" conclusion comes from Google's first-party paper — network topology differences (roughly 100ms between datacenters) directly change the latency numbers when reproduced externally.
相关产品:Google Spanner、MySQL 相关能力:交错表(Interleaved Tables)—— schema 设计即物理布局 最后核验:2026-10-02
Mercari:核心 MySQL 选型评估 Spanner 后弃选——栽在 MySQL 兼容性与迁移路径(2022) 失败教训
选型落选
MySQL 兼容性
迁移成本
TiDB 对比
- 场景
Mercari 核心 MySQL 库随业务膨胀到"难以可靠运维"(故障、漏洞响应吃力);垂直拆分走到头(剩下几张紧耦合大表无法再拆),水平拆分又需要大改应用 + 微服务化改造。Core SRE 团队立项迁往可扩展 SQL,对 Cloud Spanner、Vitess、TiDB 按六个维度打分:配置(稳定性/failover)、查询(MySQL 兼容)、运维、安全、迁移、上下游联通。
Mercari's core MySQL database had grown too large to "operate reliably" (incidents and vulnerability response were painful); vertical sharding had hit its limit (a few tightly coupled large tables couldn't be split further), while horizontal sharding required major application rewrites plus a microservices overhaul. The Core SRE team launched a move to scalable SQL, scoring Cloud Spanner, Vitess, and TiDB on six dimensions: configuration (stability/failover), query (MySQL compatibility), operations, security, migration, and upstream/downstream connectivity.
- 决策
Spanner 在"查询"与"迁移"两项得零分:不支持 AUTO_INCREMENT(且单调递增主键不被推荐)、读 MySQL binlog 做迁移需要改造、表结构要按交错表重写。TiDB 在迁移(双向复制、可回滚)与查询(MySQL 5.7 兼容)占优,被定为"最有希望候选",进入业务验证。
Spanner scored zero on "query" and "migration": no AUTO_INCREMENT support (and monotonically increasing primary keys are discouraged), migration via MySQL binlog reads required rework, and table structures had to be rewritten for interleaved tables. TiDB won on migration (bidirectional replication, rollback) and query (MySQL 5.7 compatibility), and was named the "most promising candidate" for business validation.
- 结果
Spanner 落选。Mercari 明确写道 TiDB 更契合其运维组织(支持 failover 的集群形态、纵向集群划分习惯)。诚实注记:这是 2022 年 2 月的评估结论;Spanner 在 2025 年补上了 auto_increment、SELECT…FOR UPDATE、repeatable read(见本批 TimeTree 案例),若重估结论可能不同——选型结论有保质期。
Spanner was rejected. Mercari explicitly wrote that TiDB fit its operations organization better (cluster shapes supporting failover, the habit of vertical cluster partitioning). Honest note: this is a February 2022 evaluation; in 2025 Spanner added auto_increment, SELECT…FOR UPDATE, and repeatable read (see this batch's TimeTree case) — a re-evaluation could reach a different conclusion. Selection conclusions have an expiry date.
- 机制根因
Spanner 的落选不是"性能不够",而是"迁移摩擦"。MySQL 生态的隐性契约(自增主键、binlog CDC、下游 Kafka)构成迁移的真实成本。单调主键在 Spanner 的 range 分片下退化为单分片写——这不是调优问题,是数据模型假设冲突:MySQL 时代"自增主键 = 最佳实践",Spanner 时代"自增主键 = 反模式",迁移意味着重写建模习惯。Vitess/TiDB 胜在"把 MySQL 方言和生态原样搬过去",Spanner 要求"按分布式重学建模"。
Spanner lost not on performance but on migration friction. MySQL's ecosystem carries implicit contracts (auto-increment keys, binlog CDC, downstream Kafka) that make up the real cost of migration. Monotonic primary keys degenerate into single-shard writes under Spanner's range sharding — not a tuning problem but a data-model assumption conflict: in the MySQL era "auto-increment primary key = best practice," in the Spanner era "auto-increment primary key = anti-pattern," and migration means relearning modeling habits. Vitess/TiDB won by "carrying the MySQL dialect and ecosystem over intact"; Spanner demands "relearning modeling for distribution."
- 教训
评估分布式数据库时,"查询兼容性"与"迁移可逆性"应与性能同权打分;binlog/CDC 生态是 MySQL 用户最沉的资产,选型前先审计上下游对 binlog 的依赖;任何 2022 年前后的 Spanner 落选结论都要用"2025 年 MySQL 兼容性补齐"重新校准——给选型报告写上有效期。
When evaluating distributed databases, "query compatibility" and "migration reversibility" deserve the same scoring weight as performance; the binlog/CDC ecosystem is a MySQL user's heaviest asset — audit upstream/downstream binlog dependencies before selecting; any Spanner rejection from around 2022 must be recalibrated against "2025 MySQL-compatibility catch-up" — put an expiry date on selection reports.
相关产品:Google Spanner、TiDB、MySQL 相关能力:贵是"时间的物理价格" + 单调主键写热点 最后核验:2026-10-02
Niantic:Pokémon GO 从 Datastore 迁到 Spanner,活动峰值 40 万冲到近百万 TPS(2021) 成功经验
游戏全球同服
峰值扩容
Datastore 迁移
- 场景
Pokémon GO 全球同服、所有玩家共享同一游戏世界。社区日、GO Fest 等活动期间,事务量"几分钟内从每秒 40 万冲到近百万"。后端主力是 GKE + Spanner,常备约 5000 个 Spanner 节点。每天产生 5–10TB 数据进入 BigQuery/Bigtable 分析管道。
Pokémon GO runs a single global realm where all players share one game world. During events like Community Day and GO Fest, transaction volume surges "from 400K per second to close to a million within a few minutes." The backend is primarily GKE + Spanner with roughly 5,000 Spanner nodes provisioned, and 5–10TB of data per day flows into BigQuery/Bigtable analytics pipelines.
- 决策
早期用 Google Datastore(免运维、快速起步);游戏成熟后需要关系模型 + 全局 ACID 事务 + 事务一致的索引(主/二级键复杂 schema),迁到 Spanner。抓宠流程:手机→负载均衡→NGINX→GKE 前端→写 Spanner;空间查询后端按地理分片缓存地图数据;用户行为以 protobuf 写 Bigtable 日志、经 Pub/Sub 进分析管道。
Early on the team used Google Datastore (zero ops, fast start); as the game matured it needed a relational model + global ACID transactions + transactionally consistent indexes (complex primary/secondary-key schema), so it migrated to Spanner. The catch flow: phone → load balancer → NGINX → GKE frontend → Spanner writes; the spatial-query backend caches map data sharded by geography; player behavior is logged as protobufs to Bigtable and streamed via Pub/Sub into analytics.
- 结果
单 realm 全球同服成为可能——强一致保证"同一地点所有玩家看到同一只宝可梦"。活动扩容时 Niantic SRE 只需保证配额,托管服务承担运维。未披露成本数字。
A single-realm global game became possible — strong consistency guarantees that "all players at one location see the same Pokémon." During event scale-ups, Niantic SRE only had to secure quota while the managed service absorbed the operations. No cost figures were disclosed.
- 机制根因
全球同服的本质要求是"写全局有序":两个玩家同时点同一只宝可梦,谁先抓到必须有唯一答案。Datastore 的实体组事务满足不了跨实体关系事务;Spanner 的外部一致性让"先发生的抓取在全局先成立",事务一致的二级索引让"按地理位置查宝可梦/道馆"不必扫全表。5000 节点常备的另一面是成本——Spanner 按节点计费,游戏波峰波谷明显,节点弹性是账单关键(博客未谈成本,此为机制推断)。
A global single-realm game's fundamental requirement is "globally ordered writes": when two players tap the same Pokémon simultaneously, who caught it first must have exactly one answer. Datastore's entity-group transactions couldn't cover cross-entity relational transactions; Spanner's external consistency makes "the earlier catch in real time wins globally," and transactionally consistent secondary indexes let "find Pokémon/gyms by location" avoid full-table scans. The flip side of 5,000 standby nodes is cost — Spanner bills per node, and a game's peaks and valleys are extreme, so node elasticity is the key to the bill (the blog never discusses cost; this is a mechanism-level inference).
- 教训
从 Datastore 这类"起步快"的存储迁出时,真正的迁移成本不在数据量而在语义升级(非关系→关系、最终一致→强一致),越早识别语义天花板越好;全球同服类需求选型时,"一致性语义"应排在"单机性能"之前验证;峰值型业务要把节点弹性的运维手册写在选型报告里,而不是上线后才发现。
When migrating off a "fast start" store like Datastore, the real migration cost is not data volume but the semantic upgrade (non-relational → relational, eventual → strong consistency) — identify the semantic ceiling as early as possible; for global single-realm requirements, validate "consistency semantics" before "single-node performance" during selection; peak-shaped businesses should write the node-elasticity runbook into the selection report, not discover it after launch.
相关产品:Google Spanner 相关能力:TrueTime + 外部一致性 —— 全球多活强一致的唯一解 最后核验:2026-10-02
Google 自身事故:Service Control 一次坏写入经 Spanner 秒级同步全球,重启风暴反压垮 Spanner(2025) 失败教训
生产事故
全球复制反噬
重启风暴
控制面
- 场景
2025 年 6 月 12 日,Google Cloud/Workspace 全球大范围 503。Service Control(API 管理控制面二进制)的配额/策略数据存放在 regional Spanner 表中,并"几乎瞬时"复制到全球。5 月 29 日上线的新配额检查代码缺少错误处理、未受 feature flag 保护。
On June 12, 2025, Google Cloud/Workspace suffered widespread global 503s. Service Control (the API-management control-plane binary) stores quota/policy data in regional Spanner tables, replicated "almost instantly" worldwide. A new quota-check code path released May 29 lacked error handling and was not feature-flag protected.
- 决策
Google 的架构决策是把 Service Control 的配额/策略元数据放在 regional Spanner 表中,并"几乎瞬时"复制到全球——用全球强一致复制来服务全球配额管理。2025 年 5 月 29 日,又上线了未经 feature flag 保护、缺少错误处理的新配额检查代码(带一个可关闭问题策略路径的"红按钮",但事故中 40 分钟才完成推送)。
Google's architectural decision was to keep Service Control's quota/policy metadata in regional Spanner tables and replicate it "almost instantly" worldwide — using globally strong-consistent replication to serve global quota management. On May 29, 2025, a new quota-check code path went live without feature-flag protection and without error handling (it did ship with a red-button to disable the faulty policy path, but the push took ~40 minutes during the incident).
- 结果
全球 API 约 3 小时大面积 503(官方口径 10:49–13:49 PDT),Cloud Service Health 自身也被波及,首份公开报告延迟约 1 小时。Google 事后冻结 Service Control 变更,整改清单包括:审计所有消费全球复制数据的系统、复制改为增量渐进并留出验证窗口、关键二进制强制 feature flag、补随机指数退避、Service Control 模块化 fail-open。
Roughly 3 hours of widespread global API 503s (official window 10:49–13:49 PDT); Cloud Service Health itself was affected, delaying the first public report by about an hour. Google froze Service Control changes afterward, with a remediation list: audit all systems consuming globally replicated data, switch replication to incremental rollout with validation windows, mandate feature flags for critical binaries, add randomized exponential backoff, and modularize Service Control to fail open.
- 机制根因
Spanner 的全球同步复制在这里是"放大器"不是"保险丝":一次区域写入在秒内变成全球毒药,而读路径(Service Control 二进制)对坏数据零容忍(空指针崩溃而非降级)。更深层的是"控制面依赖数据面"的循环:Service Control 重启风暴打垮了 Spanner 表,而 Spanner 表正是 Service Control 恢复所依赖的——恢复路径与故障路径共享同一基础设施。Google 自己的结论点名了机制:"无论业务多需要全球近瞬时一致,数据复制都必须增量传播并留出验证时间。"
Spanner's global synchronous replication acted as an amplifier, not a fuse: one regional write became global poison within seconds, while the read path (the Service Control binary) had zero tolerance for bad data (null-pointer crash instead of graceful degradation). Deeper still was the control-plane-depends-on-data-plane loop: the Service Control restart storm knocked over the Spanner tables that Service Control itself needed to recover — the recovery path shared the same infrastructure as the failure path. Google's own conclusion named the mechanism: "no matter how much the business needs global near-instant consistency, data replication must propagate incrementally and leave time for validation."
- 教训
把 Spanner 的全球复制用在控制面/策略数据上时,写入侧必须有 schema 校验与灰度(feature flag),不能信任"写入即正确";任何大规模重启路径都要有随机指数退避,否则恢复流量本身就是第二次故障;监控与状态页基础设施要与被监控系统做故障域隔离——否则出事时你连"出事了"都发布不出去。
When using Spanner's global replication for control-plane/policy data, the write side must have schema validation and gating (feature flags) — never trust "written means correct"; any large-scale restart path needs randomized exponential backoff, otherwise recovery traffic is a second incident; monitoring and status-page infrastructure must be fault-domain-isolated from the systems being monitored — otherwise you can't even publish "we are down" during an outage.
相关产品:Google Spanner 相关能力:TrueTime + 外部一致性 —— 全球多活强一致的唯一解 最后核验:2026-10-02
ShareChat:1.6 亿月活社交平台从 NoSQL 迁到 Spanner,流量 5 倍暴涨零改代码(2021) 成功经验
社交出海
零停机迁移
弹性扩容
成本优化
- 场景
印度社交平台 ShareChat,1.6 亿月活、15 种语言,日均百万级帖子;同期上线短视频应用 Moj(8000 万月活)。原跑在某云 NoSQL 上(客户未具名),为应对不可预测流量长期超配计算与存储。正式迁移前做了 4 个多月的 PoC,用网关复制生产流量做影子压测,验证超百万 QPS。
ShareChat, an Indian social platform with 160 million monthly active users across 15 languages and millions of daily posts, concurrently launched the short-video app Moj (80M MAU). It ran on an unnamed cloud NoSQL database, chronically overprovisioning compute and storage to cope with unpredictable traffic. Before the real migration, a 4+ month PoC shadow-tested production traffic replayed through a gateway, validating over a million queries per second.
- 决策
整体迁往 Google Cloud,实时 serving 层选 Spanner。理由:全球一致性 + 事务一致的二级索引;"不像 legacy NoSQL,扩容不用重新思考表结构"。迁移用 wrapper 封装隔离,应用代码零改动;6000 万用户 5 小时迁完,零数据丢失、零停机。
Migrate wholesale to Google Cloud with Spanner as the real-time serving layer. Rationale: global consistency plus transactionally consistent secondary indexes; "unlike our legacy NoSQL database, we could scale without having to rethink existing tables or schema definitions." The migration was wrapped in an isolation layer so application code needed zero changes; 60 million users moved in five hours with no data loss and no downtime.
- 结果
120 张表 + 17 个索引迁入 Spanner;ShareChat 联创兼 CTO Bhanu Singh 自称成本降 30%(客户口径,引自 Google Cloud 博客客座文章,无第三方复现)。一次流量几天内暴涨 500%,水平扩展"零行代码改动";平均 8 万 RPS,推送通知可致几秒内冲到 13 万 RPS。诚实备注:70TB/220 表、50 亿行大表等数字同样来自该客座文章。
120 tables with 17 indexes moved into Spanner; co-founder and CTO Bhanu Singh claims costs fell 30% (customer claim, from a Google Cloud blog guest post, no independent reproduction). When traffic spiked 500% within days, horizontal scaling required "zero lines of code change"; baseline 80,000 requests per second, with push notifications driving bursts to 130,000 RPS within seconds. Honest note: the 70TB/220-table and 5-billion-row figures come from the same guest post.
- 机制根因
NoSQL 时代的扩容税在 ShareChat 身上表现为"超配":流量不可预测 → 为峰值预留 → 大部分时间空转。Spanner 的"计算存储分离 + split 自动分裂"把扩容从容量规划问题变成 API 调用;事务一致的二级索引让"按人查帖、按话题查帖"不用另建索引表(Cassandra 式架构要维护单独索引表并处理双写不一致)。500% 暴涨零改代码的本质:分片逻辑收在数据库内部,应用层没有分片键概念。
The NoSQL-era scaling tax at ShareChat took the form of overprovisioning: unpredictable traffic → reserve for peaks → idle most of the time. Spanner's compute-storage separation plus automatic split-splitting turned scaling from a capacity-planning exercise into an API call; transactionally consistent secondary indexes eliminated the separate index tables (and their dual-write inconsistency) that Cassandra-style architectures require for "posts by user / posts by topic" queries. The essence of the zero-code-change 5x surge: sharding logic lives inside the database, so the application has no shard-key concept at all.
- 教训
迁移前用生产流量做影子压测(PoC 4 个月、百万 QPS)是零停机迁移的前提;wrapper 隔离层让"换数据库"与"改业务代码"解耦;为不可预测流量设计的系统,选型时应把"超配浪费"而非"峰值单价"作为成本比较基准——这正是 ShareChat 算出降本 30% 的口径背景。
Shadow-testing with production traffic before migration (4-month PoC at 1M QPS) is the precondition for zero-downtime migration; a wrapper isolation layer decouples "changing the database" from "changing business code"; for systems designed around unpredictable traffic, the cost benchmark should be "overprovisioning waste," not "peak unit price" — that is the accounting context behind ShareChat's claimed 30% saving.
相关产品:Google Spanner 相关能力:TrueTime + 外部一致性 —— 全球多活强一致的唯一解 最后核验:2026-10-02
TimeTree:5500 万用户时撞上 Aurora MySQL 上限,迁 Spanner 降本增效(2025) 成功经验
Aurora 扩容瓶颈
MySQL 生态迁移
日历社交
- 场景
日本日历共享应用 TimeTree,用户达 5500 万时,"在数据量与活跃连接数两方面撞上 Aurora MySQL 的扩展上限"(SRE 经理 Eiki Kanai)。同时应用团队把过多时间花在管理数据库上,挤占功能开发。
TimeTree, a Japanese shared-calendar app, "hit Aurora MySQL's scalability limits for both data volume and active connections" at 55 million users (SRE Manager Eiki Kanai). Meanwhile the application team was spending too much time managing the database, crowding out feature development.
- 决策
迁往全托管 Spanner,用 Spanner migration tool(SMT)做在线迁移。2025 年发布的能力降低了迁移摩擦:repeatable read 隔离级别(preview,常见负载延迟最高降 5 倍)、auto_increment 主键、SELECT…FOR UPDATE、近 80 个 MySQL 函数。
Move to fully managed Spanner with an online migration via the Spanner migration tool (SMT). Capabilities released in 2025 lowered migration friction: repeatable read isolation (preview, up to 5x latency reduction on common workloads), auto_increment primary keys, SELECT…FOR UPDATE, and nearly 80 new MySQL functions.
- 结果
Eiki Kanai 称"显著降本、支撑未来增长","在 Google Cloud 支持与迁移工具帮助下,以最小停机完成迁移"(客户口径)。Google 引用的 Forrester TEI 称复合组织三年 ROI 132%、总收益 $7.74M(第三方咨询机构口径,样本为"代表性复合组织"而非 TimeTree 本身)。
Kanai reports "significant cost reduction, supporting future growth," with "minimal downtime with the help of Google Cloud support and migration tools" (customer claim). A Forrester TEI study cited by Google reports 132% three-year ROI and $7.74M total benefits for a composite organization (third-party analyst claim, modeled on a "representative composite organization," not TimeTree itself).
- 机制根因
Aurora 的写上限是单 writer 实例上限 + 连接数上限的双重天花板;5500 万用户的日历共享是典型的"高扇出读 + 热点写"(同一家庭/团队日历被多人同时改)。Spanner 的多写 + 水平扩展解开写上限;repeatable read 让 MySQL 习惯的默认隔离语义得以保留,避免全量改写事务代码——这正是 2022 年 Mercari 评估时缺失的能力(见本批反例),选型结论随产品演进而过期。
Aurora's write ceiling is a double cap — single-writer instance limits plus connection limits; a 55M-user shared calendar is classic "high-fanout reads + hotspot writes" (one family/team calendar edited by many people at once). Spanner's multi-writer plus horizontal scaling unlocks the write ceiling; repeatable read preserves the default isolation semantics MySQL users rely on, avoiding a full rewrite of transaction code — exactly the capability missing in Mercari's 2022 evaluation (see this batch's failure lesson): selection conclusions expire as products evolve.
- 教训
Aurora 用户的扩容天花板是"可预期的",应在连接数/存储达上限 70% 时就开始评估 NewSQL,而不是撞墙后;迁移工具链(SMT 反向复制、回滚)与语义兼容(隔离级别、auto_increment)同等重要;引用 Forrester 类报告时必须说明是"复合组织"建模而非被访客户实测。
An Aurora user's scaling ceiling is predictable — start evaluating NewSQL when connections/storage reach 70% of limits, not after hitting the wall; the migration toolchain (SMT reverse replication, rollback) matters as much as semantic compatibility (isolation levels, auto_increment); when citing Forrester-style reports, always note they model a "composite organization," not the interviewed customer's measured results.
相关产品:Google Spanner、Amazon Aurora 相关能力:— 最后核验:2026-10-02
Uber:履约平台从 Cassandra + Saga 重构到 Spanner,撑起每日数十亿事务 成功经验
跨城履约
NoSQL 迁 NewSQL
强一致重构
补偿事务消除
- 场景
Uber 履约平台是"去任何地方、得任何东西"战略的地基,承载出行与外卖全品类的订单生命周期:超百万并发用户、每年数十亿行程、每天数十亿数据库事务,数百个微服务以它为订单与司机状态的可信源。旧架构(2014 年)是 Cassandra + Redis + Ringpop + 应用层 Saga:可用性优先、一致性靠尽力而为,跨实体写由应用层 Saga 拆成 propose/commit/cancel 三阶段协调。
Uber's Fulfillment Platform is the foundation of its "go anywhere, get anything" strategy, carrying the order lifecycle across rides and delivery: more than a million concurrent users, billions of trips per year, billions of database transactions per day, with hundreds of microservices treating it as the source of truth for order and driver state. The old architecture (2014) was Cassandra + Redis + Ringpop + application-layer Saga: availability first, consistency best-effort, cross-entity writes coordinated by application Sagas split into propose/commit/cancel phases.
- 决策
Uber 下注彻底重写(博客原话"two years ago, we made a bold bet"),存储层三选一:修补 NoSQL、自建 MySQL 分片、换 NewSQL。按可用性 SLA、运维开销、事务能力、schema 管理、分片管理、自动扩展、水平扩展等数十项维度打分并做基准测试后,选定 Google Cloud Spanner 为主存储引擎,采用北美多 region 配置 nam3。核心逻辑:把事务协调从应用层下沉到数据库层。
Uber bet on a ground-up rewrite ("two years ago, we made a bold bet"), with three storage options: patch the NoSQL stack, build sharded MySQL in-house, or switch to NewSQL. After scoring dozens of dimensions (availability SLA, operational overhead, transactional power, schema management, shard management, auto-scaling, horizontal scaling) and running benchmarks, they chose Google Cloud Spanner as the primary storage engine with the North America multi-region configuration nam3. The core logic: push transaction coordination from the application layer down into the database layer.
- 结果
约两年的重写与迁移后,"所有 Uber 产品与城市"切到新栈,100+ 工程师、30+ 团队参与(Uber 官方口径)。未披露延迟与成本对比数字。诚实细节:Uber 履约服务跑在自有机房,每笔事务都要跨网调用 Google Cloud 上的 Spanner,强一致的代价是每笔写的跨云 RTT,博客未量化。
After roughly two years of rewriting and migration, "every Uber product and city" moved to the new stack, with 100+ engineers across 30+ teams (Uber's own account). No latency or cost comparison numbers were disclosed. An honest detail: Uber's fulfillment services run in its own datacenters, so every transaction crosses the network to Spanner on Google Cloud — the price of strong consistency is a cross-cloud RTT per write, which the blog never quantified.
- 机制根因
旧架构的病根是"一致性外包给应用层"。Cassandra 的 last-write-wins 语义下,部署与 region 切换时的脑裂会产生并发写互相覆盖;Saga 让逻辑事务长期处于内部不一致态,排查一次跨实体问题要跨多个服务。Spanner 的外部一致性 + 跨表跨分片事务把这套复杂性收回数据库内核:DML 服务端事务缓冲、锁冲突检测与死锁避免、stale read 支持,让应用层从"编排一致性"退化为"声明事务边界"。本质是把分布式协调从业务代码移到了经过严格验证的数据库内核。
The old architecture's disease was "consistency outsourced to the application layer." Under Cassandra's last-write-wins semantics, split-brain during deployments and region failovers caused concurrent writes to clobber each other; Sagas left logical transactions in internally inconsistent states for long stretches, and debugging one cross-entity inconsistency spanned multiple teams. Spanner's external consistency plus cross-table, cross-shard transactions pulled that complexity back into the database kernel: server-side DML transaction buffering, lock-conflict detection with deadlock avoidance, and stale reads demoted the application layer from "orchestrating consistency" to "declaring transaction boundaries." In essence, distributed coordination moved from business code into a rigorously verified database kernel.
- 教训
当 Saga 补偿代码超过业务代码、查一次数据不一致要跨三个团队时,就是 NoSQL 语义不够用的信号;选型要把"为弥补数据库语义缺失而写的应用层胶水代码"计入 TCO,Spanner 的账单应与这部分人力成本一起算;跨云调用 Spanner 时 region 拓扑与自有机房的距离直接决定写延迟,网络架构要先于 schema 设计。
When Saga compensation code outweighs business code, and one inconsistency investigation spans three teams, the NoSQL semantics are no longer enough; TCO calculations should include the application glue code written to compensate for missing database semantics — Spanner's bill must be weighed against that headcount cost; when calling Spanner across clouds, the distance between the region topology and your own datacenters directly determines write latency, so network architecture must come before schema design.
相关产品:Google Spanner、Apache Cassandra / ScyllaDB、Redis / Valkey 相关能力:TrueTime + 外部一致性 —— 全球多活强一致的唯一解 最后核验:2026-10-02
bet365:SQL Server 纵向扩展到 160 核见顶,分片被否决后选 Riak 失败教训
纵向扩展天花板
分片
NoSQL 选型
在线博彩
- 决策
否决分片——"分片自带问题,我们的应用特性与需求不适合";评估约 10 款产品(主流 NoSQL 厂商 + 实质卖分布式缓存的厂商)后选定 Basho Riak,把大块核心功能从 SQL Server OLTP 迁往 Riak OLTP。
Sharding was rejected — "sharding brings its own problems" given the applications' nature and requirements; after evaluating about 10 products (the main NoSQL vendors plus vendors effectively selling distributed caches), bet365 picked Basho's Riak and started moving large pieces of core functionality from SQL Server-based OLTP to Riak OLTP.
- 结果
Riak 在故意制造故障的测试里最稳、恢复最快——"不是最快,但足够快",且可接近线性扩展(取决于复制因子与一致性配置);两大核心功能已在 Riak 上线,其余持续迁移中。
Riak proved the most stable and fastest to recover under deliberately injected failures — "not the fastest but fast enough" — and scaled near-linearly (depending on replication factor and consistency configuration); two large pieces of core functionality went live on Riak, with more migrating over.
- 机制根因
博彩是写密集 + 强一致 + 低延迟的 OLTP:纵向扩展的天花板不是 CPU 核数,而是单机锁/闩锁的扩展性与故障域——160 核 SMP 上跨核同步开销吃掉扩展收益,而单机始终是单点;分片被否决是因为投注业务的跨分片事务与实时结算让应用层分片代价过高。于是只剩换架构一条路。
Betting is write-intensive, strongly consistent, low-latency OLTP: the vertical-scaling ceiling is not CPU core count but single-box lock/latch scalability and failure domains — on a 160-core SMP box, cross-core synchronization overhead eats the scaling gains, and one box remains one failure domain; sharding was off the table because cross-shard transactions and real-time settlement in betting make application-level sharding prohibitively expensive. That left only a change of architecture.
- 教训
纵向扩展有两层天花板:核数是看得见的那层,锁扩展性与故障域是看不见的那层,后者先到。bet365 的选型标准值得抄:在故障注入下最稳、恢复最快的赢,而不是 benchmark 最快的赢——"足够快 + 接近线性扩展(取决于拓扑与一致性等级)+ 故障时反应最快"才是生产选型公式。本案例与 PlentyOfFish、Stack Overflow 构成对称:纵向扩展的适用条件是负载形状,不是信仰。
Vertical scaling has two ceilings: core count is the visible one, lock scalability and failure domains the invisible one — the invisible one arrives first. bet365's selection criterion is worth copying: the winner under fault injection, not the benchmark champion — "fast enough + near-linear scaling (depending on topology and consistency level) + fastest recovery under failure" is the production selection formula. This case mirrors PlentyOfFish and Stack Overflow symmetrically: vertical scaling's fit is a function of workload shape, not faith.
相关产品:Microsoft SQL Server 相关能力:— 最后核验:2026-10-02
bwin.party:SQL Server 2014 内存 OLTP,同硬件吞吐 16 倍、18 台并 1 台 成功经验
纵向扩展
内存计算
在线博彩
- 决策
不换库、不分片、不加机器,用 SQL Server 2014 的内存 OLTP 等新能力扛在线博彩的请求洪峰,硬件保持不变。
No new database, no sharding, no new hardware — absorb the betting traffic spike with SQL Server 2014's in-memory OLTP and related engine features on the same boxes.
- 结果
吞吐从旧版 SQL Server 的 1.6 万 req/sec 提到 25 万 req/sec(16 倍,同硬件);跑 SQL Server 的服务器从 18 台合并为 1 台,数据基础设施大幅简化。
Throughput rose from 16k req/sec on the previous SQL Server version to 250k req/sec (16x, same hardware); the SQL Server fleet shrank from 18 servers to 1, simplifying the data infrastructure significantly.
- 机制根因
内存 OLTP 把热点行操作从"磁盘页 + 闩锁 + 锁管理器"的重路径搬进内存、用免锁数据结构,单核做的有效功跃升;同样的 CPU 核数直接换成吞吐,少机器又等于少故障域、少复制延迟。这类"换内核不换架构"的红利只吃得到一次,但一次就很肥。
In-memory OLTP moves hot row operations off the heavy "disk pages + latches + lock manager" path into memory with lock-free data structures, so each core does far more useful work; the same CPU count converts directly into throughput, and fewer boxes mean fewer failure domains and less replication lag. This "new engine, same architecture" dividend can only be collected once — but once is plenty.
- 教训
大版本内核能力(如内存 OLTP、列存储)有时一次吃掉多年"要不要换库"的争论,升级前先算新版本能给你什么。但本例数字来自微软口径(Azure 客户咨询团队总经理 Mark Souza),且是在线博彩特定的读写模式下的结果——照搬前先确认自己的热点模式是否同类。
A major version's engine features (in-memory OLTP, columnstore) can settle years of "should we switch databases" debate in one move — price in what the new version gives you before migrating. Caveat: these numbers come from Microsoft (Mark Souza, GM of the Azure Customer Advisory Team) for an online-betting read/write pattern — confirm your own hot path looks similar before extrapolating.
来源
eWeek(转述微软 Azure 客户咨询团队总经理 Mark Souza
eWeek (quoting Microsoft's Mark Souza, GM of the Azure Customer Advisory Team
相关产品:Microsoft SQL Server 相关能力:— 最后核验:2026-10-02
DocuSign:SQL Server 分析侧"力不从心",迁 Snowflake 失败教训
数据仓库
上云迁移
BI
SaaS
- 决策
分析负载从 SQL Server 迁到 Snowflake(存算分离、弹性扩展),再用 Fivetran 把数据源管道自动化,不再手写维护 ETL。
Move the analytics workload from SQL Server to Snowflake (separated storage and compute, elastic scaling), then automate source pipelines with Fivetran instead of hand-building and maintaining ETL.
- 结果
可用数据源从 SQL Server 时代的 6 个变成 3 倍(新增十几个);100+ Qlik 仪表盘被全公司日常使用;BI 高级经理 Marcus Laanen:"Snowflake has been a game-changer for us……任何 SQL 查询都比以前快得多。"
Usable data sources tripled from 6 in the SQL Server era (a dozen more added); 100+ Qlik dashboards in daily use company-wide; Marcus Laanen, Senior Manager of Business Intelligence: "Snowflake has been a game-changer for us... You run any SQL query as you would with any other system, but with Snowflake it performs so much faster."
- 机制根因
"一个库既跑业务又跑分析"是典型反模式:分析查询抢缓冲池、锁阻塞 OLTP,还得按峰值预配硬件;Snowflake 的存算分离让分析并发与数据量解耦——这不是"SQL Server 慢",是把 OLAP 负载从行存 OLTP 引擎上卸载。
"One database for both operations and analytics" is the classic anti-pattern: analytical queries fight for buffer pool, locks block OLTP, and hardware must be provisioned for peak; Snowflake's storage-compute separation decouples analytical concurrency from data volume — this is not "SQL Server is slow," it is offloading OLAP from a row-based OLTP engine.
- 教训
当"SQL Server 不够用"的抱怨来自 BI/分析侧时,先问是不是把 OLAP 负载压在了 OLTP 引擎上——Xero(Redshift)是同一剧本的另一版本。注意本案例出自 Fivetran 厂商稿,"falling short"的具体指标未披露,证据强度弱于 eHarmony、VSTS 的第一人称事故记录,引用时应标注口径。
When "SQL Server isn't enough" complaints come from the BI/analytics side, first ask whether OLAP load is sitting on an OLTP engine — Xero (Redshift) is another telling of the same story. Note this case comes from a Fivetran vendor publication, and the specific metrics behind "falling short" were not disclosed — weaker evidence than the first-person incident records of eHarmony and VSTS, so cite it with the provenance labeled.
相关产品:Microsoft SQL Server、Snowflake 相关能力:— 最后核验:2026-10-02
eHarmony:SQL Server 2000 扛不住 1200 TPS,14 个月迁到 Oracle RAC 失败教训
纵向扩展天花板
迁移
集群
高并发写入
- 决策
不等 SQL Server 2005(决策时未发布),用 14 个月把 OLTP 迁到 Oracle Database 10g + RAC + Clusterware(Sun Fire X4600,Windows);数据仓库也一并迁到 Oracle,"换了生产源库,跟一个供应商走最省事"。
Rather than wait for SQL Server 2005 (unreleased when the call was made), eHarmony spent 14 months moving OLTP to Oracle Database 10g + RAC + Clusterware on Sun Fire X4600 servers (Windows); the data warehouse moved to Oracle too — once you change the production source, single-vendor is simplest.
- 结果
切换后性能提升 30%,达到亚秒级响应;代价是 14 个月迁移工程 + Oracle 许可与硬件账单。
Performance improved 30% after the switch, reaching sub-second response; the price was a 14-month migration plus Oracle licensing and hardware bills.
- 机制根因
瓶颈不是 4TB 数据量("这点量不算什么"),而是 1200 TPS 的事务并发:单机 SQL Server 2000 的锁/闩锁与调度器在高峰期把 8 核吃满,且当时的 SQL Server 没有真正的多活集群能力(AlwaysOn 要到 2012 年才来),"加机器"无处可加;Oracle RAC 恰好把集群做进数据库内核,"可以往集群里加机器"是打动 eHarmony 的关键。
The bottleneck was not the 4TB data volume but 1,200 TPS of transactional concurrency: lock/latch pressure and scheduler contention on single-box SQL Server 2000 saturated 8 cores at peak, and SQL Server at the time had no true active-active clustering (AlwaysOn would not arrive until 2012) — there was nowhere to add machines. Oracle RAC put clustering inside the database kernel, and "we can put it against more machines in a cluster" was what won eHarmony over.
- 教训
区分"数据量天花板"和"事务并发天花板":4TB 对 SQL Server 毫无压力,1200 TPS 才是杀手。eHarmony 技术 VP Mark Douglas 的原话值得钉在墙上——"SQL Server 多年表现很好,只是到了需要集群这类能力的那一刻,微软给不出"。选型时要把"未来三年可能需要的能力"(集群、分区)写进决策表,而不是只看当下够用。
Distinguish the data-volume ceiling from the transaction-concurrency ceiling: 4TB was nothing to SQL Server, 1,200 TPS was the killer. Mark Douglas, eHarmony's VP of Technology, put it memorably: "SQL Server performed very well for the company for many years. It was just getting to a point where we needed certain features, like clustering, that Microsoft couldn't offer." Write the capabilities you might need in three years (clustering, partitioning) into the decision sheet — don't evaluate only what suffices today.
相关产品:Microsoft SQL Server、Oracle Database(甲骨文) 相关能力:AlwaysOn 可用性组 —— "大铁块"的高可用 最后核验:2026-10-02
PlentyOfFish:1 人运维,SQL Server 纵向扩展扛下 60 亿月 PV 成功经验
纵向扩展
内存优先
运维简单
个人开发者
- 决策
不分片、不换 NoSQL:核心库 SQL Server 2008,直接升级到 512GB 内存、32 CPU 的"大铁块"(Windows 2008);读写分离、反范式化、把工作集塞进内存;图片走 Akamai CDN。
No sharding, no NoSQL: the core SQL Server 2008 database was scaled up to a 512GB-RAM, 32-CPU "big iron box" (Windows 2008); reads separated from writes, data denormalized, working set kept in memory; images served through the Akamai CDN.
- 结果
单人运维撑住数十亿级月访问;Google 广告年收入 600 万–1000 万美元。
One person operated the site at billions of monthly pageviews; Google ads brought in $6M-$10M a year.
- 机制根因
工作集能装进内存时,SQL Server 的缓冲池就是索引:读走内存、热点写走内存页,IO 从关键路径上消失;读写分离 + 反范式把锁等待和 join 开销砍到最低;瓶颈收敛到单机后,运维复杂度是 O(1)——而人是最贵的变量。创始人原话:"RAM solves all problems. After that it's just growing using bigger machines."
When the working set fits in memory, SQL Server's buffer pool is the index: reads come from RAM, hot writes hit memory pages, and I/O disappears from the critical path; separating reads from writes plus denormalization cuts lock waits and join cost to the bone; with the bottleneck converged on one box, operational complexity is O(1) — and people are the most expensive variable. In the founder's words: "RAM solves all problems. After that it's just growing using bigger machines."
- 教训
读多写少、热点可缓存、数据装得进内存,是纵向扩展的"甜区":先算工作集/内存比,再谈分片。对称面是 bet365——同一公式在写密集型博彩负载下会撞墙,说明纵向扩展的适用条件是负载形状,不是信仰。
Read-heavy, cacheable hot spots, and a dataset that fits in RAM is vertical scaling's sweet spot: compute the working-set-to-RAM ratio before talking about sharding. The symmetric counterpoint is bet365 — the same formula hits a wall under write-intensive betting load, which shows vertical scaling's fit depends on workload shape, not belief.
相关产品:Microsoft SQL Server 相关能力:— 最后核验:2026-10-02
VSTS 2016 年故障:一个存储过程拖死 SQL 后端,故障转移也救不了 失败教训
运维事故
内存泄漏
故障转移
高可用
- 决策
按标准剧本,工程师先把 SQL 数据库故障转移(failover)到新主,指望新主恢复服务。
Engineers followed the standard playbook and failed over the SQL database to a new primary, expecting service to recover.
- 结果
故障转移只换来短暂缓解——同一个存储过程在新主上继续疯狂分配内存,新主很快也被拖进无响应;最终靠"给该存储过程手动分配内存上限"才止血。
The failover bought only brief relief — the same stored procedure kept allocating memory aggressively on the new primary, which soon became unresponsive too; the bleeding stopped only when engineers "manually assign[ed] allocation limits for the procedure."
- 机制根因
故障转移换的是"哪台机器当主",不换"谁在干坏事":泄漏内存的存储过程是负载本身,跟着连接一起漂到新主;SQL Server 的资源调控器没有给它设上限时,单个坏查询可以饿死整个实例。高可用解决的是机器故障,不解决"有毒负载"——后者是负载问题,不是拓扑问题。
Failover changes which machine is primary, not what is misbehaving: the memory-leaking stored procedure was the workload itself, and it floated to the new primary along with the connections; with no resource cap on it, a single bad query can starve an entire SQL Server instance. High availability solves machine failure, not toxic load — the latter is a workload problem, not a topology problem.
- 教训
故障转移 ≠ 故障隔离:有毒负载会跟着流量走,新主一样死。生产必备:给关键存储过程/工作负载设内存与 CPU 上限(Resource Governor 或等价手段);监控要盯"单个过程的内存分配增速",而不是只看实例级水位——等实例水位报警时,转移已经来不及了。
Failover is not fault isolation: toxic load follows the traffic and kills the new primary just the same. Production hygiene: cap memory and CPU per critical stored procedure or workload (Resource Governor or equivalent); monitor per-procedure memory allocation growth, not just instance-level watermarks — by the time the instance alarm fires, failover is already too late.
相关产品:Microsoft SQL Server 相关能力:AlwaysOn 可用性组 —— "大铁块"的高可用 最后核验:2026-10-02
Xero:30TB SQL Server 数仓库迁 Amazon Redshift 失败教训
数据仓库
上云迁移
列存储
SaaS
- 决策
数仓的数据库平台从 SQL Server 换成 Amazon Redshift;用 WhereScape 自动化把 30TB SQL Server 数仓整体迁移上 Redshift,保留原有逻辑与流程。
Switch the warehouse database platform from SQL Server to Amazon Redshift; migrate the 30TB SQL Server warehouse onto Redshift with WhereScape automation, preserving existing logic and processes.
- 结果
一名开发人员、6 个月工作量完成迁移,赶上公司 deadline;之后 BI 团队得以利用云的弹性做大数据与数据科学。
A single developer completed the migration in six months of effort, meeting the company deadline; the BI team could then exploit cloud elasticity for big data and data science work.
- 机制根因
行式 OLTP 数据库做分析是结构性错配:590 亿行的扫描在行存下是全表 IO 地狱,而列存(Redshift)的压缩 + 列裁剪 + MPP 把分析查询降 1–2 个数量级(该迁移的实测对比:行存全表扫描 vs 列存+MPP,非通用加速比)。SQL Server 不是"不够好",是赛道不对——数仓场景下换列存不是优化,是换赛道。
Running analytics on a row-based OLTP database is a structural mismatch: scanning 59 billion rows in row storage is full-table I/O hell, while columnar storage (Redshift) with compression, column pruning and MPP cuts analytical queries by one to two orders of magnitude (a measured comparison from this migration: row-store full scans vs columnar + MPP, not a universal speedup ratio). SQL Server was not "not good enough" — it was the wrong track; switching to columnar for warehousing is a change of track, not an optimization.
- 教训
区分"OLTP 够用"与"OLAP 错配":Xero 的 SQL Server 数仓"成功运行多年"恰恰说明行存也能跑分析——直到数据量和并发分析需求把它变成成本黑洞。迁移的最大成本往往不在目标库,而在重写 ETL 逻辑:自动化工具(WhereScape 模板化代码生成)把 30TB 迁移压缩到 1 人 6 个月,这是本案例真正可复制的经验。注意本案例出自 WhereScape 厂商稿。
Distinguish "OLTP suffices" from "OLAP mismatch": the fact that Xero's SQL Server warehouse "ran successfully for years" shows row stores can do analytics — until data volume and concurrent analytical demand turn them into cost black holes. The biggest migration cost is usually not the target database but rewriting ETL logic: automation (WhereScape's template-driven code generation) compressed a 30TB migration into one person and six months — the genuinely replicable lesson here. Note this case comes from a WhereScape vendor publication.
相关产品:Microsoft SQL Server、Amazon Redshift 相关能力:— 最后核验:2026-10-02
Stack Overflow:不分片,纵向扩展 SQL Server 扛下 15 亿月 PV 成功经验
纵向扩展
运维简单
缓存
查询优化
- 决策
不分片、不换 NoSQL,选择纵向扩展:单主 + SQL Server AlwaysOn 异步副本;Dell R720xd(384GB 内存、4TB PCIe SSD)少数几台机器扛下全站;全站只有一个存储过程,数据访问走 Dapper 微 ORM。
No sharding, no NoSQL — scale vertically instead: a single primary plus SQL Server AlwaysOn async replicas; a handful of Dell R720xd boxes (384GB RAM, 4TB PCIe SSD) carry the entire site; exactly one stored procedure site-wide, with data access through the Dapper micro-ORM.
- 结果
少数几台机器扛下全站流量,架构极简,运维负担极低。
A few machines carry all site traffic; the architecture stays radically simple and operationally cheap.
- 机制根因
"Our usage of SQL is pretty simple. Simple is fast." 查询简单意味着索引覆盖 + Redis 缓存就能解决绝大多数压力;单机纵向扩展的天花板远比直觉高;避开重型 ORM 在热点路径上生成低效 SQL,等于把单机性能吃满。
"Our usage of SQL is pretty simple. Simple is fast." Simple queries mean covering indexes plus Redis caching absorb almost all pressure; the ceiling of vertical scaling is far higher than intuition suggests; avoiding a heavy ORM generating bad SQL on hot paths extracts the full performance of each box.
- 教训
"不换"的机制条件:查询模式简单且热点可缓存、写入量未触及单机天花板。顺序应该是先优化查询、再加硬件、最后才重构架构——纵向扩展 + 激进缓存是性价比最高的"不选型"。一旦查询复杂度或写入并发突破单机天花板,这个公式立即失效,届时再谈分片。
The mechanism conditions for "not switching": query patterns are simple and cacheable, and write volume hasn't hit the single-machine ceiling. The order of operations should be: optimize queries first, add hardware second, re-architect last — vertical scaling plus aggressive caching is the highest-ROI "non-decision." Once query complexity or write concurrency breaks through the single-machine ceiling, the formula stops working, and only then is sharding on the table.
相关产品:Microsoft SQL Server 相关能力:AlwaysOn 可用性组 —— "大铁块"的高可用 最后核验:2026-10-01
Airbnb:指标平台 Minerva 与风控分析从 Druid/Presto 迁往 StarRocks(2023–2025) 成功经验
指标平台
实时风控
BI 加速
Druid 替代
- 场景
Airbnb 的 Minerva 指标平台管理 12,000+ 指标、4,000 维度(CelerData 案例口径);原架构是每天把数万个指标反范式化成宽表,再灌入 Druid/Presto/Spark/Hive 做查询。痛点有三:Druid 对高基数维度聚合能力弱;维度定义一变,数年历史数据要重算;Tableau 无法直连 Druid,复杂查询走 Presto 要 10 分钟以上。与此同时,Trust(信任与风控)团队需要对实时更新的业务数据做秒级分析。(CelerData 官方案例,基于 Airbnb 数据基础设施工程师 Jingwei Lu 2025-03-15 直播分享整理
https://celerdata.com/hubfs/Airbnb_Case_Study.pdf?hsLang=en)
Airbnb's Minerva metrics platform manages 12,000+ metrics and 4,000 dimensions (CelerData case figures). The old architecture denormalized tens of thousands of metrics into wide tables every day, then loaded them into Druid/Presto/Spark/Hive for querying. Three pain points: Druid was weak at high-cardinality dimension aggregation; any dimension-definition change forced recomputation of years of history; Tableau could not connect to Druid directly, and complex queries via Presto took 10+ minutes. Meanwhile the Trust (trust and safety) team needed second-level analytics over real-time-updated business data. (CelerData official case study, compiled from a 2025-03-15 live session by Airbnb data infrastructure engineer Jingwei Lu
https://celerdata.com/hubfs/Airbnb_Case_Study.pdf?hsLang=en)
- 决策
Minerva 与 Trust 分析统一迁往 StarRocks。选型关键原因是 CBO + 向量化执行引擎对复杂 join/聚合的加速能力,以及 MySQL 协议带来的生态兼容(Tableau 直连、98% SQL 兼容)。Trust 场景用了 StarRocks 的主键模型做实时更新,替代原来"天级产出"的风控分析链路。
Minerva and Trust analytics both moved to StarRocks. The key selection reasons were the CBO plus vectorized execution engine accelerating complex joins/aggregations, and MySQL-protocol ecosystem compatibility (direct Tableau connection, 98% SQL compatibility). The Trust scenario uses StarRocks' primary-key model for real-time updates, replacing a risk-analysis pipeline that previously produced results daily.
- 结果
案例中给出的标杆查询——20 台 EC2 i3.8xlarge(32 vCPU/240GiB 内存/3TB SSD),3 张表(5 亿/60 亿/1 亿行)、4 个 join、3 个 distinct count,外加 JSON 解析与正则过滤——StarRocks 3.6 秒出结果,同一查询 Presto 要 3–4 分钟;TB 级数据 1 小时内完成导入。Tableau 侧从"查 10 分钟"变为秒级交互;Trust 风控分析从天级降到秒级。以上数字为 CelerData 厂商案例口径,未找到 Airbnb 官方工程博客的对应数字。
The benchmark query in the case -- 20 EC2 i3.8xlarge instances (32 vCPU / 240GiB RAM / 3TB SSD), three tables (500M / 6B / 100M rows), 4 joins, 3 distinct counts, plus JSON parsing and regex filtering -- returned in 3.6 seconds on StarRocks vs 3-4 minutes on Presto for the same query; terabyte-scale data loaded within one hour. Tableau went from "10-minute queries" to second-level interactivity; Trust risk analytics went from daily to seconds. These figures are CelerData's vendor-case numbers; no matching figures were found on Airbnb's official engineering blog.
- 机制根因
按案例描述,Druid 路线的代价是"预聚合 + 反范式化":查询快的前提是 ETL 时把 join 做完、维度拍平,所以维度一变就要重算历史,灵活性很低;同一标杆查询 Presto 需要 3–4 分钟(案例口径),案例把选型理由归于 CBO 与向量化执行引擎。StarRocks 的做法是"实时更新的主键表 + 运行时的分布式 join":明细数据直接入,主键模型支持维度变更实时可见,CBO 决定 join 策略,分析师不再需要为性能预先 join。代价是把原来分散在 ETL、Druid、Presto 三处的调优工作,集中到了 StarRocks 集群的运维与 SQL 治理上。
Per the case, the cost of the Druid route is "pre-aggregation plus denormalization": fast queries presuppose joins done at ETL time and flattened dimensions, so any dimension change means recomputing history -- very low flexibility; the same benchmark query took Presto 3-4 minutes (case figures), and the case credits the selection to the CBO and vectorized execution engine. StarRocks' approach is "real-time-updating primary-key tables plus runtime distributed joins": detail data lands directly, the primary-key model keeps dimension changes visible in real time, and the CBO picks join strategies, so analysts no longer pre-join for performance. The price is concentrating tuning work previously spread across ETL, Druid, and Presto into StarRocks cluster operations and SQL governance.
- 教训
在 Airbnb 的配置下(12,000+ 指标、维度频繁变更,CelerData 口径),每天反范式化数万指标的 ETL 负担超过了查询本身,这时评估能实时 join 的引擎是合理的——这是该样本的经验法则,非定律;对 Airbnb 而言,BI 工具兼容性是硬约束:Tableau 直连 + MySQL 协议把他们的迁移摩擦降到最低;读厂商案例的 benchmark 查询要看"像不像自己的查询":4 join + distinct count + JSON/正则是 Airbnb 真实 workload 的样子,对有类似负载的团队比 TPC-H 更有参考价值。
Under Airbnb's configuration (12,000+ metrics, frequently changing dimensions, CelerData figures), the daily ETL burden of denormalizing tens of thousands of metrics exceeded the queries themselves, making it reasonable to evaluate an engine that joins in real time -- a rule of thumb from this sample, not a law; for Airbnb, BI-tool compatibility was a hard constraint: direct Tableau connectivity plus the MySQL protocol minimized their migration friction; judge a vendor case's benchmark query by "does it look like my queries": 4 joins plus distinct counts plus JSON/regex is what Airbnb's real workload looks like, more instructive than TPC-H for teams with similar loads.
相关产品:StarRocks 相关能力:实时更新 + 多表 JOIN:终结"反范式化噩梦" 最后核验:2026-10-02
携程 UBT:用户行为追踪从 ClickHouse 迁往 StarRocks 存算分离(2024–2025) 成功经验
用户行为分析
存算分离
ClickHouse 替代
写入稳定性
- 场景
携程用户行为追踪(UBT)系统日增 30TB 数据,保留 30 天(部分表 1 年),总量约 1PB,最大单表 400TB、1.8 万亿行(以上规模数字均为携程工程师在原文中的自述)。原架构 gohangout + ClickHouse:历史数据回补会触发 ClickHouse 分区合并,CPU/IO 瞬间飙高,进而写入丢失、消费积压;扩容要做数据迁移;三副本存储开销大;大时间跨度查询慢。(StarRocks 官方 CSDN 账号,作者魏宁/携程大数据平台开发专家,2025-10-18
https://blog.csdn.net/StarRocks/article/details/153530204)
Ctrip's user behavior tracking (UBT) system ingests 30TB of new data per day, retains 30 days (one year for some tables), totals about 1PB, with the largest single table at 400TB and 1.8 trillion rows (all scale figures are the Ctrip engineer's own statements in the article). The old stack was gohangout plus ClickHouse: historical backfills triggered ClickHouse partition merges, spiking CPU/IO, which then caused write loss and consumer lag; scaling out required data migration; three-replica storage was expensive; queries over large time ranges were slow. (StarRocks official CSDN account, by Wei Ning, Ctrip big-data platform development expert, 2025-10-18
https://blog.csdn.net/StarRocks/article/details/153530204)
- 决策
链路换成 Flink + StarRocks shared-data(存算分离):计算节点不持久化数据(shared-data 架构的无状态计算层),扩缩容时数据不动,数据存对象存储;分区策略改为小时分区 + 128 buckets + zlib 压缩(省约 30% 存储);对小时级聚合查询建了分区级物化视图。
The pipeline moved to Flink plus StarRocks shared-data (storage-compute separation): compute nodes persist no data (the stateless compute layer of the shared-data architecture) so scaling never moves data, while data lives on object storage; partitioning changed to hourly partitions with 128 buckets plus zlib compression (saving about 30% storage); partition-level materialized views were built for hourly aggregation queries.
- 结果
存储总量 2.6PB 降到 1.2PB;节点数 50 降到 40;查询平均耗时 1.4s 降到 203ms(约 1/7);P95 从 8s 降到 800ms(约 1/10);写入吞吐 300 万行/秒(10GB/s);MergeCommit 优化后 IOPS 从 140 降到 10 以下(约 10 倍)、I/O size 降约 10 倍。以上数字来自携程工程师署名的实践文章(发在 StarRocks 官方渠道),非独立第三方复测。
Total storage fell from 2.6PB to 1.2PB; nodes from 50 to 40; average query latency from 1.4s to 203ms (about one-seventh); P95 from 8s to 800ms (about one-tenth); write throughput of 3 million rows per second (10GB/s); after MergeCommit optimization, IOPS dropped from 140 to under 10 (about 10x) and I/O size fell about 10x. These figures come from a practice article signed by a Ctrip engineer (published via StarRocks' official channel), not an independent third-party re-test.
- 机制根因
按原文描述,ClickHouse 部署中存储与计算在同一批节点:历史回补会触发分区合并,合并时的 CPU/IO 飙升直接打爆在线写入,导致写入丢失与消费积压;扩容则要做数据迁移(此处仅复述原文的现象描述,不展开 ClickHouse 内部存储机制的归因)。存算分离把这两件事解耦:写入抖动只影响计算节点,扩容加计算节点即可,数据不动;对象存储 + 副本数下降带来存储成本下降(原文:2.6PB 降到 1.2PB,携程口径)。代价是查询要从远端拉数据,对缓存(data cache)和网络提出了更高要求,携程为此做了大量 I/O 路径优化(MergeCommit 就是一例)。
Per the article, in the ClickHouse deployment storage and compute shared the same nodes: historical backfills triggered partition merges, and the CPU/IO spikes during merges directly killed online writes, causing write loss and consumer lag; scaling out required data migration (this restates the article's observed phenomena only, without attributing ClickHouse's internal storage mechanics). Storage-compute separation decouples the two: write jitter only affects compute nodes, scaling adds compute nodes without touching data, and object storage plus fewer replicas cut storage cost (article: 2.6PB down to 1.2PB, Ctrip figures). The price is queries fetching data remotely, which demands more from the cache (data cache) and the network -- Ctrip did extensive I/O-path optimization for this (MergeCommit being one example).
- 教训
在携程 UBT 的规模下(日增 30TB、总量约 1PB,携程口径),shared-nothing 部署的运维负担(回补、扩容、副本)随数据量放大,存算分离是他们选择的规模化路径——这是单一样本的结论,不推导为普遍规律;在该案例中,分区粒度匹配查询模式——小时分区 + 高 bucket 数对应"按小时查行为";携程把物化视图建在查询模式稳定的小时聚合层,这个做法是否适用于其他场景,取决于查询模式是否同样稳定。
At Ctrip UBT's scale (30TB new per day, about 1PB total, Ctrip figures), the ops burden of shared-nothing deployment (backfills, scaling, replicas) grew with data volume, and storage-compute separation was the scaling path they chose -- a single-sample conclusion, not a universal rule; in this case, partition granularity mirrored query patterns -- hourly partitions plus high bucket counts matching "query behavior by the hour"; Ctrip built materialized views on the hourly aggregation layer where query patterns were stable, and whether that fits other scenarios depends on whether their query patterns are equally stable.
相关产品:StarRocks、ClickHouse 相关能力:异步物化视图 + CBO 自动改写:查询加速的"隐形外挂" 最后核验:2026-10-02
得物:运营 B 端 StarRocks 流量切往 OceanBase(部分流量 POC 进行中)(2026) 失败教训
HTAP 替代
应用层分流之痛
行列混合
- 场景
In Dewu's data architecture, StarRocks served B-side analytics traffic: merchant B-side traffic went to StarRocks for big campaigns and to MySQL for small ones, with a sync tool replicating MySQL data into StarRocks; the operations B-side originally ran entirely on StarRocks. The application layer had to manually decide which database each request should hit -- a wrong route meant wrong data; MySQL and StarRocks stored two copies of the data; and there was an extra sync pipeline to operate. (Third-party study notes transcribing a Dewu tech public-account article
https://github.com/ygrowly/marvis/blob/HEAD/docs/%E6%95%B0%E6%8D%AE%E5%BA%93/OceanBase/%E5%BE%97%E7%89%A9%20-%20OceanBase%20%E5%AE%9E%E8%B7%B5.md ; the notes cite the original as the Dewu tech article "Dewu OceanBase Practice" by Feilisi and Weiwen, 2026-07-08)
- 决策
引入 OceanBase(行列混合的 HTAP),把 StarRocks 的部分流量切到 OceanBase 做 POC;运营 B 端已全部切到 OceanBase。选型逻辑:OceanBase 一套系统同时扛 TP 写入与 AP 查询,应用层不再需要手动分流,同步链路也可以砍掉。
Adopt OceanBase (row-column hybrid HTAP), shifting some StarRocks traffic to OceanBase as a POC; the operations B-side has since moved entirely to OceanBase. Selection logic: one OceanBase system handles both TP writes and AP queries, so the application layer no longer routes manually and the sync pipeline can be removed.
- 结果
据笔记转述的原文口径:存储从 TB 级降到百 GB 级;压缩率比 StarRocks 高 30%;聚合与范围查询平均快 65%(各场景 10%–90% 不等);高 QPS 点查与 MySQL 相当、显著优于 StarRocks;OceanBase 写入后完全实时可见,而 StarRocks 有秒级写读延迟。——以上数字全部来自第三方笔记对得物技术文章的转述,未找到得物官方或 OceanBase 官方的对应发布,引用时必须注明口径;原文微信链接因平台验证限制未能直接打开核验。另据笔记,截至 2026-07-08 文章发表时,运营 B 端已全量切换,但部分流量的 OceanBase POC 仍在进行中(未完成)。
Per the figures quoted in the notes from the original article: storage fell from terabyte-scale to hundreds of gigabytes; compression 30% better than StarRocks; aggregation and range queries 65% faster on average (10%-90% across scenarios); high-QPS point lookups on par with MySQL and significantly better than StarRocks; OceanBase reads-your-writes fully real-time, while StarRocks has second-level write-read delay. -- All figures above come from third-party notes transcribing Dewu's tech article; no matching publication was found from Dewu officially or OceanBase officially, so the attribution must be stated when quoting; the original WeChat link could not be opened directly for verification due to platform verification limits. Per the notes, as of the article's publication on 2026-07-08, the operations B-side had fully migrated, but the partial-traffic OceanBase POC was still incomplete.
- 机制根因
StarRocks 是为 OLAP 生的:列存、批量写入、高吞吐扫描;但 B 端业务的查询画像是"高 QPS 点查 + 频繁小批量写入 + 事务语义",这是 OLTP/HTAP 的主场。在得物 B 端这类"高 QPS 点查 + 频繁小批量写入 + 事务语义"的负载下,StarRocks 的秒级写读延迟与点查性能成为瓶颈(得物文章口径);这是否是 OLAP 引擎的普遍短板,单一样本无法证明。OceanBase 的行列混合存储让同一份数据既能做事务写入点查,又能做列存分析,替代了"MySQL + 同步 + StarRocks"的三段式架构。故事的另一面是:当初"小活动走 MySQL、大活动走 StarRocks"的应用层手动分流,本质是把本该由数据库解决的路由问题推给了业务代码,这是典型的架构债。
StarRocks was born for OLAP: columnar storage, batch writes, high-throughput scans; but B-side business query patterns are "high-QPS point lookups plus frequent small-batch writes plus transactional semantics" -- home turf for OLTP/HTAP. Under Dewu B-side's "high-QPS point lookups plus frequent small-batch writes plus transactional semantics" load profile, StarRocks' second-level write-read delay and point-lookup performance became bottlenecks (Dewu article claims); whether this is a universal shortcoming of OLAP engines cannot be proven from a single sample. OceanBase's row-column hybrid storage lets one copy of data serve both transactional writes/point lookups and columnar analytics, replacing the three-stage "MySQL plus sync plus StarRocks" architecture. The other side of the story: the original "small campaigns on MySQL, big campaigns on StarRocks" manual application-layer routing pushed a routing problem that belonged in the database up into business code -- classic architectural debt.
- 教训
在得物 B 端这类"高 QPS 点查 + 频繁小批量写入 + 事务语义"的负载下,继续用 OLAP 引擎扛 TP 负载被证明走不通(得物口径);是否推广为一般规律,需要更多样本;应用层手动分流双写是危险信号——在该案例中它意味着数据架构缺了一块(能同时做 TP 和 AP 的引擎),其他场景是否同样成立需单独评估;看厂商/用户自述的替换案例时,注意"被替换方"的视角缺失:这篇是得物单方面口径,StarRocks 侧没有回应,数字不可直接当作两种产品的客观对比。
Under Dewu B-side's "high-QPS point lookups plus frequent small-batch writes plus transactional semantics" load profile, continuing to carry TP load on an OLAP engine proved unworkable (Dewu claims); generalizing this needs more samples; manual application-layer dual-write routing is a danger signal -- in this case it meant the data architecture was missing a piece (an engine that does TP and AP together), and whether that holds elsewhere needs separate evaluation; when reading vendor/user self-reported replacement stories, watch for the missing perspective of the replaced party: this is Dewu's one-sided account with no response from the StarRocks side, so the numbers cannot be treated as an objective product comparison.
来源
第三方学习笔记 ygrowly/marvis《得物 - OceanBase 实践》(转述得物技术公众号文章,作者菲利斯、未文,2026-07-08
Third-party study notes ygrowly/marvis "Dewu - OceanBase Practice" (transcribing a Dewu tech public-account article by Feilisi and Weiwen, 2026-07-08
相关产品:StarRocks、OceanBase 相关能力:— 最后核验:2026-10-02
Didi:多 OLAP 引擎(Druid/Kylin/Presto/ClickHouse)统一到 StarRocks(2023) 成功经验
多引擎整合
成本优化
Druid 替代
实时监控
- 场景
Didi's data infrastructure generates petabytes of data daily; its OLAP layer started as a mix of Apache Druid, Apache Kylin, Presto, and ClickHouse. As data volume and analytics demand grew, multi-engine problems erupted across the board: high operational complexity; big feature differences between engines left users confused about which to use, with SQL dialects not even mutually intelligible; no real-time streaming inserts/updates/deletes; no open table formats, so scalability could not keep up. (StarRocks official blog, 2023-12-01
https://www.starrocks.io/blog/reduced-80-cost-didis-journey-from-multiple-olap-engines-to-starrocks/)
- 决策
把 OLAP 统一迁往 StarRocks,看中的四点:统一平台整合实时与历史数据、高性能(列存 + 向量化 + 分布式)、弹性扩展、原生支持 Iceberg/Hudi/Hive/Delta Lake 的湖仓架构。
Consolidate OLAP onto StarRocks, chosen for four reasons: a unified platform merging real-time and historical data, high performance (columnar plus vectorized plus distributed), elastic scaling, and a lakehouse architecture natively supporting Iceberg/Hudi/Hive/Delta Lake.
- 结果
监控告警业务(每秒 45 万条写入、日增 12TB):查询性能提升 4 倍,P90 从 500ms 降到 150ms;集群从 60+ 节点缩到不足 10 个;存储量减少约 40%;总体成本降低 80% 以上。金融数据门户(日增约 10 亿条):靠分区分桶、前缀/ZoneMap/布隆/Bitmap 索引与物化视图做到秒级响应,支持自定义时间维度和跨数据集 join。截至 2023 年,滴滴全公司 30+ StarRocks 集群、300+ TB 数据、日均 QPS 400 万以上,覆盖网约车、顺风车、两轮车、金融、能源多条业务线。以上为 StarRocks 官方博客口径(厂商渠道)。
Monitoring and alerting workload (450,000 writes per second, 12TB per day): query performance up 4x, P90 from 500ms down to 150ms; cluster shrunk from 60+ nodes to fewer than 10; storage volume down about 40%; overall cost down more than 80%. Financial Data Portal (about 1 billion new rows per day): second-level responses via partitioning and bucketing, prefix/ZoneMap/Bloom/Bitmap indices, and materialized views, supporting custom time dimensions and cross-dataset joins. As of 2023, Didi ran 30+ StarRocks clusters company-wide with 300+ TB of data and 4,000,000+ daily average QPS, serving ride-hailing, carpooling, two-wheelers, finance, and energy business lines. These are StarRocks official blog figures (vendor channel).
- 机制根因
多引擎的本质成本是"能力交集":每个引擎只擅长一种场景(Druid 预聚合、Kylin Cube、Presto 联邦查询、ClickHouse 单表扫描),业务稍复杂就要跨引擎拼链路,数据搬运、口径对齐、运维排障的成本全是隐性的。按官方博客的说法,StarRocks 的"全场景"定位(实时写入 + join + 湖上查询)让一条链路覆盖原来四种引擎的场景,节点数从 60+ 降到 10 以内;省下的不只是机器,还有多引擎并存时的运维与排障人力(原文列为迁移动因之一)。代价是单引擎依赖:所有 OLAP 负载的稳定性都押在 StarRocks 一个系统上——这是架构集中化的固有 trade-off,原文未量化其风险。
The true cost of multi-engine is the "capability intersection": each engine excels at one scenario (Druid pre-aggregation, Kylin cubes, Presto federated queries, ClickHouse single-table scans), so anything slightly complex means stitching pipelines across engines, with data movement, metric alignment, and troubleshooting costs all hidden. Per the official blog, StarRocks' "all-scenario" positioning (real-time writes plus joins plus lake queries) lets one pipeline cover what four engines used to do, with nodes falling from 60+ to under 10; the savings were not just machines but also the operations and troubleshooting headcount of running multiple engines (listed in the article as a migration reason). The price is single-engine dependence: all OLAP workload stability now rests on StarRocks alone -- an inherent trade-off of architectural consolidation, whose risk the article does not quantify.
- 教训
在滴滴 2023 年的规模下,"一个场景一个引擎"从早期的敏捷变成了技术债——这是该样本在该时代的结论;在滴滴的案例中,多引擎的隐性成本(人力、数据搬运、口径对齐)是迁移的主要动因之一,是否"经常超过"机器成本,单一样本无法证明;滴滴的选型标准是"场景覆盖度"而非单项性能冠军——这是该案例的选型逻辑;从数字看,节点数从 60+ 降到 10 以内是成本下降 80%+ 的重要组成部分(官方博客口径),各因素的具体贡献比例原文未披露。
At Didi's 2023 scale, "one engine per scenario" turned from early agility into tech debt -- a conclusion from this sample in that era; in Didi's case, the hidden costs of multi-engine (headcount, data movement, metric alignment) were a major migration reason, though whether they "often exceed" machine costs cannot be proven from a single sample; Didi's selection criterion was scenario coverage, not a single benchmark crown -- that is this case's selection logic; judging from the figures, nodes falling from 60+ to under 10 was a major component of the 80%+ cost reduction (official blog figures), and the article does not break down each factor's contribution.
相关产品:StarRocks、ClickHouse 相关能力:实时更新 + 多表 JOIN:终结"反范式化噩梦" 最后核验:2026-10-02
汽车之家:多引擎对比测试后选 StarRocks,120 亿数据下 HLL 查询明显优于 ClickHouse(2021) 成功经验
PoC
选型评估
OLAP 选型
实时分析
多引擎对比
- 场景
汽车之家(NYSE: ATHM)智能推荐效果分析、物料点击曝光、流量宽表等场景需要秒级实时分析。原有方案痛点:Flink 聚合不灵活、需求多变时重复开发成本高;Kylin 预计算模型并发好但不支持明细下钻;TiDB 不支持预聚合模型,大数据量下在线聚合导致服务器压力骤增、查询性能不稳定。2021 年实时计算平台负责人邸星星牵头做多引擎选型。
Autohome (NYSE: ATHM) needed second-level real-time analytics for intelligent recommendation analysis, material click/exposure stats, and traffic wide tables. Pain points in the incumbent setup: Flink aggregations were inflexible with high re-development cost for changing requirements; Kylin's pre-computation model had good concurrency but couldn't drill into detail rows; TiDB lacked a pre-aggregation model, so large-volume online aggregation spiked server pressure and made query performance unstable. In 2021, Di Xingxing (head of the real-time computing platform) led a multi-engine evaluation.
- 决策
对 StarRocks 与常用 OLAP 引擎做对比测试,测试注明数据量/机器数/用例:vs Kylin(6 亿数据/2 台):命中物化视图场景下 StarRocks 媲美甚至超过 Kylin;vs ClickHouse(120 亿数据/4 台):Count 场景接近,HLL 非精确去重场景 StarRocks 明显更优(且 1.18 版 HLL 比 1.17 提升 3–4 倍);vs Doris(6 亿/2 台):向量化引擎带来 2–7 倍提升;vs Presto/Spark(10 亿/8 台对等资源):StarRocks 最优。ClickHouse 落选理由:运维成本高、多表关联弱。
Head-to-head tests of StarRocks against common OLAP engines, each documented with data volume/machine count/cases: vs Kylin (600M rows/2 nodes): StarRocks matched or beat Kylin on materialized-view-hitting queries; vs ClickHouse (12B rows/4 nodes): Count queries were close, StarRocks clearly won on HLL approximate-distinct queries (and v1.18 improved HLL 3–4x over v1.17); vs Doris (600M/2 nodes): 2–7x gains from the vectorized engine; vs Presto/Spark (1B rows/8 nodes, equal resources): StarRocks fastest. ClickHouse was rejected for high ops cost and weak multi-table joins.
- 结果
选定 StarRocks 为实时 OLAP 引擎。按汽车之家自述口径:统一了明细查询与预聚合两种模型、流批统一(一份引擎同时服务实时与离线)、MySQL 协议运维简单。生产落地两个业务:推荐服务实时监控(TP95 约 1 秒)、搜索实时效果分析(1.17 十几秒→升级 1.18 后四秒内)。业务规模口径:3000+ 协议/维度表、峰值 33 亿条/分钟、3 万亿/天、并发 33w/分钟。
StarRocks was selected as the real-time OLAP engine. Per Autohome's own account: it unified detail queries and pre-aggregation models, unified streaming and batch (one engine serving both real-time and offline), and the MySQL protocol kept ops simple. Two production use cases: real-time recommendation monitoring (TP95 ~1s) and real-time search effectiveness (over ten seconds on 1.17 → under four seconds after upgrading to 1.18). Business scale per Autohome: 3,000+ protocols/dimension tables, peak 3.3B rows/min, 30T rows/day, 330K calls/min concurrency.
- 机制根因
汽车之家的决策逻辑是"模型覆盖度"优先于"单项跑分":Kylin 赢在固定报表、输在明细下钻;ClickHouse 赢在明细扫描、输在关联与运维;StarRocks 的"明细 + 预聚合 + 向量化"三合一恰好覆盖其"既要下钻又要报表"的 workload。HLL 场景的胜出有版本注脚——1.18 针对 HLL 的优化是测试当时的增量红利,说明"拿哪个版本测"本身就是评估变量。
Autohome's decision logic prioritized "model coverage" over "single-benchmark wins": Kylin won fixed reports but lost detail drill-down; ClickHouse won detail scans but lost joins and operability; StarRocks' "detail + pre-aggregation + vectorization" trio exactly covered their "drill-down plus reporting" workload. The HLL win came with a version footnote — v1.18's HLL optimization was the incremental dividend at test time, a reminder that "which version you test" is itself an evaluation variable.
- 教训
多引擎对比测试必须注明数据量/机器数/用例三要素,否则倍数没有意义;"运维成本"要进打分项(ClickHouse 就是栽在这里);版本演进快的引擎,选型报告要写清测试版本号。**渠道注记:本文托管于镜舟官网(StarRocks 商业公司渠道),但为汽车之家工程师第一人称撰写的会议演讲实录,引用时须标注渠道。**
Multi-engine benchmarks must document data volume, machine count, and test cases, or the multiples are meaningless; "operational cost" belongs on the scorecard (that's where ClickHouse lost); for fast-evolving engines, record the tested version in the evaluation report. **Channel note: hosted on mirrorship.cn (the StarRocks commercial vendor's site), but written in first person by the Autohome engineer as a conference talk — disclose the channel when citing.**
相关产品:StarRocks、ClickHouse、Apache Doris 相关能力:明细与预聚合统一模型 最后核验:2026-10-02
TRM Labs:PB 级区块链分析从 BigQuery 迁往 Iceberg + StarRocks 湖仓(2024–2025) 成功经验
湖仓架构
BigQuery 替代
区块链数据分析
Iceberg
- 场景
TRM Labs does blockchain intelligence and anti-money-laundering analytics. Its data platform serves petabyte-scale on-chain data across 30+ blockchains at 500+ queries per minute, with the largest single workload above 115TB and growing 2-3% per month. Queries are complex multi-layer joins with array filtering, and the internal target is P95 under 3 seconds. The old stack was BigQuery plus distributed Postgres (Citus): BigQuery could not satisfy "localized multi-site deployment" (data residency / multi-region requirements), while Citus scaling costs rose as data volume and query complexity grew (both are migration reasons TRM stated in the blog). (
https://www.trmlabs.com/trm-tech-blog/from-bigquery-to-lakehouse-how-we-built-a-petabyte-scale-data-analytics-platform-part-1)
- 决策
自建湖仓:存储层统一到 Iceberg(开放表格式,摆脱厂商锁定),查询层选 StarRocks 做高性能 serving 引擎。选型时 TRM 用自家负载做了 2024 年的 benchmark(2.57TB 数据集,TRM 自测口径):StarRocks 带缓存 470ms,对 Trino 1,030–1,410ms、DuckDB 2–3 秒;复杂聚合查询 StarRocks 无缓存约 2 秒、有缓存约 500ms,对 Trino 上限约 2.5 秒。TRM 在博客中明确写道,他们已经 "moved beyond traditional OLAP data stores (e.g., Clickhouse)"——即 ClickHouse 这类传统 OLAP 也在评估中被排除。
Build their own lakehouse: Iceberg as the unified storage layer (open table format, no vendor lock-in) and StarRocks as the high-performance serving engine. Their 2024 benchmark on their own workload (2.57TB dataset, TRM self-tested figures): StarRocks 470ms with cache vs Trino at 1,030-1,410ms and DuckDB at 2-3 seconds; on complex aggregations StarRocks took about 2s uncached / 500ms cached vs Trino capped around 2.5s. TRM wrote explicitly that they had "moved beyond traditional OLAP data stores (e.g., Clickhouse)" -- ClickHouse-style traditional OLAP was also ruled out in the evaluation.
- 结果
TRM 自述迁移后查询 P95 降低 50%、查询超时降低 54%——两个数字均为 TRM 企业技术博客自述口径,未找到第三方独立复现。Iceberg 统一存储后,同一份数据可同时被 StarRocks(serving)、Trino/Spark(离线)消费,不再为每个引擎维护一份拷贝。
TRM reports query P95 down 50% and query timeouts down 54% after migration -- both figures are TRM's own claims from their engineering blog, with no independent third-party reproduction found. With Iceberg as unified storage, one copy of the data is now consumed by StarRocks (serving) and Trino/Spark (offline) alike, instead of maintaining a copy per engine.
- 机制根因
BigQuery 的问题不在性能而在部署模型:serverless 按量计费 + Google 托管的公有云服务,从部署模型上就不支持 TRM 要求的本地化多站点;Citus 的问题在扩展模型:分布式 Postgres 的 coordinator 瓶颈与分片运维成本随数据量增加而放大(取决于分片策略与自动化程度)。StarRocks 在该 workload 上占优的可能原因有三(基于 TRM 公布的 benchmark 结果的分析,非 TRM 原文表述):列存向量化执行对"多层 join + array 过滤"友好、数据缓存让热查询进入亚秒(benchmark 明确对比了带缓存与无缓存)、CBO 能处理分析师手写的复杂 SQL。Iceberg 则解决了"一份数据多引擎读"的存储统一问题。代价是团队要自己运维整套湖仓(对象存储、表格式、引擎),把云厂商的托管红利换成了架构自主权——这笔账只在团队具备运维分布式系统能力时才划算。
BigQuery's problem was not performance but the deployment model: serverless pay-as-you-go plus Google hosting is incompatible with TRM's localized multi-site requirement by deployment model; Citus's problem was the scaling model: the coordinator bottleneck and sharding ops cost of distributed Postgres grow with data volume (depending on sharding strategy and automation maturity). StarRocks' edge on this workload plausibly came from three factors (analysis based on TRM's published benchmark results, not TRM's own wording): columnar vectorized execution friendly to "multi-layer joins + array filtering", a data cache pushing hot queries sub-second (the benchmark explicitly compared cached vs uncached), and a CBO handling analyst-written complex SQL. Iceberg solved the "one copy of data, many engines" storage unification problem, at the price of the team operating the whole lakehouse stack (object storage, table format, engines) themselves -- trading cloud-vendor managed convenience for architectural autonomy, a trade that only pays off if the team can operate distributed systems.
- 教训
对 TRM 这类有数据本地化硬约束的公司,选 OLAP 引擎要先看部署约束再看 benchmark:在该约束下,BigQuery 得零分与其性能无关;"一份存储、多种消费"(Iceberg + serving 引擎)是 TRM 验证可行的架构,但前提是团队有运维分布式系统的能力——这是单一样本的结论,不推导为普适规律;benchmark 测自己的真实查询模式值得借鉴——TRM 的 2.57TB 数据集测的是他们自己的 join/array 负载而非 TPC-H 通用集,但结论只适用于 TRM 的负载特征。
For companies like TRM with hard data-localization constraints, check deployment constraints before benchmarks when picking an OLAP engine: under that constraint, BigQuery scores zero regardless of its performance; "one copy of storage, many consumers" (Iceberg plus a serving engine) is an architecture TRM proved workable, but only if the team can operate distributed systems -- a single-sample conclusion, not a universal rule; benchmarking your own real query patterns is worth borrowing -- TRM's 2.57TB dataset tested their own join/array workload rather than the generic TPC-H suite, but the conclusions only apply to TRM's workload profile.
相关产品:StarRocks、Google BigQuery、DuckDB 相关能力:— 最后核验:2026-10-02
腾讯微信:基于 StarRocks 的湖仓一体与实时增量物化视图(2024–2025) 成功经验
湖仓一体
实时物化视图
降本
离线加工
- 场景
Multiple WeChat business lines (Channels live streaming, WeChat keyboard, WeRead, Official Accounts) run analytics on StarRocks, with clusters of hundreds of nodes and nearly a trillion ingested rows. The WeChat team positions StarRocks deliberately: ClickHouse is the "real-time warehouse" (sub-second / real-time / high QPS / OLAP-first), StarRocks is the "lakehouse" (seconds-level / near-real-time / low QPS / offline-processing-first) -- the two are a division of labor, not substitutes. (Conference slide deck, by Feng Lyu, Tencent WeChat OLAP kernel R&D engineer
https://ucasfl.github.io/slides/Lakehouse-practice-in-WeChat-based-on-StarRocks.pdf)
- 决策
在 StarRocks 上建湖仓一体架构:湖上直接建仓,离线加工任务跑在 StarRocks 而非 Spark;与社区共建"实时增量物化视图"能力;用基于全局字典的维表关联替代部分 Flink 实时任务。
Build a lakehouse architecture on StarRocks: build the warehouse directly on the lake, running offline processing jobs on StarRocks instead of Spark; co-build "real-time incremental materialized views" with the community; replace some Flink streaming jobs with global-dictionary-based dimension joins.
- 结果
存储成本降低 65% 以上;直播业务在湖上建仓改造后,运维任务数减半;离线任务产出时间缩短 2 小时。以上为演讲人口径(腾讯微信工程师大会分享),未找到独立复测。
Storage cost down more than 65%; after the lakehouse migration, the live-streaming business halved its ops task count; offline job output arrived 2 hours earlier. These are the speaker's figures (Tencent WeChat engineer conference talk); no independent re-test found.
- 机制根因
传统"湖 + 仓"两套系统要维护两份数据和两套任务。按演讲人描述,StarRocks 的外表能力让"湖上建仓"成为可能:离线加工直接对湖里的数据做 SQL,产出 StarRocks 内表供查询,省掉 Spark 集群与数据搬运;"实时增量物化视图"(微信与社区共建的能力)用于替代部分 Flink 实时任务,维表关联基于全局字典。代价是把原来 Spark/Flink 生态承担的调度、容错,换成了对 StarRocks 单一引擎的深度依赖,对内核能力(如物化视图的正确性)要求极高——这部分风险演讲人未展开,此处不做推断。
A traditional "lake plus warehouse" means maintaining two copies of data and two sets of jobs. Per the speaker, StarRocks' external-table capability makes "building the warehouse on the lake" possible: offline processing runs SQL directly against lake data and produces StarRocks internal tables for serving, eliminating the Spark cluster and data movement; "real-time incremental materialized views" (a capability WeChat co-built with the community) replace some Flink streaming jobs, with dimension joins based on the global dictionary. The price is trading the scheduling and fault tolerance previously provided by the Spark/Flink ecosystem for deep dependence on a single engine -- StarRocks -- which demands extremely high kernel capability (e.g., materialized-view correctness); the speaker did not elaborate on these risks, so no inference is made here.
- 教训
这是微信在该时期的分工决策:ClickHouse 管实时高 QPS,StarRocks 管湖仓离线加工——是否照搬取决于自家 workload 画像,而非唯性能论;在该案例中,湖仓一体的降本主要来自"少维护一套系统 + 少存一份数据",而不是查询更快;微信把实时增量物化视图的需求做成了上游特性——这是该案例的做法,有内核研发能力的团队可以借鉴,不推导为普适方法。
This was WeChat's division of labor at the time: ClickHouse handles real-time high QPS, StarRocks handles lakehouse offline processing -- whether to copy it depends on your own workload profile, not on performance alone; in this case, lakehouse cost savings came mostly from "one fewer system to maintain plus one fewer copy of data", not from faster queries; WeChat turned its real-time incremental materialized view requirement into an upstream feature -- that is this case's approach, worth borrowing for teams with kernel R&D capacity, not a universal method.
相关产品:StarRocks、ClickHouse 相关能力:异步物化视图 + CBO 自动改写:查询加速的"隐形外挂" 最后核验:2026-10-02
富融银行:香港持牌数字银行 10 个月换"心",15 小时完成新核心切换(2024) 成功经验
银行核心系统升级
数字银行
核心切换
成本优化
出海
- 场景
富融银行是香港持牌的数字银行,2020 年 12 月正式开业。近几年业务快速发展,原有银行核心系统已不能完全满足日渐增长的服务需求。为给客户提供更优质高效的金融服务,富融银行 2023 年 12 月开启新一代银行核心系统切换项目。一个成熟的银行切换核心如同"飞机在空中换引擎":一方面要在符合监管要求和市场标准的前提下,完成涵盖零售存款、零售贷款、公司存款、公司贷款和外汇服务五个核心业务领域的 150 多个子系统的整合改造;另一方面从数据中心选址到上层应用,涉及大量数据迁移与系统切换,还要尽可能减少对用户的影响。
Fusion Bank is a Hong Kong licensed digital bank that opened in December 2020. Rapid business growth meant its legacy core system could no longer fully meet rising service demand. To deliver better, more efficient financial services, Fusion Bank kicked off a new-generation core cutover project in December 2023. Switching a mature bank's core is like "changing the engine mid-flight": on one hand, it had to integrate and rework 150+ subsystems across five core banking domains (retail deposits, retail lending, corporate deposits, corporate lending, FX services) while meeting regulatory and market standards; on the other hand, everything from data-center site selection to upper-layer applications involved massive data migration and system cutover, with minimal impact on customers.
- 决策
依托微众银行自主开发的数字银行底座技术体系 + 腾讯云 TCE 专有云 + 腾讯云数据库 TDSQL,构建全新云计算环境并切换新一代核心。数据库选 TDSQL 的关键原因是其 100% 兼容 MySQL 的特性,可与富融银行基于 MySQL 的业务系统无缝整合;TDSQL 的批量迁移能力助力实现无感迁移,减少切换对用户的影响;TCE 则通过容器化部署提升资源使用效率。
Built on WeBank's self-developed digital-banking technology stack plus Tencent Cloud TCE (dedicated cloud) plus Tencent Cloud TDSQL, Fusion Bank constructed a brand-new cloud environment and cut over to the new core. The key reason for choosing TDSQL was its 100% MySQL compatibility, seamlessly integrating with Fusion Bank's MySQL-based business systems; TDSQL's bulk-migration capability enabled imperceptible migration, minimizing cutover impact on customers; TCE's containerized deployment raised resource efficiency.
- 结果
2024 年 10 月新一代核心系统成功上线,切换仅用时 10 个月,为香港银行核心系统升级树立新标杆;上线切换仅用 15 小时,在 6 小时的数据迁移窗口期内顺利将来自不同厂商的系统和数据迁移至新核心,实现无缝升级。富融银行副行政总裁、首席技术官邱家骅表示:预计与 2024 年相比,2027 年的 IT 非人力成本 3 年内将显著减少 53%,新需求开发时间由平均 6 个月下降至 3 个月。TCE 容器化使资源成本降低 40%,业务恢复时间 RTO 缩短至 30 分钟内。诚实标注:53% 降本与开发提速均为银行管理层的"预计"口径(2027 年对 2024 年的预测值),不是已实测结果;其余数字来自腾讯云厂商口径,未找到第三方独立复现。
The new-generation core went live in October 2024 with the cutover taking only 10 months - a new benchmark for Hong Kong bank core upgrades; the go-live cutover took just 15 hours, migrating systems and data from multiple vendors into the new core within a 6-hour data-migration window for a seamless upgrade. Fusion Bank Deputy CEO and CTO Qiu Jiahua stated: compared with 2024, IT non-labor costs in 2027 are expected to drop significantly by 53% within three years, and new-feature development time from 6 months to 3 months on average. TCE containerization cut resource costs by 40% and shrank business RTO to within 30 minutes. Honest note: the 53% cost reduction and faster development are management "expectations" (2027-vs-2024 forecasts), not measured results; remaining figures are Tencent/Fusion Bank figures with no independent third-party reproduction found.
- 机制根因
MySQL 100% 兼容是迁移可行性的前提——业务系统基于 MySQL 构建,无需重写即可无缝整合,这是"换引擎"风险可控的第一道保险;TDSQL 的批量迁移能力把 150+ 子系统、多厂商来源的数据在 6 小时窗口内搬完,迁移工具链的成熟度直接决定了切换窗口能否压缩到 15 小时;复用微众银行的数字银行底座技术体系,等于把一家互联网银行十年沉淀的架构拿来即用,避免了从零自研;TCE 容器化部署同时解决了资源效率(降本 40%)与 RTO(30 分钟内)两个问题。
100% MySQL compatibility was the precondition for migration feasibility - business systems were MySQL-based, so they integrated seamlessly without rewrites, the first insurance policy for a controllable "engine change"; TDSQL's bulk-migration capability moved 150+ subsystems and multi-vendor data within the 6-hour window - migration-toolchain maturity directly determined whether the cutover window could be compressed to 15 hours; reusing WeBank's digital-banking technology stack meant adopting a decade of an internet bank's architecture off the shelf instead of building from zero; TCE containerized deployment solved resource efficiency (40% cost cut) and RTO (under 30 minutes) together.
- 教训
数字银行换"心"可以借力——"买成熟底座 + 换数据库"比从零自研快数倍到一个数量级(该项目 10 个月上线 vs 行业常见的数年周期——单一样本对比,不可推广为通用倍数),富融 10 个月上线 vs 行业常见的数年周期,差距主要来自底座复用而非数据库本身;把"切换时长"(10 个月、15 小时、6 小时窗口)作为公开承诺指标反向倒逼迁移工具链成熟,是项目管理的有效手段;成本数字凡是"预计"必须注明预测口径,53% 降本在 2027 年兑现前都只是管理层预期;香港持牌数字银行的案例说明 TDSQL 的金融叙事已经走出内地监管语境,在境外合规框架下同样成立,这是"政策红利型招牌"出海的一个实证注脚。
A digital bank can borrow strength for its "heart transplant" - "buy a mature stack plus change the database" is several times to an order of magnitude faster than building from scratch (a single-sample comparison: this project's 10-month go-live versus the industry's typical multi-year cycles; not a generalizable multiplier); Fusion's 10-month go-live versus the industry's typical multi-year cycles owes mostly to stack reuse, not the database itself; publishing cutover durations (10 months, 15 hours, 6-hour window) as public commitments that force migration-toolchain maturity is an effective project-management lever; any "expected" cost figure must be labeled as forecast - the 53% reduction remains a management expectation until 2027 delivers it; a Hong Kong licensed digital bank case shows TDSQL's finance narrative has traveled beyond mainland regulatory contexts and holds under overseas compliance frameworks - an empirical footnote for the "policy-dividend" story going global.
来源
腾讯云开发者社区《10个月换"心"!腾讯云数据库TDSQL助力富融银行核心系统升级》(腾讯云数据库 TencentDB 官方专栏,发表于 2025-02-20,原始发表 2025-02-19
Tencent Cloud Developer Community "Changing the 'Heart' in 10 Months: Tencent Cloud TDSQL Powers Fusion Bank's Core System Upgrade" (TencentDB official column, published 2025-02-20, originally published 2025-02-19
引述富融银行副行政总裁、首席技术官邱家骅、腾讯云副总裁胡利明
quotes Qiu Jiahua, Deputy CEO and CTO of Fusion Bank, and Hu Liming, Vice President of Tencent Cloud
相关产品:腾讯云 TDSQL、MySQL 相关能力:微众银行全国首个分布式银行核心 —— 金融级生产实证 最后核验:2026-10-02
海峡银行:资产 2900 亿中小银行"小步快跑"去 IOE,新核心全面采用微服务 + TDSQL(2019–2023) 成功经验
去 IOE
银行核心系统
微服务改造
信创合规
两地三中心
- 场景
福建海峡银行是资产规模超 2900 亿元的中小银行典型代表。原有系统采用集中式 IOE 架构(IBM 小型机、Oracle 数据库、EMC 存储),在互联网金融浪潮冲击、用户线上化与个性化需求激增下,面临单点故障风险、高昂的更新维护成本(商用数据库与硬件被国外厂商垄断)、业务连续性保障压力,难以满足移动支付与线上金融业务爆发式增长所需的高并发、低延迟与弹性扩展需求。
Fujian Haixia Bank is a representative small-to-mid-size bank with assets exceeding 290 billion yuan. Its legacy systems ran a centralized IOE architecture (IBM mainframes, Oracle Database, EMC storage). Under the impact of internet finance and surging online, personalized customer demand, it faced single-point-of-failure risk, high maintenance costs (commercial databases and hardware monopolized by foreign vendors), and business-continuity pressure - unable to meet the high-concurrency, low-latency, elastic-scaling needs of booming mobile payments and online finance.
- 决策
2019 年启动数据库应用研究,遵循"小步快跑、逐步推进、先试先行、真试真用"策略:2020 年在互联网电子渠道与信贷管理系统完成试点,验证 TDSQL 在日均百万级交易量下的稳定性(秒级响应)、数据一致性(多副本与跨 IDC 强同步)及灾备能力(双中心双活,异地准实时灾备,RPO=0、RTO<120 秒);2021 年开展微服务架构与分布式数据库适配攻关;2022 年新核心系统全面采用微服务 + TDSQL;2023 年进入全面推广阶段。选型时从自主可控度、切换耗时、成本投入、成熟度四个维度综合评估三个方案,放弃"沿用集中式架构"(方案 01)与"双轨并行"(方案 03),最终选择"国产分布式数据库替代集中式数据库"(方案 02),彻底摆脱 IOE 依赖。部署上综合规模体量与实施复杂度,选择微服务模式(非单元化模式),实现应用解耦与独立部署扩展。
Starting database research in 2019, the bank followed a "small steps, fast pace; pilot first, real trials" strategy: in 2020 it piloted TDSQL on internet e-channels and the credit management system, verifying stability at millions of daily transactions (second-level response), data consistency (multi-replica cross-IDC strong sync), and DR capability (dual-center active-active, near-real-time offsite DR with RPO=0 and RTO under 120 seconds); in 2021 it tackled microservices-plus-distributed-database adaptation; in 2022 the new core system went fully microservices plus TDSQL; in 2023 it entered full rollout. Selection weighed four dimensions - self-controllability, cutover time, cost, and maturity - across three options, rejecting "keep the centralized architecture" (Option 01) and "dual-track parallel run" (Option 03) for "replace centralized databases with a domestic distributed database" (Option 02), fully shedding IOE dependence. For deployment it chose the microservices model (non-cellular) given its scale and implementation complexity, achieving application decoupling with independently deployable, scalable services.
- 结果
新核心系统、厅堂系统、企业服务总线等新建系统全面落地微服务应用与 TDSQL,替换 Oracle 数据库及小型机。高可用与容灾:两地三中心部署,同城 RPO=0、RTO<30 秒,异地 RTO<10 分钟,系统可用率 99.999%。性能与扩展:支持亿级账户,每日交易量 >5000 万,TPS≥5000,简单交易响应 <50ms,复杂交易 <200ms,日终批处理 <35 分钟。业务效率:热点账户机制使并发效率提升达数十倍。成本与自主可控:采用国产 X86 服务器与国芯服务器降低软硬件采购成本,实现数据库、应用层、基础架构自主可控。行业认可:2022 年入围工信部创新应用示范案例;新核心项目群获 2022 年度中国人民银行金融科技发展奖二等奖;2023 年度全国金融行业信创考核获"优等";连续四年高质量通过人行信创验收;2024 年发布企业标准《分布式数据库选型规范》(Q/FJHXB 00003—2025)。以上数字来源标注为"海峡银行部署架构指标/新核心系统性能指标"(银行与厂商联合口径),未找到第三方独立实测,引用时须注明口径。
The new core system, lobby systems, and enterprise service bus all went live on microservices plus TDSQL, replacing Oracle databases and mainframes. HA and DR: two-city three-center deployment, in-city RPO=0 with RTO under 30 seconds, offsite RTO under 10 minutes, 99.999% availability. Performance and scale: hundreds of millions of accounts, 50M+ daily transactions, TPS at or above 5,000, simple transactions under 50ms, complex transactions under 200ms, end-of-day batch under 35 minutes. Business efficiency: a hotspot-account mechanism multiplied concurrent efficiency by tens of times. Cost and self-controllability: domestic x86 and domestic-CPU servers cut hardware/software procurement costs, achieving self-controllable database, application, and infrastructure layers. Industry recognition: shortlisted for the MIIT innovation-application demonstration cases in 2022; the new-core program won second prize in the 2022 PBOC FinTech Development Awards; rated "Excellent" in the 2023 national finance-industry compliance evaluation; passed PBOC compliance acceptance at high quality four years running; in 2024 published the enterprise standard "Distributed Database Selection Specification" (Q/FJHXB 00003-2025). All figures are labeled as "Haixia Bank deployment-architecture metrics / new-core performance metrics" (joint bank-vendor figures); no independent third-party measurement was found, so cite with the source noted.
- 机制根因
数据分片减轻单节点负担,支持故障自动转移与弹性伸缩,这是替换集中式架构的核心理由;多副本容灾与自动故障切换满足 RTO<30 秒、RPO=0 的选型硬标准;强一致性分布式协议保障分布式事务;热点账户机制专项解决银行核心的高频账户并发瓶颈(这正是第 47 维"热点数据更新能力"的真实生产注脚——海峡银行的实践说明热点问题在银行核心选型中是独立评估项);微服务模式(非单元化)在应用层解耦,让数据库分片与业务拆分同构演进。
Data sharding relieves single-node burden and supports automatic failover plus elastic scaling - the core argument for replacing the centralized architecture; multi-replica DR and automatic failover meet the hard selection bar of RTO under 30 seconds and RPO=0; a strong-consistency distributed protocol guarantees distributed transactions; the hotspot-account mechanism specifically attacks the high-frequency-account concurrency bottleneck of bank cores (a real production footnote to dimension 47, "hotspot update handling" - Haixia's practice shows hotspot handling is an independently evaluated item in bank-core selection); the microservices (non-cellular) model decouples the application layer so database sharding and business decomposition evolve in step.
- 教训
中小银行去 IOE 的"小步快跑"节奏值得复制——先外围系统试点、再核心攻关、最后全面推广,全程 4 年,没有"一步到位"的赌博;选型评估把"自主可控度、切换耗时、成本投入、成熟度"并列,而不是只看性能跑分,切换耗时和成熟度在银行场景里经常是否决项;试点阶段就定量验证 RPO/RTO,不要等到核心切换才发现容灾指标不达标;把内部选型经验沉淀为企业标准《分布式数据库选型规范》对外发布,既是信创考核"优等"的加分项,也是把一次性项目成本转化为行业影响力的做法。
A small-to-mid-size bank's de-IOE rhythm is worth copying - pilot on peripheral systems first, attack the core second, roll out fully last, four years total, no all-in gamble; selection should weigh "self-controllability, cutover time, cost, maturity" side by side rather than benchmark scores alone - cutover time and maturity are often veto items in banking; verify RPO/RTO quantitatively during the pilot phase, not at core cutover; publishing internal selection experience as the enterprise standard "Distributed Database Selection Specification" turns one-off project cost into industry influence - and bonus points in compliance evaluations.
来源
腾讯云开发者社区《海峡银行分布式数据库转型:以TDSQL实现核心系统自主可控与性能跃升》(作者授权原创,发表于 2026-04-17
Tencent Cloud Developer Community "Haixia Bank's Distributed Database Transformation: Autonomous Control and Performance Leap for the Core System with TDSQL" (authorized original, published 2026-04-17
引述海峡银行数据库架构师朱正珊
quotes Zhu Zhengshan, database architect at Haixia Bank
相关产品:腾讯云 TDSQL、Oracle Database(甲骨文) 相关能力:"去 O"政策 + 金融合规 —— 政策红利型招牌、微众银行全国首个分布式银行核心 —— 金融级生产实证 最后核验:2026-10-02
昆山农商行:同类银行首个"微服务 + 国产分布式数据库"核心,一套 TDSQL 集群跑三个微服务集群(2021) 成功经验
银行核心系统
微服务架构
国产化替代
区域银行标杆
- 场景
昆山农商银行扎根全国百强县之首的昆山,是当地营业网点最多、服务覆盖面最广的银行。在数据库国产化的大背景下,"微服务如何跑在国产分布式数据库上、破除原先的集群模式"一直是技术难点——数据库的开发应用牵涉多项服务,要满足微服务架构、做到多个服务间数据一致性并非易事,跨服务数据查询也充满挑战。在此之前,TDSQL 已落地张家港农商银行,完成银行传统核心数据库首次国产化,但"微服务 + 国产分布式数据库"的组合在同类银行中尚无先例。
Kunshan Rural Commercial Bank is rooted in Kunshan, China's top-ranked county-level city, with the most local outlets and widest service coverage. Against the backdrop of database domestic substitution, "how microservices run on a domestic distributed database and break the old cluster model" had long been a technical hard problem - database development touches many services, guaranteeing data consistency across microservices is nontrivial, and cross-service queries are full of challenges. TDSQL had previously landed at Zhangjiagang Rural Commercial Bank, completing the first domestic substitution of a traditional bank-core database, but the "microservices plus domestic distributed database" combination had no precedent among peer banks.
- 决策
历经 300 多个日夜,2021 年 8 月基于腾讯云 TDSQL 打造的新一代核心系统成功投产上线,采用"微服务应用 + 国产分布式数据库"架构(同类银行中尚属首次)。新核心采用长亮 V8 技术,无缝衔接国产分布式数据库 TDSQL,并融入微服务、读写分离、多源同步等技术,在保证金融级数据全局一致性的基础上,把大系统拆分成小型微服务。架构上设三个微服务集群(公共服务微服务集群、账务微服务集群、历史微服务集群),每个集群由功能职责单一、高度聚合的服务组成,可灵活部署——三个微服务集群运行在同一套 TDSQL 集群中。部署上采用两地三中心,数据库一主三备,中心间数据强同步,实现中心级别灾难快速自动恢复且数据零丢失。
After 300+ days and nights, the new-generation core system built on Tencent Cloud TDSQL went live in August 2021, adopting a "microservices application plus domestic distributed database" architecture - a first among peer banks. The new core uses Sunline V8 technology seamlessly connected to TDSQL, integrating microservices, read-write splitting, and multi-source synchronization to split the monolith into small microservices while guaranteeing financial-grade global data consistency. Three microservice clusters (public-services, ledger, and history clusters), each composed of single-responsibility, highly cohesive services for flexible deployment - all three run on a single TDSQL cluster. Deployment is two-city three-center with one-primary three-standby databases and strong cross-center data sync, enabling fast automatic recovery from center-level disasters with zero data loss.
- 结果
新核心系统整体处理能力 6300TPS,可支持每日亿级交易量;高频账户类交易平均响应时间 300 毫秒之内,查询类交易平均响应 100 毫秒之内;日终批量时间缩短至 8 分钟左右,季度结息 17 分钟左右;96 秒完成 10 万笔社保代发;性能远超原核心系统,在全国同类型银行中处于领先地位。以上数字全部来自腾讯云厂商口径,未找到第三方独立复现。昆山农商行的创新转型,被腾讯云定位为"微服务 + 国产分布式数据库"架构在银行业应用的标杆。
The new core handles 6,300 TPS overall with hundreds of millions of daily transactions; high-frequency account transactions average under 300ms, query transactions under 100ms; end-of-day batch shrank to about 8 minutes, quarterly interest settlement to about 17 minutes; 100,000 social-security disbursement transactions completed in 96 seconds; performance far exceeds the legacy core and leads peer banks nationwide. All figures are Tencent vendor figures; no independent third-party reproduction was found. Tencent Cloud positions Kunshan Rural Commercial Bank's transformation as the banking benchmark for the "microservices plus domestic distributed database" architecture.
- 机制根因
微服务的横向扩展能力与场景化数据切分,契合金融科技创新对敏捷性的需求;TDSQL 的分布式能力解决了传统集中式核心的并发量瓶颈;多服务间的数据一致性难题由 TDSQL 的分布式事务与全局一致性机制兜底,这是"微服务 + 国产库"组合真正的技术门槛;一主三备 + 中心间强同步把 RPO 压到 0,中心级灾难可快速自动恢复;三个微服务集群共享一套 TDSQL 集群,说明分布式数据库的多租户承载能力足以支撑整行核心业务,而不需要"一个业务一套库"。
Microservices' horizontal scaling and scenario-based data partitioning match fintech innovation's agility needs; TDSQL's distributed capability breaks the concurrency bottleneck of the traditional centralized core; the cross-service data-consistency problem is backstopped by TDSQL's distributed-transaction and global-consistency mechanisms - the real technical bar for the "microservices plus domestic database" combination; one-primary three-standby plus cross-center strong sync pushes RPO to zero with fast automatic center-level disaster recovery; three microservice clusters sharing one TDSQL cluster shows a distributed database's multi-tenant capacity can carry an entire bank's core, no "one database per business" needed.
- 教训
2021 年对同类银行是"首例"——架构创新的窗口期红利真实存在,第一个跑通的银行拿到了标杆话语权;长亮 V8 这类银行应用厂商的适配是国产化落地的关键拼图,数据库厂商 + ISV 的联合交付模式比数据库单打独斗更重要;把公共服务、账务、历史三个集群跑在"一套 TDSQL 集群"上,验证了用一套分布式集群承载整行核心的可行性,这是后续银行选型时"合库还是分库"决策的直接参照;区域银行(农商行)反而比大行更早吃到"微服务 + 国产库"的红利——船小好调头,在核心系统换代窗口期敢下注,是中小银行的相对优势。
2021 was a "first" for peer banks - the window-period dividend of architecture innovation is real, and the first bank through the door captured the benchmark narrative; adaptation by banking application vendors like Sunline V8 is a key puzzle piece for domestic-substitution landing - the database-vendor-plus-ISV joint delivery model beats the database going it alone; running public-services, ledger, and history clusters on "one TDSQL cluster" validates carrying a whole bank's core on a single distributed cluster - a direct reference for the "consolidate or split databases" decision in later bank selections; regional banks (rural commercial banks) ate the "microservices plus domestic database" dividend earlier than large banks - smaller ships turn faster, and daring to bet during a core-system replacement window is the relative advantage of small-to-mid-size banks.
来源
腾讯云开发者社区《首例"微服务+国产分布式数据库"架构,TDSQL助力昆山农商行换"心"》(腾讯云数据库 TencentDB 官方专栏,发表于 2021-10-11,原始发表 2021-10-09
Tencent Cloud Developer Community "First 'Microservices plus Domestic Distributed Database' Architecture: TDSQL Helps Kunshan Rural Commercial Bank Change Its 'Heart'" (TencentDB official column, published 2021-10-11, originally published 2021-10-09
相关产品:腾讯云 TDSQL 相关能力:"去 O"政策 + 金融合规 —— 政策红利型招牌 最后核验:2026-10-02
微众银行:全国首个分布式银行核心,应用层分片 + TDSQL noshard 跑出银行账务(2014–) 成功经验
银行核心账务
分布式架构
应用层分片
同城多活
去 IOE
- 场景
2014 年微众银行成立之时,前瞻性地确立了分布式 IT 架构方向:摒弃传统银行依赖商业数据库、商业存储、大中型服务器的集中式架构,走互联网模式的分布式架构。当时除了 Oracle 等少数传统商业数据库,能满足金融级银行场景的数据库产品并不多。腾讯内部有一款金融级分布式数据库 TDSQL,主要承载腾讯内部计费和支付业务,其业务场景与可靠性要求和银行场景非常类似,且经受了腾讯海量计费业务的验证。微众银行基础架构团队经过多轮评估测试,决定与腾讯 TDSQL 团队合作,共同把 TDSQL 打造成适合银行核心场景的金融级分布式数据库。
When WeBank was founded in 2014, it set a forward-looking distributed IT architecture direction: abandoning the traditional centralized model of commercial databases, commercial storage, and mainframes for an internet-style distributed architecture. At the time, few database products besides Oracle-class commercial offerings could meet financial-grade banking demands. Tencent had an in-house financial-grade distributed database, TDSQL, carrying Tencent's internal billing and payment services with reliability requirements very similar to banking scenarios, battle-tested by Tencent's massive billing workloads. After multiple rounds of evaluation and testing, WeBank's infrastructure team partnered with the Tencent TDSQL team to jointly shape TDSQL into a financial-grade distributed database fit for bank-core scenarios.
- 决策
微众银行把 TDSQL 用于核心系统数据库,但在架构路线上做了一个关键抉择:是采用 TDSQL 的 shard 模式(数据库做 AutoSharding),还是在应用层做分布式、数据库用 TDSQL noshard 模式?经过大量调研分析,微众认为应用做分布式是最可控、最安全、最灵活、最可扩展的模式,于是设计了基于 DCN(Data Center Node,数据中心节点)的分布式可扩展架构——一个 DCN 是包含完整应用层、接入层和数据库的自包含逻辑单元,可通俗理解为"线上的虚拟分行",按账户分段把客户划分到不同 DCN;GNS(Global Name Service)保留全局路由信息(用 Redis 缓存 + TDSQL 持久化);RMB(Reliable Message Bus)负责业务系统间可靠消息通信。数据库层采用 TDSQL noshard 单实例模式,保证数据库架构的简洁性与业务层 MySQL 兼容性,把分片复杂度留在应用层。
WeBank put TDSQL into its core banking systems, but made a pivotal architecture choice: TDSQL's shard mode (AutoSharding inside the database) versus application-layer distribution with TDSQL in noshard mode. After extensive research, WeBank judged application-layer distribution the most controllable, safest, most flexible, and most scalable option, and designed a DCN (Data Center Node) based distributed scaling architecture - a DCN is a self-contained logical unit with a complete application layer, access layer, and database, roughly a "virtual branch" online, with customers assigned to DCNs by account segmentation; GNS (Global Name Service) keeps global routing information (Redis cache plus TDSQL persistence); RMB (Reliable Message Bus) handles reliable inter-system messaging. The database layer uses TDSQL noshard single-instance mode to keep the database architecture simple and fully MySQL-compatible, leaving sharding complexity in the application layer.
- 结果
微众银行建成两地六中心架构(深圳 5 个 IDC 为生产中心,上海 1 个 IDC 为跨城异地容灾;深圳同城 IDC 间距离控制在 10–50 公里,ping 延迟 2ms 左右);数据库采用"同城 3 副本 + 跨城 2 副本"的 3+2 noshard 部署,同城 RPO=0、RTO 秒级恢复(主备切换 30 秒内完成),跨城 2 副本经同城 slave 异步复制做容灾。规模上:TDSQL SET 个数 350+(生产+容灾),数据库实例 1700+,整体数据规模 PB 级,承载数百个核心系统;业务高峰日金融交易量 3.6 亿+,最高 TPS 10 万+;有效客户数过亿级。4 年多运营中 TDSQL 未出现大的系统故障或数据安全问题,TDSQL 承载了微众银行 99% 以上线上数据库业务。以上数字全部来自腾讯云/微众银行联合口径厂商口径,未找到第三方独立复现。顺带诚实记录:微众后来也引入了 TiDB,解决"无法拆分 DCN、但单库需要超大容量或超大吞吐"的非联机场景——TDSQL noshard 单库容量上限是这套架构的已知边界。
WeBank built a two-city six-center architecture (five Shenzhen IDCs as production, one Shanghai IDC for cross-city DR; intra-city Shenzhen IDC distances kept at 10-50 km with ping latency around 2ms); the database runs a 3+2 noshard deployment (3 replicas in-city plus 2 cross-city), with in-city RPO=0 and second-level RTO recovery (primary-standby failover completes within 30 seconds), the 2 cross-city replicas asynchronously replicating via an in-city slave for DR. Scale: 350+ TDSQL SETs (production plus DR), 1,700+ database instances, PB-scale data, carrying hundreds of core systems; peak days saw 360M+ financial transactions with TPS above 100,000; effective customers exceeded 100 million. Over 4+ years of operation TDSQL saw no major system failures or data-safety incidents, and TDSQL carries 99%+ of WeBank's online database business. All figures are joint Tencent/WeBank figures (vendor figures); no independent third-party reproduction was found. For honesty: WeBank later also adopted TiDB for non-online scenarios that cannot be split by DCN yet need very large single-database capacity or throughput - the single-database capacity ceiling of TDSQL noshard is a known boundary of this architecture.
- 机制根因
TDSQL 在 MySQL/MariaDB 内核复制模块做了系统级优化,实现多副本强一致同步(RPO=0,数据 0 丢失);Agent 上报监控信息到 ZooKeeper,Scheduler 按集群状态启动调度任务,实现 30 秒内的自动化主备强一致切换;针对同城跨 IDC 强同步场景做了内核级优化(队列异步化、并发复制),基准测试显示跨 IDC 强同步对联机 OLTP 的性能影响仅在 10% 左右;批量场景则用 WATCH 节点机制规避跨 IDC 延迟放大——在原来一主两备之外,额外部署一个与主节点同 IDC 的 WATCH 节点(异步同步、不参与选举),批量 APP 与主节点同 IDC 部署,避免跨 IDC 访问的时延累积,同时主节点 binlog 仍需同步到跨 IDC 的两个备节点才算事务成功,容灾特性不受影响。架构哲学的差异值得标注:与 OceanBase/TiDB"数据库内部做分片"不同,微众选择了"应用层 DCN 分片 + 数据库 noshard",用应用层的路由规则换取数据库层的简洁与 MySQL 完全兼容。
TDSQL's kernel-level replication-module optimization on MySQL/MariaDB delivers multi-replica strong-consistency sync (RPO=0, zero data loss); Agents report monitoring data to ZooKeeper and the Scheduler launches scheduling tasks by cluster state, automating primary-standby strong-consistency failover within 30 seconds; for cross-IDC in-city strong sync, kernel-level optimizations (queue asyncing, concurrent replication) keep the OLTP impact around 10% per benchmarks; batch scenarios use a WATCH-node mechanism to avoid cross-IDC latency amplification - an extra WATCH node (async, non-voting) is deployed in the same IDC as the primary so batch apps colocate with the primary, while the primary's binlog must still reach the two cross-IDC standbys before a transaction commits, preserving DR guarantees. The architectural philosophy is worth noting: unlike OceanBase/TiDB-style "sharding inside the database", WeBank chose "application-layer DCN sharding plus noshard database", trading application routing rules for database-layer simplicity and full MySQL compatibility.
- 教训
"用分布式数据库"不等于"用数据库的分片"——微众证明了应用层分片 + noshard 同样能跑出银行核心,选型时应把"分片放在哪一层"作为独立决策项;同城多活的真正门槛是网络(10–50 公里、2ms、IDC 间多条专线),不是数据库功能清单,网络不达标时强同步的代价会吃掉所有理论收益;联机与批量要分开部署拓扑(WATCH 节点),批量 APP 的跨 IDC 访问延迟会被"累积放大",这是容易被忽视的隐性成本;2014 年就敢把核心放在分布式数据库上,"第一个吃螃蟹"的信任状本身成了 TDSQL 后续 4000+ 客户拓展中最硬的招牌——先行者红利在金融业是真实存在的销售杠杆。
"Using a distributed database" is not the same as "using the database's sharding" - WeBank proved application-layer sharding plus noshard can run a bank core, so "which layer owns sharding" should be an explicit decision item in any selection; the real bar for in-city active-active is the network (10-50 km, 2ms, multiple dedicated lines between IDCs), not the database feature list - without the network, strong sync costs eat all theoretical gains; online and batch topologies should be separated (WATCH nodes), because cross-IDC access latency for batch apps gets cumulatively amplified - an easily missed hidden cost; daring to put the core on a distributed database in 2014 gave TDSQL the hardest trust credential for its later 4,000+ customer expansion - first-mover trust is a real sales lever in finance.
来源
腾讯云开发者社区《金融级分布式数据库打造!TDSQL 在微众银行的大规模实践》(2023
Tencent Cloud Developer Community "Building a Financial-Grade Distributed Database: TDSQL's Large-Scale Practice at WeBank" (2023
作者胡盼盼(微众银行数据库平台负责人)、黄德志(微众银行数据库平台高级 DBA)
authors Hu Panpan (head of WeBank database platform) and Huang Dezhi (senior DBA, WeBank database platform)
相关产品:腾讯云 TDSQL、Oracle Database(甲骨文)、TiDB 相关能力:微众银行全国首个分布式银行核心 —— 金融级生产实证 最后核验:2026-10-02
微信支付:TDSQL PG 版承载数据密集型应用,百亿级模糊检索从 17 秒降到 50 毫秒(2021) 成功经验
数据密集型应用
报表系统
维表系统
数据倾斜治理
PostgreSQL 兼容
- 场景
微信支付的商户服务平台为千万级商家提供账单明细下载、复杂条件查询与统计分析。平台最初以开源 MySQL 作为底层存储,随着京东等大商户接入、交易笔数持续提升,单机存储容量受限,微信支付遇到严重的容量瓶颈与性能瓶颈。在当时的技术背景下,团队需要一个能同时解决"存不下"和"查不动"的方案,于是选择了 TDSQL PG 版(开源代号 TBase)。
WeChat Pay's merchant service platform serves tens of millions of merchants with bill-detail downloads, complex conditional queries, and statistical analysis. It originally ran on open-source MySQL. As large merchants such as JD.com joined and transaction volume kept growing, single-node storage capacity hit its limit and WeChat Pay ran into severe capacity and performance bottlenecks. The team needed a solution for both "cannot store it" and "cannot query it", and chose TDSQL for PostgreSQL (open-source codename TBase).
- 决策
微信支付把商户服务平台的底层存储从开源 MySQL 迁移到 TDSQL PG 版;2021 年又进一步基于 TDSQL PG 版搭建数据仓库的维表管理系统(管理 2700+ 枚举值),让它成为大数据生态中的重要组件。选型逻辑是:商户平台的痛点不是单纯的 OLTP 写入,而是"海量存储 + 复杂查询 + 实时报表"的混合负载,PG 系的索引类型与并行能力比 MySQL 系更对味。
WeChat Pay migrated the merchant platform's underlying storage from open-source MySQL to TDSQL for PostgreSQL; in 2021 it went further and built its data-warehouse dimension-table management system on TDSQL for PostgreSQL (managing 2,700+ enum values), making it a key component of the big-data ecosystem. The selection logic: the merchant platform's pain was not plain OLTP writes but a mixed workload of massive storage plus complex queries plus real-time reporting, where the PG family's index variety and parallel capabilities fit better than the MySQL family.
- 结果
报表系统累计承载微信支付 3600+ 报表的数据写入、存储与读取,报表打开时间稳定控制在 3 秒以内;百亿级数据中模糊检索商户名称的场景,查询耗时从接近 17 秒降到 50 毫秒以内;维表系统打通后,OLTP 录入的枚举值变更能被 Spark 计算、报表系统实时引用。整体运营规模上,微信支付在 TDSQL PG 版的存储量达到 400TB+,每秒请求量超过 24 万次,99.6% 的请求耗时控制在 10 毫秒以内。以上数字全部来自腾讯云数据库官方博客厂商口径,未找到第三方独立复现。
The reporting system now carries 3,600+ WeChat Pay reports for writes, storage, and reads, with report open times stably under 3 seconds; fuzzy merchant-name search over tens of billions of rows dropped from nearly 17 seconds to under 50 milliseconds; after the dimension-table system went live, enum changes entered in the OLTP dimension-table manager are referenced in real time by Spark jobs and reporting systems. At overall operating scale, WeChat Pay stores 400TB+ on TDSQL for PostgreSQL, serving 240,000+ requests per second with 99.6% of requests under 10ms. All figures come from the official Tencent Cloud database blog (vendor figures); no independent third-party reproduction was found.
- 机制根因
容量问题靠 TDSQL 的海量数据在线线性扩容解决;大商户的数据倾斜问题靠双 KEY 分布机制让数据均匀分布到多个分片——这是微信支付商户平台实践沉淀下来的机制,后来成为 TDSQL 分片文档里的标准解法;分页查询性能问题靠基于 Index only scan 的索引优化,解决了传统 Web 应用分页场景"查总条数耗时高"的顽疾;报表系统的双写入模型(Spark 离线周期性写入 + 消息队列实时写入,单次写入可达十亿/百亿级)对底层并行写入能力要求极高,TDSQL PG 版的并行写入性能明显优于开源 MySQL,大幅降低了数据入库完成时间。维表系统的本质是 OLTP 与 OLAP 能力的融合:在 OLTP 维表管理系统中录入或更新枚举值后,在线业务、Spark 计算、报表系统都能实时引用同一份数据,避免了"上游改了枚举、下游无感知"的质量风险。
Capacity was solved by TDSQL's online linear scaling for massive data; large-merchant data skew was solved by a dual-key distribution mechanism that spreads data evenly across shards - a mechanism distilled from this very WeChat Pay merchant-platform practice that later became a standard recipe in TDSQL sharding documentation; pagination performance was fixed with Index-only-scan-based index optimization, curing the classic web-app pain of expensive total-count queries; the reporting system's dual-write model (Spark offline periodic writes plus message-queue real-time writes, with single writes reaching billions or tens of billions of rows) demands extreme parallel write throughput, where TDSQL for PostgreSQL's parallel writes clearly beat open-source MySQL and sharply cut data-ingestion completion time. The dimension-table system is essentially OLTP/OLAP fusion: enum values entered or updated in the OLTP dimension-table manager are immediately referenceable by online services, Spark compute, and reporting - eliminating the quality risk of "upstream changed the enum, downstream never noticed".
- 教训
17 秒到 50 毫秒的差距主要来自索引类型与并行写入能力,不是靠加机器堆出来的——"查不动"的问题要先看执行机制再谈扩容;"大客户数据倾斜"是分布式选型的典型信号,商户/租户平台类业务选型时应把倾斜治理机制列为必查项;报表 + 维表这类"数据密集型"场景是 PG 系分布式数据库的甜点区,MySQL 系方案在复杂查询和索引丰富度上天然吃亏;腾讯高级工程师万志颖把微信支付与 TDSQL PG 版的关系形容为"你侬我侬"——超大规模业务是最好的产品试炼场,但这也意味着案例里的优化(如双 KEY 分布)高度贴合微信支付的业务形状,照搬前要先确认自己的倾斜模式是否同构。
The 17s-to-50ms gap came mostly from index types and parallel write capability, not from adding machines - when queries are slow, examine the execution mechanism before scaling out; "large-customer data skew" is a classic signal for distributed-database selection, and merchant/tenant-platform workloads should list skew-handling mechanisms as a mandatory checklist item; reporting plus dimension tables - the "data-intensive" pattern - is the sweet spot for PG-family distributed databases, where MySQL-family options are naturally disadvantaged on complex queries and index richness; Tencent senior engineer Wan Zhiying described the WeChat Pay / TDSQL-for-PostgreSQL relationship as inseparable - hyperscale business is the best proving ground for a product, but it also means the optimizations (like dual-key distribution) are tightly shaped to WeChat Pay's workload, so adopters should first confirm their own skew patterns are isomorphic before copying.
来源
腾讯云数据库官方博客园《TDSQL 在微信支付数据密集型应用落地实践》(发表于 2021-09-02
Official Tencent Cloud database blog (cnblogs) "TDSQL in WeChat Pay Data-Intensive Application Practice" (published 2021-09-02
案例由腾讯高级工程师万志颖介绍
case presented by Tencent senior engineer Wan Zhiying
相关产品:腾讯云 TDSQL、MySQL 相关能力:— 最后核验:2026-10-02
BookMyShow:Galera 脑裂后换 TiDB,6 人团队卸下专职 DBA(2019 前后) 成功经验
Galera 脑裂
微服务数据管道
运维减负
评估选型
- 场景
BookMyShow(印度最大在线票务,2007 年成立,5000 万+用户、650+ 城市,每周数百万交易)走微服务架构,需要一条统一的数据管道把各服务数据送进大数据平台,担子落在 6 人数据运维团队肩上。当时用 Galera:主主复制、每节点全量数据,数据量上 TB 后扩缩容与资源利用率双双恶化,还要搭进去一个工程师专职看库。一次严重脑裂——两个主库完全失步,全队花 2–3 天才恢复——成为换架构的导火索。(公司与业务数字为 PingCAP 刊载的 BookMyShow 团队自述。)
BookMyShow (India's largest online entertainment ticketing site, founded 2007, 50M+ users across 650+ cities and towns, millions of transactions a week) runs a microservices architecture and needed a unified data pipeline feeding its Big Data platform — a load carried by a 6-person Data Operations team. It was on Galera: primary-primary replication with every node holding a full copy of the data, so once data reached multiple terabytes, scaling and resource utilization both deteriorated, and one engineer was fully dedicated to babysitting the database. A severe split-brain — the two primaries completely out of sync, taking the whole team 2–3 days to recover — was the trigger to change architecture. (Company and business figures are the BookMyShow team's own account as published by PingCAP.)
- 决策
候选 TiDB、Greenplum、Vitess,基于论文、博客、案例研究一轮筛选后重点评估 TiDB 与 Greenplum。TiDB 胜出理由很朴素:部署快、灌数据快——"负责评估 Greenplum 的同事还没搞定部署,TiDB 这边已经灌上数据了"。历史数据用 Spark 流式导入;最终架构为 MS SQL → Kafka → TiDB,交易数据近实时同步。
The shortlist was TiDB, Greenplum, and Vitess; after a research round over papers, blogs, and case studies, the team evaluated TiDB and Greenplum in depth. TiDB won for a plain reason: fast to deploy, fast to load data — "before my teammate evaluating Greenplum could figure out deployment, the TiDB evaluator was done and already pushing data." Historical data was streamed in with Apache Spark; the final architecture streams transactional data near real-time from MS SQL to TiDB via Kafka.
- 结果
TiDB 中约 5TB 数据、预计涨到 10TB;可用性提升;运维成本降 30%(BookMyShow 工程师原话,经 PingCAP 刊载,未找到 BookMyShow 官方独立披露);不再需要专职 DBA,只有慢查询和扩容时才需要管库。踩过的坑也如实记录:曾自定义调大 Region 大小(默认 96MB)与 gRPC 连接数导致性能变差,找 PingCAP 才解决;Grafana 指标太多却无指引,"信息过载、无从下手"。
About 5TB in TiDB, expected to reach 10TB; uptime improved; operational and maintenance cost reduced by 30% (the BookMyShow engineer's words, as published by PingCAP; no independent disclosure by BookMyShow found). No engineer needs to be fully dedicated to database operations anymore — the team only touches TiDB for slow queries and scale-outs. The stumbles were recorded honestly too: a custom increase of Region size (default 96MB) and gRPC connections caused bad performance until PingCAP helped resolve it, and the Grafana dashboards had "too much information with no guidance" — metric overload with no troubleshooting map.
- 机制根因
Galera 的病根是"主主 + 全量复制":每加一个节点就多一份全量数据,TB 级后扩容模型失效;一次脑裂的恢复成本是全队 2–3 天。TiDB 的自动分片 + Raft 把"数据怎么切、副本放哪"自动化,RocksDB 底层又给了单机性能底子,小团队才敢把库交出去。Spark/Kafka 的导入与同步链路说明:换库不只是换引擎,还要重搭数据管道,管道成本要计入迁移预算。
Galera's disease is "primary-primary + full replication": every added node adds another full copy of the data, so the scaling model breaks at terabyte scale, and one split-brain costs the whole team 2–3 days. TiDB's auto-sharding + Raft automate "how data is split and where replicas live," and RocksDB underneath provides solid single-node performance — which is why a small team dared to hand over the database. The Spark/Kafka ingestion and sync pipeline is a reminder: switching databases is not just switching engines; the pipeline rebuild cost belongs in the migration budget.
- 教训
把"脑裂恢复的人力成本"计入 TCO,主主架构的账不能只算硬件;小团队选型应把"是否需要专职 DBA"作为硬指标——BookMyShow 用 6 人团队证明了"无人值守"是可验证的选型收益;默认参数先跑起来,Region 大小这类"优化"反而可能踩坑,调优前先读文档、调优后留基线对比。
Count "split-brain recovery labor" in TCO — primary-primary math can't stop at hardware. Small teams should make "do we need a dedicated DBA" a hard selection criterion — BookMyShow's 6-person team proved "no babysitting" is a verifiable selection outcome. Start on default parameters; "optimizations" like Region sizing can backfire — read the docs before tuning and keep a baseline for comparison after.
相关产品:TiDB、Microsoft SQL Server、MySQL 相关能力:— 最后核验:2026-10-02
Jepsen 测试 TiDB 2.1.7:默认开启的事务自动重试让快照隔离"默认即违反",3.0 关默认后通过(2019) 失败教训
选型评估
独立测试
快照隔离
自动重试
版本选型
- 场景
TiDB 宣称提供快照隔离(文档里称之为 repeatable read)。2019 年 6 月 Jepsen 测试 TiDB 2.1.7 至 3.0.0-rc.2(5 节点,PD+TiKV+TiDB 同机,Region 复制因子 3)。**注:本次为 PingCAP 付费委托**(报告明确声明 "This work was funded by PingCAP"),结论引用时须标注资助关系。
TiDB claimed snapshot isolation (documented as "repeatable read"). In June 2019, Jepsen tested TiDB 2.1.7 through 3.0.0-rc.2 (5 nodes, PD+TiKV+TiDB colocated, Region replication factor 3). **Note: this analysis was funded by PingCAP** (the report states "This work was funded by PingCAP, the makers of TiDB") — cite with the funding relationship disclosed.
- 决策
bank(转账总额守恒)/long-fork/register/append 等事务工作负载;故障注入包括 SIGSTOP/SIGKILL、iptables 网络分区、最大 232 秒的指数分布时钟偏移、leader/region 洗牌与合并。
Transactional workloads including bank (conservation of total balances), long-fork, register, and append; fault injection with SIGSTOP/SIGKILL, iptables network partitions, exponentially distributed clock skew up to 232 seconds, and leader/region shuffling and merging.
- 结果
2.1.7 起默认开启的两套 auto-retry 机制在事务冲突时盲目重放写,导致**默认配置即违反快照隔离**(read skew、lost update);关闭重试后,2.1.8+ 通过快照隔离与单键线性一致性测试;3.0.0-rc.2 起默认关闭 auto-retry 并通过测试。另发现建表竞态(3.0.0-rc.2 修复)、新集群 durability 不足、启动崩溃等问题(部分 PingCAP 标为 by design 不修)。PingCAP 发布 companion blog 回应。
Two auto-retry mechanisms, enabled by default since 2.1.7, blindly replayed writes on transaction conflicts, causing snapshot isolation to be **violated under default configuration** (read skew, lost updates); with retries disabled, 2.1.8+ passed snapshot isolation and single-key linearizability tests; 3.0.0-rc.2 disables auto-retry by default and passes. Also found: a table-creation race (fixed in 3.0.0-rc.2), reduced durability in fresh clusters, and startup crashes (some marked by-design by PingCAP and left unfixed). PingCAP published a companion blog post in response.
- 机制根因
auto-retry 的设计初衷是"让应用少写重试代码",但把"重试"做进数据库内核且默认开启,就把应用层的幂等假设偷换成了"数据库替你重放写"——而重放的写基于过期快照,隔离性自然破。这是一个"便利性默认"吃掉"正确性"的典型:默认值是最重要的 API,错默认比缺功能更危险。
Auto-retry was designed so "applications write less retry code," but baking retry into the database kernel — and enabling it by default — silently swapped the application layer's idempotency assumption for "the database replays your writes," and replayed writes are based on stale snapshots, so isolation breaks. A textbook case of a "convenience default" eating correctness: the default is the most important API, and a wrong default is more dangerous than a missing feature.
- 教训
版本选型与配置选型同等重要——同一产品的"默认配置"和"调对配置"是两个产品;永远不要重新打开事务自动重试;厂商付费委托的测试报告仍可引用,但必须标注资助关系。诚实注记:这是 2019 年 2.x/3.0-rc 时代的结论,3.0.0-rc.2 已更改默认;TiDB 此后演进了 7 年,不可直接套用到当前版本,选型时以新版测试为准。
Version selection and configuration selection matter equally — the "default configuration" and the "correctly configured" version of the same product are two different products; never re-enable transaction auto-retry; vendor-funded test reports remain citable but the funding relationship must be disclosed. Honesty note: these conclusions date from the 2.x/3.0-rc era of 2019; 3.0.0-rc.2 changed the default and TiDB has evolved for 7 years since — do not apply directly to current releases; evaluate against newer tests during selection.
来源
Jepsen《Jepsen: TiDB 2.1.7》(2019-06-12,PingCAP 付费委托
Jepsen, "Jepsen: TiDB 2.1.7" (2019-06-12, funded by PingCAP
相关产品:TiDB 相关能力:Percolator 事务模型的快照隔离与自动重试默认 最后核验:2026-10-02
Ninja Van:评估 Vitess/CockroachDB/TiDB 后选 TiDB,OLAP 真实查询最高快 111.92 倍(2021) 成功经验
PoC
选型评估
MySQL 兼容
K8s
OLAP 加速
- 场景
Ninja Van(东南亚物流,日均 150 万+包裹,6 国运营)100+ 微服务跑在 MySQL(Galera)+ProxySQL 上:ProxySQL 写扩展性差、分库分表要改应用且不可逆(跨分片 JOIN 下沉应用层)、Galera 写不可扩展且流控会阻塞写。2020 年 7 月起团队调研 Vitess、CockroachDB、TiDB。
Ninja Van (Southeast Asian logistics, 1.5M+ parcels/day, 6 countries) ran 100+ microservices on MySQL (Galera) + ProxySQL: ProxySQL wasn't write-scalable, sharding required application changes and was irreversible (cross-shard JOINs pushed to the app layer), and Galera wasn't write-scalable with flow control blocking writes. Starting July 2020, the team investigated Vitess, CockroachDB, and TiDB.
- 决策
先列 7 条硬需求:MySQL 兼容、水平扩展、高可用、易运维(含在线 DDL)、CDC(核心需求:灌数据湖、建 ES 索引、更新缓存)、可观测、云原生(K8s)。Vitess 落选:破坏性 schema 变更(清外键)、跨分片默认 READ COMMITTED、长串不支持的查询、跨分片 2PC 官方不推荐开、2020 年时 Vitess Operator 不稳定。CockroachDB 落选:PG 协议意味着迁移成本高、工具链弱(TiDB 有 DM/Lightning/TiCDC 生态)。随后在 GCP 自建 ~6TB 测试集群(PD 3×n1-standard-4、TiKV 3×n1-standard-16 挂 5 块 local SSD、TiDB 2×n1-standard-16),Sysbench 测 OLTP + 5 个真实业务查询测 OLAP。
Seven hard requirements came first: MySQL compatibility, horizontal scalability, high availability, operational ease (including online DDL), CDC (a core need: feeding the data lake, building Elasticsearch indexes, updating caches), observability, and cloud-native/K8s. Vitess was rejected: disruptive schema changes (foreign-key cleanup), cross-shard queries defaulting to READ COMMITTED, a long list of unsupported queries, cross-shard 2PC officially discouraged, and an unstable Vitess Operator in 2020. CockroachDB was rejected: the PostgreSQL wire protocol meant high migration cost, and weaker tooling (TiDB's DM/Lightning/TiCDC ecosystem). They then built a ~6TB test cluster on GCP (PD 3×n1-standard-4, TiKV 3×n1-standard-16 with 5 local SSDs, TiDB 2×n1-standard-16), running Sysbench for OLTP plus 5 real business queries for OLAP.
- 结果
选定 TiDB on K8s。按 Ninja Van 自述口径:OLTP 场景 Sysbench 表现满意;OLAP 5 个真实查询相对 MySQL 最高快 111.92 倍(223.85s→2s),其余 60.37x/13.79x/12.96x/6.22x。文章由联合创始人兼 CTO Shaun Chong 与 Sr. SWE Mengnan Gong 署名撰写,记录的是客户自己的真实评估。**渠道注记:发表于 PingCAP 官网(厂商渠道),引用时须标注。**
TiDB on K8s was selected. Per Ninja Van's own account: Sysbench OLTP results were satisfactory; the 5 real OLAP queries ran up to 111.92x faster than MySQL (223.85s→2s), with the others at 60.37x/13.79x/12.96x/6.22x. The article is bylined by co-founder & CTO Shaun Chong and Sr. SWE Mengnan Gong, documenting the customer's own real evaluation. **Channel note: published on pingcap.com (vendor channel) — disclose when citing.**
- 机制根因
Ninja Van 的决策函数里"MySQL 兼容"的权重压倒一切——100+ 微服务的 SQL 方言、ORM、工具链是沉没成本,换 PG 协议等于重写数据访问层。Vitess 输在"分片中间件"的原罪:它把分布式复杂性推回给应用(schema 改造、跨分片语义降级),而 TiDB 把复杂性吃进内核(自动分片、快照隔离)。CDC 是隐藏的否决项:TiCDC 的开源协议与多 sink 支持对数据湖/ES/缓存三件套是刚需。
In Ninja Van's decision function, "MySQL compatibility" outweighed everything — the SQL dialects, ORMs, and tooling across 100+ microservices were sunk costs, and switching to the PG protocol meant rewriting the data-access layer. Vitess lost on the original sin of sharding middleware: it pushes distributed complexity back onto the application (schema surgery, degraded cross-shard semantics), while TiDB absorbs it into the kernel (auto-sharding, snapshot isolation). CDC was the hidden veto: TiCDC's open protocols and multi-sink support were non-negotiable for the data-lake/ES/cache trio.
- 教训
分布式 SQL 选型先审计"生态沉没成本"(方言、工具、CDC 下游),再谈性能;PoC 必须用真实业务查询而非只跑 Sysbench——Ninja Van 的 111.92x 来自真实 OLAP 查询;K8s Operator 的成熟度要单独验证(2020 年的 Vitess Operator 就是前车之鉴)。诚实注记:Vitess/CockroachDB 此后均有大版本演进,2021 年的落选理由不可直接套用今天;111.92x 为客户自测口径(~6TB 测试集群),未见第三方复现。
In distributed-SQL selection, audit the "ecosystem sunk costs" (dialect, tooling, CDC downstream) before talking performance; PoCs must use real business queries, not Sysbench alone — Ninja Van's 111.92x came from real OLAP queries; verify K8s Operator maturity separately (2020's Vitess Operator is the cautionary tale). Honesty note: Vitess and CockroachDB have both had major releases since; 2021's rejection reasons don't transfer directly to today; the 111.92x is the customer's self-tested claim (~6TB test cluster) with no independent reproduction found.
相关产品:TiDB、CockroachDB、MySQL 相关能力:MySQL 兼容与 TiCDC 生态 最后核验:2026-10-02
SB Payment Service:支付新系统 PoC 验证后切 TiDB,2023 年 10 月投产 成功经验
支付系统
零停机扩容
PoC 验证
MySQL 兼容
- 场景
SB Payment Service(软银集团,B2B 支付:商户入网、电商在线支付、店内终端、预付卡、运营商计费)在线支付每秒数百笔,系统故障会冲击经济活动。2021 年 11 月启动分布式 SQL 调研:最初看中的某分布式数据库与现有引擎不兼容、日本国内案例太少而放弃;2022 年初 TiDB 进入视野——MySQL 兼容 + 软银集团内已有成功案例;TiDB User Day 2022 上用户们的热情推荐("不是厂商,是用户在夸支持好")成为临门一脚。(人物与过程引自经 PingCAP 整理的 EnterpriseZine 日文报道。)
SB Payment Service (SoftBank Group; B2B payments: merchant onboarding, e-commerce online payments, in-store terminals, prepaid cards, carrier billing) handles hundreds of online payment transactions per second — a system failure would ripple into real economic activity. It began exploring distributed SQL in November 2021: the initial top choice was rejected for incompatibility with the existing engine and too few domestic use cases in Japan; TiDB entered the picture in early 2022 — MySQL-compatible with a track record inside SoftBank Group; and enthusiastic user recommendations at TiDB User Day 2022 ("not vendors, but fellow users praising the support") sealed it. (People and process quoted from the EnterpriseZine Japan report as compiled by PingCAP.)
- 决策
2022 年底到 2023 年做 PoC:在现有开发环境搭 TiDB,用生产级 workload 加压测试,并做故障测试——在 PingCAP 配合下故意 kill 实例验证恢复。结论:同等配置下 TiDB 性能超过现有库,某应用迁移后处理性能约 1.7 倍;原库升级/故障时"读副本切换"造成的停机被大幅压缩;MySQL 兼容度上,"只改连接串、核对执行计划,应用代码零改动"。落点选择很讲究:不替换现有稳定老系统(风险太高),也不做小系统(体现不出优势),而是放在"预期大幅增长"的新支付系统上。
A PoC ran from late 2022 into 2023: TiDB was set up in the existing development environment, tested with production-level workloads plus extra load, and put through failure tests — with PingCAP's cooperation, instances were intentionally killed to verify recovery. Findings: TiDB outperformed the existing database on equivalent specs, one migrated application ran about 1.7x faster, and downtime from read-replica switching during upgrades or failures was drastically reduced. On MySQL compatibility: "we only had to change the database connection and verify the SQL execution plans — no application code changes." The landing spot was chosen carefully: not the existing stable legacy system (too risky), not a small system (strengths wouldn't shine), but a new payment system expected to grow significantly.
- 结果
新支付系统 2023 年 10 月成功切换到 TiDB 生产,过程平稳、无重大问题;运维难度预计降到原来的三分之一;分区表数据膨胀问题经 PingCAP 指点用 Dumpling 命令解决;TiDB 被定位为公司核心数据库之一,后续走 TiDB Cloud(多 AZ 高可用,未来多 Region)。以上性能与运维数字为 SB Payment Service 团队在采访中的自述,经 PingCAP 转述整理,未找到第三方独立复现。
The new payment system cut over to TiDB in production in October 2023, smoothly and without significant issues; operational difficulty is expected to fall to about a third; a partitioned-table bloat problem was solved with the Dumpling command on PingCAP's advice; TiDB was positioned as one of the company's core databases, going forward on TiDB Cloud (multi-AZ availability, multi-region later). Performance and ops figures are the SB Payment Service team's own interview statements via PingCAP; no independent third-party reproduction found.
- 机制根因
计算(TiDB Server)与存储(TiKV)独立扩缩容——SQL 压力大只扩计算、容量不够只扩存储,这是分布式 SQL 相对"主从一体"最实在的运维红利;Raft 多副本让故障/升级时的切换停机从"分钟级人工操作"变成"秒级自动选举"。PoC 包含故障注入是关键:支付系统要的不是峰值 TPS 数字,而是"坏了能不能自己好",这正是读副本切换痛点被选为验证重点的原因。
Compute (TiDB Server) and storage (TiKV) scale independently — add SQL nodes when query load is high, add storage nodes when capacity runs short — the most practical ops dividend of distributed SQL over "monolithic primary-replica." Raft multi-replica turns failover/upgrade switching from "minutes of manual work" into "seconds of automatic election." Including fault injection in the PoC was the key: a payment system doesn't need a peak-TPS number, it needs "can it heal itself" — which is why the read-replica switching pain point was the focus of validation.
- 教训
选型落点比选型本身重要——把新数据库放在增长曲线最陡的新系统上,老系统"稳定压倒一切";PoC 必须包含 kill 实例的故障测试,只跑性能测试等于只验了 happy path;社区用户口碑是厂商材料之外的有效信号源(User Day 的临门一脚);分区表的数据清理/TTL 策略要在上线前规划好,别等膨胀了再救火。
Where you land a new database matters more than which one you pick — put it on the new system with the steepest growth curve and let the legacy system prioritize stability. A PoC must include kill-the-instance failure tests; performance tests alone only validate the happy path. Community user reputation is a valid signal source beyond vendor material (the User Day clincher). And plan partitioned-table cleanup/TTL strategy before go-live — don't wait for bloat to force a rescue.
相关产品:TiDB、MySQL 相关能力:— 最后核验:2026-10-02
Shopee:风控系统从 MySQL 百表分片迁到 TiDB,双写平滑迁移(2018–2019) 成功经验
风控系统
分片迁移
双写迁移
促销峰值
- 场景
Shopee(东南亚头部电商,Sea 旗下)风控系统从订单与用户行为日志中识别异常与欺诈交易,日志存在 MySQL 里、按 USER_ID 分成 100 张表。2018 年大促订单破 1100 万(上年 4.5 倍),双十二达 1200 万单、创历史纪录。团队用 InnoDB 透明页压缩把数据压掉一半、存储从 2.5TB 扩到 6TB 应急,但这只是缓兵之计,写入天花板迟早撞上。(以上业务数字为 Shopee DBA 自述,未找到第三方独立复现。)
Shopee (a leading Southeast Asia e-commerce platform under Sea) runs a risk-control system that detects abnormal and fraudulent transactions from order and user-behavior logs. The logs lived in MySQL, sharded into 100 tables by USER_ID. A 2018 promo pushed orders past 11 million (4.5x the prior year), and Double 12 hit 12 million orders, an all-time record. The team bought time with InnoDB Transparent Page Compression (halving data size) and growing storage from 2.5TB to 6TB — but that was a stopgap; the write ceiling was coming. (Business figures are the Shopee DBAs' own account; no independent third-party reproduction found.)
- 决策
团队对比了"继续分片"(100 张表重切成 1000 甚至 10000 张)与换 TiDB 两条路。继续分片意味着分片逻辑散在 Golang/Python 应用代码里:换分片键极麻烦、跨分片不支持分布式事务、挂载数据时频繁 DDL 会把库 hang 住乃至数据不一致。TiDB 的自动分片、Raft 强一致、MySQL 协议兼容和在线 DDL 正好对症。迁移采用应用层双写:双写 MySQL 与 TiDB → 历史数据迁移并校验 → 读流量渐进切到 TiDB → 停双写,历时数月、可随时回滚;还顺手按 7 个地区把风控日志拆成 7 个逻辑库(rc_sg、rc_my …… rc_tw)。
The team compared "keep sharding" (re-sharding 100 tables into 1,000 or even 10,000) against switching to TiDB. Keeping sharding meant sharding logic scattered across Golang/Python application code: changing a sharding key was painful, cross-shard distributed transactions were unsupported, and mounting data with frequent DDL could hang the database or cause inconsistency. TiDB's auto-sharding, Raft-based strong consistency, MySQL protocol compatibility, and online DDL matched the pain points. Migration used application-layer dual-write: write to both MySQL and TiDB → migrate and verify historical data → shift read traffic to TiDB gradually → stop dual-write. The process took months with rollback available at any point, and the team also split risk-control logs into 7 regional logical databases (rc_sg, rc_my, …, rc_tw).
- 结果
初始 4TB 数据落在 14 节点集群(3 PD + 3 TiDB + 8 TiKV);数据涨到 35TB 时两次扩容到 42 节点,风控与审计日志两个集群合计 60 节点。日常 QPS 不到 20000,双十二峰值冲过 100000,P99 延迟稳定在 60ms 以内。代价同样真实:团队此前零 TiDB 经验;扩容后数据重平衡很慢——约 24 小时才搬完 1TB,大促前必须提前数天扩容并调大 PD 调度参数;MySQL 兼容也不彻底,如不支持 SHOW CREATE USER,只能去读 mysql.user 系统表查账号信息。
The initial 4TB landed on a 14-node cluster (3 PD + 3 TiDB + 8 TiKV); by the time data reached 35TB the cluster had scaled twice to 42 nodes, with the risk-control and audit-log clusters totaling 60 nodes. Normal QPS stayed under 20,000, Double 12 peaked above 100,000, and P99 latency held under 60ms. The costs were real too: the team had zero prior TiDB experience; post-scale-out data rebalancing was slow — roughly 24 hours per 1TB — so scaling had to start days before promos with PD scheduling parameters raised; and MySQL compatibility was incomplete (e.g., no SHOW CREATE USER — account info had to be read from the mysql.user system table).
- 机制根因
这是"分片逻辑从应用层下沉到数据库层"的标准剧本。MySQL 分片把"数据怎么切"耦合进应用代码,每次重分片都要改代码、改分片键;TiDB 用 Region 自动分裂 + Raft 多副本把切分收进数据库内部,扩容变成加 TiKV 节点。双写迁移的价值在于把"学习 TiDB"与"切流量"解耦:数月双写期即培训期,又保留一键回滚的安全绳。
This is the textbook "push sharding logic down from the application layer into the database" play. MySQL sharding couples "how data is split" into application code, so every re-sharding means code changes; TiDB's automatic Region splitting + Raft replication pull partitioning inside the database, turning scale-out into adding TiKV nodes. Dual-write migration decouples "learning TiDB" from "cutting traffic over": the months of dual-write doubled as training, with a one-click rollback safety rope.
- 教训
当"下一次重分片"变成季度性工作,就是换分布式数据库的时机——应用层分片的债迟早要还;大促备战的扩容窗口应按"重平衡完成"而非"扩容动作"计算(约 24h/TB);边角语法的验证清单值得抄:Shopee 逐条核对并给 PingCAP 提 PR 修 SHOW CREATE USER 缺失。
When "the next re-sharding" becomes quarterly work, it's time to switch to a distributed database — application-layer sharding debt always comes due. For promo readiness, count the scaling window to "rebalancing complete," not to "scale-out action" (~24h/TB). And copy Shopee's verification checklist: they checked edge syntax one by one and filed a PR with PingCAP for the missing SHOW CREATE USER.
相关产品:TiDB、MySQL 相关能力:— 最后核验:2026-10-02
微众银行:双引擎架构,TiDB 六年扩到 80+ 集群、PB 级(2019–2025) 成功经验
银行核心账务
金融级高可用
双引擎架构
智能运维
- 场景
微众银行(中国首家互联网银行,腾讯系,2014 年成立,服务 3 亿+用户)原有 TDSQL 架构是单机垂直扩展:3TB 以下的小业务跑得很好,但容量、并发一上来就捉襟见肘,扩容往往要停机迁移、排障工具分散。银行同时立了硬指标:可用性 99.999%、RPO=0、故障恢复秒级,老架构给不出这种保证。核心账务、批量归档、合规报表等关键场景需要一条能走十年的路。
WeBank (China's first digital-only bank, backed by Tencent, founded 2014, serving 300M+ users) ran on a TDSQL architecture of vertically scaled single instances: fine for small workloads under 3TB, but capacity and concurrency hit walls fast, scaling usually meant disruptive manual migrations, and troubleshooting tools were scattered. The bank set hard targets — 99.999% availability, RPO=0, recovery in seconds — which the legacy stack could not guarantee. Core accounting, batch archiving, and compliance reporting needed a ten-year road.
- 决策
采用"双引擎"架构:TDSQL 继续承载小而简单的业务,TiDB 承接大流量高并发(核心账务、运营分析、风控)。配套自建运维平台,把分布式运维的复杂度收敛掉:集群管理、慢查询追踪、实时诊断、容量预测;容量中心用预测算法做扩容规划;智能诊断引擎在告警时自动采集日志、指标、性能快照并直接给出根因;慢日志经 ELK 实时聚合;50+ 自动化巡检覆盖复制、备份、安全配置、容量阈值。
A "dual-engine" architecture: TDSQL keeps small, simple workloads; TiDB takes high-volume, high-concurrency systems (core accounting, operational analytics, risk management). Paired with an in-house operations platform to contain distributed-ops complexity: cluster management, slow-query tracking, real-time diagnostics, capacity forecasting. A capacity center uses predictive algorithms for expansion planning; an intelligent diagnostics engine auto-collects logs, metrics, and performance snapshots on alert and delivers root causes directly; slow-query logs aggregate in real time via ELK; 50+ automated inspections cover replication, backup, security configs, and capacity thresholds.
- 结果
六年从 20 个集群扩到 80+,数据 1.3PB+,近 1000 台服务器;最大单集群 200TB+、扛 237K QPS。根因定位时间降 60% 以上;多 IDC 双活 + 智能副本放置做到 99.999% SLA、亚秒级 RTO;运维成本降 20–30%(标题作 30%,正文为 20–30%,均为 PingCAP 整理口径);扩缩容零停机,不再需要手动分片。(以上规模与成本数字来自 PingCAP 整理的分享摘录,未找到微众官方独立披露与第三方复现,引用须注口径。)
Six years, 20 → 80+ clusters, 1.3PB+ of data, nearly 1,000 servers; the largest single cluster holds 200TB+ and sustains 237K QPS. Root-cause analysis time fell by over 60%; multi-IDC active-active deployment plus intelligent replica placement deliver 99.999% SLA with sub-second RTO; operational costs fell 20–30% (the headline says 30%; both figures are PingCAP's compiled account); scaling is zero-downtime with no manual sharding. (Scale and cost figures come from PingCAP's compiled sharing notes; no independent disclosure by WeBank or third-party reproduction found — quote with the source caveat.)
- 机制根因
规模红利来自两层:TiDB 的水平扩展 + Raft 强一致解决"能不能扩",自建智能运维解决"敢不敢扩"——没有容量预测和自动诊断,PB 级分布式集群的每一次变更都是赌博。双引擎按约 3TB 分界分流,避免"一刀切":小业务继续享受单机的简单,大业务才付分布式的税。MySQL 协议兼容 + DM 工具让迁移路径平滑,开发不用重写分片逻辑。
The scale dividend came in two layers: TiDB's horizontal scalability + Raft strong consistency answered "can we scale," while the in-house intelligent operations answered "dare we scale" — without capacity forecasting and auto-diagnostics, every change on a petabyte-scale distributed cluster is a gamble. The dual engine splits at roughly 3TB to avoid one-size-fits-all: small workloads keep single-node simplicity, only large ones pay the distributed tax. MySQL protocol compatibility plus DM tooling made migration smooth, with no sharding logic to rewrite.
- 教训
分布式数据库的"能扩"不等于"敢扩",可观测性、容量预测、自动化巡检是规模化的前置条件,不是事后补课;按 workload 规模做双引擎分流,比"全上分布式"或"全守单机"都务实;金融级可用性来自"多活部署 + 副本策略 + 运维工程"的组合拳,不只是选了个分布式数据库。
A distributed database's "can scale" is not "dare scale" — observability, capacity forecasting, and automated inspection are preconditions for scale, not afterthoughts. Splitting engines by workload size beats both "all distributed" and "all single-node." Financial-grade availability comes from the combination of multi-active deployment, replica strategy, and operations engineering — not merely from picking a distributed database.
相关产品:TiDB、腾讯云 TDSQL 相关能力:— 最后核验:2026-10-02
TimescaleDB 在 PostgreSQL 上构建时序库(2017):用抽象层代替自研引擎 成功经验
时序/append-mostly
HTAP
运维简单
构建 vs 自研
- 决策
不在 PostgreSQL 之外自研存储引擎,而是在其上构建时序库(hypertable 分块抽象)。
Build the time-series database on PostgreSQL (the hypertable chunking abstraction) instead of writing a new storage engine.
- 结果
以下数字均为厂商自报口径,需打折:自称插入速度 20 倍于原生 PG;某客户场景 5 节点 TimescaleDB 对 30 节点 Cassandra。
All figures below are vendor-reported — discount heavily: claimed 20x insert speed over stock Postgres; one customer case of 5 TimescaleDB nodes versus 30 Cassandra nodes.
- 机制根因
时序负载特征鲜明(只追加、按时间查)→ 按时间分块(chunk)后,批量加载可绕开 MVCC 开销,每块自带 min/max 可做分区裁剪;元数据与时序数据同库,避免"PG + 时序库"双系统运维;复用 PG 的 SQL、事务与生态。
Time-series workloads have a distinctive exploitable shape (append-only, queried by time) → chunking by time lets bulk loads bypass MVCC overhead, and per-chunk min/max enables partition pruning; metadata and time-series data live in one database, avoiding dual-system operations; it reuses Postgres's SQL, transactions, and ecosystem.
- 教训
当工作负载有鲜明可利用的特征时,在成熟内核上加一层"利用该特征"的抽象,比自研存储引擎划算一个数量级;但厂商自报的性能对比数字一律视为营销口径,做决策前必须打折。
When a workload has a distinctive exploitable characteristic, adding a "characteristic-exploiting" abstraction on a mature kernel beats building a storage engine from scratch by an order of magnitude; but vendor-reported performance comparisons are marketing figures by default — discount before deciding.
相关产品:PostgreSQL(社区版) 相关能力:— 最后核验:2026-10-01
Uber 出走 PostgreSQL:写扩展天花板 失败教训
高并发写入
连接扩展性
零停机升级
运维负担
- 场景
Around 2016, Uber's ride-hailing business was exploding, and its core order system had to sustain high-concurrency writes across multiple data centers. Uber was on PostgreSQL 9.2 and hit several production walls: physical-replication WAL was verbose and wasteful across DCs; a replication-related bug corrupted a table in production; replicas couldn't serve consistent long-running reads; major-version upgrades required full downtime; and per-connection processes got expensive at high connection counts. (
https://www.uber.com/in/en/blog/postgres-to-mysql-migration/; the original eng.uber.com post has since moved to this URL)
- 决策
早期选 PostgreSQL(功能丰富、贴合 SQL 标准);2016 年迁往 MySQL(InnoDB),并自研 Schemaless 抽象层统一数据访问。
PostgreSQL early on (rich features, standards-compliant SQL); in 2016 Uber moved to MySQL (InnoDB) and built the Schemaless abstraction layer for unified data access.
- 结果
迁移后获得在线升级与跨 DC 高效复制能力;Schemaless 成为此后多年的标准数据访问层。草稿无量化数字,效果以工程博客自述为准。
After the migration Uber gained online upgrades and efficient cross-DC replication; Schemaless became its standard data-access layer for years. No quantified figures were published; the above rests on Uber's own engineering blog (self-reported).
- 机制根因
2016 年 PG 9.2 时代。MVCC 写放大——PG 元组不可变,UPDATE 重写整行且二级索引全指向物理位置(ctid),高写入下索引维护成本爆炸;InnoDB 二级索引指向主键(逻辑位置),更新代价小。此外:WAL 物理复制低效、备库无 MVCC、连接进程模型昂贵、大版本升级只能全停机。注:PostgreSQL 10 起引入的逻辑复制(可跨版本、表级选择复制)已解决当年的部分复制痛点;本案例的复制与升级论点仅适用于 2016 年 PG 9.2,不构成对现代 PG 复制能力的评价。
The 2016, PG 9.2 era. MVCC write amplification — PG tuples are immutable, so UPDATE rewrites the whole row and every secondary index points at a physical location (ctid); index maintenance explodes under heavy writes, while InnoDB secondary indexes point at the primary key (a logical location), making updates cheap. On top of that: inefficient physical WAL replication (MySQL's logical binlog is more compact), replicas without MVCC, a costly process-per-connection model, and major upgrades that meant full downtime. Note: logical replication introduced in PostgreSQL 10 (cross-version, table-selectable) has since addressed some of these replication pain points; the replication and upgrade arguments here apply to the 2016 PG 9.2 era only and are not an assessment of modern PostgreSQL replication.
- 教训
高写入加多 DC 场景下,复制效率、在线升级和连接成本是比"功能丰富"更硬的选型指标;但这是 2016 年 PG 9.2 时代的结论——现代 PG 已有成熟逻辑复制与更平滑的升级工具,不可直接套用今天。
For high-write, multi-DC workloads, replication efficiency, the online-upgrade path, and connection cost are harder selection criteria than "feature richness." But this is a 2016, PG-9.2-era conclusion — modern PostgreSQL has logical replication and smoother upgrade tooling, so don't apply it to today verbatim.
相关产品:PostgreSQL(社区版)、MySQL 相关能力:MVCC + 查询优化器 + 事务性 DDL —— 复杂分析直接跑在 OLTP 主库上、简单 OLTP 下"最不折腾周末"的运维体感 最后核验:2026-10-01
Booking.com:中央 embedding 服务从 OpenSearch 迁到 Weaviate,基准快 20x、成本降 40%(2024) 成功经验
向量基础设施替换
基准选型
成本优化
集中式 AI 平台
- 场景
Booking.com 的机器学习与数据科学团队运营一个中央 embedding 服务,为全公司机器学习、Agent 和 GenAI 项目提供向量检索与数据管理。服务最初建在 OpenSearch 上(团队熟悉它)。随着接入团队增多,数据集涨到数亿 embedding,查询变复杂、并发与低延迟要求上升(尤其面向用户的应用)——想在 OpenSearch 上稳住性能意味着无休止的调参与越来越大的集群,运维开销和成本同步膨胀。不同团队的需求也开始分化:有的要混合检索、多向量支持,有的要更大向量容量和更高 RPS。
Booking.com's Machine Learning & Data Science team operates a centralized embedding service providing vector search and data management for machine-learning, agent, and GenAI projects across the company. The service originally launched on OpenSearch (the team already knew it). As more teams adopted it, datasets grew to hundreds of millions of embeddings, queries grew more complex, and concurrency and low-latency expectations rose (especially for user-facing applications) -- holding OpenSearch performance steady meant endless tuning and ever-larger clusters, with operational overhead and cost swelling in step. Team requirements also diverged: some needed hybrid search and multi-vector support, others larger vector capacity and higher RPS.
- 决策
团队没有拍脑袋换库,而是按生产 workload 搭了一个贴近真实的基准:1 亿 embedding、递增并发线程,覆盖近邻检索、带过滤 KNN、混合读写。Weaviate 在被评估的数据库里表现最稳定,被选为新的共享 embedding 服务后端。具名工程师 Başak Tuğçe Eskili 的原话:"Our evaluation confirmed that systems built specifically for vector search behave better than general-purpose search engines with vector capabilities added on."(为向量搜索而生的系统,表现好于"通用搜索引擎外挂向量能力"。)
Rather than swapping on instinct, the team built a benchmark close to its production workload: 100 million embeddings, increasing concurrent threads, covering nearest-neighbor search, filtered KNN, and mixed read/write queries. Weaviate delivered the most consistent performance among the evaluated databases and was selected as the new backend for the shared embedding services platform. In the words of named engineer Basak Tugce Eskili: "Our evaluation confirmed that systems built specifically for vector search behave better than general-purpose search engines with vector capabilities added on."
- 结果
Weaviate 官网案例页公布的基准结论(厂商发布口径,基准方法与数字均未见第三方独立复现):同等规模下比 OpenSearch 快 20 倍、使用成本降 40 倍(usage cost 口径);实际运营成本降约 40%(更小的计算与内存 footprint,而 OpenSearch 那边是重度调优后的大集群);运维成熟度、部署灵活性、成本可预测性、与 Booking ML 生态的集成均占优。关键细节:因为各团队的访问都抽象在中央 embedding 服务之后,从 OpenSearch 迁到 Weaviate 对大多数团队"只是一次配置变更",无需客户端重写——迁移成本被架构设计提前消化了。
The benchmark conclusions published on Weaviate's case-study page (vendor-published; neither the benchmark methodology nor the numbers have independent third-party reproduction): 20x faster than OpenSearch at scale with 40x lower usage cost (usage-cost basis); roughly 40% lower operating cost in practice (a substantially smaller compute and memory footprint, versus a heavily tuned OpenSearch cluster); and advantages in operational maturity, deployment flexibility, cost predictability, and integration with Booking's ML ecosystem. A key detail: because every team accessed storage through the abstracted central embedding service, moving from OpenSearch to Weaviate was "little more than a configuration change" for most teams, with no client-side rewrites -- the migration cost had been absorbed by the architecture design in advance.
- 机制根因
这是一个"通用搜索外挂向量 vs 原生向量库"的对照实验:OpenSearch 的向量能力是倒排索引引擎上叠加的 ANN,过滤、混合读写、高并发下的性能曲线是后天补的;Weaviate 的 HNSW + 过滤执行路径是一体设计的,1 亿 embedding + 过滤 KNN 的 workload 正好打在它的设计点上。40% 成本下降的本质是"更小的 footprint 跑出更高的有效吞吐",即单位资源的有效 QPS 更高。但诚实地说:20x/40x 是 Booking 自家 workload 下的基准结论,换 workload、换调优水平,数字会变——能迁移的是"原生向量库在过滤+高并发下更稳"这个方向性结论,不是数字本身。
This is a controlled experiment of "general-purpose search with vector bolted on" versus "native vector database": OpenSearch's vector capability is ANN layered onto an inverted-index engine, so its performance curve under filters, mixed reads/writes, and high concurrency is retrofitted; Weaviate's HNSW plus filter execution path is designed as one, and the 100M-embedding plus filtered-KNN workload hits its design point exactly. The 40% cost reduction is essentially "higher effective throughput from a smaller footprint", i.e. higher effective QPS per unit of resource. Honestly, though: the 20x/40x figures are benchmark conclusions under Booking's own workload -- change the workload or the tuning level and the numbers change. What transfers is the directional conclusion that native vector stores are more consistent under filtered, high-concurrency workloads, not the numbers themselves.
- 教训
中央 AI 基础设施选型,基准必须按自己的生产 workload 搭(Booking 的 1 亿 embedding + 过滤 KNN + 混合读写三件套就是模板),通用基准的排名没有决策价值;"通用引擎外挂向量能力"在规模上去后会系统性输给原生向量库,这是架构代差不是调参差距;提前把存储访问抽象成内部服务,是迁移成本最低的架构投资——换库那天你会感谢当初多写的那层抽象。
For central AI infrastructure selection, the benchmark must mirror your own production workload (Booking's trio of 100M embeddings plus filtered KNN plus mixed reads/writes is a template); generic benchmark rankings have no decision value; "general-purpose engines with vector capability bolted on" systematically lose to native vector databases at scale -- that is an architecture gap, not a tuning gap; and abstracting storage access into an internal service up front is the cheapest architecture investment for migrations -- you will thank that extra abstraction layer on the day you swap databases.
来源
Weaviate 官网具名案例《Case Study - Booking.com》(厂商发布,Booking.com 机器学习工程师 Başak Tuğçe Eskili 实名 quote
Weaviate official named case study "Case Study - Booking.com" (vendor-published, with Booking.com ML engineer Basak Tugce Eskili's on-record quote
相关产品:Weaviate 相关能力:过滤优先的混合检索架构——AllowList + ACORN + 原生 BM25 最后核验:2026-10-02
DocsBot:solo founder 靠多租户隔离撑起 5 万租户、一年 610 万+ 问答(2024) 成功经验
多租户隔离
RAG 规模化
云上托管
AI 客服
- 场景
DocsBot(创始人 Aaron Edwards,前 CTO,单人创业)做"用自家文档训练的 AI 客服/文档问答"产品:每个客户都有一份隔离的知识库,跨租户数据泄露是信任红线。产品因日本一位科技博主的推文一夜爆火(230 万次展示),注册量隔夜暴增——MVP 的检索架构直接暴露生存危机:要么撑住"数万个并行小索引、彼此严格隔离",要么崩掉丢客户。早期尝试的路线要么扩展不到"数万个并行小索引",要么运维开销一个人扛不住,要么成本爆炸。
DocsBot (founder Aaron Edwards, a former CTO running a one-person company) sells AI support and documentation Q&A bots trained on each customer's own docs: every customer gets an isolated knowledge base, and cross-tenant data leakage is an existential trust line. The product went viral overnight after a tweet by a Japanese tech influencer (2.3 million impressions), and sign-ups poured in -- the MVP retrieval architecture immediately hit a survival crisis: either survive tens of thousands of parallel small indexes, strictly isolated from each other, or crash and lose customers. Early approaches either could not scale to tens of thousands of parallel indexes, demanded operational overhead one person could not carry, or were simply too costly.
- 决策
需求清单:可靠语义检索 + 混合检索 + 元数据过滤 + 干净的多租户隔离,同时运维复杂度必须接近零(没有专职 infra 团队)。评估市场后选 Weaviate Cloud:核心是其原生多租户——每个租户在 collection 内自动获得隔离分片,查询必须带租户上下文,物理隔离而非应用层的 WHERE tenant_id。创始人原话:"Weaviate stood out because it's clearly designed for real production use cases, not just experimentation. It was the only solution with an efficient tenant-based system that scaled to our unique workload of tens of thousands of distinct segmented indexes."(Weaviate 官网发布的具名案例,创始人实名 quote。)
The shortlist needed reliable semantic search plus hybrid search plus metadata filtering plus clean multi-tenant isolation, with operational complexity near zero (no dedicated infra team). After evaluating the market, he chose Weaviate Cloud, above all for native multi-tenancy: each tenant automatically gets an isolated shard inside a collection, queries must carry tenant context, and isolation is physical rather than an application-layer WHERE tenant_id. In the founder's own words: "Weaviate stood out because it's clearly designed for real production use cases, not just experimentation. It was the only solution with an efficient tenant-based system that scaled to our unique workload of tens of thousands of distinct segmented indexes." (Vendor-published named case study with the founder's on-record quote.)
- 结果
切到 Weaviate Cloud 后,单集群支撑 50,000+ 租户,一年回答 610 万+ 客户问题(以上数字均为 Weaviate 官网案例页发布的厂商口径,未见第三方独立复现)。文档摄入、分块、embedding 生成在入库前完成,查询时 Weaviate 做语义 + 混合检索取上下文,再喂给大模型生成有依据的回答。产品已从"问答引擎"向能执行动作(获客、工单路由)的 AI Agent 演进。
After moving to Weaviate Cloud, a single cluster serves 50,000+ tenants and answered 6.1 million+ customer questions in a year (all figures are vendor-published on Weaviate's case-study page; no independent third-party reproduction found). Document ingestion, chunking, and embedding generation happen before storage; at query time Weaviate performs semantic plus hybrid retrieval for context, which is then fed to large language models for grounded answers. The product has since evolved from an answer engine toward an AI agent that takes actions such as capturing leads and routing support requests.
- 机制根因
这个案例是 Weaviate 多租户机制最极端的公开验证:租户数万、单租户数据小、租户有潮汐——正是租户分片 + 状态机(ACTIVE→INACTIVE→Offloaded 到 S3)设计要解决的形状。替代路线的代价很直白:"一租户一 collection/一库"在 5 万量级下元数据和运维爆炸;"单库 + tenant_id 字段"是逻辑隔离,过滤正确性和性能随租户数退化。一个 solo founder 能把基础设施外包给云托管、只写产品逻辑,前提是租户隔离是数据库原生原语而不是应用层约定——这是该案例能成立的机制根因。诚实备注:这是 Weaviate Cloud 的案例,自托管能否复现同等规模未见独立验证。
This case is the most extreme public validation of Weaviate's multi-tenant mechanism: tens of thousands of tenants, small per-tenant data, tidal tenant activity -- exactly the shape tenant shards plus the tenant state machine (ACTIVE to INACTIVE to offloaded to S3) was designed for. The cost of the alternatives is blunt: one-collection-per-tenant or one-database-per-tenant explodes metadata and operations at 50K scale; a single store with a tenant_id field is logical isolation whose filtering correctness and performance degrade as tenant count grows. A solo founder can outsource infrastructure to managed cloud and write only product logic only when tenant isolation is a native database primitive rather than an application-layer convention -- that is the mechanism that makes this case possible. Honest caveat: this is a Weaviate Cloud case; no independent verification exists that self-hosted deployments reproduce the same scale.
- 教训
多租户 RAG SaaS 选型,第一问题不是"向量检索快不快",而是"租户隔离是数据库的原生机制还是我的应用层约定"——前者决定你能不能一个人运维 5 万租户;病毒式增长是选型的放大器:MVP 阶段"能跑就行"的检索架构,在增长 10 倍后会变成生存问题,租户模型要在第一天就选对;云托管的价值在小团队身上最大:把 infra 外包出去换来的是"创始人时间",这对 solo founder 是生死资源。
For multi-tenant RAG SaaS selection, the first question is not "how fast is vector retrieval" but "is tenant isolation a native database mechanism or my application-layer convention" -- that decides whether one person can operate 50K tenants; viral growth is a selection amplifier: a retrieval architecture that merely "works" at MVP becomes a survival problem at 10x growth, so the tenant model must be chosen correctly on day one; and managed cloud's value is largest for small teams: outsourcing infra buys "founder time", which for a solo founder is a life-or-death resource.
来源
Weaviate official named case study "Case Study - DocsBot" (vendor-published, includes founder Aaron Edwards's on-record quote and specific figures
—
相关产品:Weaviate 相关能力:多租户 SaaS 的租户分片隔离——5 万租户单集群 最后核验:2026-10-02
Finster:金融投研 AI 在生产跑 4200 万向量,从 Serverless 长到 Enterprise(2024) 成功经验
金融 RAG
混合检索
规模增长
成本优化
- 场景
Finster(2023 年成立的金融科技创业公司)给投资银行和资管机构做投研自动化:分析师要在财报季同时处理多家公司的多源文档,手动处理既慢又容易漏关键信息。金融客户对 AI 平台的要求是三条硬线:结果极准、够快、数据安全。数据源是 FactSet、Morningstar 的金融数据流,外加 SEC filing 实时摄入和自建的数千家公司数据管线,平台要支持"盈利分析、公司画像"等复杂研究任务的端到端自动化,并给出句子/单元格级的可验证引用。
Finster (a fintech startup founded in 2023) automates investment research for investment banks and asset managers: analysts must process multi-source documents for many companies simultaneously during earnings season, and manual work is both slow and prone to missing critical information. Financial clients draw three hard lines for an AI platform: extremely accurate results, fast, and secure data. Data sources are FactSet and Morningstar financial data streams plus real-time SEC filing ingestion and a custom pipeline covering thousands of companies; the platform automates complex research end-to-end (earnings analysis, company profiles) and returns sentence- and cell-level verifiable citations.
- 决策
选型理由(Weaviate 官网具名案例,厂商发布):强预过滤 + rerank + 混合检索(金融概念理解需要语义与关键词兼得);企业级能力——多租户 + VPC 部署满足银行客户的数据隔离;可扩展到数百万向量;以及"成长伙伴"式的技术支持。创始团队成员 Seán Kilgarriff 实名表示,Weaviate 的 Field CTO Byron Voorbach 早期帮他们"省了数周在各种检索方法上的迭代,直接指到了一套很有效的方案"。
Selection reasons (vendor-published named case study): strong pre-filtering plus reranking plus hybrid search (finance-specific concepts need both semantic understanding and keyword precision); enterprise readiness -- multi-tenancy and VPC deployments for bank-grade data isolation; scalability to millions of vectors; and "growth partner"-style technical support. Founding team member Sean Kilgarriff is quoted saying Weaviate's Field CTO Byron Voorbach "saved us several weeks of iterating on various retrieval methods and was able to guide us towards a specific solution that worked really well" early on.
- 结果
生产运行 4200 万向量(厂商口径,未见第三方独立复现);用 Weaviate 指导省下 4 周+ 的检索方法试错时间;为一家全球一级投行做单租户部署时一天内就进入测试(帮他们把银行漫长的销售周期走快了一步);业务增长后从 Weaviate Serverless 平滑迁到 Enterprise(高可用、压缩选项、SLA、更高 QPS),"按我们扩张的速度,Enterprise 反而更划算"。下一步是与 Weaviate 团队一起做 hot/warm/cold 存储分层与量化降本——连标杆客户都需要厂商手把手做成本优化,侧面印证向量规模上去后内存/成本是真实痛点(与"Weaviate Go 内存天花板"招牌互为表里)。
42M vectors in production (vendor-reported; no independent third-party reproduction found); 4+ weeks of retrieval-method trial-and-error saved with Weaviate's guidance; a single-tenant deployment for a global tier-one investment bank reached testing within a day (accelerating banking's notoriously long sales cycle); and after business growth the team moved smoothly from Weaviate Serverless to Enterprise (high availability, compression options, SLAs, higher QPS) -- "at the rate we were scaling, Enterprise was also more cost efficient". Next steps include hot/warm/cold storage tiering and quantization for cost optimization with the Weaviate team's guidance -- even the flagship customer needs the vendor's hands-on cost optimization, which confirms that memory and cost are real pain points once vector scale grows (the mirror image of the "Go memory ceiling" capability).
- 机制根因
金融 RAG 的检索正确性 = 预过滤(权限、时间窗口、数据源)+ 混合检索(金融术语的精确命中 + 语义理解)+ rerank 三件套,Weaviate 把三者做进同一执行路径是选型成立的技术根因;Serverless→Enterprise 的迁移路径验证了"先跑起来再长大"的可行性——但注意这是在 Weaviate 自家云产品线内部的迁移,不是跨厂商迁移。42M 向量量级下成本优化成为显性工作项,说明向量数据库的 TCO 拐点真实存在:检索能力买回来之后,内存账迟早要算。
Finance RAG retrieval correctness equals pre-filtering (permissions, time windows, data sources) plus hybrid retrieval (exact hits on finance terminology plus semantic understanding) plus reranking; Weaviate putting all three inside one execution path is the technical root of the selection. The Serverless-to-Enterprise migration validates the "get running first, grow later" path -- with the caveat that this migration stayed inside Weaviate's own cloud product line, not across vendors. At 42M vectors, cost optimization becomes explicit work, which confirms the TCO inflection point of vector databases is real: after buying retrieval capability, the memory bill eventually comes due.
- 教训
强监管行业的 RAG 选型,"数据隔离部署形态(VPC/单租户)"和"引用可验证性"是与检索质量同等权重的决策项;创业团队早期用 Serverless/托管把 infra 外包、把时间花在检索方法上,是比自建集群更划算的起步姿势;把"厂商技术支持质量"写进选型标准是务实的——Finster 省下的 4 周试错时间就是支持团队的价值量化;规模上去后主动做存储分层与量化,而不是等到账单爆炸。
For RAG selection in heavily regulated industries, "data-isolated deployment shapes (VPC / single-tenant)" and "verifiable citations" carry the same decision weight as retrieval quality; a startup team outsourcing infra to serverless/managed early and spending time on retrieval methods is a cheaper starting posture than building its own clusters; writing "vendor support quality" into the selection criteria is pragmatic -- Finster's 4 saved weeks of trial-and-error is the quantified value of the support team; and proactively tiering storage and quantizing at scale, instead of waiting for the bill to explode.
来源
Weaviate official named case study "Case Study - Finster AI" (vendor-published, with founding team member Sean Kilgarriff's on-record quote
—
相关产品:Weaviate 相关能力:过滤优先的混合检索架构——AllowList + ACORN + 原生 BM25、Go 内存模型与单机内存天花板——50M+ 向量后成本陡增 最后核验:2026-10-02
Moonsift:电商多模态发现引擎,几周内上线生产级 AI 搜索(2023) 成功经验
多模态检索
电商搜索
购物助手
快速上线
- 场景
英国创业公司 Moonsift 做电商浏览器插件,让用户从全网零售商处收藏商品、做可购物的合集。创始团队(David Wood、Alex Reed)发现零售商为搜索引擎而非用户写商品关键词,再华丽的描述也解决不了"搜不到想要的东西"——他们的愿景是用 AI 做"品味驱动的发现"。几年的插件运营给他们攒下了训练数据:6000 万+ 商品、2.5 亿次交互、4 万家零售商(Weaviate 博客原文数字,厂商发布口径),现在要把这些数据变成一个能理解用户意图的 AI 购物 Copilot。
UK startup Moonsift builds an ecommerce browser extension that lets users collect products from retailers across the internet into shoppable boards. The founding team (David Wood, Alex Reed) saw that retailers write product keywords for search engines rather than for shoppers -- even the most polished descriptions cannot fix "I can't find what I want" -- and their vision is taste-driven discovery powered by AI. Years of extension usage gave them training data: 60M+ products, 250M interactions, 40K retailers (figures from Weaviate's blog post, vendor-published), which they now want to turn into an AI shopping Copilot that understands user intent.
- 决策
机器学习负责人 Marcel Marais 先试了 BM25 + 重排的关键词方案,很快判断它撑不起推荐引擎——需要的是理解意图的语义搜索,以及能索引/检索多模态(文本+图像)、数百万级商品的系统。向量数据库的选型清单:易用排第一(要快)、开源、有活跃社区和托管选项、数百万商品规模下高性能且成本可控。评估了几个开源/闭源向量库后选定 Weaviate,理由包括:模块系统可直接集成 ML 模型并方便换模型实验、监控与复制能力、高并发查询吞吐、关键词+向量兼得的独特检索特性。
Lead ML engineer Marcel Marais first tried a keyword-based approach with BM25 plus reranking, but quickly judged it insufficient for a recommendation engine -- what they needed was semantic search that interprets user intent, plus a system that could index and search multimodal (text and image) data across millions of products. The vector-database shortlist put ease of use first (move fast), plus open-source with an active community and a managed offering, and high performance at cost-efficient scale for millions of products. After evaluating several open- and closed-source vector databases, the team chose Weaviate, citing: a module system that integrates ML models directly and makes swapping models for experiments easy, monitoring and replication capabilities, high query throughput at scale, and unique search features combining keyword and vector retrieval.
- 结果
CTO David Wood 实名表示:"Weaviate was exactly what we needed. Within a couple of weeks we had a production-ready AI-powered search engine."("几周内我们就有了生产就绪的 AI 搜索引擎",Weaviate 博客引用的创始人原话,属厂商渠道口径。)早期效果的例子很直观:"shirt that looks like a Caipirinha"(像卡琵莉亚鸡尾酒的衬衫)、"skirt with a pattern inspired by ocean waves"(海浪纹半身裙)这类意图式查询能返回合理结果。团队下一步在看 Product Quantization(压缩向量降内存 footprint、重排保相关性)和多租户——为数百万用户的个性化向量检索做准备,后者正好对应 Weaviate 的租户分片机制。
CTO David Wood is quoted: "Weaviate was exactly what we needed. Within a couple of weeks we had a production-ready AI-powered search engine." (Founder quote via Weaviate's blog, vendor channel.) Early results are intuitive: intent-style queries like "shirt that looks like a Caipirinha" or "skirt with a pattern inspired by ocean waves" return sensible results. Next on the roadmap: Product Quantization (compress vectors to cut the memory footprint, rescore to keep relevance) and multi-tenancy -- in preparation for personalized vector search for millions of customers, the latter mapping directly onto Weaviate's tenant-shard mechanism.
- 机制根因
电商发现是"语义 + 多模态 + 规模成本"的三重 workload:纯关键词在"品味"类查询上直接失效(没有关键词能描述"像卡琵莉亚的衬衫");图像向量和文本向量要进同一检索路径;6000 万商品意味着内存账必须提前算(所以 PQ 和多租户出现在路线图里)。Weaviate 的模块系统(vectorizer/model 直接集成、可换模型实验)降低了"把 ML 模型接到检索管线"的工程摩擦——这是小团队几周上线的关键:选型时"实验速度"和"检索质量"同等重要。
Ecommerce discovery is a triple workload of semantics plus multimodality plus scale-cost: pure keyword search fails outright on taste queries (no keyword describes "a shirt like a Caipirinha"); image and text vectors must enter the same retrieval path; and 60M products mean the memory bill has to be planned up front (hence PQ and multi-tenancy on the roadmap). Weaviate's module system -- vectorizers and models integrated directly, easy to swap for experiments -- lowered the engineering friction of wiring ML models into the retrieval pipeline, which is the key to a small team shipping in weeks: at selection time, "experiment velocity" matters as much as "retrieval quality".
- 教训
电商/内容发现类场景,向量库选型的第一性问题是"能不能表达用户意图"而不是"ANN 延迟",BM25 + 重排的天花板在品味类查询上很低;小团队的选型权重里,"几周内生产就绪"的工程速度可以压倒 10% 的性能差距——Moonsift 的清单把易用性放第一是理性的;提前把内存优化(PQ/量化)和多租户写进路线图:在商品规模下,检索能力上线只是上半场,成本优化是确定的下半场。
For ecommerce and content-discovery scenarios, the first-order question in vector-store selection is "can it express user intent", not "what is the ANN latency" -- BM25 plus reranking has a low ceiling on taste queries; for small teams, "production-ready in weeks" can outweigh a 10% performance gap -- Moonsift putting ease of use first was rational; and writing memory optimization (PQ/quantization) and multi-tenancy into the roadmap early: at product-catalog scale, shipping retrieval is only the first half, cost optimization is the certain second half.
来源
Weaviate 官方博客《Building an AI-Powered Shopping Copilot with Weaviate》(Moonsift 故事,厂商发布,含 CTO David Wood 实名 quote
Weaviate official blog "Building an AI-Powered Shopping Copilot with Weaviate" (the Moonsift story, vendor-published, with CTO David Wood's on-record quote
相关产品:Weaviate 相关能力:过滤优先的混合检索架构——AllowList + ACORN + 原生 BM25、Go 内存模型与单机内存天花板——50M+ 向量后成本陡增 最后核验:2026-10-02
Morningstar:Intelligence Engine 平台,内部孵化出数百个 RAG 应用(2023–2024) 成功经验
金融 RAG
自助式 AI 平台
企业级数据安全
- 场景
Morningstar 40 年积累了海量专有金融数据。2023 年初团队在自家数据片段上做 LLM 实验初见成效,意识到可以用 RAG 把几十年的长研究报告和实时数据变成 AI 应用,于是开始找向量数据库。技术负责人 Benjamin Barrett(Head of Technology, Research Products)点出的核心问题很诚实:建 chatbot "看起来像魔法,但剥开洋葱要问:它真的准吗?拉的是最新、最相关的数据吗?答案 robust 且完整吗?"——金融数据公司的命根子是"可信",准确率是信任问题不是技术指标。
Morningstar has amassed a vast proprietary financial dataset over 40 years. In early 2023 the team saw early success experimenting with LLMs on snippets of its own data and realized RAG could turn decades of long-form research content and real-time data into AI applications -- so the search for a vector database began. Benjamin Barrett (Head of Technology, Research Products) framed the core question honestly: building a chatbot "looks like magic, but when you start peeling back the layers of the onion, you have to ask, is it actually accurate? Is it pulling the latest, greatest, most relevant data? Are our answers robust and complete?" -- for a financial data company, trustworthiness is the business, and accuracy is a trust question, not a technical metric.
- 决策
选 Weaviate 的四条理由(Weaviate 官网具名案例,厂商发布):易用——开源版 Docker 一把梭就能本地起起来做实验;数据隐私与安全——灵活部署 + 多租户架构满足严格合规;灵活可扩展——从搜索引擎到定制 AI 应用、大小数据集都接得住;支持——从本地开发到生产都有人管。
Four reasons for choosing Weaviate (vendor-published named case study): ease of use -- the open-source database spins up locally in a Docker container for instant experimentation; data privacy and security -- flexible deployment options plus multi-tenant architecture for strict compliance; flexibility and scalability -- from search engines to tailored AI applications, large and diverse datasets; and support -- from local development all the way to production.
- 结果
建成了 Intelligence Engine Platform:一个"在可信金融数据和研究之上轻松创建、定制 AI 应用"的底座,内部孵化出数百个应用,覆盖内部用例和外部产品线;投研助手 Mo(Weaviate 驱动)在几周内上线,服务专业与个人投资者;RAG 管线强调动态上下文感知的文档切分 + 引用来源透明;更关键的是"自助式 RAG":内部用户通过 Corpus API(接 Weaviate)"几分钟、低代码就能建出很强的低延迟搜索引擎,还能换检索算法而不必重建索引、不碰基础设施"(高级软件工程师 Aisis Julian 实名 quote)。以上规模与速度描述均为厂商发布口径,未见第三方独立复现。
The result was the Intelligence Engine Platform: a foundation for "easily creating and customizing AI applications built on trusted financial data and research", which has allowed hundreds of applications to be created in-house across internal use cases and external product lines; the Weaviate-powered investment research assistant Mo launched within weeks for both professionals and individual investors; RAG pipelines emphasize dynamic, context-aware document chunking plus cited-source transparency; and most importantly "self-serve RAG": internal users can, via the Corpus API connected to Weaviate, "build very powerful, low latency search engines in minutes with little to no code" and "test different search algorithms without having to worry about re-indexing their data or that infrastructure at all" (senior software engineer Aisis Julian, on record). All scale and speed claims are vendor-published; no independent third-party reproduction found.
- 机制根因
这个案例的价值不在"选了 Weaviate",而在"把向量检索做成内部平台":Corpus API 把"建语料→切分→索引→检索→换算法实验"封装成自助服务,数百个内部应用是平台化的复利。能成立的前提:向量库支持多租户隔离(不同业务线数据不串)、部署灵活(合规)、检索算法可换而不重建索引(实验速度)。Weaviate 在这里的角色是"平台的检索底座"而非"某个应用的组件"——这是企业级向量库和"demo 库"的分水岭。
This case's value is not "they chose Weaviate" but "they turned vector retrieval into an internal platform": the Corpus API packages corpus-building, chunking, indexing, retrieval, and algorithm experimentation into a self-serve offering, and the hundreds of in-house applications are the compounding interest of platformization. The prerequisites: the vector store supports multi-tenant isolation (business lines' data never mixes), flexible deployment (compliance), and swappable retrieval algorithms without reindexing (experiment velocity). Weaviate's role here is "the platform's retrieval foundation", not "a component of one app" -- that is the dividing line between enterprise-grade vector stores and demo-grade ones.
- 教训
大企业用向量数据库的正确姿势往往是"先做成内部检索平台,再让业务线自助长应用",而不是每个团队各搭一套——Morningstar 的数百个应用就是平台化的证据;金融场景下"引用透明"和"数据新鲜度"是 RAG 的生死线,选型时要当一等需求验收;"换检索算法不重建索引"这类实验速度特性,决定了平台上线后能不能持续进化,而不是上线即巅峰。
The right posture for large enterprises is often "build an internal retrieval platform first, then let business lines self-serve applications", not one bespoke stack per team -- Morningstar's hundreds of apps are the evidence of platformization; in finance, "citation transparency" and "data freshness" are RAG's life-or-death lines and should be accepted as first-class requirements at selection; and experiment-velocity features like "swap retrieval algorithms without reindexing" decide whether the platform keeps evolving after launch instead of peaking on day one.
来源
Weaviate 官网具名案例《Case Study - Morningstar》(厂商发布,技术负责人 Benjamin Barrett、高级软件工程师 Aisis Julian 实名 quote
Weaviate official named case study "Case Study - Morningstar" (vendor-published, with Head of Technology Benjamin Barrett and senior software engineer Aisis Julian on record
相关产品:Weaviate 相关能力:多租户 SaaS 的租户分片隔离——5 万租户单集群、过滤优先的混合检索架构——AllowList + ACORN + 原生 BM25 最后核验:2026-10-02
Stack Overflow:在自家 Azure 上跑 Weaviate,给站内搜索加上语义检索(2023) 成功经验
语义搜索
混合检索
自托管
企业基础设施选型
- 场景
Stack Overflow 的站内搜索是核心体验(全站超过 90% 流量来自搜索引擎跳转),长期跑在 Elasticsearch 词法检索(TF-IDF + bi-gram)上。2023 年团队决定把语义搜索做进 overflow.com:用户可以用自然语言提问而不是关键词拼凑,错误码、人名等精确词仍需要词法检索托底。语义搜索原型显示效果明显("how to sort list of integers in python" 返回的相关度完胜旧搜索),但他们需要的是能跑在自己 Azure 基础设施上、同时承载词法和语义的检索系统。
Stack Overflow's on-site search is a core product experience (over 90% of site traffic arrives via search-engine referrals) and had long run on Elasticsearch lexical retrieval (TF-IDF + bi-gram shingles). In 2023 the team set out to add semantic search to overflow.com: users should be able to ask questions in natural language instead of assembling keyword incantations, while exact tokens like error codes and names still needed lexical retrieval as a backstop. A semantic-search prototype beat the old search convincingly ("how to sort list of integers in python" returned far more relevant results), but the team needed a retrieval system that could run on its own Azure infrastructure and carry both lexical and semantic retrieval on the same data.
- 决策
选型给出了三条不可谈判的需求:必须是开源、可自托管(跑在现有 Azure 基础设施上);必须同一份数据上同时支持词法与语义检索(混合检索);必须有原生 Spark 连接(团队的数据科学管线重度依赖 PySpark/Azure Databricks)。Weaviate 同时满足三条。向量管线沿用已有 Azure Databricks 平台,embedding 用 SentenceTransformers 的预训练 BERT(768 维),语料是全站数千万问答("tens of millions of questions and answers",Stack Overflow 工程博客原话),另叠加投票、浏览量、编辑信号做相关性调优。
Selection came with three non-negotiable requirements: open-source and not hosted, so it could run on the existing Azure infrastructure; hybrid search -- lexical and semantic on the same data; and a native Spark connection, because the team's data-science pipelines lean heavily on the PySpark ecosystem. Weaviate satisfied all three. The embedding pipeline reused the existing Azure Databricks platform, using a pre-trained BERT from SentenceTransformers at 768 dimensions over the site's tens of millions of questions and answers ("tens of millions of questions and answers", in the engineering blog's own words), plus votes, page views, and edit signals for relevance tuning.
- 结果
团队先在博客公开了选型与原型方法论(2023 年 7 月),随后在其"从原型到生产:GenAI 应用中的向量数据库"一文中确认 Stack Overflow 已用 Weaviate 实现混合搜索、改善了搜索结果质量("Stack Overflow has implemented hybrid search with Weaviate to achieve better search results"),并把"必须开源、不能托管、跑在自己 Azure 上"作为生产部署的硬性案例写入其中。后续规划是 RAG + LLM 摘要("OverflowAI"方向),检索质量被视为 RAG 效果的上限——"RAG 只在检索够好时才够好"。
The team first published its selection methodology and prototype in July 2023, then confirmed in a follow-up article that Stack Overflow had implemented hybrid search with Weaviate and improved search result quality ("Stack Overflow has implemented hybrid search with Weaviate to achieve better search results"), citing the open-source, non-hosted, run-on-our-Azure requirement as a hard production case. The roadmap points at RAG with LLM summaries (the "OverflowAI" direction), with retrieval quality treated as the ceiling on RAG effectiveness -- RAG is only as good as its retrieval.
- 机制根因
这个案例的选型逻辑是"部署形态优先于基准数字":企业把向量数据库当基础设施组件买断,数据主权和现有云栈(Azure + Spark)一票否决托管方案;开源 + 自托管是进门票。混合检索则是产品正确性问题:语义检索在短查询、错误码、精确词上天然弱于词法(团队自己承认),纯向量方案上线会被站内高级搜索场景打回来——Weaviate 的混合检索是"必须有"而非"加分项"。教训是 CTO 们选型向量库时,第一优先级往往是"能不能塞进我现在的云和数据栈",而不是"谁 ANN 快 5ms"。
The selection logic here is "deployment shape before benchmark numbers": enterprises buy vector databases as infrastructure components, and data sovereignty plus the existing cloud stack (Azure + Spark) veto hosted-only options -- open-source plus self-hosted is the ticket to enter. Hybrid search was a product-correctness requirement, not a bonus: the team itself admitted semantic search is weaker than lexical on short queries, error codes, and exact terms, so a pure-vector deployment would have been defeated by the site's advanced-search scenarios. The lesson for CTOs: the first priority in vector-store selection is usually "can it fit inside my current cloud and data stack", not "whose ANN is 5ms faster".
- 教训
基础设施型选型里,部署约束(自托管/云/开源)是硬门槛,性能只是门槛后的排序项;混合检索不是"向量库附赠功能",而是真实产品场景(精确词 + 自然语言并存)的必需品——纯语义方案在错误码类查询上会输给 20 年前的 TF-IDF;embedding 生成管线的可实验性(换模型、换 chunk 策略、多向量)比第一版选哪个模型更重要,Stack Overflow 把"能快速试错"做成了管线的一等特性。
In infrastructure-style selection, deployment constraints (self-hosted vs cloud vs open-source) are hard gates and performance is only a ranking item after the gate; hybrid retrieval is not a bonus feature of vector stores but a necessity for real product scenarios where exact tokens and natural language coexist -- pure-semantic search loses to twenty-year-old TF-IDF on error-code queries; and experimentability of the embedding pipeline (swap models, change chunking, multiple vectors per post) matters more than which model you pick first -- Stack Overflow made rapid experimentation a first-class property of the pipeline.
相关产品:Weaviate 相关能力:过滤优先的混合检索架构——AllowList + ACORN + 原生 BM25 最后核验:2026-10-02
小米:手机负一屏与广告业务从单机 MySQL 迁到 TiDB(2018) 成功经验
分库分表替代
MySQL 平滑迁移
高频写入
OLTP 扩容
- 场景
2018 年,小米手机桌面负一屏的快递业务与商业广告交易平台的素材抽审平台接入 TiDB。背景是 MIUI 负一屏用户量快速增长,MySQL 单机(2.6T 磁盘)出现性能明显下降、可用存储空间不断降低、大表 DDL 无法执行;智能终端业务需定时从几千万级设备高频写入监控采集数据,MySQL 的 Binlog 单线程复制导致从库延迟持续堆积。据小米 DBA 团队称,两业务每天读、写量各自达到上亿级别。(
https://cloud.tencent.cn/developer/article/1367777)
In 2018, Xiaomi's phone "minus one screen" courier business and its commercial ads trading platform's creative-review platform adopted TiDB. The backdrop: rapid MIUI minus-one-screen growth had pushed single-node MySQL (2.6TB disk) into visible performance degradation, shrinking free storage, and un-executable DDL on large tables; the smart-device business needed to ingest monitoring data from tens of millions of devices at high frequency, where MySQL's single-threaded binlog replication caused ever-growing replica lag. Per Xiaomi's DBA team, each of the two businesses reached hundreds of millions of daily reads and writes. (
https://cloud.tencent.cn/developer/article/1367777)
- 决策
DBA 团队评估后否决了分库分表:对业务代码侵入大、DBA 运维成本随拆分持续攀升、中间件有局限性,且智能终端业务需要多维度、维度可随时扩展的查询,分表方案"基本不能满足"。经兼容性对比与业务压测(TiDB 2.0.3)后决定迁移:增量用 Syncer(DM 前身)同步,读流量先灰度 1~2 周再全量切换,写流量随后迁移,全程基本无需改业务代码。
The DBA team ruled out sharding: it would invade application code, DBA operational costs would keep climbing with every re-shard, middleware had its own limits, and the device business needed multi-dimensional queries whose dimensions could expand at any time — sharding "basically couldn't satisfy" that. After compatibility comparison and workload benchmarking (TiDB 2.0.3), they migrated: incremental sync via Syncer (DM's predecessor), read traffic gray-released for 1–2 weeks before full cutover, write traffic migrated after — with essentially no application code changes.
- 结果
据小米 DBA 团队在文章中称,两业务上线数月后"整个服务稳定运行";生产集群为 3 TiDB + 3 PD + 4 TiKV(TiDB 与 PD 两两共机,共 7 台物理机)。压测数据(文章自带压测表,测试机 Xeon E5-2620 v3/128G/SSD Raid 5):标准 Select 峰值约 40920 QPS(256 并发)、标准 OLTP 峰值约 23108 QPS(64 并发)、标准 Insert 峰值约 17286 行/秒(256 并发)。生产 QPS 峰值、总数据量未找到公开数据。
Per the Xiaomi DBA team's write-up, both businesses were "running stably" months after launch; the production cluster was 3 TiDB + 3 PD + 4 TiKV nodes (TiDB and PD co-located pairwise, 7 physical machines total). Benchmark figures (from the article's own test tables, Xeon E5-2620 v3 / 128GB / SSD RAID 5): standard SELECT peaked around 40,920 QPS (256 concurrency), standard OLTP around 23,108 QPS (64 concurrency), standard INSERT around 17,286 rows/sec (256 concurrency). No public data found for production peak QPS or total data volume.
- 机制根因
TiDB 的存算分离让扩容变成"加 TiKV 节点"而非"拆库拆表",从机制上消除了"越拆越多"的运维死循环;基于 Raft 的多副本替代了 MySQL 主从 Binlog 单线程复制,从根上消除了从库延迟堆积——这是小米智能终端高频写入场景选它的决定性原因;MySQL 协议兼容使迁移无需重写业务 SQL,分布式事务保证跨分片写入 ACID,避免了分库分表下柔性事务对业务的侵入。代价是引入 PD/TiDB/TiKV 三组件的运维复杂度,以及早期版本(2.x)的成熟度风险,小米用"压测 + 灰度读流量 1~2 周"的节奏对冲。
TiDB's disaggregated compute/storage turns scaling into "add TiKV nodes" instead of "shard again," eliminating the operational death spiral of endless re-sharding; Raft-based multi-replica replaces MySQL's single-threaded binlog replication, removing replica-lag pileup at the root — the decisive reason for Xiaomi's high-frequency device-ingest workload; MySQL protocol compatibility meant no business SQL rewrites, and distributed transactions guaranteed cross-shard ACID, avoiding the flexible-transaction invasion that sharding forces on applications. The cost: operational complexity of three components (PD/TiDB/TiKV) plus the maturity risk of an early (2.x) version, hedged by Xiaomi with a "benchmark first, gray-release reads for 1–2 weeks" cadence.
- 教训
当单机 MySQL 出现"存储告罄 + 大表 DDL 做不动 + 从库延迟堆积"三联征时,是评估分布式 NewSQL 的明确信号,而非继续加分片;否决分库分表的关键判据不是性能,而是查询维度是否固定——维度随时扩展的业务,分片键会迅速失效;迁移节奏"先压测、再灰度读、最后切写",每次观察 1~2 周并备好回滚。
When single-node MySQL shows the triad of "disk exhaustion + un-runnable large-table DDL + piling replica lag," that's a clear signal to evaluate distributed NewSQL — not to add more shards. The key criterion for rejecting sharding isn't performance but whether query dimensions are fixed: businesses whose dimensions expand at any time will quickly invalidate any shard key. Migration rhythm — benchmark, gray-release reads, then cut writes — with 1–2 weeks of observation and a rollback plan at each step.
相关产品:TiDB、MySQL 相关能力:AUTO_RANDOM / SHARD_ROW_ID_BITS —— 热点打散,迁移第一课 最后核验:2026-10-01
YouTube 用 Vitess 分片 MySQL(2010s) 成功经验
水平分片
代理层路由
在线扩容
- 场景
Around 2010, YouTube's MySQL hit the single-node ceiling: first read/write splitting (primary for writes, replicas for reads), then the read replicas themselves got overloaded, so more replicas were added as a stopgap; eventually write traffic exceeded the primary's capacity and sharding became unavoidable. The first sharding logic lived in the application layer — before every database operation, code had to compute which shard to talk to. (
https://vitess.io/docs/overview/history/)
- 决策
2010 年启动 Vitess 项目,在应用与 MySQL 之间引入代理层(VTGate),把查询路由、连接池、主从切换、备份等逻辑从应用代码里抽出来;MySQL 本体不动,继续做它最擅长的单分片事务与存储。
In 2010 the Vitess project was started, introducing a proxy layer (VTGate) between the application and MySQL that took over query routing, connection pooling, failover, and backups; MySQL itself stayed untouched, doing what it does best — ACID transactions and storage within a single shard.
- 结果
据 vitess.io 官方历史文档,此后 YouTube 用户规模增长 50 倍以上;Vitess 2018 年 2 月进入 CNCF 孵化、2019 年 11 月毕业,成为 Slack、GitHub、Square 等公司也在用的通用 MySQL 分片方案。
Per the official Vitess history docs, YouTube's user base has since grown by more than 50x; Vitess entered CNCF incubation in February 2018 and graduated in November 2019, becoming the general-purpose MySQL sharding solution also used by Slack, GitHub, Square, and others.
- 机制根因
应用层分片的最大成本不是性能,而是"分片逻辑与业务代码耦合"——每次扩容、改分片键都要改代码、走发版。代理层把"数据放在哪"的决策从业务代码中剥离:MySQL 只负责单分片内的 ACID(它最成熟的能力),扩展性问题上浮到无状态的代理层解决,在线 reshard 不用停机。代价是多一跳网络延迟,以及代理层自身的高可用与运维复杂度。
The biggest cost of application-layer sharding isn't performance — it's the coupling of sharding logic to business code: every reshard or shard-key change requires code changes and a deploy. The proxy layer extracts the "where does this data live" decision out of business code: MySQL handles only single-shard ACID (its most mature capability), while scalability moves up to a stateless proxy layer with online resharding and no downtime. The price is one extra network hop and the HA/operational complexity of the proxy layer itself.
- 教训
当分片逻辑开始污染业务代码、扩容需要改代码发版时,就是引入代理层的时机;"不换库、只加一层"往往比换分布式数据库便宜一个数量级。
When sharding logic starts polluting business code and scaling requires code changes plus deploys, it's time for a proxy layer; "don't switch databases, just add a layer" is often an order of magnitude cheaper than adopting a distributed database.
相关产品:MySQL 相关能力:Vitess 水平分片 —— "第二增长曲线" 最后核验:2026-10-01
Bed Bath & Beyond:电商核心链路选型,YugabyteDB 通过故障注入与性能实测 成功经验
电商现代化
选型实测
故障注入
微服务
- 场景
这家美国全渠道零售商(Yugabyte 官方博客以 Bed Bath & Beyond 命名该访谈,见 URL slug;正文本称其为"a leading US-based omnichannel retailer",公司名未在正文中出现)的电商平台历经多轮改造:2013 年从 ASP.NET 老应用迁到 Oracle ATG Commerce,2018 年转 headless 架构,同期把数字平台改造成跑在 Kubernetes 上的 Spring Boot 微服务,数据库是 Firestore、Cloud SQL、Cassandra、Redis 的混合体。如今要把核心电商链路——促销引擎、商品目录、客户账户、registry、卡与结算——全部现代化,数据库是关键选型。
This US-based omnichannel retailer (Yugabyte's official blog names the interview after Bed Bath & Beyond in the URL slug; the body only calls it "a leading US-based omnichannel retailer" and never names the company) has rebuilt its e-commerce platform through several transformations: from a legacy ASP.NET app to Oracle ATG Commerce in 2013, to a headless architecture in 2018, and concurrently to Spring Boot microservices on Kubernetes backed by a mix of Firestore, Cloud SQL, Cassandra, and Redis. The current modernization covers the core commerce paths — promotions engine, catalog, customer accounts, registries, card and checkout — with the database as the key selection.
- 决策
Digital Engineering and Architecture 负责人带队做了完整选型:除 YugabyteDB 外还 shortlist 了另外两款数据库,按"易维护性、韧性"评估并跑了性能测试。测试路径很扎实:先用全托管的 YugabyteDB Managed 快速起库做基础验证;在 Yugabyte 团队协助下做故障注入(一边灌数据一边 kill 一个节点,看其他节点多久接管);再装到自家 VPC 里用真实电商负载跑 benchmark;还做了 COPY 批量导数(500–800 万条几分钟内完成)等 POC。最终选定 YugabyteDB,并计划用自运维的 YugabyteDB Anywhere 落地。
Led by their Director of Digital Engineering and Architecture, the team ran a full selection: besides YugabyteDB, two other databases were shortlisted and evaluated on ease of maintenance and resiliency, plus performance tests. The test path was solid: first spin up the fully managed YugabyteDB Managed for quick baseline validation; then, with Yugabyte's team, run fault injection (loading records on one side while killing a node on the other, watching how fast peers take over); then install in their own VPC and benchmark with real e-commerce load; plus POCs like bulk COPY loads (5–8 million records in a few minutes). YugabyteDB won, to be landed self-managed via YugabyteDB Anywhere.
- 结果
几个"aha 时刻"被具名记录:SQL 与 NoSQL 双接口让评估环境搭建很容易;故障注入中 peer 节点几秒内检测到 master leader 失效并自行选出新 leader;某卡相关操作的平均响应 <5ms。组织侧收益:部署频率从每月一次提到每月两次再到每周。以上均为受访人口述、经 Yugabyte 官方博客发布,属厂商口径;"另外两款数据库"的名称与测试数据未公开。
Several "aha moments" were recorded on the record: the dual SQL/NoSQL interfaces made the evaluation setup easy; during fault injection, peer nodes detected the master leader's failure within seconds and elected a new leader themselves; one card-related operation averaged under 5ms response time. Organizational payoff: deployment frequency went from monthly to twice-monthly to weekly. All of the above is the interviewee's account published on Yugabyte's official blog — vendor-sourced; the names of the "two other databases" and the test data were never published.
- 机制根因
这个案例的价值不在结论而在过程——它展示了分布式 SQL 选型的"标准动作":先全托管验证功能假设,再故障注入验证 Raft 选主与恢复时间(这是分布式数据库相对主从架构的核心差异点,必须实测),最后在自家 VPC 用真实负载跑 benchmark 验证延迟。YSQL 的 PG 兼容(含 recursive CTE、多 schema)让 Oracle ATG 时代沉淀的 SQL 资产得以复用,这是它赢过另外两款候选者的隐性加分。ACID + 地理分布 + 可扛整区故障,则对应了"价格库存必须强一致、profile 会话可接受最终一致"的分级一致性需求。
This case's value is in the process, not the verdict — it demonstrates the standard playbook for distributed SQL selection: validate functional assumptions on fully managed first, then fault-inject to verify Raft leader election and recovery time (the core differentiator of distributed databases versus primary-replica architectures — it must be measured), and finally benchmark latency in your own VPC under real load. YSQL's Postgres compatibility (including recursive CTEs, multiple schemas) let SQL assets accumulated since the Oracle ATG era be reused — a hidden bonus over the other two finalists. ACID + geo-distribution + surviving a full region outage map to the tiered consistency requirement: "price and inventory must be strongly consistent, profile sessions can be eventually consistent."
- 教训
选型不要只看功能矩阵:故障注入(kill 节点看选主时间)和真实负载 benchmark 是分布式数据库的两道必考题,不过这两关的"多活""高可用"都是纸面参数;全托管先行、自运维落地的两段式路径值得抄——评估期用云上托管把时间花在验证假设上,而不是搭集群;受访人口中的"<5ms"是特定操作的平均值,引用时不要泛化成"全链路 P99 <5ms"。
Don't select on feature matrices alone: fault injection (kill a node, time the election) and real-load benchmarking are the two mandatory exams for distributed databases — "multi-active" and "high availability" that skip them are paper specs. The two-stage path — fully managed for evaluation, self-operated for landing — is worth copying: spend evaluation time validating hypotheses on hosted cloud, not building clusters. The interviewee's "under 5ms" is an average for one specific operation; don't generalize it into "whole-path P99 under 5ms" when quoting.
来源
Yugabyte 官方博客《Omnichannel Retailer Improves Digital Customer Experience with YugabyteDB》(客户访谈实录,厂商口径
Yugabyte official blog "Omnichannel Retailer Improves Digital Customer Experience with YugabyteDB" (customer interview transcript, vendor-sourced
相关产品:YugabyteDB、Oracle Database(甲骨文)、Apache Cassandra / ScyllaDB、Redis / Valkey 相关能力:复用真实 PG 查询层 —— 迁移成本最低的分布式 PG 最后核验:2026-10-02
Jepsen 测试 YugaByte DB 1.1.9:健康集群下读偏斜致逻辑损坏,1.2.0 修复三项安全问题(2019) 失败教训
选型评估
独立测试
读偏斜
混合逻辑时钟
版本选型
- 场景
YugaByte DB 宣称"full spectrum of ACID compliance"。2019 年 3 月 Jepsen 测试 YugaByte DB 1.1.9 至 1.2.0.0-b7(5 节点,复制因子 3;测的是 YCQL 接口,YSQL 当时仍为 beta)。**注:本次为 YugaByte 付费委托**(报告明确声明 "This work was funded by YugaByte"),结论引用时须标注资助关系。
YugaByte DB claimed "full spectrum of ACID compliance." In March 2019, Jepsen tested YugaByte DB 1.1.9 through 1.2.0.0-b7 (5 nodes, replication factor 3; the YCQL interface — YSQL was still in beta at the time). **Note: this analysis was funded by YugaByte** (the report states "This work was funded by YugaByte") — cite with the funding relationship disclosed.
- 决策
Knossos 单键/多键线性一致性(counter/set/register)+ bank 快照隔离测试;故障注入包括 tablet server/master 崩溃、单节点/多数-少数/非传递网络分区、进程暂停、毫秒到数百秒的时钟偏移。
Knossos single-key/multi-key linearizability (counter/set/register) plus bank snapshot-isolation tests; fault injection including tablet server/master crashes, single-node/majority-minority/non-transitive network partitions, process pauses, and clock skew from milliseconds to hundreds of seconds.
- 结果
3 项安全问题——① 健康集群、无故障时 read skew 致逻辑数据损坏(bank 测试可复现);② 时钟偏移下 read skew;③ 多网络分区下少量确认插入丢失。另有内存泄漏(节点快速耗尽内存)、leader 选举竞态可致全集群无限期宕机等可用性问题。1.2.0 修复全部 3 项安全问题(read skew 在 1.1.10 已先修,GitHub issue #894 有完整复现与修复记录);YugaByte 官方发布 companion piece 回应。勘误:2019-04-10 发现 long fork 检查器自身有 bug,用修正后的检查器复测 1.1.15.0-b16 仍通过。
Three safety issues — (1) read skew causing logical data corruption in healthy clusters with no faults (reproducible in the bank test); (2) read skew under clock skew; (3) occasional loss of small numbers of acknowledged inserts during network partitions. Also: a memory leak letting nodes rapidly exhaust all memory, and a leader-election race that could take down the entire cluster indefinitely. Version 1.2.0 fixed all three safety issues (read skew was fixed earlier in 1.1.10, with full reproduction and fix history in GitHub issue #894); YugaByte published an official companion piece in response. Errata: on 2019-04-10 the long-fork checker itself was found buggy; re-testing 1.1.15.0-b16 with the corrected checker still passed.
- 机制根因
YugaByte 为性能让读绕过 Raft(用 leader lease 做本地读)+ 跨分片用混合逻辑时钟(HLC)定序——读快照取时间戳 t,依赖各分片"已收敛到 t"的等待,而 HLC 本质是"用时钟换性能"的折中。时钟一偏,读到的就是错乱的时间线,read skew 于是发生。这是"为性能绕过共识"的架构税:单分片内 Raft 保线性一致,多分片间 HLC 只给尽力而为。
For performance, YugaByte let reads bypass Raft (local reads via leader leases) and ordered cross-shard operations with hybrid logical clocks (HLC) — a read snapshot at timestamp *t* depends on every shard having "converged to *t*," and HLC is fundamentally a "clocks for performance" trade-off. Once clocks skew, reads observe a scrambled timeline and read skew follows. This is the architecture tax of bypassing consensus for speed: linearizability within a shard via Raft, best-effort only across shards via HLC.
- 教训
版本选型结论明确——1.1.x 别上生产,1.2.0 起才通过安全测试;任何依赖时钟同步的一致性承诺,评估时必须把"时钟异常时降级成什么"写进测试用例;厂商付费委托报告引用时标注资助关系。诚实注记:这是 2019 年 YCQL 接口的结论,YSQL 当时还是 beta;Yugabyte 此后演进多年,不可直接套用当前版本。
The version-selection conclusion is crisp — don't run 1.1.x in production; 1.2.0 onward passes the safety tests. Any consistency promise that depends on clock synchronization must include "what does it degrade to when clocks misbehave" in the test plan; disclose the funding relationship when citing vendor-funded reports. Honesty note: these conclusions concern the 2019 YCQL interface, when YSQL was still beta; YugabyteDB has evolved for years since — do not apply directly to current releases.
来源
Jepsen《Jepsen: YugaByte DB 1.1.9》(2019-03-26,YugaByte 付费委托
Jepsen, "Jepsen: YugaByte DB 1.1.9" (2019-03-26, funded by YugaByte
相关产品:YugabyteDB 相关能力:混合逻辑时钟与跨分片事务隔离 最后核验:2026-10-02
Justuno:Cassandra、Neo4j、SQL Server、CockroachDB 四库并入一个 YugabyteDB 集群 成功经验
数据库整合
多模统一
成本优化
SaaS 增长
- 场景
Justuno 是云原生的网站访客转化优化(CRO)平台,为数万家电商品牌(如 Pura Vida、Blenders Eyewear、Rothy's)做个性化站内营销。访客画像数据要求极高可用与极低延迟,且流量有明显季节性波峰。历史上他们为此堆了四套系统:Cassandra、Neo4j、Microsoft SQL Server 和 CockroachDB。Yugabyte 官方成功故事页写道:"Justuno previously relied on a variety of databases including CockroachDB, but its poor performance became a liability."厂商口径
Justuno is a cloud-native website visitor conversion optimization (CRO) platform doing personalized onsite marketing for tens of thousands of e-commerce brands (Pura Vida, Blenders Eyewear, Rothy's, and others). Visitor-profile data demands extreme availability and very low latency, with pronounced seasonal traffic spikes. Historically they stacked four systems for this: Cassandra, Neo4j, Microsoft SQL Server, and CockroachDB. Yugabyte's official success-story page states: "Justuno previously relied on a variety of databases including CockroachDB, but its poor performance became a liability." 厂商口径
- 决策
CTO/联合创始人 Travis Logan 决定把四套系统合并进单个 YugabyteDB 集群,并用 YugabyteDB Anywhere 自运维。选择逻辑:与其为四种数据模型维护四个运维面,不如用一个同时支持 SQL 与 NoSQL 接口的分布式数据库收敛;且 CockroachDB 的实测性能已成为瓶颈,合并本身就是一次升级。
CTO/Co-Founder Travis Logan decided to consolidate all four systems into a single YugabyteDB cluster, self-operated with YugabyteDB Anywhere. The logic: instead of maintaining four operational surfaces for four data models, converge on one distributed database that serves both SQL and NoSQL interfaces; and since CockroachDB's measured performance had already become the bottleneck, the consolidation was itself an upgrade.
- 结果
Yugabyte 官方页宣称:16 节点横跨 3 个 GCP 可用区,承载 20,000 SQL QPS,读延迟 <3ms,节点数仅为原 CockroachDB 方案的三分之一("3x fewer nodes than CockroachDB"),季节性扩缩容变成点选操作。Travis Logan 引述:"By consolidating our Cassandra, Neo4j, Microsoft SQL Server, and CockroachDB systems into a single YugabyteDB cluster we were able to radically simplify our operations…"以上数字均为厂商口径,未找到 Justuno 官方工程博客或第三方独立复现。
Yugabyte's official page claims: 16 nodes across 3 GCP availability zones, 20,000 SQL QPS, sub-3ms read latency, on one-third the nodes of the previous CockroachDB deployment ("3x fewer nodes than CockroachDB"), with seasonal scaling reduced to point-and-click. Travis Logan is quoted: "By consolidating our Cassandra, Neo4j, Microsoft SQL Server, and CockroachDB systems into a single YugabyteDB cluster we were able to radically simplify our operations…" All figures are vendor-claimed; no Justuno engineering blog or independent third-party reproduction found.
- 机制根因
多库并存的最大成本不是 license,而是四个备份、四个监控、四套故障排查手册和四套扩容剧本——运维面随系统数增加而叠加(接近线性,取决于自动化程度)。YugabyteDB 的双 API(YSQL + YCQL)架构让关系查询与 KV/宽列跑在同一份 Raft 复制的 tablet 存储上,合并后故障域与运维面合一。值得注意的是同一页承认他们之前用过 CockroachDB 且"性能成为负担":同为分布式 SQL,YugabyteDB 在这套访客画像 workload 上用更少节点跑出更好性能——但 workload 细节未公开,无法判断是架构差还是调优差,引用时须保留这一不确定性。
The biggest cost of running four databases side by side isn't licensing — it's four backup regimes, four monitoring stacks, four incident runbooks, and four scaling playbooks; the operational surface stacks up as system count grows (near-linear, depending on automation maturity). YugabyteDB's dual-API (YSQL + YCQL) architecture runs relational queries and KV/wide-column workloads on the same Raft-replicated tablet storage, so post-consolidation there is one failure domain and one operational surface. Notably, the same page admits they previously ran CockroachDB and its "performance became a liability": two distributed SQL engines, and YugabyteDB ran this visitor-profile workload better on fewer nodes — but workload details were never published, so whether that's an architecture gap or a tuning gap is unverifiable; keep that uncertainty when quoting.
- 教训
数据库整合本身就是降本:四个系统的运维税经常超过任何单一系统的性能税;从竞品分布式 SQL 迁出的案例提醒我们,"分布式 SQL"不是单一性能档位,同一 workload 在不同实现上可差出数倍,选型必须拿自己的 workload 实测;季节性业务选数据库时,把"弹性扩缩的操作成本"和峰值性能放在同一张表里打分。
Database consolidation is itself a cost cut: the ops tax of four systems usually exceeds any single system's performance tax. A migration *off* a competing distributed SQL engine is a reminder that "distributed SQL" is not one performance tier — the same workload can differ several-fold across implementations, so selection must be benchmarked on your own workload. For seasonal businesses, score "operational cost of elastic scaling" on the same sheet as peak performance.
来源
Yugabyte official success story "A Customer Success Story: Justuno" (vendor claim
—
相关产品:YugabyteDB、Apache Cassandra / ScyllaDB、CockroachDB、Microsoft SQL Server 相关能力:— 最后核验:2026-10-02
Kroger:全渠道数字化,YugabyteDB 支撑多地域微服务(2020–2022) 成功经验
全渠道零售
微服务
多地域
混合云
- 场景
Kroger 是美国营收第一的连锁超市(约 3000 家门店、42 个州),数字渠道是增长最快的部分。2020 年前后 Kroger 启动数字化转型:技术栈老化、point solution 太多,开始转向微服务 + 混合云(on-prem、门店边缘、公有云)。数据层需求很明确:同时支持 SQL 与 NoSQL 负载、分布式 ACID 事务、自动 geo-distribute、自动分片、开源且有商业支持。VP Engineering Mahesh Tyagarajan 在 2020 年 Distributed SQL Summit 上公开分享了这一实践。
Kroger is the largest US supermarket by revenue (~3,000 stores across 42 states), with its digital channel the fastest-growing part of the business. Around 2020 Kroger launched a digital transformation: an aging tech stack and too many point solutions drove a shift to microservices plus hybrid cloud (on-prem, in-store edge, public cloud). The data-layer requirements were explicit: support both SQL and NoSQL workloads, distributed ACID transactions, automatic geo-distribution, auto-sharding, and open source with commercial backing. VP Engineering Mahesh Tyagarajan publicly shared the practice at the 2020 Distributed SQL Summit.
- 决策
Kroger 选择 YugabyteDB 作为微服务的数据层,跑在 Kubernetes 里。部署拓扑有两种:三地域同步复制的 geo-distributed 单集群,以及两地域双集群双向异步复制(xCluster)——后者服务于 shopping list 应用,把数据放得离用户更近以换取低延迟。Sriram Samu(VP Engineering, Customer Technology)引述:"We have been leveraging YugabyteDB as the distributed SQL database running natively inside Kubernetes to power the business-critical apps that require scale and high availability."
Kroger chose YugabyteDB as the data layer for its microservices, running natively inside Kubernetes. Two deployment topologies: a geo-distributed single cluster with synchronous replication across three regions, and bidirectional asynchronous replication (xCluster) between two clusters in different regions — the latter serving the shopping-list application, placing data closer to users for low latency. Sriram Samu (VP Engineering, Customer Technology) is quoted: "We have been leveraging YugabyteDB as the distributed SQL database running natively inside Kubernetes to power the business-critical apps that require scale and high availability."
- 结果
Yugabyte 官方材料称 Kroger 集群达到数据层 single-digit 毫秒延迟、多地域 active-active 部署,生产环境超过 5000 核(">5,000 Cores in production",Mahesh Tyagarajan 口径)。2022 年 Distributed SQL Summit 上 Sriram Samu 以"Fireside Chat with Kroger: Examining a Two-Year Journey with Distributed SQL and What's Next"为题回顾了两年实践。以上数字与进展均为厂商口径,未找到 Kroger 官方工程博客披露对应数字。
Yugabyte's official material claims single-digit millisecond data-tier latency, multi-region active-active deployments, and over 5,000 cores in production (">5,000 Cores in production," Mahesh Tyagarajan's figure). At the 2022 Distributed SQL Summit, Sriram Samu revisited two years of practice in "Fireside Chat with Kroger: Examining a Two-Year Journey with Distributed SQL and What's Next." All figures and progress are vendor-claimed; no Kroger engineering blog discloses corresponding numbers.
- 机制根因
Kroger 的解法是"按一致性需求选拓扑":需要强一致的多地域写走同步复制的三地集群(Raft quorum 跨地域,写延迟换零 RPO);shopping list 这类可接受异步的场景走双集群 xCluster 双向复制,用最终一致性换本地低延迟。同一个数据库产品同时提供同步多活与异步双活两种拓扑,避免了"一个需求一种数据库"的 sprawl。这是分布式 SQL 相对 Cassandra(无分布式事务)与单机 PG(无多活)的结构性优势:把"一致性—延迟" tradeoff 做成部署选项,而不是产品选型题。
Kroger's answer is "pick the topology by consistency requirement": multi-region writes needing strong consistency run on the synchronously replicated three-region cluster (cross-region Raft quorum trades write latency for zero RPO); workloads like the shopping list that tolerate async run on dual-cluster bidirectional xCluster replication, trading eventual consistency for local low latency. One database product offering both synchronous active-active and asynchronous active-active topologies avoids the "one requirement, one database" sprawl. That's distributed SQL's structural edge over Cassandra (no distributed transactions) and single-node Postgres (no multi-active): the consistency-latency tradeoff becomes a deployment option, not a product-selection question.
- 教训
多活没有银弹:先给每个 workload 标注它能接受的一致性等级,再选拓扑,Kroger 的两种拓扑并存就是范本;零售业的"全渠道"本质是数据层问题——库存、价格、购物车在所有触点看到同一份真相,收银台和 App 不能各说各话;厂商 summit 演讲是可信度较高的厂商口径(真人真名出镜),但数字仍要标注口径,不宜当作独立第三方数据引用。
There is no silver bullet for multi-active: label each workload with the consistency level it can accept first, then pick the topology — Kroger running both topologies side by side is the template. Retail "omnichannel" is at heart a data-layer problem: inventory, price, and carts must show one truth at every touchpoint; the register and the app can't disagree. Vendor summit talks are higher-credibility vendor claims (real people, real names on stage), but figures still need source labeling and shouldn't be quoted as independent third-party data.
来源
Yugabyte official blog "Transforming the Omnichannel Experience at Kroger" (Nov 20, 2020
—
相关产品:YugabyteDB 相关能力:— 最后核验:2026-10-02
Narvar:DynamoDB 账单失控,迁 YugabyteDB 拿下 4 倍 TCO 下降 成功经验
成本优化
云厂商锁定
GDPR
多云
- 场景
Narvar 是客户体验 SaaS 平台,服务 800 余家零售商(Sephora、Patagonia、Home Depot 等),覆盖从售前到店内的个性化体验。随着业务增长,其 AWS 数据库(DynamoDB)在规模上来后"quickly became extremely expensive":DynamoDB 按吞吐量计费的模式在零售旺季流量洪峰下把成本推到不可接受的水平;同时全球客户的数据隐私法规要求多云部署能力。CTO Ram Ravichandran 原话:"Yugabyte helped Narvar avoid cloud lock-in, stay GDPR compliant, and save money in the process."
Narvar is a customer-experience SaaS platform serving 800+ retailers (Sephora, Patagonia, Home Depot, and others), covering personalized experiences from pre-purchase to in-store. As the business grew, its AWS databases (DynamoDB) "quickly became extremely expensive" at scale: DynamoDB's throughput-based pricing pushed costs to unacceptable levels under retail peak-season traffic spikes, while data-privacy regulations from global customers demanded multi-cloud deployment capability. CTO Ram Ravichandran's own words: "Yugabyte helped Narvar avoid cloud lock-in, stay GDPR compliant, and save money in the process."
- 决策
Narvar 寻找开源、云原生的数据库,要求支持多云/多地域部署、简单可自运维的 DBaaS 形态,关键标准是高性能、易扩展、"consistent cost-per-query"(每查询成本可预测)。最终从 AWS DynamoDB 与 PostgreSQL RDS 迁到 YugabyteDB,一次性解决"账单随吞吐量失控"和"被锁定在单一云"两个问题。
Narvar looked for an open-source, cloud-native database supporting multi-cloud/multi-region deployments with simple self-managed DBaaS operations; key criteria were high performance, easy scaling, and "consistent cost-per-query." It ultimately moved from AWS DynamoDB and PostgreSQL RDS to YugabyteDB, solving "bills spiraling with throughput" and "locked into a single cloud" in one move.
- 结果
Yugabyte 官方页宣称 Narvar 实现 4 倍 TCO 下降("reducing their total cost of ownership (TCO) by 4x"),同时扩展性与性能提升、迁移过程零停机、流量洪峰期间无服务中断,并可按客户需求支持多云。该数字与"零停机"均为厂商口径:未找到 Narvar 官方工程博客披露迁移细节,4x 的计算口径(是否含人力、是否同 workload 对比)未公开,引用须注明。
Yugabyte's official page claims Narvar achieved 4x lower TCO ("reducing their total cost of ownership (TCO) by 4x") while improving scale and performance, with zero downtime during migration, no service disruptions during traffic spikes, and multi-cloud support per customer needs. Both the figure and "zero downtime" are vendor-claimed: no Narvar engineering blog discloses migration details, and the 4x calculation basis (whether it includes labor, whether it's a like-for-like workload comparison) was never published — label the source when quoting.
- 机制根因
DynamoDB 的按请求/吞吐量计费在"流量可预测且平稳"时很划算,但在"季节性尖峰 + 快速增长"下变成累退税:你为峰值吞吐能力付费,而平时用不满。自运维的分布式 SQL 把成本结构从"按吞吐量交税"变回"按机器付费",峰值只需加节点,且节点可复用于所有 workload。GDPR 与多云是第二根因:DynamoDB 把你绑在 AWS,而 YugabyteDB 的多云部署让 Narvar 可以按客户属地选云——合规需求直接变成了架构需求。
DynamoDB's per-request/throughput pricing is economical when "traffic is predictable and flat," but turns into a regressive tax under "seasonal spikes + fast growth": you pay for peak throughput capacity you don't use most of the time. Self-operated distributed SQL flips the cost structure from "pay tax per throughput" back to "pay per machine" — peaks just need more nodes, and nodes are reusable across all workloads. GDPR and multi-cloud are the second root cause: DynamoDB ties you to AWS, while YugabyteDB's multi-cloud deployment lets Narvar pick clouds by customer domicile — a compliance requirement that became an architecture requirement directly.
- 教训
Serverless 按量计费不是永远便宜:把"旺季峰值倍数 × 增长斜率"代入计费模型再做 TCO,不要只看日常账单;云锁定是成本问题也是合规问题,GDPR 属地要求会让"换云"从可选项变成必答题;迁移零停机的关键是双写/灰度而非停机窗口,Narvar 案例里"流量洪峰无中断"说明切换是在真实峰值下完成的——这是比实验室压测更硬的证据(仍是厂商口径)。
Serverless pay-as-you-go is not forever cheap: plug "peak-season spike multiple × growth slope" into the pricing model before doing TCO; don't look at the everyday bill alone. Cloud lock-in is a cost problem and a compliance problem — GDPR data-domicile rules turn "switch clouds" from optional into mandatory. The key to zero-downtime migration is dual-write/canary, not maintenance windows; Narvar's "no interruptions during traffic peaks" means the cutover happened under real peak load — harder evidence than any lab benchmark (still vendor-claimed).
来源
Yugabyte official industry page "YugabyteDB is the Choice of Modern Retailers" (vendor claim
—
相关产品:YugabyteDB、Amazon DynamoDB、PostgreSQL(社区版) 相关能力:— 最后核验:2026-10-02
pidgeiot(反例):单节点 YugabyteDB 合并回自建 PostgreSQL——分布式 SQL 的"起步税" 失败教训
起步税
生态兼容
资源下限
小团队
- 场景
pidgeiot 是独立开源项目 justins-engineering/pidgeiot(作者自述为 pre-revenue 的个人/小团队项目),技术栈是 Cloudflare Workers + Hyperdrive + Ory Kratos(身份系统)+ GreptimeDB(时序)。项目早期在 staging/prod 跑单节点 YugabyteDB(RF1),dev 环境则一直用 plain postgres:18-alpine。2026 年 7 月作者做了完整复盘并写入项目文档,结论是把单节点 YugabyteDB 换成同一台机器上的自建 PostgreSQL 18,infra 增量成本 $0/月。证据等级说明:这是独立项目的文档化决策复盘(hobby 规模),不是大厂生产事故,引用时须说明量级。
pidgeiot is the independent open-source project justins-engineering/pidgeiot (the author's own words: a pre-revenue solo/small-team project) built on Cloudflare Workers + Hyperdrive + Ory Kratos (identity) + GreptimeDB (time series). Early on, staging/prod ran single-node YugabyteDB (RF1) while dev always ran plain postgres:18-alpine. In July 2026 the author did a full review, written into the project's docs: replace the single-node YugabyteDB with self-hosted PostgreSQL 18 on the same box, at $0/mo incremental infra cost. Evidence-grade note: this is a documented decision review from an independent project at hobby scale — not a big-company production incident; state the scale when quoting.
- 决策
合并而非扩容。复盘文档列了三条硬理由:第一,workload 根本不需要水平写扩展(5 分钟 cron + 设备遥测 + 人工看板,单核 PG 都吃不饱),RF1 的 YugabyteDB 是"付着分布式运维税、拿不到 HA 收益";第二,Kratos 官方支持列表只有 PostgreSQL、MySQL、CockroachDB 和 SQLite,YugabyteDB 不在其中——Kratos 官方 tracker 上有迁移死于 `ALTER TABLE ... ALTER COLUMN TYPE` 的 closed-unresolved issue(kratos#715),Hydra 有同类报告;第三,YugabyteDB 官方生产指导是 3 节点 × 每节点 4–8 vCPU,而 PG + Patroni 方案每节点 1–2 核就能跑——在共享小机器上,Yugabyte 一个就吃掉半台机器。
Consolidate, don't scale out. The review doc lists three hard reasons: first, the workload never needed horizontal write scaling (a 5-minute cron + device telemetry + human dashboard traffic wouldn't stress one Postgres core) — RF1 YugabyteDB meant "paying distributed-ops tax with zero HA payoff"; second, Kratos's officially supported list is exactly PostgreSQL, MySQL, CockroachDB, and SQLite — YugabyteDB is absent, and Kratos's own tracker has a closed-unresolved issue where migrations die on `ALTER TABLE ... ALTER COLUMN TYPE` (kratos#715), with the same class reported against Hydra; third, YugabyteDB's own production guidance is 3 nodes × 4–8 vCPU each, while a PG + Patroni setup runs at 1–2 cores per node — on shared small boxes, YugabyteDB alone budgets half the machine.
- 结果
近-term 方向是单节点 YugabyteDB → plain PostgreSQL 18(dump/restore + 改两个配置),作者估计 1–2 个专注 session(6–10 小时)可完成,因为代码里没有 Yugabyte-specific 假设,dev 环境一直在"dry-run"同一套 schema。3 节点 RF3 的 HA plan 被 shelve:等真有 HA 需求时,方案是 PG + Patroni(etcd colocated ×3)+ HAProxy,而不是 YugabyteDB——作者测算这能省约 $80/月(~$960/年)。诚实备注:作者明确写道"YugabyteDB itself is healthier than ever as a product",PG15 rebase 后 ALTER TYPE 已支持,只是上游没人替 Kratos 测——它仍是"真需要水平写扩展"那天的 conditional fallback。
The near-term direction is single-node YugabyteDB → plain PostgreSQL 18 (dump/restore + two config values), estimated at 1–2 focused sessions (6–10 hours) because the code carries no Yugabyte-specific assumptions and dev has effectively been dry-running the same schema all along. The 3-node RF3 HA plan is shelved: when real HA need arrives, the plan is PG + Patroni (etcd colocated ×3) + HAProxy, not YugabyteDB — the author estimates ~$80/mo (~$960/yr) in savings. Honest caveat: the author explicitly writes "YugabyteDB itself is healthier than ever as a product" — the PG15 rebase added ALTER TYPE support, it's just that nobody upstream tests it for Kratos — it stays the conditional fallback for the day horizontal write scaling becomes real.
- 机制根因
分布式 SQL 的"起步税"有三笔:资源税——每个节点跑完整分布式栈(tablet server、Raft、双 API),官方生产下限 4 vCPU/节点,3 节点 RF3 起步就是 12 核固定开销,而 PG 同等"单节点故障可活"(Patroni + 同步备)1–2 核起步;生态税——"兼容 PG"不等于"被 PG 生态承认",Kratos 这类上游只测官方列表里的数据库,每次上游升级都要重新验证一遍 YugabyteDB 的 DDL 兼容性;语义税——Yugabyte 的 DDL 行为与 PG 有细微差异(如 CREATE INDEX 在线构建导致不能放在显式事务里,作者真实踩过 yugabyte-db#6240),小团队为这些差异付的排查时间远超大厂。
Distributed SQL's "starter tax" has three line items: the resource tax — every node runs the full distributed stack (tablet server, Raft, dual APIs) with a documented 4 vCPU/node production floor, so 3-node RF3 starts at 12 cores of fixed overhead, while PG with equivalent "survive one node failure" (Patroni + sync standby) starts at 1–2 cores; the ecosystem tax — "Postgres-compatible" is not "acknowledged by the Postgres ecosystem," and upstreams like Kratos only test their officially listed databases, so every upstream upgrade re-gambles YugabyteDB's DDL compatibility; the semantics tax — YugabyteDB's DDL behaves subtly differently from PG (e.g., CREATE INDEX builds online so it can't go inside an explicit transaction — the author really hit yugabyte-db#6240), and a small team pays more debugging hours for those deltas than a large one.
- 教训
先回答"你要的是 failover 还是水平写扩展":前者 PG 原生 HA(Patroni)更便宜、生态零风险,后者才是分布式 SQL 的甜蜜区;"PG 兼容"要拆成三问——SQL 方言兼容、DDL 语义兼容、上游生态承认,pidgeiot 恰恰死在第三问;小团队选分布式数据库前,先算"起步税":3 节点 × 4 vCPU 的固定开销在你当前流量下分摊到每查询是多少钱,算完再决定。
Answer "do you need failover or horizontal write scaling" first: for the former, native PG HA (Patroni) is cheaper with zero ecosystem risk; the latter is distributed SQL's sweet spot. Split "PG compatibility" into three questions — SQL dialect compatibility, DDL semantics compatibility, and upstream ecosystem acknowledgment; pidgeiot died on the third. Before a small team picks a distributed database, compute the "starter tax": what the fixed 3-node × 4-vCPU overhead amortizes to per query at your current traffic — then decide.
来源
"Distributed SQL comparison: is YugabyteDB still the right pick for the HA plan's database slot?" (researched 2026-07-26
—
相关产品:YugabyteDB、PostgreSQL(社区版) 相关能力:pidgeiot 迁回自建 PG —— 分布式 SQL 的"起步税"实录 最后核验:2026-10-02
Shopify:数万 MySQL 手工分片并入 YugabyteDB,剑指 2000 万 QPS 成功经验
去分片
全球多活
数据主权
OLTP 扩展
- 场景
Shopify 服务数百万商家、10 亿在线买家,关系型数据库整体流量达 2000 万 QPS、1500 余张表。这套流量长期跑在"数万个定制分片的 MySQL 节点 + PB 级存储"上:分片由应用层手工维护,跨地域故障切换依赖 bespoke 的主从提升流程,单点故障与运维复杂度随规模放大。Yugabyte 官方成功故事页称旧架构带来了"operational complexity, bespoke failover mechanisms, and forced application-level workarounds"。以下规模数字均为厂商口径。
Shopify serves millions of merchants and 1 billion online buyers, with 20 million QPS across its relational database footprint and 1,500+ tables. That traffic long ran on "tens of thousands of custom-sharded MySQL nodes with petabyte-scale storage": sharding was maintained by hand at the application layer, cross-region failover depended on bespoke master-promotion runbooks, and single points of failure plus operational complexity grew with scale. Yugabyte's official success-story page says the legacy architecture brought "operational complexity, bespoke failover mechanisms, and forced application-level workarounds." All scale figures below are vendor-claimed.
- 决策
Shopify 选择 YugabyteDB 的分布式 SQL 架构重建数据层,目标是单一全局命名空间 + 多地域 quorum 写 + 按大洲/区域划分的 tablespace 做数据主权隔离。官方引述选型理由:"We chose Yugabyte for its Raft consensus-based distributed SQL with Postgres semantics, geo-partitioning capabilities, and the ability to self-operate at our scale while maintaining sovereignty requirements."(
https://www.yugabyte.com/success-stories/shopify/)
Shopify chose YugabyteDB's distributed SQL architecture to rebuild its data layer, targeting a single global namespace, multi-region quorum writes, and continent/region-scoped tablespaces for data-sovereignty isolation. The official selection quote: "We chose Yugabyte for its Raft consensus-based distributed SQL with Postgres semantics, geo-partitioning capabilities, and the ability to self-operate at our scale while maintaining sovereignty requirements." (
https://www.yugabyte.com/success-stories/shopify/)
- 结果
Yugabyte 官方页宣称已迁移 200,000 QPS 生产流量,部署规模为 160 节点、7000 CPU 核、1.4 PB 裸存储,集群横跨美国与欧盟。迁移仍在进行中:目标是 20 倍扩容以承接全部 2000 万 QPS、1500 余张表,后续能力包括计算存储分离扩缩容、非投票读副本与分析链路集成。以上数字均未找到 Shopify 官方工程博客或第三方独立复现,引用须注明厂商口径。
Yugabyte's official page claims 200,000 QPS of production traffic migrated, on 160 nodes, 7,000 CPU cores, and 1.4 PB of raw storage, in a cluster stretched across the US and the EU. Migration is ongoing: the goal is 20x scale-out to carry the full 20M QPS and 1,500+ tables, with compute-storage decoupled scaling, non-voting read replicas, and analytics integration to follow. None of these figures appear in a Shopify engineering blog or independent third-party reproduction — quote them as vendor claims.
- 机制根因
手工分片的本质是把"数据分布"推给应用层,代价是跨分片无事务、扩容等于重分片工程、故障切换靠定制脚本——节点数上万后,每一次主从提升都是一次高风险手工操作。YugabyteDB 用 Raft 共识把副本与故障切换收回数据库内部:tablet 自动分裂与负载均衡替代手工分片,多地域 quorum 写替代主从架构,tablespace 级 geo-partitioning 让"数据放哪"变成 DDL 而非运维项目,PG 语义兼容则把应用改造成本压到最低。代价是把核心交易链路押注在相对年轻的分布式系统上,且自运维超大规模集群对团队的内核能力要求极高。
Manual sharding pushes "data distribution" onto the application layer, and the price is no cross-shard transactions, scaling that equals re-sharding projects, and failover via custom scripts — with tens of thousands of nodes, every master promotion is a high-risk manual operation. YugabyteDB pulls replication and failover back inside the database with Raft consensus: automatic tablet splitting and load balancing replace hand sharding, multi-region quorum writes replace the primary-replica architecture, tablespace-level geo-partitioning turns "where data lives" into DDL instead of an ops project, and Postgres-semantics compatibility minimizes application rewrite cost. The price is betting the core commerce path on a relatively young distributed system, and self-operating a hyperscale cluster demands serious database-internals capability from the team.
- 教训
当分片数以万计、故障切换还要靠手工剧本时,"换数据库"比"继续修分片工具链"更便宜;数据主权合规(175+ 国家)这类需求越早进入架构选型越好,事后补 geo-partitioning 等于重写分布层;去分片迁移是马拉松——先迁 20 万 QPS 验证架构,再谈 2000 万 QPS 全量,阶段性目标比一次性切换更现实。
When shard counts reach five figures and failover still runs on hand-written runbooks, "change the database" is cheaper than "keep fixing the sharding toolchain." Data-sovereignty compliance (175+ countries) belongs in architecture selection as early as possible — retrofitting geo-partitioning later means rewriting the distribution layer. De-sharding migration is a marathon: migrate 200K QPS to validate the architecture first, then talk about the full 20M QPS; staged goals beat big-bang cutovers.
来源
Yugabyte official success story "Shopify Counts on YugabyteDB for its AI-Ready Global Commerce Infrastructure" (vendor claim
—
相关产品:YugabyteDB、MySQL 相关能力:替 Shopify 拆掉数万手工分片 —— "去分片"本身就是招牌 最后核验:2026-10-02