先说清楚:这个页面是故意只收负面的。
这里只收录客户自发分享的负面体验与吐槽,不收录正面评价——所以它天然带有负面偏见,不代表某款数据库的全面评价,也不构成选型建议。
每条都附了可打开核实的原始引用源,并标注了印证情况:
- 多方印证 多个相互独立的来源报告过同一问题;
- 单方声音 只有一个来源,但细节充分、具名可核。
请你自己判断:这是真实存在的典型问题,还是偏见下的个例。
另外注意时效:问题会被版本修复。每条卡片都标了年份和版本,已在后续版本修复的会注明"已修复于 X 版本"。
Heads up: this page is negative by design.
It collects only customer-volunteered complaints and war stories — no praise. It therefore carries an inherent negativity bias and is not a balanced review of any database, nor a selection recommendation.
Every entry links to its original, verifiable source and is labeled for corroboration:
- Corroborated multiple independent sources report the same issue;
- Single voice one source, but detailed and attributable.
Judge for yourself whether each entry is a real, typical problem or a biased anecdote.
Mind the timestamps: issues get fixed. Each card carries its year and version; issues fixed in later versions are marked "fixed in X".
PG 兼容的 premium 定价:纯 OLTP 场景下溢价难 justify
多方印证
成本账单
- 一句话
同等规格下计算贵约 39%、存储贵约 2 倍——除非你真的用上列式引擎或向量能力,否则这笔 premium 就是白交的。
At equal specs, compute costs ~39% more and storage ~2x more than Cloud SQL — unless you genuinely use the columnar engine or vector features, the premium buys you nothing.
- 窄场景
从 Cloud SQL for PostgreSQL 评估升级到 AlloyDB 的团队;纯 OLTP、无分析/向量负载的工作负载。
Teams evaluating a move from Cloud SQL for PostgreSQL to AlloyDB; pure-OLTP workloads with no analytical or vector load.
- 机制
AlloyDB 按 vCPU/内存 + 区域分布式存储 + 网络三项计费,HA 配置计算翻倍、读池每个节点都是完整计费实例。分布式存储多副本、高可用 standby、列式引擎内存都是成本项——架构上它就是比单机形态的 Cloud SQL 贵,premium 买的是列式引擎、ScaNN 向量索引和计算存储分离,不是免费午餐。
AlloyDB bills on three axes — vCPU/memory, regional distributed storage, and networking; HA doubles compute, and every read-pool node is a fully billed instance. Multi-copy distributed storage, HA standbys, and columnar-engine memory are real cost items — architecturally it is simply more expensive than single-node-shaped Cloud SQL. The premium pays for the columnar engine, ScaNN vector indexes, and compute-storage separation, not a free lunch.
- 生产验证
来源 1:2026-04 独立 field notes 实测拆解——8 vCPU 主(HA)+ 读池 + 500GB 存储约 $1,478/月,同等 Cloud SQL Enterprise Plus 约 $1,060、Enterprise 约 $810;原话 "For pure OLTP without vector workloads, the premium is hard to justify";
来源 2:Bytebase 2025-04 定价分析——vCPU 单价 $54.51/月,比 Cloud SQL Enterprise Plus 贵 39%;数据存储 $0.339/GB/月,约 2 倍于 Cloud SQL SSD($0.17);
来源 3:2025-07 独立博客 "What to Watch Out For"——"Cost: It's not the cheapest option. You're paying for performance and flexibility.";
来源 4:2026-05 独立评测——"AlloyDB's cost and complexity are overkill for workloads that fit comfortably in a Cloud SQL Enterprise instance";且读池加节点按整实例计费("Adding a node to the read pool costs the same as adding an instance")。
Source 1: Apr 2026 independent field notes with a worked breakdown — 8-vCPU HA primary + read pool + 500GB storage ≈ $1,478/month vs ≈ $1,060 for Cloud SQL Enterprise Plus and ≈ $810 for Enterprise at comparable specs; verbatim: "For pure OLTP without vector workloads, the premium is hard to justify";
Source 2: Bytebase Apr 2025 pricing analysis — $54.51/vCPU/month, a 39% markup over Cloud SQL Enterprise Plus; data storage $0.339/GB/month, ~2x Cloud SQL SSD ($0.17);
Source 3: Jul 2025 independent blog "What to Watch Out For" — "Cost: It's not the cheapest option. You're paying for performance and flexibility.";
Source 4: May 2026 independent review — "AlloyDB's cost and complexity are overkill for workloads that fit comfortably in a Cloud SQL Enterprise instance"; read-pool nodes are full compute instances ("Adding a node to the read pool costs the same as adding an instance").
- 证据等级
`多方印证(4 个独立来源)`,独立博客 field notes ×1、厂商定价分析 ×1、独立数据工程博客 ×1、独立技术评测 ×1。
`Corroborated (4 independent sources)`, independent blog field notes x1, vendor pricing analysis x1, independent data-engineering blog x1, independent technical review x1.
- 备注
本卡主题可能与本站 [避坑] 卡重叠(成本相关)。
this card's topic may overlap an existing [Pitfall] card on this site on this site (cost-related).
Google AlloyDB(AlloyDB for PostgreSQL) 年份:2026
列式引擎 100x:宣传数字只在"对的查询"上成立
多方印证
性能问题
- 一句话
100x 是"为展示收益而精心设计的窄查询"上跑出来的——你的查询形状不对、列没进内存,收益直接归零,还要额外交一份内存税。
The 100x was measured on narrow queries designed to show off the benefit — get the query shape wrong or miss the memory fit, and the gain drops to zero while you still pay a second copy's worth of memory tax.
- 窄场景
冲着 HTAP/100x 宣传来选型的团队;OLTP 为主、偶发分析查询的工作负载。
Teams choosing AlloyDB for its HTAP/100x marketing; OLTP-dominant workloads with occasional analytical queries.
- 机制
列式引擎是计算节点上独立于行式堆的内存列存,规划器决定是否走列存路径。收益取决于三件事:查询形状是否适合列存(宽表、窄投影、聚合扫描)、相关列是否被物化进列存、数据是否装得进内存。列存是**额外于**行堆物化的——同一张表占两份内存,内存装不下时收益衰减。启用本身也是多步操作:改 flag(要重启实例)、建扩展、手动加载/训练推荐引擎。
The columnar engine is an in-memory column store on the compute node, separate from the row heap; the planner decides per query whether to take the columnar path. The payoff depends on three things: whether the query shape fits columnar access (wide tables, narrow projections, scan aggregations), whether the relevant columns are materialized into the column store, and whether the data fits in memory. Columnar data is materialized **in addition to** the row heap — the same table consumes memory twice, and the benefit degrades when it does not fit. Enabling it is itself multi-step: flip a flag (instance restart), create the extension, load data manually or train the recommendations engine.
- 生产验证
来源 4:2026-05 独立评测——100x 是 "product-page-shaped",独立测试结论是 "'10x to 100x, for the right query'";列存数据是行堆之外的额外物化,"A table in the column store consumes memory for both representations";
来源 5:Pythian 2022 年实测——作者自承测试用的是 "admittedly narrow scoped queries designed specifically to illustrate the potential benefits",并列出 caveat:列存只对特定分析查询最优、相关列必须被正确缓存、需要前期工作(sizing 列存、训练推荐引擎、确认列已加载)。
Source 4: May 2026 independent review — the 100x is "product-page-shaped"; independent testing lands at "'10x to 100x, for the right query'"; columnar data is an extra materialization on top of the heap ("A table in the column store consumes memory for both representations");
Source 5: Pythian's 2022 hands-on test — the author admits to "admittedly narrow scoped queries designed specifically to illustrate the potential benefits," with caveats: the columnar store is optimal only for specific analytical queries, the right columns must be properly cached, and upfront work is required (sizing the column store, training the recommendations engine, confirming columns are loaded).
- 证据等级
`多方印证(2 个独立来源)`,独立技术评测(2026)+ 独立咨询公司实测(2022,核心结论被 2026 年评测印证;旧测试的方法论 caveat 至今成立)。
`Corroborated (2 independent sources)`, independent technical review (2026) + independent consultancy hands-on test (2022; core conclusion confirmed by the 2026 review, and the old test's methodology caveats still hold).
- 备注
本卡主题可能与本站 [避坑] 卡重叠(列式引擎相关)。
this card's topic may overlap an existing [Pitfall] card on this site on this site (columnar-engine related).
Google AlloyDB(AlloyDB for PostgreSQL) 年份:2026
I/O 计费惊魂:读一次页面,账单记一笔
多方印证
成本账单
- 一句话
Aurora Standard 把每一次存储 I/O 都变成账单事件——Graphite 的 Aurora 花费曾占到整个 AWS 账单的八成以上,不是因为实例大,而是因为 I/O 多。
Aurora Standard turns every storage I/O into a billable event — at one point Aurora cost Graphite over 80% of its entire AWS bill, not because the instances were large, but because the I/O was.
- 窄场景
I/O 密集型负载(高频小写、同步类业务如 Graphite 的 GitHub 双向同步);按默认 Standard 计费开通的集群。
I/O-heavy workloads (frequent small writes, sync-heavy businesses like Graphite's bidirectional GitHub sync); clusters opened on the default Standard billing.
- 机制
Aurora Standard 存储与 I/O 分开计费:存储按 GB-月,I/O 按百万次请求数。缓存未命中、顺序扫描、大范围回表等读放大直接转化为费用;I/O-Optimized 则把 I/O 打包进更高的实例+存储单价。计费模式本身会改变"什么查询算贵"的定义。
Aurora Standard bills storage and I/O separately: storage per GB-month, I/O per million requests. Read amplification (cache misses, sequential scans, wide lookups) converts directly into money; I/O-Optimized instead bundles I/O into higher instance+storage rates. The billing model itself redefines which queries count as "expensive."
- 生产验证
来源 1:Graphite CTO Greg Foster 2023-08-09 具名复盘——全公司数据跑在 Aurora Postgres 上(约 4000 qps),Aurora 成本一度占 AWS 总账单 80% 以上;试过 Serverless,反而比固定实例更贵;切到 I/O-Optimized 后省了 90%;
来源 2:The Build 2026 年独立分析——Standard 下"一个读很多页的查询就是一次成本事件";预置/serverless × Standard/I-Optimized × 跨区流量 × 备份 × Performance Insights 留存的计费矩阵,让团队在立项时系统性低估真实月账单。
Source 1: Graphite CTO Greg Foster's named postmortem, Aug 9 2023 — all company data on Aurora Postgres (~4,000 qps); Aurora at one point exceeded 80% of the total AWS bill; Serverless cost more than fixed instances; switching to I/O-Optimized saved 90%;
Source 2: The Build's 2026 independent analysis — under Standard, "a query that reads a lot of pages is a cost event"; the matrix of provisioned/serverless x Standard/I-Optimized x cross-region transfer x backups x Performance Insights retention leads teams to systematically underestimate the true monthly bill at commit time.
- 证据等级
`多方印证(2 个独立来源)`,具名公司 CTO 生产复盘 + 独立技术分析。
`Corroborated (2 independent sources)`, named company CTO production postmortem + independent technical analysis.
Amazon Aurora 年份:2026
版本滞后:社区 PG 发版日,Aurora 用户只能看
多方印证
生态与信任
- 一句话
Aurora 的 PG 大版本支持比社区晚好几个月——想要发版当天就用上新特性,Aurora 不是你的选择。
Aurora's support for new Postgres major versions trails the community by many months — if you need a release-day feature, Aurora is not your database.
- 窄场景
对 PG 新大版本有强依赖的团队(如新版本的新特性);"必须跑最新版"的合规/功能约束。
Teams depending on a new Postgres major version's features; "must run the latest" compliance or feature constraints.
- 机制
Aurora 要把存储层集成向前移植到每个大版本,工作量决定了它永远慢半拍:社区 → RDS(慢几个月)→ Aurora(再慢几个月)。
Aurora must forward-port its storage-layer integration to every major version, so it is always a beat behind: community, then RDS (months later), then Aurora (months after that).
- 生产验证
来源 2:The Build——"Aurora 对新 PG 大版本的支持比社区晚好几个月;想发版当周就用上最新大版本的组织会失望";
来源 13:Harshith(2026-09)——"你放弃了版本曲线的前端(top of the version curve);如果发版日就需要新大版本的特性,Aurora 给不了"。
Source 2: The Build — "Aurora's support for new PostgreSQL major versions lags community PostgreSQL, often by many months. Teams with a 'must be on the newest version' constraint will be disappointed";
Source 13: Harshith (Sep 2026) — "You give up the top of the version curve. If you need a feature from the newest major version on release day, Aurora will not have it."
- 证据等级
`多方印证(2 个独立来源)`,独立技术分析 ×2。
`Corroborated (2 independent sources)`, independent technical analyses x2.
Amazon Aurora 年份:2026
Aurora MySQL 的 DDL:看着像元数据变更,实际全表重建 30 分钟
多方印证
性能问题运维复杂度
- 一句话
在 Aurora MySQL 上加个外键,默认走 `ALGORITHM=COPY`——全表重写,所有写被锁 30 分钟;更坑的是 Aurora 的版本号和社区 MySQL 对不上,你背的"哪个版本支持 INSTANT"口诀可能是错的。
Adding a foreign key on Aurora MySQL defaults to `ALGORITHM=COPY` — a full table rebuild with all writes locked for 30 minutes; worse, Aurora's version numbers don't map to community MySQL, so your memorized "which version supports INSTANT" rules may be wrong.
- 窄场景
Aurora MySQL 3.04(8.0.28)及更早版本的大表 DDL;照着社区 MySQL 8.0.29+ 文档做变更评审的团队。
Large-table DDL on Aurora MySQL 3.04 (8.0.28) and earlier; teams reviewing migrations against community MySQL 8.0.29+ docs.
- 机制
InnoDB 加外键默认 `foreign_key_checks=1` 时只能用 COPY 算法(全表重建+全程写锁);`INSTANT` 算法的能力边界随版本变化,而 Aurora 从 3.04(8.0.28)直接跳到 3.05(8.0.32),跳过了 8.0.29–8.0.31——社区文档说 8.0.29 支持的 instant 加列/删列,在 Aurora 3.04 上依然回退到重建。
InnoDB can only add a foreign key with the default `foreign_key_checks=1` via the COPY algorithm (full rebuild, writes locked throughout); INSTANT eligibility shifts by version, and Aurora jumped from 3.04 (8.0.28) straight to 3.05 (8.0.32), skipping 8.0.29-8.0.31 — the instant add/drop column support community docs attribute to 8.0.29 still falls back to a rebuild on Aurora 3.04.
- 生产验证
来源 14:dev.to 独立作者 2026-09-30 生产事故记录——一条加外键的迁移在 Aurora MySQL 上触发 COPY 重建,大表写被锁约 30 分钟、连接堆积、功能下线;作者总结"危险的迁移和安全的迁移在 PR 里长得一模一样";
来源 15:Remitly 后端工程师 Dimitrios Sołtysiak(客户执笔,2026-02-25)——生产大表的 schema 变更必须"小心部署":显式指定 `ALGORITHM=INPLACE, LOCK=NONE`、用不可见索引分阶段上线、不确定的变更上 gh-ost/pt-osc 级别的外部工具,迁移窗口全程盯锁等待与副本延迟。
Source 14: dev.to independent author's Sep 30 2026 incident log — a one-line add-foreign-key migration triggered a COPY rebuild on Aurora MySQL; writes to the large table blocked ~30 minutes, connections backed up, the feature went down; "the dangerous migration and the safe one look the same in the pull request";
Source 15: Remitly backend engineer Dimitrios Soltysiak (customer-authored, Feb 25 2026) — production schema changes on large tables must be "deployed carefully": explicit `ALGORITHM=INPLACE, LOCK=NONE`, invisible-index staged rollouts, gh-ost/pt-osc-class external tools when in doubt, watching lock waits and replica lag through the migration window.
- 证据等级
`多方印证(2 个独立来源)`,生产事故记录 + 客户工程团队调优实录。
`Corroborated (2 independent sources)`, production incident log + customer engineering team's tuning account.
- 备注
来源 14 作者在文末推广自研的迁移检查工具 migracheck,事故记录本身为第一手生产叙述,采信事实部分。
the Source 14 author promotes his own migration-check tool migracheck at the end; the incident log itself is a first-hand production account — only the factual parts are relied upon.
Amazon Aurora 年份:2026
LIMIT 100 不省一分钱:按"引用数据"而非结果计费
多方印证
成本账单
- 一句话
BigQuery 的账单看的是查询引用了多少列的数据,而不是返回了多少行——LIMIT 只裁输出,不裁扫描。
BigQuery bills on how much column data a query references, not how many rows come back — LIMIT trims the output, not the scan.
- 窄场景
习惯用 `SELECT *` 探索数据的分析师团队;对大表做抽样、分页、LIMIT 调试的日常查询;表未分区、无聚簇的"裸奔"表。
Analyst teams in the habit of `SELECT *` exploration; sampling, paging, and LIMIT-based debugging against large tables; "naked" tables with no partitioning or clustering.
- 机制
列式存储按"读取的列数据量"计费,与返回行数、查询耗时无关;LIMIT 作用于输出阶段,执行仍需扫描全部引用列;未分区表上 WHERE 过滤发生在读取之后(先读后滤),过滤条件本身不减少扫描量。一位评论者总结:"charges based on referenced data, not processed data"。
Columnar storage bills on bytes read in referenced columns — independent of rows returned or query duration; LIMIT applies at the output stage while execution still scans every referenced column; on unpartitioned tables a WHERE filter runs after the read (read-then-filter), so the filter itself reduces nothing. As one commenter put it: "charges based on referenced data, not processed data."
- 生产验证
来源 3:HN 讨论(2025-03-25,原帖标题"BigQuery pricing model cost us $10k in 22 seconds")——发帖人 22 秒跑出 $10,000 账单;15 条评论中多位用户交叉验证了"LIMIT 不影响价格"这一计费机制(jerrygenser:"obvious from the documentation…LIMIT does not change the price");
来源 4:Balaguru Sivasambagupta 2026-03 生产复盘——银行交易表 1.2TB/40 列,分析师每天多次 `SELECT *`,每次全表扫描 1.2TB;`WHERE country='India'` 仍扫描全部 1.2TB("filter runs after the data is read");`LIMIT 100` 照样全表扫描;改列选择 + 分区 + 聚簇后单次查询降到 3–8GB,成本下降 90% 以上。
Source 3: HN discussion (Mar 25 2025, original post titled "BigQuery pricing model cost us $10k in 22 seconds") — the poster burned $10,000 in 22 seconds; across 15 comments multiple users cross-verified the "LIMIT doesn't change the price" mechanism (jerrygenser: "obvious from the documentation…LIMIT does not change the price");
Source 4: Balaguru Sivasambagupta's Mar 2026 production postmortem — a 1.2TB / 40-column banking transactions table; analysts ran `SELECT *` several times a day, each scanning the full 1.2TB; `WHERE country='India'` still scanned all 1.2TB ("filter runs after the data is read"); `LIMIT 100` scanned the full table too; after column selection + partitioning + clustering, a typical query dropped to 3–8GB — over 90% cheaper.
- 证据等级
`多方印证(2 个独立来源)`,HN 社区讨论 + 独立生产复盘。
`Corroborated (2 independent sources)`, HN community discussion + independent production postmortem.
- 备注
原始 LinkedIn 发帖人为 Yingjun Wu(RisingWave 创始人,竞品厂商背景),本卡以 HN 社区多方验证 + Balaguru 独立复盘为证据主体;"LIMIT 不影响计费"与官方文档一致,机制本身无争议。
the original LinkedIn poster was Yingjun Wu (RisingWave founder — competing-vendor background); this card rests on the HN community's multi-party verification plus Balaguru's independent postmortem. "LIMIT doesn't affect billing" matches the official documentation, so the mechanism itself is undisputed.
Google BigQuery 年份:2026
Repair 是硬性期限:错过 gc_grace,删掉的数据会"复活"
多方印证
稳定与故障运维复杂度
- 一句话
repair 不是卫生习惯,是每一次删除的下半场——任一节点错过 `gc_grace_seconds` 窗口没跑完 repair,已删除的数据会静默复活成 zombie rows,且无任何报错。
Repair is not hygiene; it is the second half of every delete — if any node misses the `gc_grace_seconds` window without a completed repair, deleted data silently resurrects as zombie rows, with no error raised.
- 窄场景
有删除/TTL 业务的集群;节点曾宕机超过 hinted handoff 默认 3 小时窗口;repair 靠手工 cron 或长期没人看的集群;任何大版本(机制为架构性设计)。
Clusters with deletes/TTL workloads; nodes that were down past the 3-hour hinted-handoff default window; clusters where repair runs from a hand-rolled cron or nobody watches it; all versions (the mechanism is architectural).
- 机制
SSTable 不可变,删除 = 写 tombstone 墓碑;墓碑要在 `gc_grace_seconds`(默认 10 天)后才允许被 compaction 清除,而清除的安全前提是**所有副本都已见过这个删除**——repair 就是让它们"听见"的方式。hinted handoff 只覆盖 3 小时内的短暂宕机。repair 本身也不便宜:Merkle 树构建要顺序读全量数据、大范围比对会 overstreaming(一个小分歧拖回一大块数据)、修完还欠一笔 compaction 债。
SSTables are immutable, so a delete is a tombstone write; tombstones may only be purged by compaction after `gc_grace_seconds` (default 10 days), and purging is only safe if **every replica has seen the delete** — repair is how replicas "hear about" it. Hinted handoff only bridges outages under 3 hours. Repair itself is not cheap: building Merkle trees means sequentially reading the entire dataset, large-range comparisons overstream (a tiny difference pulls back a big chunk), and a completed repair leaves a compaction debt behind.
- 生产验证
来源 1:2026-09,千节点运维者实录——"Repair is the tax on eventual consistency"(repair 是最终一致性的税);运营策略是 10 天 grace 对 7 天 repair 周期,告警看的不是"repair 失败"而是"每张表距上次完整 repair 过去了多久"(staleness 才是真风险);
来源 2:2022-12,HN 一线运维者评论——"Operationally…the trickiest part is repairs"(运维上最棘手的就是 repair)。
Source 1: 2026-09, thousand-node operator account — "Repair is the tax on eventual consistency"; the operational policy is a 7-day repair cycle against a 10-day grace, and the alert that matters is not "repair failed" but "how long since this table's last completed repair" (staleness is the real risk);
Source 2: 2022-12, HN frontline operator comment — "Operationally…the trickiest part is repairs."
- 证据等级
`多方印证(2 个独立来源)`,千节点运维实录 + 一线运维者评论。
`Corroborated (2 independent sources)`, thousand-node operator account + frontline operator comment.
- 备注
与现有 [避坑] 卡主题重叠(repair 强制周期运维、gc_grace 配错致 zombie rows),双方保留。据本站档案页记录,Cassandra 5.0.8+ 已将自动 repair 调度 backport 为可选项(opt-in),属"已部分改善",但默认仍需人工排期,故保留收录。
Topic overlaps with existing [Pitfall-avoidance] cards (mandatory periodic repair; misconfigured gc_grace resurrecting zombie rows); both are kept. Per this site's profile page, Cassandra 5.0.8+ backported automatic repair scheduling as an opt-in, marked "partially improved" — manual scheduling remains the default, hence the card is retained.
Apache Cassandra / ScyllaDB 年份:2026
Tombstone 读放大:读 600 行活数据,要先扫 1.6 万个墓碑
多方印证
性能问题稳定与故障
- 一句话
删除在 Cassandra 里是"逻辑先行、物理靠后"——读请求要穿越海量墓碑才能找到活数据,墓碑一多,读延迟、CPU、堆内存一起恶化;TTL 误删后想捞数据,得 dump 墓碑外加回拨系统时间。
Deletes in Cassandra are "logical now, physical later" — read requests must traverse seas of tombstones to find live data, so read latency, CPU and heap all degrade together once tombstones pile up; recovering accidentally TTL-deleted data means dumping tombstones and rewinding the system clock.
- 窄场景
频繁删除/更新、TTL 大量过期、集合整列替换的表;墓碑堆积叠加大分区的集群。
Tables with frequent deletes/updates, heavy TTL expiry, or whole-collection replacement; clusters where tombstone buildup compounds with large partitions.
- 机制
SSTable 不可变,删除只写墓碑标记;读路径要合并多个 SSTable 并跳过墓碑,墓碑数量直接计入读放大;墓碑靠 compaction 在 gc_grace 后清除,清除前常驻。墓碑扫描量超过阈值(本站档案页记录默认 10 万)时查询直接失败。TTL 过期本质也是墓碑:墓碑数据不可查、不能直接重插(重插会因同一 TTL 立即再过期)。
SSTables are immutable, so a delete is only a tombstone marker; the read path merges multiple SSTables and skips tombstones, and the tombstone count directly multiplies read amplification. Tombstones persist until compaction purges them after gc_grace. Queries whose tombstone scan count crosses the threshold (this site's profile records a default of 100,000) fail outright. TTL expiry is tombstones underneath: tombstoned data is not queryable and cannot simply be re-inserted (a re-insert under the same TTL expires again immediately).
- 生产验证
来源 3:2025-05,Krishna Alapati 客户生产根因分析——日志频繁出现 "Read live 600 rows and 16,526 tombstone cells for query",即 16600 个 cell 里 96% 是已删除数据;读延迟飙升、CPU 与内存压力齐涨;
来源 4:2026-02,Pandu Boyina 生产复盘——TTL 策略误删关键记录导致业务中断,恢复时被迫 dump 墓碑、挪到独立服务器、回拨系统日期、改写时间戳与 TTL 字段后重插,"Tombstones exist for consistency, not for easy recovery"(墓碑为一致性而存在,不是为方便恢复)。
Source 3: 2025-05, Krishna Alapati's client production root-cause analysis — logs repeatedly showed "Read live 600 rows and 16,526 tombstone cells for query", i.e. 96% of the ~16,600 cells scanned were already-deleted data; read latency spiked while CPU and memory pressure rose together;
Source 4: 2026-02, Pandu Boyina production postmortem — a TTL policy accidentally deleted critical records and caused a business outage; recovery meant dumping tombstones, moving them to a separate server, rewinding the system date, rewriting timestamps and TTL fields before re-insertion: "Tombstones exist for consistency, not for easy recovery."
- 证据等级
`多方印证(2 个独立来源)`,客户生产根因分析 + 生产复盘。
`Corroborated (2 independent sources)`, client production root-cause analysis + production postmortem.
- 备注
与现有 [避坑] 卡主题重叠(tombstone 风暴致读放大、查询超时),双方保留。
Topic overlaps with existing [Pitfall-avoidance] cards (tombstone storms causing read amplification and query timeouts); both are kept.
Apache Cassandra / ScyllaDB 年份:2026
ScyllaDB 许可证:2025 年起,它不再是开源软件
多方印证
生态与信任
- 一句话
2024 年 12 月官方宣布、2025.1(2025 年 4 月发布)起,ScyllaDB OSS 与 Enterprise 合并为统一的 source-available 版本——最后一个纯开源版本是 AGPL 的 6.2;免费额度是每组织 10TB 磁盘 + 50 vCPU,超了就要谈商业许可;官方仓库甚至删掉了 Contribute 文档页,理由是"ScyllaDB is no longer open source"。
Announced December 2024 and effective with 2025.1 (released April 2025), ScyllaDB OSS and Enterprise merged into a single source-available edition — the last purely open-source release is AGPL 6.2; the free tier is 10TB of disk + 50 vCPUs per organization, beyond which you negotiate a commercial license; the official repository even deleted its Contribute docs page with the rationale "ScyllaDB is no longer open source."
- 窄场景
此前按"开源免费"做 TCO 预算与法务评估的团队;规模超过免费额度的自建集群;依赖社区分叉延续开源路线的用户。
Teams that budgeted TCO and legal review on "open source and free" assumptions; self-hosted clusters exceeding the free tier; users counting on a community fork to continue the open-source line.
- 机制
Source-available ≠ 开源(不符合 OSI 开源定义):源码可见但使用受商业条款限制;免费 tier 按组织维度封顶(10TB 配置磁盘、50 vCPU,跨所有集群);旧版本(6.2.x 及更早)永久保持 AGPL,但不再获得新功能与 bug 修复;外部贡献者需签 CLA,而核心引擎历史上几乎只有 ScyllaDB 公司自己在贡献。
Source-available is not open source (incompatible with the OSI Open Source Definition): source is visible but usage is bound by commercial terms; the free tier is capped per organization (10TB provisioned disk, 50 vCPUs, across all clusters); old releases (6.2.x and earlier) stay AGPL forever but receive no new features or bug fixes; external contributors must sign a CLA while the core engine has historically been contributed almost exclusively by ScyllaDB itself.
- 生产验证
来源 8:2025-06,The Register 独立报道——公司从 AGPL 转向 source-available,6.2 为最后一个 AGPL 版本,首个 source-available 产品为 Enterprise 2025.1(4 月发布);CEO 称这是为了解决"以开源为核心产品的厂商永恒的难题";
来源 9:2024-12,LinuxIAC 独立报道——确认 6.2 为"历史上最后一个 OSS AGPL 版本",并指出"一些传统 OSS 用户可能对纯开源替代品的消失感到失望";
来源 10:2025-05,ScyllaDB 官方仓库 issue #24060——文档团队提议删除 Contribute 页面,"it can be confusing since ScyllaDB is no longer open source"(这会让人困惑,因为 ScyllaDB 已不再是开源软件),经维护者 @tzach 批准、列入 2025.1;
来源 11:官方 FAQ(2026-10-03 亲自打开核验)——免费 tier 口径为"每组织 10TB 磁盘 + 50 vCPU(跨所有集群)","Old releases will not receive bug fixes or new functionality"。
Source 8: 2025-06, The Register, independent reporting — the company moved from AGPL to source-available, 6.2 the last AGPL release, Enterprise 2025.1 (April) the first source-available product; the CEO framed it as addressing the "constant challenge for vendors whose core product is open source";
Source 9: 2024-12, LinuxIAC, independent reporting — confirmed 6.2 as the "last OSS AGPL release ever" and noted that "some traditional OSS users might feel disappointed by the removal of a purely open-source alternative";
Source 10: 2025-05, ScyllaDB official repository issue #24060 — the docs team proposed removing the Contribute page because "it can be confusing since ScyllaDB is no longer open source", approved by maintainer @tzach and scheduled for 2025.1;
Source 11: Official FAQ (personally opened and verified 2026-10-03) — free tier stated as "10TB of disk and 50 vCPUs per organization (across all clusters)"; "Old releases will not receive bug fixes or new functionality."
- 证据等级
`多方印证(4 个独立来源)`,独立媒体 ×2 + 官方仓库 issue + 官方 FAQ 事实口径。
`Corroborated (4 independent sources)`, independent press ×2 + official repository issue + official FAQ factual parameters.
- 备注
与现有 [避坑] 卡主题重叠(授权坑:ScyllaDB source-available),双方保留。任务提示中"2025 年起换许可证"口径与核验一致:2024-12 宣布,2025.1(2025-04)为首个 source-available 版本。
Topic overlaps with existing [Pitfall-avoidance] cards (licensing pitfall: ScyllaDB source-available); both are kept. The task brief's "license change from 2025" matches verification: announced 2024-12, first source-available release 2025.1 (2025-04).
Apache Cassandra / ScyllaDB 年份:2026
ScyllaDB 硬件挑剔:seastar 的"贴近硬件"是有代价的
多方印证
运维复杂度成本账单
- 一句话
ScyllaDB 的 shard-per-core 架构把 GC 问题连根拔了,但换来的是对硬件的挑剔——官方口径生产节点建议 2GB/逻辑核、第三方自建指南的生产起点是 8 vCPU + 16GB + 高速 SSD;想在小机器或超售云主机上"先跑起来",体验并不友好。
ScyllaDB's shard-per-core architecture pulls the GC problem out by the root, but demands picky hardware in return — official guidance suggests 2GB per logical core for production, and an independent self-hosting guide's sensible starting point is 8 vCPU + 16GB RAM + fast SSDs; "just get it running" on small boxes or oversubscribed cloud instances is not a friendly experience.
- 窄场景
自建生产集群的选型与预算;小规格起步、超售 vCPU、远端盘的云环境。
Sizing and budgeting a self-hosted production cluster; starting small, oversubscribed vCPUs, or remote-disk cloud environments.
- 机制
Seastar 框架每 CPU 核一个 shard,各自独占内存、缓存分区与 IO 调度——设计假设是"独占高性能硬件";核数、内存配比、磁盘类型直接决定单节点吞吐上限,配比失衡时加机器不如加配置。容量规划要按 shard 而不是按节点思考(本站档案页深水区四已有论述)。
The Seastar framework runs one shard per CPU core, each with its own memory, cache partitions and IO scheduler — the design assumes "dedicated high-performance hardware"; core count, memory ratio and disk type directly bound per-node throughput, so adding the wrong configuration adds cost without capacity. Capacity planning must think in shards, not nodes (this site's profile page covers this in its shard-per-core deep dive).
- 生产验证
来源 2:2022-12,从 Cassandra 迁往 ScyllaDB 的一线运维者——"Scylla…is not without its own set of issues - and somewhat strict hardware requirements thanks to the seastar engine it is built on top of";
来源 12:2026-09,Plainwire 开源项目自建指南——"A sensible operator starting point for a real Scylla node is 8 vCPU and 16 GB RAM with fast SSD storage";引官方文档口径:生产下限 4GB,推荐 16GB 或 2GB/逻辑核取高者。
Source 2: 2022-12, a frontline operator who migrated from Cassandra to ScyllaDB — "Scylla…is not without its own set of issues - and somewhat strict hardware requirements thanks to the seastar engine it is built on top of";
Source 12: 2026-09, Plainwire open-source self-hosting guide — "A sensible operator starting point for a real Scylla node is 8 vCPU and 16 GB RAM with fast SSD storage"; quoting official docs: 4GB production floor, 16GB or 2GB per logical core whichever is higher recommended.
- 证据等级
`多方印证(2 个独立来源)`,迁移用户评论 + 独立第三方自建指南。
`Corroborated (2 independent sources)`, migrating-user comment + independent third-party self-hosting guide.
- 备注
与现有 [避坑] 卡深水区四(shard-per-core)部分重叠——该卡讲架构原理,此卡聚焦自建硬件门槛与预算影响,双方保留。
Partially overlaps with existing [Pitfall-avoidance] deep dive on shard-per-core — that card covers the architecture, this one focuses on the self-hosting hardware bar and budget impact; both are kept.
Apache Cassandra / ScyllaDB 年份:2026
运维需要专职团队:集群在干什么,很难看清
多方印证
运维复杂度成本账单
- 一句话
compaction 和 repair 的内存占用与耗时"很难建模、很难容纳",集群此刻在干什么"很难看清"——有用户直言,没有高薪专职 Cassandra DBA 的团队,体验是另一个世界;DataStax 的商业支持也被吐槽"相当惨淡"。
Compaction and repair memory usage and duration are "hard to model and hard to accommodate," and what the cluster is doing at any given moment is "hard to tell" — one user put it plainly: without very highly paid dedicated Cassandra DBAs, the experience is a different world; DataStax's commercial support was called "pretty dismal."
- 窄场景
没有专职 DBA/平台团队的中小团队;把 Cassandra 当"装好即用"型数据库的团队。
Small and mid-size teams without dedicated DBA/platform engineers; teams treating Cassandra as an "install and forget" database.
- 机制
后台任务(compaction、repair、hinted handoff、gossip)与前台读写争抢同一批 CPU/内存/IO,且行为与数据模型、墓碑量、分区大小强相关——"黑盒感"来自可观测性与行为建模的双重缺失。出问题时排查链条长(JVM、compaction 策略、repair 进度、墓碑、数据模型逐项排查),对人员经验要求陡峭。
Background work (compaction, repair, hinted handoff, gossip) contends with foreground reads/writes for the same CPU, memory and IO, and its behavior depends heavily on the data model, tombstone volume and partition sizes — the "black box" feeling comes from the twin absence of observability and behavioral modeling. Debugging chains are long (JVM, compaction strategy, repair progress, tombstones, data model, one by one), so the experience bar is steep.
- 生产验证
来源 5:HN 用户——"The memory consumption and duration of compactions and column/node repair jobs is hard to model and accommodate. It's hard to tell what the cluster is doing at any given moment. Our experience with support plans from Datastax was also pretty dismal";同一讨论串另一用户:"At a much bigger company that had some very, very highly paid Cassandra DBAs it was actually a relatively smooth experience"(有高薪专职 DBA 才顺滑);
来源 1:2026-09,千节点运维者——"Repair is the tax on eventual consistency. You can pay it on a schedule you chose, or all at once with interest",即运维投入是刚性的,不是可选的。
Source 5: HN users — "The memory consumption and duration of compactions and column/node repair jobs is hard to model and accommodate. It's hard to tell what the cluster is doing at any given moment. Our experience with support plans from Datastax was also pretty dismal"; another user in the same thread: "At a much bigger company that had some very, very highly paid Cassandra DBAs it was actually a relatively smooth experience";
Source 1: 2026-09, thousand-node operator — "Repair is the tax on eventual consistency. You can pay it on a schedule you chose, or all at once with interest" — the operational spend is mandatory, not optional.
- 证据等级
`多方印证(3 个独立来源)`,HN 用户 ×2(含专职 DBA 对比)+ 千节点运维实录。
`Corroborated (3 independent sources)`, HN users ×2 (including the dedicated-DBA contrast) + thousand-node operator account.
Apache Cassandra / ScyllaDB 年份:2026
ClickHouse 自己把自己写满:system 日志表默认无 TTL,11.8GB 日志 vs 20MB 真实数据
多方印证
稳定与故障运维复杂度
- 一句话
`system.*_log` 默认无上限——uptimepage 磁盘 100% 写满、Postgres 崩溃,查下来真实数据仅 20MB,`system` 库 11.8 GiB(text_log 5.29G + trace_log 3.02G);Langfuse 自建曾攒出 59GB text_log。
`system.*_log` tables are unbounded by default — uptimepage's disk hit 100% and Postgres crashed; real data was 20MB while the `system` database held 11.8 GiB (text_log 5.29G + trace_log 3.02G). A self-hosted Langfuse once accumulated 59GB of text_log.
- 窄场景
自建小规格实例;Langfuse/SigNoz/ClickStack 等"开箱自带 ClickHouse"的部署。
Self-hosted small instances; "batteries-included ClickHouse" deployments like Langfuse/SigNoz/ClickStack.
- 机制
系统日志表默认无 TTL、无大小上限;小机器上空载还烧 CPU(关日志后 CPU 从 40-70% 降到 1.5%)。TRUNCATE 后磁盘从 74% 降到 41%。另有坑:`opentelemetry_span_log` 的 TTL 必须写在 `<engine>` 里否则服务起不来。
System log tables ship with no TTL and no size cap; on small boxes they even burn CPU at idle (disabling logging dropped CPU from 40–70% to 1.5%). TRUNCATE took disk from 74% to 41%. Bonus pit: `opentelemetry_span_log`'s TTL must be written inside `<engine>` or the service won't start.
- 生产验证
—
Source 13: uptimepage official blog postmortem, Jul 2026;
Source 14: dev.to, protemir, circa Sep 2026 — citing Langfuse discussion #15024 (59.2 GiB text_log), SigNoz #12050 (80GB of system logs vs <500MB real data), ClickStack PR #275 (10Gi volume filled in 10 days).
- 证据等级
`多方印证(3 个独立来源)`。
`Corroborated (3 independent sources)`.
ClickHouse 年份:2026
没有 MERGE 语句:Snowflake/BigQuery/PG15+ 都有,ClickHouse 没有
多方印证
生态与信任
- 一句话
Snowflake、BigQuery、PostgreSQL(15+)都有 `MERGE INTO`,ClickHouse 至今没有——官方自己的 Snowflake 迁移课程都单列"Gap 4: MERGE INTO":"ClickHouse has no `MERGE` statement."
Snowflake, BigQuery, and PostgreSQL (15+) all have `MERGE INTO`; ClickHouse doesn't — its own Snowflake migration course lists "Gap 4: MERGE INTO": "ClickHouse has no `MERGE` statement."
- 窄场景
CDC upsert、SCD2 增量;从 Snowflake 迁来的 ETL。
CDC upserts, SCD2 incrementals; ETL ported from Snowflake.
- 机制
替代方案都有代价:ReplacingMergeTree 的去重是异步的(merge 前新旧行并存,查询必须加 FINAL);dbt 的 `delete_insert` 是 delete+insert 两步、非真原子 upsert;第三方工具(Bruin)的 merge/scd2 增量策略在 ClickHouse 上直接标注不支持。
Every alternative costs: ReplacingMergeTree dedup is async (old and new rows coexist pre-merge; queries need FINAL); dbt's `delete_insert` is a two-step delete+insert, not a true atomic upsert; third-party tools (Bruin) mark their merge/scd2 incremental strategies unsupported on ClickHouse.
- 生产验证
—
Source 20: ClickHouse official Snowflake migration course (circa Aug 2026), "Gap 4: MERGE INTO";
Source 21: third-party ELT tool Bruin docs — "ClickHouse has no `MERGE INTO` statement, so Bruin implements the `merge` strategy with a delete+insert pattern";
Source 22: ORM NextORM capability matrix — ClickHouse key upsert and full MERGE both NotSupportedException ("no engine-level upsert").
- 证据等级
`多方印证(3 个独立来源)`,官方课程 + 两款第三方工具。
`Corroborated (3 independent sources)`, official course + two third-party tools.
ClickHouse 年份:2026
没有 PIVOT/UNPIVOT:标准语法缺失,50 个透视值就得手写 50 个表达式
多方印证
生态与信任
- 一句话
Snowflake/BigQuery/DuckDB/Oracle 都有标准 `PIVOT`/`UNPIVOT`,ClickHouse 没有——只能用 `sumIf`/`groupArray` 手工拼,透视值必须手写枚举,且没有动态列名能力。
Snowflake, BigQuery, DuckDB, and Oracle all have standard `PIVOT`/`UNPIVOT`; ClickHouse doesn't — you're left hand-assembling `sumIf`/`groupArray`, enumerating pivot values manually, with no dynamic column names.
- 窄场景
行转列报表、透视分析;从 Oracle/SQL Server 迁移的 SQL。
Crosstab reports, pivot analysis; SQL ported from Oracle/SQL Server.
- 机制
官方知识库自认:"ClickHouse doesn't have a PIVOT clause... ClickHouse has no pivot operator"。第三方基准项目为此跳过测试项:"Re-enable when: Standard `PIVOT`/`UNPIVOT` clause syntax appears in the ClickHouse SELECT reference"。
The official knowledge base admits it: "ClickHouse doesn't have a PIVOT clause... ClickHouse has no pivot operator." A third-party benchmark project skips test items over it: "Re-enable when: Standard `PIVOT`/`UNPIVOT` clause syntax appears in the ClickHouse SELECT reference."
- 生产验证
—
Source 23: official repo feature request (2023), with follow-ups — "Bumping as this is super useful and painful without";
Source 24: official knowledge base (circa Aug 2026) admits the missing operator.
- 证据等级
`多方印证(2 个独立来源)`,官方承认 + 用户请求。
`Corroborated (2 independent sources)`, official admission + user request.
ClickHouse 年份:2026
没有生产级多语句事务:迁移工具明示"部分迁移自己修"
多方印证
运维复杂度生态与信任
- 一句话
MergeTree 没有 PG/MySQL 式的多语句事务——golang-migrate 官方文档明示:"The queries are not executed in any sort of transaction/batch, meaning you are responsible for fixing partial migrations."
MergeTree has no PG/MySQL-style multi-statement transactions — the golang-migrate docs state it plainly: "The queries are not executed in any sort of transaction/batch, meaning you are responsible for fixing partial migrations."
- 窄场景
schema 迁移工具(golang-migrate 等);需要跨表原子提交的写入。
Schema migration tooling (golang-migrate et al.); writes needing cross-table atomic commits.
- 机制
无 `SELECT FOR UPDATE`、无跨表原子提交。官方文档承认有实验性事务,但限制极多:需 Keeper、仅限 Atomic 库、**仅限非复制 MergeTree**、Cloud 完全不支持——生产基本不可用。注意这与 DDL 原子性(Atomic database engine)是两回事。
No `SELECT FOR UPDATE`, no cross-table atomic commit. The docs admit experimental transactions, but the restrictions are severe: Keeper required, Atomic databases only, **non-replicated MergeTree only**, Cloud entirely unsupported — effectively unusable in production. (This is distinct from DDL atomicity via the Atomic database engine.)
- 生产验证
—
Source 25: mainstream open-source migration tool golang-migrate's official docs warning;
Official docs (updated Sep 2026): experimental transactions carry Experimental + CloudNotSupported badges; requirements state "Non-Replicated MergeTree table engine only."
- 证据等级
`多方印证(2 个独立来源)`,主流工具文档 + 官方文档限制。
`Corroborated (2 independent sources)`, mainstream tooling docs + official doc restrictions.
ClickHouse 年份:2026
MySQL wire 协议是"残血版":预编译查询至今不支持
多方印证
生态与信任
- 一句话
ClickHouse 的 MySQL 协议端口能连上,但官方文档自己承认"not guaranteed to be a drop-in replacement"——**prepared queries are not supported**,部分类型按字符串返回;PG 的 `mysql_fdw` 直接不兼容。
ClickHouse's MySQL protocol port connects — but the official docs admit it's "not guaranteed to be a drop-in replacement": **prepared queries are not supported**, some types come back as strings. PG's `mysql_fdw` simply doesn't work with it.
- 窄场景
从 MySQL 生态迁移、依赖预编译的工具链;想拿 MySQL 客户端直连的团队。
Migrating from the MySQL ecosystem with prepared-statement-dependent tooling; teams wanting to point a MySQL client straight at it.
- 机制
Restrictions 白纸黑字:"prepared queries are not supported"、"some data types are sent as strings"。依赖预编译的工具链直接不兼容,相关功能请求长期 open。
The Restrictions say it in black and white: "prepared queries are not supported," "some data types are sent as strings." Toolchains depending on prepared statements break; the feature request stays open.
- 生产验证
—
Source 27: official docs (updated Sep 2026), Restrictions;
The long-open user request aggregated on CodeTriage: "MySQL interface not compatible with Postgres' `mysql_fdw` foreign data wrapper."
- 证据等级
`多方印证(2 个独立来源)`,官方自认 + 用户 issue。
`Corroborated (2 independent sources)`, official self-admission + user issue.
ClickHouse 年份:2026
全文检索有了,但没有 BM25:ES/DuckDB 标配的相关性排序,ClickHouse 没有
多方印证
生态与信任
- 一句话
ClickHouse 的全文检索(倒排索引)26.2 才 GA,但 BM25 相关性排序至今缺失——ES、DuckDB 都是标配;需求强烈到催生出专门为此而生的 fork(MyScale)。
ClickHouse full-text search (inverted index) went GA in 26.2 — but BM25 relevance ranking is still missing, a standard on Elasticsearch and DuckDB; demand was strong enough to spawn a purpose-built fork (MyScale).
- 窄场景
在日志索引之上做搜索产品(而非纯日志分析);从 ES 迁移的团队。
Search products built on log indexes (not pure log analytics); teams migrating from ES.
- 机制
倒排索引本身 26.2 GA(此前 23.2→26.1 全程实验性),但"按词频赋予 token 相关性权重"的能力没有。Logchef 从 ES 迁移文档明示:ClickHouse 在其场景是"快速过滤而非排序检索"——"如果你在日志索引之上跑的是搜索产品,Elasticsearch 仍然是更合适的选择。"
The inverted index itself went GA in 26.2 (experimental all the way from 23.2 to 26.1), but "weighting tokens by term frequency" doesn't exist. Logchef's Elasticsearch migration docs put it plainly: ClickHouse there is "fast filtering, not ranked retrieval" — "if you're running a search product on top of a log index, Elasticsearch is still the better fit."
- 生产验证
—
Source 30: Logchef official "From Elasticsearch" migration docs, "What you lose" section;
The Tinybird comparison blog (cited by the official issue): "ClickHouse® lacks built-in relevance algorithms like BM25 that Elasticsearch uses for ranking search results.";
Source 31: official repo user request (circa Jan 2026).
- 证据等级
`多方印证(3 个独立来源)`。
`Corroborated (3 independent sources)`.
ClickHouse 年份:2026
按 vCPU 计费:小集群起步价就劝退
多方印证
成本账单
- 一句话
云上按 vCPU 小时、存储、备份、流量逐项计费——最小可用集群的月账单是单机 Postgres 的十倍量级起跳,"免费试用"之后就是真金白银。
The cloud meters vCPU-hours, storage, backups, and transfer separately — the smallest viable cluster bills at an order of magnitude above a single Postgres box, and the "free trial" ends in a real invoice.
- 窄场景
从 Serverless/Basic 免费档长大的小团队;需要多可用区/多 region 的生产集群;自托管但营收超 1000 万美元门槛的公司。
Small teams outgrowing the Serverless/Basic free tier; production clusters needing multi-AZ or multi-region; self-hosters above the $10M revenue threshold.
- 机制
分布式架构的成本是乘法:3 副本 × 节点数 × 单价。2024 年定价体系里 Advanced(原 Dedicated)$295/月/2vCPU(来源 7),3 节点最小集群 ≈ $885/月(6 vCPU)起;2026 年 9 月价目表:Standard $0.092/vCPU-hour(厂商估算 2vCPU+100GB ≈ $203/月),Mission Critical $0.162/vCPU-hour(12vCPU ≈ $1,528/月),存储/备份/流量另计(来源 6);自托管 Enterprise 按 $1,400–2,200/vCPU/年订阅(来源 14)。账单没有"小"字选项。
Distributed cost is multiplicative: 3 replicas x node count x unit price. In the 2024 pricing, Advanced (formerly Dedicated) was $295/month per 2 vCPUs (Source 7), so a 3-node minimum cluster started around $885/month (6 vCPUs). The Sep 2026 rate card: Standard at $0.092/vCPU-hour (vendor estimate ~$203/month for 2 vCPUs + 100 GB), Mission Critical at $0.162/vCPU-hour (~$1,528/month for 12 vCPUs), storage/backups/transfer billed on top (Source 6); self-hosted Enterprise is a $1,400–2,200/vCPU/year subscription (Source 14). There is no "small" line item.
- 生产验证
来源 6:独立评测 2026-09-23 对照厂商价目表核对——Standard 2vCPU+100GB 约 $203/月,Mission Critical 12vCPU 约 $1,528/月,且存储备份流量另算;对比 Supabase Pro $25/月、Neon 按用付费无月最低消费;
来源 7:2024-11 定价:Standard $146/月/2vCPU 起,Advanced $295/月/2vCPU 起(TechTarget 报道);
来源 14:2026-09 对比评测——自托管 Enterprise 授权 $1,400–2,200/vCPU/年。
Source 6: independent review checked against the vendor rate card on Sep 23, 2026 — Standard 2 vCPUs + 100 GB ~$203/month, Mission Critical 12 vCPUs ~$1,528/month, storage/backups/transfer extra; versus Supabase Pro at $25/month and Neon with usage billing and no monthly minimum;
Source 7: Nov 2024 pricing — Standard from $146/month per 2 vCPUs, Advanced from $295/month per 2 vCPUs (TechTarget);
Source 14: Sep 2026 comparison — self-hosted Enterprise licensed at $1,400–2,200/vCPU/year.
- 证据等级
`多方印证(3 个独立来源)`,独立评测/媒体 ×3(2024-2026)。
`Corroborated (3 independent sources)`, independent reviews/media x3 (2024–2026).
- 备注
来源 6 与来源 14 的 per-vCPU-hour 数字口径不一致($0.092/$0.162 vs $0.50/$0.95),卡片采用 2026-09-23 最新对照厂商价目表核对的来源 6 数字,差异如实标注。
Sources 6 and 14 disagree on per-vCPU-hour figures ($0.092/$0.162 vs $0.50/$0.95). The card uses Source 6's numbers (checked Sep 23, 2026 against the vendor pricing page); the discrepancy is stated as-is.
CockroachDB 年份:2026
免费档又缩水:Basic 的 $15/月免费额度下线了
多方印证
成本账单生态与信任
- 一句话
2024 年 Serverless 改名 Basic、保留 10GB+5000 万 RU/月免费档;2026 年 9 月定价页一改,Basic 从新用户页面消失——新组织只剩 30 天 $400 试用,之后直接出账单。
In 2024 Serverless was renamed Basic and kept a 10 GB + 50M RU/month free tier; in Sep 2026 a pricing-page redesign removed Basic from the new-customer page — new organizations get a 30-day $400 trial, then the meter runs.
- 窄场景
靠免费档跑 side project、做 PoC 的独立开发者与小团队;照着旧教程("免费 10GB+5000 万 RU")来注册的新用户。
Indie developers and small teams running side projects or PoCs on the free tier; new users signing up from old tutorials promising "free 10 GB + 50M RUs."
- 机制
免费档是获客漏斗,也是信任承诺。2024-11 的三档体系里 Basic 明确"从免费起步,超 10GB 存储和 5000 万 RU/月才收费"(来源 7);2026-09-15 随 Continuum 发布改版定价页:只剩 Standard 与 Mission Critical 两档,Basic($15/月免费额度)不在新页面上,老客户保留原计划,新组织只有 30 天 $400 试用(来源 6)。教程与现实脱节——旧版定价指南一夜之间过时。
The free tier is the acquisition funnel and a trust promise. The Nov 2024 three-tier lineup had Basic explicitly "starting at no cost, charges only past 10 GB storage and 50 million request units per month" (Source 7). With the Sep 15, 2026 Continuum launch, the pricing page listed only Standard and Mission Critical — Basic ($15/month free usage) was gone from the new-customer page; existing customers kept their plans, new organizations got only the 30-day $400 trial (Source 6). Tutorials went stale overnight.
- 生产验证
来源 7:2024-11 报道——Basic 从免费起步,超 10GB 存储 + 5000 万 RU/月后收费(TechTarget);
来源 6:2026-09 实测记录——2026-09-15 定价页改版后 Basic 不在新页面上,新组织无免费档,只有 30 天 $400 试用;
来源 14(2026-09-14,改版前一天)还写着 Basic 免费档含 10GB + 5000 万 RU/月——前后脚互相印证了这次收缩。
Source 7: Nov 2024 — Basic starts free, billing past 10 GB storage + 50M RU/month (TechTarget);
Source 6: Sep 2026 field check — after the Sep 15, 2026 pricing-page change, Basic is not on the new page; new organizations get no free tier, only a 30-day $400 trial;
Source 14 (Sep 14, 2026, the day before the change) still documented Basic's free tier at 10 GB + 50M RU/month — the before/after pair corroborates the contraction.
- 证据等级
`多方印证(3 个独立来源)`,独立评测/媒体 ×3(改版前后对照)。
`Corroborated (3 independent sources)`, independent reviews/media x3 (before/after the change).
CockroachDB 年份:2026
跨区写入的延迟税:region survival 是每笔写入在交税
多方印证
性能问题
- 一句话
想要"挂掉一个 region 还能写",每一笔健康日的写入都要先凑齐跨区 quorum——地理距离变成写延迟的下限,光速税天天交。
Want to survive losing a region and keep writing? Every write on every healthy day must first assemble a cross-region quorum — geography becomes the floor under write latency, a speed-of-light tax paid daily.
- 窄场景
SURVIVE REGION FAILURE 的多区集群;对写入 p99 有预算(如 <50ms)的 OLTP 业务;把"多活"当默认架构的团队。
Multi-region clusters with SURVIVE REGION FAILURE; OLTP workloads with a write p99 budget (e.g. under 50ms); teams treating "multi-active" as the default architecture.
- 机制
强一致 + 跨区存活 = 写入必须等远端副本确认。最近的必需副本决定了写延迟的下限,再怎么调优也越不过物理距离(来源 8 引述 Cockroach Labs 自己的说法:region survival 至少增加一个跨区 RTT 的写延迟)。这不是实现缺陷,是 CAP 的账单——只是这张账单是按笔收的,不是灾难日才收。
Strong consistency plus region survival means a write is not done until replicas in multiple regions agree. The nearest required replica sets the latency floor — no tuning beats physics (Source 8 quotes Cockroach Labs itself: region survival adds at least one cross-region round trip to write latency). This is not an implementation defect; it is the CAP bill — except it is collected per write, not on disaster day.
- 生产验证
来源 8:2026-08 独立分析——"CockroachDB charges consensus on the write path";quorum 需要远端响应,最近的必需副本给写延迟定下限;作者结论:问题不是"多区好不好",而是"你想在哪天为故障付费——每笔健康写入,还是灾难恢复时";
来源 9:2026 架构剖析——需要 Raft 共识的写入"比只需要落一个单机 WAL+fsync 的写入慢";range 跨区部署做存活,等于"在每笔写入 quorum 上支付光速往返成本"。
Source 8: Aug 2026 independent analysis — "CockroachDB charges consensus on the write path"; a quorum requires a remote response, and the nearest required replica places a floor under write latency. The author's conclusion: the question is not "is multi-region better" but "when do you want to pay for failure — during every healthy write, or during recovery";
Source 9: 2026 architecture teardown — a write needing Raft consensus across nodes, potentially across regions, is slower than one that only needs a single machine's WAL and fsync; spanning regions for survivability means "paying speed-of-light round-trip costs on every write quorum."
- 证据等级
`多方印证(2 个独立来源)`,独立工程分析 ×2(2026)。
`Corroborated (2 independent sources)`, independent engineering analyses x2 (2026).
CockroachDB 年份:2026
单区也慢:Raft 共识 vs 单机 fsync
多方印证
性能问题
- 一句话
不跨区也别指望跟单机 Postgres 一样快——单区集群每笔提交大约慢 1.5–3 倍,按 id 查单行、做聚合都慢,这是架构税不是 bug。
Don't expect single-node Postgres speed even without leaving the region — a single-region cluster commits roughly 1.5–3x slower per transaction; point lookups and aggregates are all slower. That is architecture tax, not a bug.
- 窄场景
单 region 部署、期望"Postgres 平替"的团队;对 p50 延迟敏感的 OLTP;拿单机 PG 做基准压测的选型 PoC。
Single-region deployments expecting a "Postgres replacement"; latency-sensitive OLTP; selection PoCs benchmarked against single-node Postgres.
- 机制
单机 PG 写 = 一次本地 WAL fsync;CockroachDB 写 = leaseholder 协调 + range 内多数派 Raft 确认,多一次网络往返。读强一致也要找 leaseholder。共识的固定开销与负载无关——集群越闲,单笔越显得贵。
Single-node Postgres writes with one local WAL fsync; CockroachDB writes with leaseholder coordination plus a Raft majority acknowledgment inside the range — an extra network round trip per commit. Strongly consistent reads must also reach the leaseholder. The consensus overhead is fixed regardless of load — the idler the cluster, the more expensive each write looks.
- 生产验证
来源 10:2026-08——"单区集群里,CockroachDB 每笔提交大约比单节点 Postgres 慢 1.5–3×(Raft 往返 vs 单次 fsync)";
来源 11:2022-11 HN 生产用户——"查询通常比 PostgreSQL 慢,不管是按 id 取一行还是做聚合;也许规模大到 Postgres 根本扛不住时表格会反转,我也不知道"。
Source 10: Aug 2026 — "in a single-region cluster, CockroachDB is roughly 1.5–3x slower than single-node Postgres per commit (Raft round-trip vs single fsync)";
Source 11: Nov 2022 HN production user — "queries will generally execute slower than PostgreSQL — whether we're talking about fetching a row by id, or doing an aggregate. Maybe the performance table flips at a huge scale where PostgreSQL can't even keep up, I don't know."
- 证据等级
`多方印证(2 个独立来源)`,2026 独立分析 + 2022 生产用户实测体感(旧源仅作印证,机制未变)。
`Corroborated (2 independent sources)`, 2026 independent analysis + 2022 production-user measurement (the older source corroborates only; the mechanism is unchanged).
CockroachDB 年份:2026
Postgres 兼容的窟窿:触发器等关键特性缺失
多方印证
生态与信任
- 一句话
线协议是 Postgres 的,`psql` 连上去毫无违和感——直到你发现触发器没有、LISTEN/NOTIFY 没有、domain/range 类型没有、FDW 没有,迁移到一半才发现"兼容"的是子集。
The wire protocol is Postgres and `psql` connects without complaint — until you discover there are no triggers, no LISTEN/NOTIFY, no domains, no range types, no foreign data wrappers, and you learn mid-migration that "compatible" meant a subset.
- 窄场景
从 Postgres 迁移、重度依赖触发器/扩展/自定义类型的老应用;用 Ecto 等 ORM 做迁移锁定的项目;全文检索用 tsvector 的团队。
Migrations from Postgres by applications heavy on triggers, extensions, or custom types; projects using Ecto-style migration locking; teams on tsvector full-text search.
- 机制
CockroachDB 是自研 SQL 层、不是复用 PG 查询层(与 YugabyteDB 路线不同),所以 PG 特性要逐个重新实现。2026 年 9 月的兼容性对照表仍列:LISTEN/NOTIFY 不支持、CREATE DOMAIN 不支持、range 类型不支持、FDW 不支持、列级权限不支持、表继承不支持、deferrable 约束不支持、DROP TRIGGER…CASCADE 不支持、tsvector 全文检索仅有限支持、PL/pgSQL 存储过程仅有限支持(来源 13)。注意:存储过程已在 23.2 重做补上(来源 15),不再列为缺失;但触发器从 2022 年生产用户抱怨至今仍无(来源 11、来源 10)。
CockroachDB built its own SQL layer rather than reusing Postgres's query layer (unlike YugabyteDB), so every Postgres feature must be reimplemented one by one. A Sep 2026 compatibility table still lists: LISTEN/NOTIFY unsupported, CREATE DOMAIN unsupported, range types unsupported, foreign data wrappers unsupported, column-level privileges unsupported, table inheritance unsupported, deferrable constraints unsupported, DROP TRIGGER ... CASCADE unsupported, tsvector full-text search only limited, PL/pgSQL stored procedures only limited (Source 13). Note: stored procedures were rebuilt and added in 23.2 (Source 15), so they are no longer listed as missing; triggers have been absent since production users complained in 2022 (Sources 11, 10).
- 生产验证
来源 11:2022-11 HN 生产用户——缺失最多的三样:扩展(逐个核对内置支持)、全文检索、触发器(以及更广义的 procedure/function);
来源 10:2026-08——"不支持:触发器、存储过程(PL/pgSQL)、LISTEN/NOTIFY、SEQUENCE、自定义类型;避开这五样的 Django/Rails 应用才能无改动运行";
来源 13:2026-09 兼容性对照表——触发器相关(UPDATE OF 列级触发器、DROP TRIGGER…CASCADE)不支持,另有 advisory lock 静默 no-op 等行为差异;
附带实录:
来源 12 的 Fly.io 用户发现 CockroachDB 不支持 `LOCK TABLE`,Ecto 迁移必须加 `migration_lock: false`(2021-07)。
Source 11: Nov 2022 HN production user — the three biggest missing features: extensions (each needs a built-in-support check), full-text search, triggers (and procedures/functions more broadly);
Source 10: Aug 2026 — "Not supported: triggers, stored procedures (PL/pgSQL), LISTEN/NOTIFY, SEQUENCE, custom types. A Django or Rails app that avoids those five constructs runs unchanged";
Source 13: Sep 2026 compatibility table — trigger-related gaps (column-list UPDATE triggers, DROP TRIGGER ... CASCADE) unsupported, plus behavioral differences like advisory locks as silent no-ops;
Field note: a
Source 12 Fly.io user found CockroachDB does not support `LOCK TABLE`, forcing Ecto migrations to set `migration_lock: false` (Jul 2021).
- 证据等级
`多方印证(4 个独立来源)`,HN 生产用户 + 独立分析 ×2 + 开源兼容性对照表。
`Corroborated (4 independent sources)`, HN production user + independent analyses x2 + open-source compatibility table.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。另:来源 10 称"存储过程不支持"与来源 15(DBTA 2024-01 报道 23.2 新增存储过程/UDF)矛盾,卡片以较新的版本事实为准,标注了口径分歧。
this card's this card's topic may overlap existing [Pitfall] cards on this site on this site. Also: Source 10's "stored procedures unsupported" contradicts Source 15 (DBTA, Jan 2024, reporting stored procedures/UDFs added in 23.2); the card follows the newer version fact and flags the disagreement.
CockroachDB 年份:2026
Serializable 默认 + 40001 重试:并发一上来应用层改代码
多方印证
性能问题运维复杂度
- 一句话
默认隔离级别是 Serializable——写冲突不排队,直接甩给你一个 40001 "restart transaction",重试逻辑、幂等、退避全得应用层自己写;PG 里跑得好好的高并发代码搬过来先加一圈 try/retry。
The default isolation level is Serializable — write conflicts don't queue, they hand you a 40001 "restart transaction." Retry loops, idempotency, and backoff all become application code; high-concurrency code that ran fine on Postgres needs a try/retry wrapper added everywhere.
- 窄场景
从 PG(默认 read committed)迁移的高并发 OLTP;事务里调外部接口(发邮件、扣款、发消息)的业务;热点行并发写的场景。
High-concurrency OLTP migrated from Postgres (default read committed); business logic calling external services inside transactions (email, charging, messaging); hot-row write contention.
- 机制
无锁并发控制:写冲突靠 MVCC 时间戳检测,失败的事务由数据库判负、应用重做。40001/40003/CockroachDB 特有的 "restart transaction" 都是"整个事务重来"的信号;重试必须带指数退避+抖动+次数上限,且 COMMIT 阶段也可能报错要重试(来源 13 的生产检查清单要求"所有写事务"实现重试)。更坑的是:重试前已发生的外部副作用(邮件已发、款已扣)数据库管不了,会重复(来源 8)。
Lock-free concurrency control: write conflicts are detected via MVCC timestamps, the database declares the loser, and the application redoes the work. 40001/40003 and CockroachDB's "restart transaction" all mean "replay the whole transaction"; retries need exponential backoff with jitter and a bounded attempt count, and even COMMIT can fail and need a retry (Source 13's production checklist requires retry logic for ALL write transactions). Worse: side effects that already happened before the retry — email sent, charge made — are outside the database's control and get duplicated (Source 8).
- 生产验证
来源 8:2026-08——"CockroachDB 默认 serializable;冲突时事务可能需要重试……如果事务在提交前发了邮件、调了外部扣款、发了消息,重试会复制副作用";
来源 13:2026-09 生产检查清单——"所有写事务实现重试逻辑;处理 40001 与 40003;COMMIT 的错误也要重试;指数退避+抖动;监控 40001 率,高即代表有 contention";
来源 9:2026——"Serializable、always" vs Postgres "Read Committed";"Transaction retries under contention" 列在营销页不写的 trade-off 清单里。
Source 8: Aug 2026 — "CockroachDB uses serializable isolation by default. Under contention, transactions may need retries... If a transaction sends an email, charges an external provider, or publishes a message before commit, retrying it can duplicate side effects";
Source 13: Sep 2026 production checklist — "transaction retry logic implemented for ALL write transactions; retry handles 40001 AND 40003; retry handles errors from COMMIT; exponential backoff with jitter; alert on high 40001 error rates (indicates contention)";
Source 9: 2026 — "Serializable, always" versus Postgres "Read Committed"; "transaction retries under contention" listed among the trade-offs absent from the marketing page.
- 证据等级
`多方印证(3 个独立来源)`,独立工程分析 ×2 + 开源生产检查清单。
`Corroborated (3 independent sources)`, independent engineering analyses x2 + open-source production checklist.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。另:23.2 已提供 read committed preview,迁移 PG 高并发应用可少写重试错误处理(来源 15)——属"已缓解于 23.2",未彻底修复(仍为 preview)。
this card's this card's topic may overlap existing [Pitfall] cards on this site on this site. Also: 23.2 added preview read-committed isolation so high-concurrency Postgres migrations need less retry error handling (Source 15) — mitigated in 23.2, not fully fixed (still preview).
CockroachDB 年份:2026
时钟即正确性:NTP 配错是真实故障模式
多方印证
稳定与故障
- 一句话
Postgres 不在乎服务器时钟漂多少;CockroachDB 的正确性压在"时钟偏移有界"上——NTP 没配好,节点直接拒绝启动,事务莫名重启。
Postgres does not care how far your server clocks drift; CockroachDB's correctness rests on bounded clock skew — with NTP misconfigured, nodes refuse to start and transactions restart for no apparent reason.
- 窄场景
自建机房/混合云、NTP 没进基线配置的集群;跨区部署但时钟源不一致的环境;虚拟机时钟漂移大的平台。
Self-built data centers or hybrid clouds where NTP is not in the baseline config; cross-region clusters with inconsistent time sources; virtualized platforms with heavy clock drift.
- 机制
HLC(混合逻辑时钟)+ 不确定性区间:默认要求节点间时钟偏移 ≤500ms。超限的节点启动时直接报错退出("clock synchronization error: this node is more than 500ms away…",exit code 7);运行中读到落在不确定性区间内的时间戳,事务被强制重启。正确性依赖项里多了一项"全集群时钟同步",这是 PG 架构里不存在的故障面。
Hybrid logical clocks plus uncertainty intervals: nodes must stay within ~500ms of each other by default. An out-of-skew node fails at startup ("clock synchronization error: this node is more than 500ms away...", exit code 7); at runtime, a read whose timestamp falls inside the uncertainty window forces a transaction restart. "All cluster clocks agree" becomes a baseline operational requirement — a failure surface that does not exist in Postgres's architecture.
- 生产验证
来源 12:2021-07 Fly.io 社区实录——用户两节点(ams/lhr)集群节点反复以 exit code 7 退出,日志里就是 "clock synchronization error: this node is more than 500ms away from at least half of the known nodes";排查发现是几台欧洲宿主机没跑 NTP,修好即恢复;
来源 9:2026 架构剖析——"Postgres 不在乎服务器时钟漂移;CockroachDB 的正确性保证依赖有界时钟偏移——配错的 NTP 是真实存在、虽罕见、但 Postgres 没有对应物的故障模式"。
Source 12: Jul 2021 Fly.io community account — a user's two-node (ams/lhr) cluster kept crash-looping with exit code 7 and the log line "clock synchronization error: this node is more than 500ms away from at least half of the known nodes"; the root cause was NTP not running on several European hosts; fixed and the cluster recovered;
Source 9: 2026 architecture teardown — "Postgres doesn't care if your server clocks drift. CockroachDB's correctness guarantees lean on bounded clock skew — badly configured NTP is a real, if rare, failure mode that has no Postgres equivalent."
- 证据等级
`多方印证(2 个独立来源)`,2021 用户故障实录 + 2026 机制分析(旧源仅作机制印证,约束至今未变)。
`Corroborated (2 independent sources)`, 2021 user incident account + 2026 mechanism analysis (the older source corroborates the mechanism only; the constraint is unchanged).
CockroachDB 年份:2026
Serverless 是个黑盒:扩到更大反而更慢,账单方差 9 倍
多方印证
性能问题成本账单
- 一句话
Serverless 仓库扩到 Large 以上不再提速只涨价,相同 workload 成本能差 9 倍——而你看不到任何实例细节。
Serverless warehouses stop getting faster past Large and only get pricier; identical workloads can vary 9x in cost — and you see none of the instance details.
- 窄场景
Serverless SQL Warehouse + dbt 并发;按"越大越快"直觉选型的团队。
Serverless SQL Warehouses + concurrent dbt; teams sizing by the "bigger is faster" intuition.
- 机制
Serverless 计算全托管,实例选型与底层变更对用户不可见;独立实验显示 Medium 之后扩容收益归零(Amdahl 定律),成本却随规格线性上涨。实验者进一步指出厂商的收入激励是"降自身成本、保持客户 runtime 不变",加速技术未必转化为客户账单下降。
Serverless compute is fully managed — instance selection and under-the-hood changes are invisible. An independent experiment showed scaling benefits hitting zero past Medium (Amdahl's law) while cost rises linearly with size. The experimenters further note the vendor's revenue incentive is to "lower its own costs while keeping customer runtimes about the same" — acceleration technology does not necessarily reach the customer's bill.
- 生产验证
—
Source 3: Towards Data Science, 2024 — Jeff Chou and Stewart Bryson spent $12K benchmarking TPC-DI: the Medium warehouse beat larger sizes on both cost and duration ("we have no clue why"); identical workloads ranged $5–$45 in cost and 2–90 minutes in runtime;
Source 4: Madhukar, Jul 2026 — Serverless ships with Photon on by default and no Spark UI to observe, making it a poor fit for routine long-running jobs.
- 证据等级
`多方印证(2 个独立来源)`,独立基准实验 + 个人实测(来源 3 作者供职于 Spark 成本优化厂商 Sync,雇主身份已在原文披露)。
`Corroborated (2 independent sources)`, independent benchmark experiment + personal measurement (source 3 authors work for Spark cost-optimization vendor Sync; employer affiliation disclosed in the article).
Databricks 年份:2026
Photon:提速 2 倍、账单也涨 2 倍的"加速"
多方印证
性能问题成本账单
- 一句话
Photon 常被宣传 2x+ 提速,但它的 DBU 费率也高 2–3 倍——没把 runtime 真砍半,等于加钱没提速;Python UDF 还会悄悄绕过它。
Photon is marketed at 2x+ speedups, but its DBU rate is also 2–3x higher — unless it truly halves your runtime, you pay more for no gain; Python UDFs silently bypass it.
- 窄场景
PySpark 为主、重度用 Python UDF、轻量短任务的作业;默认勾选 Photon 的团队。
PySpark-heavy jobs, heavy Python UDF use, lightweight short tasks; teams that leave Photon checked on by default.
- 机制
Photon 是向量化 C++ 引擎,job compute 上费率约为标准计算 2.9 倍(all-purpose 上约 2 倍);提速收益高度 workload 相关——重 join/窗口函数受益,Python UDF 会静默回退到非 Photon 路径,轻量任务还有引擎启动开销。TDS 实验者指出厂商曾用"2x 提速配 2x DBU 涨价"实现收入中性。
Photon is a vectorized C++ engine priced ~2.9x standard compute on job clusters (~2x on all-purpose). Speedup is highly workload-dependent — heavy joins/window functions benefit, Python UDFs silently fall back to the non-Photon path, and lightweight jobs pay engine startup overhead. The TDS experimenters note the vendor once achieved revenue neutrality via "2x faster paired with 2x DBU price increase."
- 生产验证
—
Source 1: AT&T Israel, Dec 2025 — rule of thumb: only use Photon if it more than halves the job's runtime, otherwise "pay more for minimal speed gains";
Source 4: Madhukar, Jul 2026 — Photon performs poorly on UDF jobs; at the 2.9x job-compute rate, negative net benefit is common;
Source 3: TDS 2024 $12K experiment — "Jobs clusters may be cheapest"; Serverless (Photon-included) acceleration does not necessarily reach the bill.
- 证据等级
`多方印证(3 个独立来源)`,具名团队博客 + 个人实测 + 独立基准实验。
`Corroborated (3 independent sources)`, named team blog + personal measurement + independent benchmark experiment.
Databricks 年份:2026
Unity Catalog 迁移:迁的不是数据,是全部访问模式
多方印证
运维复杂度升级迁移
- 一句话
上 Unity Catalog 不是"打开开关",而是全仓库扫描替换 mount point、重构权限模型、接受 shared 集群的功能阉割——还要先搞清哪些配额是硬上限。
Adopting Unity Catalog is not "flipping a switch" — it means repo-wide scans to replace mount points, rebuilding the permission model, accepting neutered shared clusters — and first figuring out which quotas are hard limits.
- 窄场景
历史包袱重的 Azure Databricks 老工作区(大量 mount point、DBFS 依赖);多订阅、多项目的大型组织。
Legacy Azure Databricks workspaces heavy on mount points and DBFS; large organizations with many subscriptions/projects.
- 机制
mount point 绕过 UC 安全控制,必须全量替换为三级命名或路径;因代码风格多样,全自动替换困难。UC 下 shared(Standard)集群不支持 ML runtime 与 RDD API,装库/init 脚本需 allowlist + metastore admin 特权。同时 UC 配额分"软上限可提"与"硬上限"(identity 类多为硬上限,storage credential 默认仅 200),设计前不摸清会直接卡死。
Mount points bypass UC security and must be fully replaced with three-level names or paths; diverse coding styles make full automation hard. Under UC, shared (Standard) clusters lose ML runtimes and RDD APIs, and installing libraries/init scripts requires an allowlist plus metastore-admin privileges. UC quotas split into "soft, raisable" and "hard" (identity-related ones are mostly hard; storage credentials default to just 200) — design blind and the migration stalls.
- 生产验证
—
Source 5: Valcon consulting, Karlo Kotarac, Jan 2026 — two client migrations: Client A did a full repo scan replacing mount points and retreated to small dedicated clusters for dev under shared-cluster limits; Client B could not fully use the official UCX migration tooling because hive_metastore had never been used;
Source 6: Adam Marczak, Dec 2025 migration series — raised the storage-credential quota from the default 200 to 10,000; identity-class hard limits cannot be raised; "workspace admins are not UC admins."
- 证据等级
`多方印证(2 个独立来源)`,咨询公司双客户实录 + 独立工程师迁移系列。
`Corroborated (2 independent sources)`, consulting firm's two-client field notes + independent engineer's migration series.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
topic may overlap the existing [Pitfall] card.
Databricks 年份:2026
"不如 Snowflake 好伺候":扩缩与调优的体感差距
多方印证
运维复杂度生态与信任
- 一句话
多位用户直言 Databricks"不如 Snowflake 好伺候"——扩缩资源慢半拍、调优要人肉,容错度低。
Multiple users put it bluntly — Databricks is "not as forgiving as Snowflake": scaling lags, tuning is manual, and the margin for error is thin.
- 窄场景
从 Snowflake 对比选型/双跑的团队;缺专职数据平台工程师的小团队。
Teams evaluating or dual-running against Snowflake; small teams without dedicated data-platform engineers.
- 机制
Databricks 把集群选型、扩缩、调优暴露给用户(实例家族、spot、Photon、warehouse 类型),灵活性高但心智负担重;Snowflake 把计算抽象成仓库,启停扩缩对用户近乎无感。PeerSpot 用户原话点出 "not as forgiving as Snowflake"。
Databricks exposes cluster choice, scaling, and tuning to the user (instance families, spot, Photon, warehouse types) — flexible but heavy cognitive load; Snowflake abstracts compute into warehouses where start/stop/scale feels instantaneous. The PeerSpot quote names the gap directly.
- 生产验证
—
Source 7a: PeerSpot — Slawomir Zablocki (Data Platform Architect, Jul 2024): "The product could be improved regarding the delay when switching to higher-performing virtual machines compared to other platforms like Snowflake. The ease and speed of managing clusters can also be enhanced" (direct quote);
Source 7b: PeerSpot — Simon Robinson (Governance And Engagement Lead, Jan 2026): "it's not as forgiving as other platforms such as Snowflake" (direct quote); warehouses left running directly inflate cost.
- 证据等级
`多方印证(2 个独立来源)`,具名用户复评 ×2。
`Corroborated (2 independent sources)`, named user reviews x2.
Databricks 年份:2026
foreachBatch 只是 at-least-once:重启就重复,幂等得自己写
多方印证
稳定与故障运维复杂度
- 一句话
Structured Streaming 的 foreachBatch 看起来像"每批处理一次",实际只保证 at-least-once——job 重启后同一 batch_id 会被再调一次,不幂等的写入就写重。
Structured Streaming's foreachBatch looks like "process each batch once" but only guarantees at-least-once — a job restart re-invokes the same batch_id, and non-idempotent writes double-write.
- 窄场景
foreachBatch 里做 DeltaTable.merge 或写外部系统;有状态算子的流作业。
DeltaTable.merge or external-system writes inside foreachBatch; stateful streaming jobs.
- 机制
API 的简洁掩盖了运维复杂度:checkpoint 只能保证 at-least-once;Delta 的 txnAppId+txnVersion 只保证 append/save 幂等,**不包裹** DeltaTable.merge(...);有状态算子后必须消费完整 batch,否则状态膨胀/过期驱逐会拖住 micro-batch。"Treat every handler as if Spark will call it again with the same batch_id."
The API's simplicity hides operational complexity: checkpoints guarantee at-least-once only; Delta's txnAppId+txnVersion make append/save idempotent but do **not** wrap DeltaTable.merge(...); after stateful operators you must consume the full batch or state bloat/expiry stalls the micro-batch. "Treat every handler as if Spark will call it again with the same batch_id."
- 生产验证
—
Source 15: DEV community, firfircelik, circa Oct 2026 — a restart re-invoked the same batch_id; a green micro-batch in the UI doesn't mean the retry was a no-op;
Source 16: Towards Data Engineering, Apr 2026 — "The API simplicity hides real operational complexity." Checkpoints alone only get you at-least-once.
- 证据等级
`多方印证(2 个独立来源)`,两篇独立从业者生产复盘。
`Corroborated (2 independent sources)`, two independent practitioner postmortems.
Databricks 年份:2026
VNet Injection 合规之路"painful":子网规划失误=全量重建,且不可逆
多方印证
运维复杂度生态与信任
- 一句话
VNet Injection 是唯一让 InfoSec 点头的组网方案,但"although painful"——subnet CIDR 规划错了只能全量重建,已有 workspace 事后无法切换。
VNet Injection is the only networking setup InfoSec actually approves, but "although painful" — get the subnet CIDR wrong and it's a full rebuild; existing workspaces can't switch after the fact.
- 窄场景
金融/医疗等强合规行业 Azure 上云;初期图省事用默认网络、后期被审计要求整改的团队。
Regulated industries (finance/healthcare) on Azure; teams that started on default networking and get flagged by audit later.
- 机制
网络配置不可逆:已有 workspace 无法事后注入 VNet;子网扩容/迁移至今不是自助操作,要维护窗口 + 支持团队协调,Terraform 尚不支持;出站流量全走防火墙例外,复杂度外溢到整个 landing zone。
The networking choice is irreversible: existing workspaces can't adopt VNet Injection retroactively; subnet expansion/migration is still not self-serve (maintenance window + support coordination, Terraform unsupported); all egress goes through firewall exceptions, spilling complexity across the landing zone.
- 生产验证
—
Source 26: Advancing Analytics (independent UK consultancy), Apr 2022 field blog — "although painful";
Source 27: official community thread (2026 update) — subnet changes still require manual coordination, not self-serve.
- 证据等级
`多方印证(2 个独立来源)`,独立咨询公司实战 + 2026 年社区最新确认。
`Corroborated (2 independent sources)`, independent consultancy field notes + 2026 community confirmation.
Databricks 年份:2026
共享仓库"钱算不清":语句级成本归因缺失多年,Query Tags 2026-09 才 GA
多方印证
成本账单
- 一句话
`system.billing.usage` 没有 query_id,join 不上 query history——共享仓库里"谁花的钱"长期只能事后反查,官方 Query Tags 2026-09 才 GA 补上打标一环。
`system.billing.usage` has no query_id and can't join query history — "who spent the money" on a shared warehouse was answerable only by after-the-fact forensics; official Query Tags went GA in Sep 2026 to close half the gap.
- 窄场景
多团队共享 warehouse 做 FinOps 分账;从"一 team 一仓库"合并降本后丢了可见性的团队。
FinOps chargeback across teams sharing warehouses; teams that consolidated from "one warehouse per team" and lost visibility.
- 机制
账单表粒度是 compute-hour,无 query_id/statement_id;实测 30 天窗口 89,264 行 usage 里 84.4% 完全无 custom_tags、usage_metadata 表名字段 0 行有值。用户只能自建 6 层归因管线。Query Tags(2026-02 preview → 2026-09 GA)补上"打标",但账单表本身的粒度问题在 2026-09 实测中依然存在。
The billing table is metered at compute-hour granularity with no query_id/statement_id; a measured 30-day window showed 84.4% of 89,264 usage rows with zero custom_tags and the usage_metadata table-name field empty on all rows. Users built 6-layer attribution pipelines themselves. Query Tags (preview Feb 2026 → GA Sep 2026) added tagging, but the billing table's granularity problem persisted in Sep 2026 measurements.
- 生产验证
—
Source 11: Ke Zhu, Mar 2026 — "The only way to attribute cost was to reverse-engineer it from query history after the money was already spent.";
Source 30: Pulse2, Jun 2026 — named customer ASOS: "With Query Tags we can finally accurately split up warehouse costs by the teams that are running dbt on it." Unit21: "We moved from one warehouse per team to shared warehouses to cut costs, but lost visibility into which team was driving spend.";
Source 31: official community thread, Sep 2026 — a user-built 6-layer attribution pipeline;
Source 32: althrussell governance field guide, Sep 2026 production measurements.
- 证据等级
`多方印证(4 个独立来源)`。
`Corroborated (4 independent sources)`.
Databricks 年份:2026
预算只有"告警"没有"熔断":跑飞的仓库,原生机制拦不住
多方印证
运维复杂度成本账单
- 一句话
账户 Budgets 对计算支出只有邮件告警(延迟可达 24 小时),"Block usage" 熔断仅限 Genie/AI Gateway——$14k 周末跑飞事件里,用户只能自建 Lambda 轮询 Billing API,连"及时发现"都做不到。
Account Budgets on compute spend send email alerts only (up to 24h delay); "Block usage" circuit-breaking exists solely for Genie/AI Gateway — in the $14k weekend runaway, the user could only build a Lambda polling the Billing API; even "timely discovery" wasn't native.
- 窄场景
Serverless 仓库 + BI 长连接;没有专人盯账单的团队。
Serverless warehouses with long-lived BI connections; teams without someone watching the bill.
- 机制
官方文档明示:"Per-user overrides and usage blocking is only available for Genie budgets.";"There could be up to a 24-hour delay between usage occurring and an email notification being sent." 仓库级语句超时(防"跑飞查询")2026-07 才进 Beta——此前连平台级熔断都没有。
Official docs: "Per-user overrides and usage blocking is only available for Genie budgets." / "There could be up to a 24-hour delay between usage occurring and an email notification being sent." Warehouse-level statement timeouts (against runaway queries) only entered Beta in Jul 2026 — before that, no platform-level tripwire existed at all.
- 生产验证
—
Source 10: aniketsoni, Sep 2026 $14k case — afterward built a Lambda polling the Billing API every 6 hours for variance alerts: "While this didn't stop the spending, it allowed me to isolate the DBU consumption…within minutes, rather than waiting for the bill to aggregate.";
Source 33: official docs confirm the mechanism;
Source 7b: PeerSpot user Simon Robinson (Jan 2026), cited on this site's existing card — "it's not as forgiving as other platforms such as Snowflake," cross-confirming the same pain.
- 证据等级
`多方印证(3 个独立来源)`,用户实案 + 官方文档 + 站内卡交叉。
`Corroborated (3 independent sources)`, user case + official docs + cross-reference.
Databricks 年份:2026
Delta Sharing 分享不了带行级安全/列掩码的表:想按 recipient 过滤?先人肉建 view
多方印证
运维复杂度生态与信任
- 一句话
UC 表一旦加上 row filter / column mask,就不能作为 Delta Sharing provider 资产分享——"这个 recipient 只能看 IN 区客户行"这种需求,得为每个 recipient 预建 view 再分享。
Put a row filter / column mask on a UC table and it can't be shared as a Delta Sharing provider asset — "this recipient only sees customer_country='IN' rows" requires a pre-built view per recipient.
- 窄场景
跨组织数据共享 + 行级权限;用 Delta Sharing 做数据产品的团队。
Cross-org data sharing with row-level entitlements; teams productizing data via Delta Sharing.
- 机制
provider 端直接拒绝带 RLS/列掩码的表(报错原文 "InvalidParameterValue: Table <fqn> has row level security or column masks, which is not supported by Delta Sharing.")。行级安全是"粗粒度"的:只能按分区过滤,不能按 recipient 动态过滤行。
The provider side flatly rejects tables with RLS/column masks (original error: "InvalidParameterValue: Table <fqn> has row level security or column masks, which is not supported by Delta Sharing."). Row-level security is coarse: partition-level filtering only, no per-recipient dynamic row filtering.
- 生产验证
—
Source 37: Manish Pansari, Apr 2026 — "you can't easily say 'this recipient sees only rows where customer_country = 'IN'' without pre-creating a view";
Source 38: a Databricks employee's UC migration tool README — provider rejects the tables; the tool falls back to "skip + audit";
Official docs, "Row filters and column masks" Limitations (verified 2026-10-07): "Delta Sharing providers cannot share tables with row-level security or column masks."
- 证据等级
`多方印证(3 个独立来源)`,独立博客 + 生产工具文档 + 官方文档。
`Corroborated (3 independent sources)`, independent blog + production tooling docs + official docs.
Databricks 年份:2026
UC 行列级安全是"孤岛治理":上了 RLS/掩码,就告别 13 项能力
多方印证
运维复杂度生态与信任
- 一句话
行过滤器/列掩码一旦启用,官方文档 Limitations 清单列出 13 项同时失效:time travel、deep/shallow clone、view、streaming 读取、Iceberg REST API、Delta Sharing 分享……安全越细,能力越少。
Enable row filters/column masks and the official docs' Limitations list kills 13 things at once: time travel, deep/shallow clone, views, streaming reads, Iceberg REST APIs, Delta Sharing… the finer the security, the fewer the capabilities.
- 窄场景
从 Snowflake 迁来、指望"行策略+clone+time travel 正交组合"的治理团队;要给测试环境 clone 生产数据的团队。
Governance teams migrating from Snowflake expecting row policies to compose orthogonally with time travel, clone, and secure views; teams cloning prod data into test.
- 机制
官方文档逐条在列:"Time travel does not work with row-level security or column masks." / "Deep and shallow clones are not supported…" / "You cannot apply row-level security or column masks to a view." / "You cannot use Iceberg REST catalog or Unity REST APIs…"。用户原声:上了行过滤器就得放弃用 clone 拷生产数据到测试环境,"will probably discourage us to use Row Filters"。
The docs list it item by item: "Time travel does not work with row-level security or column masks." / "Deep and shallow clones are not supported…" / "You cannot apply row-level security or column masks to a view." / "You cannot use Iceberg REST catalog or Unity REST APIs…" User voice: adopting row filters meant giving up cloning prod into test — "will probably discourage us to use Row Filters."
- 生产验证
—
Source 40: community user Antoine_B, Aug 2024 — filed Idea DBE-I-1461 for clone support; staff confirmed it's not on the roadmap; the docs still list the limitation in Oct 2026;
Source 39: bhavink's governance field repo gotchas table;
Source 32's repo troubleshooting table: time-travel on a row-filtered table throws INVALID_PARAMETER_VALUE.
- 证据等级
`多方印证(3 个独立来源)`,官方文档 + 独立仓库 + 社区实案。
`Corroborated (3 independent sources)`, official docs + independent repos + community case.
Databricks 年份:2026
Ranger 线程泄漏:每次元数据 checkpoint 漏两个线程,拖垮 Ranger Admin
多方印证
稳定与故障生态与信任
- 一句话
每次元数据 checkpoint 都会永久残留两个 Ranger PolicyRefresher 线程,线程数随 checkpoint 次数线性增长,最终把 Ranger Admin 的请求打爆、拖到 OOM。
Every metadata checkpoint permanently leaves behind two Ranger PolicyRefresher threads; the thread count grows linearly with checkpoint count until Ranger Admin is request-flooded into OOM.
- 窄场景
云模式(存算分离)+ 接入 Apache Ranger 做权限管控的 Doris 4.1.0 集群;默认 1 小时一次 cloud checkpoint。
Doris 4.1.0 clusters in cloud (compute-storage decoupled) mode with Apache Ranger authorization; default hourly cloud checkpoint.
- 机制
checkpoint 会新建一个临时 Env 做元数据镜像,Env 初始化时经 RangerDorisAccessControllerFactory 新建 Ranger 权限控制器,每个控制器启动 PolicyRefresher 后台线程;checkpoint Env 销毁时只清掉了静态引用,没有关闭 AccessControllerManager、没有停掉 PolicyRefresher 线程。于是每次 checkpoint 永久残留 2 个线程;每个线程按默认 30 秒轮询向 Ranger Admin 拉取策略与角色,请求量随运行时间线性增长。2.1 时期曾有人报过同类问题并修成单例(PR #45645),4.1 重构引入 Factory 后回归。
Each checkpoint creates a temporary Env to build the metadata image; Env initialization creates a new Ranger access controller via RangerDorisAccessControllerFactory, and each controller spawns PolicyRefresher background threads. Destroying the checkpoint Env only clears the static reference — it never closes the AccessControllerManager or stops the PolicyRefresher threads. So every checkpoint leaks exactly 2 threads, each polling Ranger Admin for policies/roles on the default 30-second interval, so request volume grows linearly with uptime. The same symptom was reported in the 2.1 era and fixed with a singleton (PR #45645); the 4.1 refactor introduced the factory and regressed it.
- 生产验证
来源 1:apache/doris#65524(2026-07),Doris 4.1.0 云模式 + Ranger 2.7.0——生产 FE 攒下 1321 个 PolicyRefresher 线程,按默认 30 秒轮询约 88 req/s 打到 Ranger Admin;Ranger Admin 积压 2000 万+ 异步任务、Full GC 5197 次累计约 18 小时、最终 `OutOfMemoryError: GC overhead limit exceeded`,Tomcat Acceptor/Poller 线程消失、6080 端口看似监听实则不可访问。报告给出完整复现步骤与根因调用链,作者表示愿意提交 PR;
来源 2:apache/doris#45641(2024,Doris 2.1 时期)——同一症状的早期报告:"每次 checkout 操作都新建 RangerDorisAccessController,每个控制器再新建 ranger policy refresher","Too many policy refreshers will cause ranger admin overload",作者同样表示愿意提交 PR(后由 PR #45645 修成单例)。
Source 1: apache/doris#65524 (Jul 2026), Doris 4.1.0 cloud mode + Ranger 2.7.0 — a production FE accumulated 1,321 PolicyRefresher threads, generating ~88 req/s against Ranger Admin; Ranger Admin accumulated 20M+ async tasks, 5,197 Full GCs totaling ~18 hours, and died with `OutOfMemoryError: GC overhead limit exceeded`; Tomcat acceptor/poller threads vanished while port 6080 stayed in LISTEN, appearing alive but unreachable. The report includes full reproduction steps and the root-cause call chain; the author offered to submit a PR;
Source 2: apache/doris#45641 (2024, Doris 2.1 era) — an earlier report of the same symptom: "doris create a new RangerDorisAccessController after every checkout operation, and every RangerDorisAccessController create a new ranger policy refresher," and "too many policy refreshers will cause ranger admin overload"; the author likewise offered a PR (later fixed as a singleton by PR #45645).
- 证据等级
`多方印证(2 个独立来源)`,GitHub 用户生产复盘/缺陷报告 ×2(#65524 具名复现步骤 + #45641 同症状早期报告)。
`Corroborated (2 independent sources)`, GitHub user production postmortem/bug reports x2 (named reproduction steps in #65524 + earlier same-symptom report in #45641).
Apache Doris 年份:2026
Tablet 版本链断裂:一次 BE 硬杀,tablet 永久不可读写
多方印证
稳定与故障
- 一句话
BE 在版本发布中途被硬杀(或磁盘迁移与 compaction 竞态),tablet 元数据的 rowset 版本链出现"洞",tablet 永久不可读写,连 `ADMIN SET REPLICA VERSION` 都修不好。
A BE hard-killed mid-publish (or a disk migration racing compaction) leaves a hole in the tablet's rowset version chain; the tablet becomes permanently unreadable and unwritable, and not even `ADMIN SET REPLICA VERSION` can repair it.
- 窄场景
K8s 部署下 BE pod 被 kubelet 硬杀(OOM/节点故障),或 BE 多磁盘间 tablet 迁移与 compaction 并发;副本数=1 时直接等价于数据丢失。
K8s deployments where the kubelet hard-kills BE pods (OOM/node failure), or tablet migration across BE disks racing with compaction; with replication factor 1 this is equivalent to data loss.
- 机制
Doris 的 tablet 数据由一串 rowset 版本组成版本图(version graph),读/写需要一条从版本 0 到可见版本的连续路径。publish 阶段的 tablet meta 更新不是崩溃原子的:硬杀落在中间状态,重启后版本图出现断档 → `fail to find path in version_graph`。FE 侧的 `ADMIN SET REPLICA VERSION` 只改 FE 侧状态,BE 下一次 tablet report 会用磁盘真实状态覆盖回来,所以"修了跟没修一样";BE 侧没有任何 salvage/repair 路径。RF1 下没有健康副本可克隆,唯一出路是删表从上游重建。
A Doris tablet's data is a chain of rowset versions forming a version graph; reads/writes need a continuous path from version 0 to the visible version. The tablet meta update during publish is not crash-atomic: a hard kill landing in the middle leaves a gap after restart → `fail to find path in version_graph`. The FE-side `ADMIN SET REPLICA VERSION` only changes FE-side state, which the BE's next tablet report overwrites with the true on-disk state — so "the repair does not stick"; and there is no BE-side salvage/repair path. With RF1 there is no healthy peer to clone from, so the only way out is dropping the table and reloading from upstream.
- 生产验证
来源 1:apache/doris#66301(2026),Doris 4.1.0-rc03 on K8s + Ceph——kubelet 硬杀 BE 落在 publish 中途,tablet 报 `fail to find path in version_graph. spec_version: 0-51946`,读写与 compaction 全部永久失败;`ADMIN SET REPLICA VERSION` 被 BE 上报覆盖;RF1 下只能删表从上游重建。报告人列出 #36832/#42021/#49524 三个同错误签名的历史 issue,横跨 2.0.x–4.1.x,均被 stale 关闭、无根因无修复;
来源 2:apache/doris#49524(2025-03),Doris 2.0.13——磁盘迁移(SSD→HDD)任务与 cumulative compaction 并发,迁移只搬了合并后的版本文件,tablet 报同样的 `fail to find path in version_graph`,读取失败。
Source 1: apache/doris#66301 (2026), Doris 4.1.0-rc03 on K8s + Ceph — the kubelet hard-killed a BE mid-publish; the tablet failed with `fail to find path in version_graph. spec_version: 0-51946`, and all reads, writes, and compaction on it failed permanently; `ADMIN SET REPLICA VERSION` was reverted by BE re-reports; with RF1 the only fix was dropping the table and rebuilding from upstream. The reporter cites three earlier issues with the same error signature (#36832/#42021/#49524), spanning versions 2.0.x–4.1.x, all closed stale with no root cause and no fix;
Source 2: apache/doris#49524 (Mar 2025), Doris 2.0.13 — a disk-migration (SSD→HDD) task raced with cumulative compaction; the migration only moved the post-merge version files, and the tablet failed reads with the same `fail to find path in version_graph` error.
- 证据等级
`多方印证(2 个独立来源)`,GitHub 用户缺陷报告 ×2(2026 K8s 生产 + 2025 磁盘迁移),同错误签名横跨 2.0.x–4.1.x。
`Corroborated (2 independent sources)`, GitHub user bug reports x2 (2026 K8s production + 2025 disk migration), same error signature across 2.0.x–4.1.x.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site on this site.
Apache Doris 年份:2026
单写者文件锁:把它当服务器库用的都会被锁教做人
多方印证
性能问题运维复杂度
- 一句话
一个进程以读写模式打开 DuckDB 文件后,其他任何进程(哪怕 `read_only=True`)都打不开同一文件——直接报锁冲突,读也要等写。
Once one process opens a DuckDB file read-write, no other process can open the same file — not even `read_only=True`. You get a lock conflict; even reads must wait for writes.
- 窄场景
把 `.duckdb` 文件当共享数仓,同时被 BI 工具(Metabase)、notebook、定时同步进程访问的团队;带着"连上就能查"的服务器数据库心智来用 DuckDB 的架构。
Teams using a `.duckdb` file as a shared warehouse accessed by BI tools (Metabase), notebooks, and scheduled sync processes; architectures arriving with a "connect and query" server-database mindset.
- 机制
DuckDB 是嵌入式库,没有 server 进程仲裁访问,文件锁按 PID 持有。一个进程持有读写连接时,其他进程的任何连接(包括只读)都会失败,报错 `IO Error: Could not set lock on file … Conflicting lock is held`。MVCC 只在同一进程内提供读写并发;跨进程的读只能在写连接释放后"错峰查询",或另行做快照拷贝。官方长期立场是多进程并发写"不是主要设计目标"。
DuckDB is an embedded library with no server process to arbitrate access; the file lock is held per PID. While one process holds a read-write connection, any connection from another process — including read-only — fails with `IO Error: Could not set lock on file … Conflicting lock is held`. MVCC only provides read/write concurrency within a single process. Cross-process reads must wait for the write connection to be released ("query between syncs") or use a snapshot copy. The long-standing upstream position is that multi-process concurrent writes are "not a primary design goal."
- 生产验证
来源 1:Dango 项目 ADR-003(2026)——选定 DuckDB 做单文件数仓后,被迫把 dlt 同步、dbt run/test 全部串行化(调度器 `max_instances=1`);Metabase(独立 Java 进程)与 Marimo notebook 在写入期间被文件锁挡住,只能在同步间隙查询或用快照拷贝,消费者连临时表都建不了;磁盘一坏就只能从源头重建,无复制;
来源 2:ChunkHound 项目的并发探针实验(2026)——1 写 + N 读的多进程实验里,所有读进程在 connect 阶段即失败(`Could not set lock on file`),实测结论是"一个 DuckDB 文件同一时间只能属于一个 OS 进程",项目被迫全站走单进程串行访问层(对照组 LanceDB 同场景可正常读写)。
Source 1: Dango project's ADR-003 (2026) — after choosing DuckDB as a single-file warehouse, they had to serialize all write jobs (dlt sync, dbt run/test) via scheduler `max_instances=1`; Metabase (a separate Java process) and Marimo notebooks get blocked by the file lock during writes and can only query between syncs or via snapshot copies; consumers cannot even create temp tables; a disk failure means rebuilding from sources, with no replication;
Source 2: ChunkHound's concurrency probe experiments (2026) — in a 1-writer + N-reader multi-process experiment, every reader process failed at connect time (`Could not set lock on file`); the measured conclusion was "one DuckDB file belongs to one OS process at a time," forcing a single-process serialized access layer across the project (with LanceDB as a working control in the same setup).
- 证据等级
`多方印证(2 个独立来源)`,独立项目工程记录 ×2(2026,均为带实测数据的记录)。
`Corroborated (2 independent sources)`, engineering records from 2 independent projects (2026, both with measured data).
- 备注
本卡主题可能与本站 [避坑] 卡重叠(单写者限制是 DuckDB 最广为人知的约束)。
this card's theme may overlap an existing [Pitfall] card on this site on this site (the single-writer limit is DuckDB's best-known constraint).
DuckDB 年份:2026
On-Demand 按量计费:一次成功的流量,就是一次账单灾难
多方印证
成本账单
- 一句话
按量计费没有上限:一个每页发 47 个查询的 dashboard 功能上线一个周末,账单从每月 $150 变成 $9,247。
On-demand has no ceiling: a dashboard feature issuing 47 queries per page load went live for a weekend, and the bill went from $150/month to $9,247.
- 窄场景
On-Demand 模式 + 查询扇出(无 JOIN 导致一个 API 拆成数十个 Query)+ 没有实时成本监控的团队。
On-demand mode + query fan-out (no JOINs means one API call becomes dozens of Query calls) + a team with no real-time cost monitoring.
- 机制
On-Demand 按请求数 × item 大小计费(读按 4KB、写按 1KB 向上取整),没有 JOIN 意味着关系查询被拆成 N 个独立 Query;与预置容量不同,按量模式没有"到顶变慢/报错"的天然熔断——流量即账单;常规监控只看延迟和错误率,看不到"单次请求的成本",出事时一切指标都是绿的。
On-demand bills per request x item size (reads in 4KB chunks, writes in 1KB chunks, rounded up); without JOINs, relational queries decompose into N independent Query calls; unlike provisioned capacity, on-demand has no built-in circuit breaker — traffic is the bill; standard monitoring watches latency and error rates, not cost per request, so everything stays green while money burns.
- 生产验证
来源 1:CodexLab 2026-01——周五下午上线用户活跃度 dashboard(每页加载 47 个 DynamoDB 查询,随后向 12,000 用户发邮件推送),周末 AWS 账单 $9,247.83(其中 DynamoDB $8,963.22/72 小时);应急加 Redis 缓存、合并查询 47→8,最终迁回预置容量(上限 $680/月);
来源 2:HN 2017 年讨论——从业者原话 "DynamoDB's billing and provisioning model is awful to deal with",另一条评论称曾见某初创公司出现约 $85K 的 DynamoDB 账单。
Source 1: CodexLab, Jan 2026 — shipped a user-activity dashboard Friday afternoon (47 DynamoDB queries per page load, then emailed 12,000 users); the weekend AWS bill was $9,247.83 ($8,963.22 of it DynamoDB over 72 hours); emergency Redis caching and query consolidation (47 to 8 queries per load), then back to provisioned capacity (capped at $680/month);
Source 2: HN discussion, 2017 — practitioner quote: "DynamoDB's billing and provisioning model is awful to deal with"; another commenter reported seeing a ~$85K DynamoDB bill at a startup.
- 证据等级
`多方印证(2 个独立来源)`,个人博客账单复盘 + HN 社区多方印证。
`Corroborated (2 independent sources)`, personal billing postmortem + HN community corroboration.
Amazon DynamoDB 年份:2026
访问模式先行:漏掉的查询没有"加个索引"那么简单
多方印证
运维复杂度成本账单
- 一句话
DynamoDB 要求你在第一天就想清楚所有查询:漏掉的访问模式,代价是回填、双写和迁移,而不是一条 DDL。
DynamoDB demands you know every query on day one: a missed access pattern costs a backfill, dual-writes, and a migration — not a DDL statement.
- 窄场景
需求还在演进的业务应用(报表、admin 后台、多条件筛选排序)。
Business applications whose requirements keep evolving (reports, admin consoles, multi-filter sorted lists).
- 机制
只能按键查询;FilterExpression 在 key 命中之后执行,不省 RCU;跨分区全局排序不支持;新增查询维度 = 新建 GSI(整表回填 + 写入成本翻倍)或反范式化冗余;改一个字段语义 = 全表 backfill(每条目一次读 + 一次写、限流管理、热点风险、双读兼容层)。
Key-only queries; FilterExpression runs after key selection and saves no RCUs; no cross-partition global sort; each new query dimension means a new GSI (full-table backfill + doubled write cost) or denormalized duplication; changing one field's semantics means a full-table backfill (one read + one write per item, throttle management, hot-partition risk, dual-read compatibility layers).
- 生产验证
来源 6:maxime 2025-11,多年 SaaS 生产经验——"small change"(1200 万条目改一个字段)= 数天设计 + backfill + 双写 + feature flag,原话 "You'll spend more time building the migration machinery than the feature that prompted it";单表设计下 Streams 变成混合事件流(每个消费者重做路由)、监控指标跨实体类型模糊、PITR 无法只恢复"某租户的订单"、一个坏写者可污染无关实体;
来源 7:Bhoos Games 2026-08——除 id 外没有任何索引,非 id 查询"既耗时又贵",用户数据分析根本做不了;为省空间把字段名编码成 `$a$1`/`$b$11b$5`(分别代表 diamond/coin/XP),外人看到的是乱码。
Source 6: maxime, Nov 2025, from years of SaaS production use — a "small change" (renaming one attribute across 12M items) meant days of design + backfill + dual-writes + feature flags: "You'll spend more time building the migration machinery than the feature that prompted it"; single-table design turns Streams into a mixed firehose (every consumer re-implements routing), blurs monitoring metrics across entity types, makes PITR unable to restore "just tenant X's orders", and lets one bad writer pollute unrelated entities;
Source 7: Bhoos Games, Aug 2026 — no indexes except id, so non-id queries were "both time-consuming and expensive" and user-data analytics was impossible; to save space, field names were encoded as `$a$1`/`$b$11b$5` (diamond/coin/XP) — gibberish to anyone reading the table.
- 证据等级
`多方印证(2 个独立来源)`,个人博客 ×2(含 1 个具名迁出复盘)。
`Corroborated (2 independent sources)`, personal blogs x2 (including 1 named migration postmortem).
Amazon DynamoDB 年份:2026
热分区:表级容量还剩 85%,请求却被限流
多方印证
性能问题稳定与故障
- 一句话
表级容量仪表盘一片健康,但某个 key 的流量打满了它所在的物理分区——你的真实上限由最热的那个 key 决定。
The table-level dashboard is all green, but one key's traffic saturated its physical partition — your real ceiling is set by your hottest key.
- 窄场景
多租户 SaaS(按 tenant 粒度做分区键)、事件型突发流量(大促、单 campaign 爆量)。
Multi-tenant SaaS keyed at tenant granularity; event-driven traffic bursts (sales events, one campaign exploding).
- 机制
DynamoDB 按分区键哈希分片,每个物理分区有固定吞吐上限(约 3000 RCU / 1000 WCU);表级容量是各分区预算之和,一个热 key 打满所在分区即返回限流(ProvisionedThroughputExceeded/ThrottlingException),其余分区完全空闲;"加总容量"只能按比例抬高所有分区,治标不治本且更贵;分区按吞吐自动分裂,但单个 key 永远分不到两个物理分区。
DynamoDB shards by hashing the partition key; each physical partition has a fixed throughput ceiling (~3,000 RCU / 1,000 WCU); table capacity is the sum of partition budgets, so one hot key exhausts its partition and gets throttled (ProvisionedThroughputExceeded/ThrottlingException) while the rest sit idle; raising total capacity only scales all partitions proportionally — treating a design problem as a capacity problem, at higher cost; partitions auto-split on throughput, but a single key can never span two physical partitions.
- 生产验证
来源 8:Rajamohan Jabbala 2026-01——广告实时竞价平台(50K 写/s、200K 读/s、2TB、on-demand 约 $18K/月),超级碗期间单个 campaign 占 80% 流量,打爆 `campaign_id` 分区,修复方案是分区键加随机后缀打散;
来源 6:maxime 2025-11——把 "fighting hot partitions and throughput tuning" 列为 DynamoDB 建模债清单之一。
Source 8: Rajamohan Jabbala, Jan 2026 — real-time ad-bidding platform (50K writes/s, 200K reads/s, 2TB, ~$18K/month on-demand): during the Super Bowl a single campaign took 80% of traffic and saturated the `campaign_id` partition; fix was a random suffix on the partition key to scatter load;
Source 6: maxime, Nov 2025 — lists "fighting hot partitions and throughput tuning" among DynamoDB's modeling debts.
- 证据等级
`多方印证(2 个独立来源)`,具名生产事故 + 独立博客印证。
`Corroborated (2 independent sources)`, named production incident + independent blog corroboration.
- 备注
分区级上限为架构性设计(文档化行为),adaptive capacity 只能缓解、不能消除单 key 上限,故保留收录,未作"已修复"标注。
per-partition caps are architectural (documented behavior); adaptive capacity only mitigates, never removes, the single-key ceiling — retained without a "fixed in" label.
Amazon DynamoDB 年份:2026
预置容量的运维税:为"4KB 涨到 8KB"开会
多方印证
运维复杂度
- 一句话
用预置容量,团队要为"数据值从 4KB 涨到 8KB"这种事认真开会:容量规划成了日常运维税。
On provisioned capacity, teams hold serious meetings about values growing "from 4KB to 8KB": capacity planning becomes a standing operations tax.
- 窄场景
流量有波峰波谷、用 provisioned + auto-scaling 控成本的中等规模团队。
Mid-size teams with spiky traffic using provisioned + auto-scaling to control costs.
- 机制
WCU 按 1KB、RCU 按 4KB 向上取整,item 平均大小翻倍 = 容量需求翻倍;auto-scaling 按分钟级调整,突发靠 burst 余额扛;新增功能(夜间分析 scan、备份)都要重调 auto-scaling 参数,否则限流 = 应用随机报错;缩容每天限次,峰值过后要为闲置容量继续付费。
WCUs bill in 1KB increments and RCUs in 4KB increments, rounded up — doubling average item size doubles capacity needs; auto-scaling adjusts on minute timescales with burst balance covering spikes; every new feature (nightly analytics scans, backups) forces auto-scaling parameter retuning, otherwise throttling means random application errors; scale-down is limited per day, so you keep paying for idle capacity after peaks.
- 生产验证
来源 2:HN 2017 年讨论——从业者原话 "having to constantly manage and change our provisioned capacity so that we could respond to usage spikes and keep our bill from becoming stratospheric…anytime we'd add new functionality…it meant revisiting the parameters for our auto-scaling or risk getting throttled";"serious conversations about the repercussions of your data values going from 4kb to 8kb";
来源 1:CodexLab 2026-01——账单惊魂后迁回预置容量,"Predictable costs: $680/month max…Auto-scaling with limits we control",即用"为闲置容量付费"换"可预测"。
Source 2: HN discussion, 2017 — practitioner quote: "having to constantly manage and change our provisioned capacity so that we could respond to usage spikes and keep our bill from becoming stratospheric…anytime we'd add new functionality…it meant revisiting the parameters for our auto-scaling or risk getting throttled"; "serious conversations about the repercussions of your data values going from 4kb to 8kb";
Source 1: CodexLab, Jan 2026 — after the billing shock, back to provisioned: "Predictable costs: $680/month max…Auto-scaling with limits we control" — i.e., paying for idle capacity in exchange for predictability.
- 证据等级
`多方印证(2 个独立来源)`,HN 社区 + 个人博客(2017 与 2026,机制九年未变,属持续性抱怨)。
`Corroborated (2 independent sources)`, HN community + personal blog (2017 and 2026; mechanism unchanged for nine years — a persistent complaint).
- 备注
HN 来源为 2017 年(预置容量时代);2026 年 CodexLab 复盘印证同一结构性抱怨仍然存在,故保留收录。
the HN source is from 2017 (the provisioned-capacity era); CodexLab's 2026 postmortem confirms the same structural complaint still holds, so it is retained.
Amazon DynamoDB 年份:2026
按 vCPU 计费:生产库、备库、分布式节点个个都要花钱,长期预算成"移动靶"
多方印证
成本账单
- 一句话
许可按 vCPU 数算——为了高可用加的备节点、为了分布式加的节点,全都要计费,扩容一次账单涨一截,提前几年做预算几乎不可能。
Licensing is priced per vCPU — standby nodes added for resilience and extra nodes added for distribution all count, so every scale-out bumps the bill and multi-year budgeting is nearly impossible.
- 窄场景
按 EDB 订阅跑生产的中大型团队;上了备库/DR、EDB Postgres Distributed 多节点、或云上按 vCPU 小时计费(BigAnimal)的架构。
Mid-to-large teams running production on EDB subscriptions; architectures with standby/DR nodes, EDB Postgres Distributed multi-node setups, or cloud vCPU-hour billing (BigAnimal).
- 机制
EDB 订阅按 CPU/vCPU 规模定价,计费的不是"数据库实例"而是"算力单元"。而企业级高可用恰恰要求更多算力单元:流复制备库、见证节点、分布式集群每个写节点都占 vCPU。于是"为了更可靠"和"为了更便宜"直接冲突——备节点提升了韧性,也同步抬高了许可费。云上 BigAnimal 按 vCPU 小时计费更把这种冲突实时化。
EDB subscriptions are priced on CPU/vCPU footprint, billing compute units rather than database instances. Enterprise-grade HA demands exactly more compute units: streaming-replication standbys, witness nodes, and every writer node in a distributed cluster consume vCPUs. So "more reliable" and "cheaper" pull in opposite directions — standby nodes buy resilience and simultaneously raise the license fee. Cloud vCPU-hour billing (BigAnimal) makes that conflict real-time.
- 生产验证
来源 1:Artee Y.(Novartis,企业用户,2026-04-20)原话:"Licensing based on vCPUs means costs increase as databases scale, especially with distributed configurations. Standby nodes improve resilience but also increase infrastructure and licensing cost. Estimating long term costs upfront is challenging."(按 vCPU 许可意味着成本随扩容上涨,备节点提升韧性也同步增加许可成本,提前估算长期成本很困难);
来源 1:Vineet K.(小企业,2026-04-24)原话:"because it is often tied to vCPU counts, costs can spiral quickly as you scale or add standby nodes for resilience, making long-term budget forecasting a bit of a moving target."(成本常与 vCPU 挂钩,扩容或为加备库加节点时成本迅速失控,长期预算预测成了移动靶)。
Source 1: Artee Y. (Novartis, enterprise user, Apr 20 2026): "Licensing based on vCPUs means costs increase as databases scale, especially with distributed configurations. Standby nodes improve resilience but also increase infrastructure and licensing cost. Estimating long term costs upfront is challenging.";
Source 1: Vineet K. (small business, Apr 24 2026): "because it is often tied to vCPU counts, costs can spiral quickly as you scale or add standby nodes for resilience, making long-term budget forecasting a bit of a moving target."
- 证据等级
`多方印证(2 个独立来源)`,具名评价 ×2。
`Corroborated (2 independent sources)`, named reviews x2.
EDB Postgres / EDB Postgres Advanced Server(EPAS) 年份:2026
Oracle 兼容的"最后一公里":最后 5–10% 的小众特性,迁移工具修不好,只能人工填坑
多方印证
升级迁移
- 一句话
PL/SQL 主体能过去,但 Advanced Queuing、专有包、特定分析函数(如 lag() ignore nulls)这类边角特性不在兼容清单里,工具链修不了,只能人工重写。
The PL/SQL bulk migrates fine, but edge features like Advanced Queuing, proprietary packages, and specific analytic functions (e.g. lag() ignore nulls) sit outside the compatibility list — tooling can't repair them, only manual rewrites can.
- 窄场景
从 Oracle 迁往 EDB Postgres Advanced Server、且原库深度使用了 Oracle 专有能力(AQ 消息队列、DBMS_* 包族、特殊分析函数语义、BFILE 等)的团队。
Teams migrating from Oracle to EDB Postgres Advanced Server whose source databases lean on Oracle-proprietary capabilities (AQ messaging, DBMS_* package families, special analytic-function semantics, BFILE, etc.).
- 机制
EPAS 的 Oracle 兼容是"语法子集"实现:PL/SQL 包、触发器、函数的主体语法被重写解释,但 Oracle 的专有运行时服务(队列、调度、全文、外部表、存储管理)没有对应实现。迁移工具只能做语法转换,遇到语义级缺口只能打报告、留人工。用户原话印证了厂商口径之外的真实缺口:lag() ignore nulls 这类单个函数语义缺失,意味着"兼容 85%"的每一份缺口都落在具体业务代码上。
EPAS Oracle compatibility is a syntax-subset implementation: PL/SQL packages, triggers, and function bodies get reinterpreted, but Oracle's proprietary runtime services (queuing, scheduling, full-text, external tables, storage management) have no counterpart implementation. Migration tooling can only do syntax conversion; semantic gaps get flagged in a report and left to humans. The reviewers' examples confirm the gap behind the vendor's headline numbers: one missing function semantic like lag() ignore nulls means every point of the "85% compatible" shortfall lands on real business code.
- 生产验证
来源 1:Vineet K.(小企业,2026-04-24)原话:"while the Oracle compatibility is excellent, it is not a 1:1 perfect match, and hitting that final 5-10% of niche features like advanced queuing or specific proprietary packages often requires frustrating manual workarounds."(Oracle 兼容很优秀但不是 1:1,最后 5–10% 的小众特性如 Advanced Queuing、专有包经常需要令人抓狂的手工绕行);
来源 3:Ahmad H.(软件研发经理,中型企业,2025-11-04,5/5 好评中仍提)原话:"there are some missing features, like 'lag() ignore nulls', that would significantly speed up development."(缺少 lag() ignore nulls 这类特性)。
Source 1: Vineet K. (small business, Apr 24 2026): "while the Oracle compatibility is excellent, it is not a 1:1 perfect match, and hitting that final 5-10% of niche features like advanced queuing or specific proprietary packages often requires frustrating manual workarounds.";
Source 3: Ahmad H. (software development manager, mid-market, Nov 4 2025, inside an otherwise 5/5 review): "there are some missing features, like 'lag() ignore nulls', that would significantly speed up development."
- 证据等级
`多方印证(2 个独立来源)`,具名评价 ×2。另有厂商自认佐证:来源 5(EDB 官方 PGD 已知问题)承认 EPAS 17 的 BFILE 数据类型在分布式复制中不支持、Eager 复制事务不能执行 DDL。
`Corroborated (2 independent sources)`, named reviews x2. Vendor-admitted corroboration: Source 5 (EDB's own PGD known issues) concedes the EPAS 17 BFILE data type is unsupported in distributed replication and eager-replication transactions cannot execute DDL.
- 备注
本卡主题可能与本站 [避坑] 卡重叠(Oracle 兼容缺口类)。
may overlap an existing [Pitfall] card on this site (Oracle compatibility gaps).
EDB Postgres / EDB Postgres Advanced Server(EPAS) 年份:2026
专有特性锁死退路:哪天想迁回社区 PG,"解绑"本身是架构级工程
多方印证
生态与信任
- 一句话
Oracle 兼容模式、EDB Failover Manager、TDE、数据脱敏视图这些"魔法功能"全是闭源专有——用得越深,迁回社区 PG 时要拆的东西越多。
The "magic features" — Oracle compatibility mode, EDB Failover Manager, TDE, redacted views — are all closed-source proprietary; the deeper you use them, the more there is to dismantle on the way back to community Postgres.
- 窄场景
已把 EDB 专有能力写进应用或运维体系的团队(用了 Oracle 兼容语法、EFM 自动故障转移、TDE、redacted views、PL/SQL 包装器);未来考虑降本迁回社区 PG 的 CTO。
Teams that have written EDB-proprietary capabilities into their applications or operations (Oracle-compatible syntax, EFM automated failover, TDE, redacted views, PL/SQL wrappers); CTOs weighing a future cost-driven move back to community Postgres.
- 机制
EDB 的商业价值恰恰来自社区 PG 没有的那一层:Oracle 兼容的 SQL 方言、专有故障转移编排、透明加密与脱敏。这些能力以闭源扩展/分支形式存在,没有社区等价物。应用代码一旦调用 EDB-only 函数、视图定义依赖脱敏列、运维依赖 EFM,迁回社区 PG 就不是"换个发行版",而是逐项重做:改 SQL 方言、换掉故障转移方案、重做加密与脱敏层。
EDB's commercial value is precisely the layer community Postgres lacks: the Oracle-compatible SQL dialect, proprietary failover orchestration, transparent encryption and redaction. These ship as closed-source extensions/forks with no community equivalents. Once application code calls EDB-only functions, view definitions depend on redacted columns, and operations depend on EFM, moving back to community Postgres is not "switching distributions" — it is redoing each item: rewriting SQL dialects, replacing the failover design, rebuilding the encryption/redaction layer.
- 生产验证
来源 2:匿名已验证用户 "II"(IT 与服务行业,中型企业,2026-05-15)原话:"many of the 'magic' features (Oracle compatibility, EDB Failover Manager, TDE) are proprietary. If you choose to move away from EDB later, your application code could end up closely tied to these EDB-only functions, making a transition more involved."(很多"魔法功能"是专有的,日后想离开 EDB,应用代码可能已与这些 EDB-only 函数深度绑定,迁移更麻烦);
来源 1:Vineet K.(小企业,2026-04-24)原话:"if you ever want to move back to community Postgres, untangling yourself from EDB-specific enhancements like redacted views or PL/SQL wrappers can be a significant architectural burden."(想迁回社区 PG 时,解开 redacted views、PL/SQL 包装器这类 EDB 特有增强是沉重的架构负担)。
Source 2: anonymous verified user "II" (IT & services, mid-market, May 15 2026): "many of the 'magic' features (Oracle compatibility, EDB Failover Manager, TDE) are proprietary. If you choose to move away from EDB later, your application code could end up closely tied to these EDB-only functions, making a transition more involved.";
Source 1: Vineet K. (small business, Apr 24 2026): "if you ever want to move back to community Postgres, untangling yourself from EDB-specific enhancements like redacted views or PL/SQL wrappers can be a significant architectural burden."
- 证据等级
`多方印证(2 个独立来源)`,G2 已验证评价 ×2(含 1 具名)。说明:本站尚无"从 EDB 迁回社区 PG"的具名完成案例(已知证据缺口),本卡收录的是用户明确感知的"迁出摩擦",非已发生的迁出复盘。
`Corroborated (2 independent sources)`, verified G2 reviews x2 (1 named). Note: this site has no named completed case of migrating from EDB back to community Postgres (a known evidence gap); this card records users' clearly perceived exit friction, not a completed migration postmortem.
- 备注
本卡主题可能与本站 [避坑] 卡重叠(供应商锁定类)。
may overlap an existing [Pitfall] card on this site (vendor lock-in).
EDB Postgres / EDB Postgres Advanced Server(EPAS) 年份:2026
文档散、社区小:EDB 特有的问题,在 StackOverflow 上几乎找不到人问
多方印证
运维复杂度
- 一句话
官方文档"很全但散在各处",而 EDB 分支特有的行为在公开社区几乎没有问答积累——出问题要么啃厂商文档,要么开支持工单。
The official docs are "comprehensive but scattered everywhere," while the fork's specific behaviors have almost no public Q&A history — when something breaks you either chew through vendor docs or open a support ticket.
- 窄场景
用 EPAS Oracle 兼容模式、EDB 专有工具链的团队;半夜排障、需要快速找到"别人踩过没有"的 DBA。
Teams using EPAS Oracle-compatibility mode or the EDB-proprietary toolchain; DBAs troubleshooting at 3 AM who need to know whether anyone else has hit the same wall.
- 机制
EDB 是 PG 的商业分支:Oracle 兼容模式的报错语义、EPAS 特有系统视图、PEM/EFM 的工具行为,都与社区 PG 有差异,社区 PG 的海量问答对这些差异无效。而 EDB 用户基数远小于社区 PG 和 Oracle,StackOverflow/论坛上 EDB 标签的问题量和回答者都很少,形成"文档是唯一指路牌"的局面;文档本身又按产品线拆成多套(EPAS、PEM、EFM、PGD 各一套),导航成本高。
EDB is a commercial Postgres fork: error semantics in Oracle-compatibility mode, EPAS-only catalog views, and PEM/EFM tool behaviors all differ from community Postgres, so community Postgres's vast Q&A corpus doesn't fully apply. Meanwhile the EDB user base is far smaller than community Postgres's or Oracle's, so EDB-tagged questions and answerers on StackOverflow/forums are scarce — leaving vendor docs as the only signpost, and those are split across product lines (EPAS, PEM, EFM, PGD each with its own set), raising navigation cost.
- 生产验证
来源 2:Kanishka R.(中型企业)原话:"Limited Community Resources: While PostgreSQL has a huge open-source community, EDB-specific features don't always have the same breadth of community-driven documentation or tutorials, so you often rely on vendor docs or support."(社区资源有限:EDB 特有功能没有社区 PG 那样丰富的社区文档/教程,经常只能依赖厂商文档或支持);
来源 1:Subhajit P.(中型企业,2026-04-25)原话:"documentation and troubleshooting resources, while comprehensive, can sometimes be difficult to navigate when trying to resolve specific issues quickly."(文档虽全,但紧急排障时很难快速导航到要找的内容);
来源 4:TrustRadius Vetted Review(在线教育 MOOC 服务商)原话:"It is sometimes hard to find a community of users on StackOverflow so a larger community, and a dedicated forum with active members to answer questions and work through issues would be nice."(在 StackOverflow 上很难找到 EDB 用户社区)。
Source 2: Kanishka R. (mid-market): "Limited Community Resources: While PostgreSQL has a huge open-source community, EDB-specific features don't always have the same breadth of community-driven documentation or tutorials, so you often rely on vendor docs or support.";
Source 1: Subhajit P. (mid-market, Apr 25 2026): "documentation and troubleshooting resources, while comprehensive, can sometimes be difficult to navigate when trying to resolve specific issues quickly.";
Source 4: TrustRadius vetted review (online-education MOOC provider): "It is sometimes hard to find a community of users on StackOverflow so a larger community, and a dedicated forum with active members to answer questions and work through issues would be nice."
- 证据等级
`多方印证(3 个独立来源)`,G2 评价 ×2 + TrustRadius Vetted Review ×1。
`Corroborated (3 independent sources)`, G2 reviews x2 + TrustRadius vetted review x1.
EDB Postgres / EDB Postgres Advanced Server(EPAS) 年份:2026
管理/监控工具 UI 欠现代:功能有,但用起来像上个时代的软件
多方印证
运维复杂度
- 一句话
PEM 等管理工具"能用",但界面与交互明显落后于新一代数据库平台,高级功能经常要回头翻文档才找得到。
PEM and friends "work," but their UI and interaction clearly lag newer database platforms — advanced features often send you back to the docs to find them.
- 窄场景
日常用 PEM 做监控、性能诊断、备份管理的 DBA;从 Datadog/PgHero 这类现代工具转过来的团队。
DBAs doing daily monitoring, performance diagnostics, and backup management in PEM; teams coming from modern tooling like Datadog or pgHero.
- 机制
EDB 的工具链(PEM 监控、性能诊断、Index Advisor 等)是随订阅附赠的"全家桶"组件,演进节奏服从数据库发行版而非独立产品迭代;且 PEM 客户端历史上是桌面应用架构,Web 化后交互包袱仍在。结果是:核心数据库能力持续加码,管理面的体验改进滞后,用户为"找一个功能"付出额外学习成本。
EDB's toolchain (PEM monitoring, performance diagnostics, Index Advisor, etc.) ships as bundled components of the subscription, evolving on the database release train rather than as independent products; PEM's client also carries desktop-app architectural heritage into its web UI. The result: core database capabilities keep advancing while management-plane experience lags, and users pay an extra learning cost just to locate a feature.
- 生产验证
来源 1:Subhajit P.(中型企业,2026-04-25)原话:"The user interface for some of the management and monitoring tools could be more intuitive and modern. Compared to some newer database platforms, navigation and usability can feel slightly less user-friendly."(部分管理/监控工具的 UI 可以更直观现代,相比新一代数据库平台,导航和易用性稍逊);
来源 2:Girishchand B.(测试工程师,中型企业,2026-02-20)原话:"There's also a bit of a learning curve because EDB adds its own tools and features on top of PostgreSQL, which takes time to get familiar with."(EDB 在 PG 之上加了自己的工具和功能,需要时间熟悉,有一定学习曲线)。
Source 1: Subhajit P. (mid-market, Apr 25 2026): "The user interface for some of the management and monitoring tools could be more intuitive and modern. Compared to some newer database platforms, navigation and usability can feel slightly less user-friendly.";
Source 2: Girishchand B. (test engineer, mid-market, Feb 20 2026): "There's also a bit of a learning curve because EDB adds its own tools and features on top of PostgreSQL, which takes time to get familiar with."
- 证据等级
`多方印证(2 个独立来源)`,具名评价 ×2。
`Corroborated (2 independent sources)`, named reviews x2.
EDB Postgres / EDB Postgres Advanced Server(EPAS) 年份:2026
企业版复杂度 vs 社区 PG:为了那层"企业",团队要多付一份学习税
多方印证
运维复杂度
- 一句话
EDB 在 PG 之上叠了自己的工具、扩展和概念,DBA 既要懂 PG 内核,又要懂 EDB 这一层——小团队常觉得"还不如直接用社区版"。
EDB stacks its own tools, extensions, and concepts on top of Postgres — DBAs must know both the Postgres engine and the EDB layer, and smaller teams often conclude community edition would have been simpler.
- 窄场景
PG 经验不深、被"企业级"卖点吸引的中小团队;本地化/培训资源有限的组织(如发展中国家公共部门团队)。
Teams without deep Postgres bench strength, drawn in by the "enterprise-grade" pitch; organizations with limited training/localization resources (e.g. public-sector teams in developing countries).
- 机制
EDB 的商业模式是"PG 内核 + 企业层":Oracle 兼容方言、专有扩展、自有工具链各自引入新概念和新参数。团队的知识负担不是 PG 的子集而是超集——排障时要先判断问题在 PG 层还是 EDB 层,再决定查哪套文档。对本来就缺 PG 资深人员的团队,这层增量直接转化为培训成本和招聘门槛。
EDB's business model is "Postgres engine + enterprise layer": the Oracle-compatible dialect, proprietary extensions, and its own toolchain each introduce new concepts and parameters. The team's knowledge burden is a superset, not a subset, of Postgres — troubleshooting first requires deciding whether the problem lives in the Postgres layer or the EDB layer, then picking the right doc set. For teams already short on senior Postgres people, that increment converts directly into training cost and hiring bar.
- 生产验证
来源 1:Suon S.(IAMS 水资源管理专家,小企业,2026-04-25,3.5/5)原话:"Its enterprise complexity can feel heavy compared to community PostgreSQL, and the training burden for local teams"(企业版复杂度相比社区 PG 显得沉重,本地团队的培训负担大),另提"proprietary features that risk vendor lock-in"(专有特性有锁定风险);
来源 2:Daniel G.(小企业,5.0/5)原话:"The main drawbacks are the higher cost and learning curve. Some advanced features require extra setup and resources"(主要缺点是更高的成本和学习曲线,高级功能需要额外配置与资源);
来源 2:匿名小企业已验证用户(4.0/5,"Reliable PostgreSQL")原话:"It can feel more complex than plain PostgreSQL, and some features require extra setup."(比纯 PG 更复杂,有些功能要额外配置)。
Source 1: Suon S. (IAMS water management specialist, small business, Apr 25 2026, 3.5/5): "Its enterprise complexity can feel heavy compared to community PostgreSQL, and the training burden for local teams" is real; he also flags "proprietary features that risk vendor lock-in";
Source 2: Daniel G. (small business, 5.0/5): "The main drawbacks are the higher cost and learning curve. Some advanced features require extra setup and resources";
Source 2: anonymous verified small-business user (4.0/5, "Reliable PostgreSQL"): "It can feel more complex than plain PostgreSQL, and some features require extra setup."
- 证据等级
`多方印证(3 个独立来源)`,具名评价 ×2 + 匿名已验证评价 ×1。
`Corroborated (3 independent sources)`, named reviews x2 + anonymous verified review x1.
EDB Postgres / EDB Postgres Advanced Server(EPAS) 年份:2026
2GB 配额打满:etcd 用 NOSPACE 把整个集群变成只读
多方印证
稳定与故障运维复杂度
- 一句话
MVCC 历史版本默认永远保留,bbolt 文件只涨不缩——到 quota 那天 etcd 直接全集群拒写,读正常、写全死,apiserver 接着 5xx。
MVCC history is kept forever by default and the bbolt file only grows — the day it hits quota, etcd rejects every write cluster-wide: reads fine, writes all dead, API server starts 5xxing.
- 窄场景
自建 Kubernetes/Patroni 集群,装机时没配自动压缩、没排 defrag、没改默认 2GB 配额;写越频繁(apiserver churn、Patroni 每几秒续租)死得越快。
Self-hosted Kubernetes/Patroni clusters set up without auto-compaction, without scheduled defrag, still on the default 2GB quota; the chattier the writers (API server churn, Patroni renewing leases every few seconds), the faster it dies.
- 机制
etcd 每个写操作(增删改)都产生一个新 revision,默认保留全部历史;compaction 只是在文件内部标记页面可复用,**文件本身不收缩**;`quota-backend-bytes` 默认 2GB(最大 8GB)。文件触及配额 → 触发 NOSPACE 告警 → 全集群拒绝一切写入。恢复必须三步按顺序:compact 丢历史 → 逐成员 defrag 真正回收 → 手工 `alarm disarm`,缺一步都继续拒写。
Every etcd write (create/update/delete) mints a new revision, and all history is retained by default; compaction only marks pages reusable **inside** the file — the file itself never shrinks; `quota-backend-bytes` defaults to 2GB (8GB max). Touching the quota raises the NOSPACE alarm and the whole cluster refuses writes. Recovery is three steps in order: compact away history, defrag each member to actually reclaim, then manually `alarm disarm` — skip one and writes stay rejected.
- 生产验证
来源 1:2026-07,Patroni 集群——leader 续租写 etcd 失败,`etcdserver: mvcc: database space exceeded`,自动 failover 机制在故障窗口内名存实亡;根因是装机以来从未配过 auto-compaction、从未 defrag,默认 2GB 配额原样保留;
来源 2:2026,自建 k8s 三 master——kubectl 先变慢后彻底无响应,三个节点的 etcd DB 文件大小竟是 1.1G/2.1G/2.4G(从未压缩/整理导致的不对称),`dbSize 2.1GB vs dbSizeInUse 829MB`,1.2GB 以上是纯碎片;compact 到 revision 98469458 + 逐节点 defrag 后 2.4GB→434MB,控制面复活。
Source 1: Jul 2026, Patroni cluster — leader lease renewals against etcd failed with `etcdserver: mvcc: database space exceeded`, leaving the automatic failover mechanism effectively decorative during the window; root cause was never configuring auto-compaction or defrag since day one, default 2GB quota untouched;
Source 2: 2026, self-hosted 3-master Kubernetes — kubectl degraded then died entirely; the three nodes' etcd DB files were 1.1G/2.1G/2.4G (asymmetric from never being compacted/defragged), `dbSize 2.1GB vs dbSizeInUse 829MB` with 1.2GB+ of pure fragmentation; compacting to revision 98469458 plus per-node defrag shrank 2.4GB to 434MB and revived the control plane.
- 证据等级
`多方印证(2 个独立来源)`,具名生产复盘 ×2。
`Corroborated (2 independent sources)`, named production postmortems x2.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site on this site.
etcd 年份:2026
慢盘即原罪:fsync 慢几毫秒,选主风暴教你做人
多方印证
性能问题稳定与故障
- 一句话
etcd 每个写都要 fdatasync,官方线是 p99 < 10ms——HDD/网络盘/被邻居挤占的盘分分钟超标,然后就是 "apply request took too long"、心跳发不出、Raft 反复重选,写在选主窗口里直接停摆。
etcd fsyncs every write, and the official bar is p99 < 10ms — HDDs, network disks, or disks shared with noisy neighbors blow past it, and then come the "apply request took too long" warnings, unsent heartbeats, and repeated Raft re-elections, with writes stalling inside every election window.
- 窄场景
etcd 跑在传统 HDD、虚拟化薄盘、云网络块存储(如 iSCSI)或与高 IO 业务混部的机器上;负载一高(备份窗口、存储重建、流量高峰)就从"看着健康"滑向 quorum 抖动。
etcd on spinning HDDs, thin-provisioned virtual disks, cloud network block storage (e.g. iSCSI), or machines shared with heavy I/O workloads; any load spike (backup windows, storage rebuilds, traffic surges) slides it from "looks healthy" to quorum flapping.
- 机制
Raft 提交前必须把 WAL 落盘,leader 还要按时发出心跳;fsync 超时 → 成员判定 leader 失联 → 发起选举 → 选举期间写暂停 → 新 leader 若同样慢盘,选举反复触发(election storm),集群长时间无主可写。时钟漂移(NTP 被拦)会叠加放大同一风险。加 CPU/加内存对此毫无作用——瓶颈在同步写路径。
Raft must persist the WAL to disk before committing, and the leader must emit heartbeats on time; slow fsync → members declare the leader dead → new election → writes pause during the election → if the new leader is equally slow, elections repeat (election storm) and the cluster goes long stretches with no writable leader. Clock drift (silently blocked NTP) compounds the same risk. More CPU/RAM does nothing — the bottleneck is the synchronous write path.
- 生产验证
来源 3:2026-08,Jason Chen 的生产 k3s 集群——三节点 quorum 明明"全健康",实测两台虚拟 HDD 节点 fsync 均值 14.2ms/13.9ms(**均值**已超 10ms 线),p99 估算 64–256ms(超标 6–25 倍),一小时内三台机器分别打出 512/1540/1223 条 "apply request took too long",单次 apply 延迟 100–600ms;另发现 NTP 被公司网络静默拦截,时钟漂移与慢盘叠加;
来源 4:2020-04,Gojek 生产 etcd 运维实录——存储延迟过线会引发 leader election storm("no leader keeps the lease for long enough"),明确建议**永不**把 etcd 放在远程块存储上,SSD 是硬性要求。
Source 3: Aug 2026, Jason Chen's production k3s cluster — quorum nominally "all healthy," yet two virtual-HDD nodes averaged 14.2ms/13.9ms fsync (**the average** past the 10ms line), estimated p99 64-256ms (6-25x over the ceiling), and the three servers logged 512/1540/1223 "apply request took too long" warnings in a single hour, with individual apply latencies of 100-600ms; separately found NTP silently blocked by the corporate network, clock drift stacking on the same risk;
Source 4: Apr 2020, Gojek's production etcd operations notes — storage latency past the line triggers leader election storms ("no leader keeps the lease for long enough"), with an explicit recommendation to **never** put etcd on remote block storage; SSD is a hard requirement.
- 证据等级
`多方印证(2 个独立来源)`,个人博客 ×2(含 1 个带完整实测数据的生产调查)。
`Corroborated (2 independent sources)`, personal blogs x2 (including 1 production investigation with full measurements).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site on this site.
etcd 年份:2026
升级 etcd:只能逐个小版本爬,滚回去要停机全量恢复
多方印证
来源存疑
升级迁移
- 一句话
etcd 不支持跨版本跳跃升级(3.4→3.6 必须经 3.5),且升级会改磁盘数据结构让老版本读不懂——没有滚动回滚,翻车=停机、从备份全量恢复;更讽刺的是,升级编排的状态往往就存在你正在升级的那个 etcd 里。
etcd does not support skipping versions (3.4 to 3.6 must go through 3.5), and upgrades rewrite the on-disk format so old binaries can't read it — there is no rolling rollback; a bad upgrade means stop-the-world and restore everything from backup. The irony: the upgrade orchestration state often lives in the very etcd being upgraded.
- 窄场景
气隙/离线环境、客户数据中心的无人值守集群(Teleport Gravity 场景);用发行版包(Ubuntu 24.04 自带 3.4.30)而非上游二进制的用户;还依赖 v2 API 的老客户端。
Air-gapped/offline customer data centers with unattended clusters (the Teleport Gravity case); users on distro packages (Ubuntu 24.04 still ships 3.4.30) instead of upstream binaries; legacy clients still on the v2 API.
- 机制
etcd 滚动升级只保证相邻小版本协议兼容,数据目录格式随版本演进且**不可逆**;一旦某成员写坏或升级后发现问题,不存在"逐个滚回"的官方路径,只能全停集群、清空数据目录、从备份恢复再重启。Teleport 的 catch-22:他们的升级流程协调状态存在 etcd 里,"把 etcd 下线一分钟做升级"等于在升级期间亲手关掉自己的协调器。另:v2store 自 3.4 起废弃、`--enable-v2` 在 3.6 被彻底移除,老客户端无退路;发行版打包严重滞后(2026 年的 Ubuntu 24.04 仍是 3.4.30,其 gRPC-gateway 的 member/list 行为与 Patroni 预期不一致导致 404)。
etcd's rolling upgrade only guarantees protocol compatibility between adjacent minors, and the data-directory format evolves irreversibly; if a member breaks mid-upgrade or the new version misbehaves, there is no supported "roll back member by member" path — only stop the whole cluster, wipe data dirs, restore from backup, restart. Teleport's catch-22: their upgrade coordination state lives in etcd, so "take etcd down for a minute to upgrade" means switching off your own coordinator mid-upgrade. Separately: v2store deprecated since 3.4, `--enable-v2` removed entirely in 3.6 — old clients have no way back; distro packaging lags badly (Ubuntu 24.04 in 2026 still at 3.4.30, whose gRPC-gateway member/list behavior 404s Patroni's expectations).
- 生产验证
来源 5:2018-07,Teleport 工程师 Kevin Nisbet 具名访谈——为 Gravity 用户做全自动无人值守升级时发现:不能跳版本(2→3.0→3.1 必须逐级走),没有干净的回滚,"might have to shut down your entire cluster, restore all of the data directories from backup",最终被迫自研"导数据→起空新集群→写回"的停机式升级方案(停 etcd 30 秒到 1 分钟);
来源 6:2026-10,Twinhull 实测——Ubuntu 24.04 的 etcd 3.4.30 上 Patroni etcd3 后端因 `/v3/cluster/member/list` 返回 404 而永远起不来,etcd 本体却显示 healthy;同机换上游 3.6 二进制立即恢复;并确认现行规则仍是"一次只能升一个小版本:3.4 → 3.5 → 3.6,逐成员"。
Source 5: Jul 2018, named interview with Teleport engineer Kevin Nisbet — building fully autonomous unattended upgrades for Gravity users, they found versions can't be skipped (2 to 3.0 to 3.1, each step mandatory), there is no clean rollback ("might have to shut down your entire cluster, restore all of the data directories from backup"), and ended up writing their own outage-style upgrade (export data, boot an empty new cluster, write everything back — etcd down 30 seconds to a minute);
Source 6: Oct 2026, Twinhull's hands-on test — Patroni's etcd3 backend on Ubuntu 24.04's etcd 3.4.30 never bootstraps because `/v3/cluster/member/list` returns 404, while etcd itself reports healthy; swapping in the upstream 3.6 binary on the same machine fixes it immediately; confirms the standing rule is still "one minor version at a time: 3.4 → 3.5 → 3.6, member by member."
- 证据等级
`多方印证(2 个独立来源)[来源存疑:来源 6 为商业 HA 套件厂商博客,文末含产品推广;其事实部分(curl 404 复现、逐级升级规则)具体可核,已如实标注]`,具名工程师访谈 + 独立实测记录。
`Corroborated (2 independent sources) [Questionable source: source 6 is a commercial HA-kit vendor blog with product promotion at the end; its factual claims (curl 404 reproduction, stepped-upgrade rule) are concrete and verifiable, flagged as-is]`, named engineer interview + independent hands-on test.
etcd 年份:2026
Galera 开源前途成疑:基金会主席说商业公司"想干嘛干嘛"
多方印证
生态与信任
- 一句话
MariaDB plc 2025 年收购 Galera 背后的 Codership 后,基金会执行主席公开表示商业公司对 Galera"想做什么都行"——集群技术的开源未来蒙上阴影。
After MariaDB plc acquired Codership (Galera's original company) in 2025, the Foundation's executive chairman publicly said the commercial company is free to do whatever it wants with Galera — clouding the cluster technology's open-source future.
- 窄场景
把 Galera Cluster 作为长期 HA 方案的重度用户。
Heavy Galera Cluster users who planned it as their long-term HA solution.
- 机制
Galera 从来不是 MariaDB 服务端核心工程团队的产品(Codership Oy 独立开发),2025 年被 MariaDB plc 收购后归商业公司所有;基金会只承诺保护 MariaDB server 本体。参照 MaxScale 的 BSL→纯商业先例,商业公司可随时调整 Galera 的许可或产品线归属,用户没有话语权。
Galera was never developed by MariaDB's core server engineering team (it came from the independent Codership Oy); since its 2025 acquisition by MariaDB plc it belongs to the commercial company, and the Foundation only commits to protecting MariaDB Server itself. Following the MaxScale BSL-to-commercial precedent, the commercial company can change Galera's licensing or product-line placement at any time — users have no say.
- 生产验证
来源 5:The Register 2026-09,基金会执行主席 Kaj Arnö 原话——Galera "attracted controversy over its open source future and its place in the company's enterprise product line","It's their freedom. It's their prerogative to decide what they want to do";
来源 4:Jepsen 2026 证实收购事实——"In 2025 MariaDB acquired Codership Oy, bringing Galera Cluster under the MariaDB umbrella"。
Source 5: The Register 2026-09, quoting Foundation executive chairman Kaj Arnö — Galera has "attracted controversy over its open source future and its place in the company's enterprise product line", and "It's their freedom. It's their prerogative to decide what they want to do" (author quotes);
Source 4: Jepsen 2026 confirms the acquisition — "In 2025 MariaDB acquired Codership Oy, bringing Galera Cluster under the MariaDB umbrella".
- 证据等级
`多方印证(2 个独立来源)`,独立媒体访谈 + 独立测试机构事实确认。
`Multi-source corroborated (2 independent sources)`, independent media interview + independent test-agency fact confirmation.
MariaDB 年份:2026
版本号迷宫:LTS/STS 交替发车,版本字符串还假装自己是 5.5.5
多方印证
运维复杂度升级迁移
- 一句话
10.11 LTS、11.0–11.3 一年支持、11.4 LTS、12.x 滚动——用户和面板厂商被版本号绕晕;驱动还得专门处理 MariaDB 假装 5.5.5 的握手字符串。
10.11 LTS, 11.0–11.3 with one year of support, 11.4 LTS, 12.x rolling — users and panel vendors get lost in version numbers; drivers still have to special-case MariaDB's handshake string that pretends to be 5.5.5.
- 窄场景
选型定版本、面板(CyberPanel/CWP)用户、写驱动/连接器的开发者。
Picking a version at selection time, panel (CyberPanel/CWP) users, and developers writing drivers/connectors.
- 机制
2023 年后 MariaDB 切到 LTS/STS 交替:STS 只支持 1 年、不建议生产用,但 11.0/11.1/11.2 这类短命版本会先出现在发行版和面板里,用户分不清该追哪个;LTS 之间(10.11→11.4)又有优化器代价模型改写等大变更。叠加历史包袱:服务端握手报 `5.5.5-10.11.2`(当年为兼容老客户端拒绝前导 `10.` 的 hack),naive 的版本解析会误判成 MySQL 5.5.5 走 legacy 路径,驱动必须特殊处理。
Since 2023 MariaDB alternates LTS/STS releases: STS gets one year of support and is not production-recommended, yet short-lived versions like 11.0/11.1/11.2 ship first in distros and panels, so users can't tell which train to ride; and between LTS trains (10.11 → 11.4) there are major changes like the optimizer cost-model rewrite. On top of that sits the historical hack: the server handshake reports `5.5.5-10.11.2` (a relic from when old clients rejected a leading `10.`), so naive version parsing mistakes it for MySQL 5.5.5 and takes legacy code paths — drivers must strip the prefix and detect the family explicitly.
- 生产验证
来源 10:CyberPanel 论坛 2024-07——用户想升 11.4 LTS,官方回复"我们还没支持这个版本,请等官方指南",用户不敢动;
来源 11:byteink/bit 驱动文档 2026-09——"Two traps that only show up against MariaDB"之首即版本字符串 hack,驱动被迫剥离前缀、显式判定家族;
来源 12:Vettabase 2024-05(Federico Razzoli)——"Short Term Support versions are not recommended for production, because support only lasts for one year"(引文为作者原话),11.4 LTS 支持到 2029-05。
Source 10: CyberPanel forum 2024-07 — a user wanting to upgrade to 11.4 LTS was told by the official account "we don't provide this version… please wait for our official guide", and backed off;
Source 11: byteink/bit driver docs 2026-09 — "Two traps that only show up against MariaDB", first being the version-string hack, forcing prefix-stripping and explicit family detection;
Source 12: Vettabase 2024-05 (Federico Razzoli) — "Short Term Support versions are not recommended for production, because support only lasts for one year" (author quote); 11.4 LTS supported until 2029-05.
- 证据等级
`多方印证(3 个独立来源)`,社区论坛 + 开源驱动文档 + 独立顾问评测。
`Multi-source corroborated (3 independent sources)`, community forum + open-source driver docs + independent consultant review.
- 备注
与现有 [避坑] 卡部分重叠(档案吐槽清单"版本坑:10.6 EOL"属同类主题)。
Partially overlaps the existing profile's "pitfall" list (the version-pitfall row on 10.6 EOL is the same theme).
MariaDB 年份:2026
Oracle 治下的社区信任危机:裁员、零提交、公开信与"换库"建议
多方印证
生态与信任
- 一句话
2025-09 Oracle MySQL 团队裁员、GitHub 公开仓库超 3 个月零提交,社区成立 OurSQL Foundation 并发表公开信要求独立治理;多位具名人物公开建议迁往 MariaDB/PostgreSQL。
—
- 窄场景
选型期评估 MySQL 长期维护风险的技术决策者;依赖社区版的企业用户。
Decision-makers evaluating MySQL's long-term maintenance risk during selection; enterprises depending on the Community Edition.
- 机制
MySQL 开发长期闭门(private code drops)、路线图不透明、安全 bug 不公开跟踪、企业版功能付费墙;2025 年公开提交量跌至 2000 年以来最低,社区贡献流程不透明导致外部无法实质参与,信任持续流失。
MySQL development happens behind closed doors (private code drops), the roadmap is opaque, security bugs are not publicly tracked, and enterprise features sit behind a paywall; public commit volume in 2025 fell to its lowest since 2000, and the opaque contribution process keeps outsiders from meaningfully participating — trust keeps draining.
- 生产验证
The Register(2026-02-17):致 Oracle 公开信(letter.3306-db.org)初期获约 100 个签名、后增至 248+,Vadim Tkachenko(Percona CTO、前 MySQL AB)、Peter Zaitsev(Percona 联合创始人)具名发声,MySQL 创始人 Michael "Monty" Widenius 对裁员表示 "heartbroken";DEVCLASS(2026-01-13):GitHub 仓库自 2025-09 零提交,Julia Vural(Percona)统计 2025 提交量为 2000 年以来最低,Otto Kekäläinen(前 AWS RDS、前 MariaDB 基金会 CEO)称 MySQL "open source only by license, but not as a project",建议改用 MariaDB/PostgreSQL;WebProNews(2026-06):Oracle 于 2026-06-25 发布治理改革(技术指导委员会初始席位为 AWS/Google Cloud/Oracle),Zaitsev 回应 "advisory capacity…better than nothing, but not PostgreSQL-type engagement",2026-05 成立的 OurSQL Foundation 继续要求有约束力的承诺(binding commitments)。
The Register (Feb 17, 2026): the open letter to Oracle (letter.3306-db.org) gathered ~100 signatures at first, later 248+; Vadim Tkachenko (Percona CTO, ex-MySQL AB) and Peter Zaitsev (Percona co-founder) spoke on the record; MySQL creator Michael "Monty" Widenius said he was "heartbroken" over the layoffs. DEVCLASS (Jan 13, 2026): the GitHub repo had zero commits since Sep 2025; Julia Vural (Percona) charted 2025 commit volume as the lowest since 2000; Otto Kekäläinen (ex-AWS RDS, ex-MariaDB Foundation CEO) called MySQL "open source only by license, but not as a project" and recommended MariaDB/PostgreSQL. WebProNews (Jun 2026): Oracle published a governance reform on Jun 25, 2026 (Technical Steering Committee initially seated with AWS/Google Cloud/Oracle); Zaitsev responded "advisory capacity…better than nothing, but not PostgreSQL-type engagement"; the OurSQL Foundation (founded May 2026) keeps demanding binding commitments.
- 证据等级
`多方印证(3 个独立来源)`,来源性质:科技媒体报道 + 具名人物观点(Percona CEO/CTO、MySQL 创始人、前 MariaDB 基金会 CEO),无阴谋论。
—
MySQL 年份:2026
8.4 LTS 默认禁用 mysql_native_password:老客户端升级即失联
多方印证
稳定与故障升级迁移
- 一句话
MySQL 8.4 LTS 默认不再加载 mysql_native_password,老驱动/老账号升级后直接认证失败;RDS 上修复还要改参数组并重启实例。
—
- 窄场景
8.0/8.3 升 8.4 LTS;PHP/PDO 老驱动、C 客户端等不支持 caching_sha2_password 的应用;账号仍用 mysql_native_password 的存量系统。
Upgrading 8.0/8.3 to 8.4 LTS; apps on old PHP/PDO drivers or C clients without caching_sha2_password support; fleets where accounts still use mysql_native_password.
- 机制
8.4 起 mysql_native_password 从内置改为需显式启用的独立插件(--mysql-native-password=ON),authentication_policy 里配了也没用——插件没加载直接报 ERROR 1524;9.x 将彻底移除。老客户端在握手阶段就被拒,不是"连上后报错",排查时容易误判为网络或账号问题。
Since 8.4, mysql_native_password moved from built-in to an explicitly-enabled standalone plugin (--mysql-native-password=ON); setting it in authentication_policy alone does nothing — with the plugin unloaded you get ERROR 1524; it will be removed entirely in 9.x. Legacy clients are rejected at the handshake stage, not "after connecting", so it is easily misdiagnosed as a network or account problem.
- 生产验证
AWS re:Post(约 2026-04):用户升到 8.4.3 后 PHP/PDO 应用报 unknown authentication method,而 DBeaver(新驱动)正常;修复=参数组 mysql_native_password=ON + 重启 RDS 实例(静态参数);GitHub kbk#13(2026-01):项目因 C 客户端不支持 caching_sha2_password 被迫把 MySQL pin 在 8.3.0,无法升级 8.4 LTS。
AWS re:Post (circa Apr 2026): after upgrading to 8.4.3, a PHP/PDO app reported "unknown authentication method" while DBeaver (new driver) worked fine; fix = set mysql_native_password=ON in the parameter group + reboot the RDS instance (static parameter). GitHub kbk#13 (Jan 2026): a project whose C client lacks caching_sha2_password support was forced to pin MySQL at 8.3.0 and cannot move to 8.4 LTS.
- 证据等级
`多方印证(2 个独立来源)`,来源性质:AWS 社区问答 + GitHub 项目 issue。
—
MySQL 年份:2026
VMware 软分区陷阱:虚拟机里只跑一个实例,按整个集群的物理核付费
多方印证
来源存疑
成本账单
- 一句话
Oracle 不承认 VMware 是"硬分区"——在 VMware 集群上跑 Oracle,哪怕只用其中几台宿主机,也可能被要求按整个集群(乃至整个 vCenter)的物理 CPU 买许可。
Oracle does not recognize VMware as "hard partitioning" — run Oracle on a VMware cluster and you can be asked to license every physical CPU in the cluster (or the whole vCenter), even if you only use a few hosts.
- 窄场景
在 VMware 上虚拟化部署 Oracle EE 的客户;这是审计中最常见的"大额缺口"来源。
Oracle EE virtualized on VMware; the most common source of large "gap" findings in audits.
- 机制
Oracle 的分区政策(partitioning policy)只认可 Solaris Containers、IBM LPAR、Fujitsu PAR 等"硬分区"技术;VMware 被归为"软分区",因此许可计量按 Oracle 可"触及"的全部物理处理器计算。独立授权顾问 2026 年的机制说明写道:"all servers within a cluster — or potentially an entire data center — must be licensed, even if Oracle is only running on a subset of hosts"(来源 13)。
Oracle's partitioning policy only recognizes "hard partitioning" technologies (Solaris Containers, IBM LPAR, Fujitsu PAR); VMware counts as "soft partitioning", so license metering covers all physical processors Oracle can "reach". An independent licensing consultancy's 2026 explainer puts it this way: "all servers within a cluster — or potentially an entire data center — must be licensed, even if Oracle is only running on a subset of hosts" (source 13).
- 生产验证
—
Source 2: former Oracle LMS manager Adi Ahuja testified the VMware issue "was in the middle of every audit... driver of half their deals";
Source 13: Nafkha Consulting's 2026-01-30 mechanics explainer (quoted above).
- 证据等级
`多方印证(2 个独立来源)`[来源存疑],具名前 Oracle 高管作证 + 授权咨询公司机制说明(咨询厂商,存在商业动机)。
`Multi-source corroboration (2 independent sources)` [Questionable source], named former Oracle executive testimony + licensing-consultancy mechanics explainer (consultancy with commercial motives).
- 备注
主题可能与现有 [避坑] 卡(虚拟化许可类)重叠。
topic may overlap with existing [Pitfall] cards (virtualization-licensing theme).
Oracle Database(甲骨文) 年份:2026
误点一下鼠标:默认启用的付费选项包,按"使用"不按"购买"计费
多方印证
来源存疑
运维复杂度成本账单
- 一句话
Oracle EE 默认安装并启用全部付费选项包,且无任何许可密钥检查——DBA 在 EM 里点一下或跑个 AWR 报告,就可能欠下按整个数据库服务器计价的账单。
Oracle EE installs and enables every paid option pack by default with no license-key check — a DBA clicking around in EM or running an AWR report can incur a bill priced on the entire database server.
- 窄场景
任何 Oracle EE 部署;审计时 DBA_FEATURE_USAGE_STATISTICS 视图里的使用记录是铁证。
Any Oracle EE deployment; DBA_FEATURE_USAGE_STATISTICS usage records are ironclad evidence at audit time.
- 机制
EE 安装时全部选项默认可用,无 license key 门槛;CONTROL_MANAGEMENT_PACK_ACCESS 参数默认值为 DIAGNOSTIC+TUNING,意味着 EM 自动启用诊断/调优包功能;一旦使用(use triggers license),DBA_FEATURE_USAGE_STATISTICS 会永久记录,"过去使用仍算数",且计费按底层数据库的全部处理器/NUP 算,而非按实际使用的那台小机器。
EE installs all options usable with no license-key gate; CONTROL_MANAGEMENT_PACK_ACCESS defaults to DIAGNOSTIC+TUNING, so EM enables Diagnostic/Tuning Pack features automatically; once used (use triggers license), DBA_FEATURE_USAGE_STATISTICS records it permanently — "past use still counts" — and billing is based on all processors/NUPs of the underlying database, not the small machine actually used.
- 生产验证
—
Source 4: a 2026 The Register forum first-hand account — during two audits in 2006–2007, a monitoring tool had accidentally enabled the Partitioning option, resulting in a $100–200k penalty; the kicker: Oracle's own sales rep initially did not understand Oracle's own SE licensing rules;
Source 5: MetrixData 360 (2026-10) statistics across 205 audited EE assets: 71% had accidental enablement — Diagnostics Pack 43%, Tuning Pack 31%, Partitioning 24%.
- 证据等级
`多方印证(2 个独立来源)`[来源存疑],亲历者故事 + 授权咨询公司统计(咨询厂商,样本与口径未经第三方复核;该公司自己也提醒审计防御厂商有夸大动机)。
`Multi-source corroboration (2 independent sources)` [Questionable source], first-hand story + licensing-consultancy statistics (consultancy sample and methodology not third-party verified; the firm itself warns audit-defense vendors have an incentive to exaggerate).
- 备注
亲历故事发生在 2006–2007 年,但"默认启用、无密钥检查、按使用计费"的机制经 2026 年来源印证仍然有效,故保留。主题可能与现有 [避坑] 卡(选项包许可类)重叠。
the first-hand story dates to 2006–2007, but the "enabled by default, no key check, billed by use" mechanism is confirmed by 2026 sources as still in force, hence retained. This card's topic may overlap with existing [Pitfall] cards on this site (option-pack licensing theme).
Oracle Database(甲骨文) 年份:2026
Oracle Support:MOS 又慢又难用,升级 SR 得靠隐藏按钮
多方印证
运维复杂度生态与信任
- 一句话
连 Oracle 社区意见领袖都公开吐槽:MOS 网站迟缓、AI 搜索难用、补丁下载报 400 错误,SR 升级路径藏在一个连客户经理都不知道的"Manager Actions"按钮里。
Even an Oracle community opinion leader publicly ranted: the MOS website is sluggish, its AI search is unusable, patch downloads throw 400 errors, and the SR escalation path hides behind a "Manager Actions" button even account managers did not know about.
- 窄场景
需要下载补丁、开 SR 解决生产问题的所有客户;越是紧急越能感受到。
Any customer needing to download patches or open SRs for production issues; the more urgent, the more it hurts.
- 机制
MOS(My Oracle Support)前端臃肿导致页面迟缓;补丁下载在浏览器端报 "400 Bad Request / Request Header Or Cookie Too Large"(cookie 过大),而 wget 或 AutoUpgrade 命令行反而正常——问题出在 Web 层而非后端;SR 升级(escalation)入口藏在 "Manager Actions" 菜单下,无文档指引。Tim Hall 的核心批评是:"it should work for everyone"——他需要 26 年建站积累的影响力、发公开 rant 才被官方约谈并解决问题,普通客户没有这条路。
The MOS (My Oracle Support) front end is bloated and slow; patch downloads fail in the browser with "400 Bad Request / Request Header Or Cookie Too Large" (oversized cookies) while wget or the AutoUpgrade CLI work fine — the problem is in the web layer, not the backend; the SR escalation entry hides under a "Manager Actions" menu with no documentation. Tim Hall's core criticism: "it should work for everyone" — he needed 26 years of building a site, an audience, and a public rant to get Oracle's attention and a resolution; ordinary customers have no such path.
- 生产验证
—
Source 7: Tim Hall's 2026-08-05 "Oracle Support: An Update" — full record of slow MOS, the 400 download errors, and the hidden escalation path; written as follow-up after his earlier public rant got him a meeting with Oracle;
Source 10: an old dbasupport independent DBA forum thread — a dbsnmp memory leak killed a database while Oracle Support deflected with template replies (status repeatedly set to "Waiting on Customer"), with multiple customers echoing in the thread.
- 证据等级
`多方印证(2 个独立来源)`,具名社区作者 2026 年记录 + 独立 DBA 论坛历史共鸣。
`Multi-source corroboration (2 independent sources)`, named community author 2026 record + independent DBA forum historical echoes.
- 备注
来源 10 为年代不明的旧帖,仅作"支持体验抱怨由来已久"的历史印证,不单独成证。
source 10 is an old thread of unknown vintage; used only as historical corroboration that Support-experience complaints predate 2026, not as standalone evidence.
Oracle Database(甲骨文) 年份:2026
锁管理器 LWLock 争用:大并发下"正确"的锁变成瓶颈
多方印证
性能问题
- 一句话
PG 的重量锁走共享锁管理器,fast-path 槽位每进程只有 16 个——超了就去挤 16 个分区的 LWLock,大并发下锁本身成为 CPU 瓶颈。
Postgres heavyweight locks go through a shared lock manager with only 16 fast-path slots per backend — exceed them and you pile onto 16 partitioned LWLocks, where the locks themselves become the CPU bottleneck under concurrency.
- 窄场景
数百上千连接、多表 join、多分区查询的集群;缩减副本数、单机承载更多查询后诱发。
Clusters with hundreds to thousands of connections, multi-table joins, or many-partition queries; triggered after consolidating replicas so each server carries more query load.
- 机制
AccessShare 等弱锁默认走 backend 本地 fast-path(FP_LOCK_SLOTS_PER_BACKEND=16,编译期常量,改不了);规划阶段要对表及全部索引加锁,分区表每个分区都要加锁,极易超 16 槽 → 回落到共享 LockManager → 争抢 16 个分区的 LWLock。PG16 之前等待队列还有二次方级退化。
Weak locks like AccessShare normally take the backend-local fast path (FP_LOCK_SLOTS_PER_BACKEND=16, a compile-time constant); planning locks every table plus all its indexes, and every partition of a partitioned table, so the 16 slots are easily exhausted → fall back to the shared LockManager → contention on its 16 partitioned LWLocks. Before PG16 the wait queues also degraded superlinearly.
- 生产验证
来源 3:GitLab 2023 年,工程师 Matt Smiley 在公开 issue 中调查——合并一台副本后 API/Web 慢约 2 小时,用 bpftrace 定位到 lock_manager LWLock 等待,根因是单机查询量上升导致锁争用(pganalyze E91 转述,2023-11);
来源 4:WebProNews 2026-07 独立报道印证了 GitLab 这一事件,并指出 AWS 官方博客 2025-07 亦记载 Aurora PostgreSQL 在高并发读多分区表时出现相同症状。
Source 3: GitLab, 2023 — engineer Matt Smiley's public issue investigation: after consolidating one replica, API/Web slowed for ~2 hours; bpftrace pinpointed lock_manager LWLock waits, root-caused to higher per-server query volume (recounted by pganalyze E91, Nov 2023);
Source 4: WebProNews, Jul 2026, independently corroborates the GitLab incident and notes AWS's Jul 2025 blog documenting the same symptoms on Aurora PostgreSQL with high-concurrency reads over many-partition tables.
- 证据等级
`多方印证(2 个独立来源)`,具名工程调查(经第三方转述)+ 独立媒体报道。
`Corroborated (2 independent sources)`, named engineering investigation (via third-party recount) + independent media coverage.
PostgreSQL(社区版) 年份:2026
复制槽失联 = 主库磁盘定时炸弹
多方印证
稳定与故障运维复杂度
- 一句话
复制槽承诺"消费确认前不删 WAL"——消费者一死,主库 pg_wal 无限堆积,直到写不进去。
A replication slot promises "no WAL deleted before the consumer confirms" — when the consumer dies, pg_wal on the primary grows without bound until writes fail.
- 窄场景
用逻辑复制/CDC(Debezium 类)或做过测试性订阅的 PG 主库;槽建完没人管、没人监控。
Primaries using logical replication/CDC (Debezium-style) or that ever had experimental subscriptions; slots created and never monitored.
- 机制
slot 的 restart_lsn 钉住 WAL 回收下限;inactive slot 的 catalog_xmin 还会在**全集群范围**挡住 autovacuum 清死元组(一个库的槽能拖住所有库的 vacuum);安全阀 max_slot_wal_keep_size(PG13+)默认 -1 即无限。
The slot's restart_lsn pins the WAL retention floor; an inactive slot's catalog_xmin additionally blocks autovacuum's dead-tuple cleanup **cluster-wide** (a slot on one database stalls vacuum on all databases); the safety valve max_slot_wal_keep_size (PG13+) defaults to -1, i.e. unlimited.
- 生产验证
来源 5:2025-03 个人复盘——2 个被遗忘的测试逻辑槽让集群 1.06TB 磁盘被 WAL 占满,删槽后瞬间降到 40GB;同时槽挡住了 autovacuum,删槽后业务时段 CPU 从 ~80% 降到 <10%;
来源 6:2026-09 实战记录——"一次性的测试订阅者忘删槽"是最常见的真实起因;
来源 7:2026-01 on-call 手册——restart_lsn 不推进是逻辑复制"最常见的生产故障模式",并给出 pg_replication_slots 监控阈值。
Source 5: Mar 2025 personal postmortem — 2 forgotten test logical slots let WAL fill 1.06TB of cluster disk; dropping them shrank usage to 40GB instantly; business-hours CPU fell from ~80% to <10% because the slots had also been blocking autovacuum;
Source 6: Sep 2026 field notes — "a one-off test subscriber, spun up and forgotten" is the most mundane real-world cause;
Source 7: Jan 2026 on-call runbook — a non-advancing restart_lsn is "the most common production failure mode" of logical replication, with pg_replication_slots alert thresholds.
- 证据等级
`多方印证(3 个独立来源)`,个人博客 ×3(含 1 个带完整数据的生产复盘)。
`Corroborated (3 independent sources)`, personal blogs x3 (including 1 production postmortem with full numbers).
PostgreSQL(社区版) 年份:2026
执行计划深夜翻转:没人改代码,查询慢了几十倍
多方印证
性能问题
- 一句话
统计信息漂移 + prepared statement 通用计划 + autovacuum 改分布——PG 规划器会在你睡觉时换计划,且无任何预警。
Statistics drift + prepared-statement generic plans + autovacuum changing the data landscape — the planner switches plans while you sleep, with zero warning.
- 窄场景
数据分布倾斜、有长尾值的表;用 prepared statement/JDBC 的应用;大表 autoanalyze 间隔长的库。
Skewed distributions with long-tail values; applications using prepared statements/JDBC; large tables with long autoanalyze intervals.
- 机制
规划器靠采样统计估行数,并假设列相互独立(相关列的选择率直接相乘 → 严重低估);prepared statement 前 5 次用定制计划、之后切通用计划(generic plan),长尾参数直接翻车;autovacuum 跑完更新统计/清理膨胀 → 成本模型突变 → 计划翻转。社区长期拒绝 hints,也没有官方的执行计划冻结/基线机制。
The planner estimates row counts from sampled statistics and assumes column independence (multiplying correlated selectivities → severe underestimates); prepared statements use custom plans for the first 5 executions then switch to a generic plan, which blows up on long-tail parameters; each autovacuum updates statistics/clears bloat → cost model shifts → plan flips. The community has long refused query hints, and there is no official plan-freezing/baseline mechanism.
- 生产验证
来源 8:2026-01 工程师亲历——按 user_id 查的长尾查询在第 5 次执行后切通用计划,"花了一整周"才定位;session 表把 5 万行估成 10 行选 nested loop,查询超时被 pager 叫醒;
来源 9:2026-05 事故叙述——跑了两年一直 2ms 的 dashboard 查询,深夜突变为 45 秒,200 连接池被打满引发级联崩溃(连接池耗尽 → 健康检查失败)。
Source 8: Jan 2026 first-hand account — a long-tail user_id query flipped to a generic plan after the 5th execution and "took a whole week" to diagnose; a sessions table misestimated 50,000 rows as 10, picked a nested loop, timed out, pager fired;
Source 9: May 2026 incident narrative — a dashboard query stable at 2ms for two years jumped to 45s at 2:47 AM, saturated a 200-connection pool, and cascaded (pool exhaustion → health-check failures).
- 证据等级
`多方印证(2 个独立来源)`,个人博客 ×2。
`Corroborated (2 independent sources)`, personal blogs x2.
PostgreSQL(社区版) 年份:2026
autovacuum 默认值:按比例触发在大表上形同虚设
多方印证
性能问题运维复杂度
- 一句话
autovacuum_vacuum_scale_factor=0.2 是按表大小百分比触发的——1000 万行表要死 200 万行才清理,之前全是膨胀。
autovacuum_vacuum_scale_factor=0.2 triggers on a percentage of table size — a 10M-row table must accumulate 2M dead rows before cleanup even starts.
- 窄场景
千万行以上大表、高频 UPDATE/DELETE 的 OLTP;默认参数直接上生产的团队。
Large tables (10M+ rows) with heavy UPDATE/DELETE in OLTP; teams that ship with defaults.
- 机制
触发条件 = threshold(50) + scale_factor × 行数,默认 0.2;默认只有 3 个 worker 全库排队,大表一次 vacuum 跑很久占住 worker,其他表继续攒死元组;cost-based 限流默认保守。调激进怕挤占业务 IO,调保守表就膨胀——"死亡之角"。
Trigger = threshold(50) + scale_factor x row count, default 0.2; only 3 workers queue for the whole cluster, and one long vacuum on a huge table starves the rest while they accumulate dead tuples; cost-based throttling defaults conservative. Tune aggressively and you risk squeezing business I/O; tune conservatively and tables bloat — the "horns of a dilemma."
- 生产验证
来源 10:2026-03 生产实录——接手 4000 万行表,已攒 800 万死元组、autovacuum 一小时没碰过,表膨胀 30%,顺序扫描拖慢、buffer cache 被垃圾页塞满;
来源 11:2026-06 事故清单——用户行为表 6 个月从 45GB 膨胀到 187GB,"六个月慢烧",最后靠 pg_repack 在线重整救回。
Source 10: Mar 2026 production account — inherited a 40M-row table with 8M dead tuples that autovacuum hadn't touched in over an hour; 30% bloat, dragging sequential scans, buffer cache full of garbage pages;
Source 11: Jun 2026 incident list — a user-activity table slow-burned from 45GB to 187GB over six months, eventually rescued with online pg_repack.
- 证据等级
`多方印证(2 个独立来源)`,个人博客 ×2。
`Corroborated (2 independent sources)`, personal blogs x2.
- 备注
与现有 [避坑] 卡主题部分重叠(该卡已含 VACUUM 运维税)。
partially overlaps the existing [Pitfall] card (which already covers the VACUUM operations tax).
PostgreSQL(社区版) 年份:2026
Sentinel:故障转移不丢数据是幻觉
多方印证
稳定与故障
- 一句话
Sentinel 的故障转移不是零数据丢失——异步复制下,主库已确认的写只要还没送到将要晋升的从库,晋升那一刻就没了;网络分区时还可能出现双主,旧主的写在分区愈合后被整体丢弃。
Sentinel failover is not zero-data-loss — with asynchronous replication, any acknowledged write on the old primary that has not yet reached the promoted replica is gone the moment promotion happens; during network partitions two primaries can coexist, and the old primary's writes are discarded wholesale when the partition heals.
- 窄场景
把 Redis 当队列/分布式锁/唯一状态源用的 Sentinel 部署;跨机房、网络分区场景。
Sentinel deployments using Redis as queues, distributed locks, or the sole source of state; cross-datacenter and network-partition scenarios.
- 机制
复制异步,主库 ack 写不等从库;Sentinel 选"最跟得上"的从库晋升,但窗口真实存在(min-replicas-to-write / WAIT 是 opt-in,非默认)。Sentinel 与 Redis 是两套分离系统:Sentinel 投票的是"它看到的"状态,脑裂时少数派分区照样晋升(被隔离的老主继续接受写),分区愈合后老主被降级、全量复制新主数据 → 分区期间的写被整体销毁。
Replication is asynchronous — the primary acks a write without waiting for replicas; Sentinel promotes the "most caught-up" replica, but the window is real (min-replicas-to-write / WAIT are opt-in, not defaults). Sentinel and Redis are two separate systems: Sentinel votes on what it sees, so in a split-brain the minority partition promotes anyway (the isolated old primary keeps accepting writes), and when the partition heals the old primary is demoted and full-syncs from the new primary — the partition-era writes are destroyed.
- 生产验证
来源 11(Kyle Kingsbury,具名专家分析):"any system which uses asynchronous primary-secondary replication, and can change which node is the primary, is inconsistent";分区场景推演:双主并存、老主写在愈合后被整体替换;
来源 12(2026-09 实测):"A write acknowledged by the old primary but not yet copied to the replica is gone when that replica is promoted";
来源 10(2026-09):"if the master dies before a replica catches up, whatever writes hadn't made it across yet are gone"。
Source 11 (Kyle Kingsbury, named expert analysis): "any system which uses asynchronous primary-secondary replication, and can change which node is the primary, is inconsistent"; the partition walkthrough shows dual primaries and wholesale replacement of the old primary's writes after healing;
Source 12 (measured, Sep 2026): "A write acknowledged by the old primary but not yet copied to the replica is gone when that replica is promoted";
Source 10 (Sep 2026): "if the master dies before a replica catches up, whatever writes hadn't made it across yet are gone".
- 证据等级
`多方印证(3 个独立来源)`,具名专家分析 + 2026 年实测 + 独立印证。
`Corroborated (3 independent sources)`, named expert analysis + 2026 measurement + independent corroboration.
- 备注
来源 11 发表于 2013 年;该结论针对异步复制架构(非已修复缺陷),且有 2026 年两个独立来源印证同一机制,故保留收录,未作"已修复"标注。
source 11 was published in 2013; the conclusion concerns the asynchronous-replication architecture (not a fixed bug), and two independent 2026 sources corroborate the same mechanism, so it is retained without a "fixed in version X" annotation.
Redis / Valkey 年份:2026
把 Redis 当主库:一次崩溃,几分钟的写就没了
多方印证
稳定与故障
- 一句话
RDB 是快照不是日志——两次快照之间崩溃,中间的写静默消失;容器里常见的"无持久化"配置下,一次重启就是一次删库。
RDB is a snapshot, not a log — a crash between two snapshots silently loses everything written in between; under the "no persistence" configs common in containers, one restart is a self-inflicted data wipe.
- 窄场景
拿 Redis 存库存/订单/队列(无上游副本)当唯一真相源的团队;容器部署沿用默认或空持久化配置。
Teams storing inventory, orders, or queues (with no upstream copy) in Redis as the single source of truth; container deployments keeping default or empty persistence settings.
- 机制
RDB 按 save 规则周期快照,崩溃丢快照间隔内全部写入;AOF everysec 最多丢 1 秒、always 才接近不丢但写吞吐腰斩(生产几乎没人用 always);容器常见配置等于无持久化。雪上加霜:maxmemory 淘汰策略配错会把"主数据"当缓存悄悄删掉。
RDB snapshots on save rules lose all writes since the last snapshot on crash; AOF everysec can lose up to 1 second, only always comes close to no-loss at roughly halved write throughput (almost nobody runs always in production); common container defaults equal no persistence. Worse: a misconfigured maxmemory eviction policy quietly deletes "primary data" as if it were cache.
- 生产验证
来源 13(具名,2026-06):秒杀库存计数器用 DECR 当主库,"It was fast, it was simple, and honestly, it just worked in staging. Then production happened.";
来源 6:AOF everysec 仍可丢 1 秒写、数据集必须永远 fit in RAM、"Eviction chaos... quietly starts deleting your 'primary' data";
来源 15(Valkey,机制同源,2026):VPS 半夜维护重启,4000 个队列 job(发票/欢迎邮件/webhook)无声消失,"The reboot didn't crash anything. It just quietly forgot."(持久化未开)。
Source 13 (named, Jun 2026): a flash-sale inventory counter decremented with DECR as the primary store — "It was fast, it was simple, and honestly, it just worked in staging. Then production happened.";
Source 6: AOF everysec can still lose 1 second of writes, the dataset must always fit in RAM, "Eviction chaos... quietly starts deleting your 'primary' data";
Source 15 (Valkey, same mechanism, 2026): a VPS maintenance reboot at night silently erased 4,000 queue jobs (invoices, welcome emails, webhooks) — "The reboot didn't crash anything. It just quietly forgot." (persistence was off).
- 证据等级
`多方印证(3 个独立来源)`,含 2 个具名生产复盘(其一为 Valkey 案例,机制与 Redis 同源)。
`Corroborated (3 independent sources)`, including 2 named production postmortems (one is a Valkey case with the same mechanism as Redis).
- 备注
来源 13 部分付费墙,仅前半可见,已如实标注。
source 13 is partially paywalled; only the first half is visible, flagged as such.
Redis / Valkey 年份:2026
换许可证:从 BSD 到 SSPL,社区信任的裂痕
多方印证
生态与信任
- 一句话
2024-03 Redis 把 BSD 换成 RSALv2/SSPLv1,"诱饵调包"(bait and switch)的指控、分叉与用户出走接踵而至——技术没变,信任变了。
In Mar 2024 Redis swapped BSD for RSALv2/SSPLv1, and "bait and switch" accusations, a fork, and a user exodus followed — the technology did not change, the trust did.
- 窄场景
需要 OSI 认证开源许可的企业/发行版(Debian/Fedora/RHEL 系);基于 Redis 做托管服务或嵌入商业产品的公司;pin 了 redis:7 这类浮动 docker tag 的团队。
Enterprises and distributions requiring OSI-approved licenses (Debian/Fedora/RHEL families); companies offering managed Redis or embedding it in commercial products; teams pinned to floating docker tags like redis:7.
- 机制
SSPL 要求"把 Redis 当服务提供就得开源整个服务栈",Debian/Fedora 认定非自由软件、拟从发行版移除;外部贡献者归零(CHAOSS 数据:之前 12 个非雇员贡献 54% commits,之后 5+ commits 的非雇员为 0);浮动标签 redis:7-alpine 在 2024-07 静默从 7.2(BSD)滑到 7.4(RSAL/SSPL),"没人改一行,标签自己漂移了"。
SSPL requires open-sourcing the entire service stack if Redis is offered as a service; Debian/Fedora deemed it non-free and prepared to remove it from distributions; external contributors dropped to zero (CHAOSS data: 12 non-employees with 54% of commits before, zero non-employees with 5+ commits after); the floating tag redis:7-alpine silently slid from 7.2 (BSD) to 7.4 (RSAL/SSPL) in Jul 2024 — nobody changed a line, the tag drifted by itself.
- 生产验证
来源 16(2025-04,独立媒体):Madelyn Olson 在 Monki Gras 具名讲述——"they very kindly informed me I was no longer a maintainer by deleting it from the governance stuff";
来源 17(2024,独立媒体):Percona 调查 70% Redis 用户另寻出路、83% 大企业已采用或评估 Valkey;AlmaLinux 基建负责人 Jonathan Wright 具名确认切换,"Valkey continues the open source legacy of Redis prior to its license change";
来源 18(2024-03 开源社区):Foreman/Pulp 维护者连夜评估 Dragonfly、memcached、Solid Queue,Debian/Fedora 法律列表跟进;
来源 19(2026-09):norite 项目被迫把 redis:7-alpine 换成 valkey,"the swap was forced rather than chosen";评估过 Redis 8(AGPLv3)但 "Valkey wins on being plainly BSD with no license to read twice"。
Source 16 (Apr 2025, independent media): Madelyn Olson telling it at Monki Gras, named — "they very kindly informed me I was no longer a maintainer by deleting it from the governance stuff";
Source 17 (2024, independent media): Percona survey — 70% of Redis users seeking alternatives, 83% of large enterprises adopting or evaluating Valkey; AlmaLinux infrastructure lead Jonathan Wright (named) confirming the switch — "Valkey continues the open source legacy of Redis prior to its license change";
Source 18 (Mar 2024, open-source community): Foreman/Pulp maintainers scrambling overnight to evaluate Dragonfly, memcached, and Solid Queue, with Debian/Fedora legal lists following;
Source 19 (Sep 2026): the norite project forced to swap redis:7-alpine for valkey — "the swap was forced rather than chosen"; Redis 8 (AGPLv3) was evaluated but "Valkey wins on being plainly BSD with no license to read twice".
- 证据等级
`多方印证(4 个独立来源)`,具名亲历 ×2 + 独立媒体 ×2。
`Corroborated (4 independent sources)`, named firsthand accounts x2 + independent media x2.
- 备注
本卡只收录事实与具名表态,不收立场争吵;Redis 8 的 AGPL 回调标注为"已部分修复(许可层面),社区分裂未愈"。
this card records facts and named statements only, not the partisan shouting; Redis 8's AGPL callback is annotated as "partially repaired (licensing level), community split unresolved".
Redis / Valkey 年份:2026
Valkey:分叉保住了协议,保不住"全家桶"
多方印证
来源存疑
升级迁移生态与信任
- 一句话
Valkey 保住了 BSD 与 RESP 协议,但没保住"全家桶"——Redis Stack 的 Search/JSON/TimeSeries 没有官方对应物,Redis 7.4 还被指破坏了与 Valkey 的数据文件兼容,"drop-in replacement"是有保质期的。
Valkey kept BSD and the RESP protocol, but not the "full suite" — Redis Stack's Search/JSON/TimeSeries have no official equivalents, Redis 7.4 is said to have broken data-file compatibility with Valkey, and "drop-in replacement" has an expiration date.
- 窄场景
用了 RediSearch/RedisJSON/RedisTimeSeries 后想迁 Valkey 的团队;从 Redis 7.4+ 向 Valkey 迁移。
Teams using RediSearch/RedisJSON/RedisTimeSeries that want to move to Valkey; migrations from Redis 7.4+ to Valkey.
- 机制
Valkey 从 7.2.4 分叉,只继承核心命令;Redis 8 把 JSON/时序/查询引擎并入核心发行版,Valkey 无对应集成(社区 valkey-search 等在追赶);RDB/AOF 二进制格式在 7.4 后被指不再互通(单方声称,见存疑标注),停机拷文件迁移的路变窄;RIOT(常用开源迁移工具)已归档。
Valkey forked from 7.2.4 and inherits only the core commands; Redis 8 folded JSON/time-series/query engines into the core distribution, with no integrated Valkey counterpart (community efforts like valkey-search are catching up); the RDB/AOF binary formats after 7.4 are said to no longer interoperate (single-vendor claim, see flags), narrowing the stop-the-world file-copy migration path; RIOT (a common open-source migration tool) is archived.
- 生产验证
来源 21(2026,真实项目工程规范):"Valkey-safe commands... Avoid depending on RediSearch / RedisJSON / Redis Stack unless ticket explicitly targets Redis node migration";"Redis node — only if Redis Stack / modules required";
来源 20(2026)[来源存疑]:"Redis 8.0 also integrates Redis Stack technologies like JSON, Time Series, and the Redis Query Engine. Valkey lacks these integrated features.";
来源 22(2026,HN)[来源存疑]:BetterDB 作者(前 Redis 工程经理)称 "Redis 7.4 broke data file compatibility with Valkey",RIOT 归档后无开源替代。
Source 21 (2026, real-project engineering spec): "Valkey-safe commands... Avoid depending on RediSearch / RedisJSON / Redis Stack unless ticket explicitly targets Redis node migration"; "Redis node — only if Redis Stack / modules required";
Source 20 (2026) [Questionable source]: "Redis 8.0 also integrates Redis Stack technologies like JSON, Time Series, and the Redis Query Engine. Valkey lacks these integrated features.";
Source 22 (2026, HN) [Questionable source]: the BetterDB author (former Redis engineering manager) claims "Redis 7.4 broke data file compatibility with Valkey", with no open-source alternative after RIOT's archival.
- 证据等级
`多方印证(3 个独立来源)[来源存疑]`——来源 20 的站点有 AI 生成痕迹(文中有无意义插句),来源 22 为迁移工具厂商的发布帖;核心事实(Valkey 无官方 Stack 模块对应物)由来源 21 的真实工程规范印证。
`Corroborated (3 independent sources) [Questionable source]` — source 20 shows signs of AI generation (meaningless filler phrases in the text), source 22 is a migration-tool vendor's launch post; the core fact (Valkey has no official Stack-module equivalents) is corroborated by the real-world engineering spec in source 21.
- 备注
Valkey 2024-03 才分叉,社区吐槽样本天然少;本卡已是诚实上限,未硬凑。
Valkey forked only in Mar 2024, so community complaint samples are naturally scarce; this card is the honest ceiling and nothing was padded.
Redis / Valkey 年份:2026
默认 10 分钟 auto-suspend:没人提醒你的空转税
多方印证
成本账单
- 一句话
仓库默认 10 分钟无查询才挂起、超配不告警——"能跑"不等于"跑得省",账单是事后才看到的慢漏。
Warehouses suspend only after 10 idle minutes by default, and nothing alerts you to over-provisioning — "it runs" is not "it runs cheap," and the bill is a slow leak you only see afterward.
- 窄场景
交互式 BI / adhoc 查询稀疏的仓库;默认配置直接上生产、之后无人复核的团队。
Sparse interactive BI / ad-hoc warehouses; teams that shipped with defaults and never reviewed them.
- 机制
`CREATE WAREHOUSE` 默认 `AUTO_SUSPEND=600`(10 分钟),每次查询结束后仓库再空转烧 9 分钟 credit;Snowflake 没有"利用率过低请降配"的反馈回路;workload 变化后当初合理的 size 不再合理,但没有任何东西会提醒你。
`CREATE WAREHOUSE` defaults to `AUTO_SUSPEND=600` (10 minutes), so every query leaves the warehouse burning credits for 9 more idle minutes; Snowflake has no feedback loop that says "utilization is low, size down"; a size that was right six months ago silently stops being right as workloads shift.
- 生产验证
来源 1:Surfalytics 2026-08-12 客户账单审计清单——"A warehouse running for 9 idle minutes after every query is pure waste",建议除 BI 仓库外一律 60 秒挂起;
来源 2:DSC 2026-02-24——仓库总 credit 中空转占比超 15% 是警告线、超 30% 必须处理;"Snowflake doesn't send you an alert that says 'hey, your ETL warehouse is running at 15% utilization, maybe size it down.' The warehouse just runs."
Source 1: Surfalytics, Aug 12 2026, client-bill audit checklist — "A warehouse running for 9 idle minutes after every query is pure waste"; recommends 60-second suspend for everything except BI warehouses;
Source 2: DSC, Feb 24 2026 — idle share above 15% of a warehouse's total credits is a warning sign, above 30% needs immediate action; "Snowflake doesn't send you an alert that says 'hey, your ETL warehouse is running at 15% utilization, maybe size it down.' The warehouse just runs."
- 证据等级
`多方印证(2 个独立来源)`,独立从业者博客 ×2。
`Corroborated (2 independent sources)`, independent practitioner blogs x2.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
账单惊魂:$2k 个人项目与"五位数意外"
多方印证
成本账单
- 一句话
Snowflake 默认没有"熔断器"——仓库配错、查询失控、dashboard 后台刷新,账单数字都是事后才知道的。
Snowflake ships with no circuit breaker — a misconfigured warehouse, a runaway query, a dashboard refreshing in a background tab, and the number on the invoice is something you learn after the fact.
- 窄场景
个人项目/小团队试用;未配置 resource monitor 的账户;BI dashboard 高频自动刷新的团队。
Side projects / small teams trialing the platform; accounts with no resource monitor; teams with aggressively auto-refreshing BI dashboards.
- 机制
resource monitor 默认不存在,必须手动创建(`credit_quota` + 触发器);仓库按运行时间计费,查询"是否必要"不影响计费;浏览器后台 tab 里开着的 dashboard 会按刷新间隔持续触发查询,无人观看也照样计费。
Resource monitors do not exist by default — you must build them yourself (`credit_quota` + triggers); warehouses bill for uptime regardless of whether the queries were necessary; a dashboard left open in a background tab keeps firing refresh queries on its interval, billed in full with zero humans watching.
- 生产验证
来源 3:Madison Schott 2025 年 LinkedIn 具名复盘——个人项目"losing $2k of my own precious money on a poorly optimized Snowflake cluster",所在公司 warehouse 年账单 $55,000;
来源 1:Surfalytics 2026-08-12——"A resource monitor will not optimize anything, but it stops a runaway query from turning into a five-figure surprise";
来源 4:Framesta Fernando 2026-07-22——某增长团队营销归因 dashboard 因后台 tab 自动刷新,一年产生五位数账单,"the refresh keeps happening in empty rooms, in background tabs, and through weekends, billing a full scan for an audience of zero"。
Source 3: Madison Schott's named 2025 LinkedIn retrospective — "losing $2k of my own precious money on a poorly optimized Snowflake cluster" on a personal project; her company's warehouse spend was $55,000;
Source 1: Surfalytics, Aug 12 2026 — "A resource monitor will not optimize anything, but it stops a runaway query from turning into a five-figure surprise";
Source 4: Framesta Fernando, Jul 22 2026 — a growth team's marketing-attribution dashboard produced a five-figure annual charge from background-tab refreshes: "the refresh keeps happening in empty rooms, in background tabs, and through weekends, billing a full scan for an audience of zero."
- 证据等级
`多方印证(3 个独立来源)`,具名个人复盘 + 独立从业者博客 ×2。
`Corroborated (3 independent sources)`, named personal retrospective + independent practitioner blogs x2.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
60 秒最低计费:小查询税
多方印证
成本账单
- 一句话
仓库每次 resume 先收 60 秒的钱——5 秒的查询按 60 秒计费,高频小查询的账单里大部分是空气。
Every warehouse resume bills a 60-second minimum — a 5-second query is billed as 60 seconds, so most of a high-frequency small-query bill is thin air.
- 窄场景
高频小查询/监控探针/轮询式集成;BI dashboard 每次加载触发几十个几秒级查询。
High-frequency small queries / monitoring probes / polling integrations; BI dashboards firing dozens of few-second queries per page load.
- 机制
按秒计费,但每次 resume 有 60 秒最低消费且 resume 即重置;频繁 suspend/resume 的小查询 workload,实际工作量远小于计费量——suspend 越激进,税越重。
Billing is per-second with a 60-second minimum on every resume, and each resume resets the minimum; for workloads that suspend and resume constantly around tiny queries, billed work dwarfs actual work — the more aggressive the suspend, the heavier the tax.
- 生产验证
来源 5:独立开源项目 detectkit 文档(2026)——"Snowflake bills each warehouse resume with a 60-second minimum — so detectkit's normal cadence of many small, frequent bookkeeping writes is disproportionately expensive against it",该项目因此被迫设计 hybrid mode:Snowflake 只许做只读数据源,状态一律写本地 DuckDB;
来源 4:Framesta Fernando 2026-07-22——"enforces a 60-second minimum charge each time a warehouse resumes, which creates an idle tax on exactly the frequent, short-running queries that dashboards generate"。
Source 5: docs of the independent open-source project detectkit (2026) — "Snowflake bills each warehouse resume with a 60-second minimum — so detectkit's normal cadence of many small, frequent bookkeeping writes is disproportionately expensive against it," forcing a hybrid-mode architecture where Snowflake is read-only source and all state lives in local DuckDB;
Source 4: Framesta Fernando, Jul 22 2026 — "enforces a 60-second minimum charge each time a warehouse resumes, which creates an idle tax on exactly the frequent, short-running queries that dashboards generate."
- 证据等级
`多方印证(2 个独立来源)`,独立开源项目文档 + 具名工程师博客。
`Corroborated (2 independent sources)`, independent open-source project docs + named engineer blog.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
Serverless 功能静默烧 credit:调优工具自己成了账单项
多方印证
成本账单
- 一句话
Automatic Clustering、Search Optimization、物化视图、Snowpipe 走 serverless credit 单独计费——省查询的钱可能不够付它们烧的钱,且只看仓库账单根本发现不了。
Automatic Clustering, Search Optimization, materialized views, and Snowpipe bill serverless credits on a separate meter — they can cost more than the queries they save, and a warehouse-only cost review never sees them.
- 窄场景
给大表开了 automatic clustering / search optimization / 物化视图的账户;只看 warehouse 账单做成本复核的团队。
Accounts with automatic clustering / search optimization / materialized views on large tables; teams doing cost reviews off the warehouse bill alone.
- 机制
serverless 功能不走用户仓库,走独立的 `SERVICE_TYPE` 计量(`METERING_HISTORY`);不适用 cloud services 10% 减免;automatic clustering 按数据变更持续烧 credit,物化视图按 base 表变更持续刷新——都是"开了就一直在后台跑"的计量项。
Serverless features bypass user warehouses and meter under separate `SERVICE_TYPE` entries (`METERING_HISTORY`); they are excluded from the cloud-services 10% adjustment; automatic clustering burns credits continuously with data churn, materialized views refresh on every base-table change — both are "always running in the background" meters once enabled.
- 生产验证
来源 1:Surfalytics 2026-08-12——"Automatic clustering is not free — it runs in the background and consumes credits","If [pruning] is already small, clustering will just cost you money";
来源 7:GroupBWT 2026-09-28——"the extra spend may sit in Automatic Clustering, Snowpipe, retention, replication, or egress instead";"Warehouse resizing does not directly control Snowpipe, Automatic Clustering, Materialized View refresh, storage, or transfer spend"。
Source 1: Surfalytics, Aug 12 2026 — "Automatic clustering is not free — it runs in the background and consumes credits"; "If [pruning] is already small, clustering will just cost you money";
Source 7: GroupBWT, Sep 28 2026 — "the extra spend may sit in Automatic Clustering, Snowpipe, retention, replication, or egress instead"; "Warehouse resizing does not directly control Snowpipe, Automatic Clustering, Materialized View refresh, storage, or transfer spend."
- 证据等级
`多方印证(2 个独立来源)`,独立从业者博客 + 咨询公司工程博客。
`Corroborated (2 independent sources)`, independent practitioner blog + consultancy engineering blog.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
Warehouse 冷启动:suspend 一丢缓存,早高峰的查询就变慢
多方印证
性能问题
- 一句话
仓库挂起即清空本地 SSD 缓存——省钱的 auto-suspend 和热乎的查询是一对矛盾,早高峰的第一波查询替你付了"冷启动税"。
Suspending a warehouse wipes its local SSD cache — the cost-saving auto-suspend and warm queries are the same knob turned opposite ways, and the first queries of the morning pay the "cold-start tax."
- 窄场景
夜间挂起、早间 BI 高峰的仓库;查询依赖本地 SSD 缓存命中的大表扫描。
Warehouses suspended overnight with a morning BI peak; large table scans that depend on local SSD cache hits.
- 机制
运行中的仓库把微分区缓存在本地 SSD(local disk cache),suspend/resize/drop 即清空;resume 后首批查询走远端对象存储重建缓存;result cache 只有 24 小时且只命中完全相同的查询文本,救不了"相似但不同"的 BI 查询。
A running warehouse caches micro-partitions on local SSDs (local disk cache); suspend/resize/drop clears it; the first queries after resume rebuild the cache over remote object storage; the result cache only lives 24 hours and only hits byte-identical query text, which does not save "similar but different" BI queries.
- 生产验证
来源 1:Surfalytics 2026-08-12——"The exception is a warehouse serving interactive BI. There, suspend kills the cache, and users feel it. Keep those at 5-10 minutes and leave the rest at 60 seconds";
来源 2:DSC 2026-02-24——"There are cases where a longer suspend makes sense (like when cache reuse is critical), but the default of 5 to 10 minutes is almost always too high"。
Source 1: Surfalytics, Aug 12 2026 — "The exception is a warehouse serving interactive BI. There, suspend kills the cache, and users feel it. Keep those at 5-10 minutes and leave the rest at 60 seconds";
Source 2: DSC, Feb 24 2026 — "There are cases where a longer suspend makes sense (like when cache reuse is critical), but the default of 5 to 10 minutes is almost always too high."
- 证据等级
`多方印证(2 个独立来源)`,独立从业者博客 ×2。
`Corroborated (2 independent sources)`, independent practitioner blogs x2.
Snowflake 年份:2026
主键外键只是"纸面约束":声明了也不拦重复行
多方印证
生态与信任
- 一句话
Snowflake 里 PRIMARY KEY / UNIQUE / FOREIGN KEY 只存元数据、不做强制——重复主键照插不误,RELY 还会让优化器基于"假唯一"消掉 join,静默算错。
In Snowflake, PRIMARY KEY / UNIQUE / FOREIGN KEY are stored metadata, not enforced — duplicate primary keys insert without error, and RELY lets the optimizer eliminate joins on a "fake unique" assumption, silently computing wrong answers.
- 窄场景
从传统 RDBMS 迁移、指望 DDL 约束保数据质量的团队;给约束加了 RELY 的表。
Teams migrating from traditional RDBMS expecting DDL constraints to guard data quality; tables with RELY on their keys.
- 机制
列式 MPP 为保批量写入吞吐,不做唯一性/引用完整性检查;约束仅用于文档与优化器改写;RELY 让优化器信任"唯一"假设做 join elimination,数据若有重复则查询结果静默错误——"wrong results, faster"。
The columnar MPP engine skips uniqueness/referential-integrity checks to protect bulk-load throughput; constraints serve documentation and optimizer rewrites only; RELY tells the optimizer to trust the uniqueness assumption for join elimination — if the data has duplicates, results are silently wrong: "wrong results, faster."
- 生产验证
来源 8:Roshan Bagde 个人 GitHub 实证项目(2026),《Snowflake Constraints — What's Enforced, What's a Lie》——可运行的 SQL demo:`01_proof_not_enforced.sql` 证明重复主键/外键违规插入成功(只有 NOT NULL 会拦),`02_rely_join_elimination.sql` 证明同一查询在 NORELY vs RELY 下返回不同结果;
来源 9:独立开源项目 dlt 文档(2026)——"`unique` and `primary_key` are not enforced and dlt does not instruct Snowflake to `RELY` on them when query planning"。
Source 8: Roshan Bagde's personal GitHub evidence project (2026), "Snowflake Constraints — What's Enforced, What's a Lie" — runnable SQL demos: `01_proof_not_enforced.sql` proves duplicate primary keys and FK violations insert fine (only NOT NULL blocks), `02_rely_join_elimination.sql` proves the same query returns different results under NORELY vs RELY;
Source 9: docs of the independent open-source project dlt (2026) — "`unique` and `primary_key` are not enforced and dlt does not instruct Snowflake to `RELY` on them when query planning."
- 证据等级
`多方印证(2 个独立来源)`,独立工程师 GitHub 实证 + 独立开源项目文档。
`Corroborated (2 independent sources)`, independent engineer's GitHub evidence + independent open-source project docs.
- 备注
部分缓解——Hybrid tables(Unistore)支持强制主键/外键/唯一约束,但标准列式表至今仍不强制,故保留收录,未作"已修复"标注。本卡主题可能与本站 [避坑] 卡重叠。
partially mitigated — Hybrid tables (Unistore) support enforced primary/foreign/unique constraints, but standard columnar tables still enforce nothing, so this is retained without a "fixed in" label. Topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
裁剪失效:84 亿行表只返回 1284 行,却扫描 1.4TB
多方印证
来源存疑
性能问题
- 一句话
选择性谓词不等于选择性扫描——CDC 持续写入让 micro-partition 的 min/max 元数据"撒得很开",优化器有谓词也裁不掉 partition,查询计划看起来一切正常,直到你去数实际扫描的分区数。
A selective predicate does not produce a selective scan — continuous CDC writes scatter each partition's min/max metadata so widely that the optimizer can't prune partitions despite having a predicate; the query plan looks perfectly fine until you count the partitions actually scanned.
- 窄场景
CDC/Kafka 持续写入的大表;写入顺序与查询访问模式长期错位、未做 clustering 的表。
Large tables with continuous CDC/Kafka writes; tables where ingestion order has drifted far from query access patterns and no clustering was ever done.
- 机制
micro-partition 裁剪依赖每个分区的 min/max 元数据;持续追加写入使同一谓词值散布在大量分区里,裁剪彻底失效,查询静默退化为接近全表扫——且执行计划里看不出异常。
Micro-partition pruning depends on per-partition min/max metadata; steady append writes spread the same predicate values across huge numbers of partitions, so pruning silently degrades into a near-full scan — invisible in the execution plan.
- 生产验证
来源 10:独立数据工程师 Sendoa Moronta 2026-09-08 复盘——ORDERS 表 84 亿行/3.7TB/约 190 万分区,一条只返回 1,284 行的查询耗时 41 秒,扫描 1.4TB、124.7 万个 partition(占全表 65%);"A selective predicate does not necessarily produce a selective scan";
来源 11:Alexandra Sampietro 2026-02-20 实验([来源存疑])——clustering key 列顺序放错会导致裁剪彻底失效:正确排序只扫 22/1600+ 个 partition,高基数列在前则扫 ~99%,第三列直接 100% 全表。
—
- 证据等级
`多方印证(2 个独立来源,含 1 个 [来源存疑])`,独立工程师生产复盘 + 机制实验。
`Corroborated (2 independent sources, including 1 [source questionable])`, independent engineer production postmortem + mechanism experiment.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
整点延迟尖刺:共享云服务层被挤爆,warehouse 是无辜的
多方印证
性能问题
- 一句话
每小时整点查询延迟翻倍,但查询执行时间只有 16ms——尖刺全在 cloud services(编译 200ms + 请求接收 350ms),根因是跨租户的共享层容量争用,调 warehouse 参数毫无用处。
Query latency doubles on the hour, yet query execution time is ~16ms — the spikes live entirely in cloud services (~200ms compilation + ~350ms request receipt); the cause is cross-tenant capacity contention on the shared layer, and tuning warehouse parameters does nothing.
- 窄场景
把 Snowflake 当低延迟 operational serving 层用的团队;对整点/半点延迟尖刺敏感的在线查询。
Teams using Snowflake as a low-latency operational serving layer; online queries sensitive to on-the-hour latency spikes.
- 机制
Cloud Services 是处理全账号会话管理、编译与路由的共享层,有独立于 warehouse size 的容量与限流;整点时刻跨租户活动叠加导致争用,延迟与数据扫描量无关。
Cloud Services is the shared layer handling session management, compilation and routing for all accounts, with capacity and throttling independent of warehouse size; top-of-hour activity across tenants collides there, and the latency is unrelated to data scanned.
- 生产验证
来源 12:Mechanical Rock 2026-10 生产复盘——目标 <300ms、10 rps 的 hybrid 表 serving,每小时整点延迟从 200–300ms 跳到 400–600ms;"Root cause: shared cloud services capacity experiencing contention at hour boundaries from other activity across the Snowflake platform. The warehouse was irrelevant.";最终靠 escalation 让 Snowflake 为该账号单独 provision dedicated cloud services tenancy——"something that requires an escalation, not a configuration change";
来源 14:Slim Baltagi 引用的 2022-09-01 Snowflake 客服回复原文——"Cloud Services servers encountering heavier than normal usage. As a result, some queries had timeouts reading metadata";同一根因跨 4 年仍在发生。
—
- 证据等级
`多方印证(2 个独立来源)`,独立咨询公司生产复盘 + Snowflake 客服工单原文。
`Corroborated (2 independent sources)`, independent consultancy production postmortem + Snowflake support ticket verbatim.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
并发一高查询就排队:2 秒的查询突然变 45 秒;spill 落盘是双重惩罚
多方印证
来源存疑
性能问题
- 一句话
warehouse 在压力下只有两条可预测的崩溃路径——并发饱和让查询在队列里等(执行时间没变,全耗在排队上),内存压力让中间结果 spill 到远端存储;spill 的小 warehouse 跑得越久烧得越多,"spill 六分钟的小 warehouse 比两分钟跑完的合适规格更贵"。
Warehouses fail under pressure in exactly two predictable ways — concurrency saturation parks queries in the queue (execution time unchanged, all of it spent waiting), and memory pressure spills intermediate results to remote storage; a small warehouse that spills runs longer and bills more: "a small warehouse that spills for six minutes can cost more than a properly sized one that finishes in two."
- 窄场景
并发突增的 BI 高峰;warehouse 规格偏小、中间结果大的复杂查询。
BI rush hours with concurrency bursts; undersized warehouses running complex queries with large intermediate results.
- 机制
CPU 饱和时查询在队列等待,可通过 `QUERY_HISTORY.queued_overload_time` 确诊;内存不足时中间结果从 RAM→本地 SSD→远端对象存储逐级 spill;Snowflake 按 size×time 计费,spill 让"小规格省钱"的直觉反转为"小规格更贵"。
CPU saturation queues queries (diagnosable via `QUERY_HISTORY.queued_overload_time`); memory pressure spills intermediates RAM → local SSD → remote object storage; per size×time billing inverts the "smaller is cheaper" intuition into "smaller spills longer and costs more."
- 生产验证
来源 13:Ismail Mezzour 2026-03——"Queries that normally take 2 seconds suddenly take 45 seconds.";"A small warehouse that spills for six minutes can cost more than a properly sized one that finishes in two.";"Performance breaks for predictable reasons: Concurrency saturation, Memory pressure.";
来源 14:Slim Baltagi——默认 `MAX_CONCURRENCY_LEVEL=8`,"Such a low query concurrency limit forces either increasing the size… or starting additional clusters… this forces you to burn even more credits";
来源 11:Alexandra Sampietro([来源存疑])——remote disk spillage 是"终极性能杀手",节点 16GB RAM / 200GB SSD,一旦 spill 到 S3 性能断崖。
—
- 证据等级
`多方印证(3 个独立来源,含 1 个 [来源存疑])`,独立作者 ×3,机制描述一致。
`Corroborated (3 independent sources, including 1 [source questionable])`, three independent authors describing the same mechanism.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
RBAC 角色层级:一座没人敢动的"承重墙"
多方印证
运维复杂度
- 一句话
角色层级腐烂有四个阶段——随意起步、复制粘贴式增长、应急授权补丁、彻底纠缠,终点是 "The mess is now load-bearing — people are relying on access paths nobody remembers granting, and everyone is afraid to touch it."(这团乱麻成了承重墙——人们依赖着没人记得授权过的访问路径,谁也不敢动它。)
Role-hierarchy decay has four stages — casual beginnings, copy-paste growth, emergency grant patches, total entanglement — ending at "The mess is now load-bearing — people are relying on access paths nobody remembers granting, and everyone is afraid to touch it."
- 窄场景
多人多团队共用账号、role 层层嵌套的大型 Snowflake 租户;做过多次"临时授权"的组织。
Large Snowflake tenants with nested roles shared across teams; organizations with a history of "temporary" grants.
- 机制
Snowflake 的 role 继承 + secondary roles 让授权路径难以审计:secondary roles 激活时用户可用所持任意角色的权限,但 query history 上记的 primary role 未必是真正放行的那个——"who can touch this"和"which specific grant let this exact query succeed"是两个不同的问题;撤销变成赌博:"Was this access load-bearing, or a leftover from an incident eighteen months ago? The only way to find out is to revoke it and see who complains."
Role inheritance plus secondary roles make grant paths unauditable: with secondary roles active a user can exercise any held role's privileges, but the primary role recorded in query history may not be the one that actually authorized the query — "who can touch this" and "which specific grant let this exact query succeed" are two different questions; revocation becomes a gamble: "Was this access load-bearing, or a leftover from an incident eighteen months ago? The only way to find out is to revoke it and see who complains."
- 生产验证
来源 20:DZone 从业者投稿 2026-10(作者在两家金融服务机构负责 Snowflake 治理)——非生产角色"临时"桥接生产原始数据,"Now your production PII exposure isn't just a function of who holds production roles. It's also a function of who holds non-production roles";
来源 21:Mayank Sethi 2025-12-26(两次 petabyte 级、七位数年 spend 落地)——Snowflake 上手太容易反而是问题,非正式授权模式一旦被人和 pipeline 依赖,"I've never seen it happen cleanly"(清理从未顺利过)。
—
- 证据等级
`多方印证(2 个独立来源)`,同一现象:RBAC 债务一旦形成几乎无法干净清理。
`Corroborated (2 independent sources)`, same phenomenon: RBAC debt, once formed, is nearly impossible to clean up.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
Snowflake Task:静默失败、没有重试、重启靠手
多方印证
运维复杂度
- 一句话
task 默认 SUSPENDED 且永不报错——"A suspended task will never fail or produce an error message because it's not running."(挂起的 task 永远不会失败也不会报错,因为它根本没在跑);无自动重试,父失败则子无限等待,"Failures happen but go unnoticed until data is missing"(失败发生时没人知道,直到发现数据没了)。
Tasks default to SUSPENDED and never error — "A suspended task will never fail or produce an error message because it's not running." No automatic retries; a failed parent leaves children waiting forever: "Failures happen but go unnoticed until data is missing."
- 窄场景
用 Snowflake Task 做数据 ingestion/调度的团队;多步 DAG 依赖链。
Teams scheduling data ingestion on Snowflake Tasks; multi-step DAG dependency chains.
- 机制
Task 的可观测性黑洞——SHOW TASKS 是唯一查状态的办法,INFORMATION_SCHEMA 无等价视图,包成 view 的路全被堵死(CREATE VIEW 不支持 SHOW、存储过程塞 FROM 子句也建不了 view);单 task 限一条 SQL,复杂逻辑被迫包进存储过程后,"Snowflake sees the procedure as one successful block even if internal steps fail, making debugging a nightmare in the QUERY_HISTORY";20 步 DAG 中间插 task 要手工重连;上一次运行超时会导致下一轮整轮被 skip,数据静默变 stale。
A Task observability black hole — SHOW TASKS is the only way to check state, INFORMATION_SCHEMA has no equivalent view, and every path to wrap it in a view is blocked (CREATE VIEW doesn't support SHOW; procedures can't go in a FROM clause either); one task runs exactly one SQL statement, so complex logic gets stuffed into stored procedures, after which "Snowflake sees the procedure as one successful block even if internal steps fail, making debugging a nightmare in the QUERY_HISTORY"; inserting a task mid-DAG (20 steps) requires manual rewiring; a timed-out run causes the next round to be skipped wholesale, silently staling data.
- 生产验证
来源 23:Greenhouse 工程团队(Matt Feeser)2025-11——"I spent much more time on this than I feel should be required. All I'm after is a list of suspended tasks. That should be easy, right?";
来源 24:Manoj Patil 2025-12——"Tasks do not retry automatically";"If a parent task fails, child tasks wait indefinitely";"Restarting a failed chain becomes manual and risky";
来源 25:Nazeer Syed 2026-01-26——单 DAG 上限 1000 个 task、单个 task 最多 100 个前驱。
—
- 证据等级
`多方印证(3 个独立来源)`,客户工程团队 + 2 个独立从业者博客,同一现象:task 失败静默、无自动重试、可观测性差、重启靠手。
`Corroborated (3 independent sources)`, customer engineering team + two independent practitioner blogs, same phenomenon: silent failures, no auto-retry, poor observability, manual restarts.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
New Relic:"深陷单一厂商的存储、计算与查询模型之锁",迁 Iceberg 年化降本 35–52%
多方印证
升级迁移
- 一句话
New Relic 数据工程团队公开承认 "we were deeply locked into a single vendor's storage, compute, and query model";Snowflake 打包存储+计算+查询的模型让成本随数据量"阶梯式跳涨",且任何架构变更都必须"经过 Snowflake"——最终把 1000+ 数据集迁到 Apache Iceberg 开放表格式,年化开支降 35–52%,"We're not tied to anyone's roadmap or pricing decisions"(不再被任何人的路线图或定价决策绑住)。
New Relic's data engineering team publicly admitted "we were deeply locked into a single vendor's storage, compute, and query model"; Snowflake's bundled storage+compute+query model made costs jump in steps with data growth, and every architecture change had to go "through Snowflake" — so they moved 1,000+ datasets to Apache Iceberg open table format, cutting annualized spend 35–52%: "We're not tied to anyone's roadmap or pricing decisions."
- 窄场景
数据量持续增长、架构自主权重要的平台团队;被 Snowflake 打包模型卡住扩展路线的大客户。
Platform teams with growing data volumes that value architectural autonomy; large customers whose expansion roadmap is blocked by Snowflake's bundled model.
- 机制
专有存储格式 + 计算绑定 + 查询引擎三位一体,迁出=重写整套管线;开放表格式(Iceberg on S3)把存储、计算、编排三权分立。
Proprietary storage format + bound compute + query engine as one package: leaving means rewriting the whole pipeline; open table format (Iceberg on S3) separates storage, compute and orchestration.
- 生产验证
来源 33:New Relic 官方工程博客(数据工程总监 Matt Olsen 与 Mehreen Tahir 合著)——并行双跑+行级校验,1000+ 数据集(批+流)迁往 Iceberg(S3 存储、Spark on K8s 计算、Airflow 编排),最终 "fully open and portable";
来源 39:fintech 数据工程师 Aniket Soni 2026-09——专有 micro-partition 是"引擎兼监狱"("we treated Snowflake as both the engine and the prison"),迁出只能 UNLOAD——"a massive export job that burns compute credits, messes with your time-travel metadata";"Stop paying for storage hostage-taking"。
—
- 证据等级
`多方印证(2 个独立来源)`,客户官方工程博客 + 独立从业者,锁定机制互相印证。
`Corroborated (2 independent sources)`, customer engineering blog + independent practitioner, lock-in mechanism mutually confirmed.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
Dagster:Python Connector 时间戳 bug,2008 年变成 52170 年,上游明确"不优先修"
多方印证
生态与信任
- 一句话
`snowflake-connector-python` 存 timestamp 列会"搅碎"数据——年份 2008 变成 52170,下游 Pandas 读取直接报错;Dagster 两种 workaround 都恶心(转字符串/强制 UTC 时区),而上游维护者明确 "will not prioritize fixing this issue because some workarounds exist"(不优先修,因为有变通办法)——"until Snowflake fixes the underlying issues in their Python connector, it is unlikely that we will be able to provide a solution that works in all cases for all users."
`snowflake-connector-python` mangles timestamp columns on write — year 2008 becomes 52170 and downstream Pandas reads blow up; both of Dagster's workarounds are ugly (stringify / force UTC), while upstream maintainers said they "will not prioritize fixing this issue because some workarounds exist": "until Snowflake fixes the underlying issues in their Python connector, it is unlikely that we will be able to provide a solution that works in all cases for all users."
- 窄场景
Python 生态(Dagster/pandas)写 Snowflake 时间列;跨时区时间数据。
Python-ecosystem (Dagster/pandas) writes of timestamp columns to Snowflake; cross-timezone temporal data.
- 机制
connector 底层类型映射 bug(`df.to_sql` + `pd_writer` 把 datetime 列误判为 VARCHAR(16777216));同讨论还列出:Pandas Timedelta 被存成纳秒整数、含 None 的整数列被存成 float 列。
A connector-level type-mapping bug (`df.to_sql` + `pd_writer` misjudging datetime columns as VARCHAR(16777216)); the same thread lists more: Pandas Timedelta stored as nanosecond integers, integer columns with None stored as float.
- 生产验证
来源 43:Dagster 维护团队技术 RFC——引用 2020 年的 connector issue #319 为已知背景,多年未修;
来源 44:ibis 社区 issue #11983(约 2026-04)——同一类 Arrow/时间戳管线问题的另一个独立实例:`cursor.fetch_arrow_batches()` 按每个批次数据"猜"时间戳精度(ns vs us),多批次合并报 `ArrowInvalid: Schema at index 6 was different`;修复参数 `force_microsecond_precision=True` 默认 False 且藏在 docstring 里。
—
- 证据等级
`多方印证(2 个独立来源)`,Dagster 维护团队 + ibis 社区用户,同一 connector 时间戳管线痛点的两个独立实例;Snowflake 官方 connector 更新日志多次出现同类时间戳转换修复,说明反复出现。
`Corroborated (2 independent sources)`, Dagster maintainers + ibis community user, two independent instances of the same connector timestamp-pipeline pain; Snowflake's own connector changelog repeatedly fixes the same class of timestamp conversion bugs.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
没有原生 MLflow:Databricks 一个 workspace 走完,Snowflake 要跨三个面
多方印证
生态与信任
- 一句话
Databricks 把 AI 当"system to engineer":Spark 做特征、MLflow 做跟踪、Model Serving 做端点、Unity Catalog 横跨表+特征+模型+向量索引——"Promoting to ML on Databricks: add a notebook that reads the same silver tables, trains a classifier, logs to MLflow, serves via Model Serving. **One workspace.**";Snowflake 侧是"AI as a function call from the warehouse",真要训自己的模型就得跨面:"either call Cortex from a SQL function, OR export features to a Snowpark Python session, train, register in Snowflake's model registry. **Two surfaces (SQL + Snowpark).**"——"The moment ML enters, Databricks pulls ahead — MLflow has no Snowflake equivalent."
Databricks treats AI as a "system to engineer": Spark for features, MLflow for tracking, Model Serving for endpoints, Unity Catalog across tables + features + models + vector indexes — "Promoting to ML on Databricks: add a notebook that reads the same silver tables, trains a classifier, logs to MLflow, serves via Model Serving. **One workspace.**" Snowflake is "AI as a function call from the warehouse" — training your own model means crossing surfaces: "either call Cortex from a SQL function, OR export features to a Snowpark Python session, train, register in Snowflake's model registry. **Two surfaces (SQL + Snowpark).**" — "The moment ML enters, Databricks pulls ahead — MLflow has no Snowflake equivalent."
- 窄场景
想在数仓内完成"特征→训练→跟踪→上线"全链路的 ML 团队。
ML teams wanting the full "features → train → track → serve" loop inside the warehouse.
- 机制
Snowflake 后来补了 Model Registry(Snowflake ML),但原生 MLflow 实验跟踪 / Model Serving 式端点 / 原生 Feature Store 仍无等价物;Celebal 实测补细节:"No in-built MLflow support, can connect to external server such as AML experiments";"Feature store not natively available";"No support for auto ML";当时模型"can only be deployed for batch inferencing and there is no support for real-time or streaming deployments"。
Snowflake later added a Model Registry (Snowflake ML), but native MLflow experiment tracking / Model-Serving-style endpoints / native Feature Store still have no equivalent; Celebal's measurements add: "No in-built MLflow support, can connect to external server such as AML experiments"; "Feature store not natively available"; "No support for auto ML"; models "can only be deployed for batch inferencing and there is no support for real-time or streaming deployments" (at the time).
- 生产验证
来源 50:btriani 实测 repo 第 2/5 问(2026-05);
来源 51:Celebal Technologies 实测对比长文(Kaggle 数据集双平台实跑,含对比表)。
—
- 证据等级
`多方印证(2 个独立来源)`,独立实测 repo + 咨询公司实测对比。
`Corroborated (2 independent sources)`, independent measurement repo + consultancy hands-on comparison.
- 备注
缺口状态:至今缺失(Model Registry 已补,MLflow/Feature Store/AutoML/实时 serving 仍无)。与现有 2 张 Snowpark 卡不重复——那两张讲"难用",本卡讲"ML 生命周期能力缺失本身"。本卡主题可能与本站 [避坑] 卡重叠。
gap status: still missing (Model Registry added; MLflow/Feature Store/AutoML/real-time serving still absent). Not a duplicate of the 2 existing Snowpark cards — those cover "hard to use," this covers "missing ML-lifecycle capabilities." This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
Git 集成:喊了多年,2025 年 11 月才 GA,至今仍要大量手工定制
多方印证
运维复杂度
- 一句话
2023 年 Summit 后咨询顾问写道,Git 集成"we don't yet have a release timeline for this"(连发布时间表都没有),但"this is a commonly-requested feature"(这是个被反复要的功能);时间线:2024 年 6 月 Summit 才宣布私有预览,2025 年 11 月 4 日官方通稿宣布 GA。但 GA 不等于好用:2026 年 DevOps 横评:"Snowflake Git Integration is worth exploring, though it currently requires significant customization and development effort"(值得一试,但目前需要大量的定制和开发工作)——一个喊了 N 年的功能,补上时已经是"能用但不好用"的状态。
After Summit 2023 a consultant wrote that Git integration had "we don't yet have a release timeline for this," yet "this is a commonly-requested feature"; timeline: private preview announced at Summit Jun 2024, GA announced Nov 4 2025 via official press release. But GA ≠ usable: a 2026 DevOps bake-off says "Snowflake Git Integration is worth exploring, though it currently requires significant customization and development effort" — a feature asked for over N years arrived as "works but not well."
- 窄场景
想把 Snowflake 对象纳入 Git 版控的团队;2025 年 11 月之前等集成的用户。
Teams wanting Snowflake objects under Git version control; anyone who waited for the integration before Nov 2025.
- 机制
版控是数据平台的刚需,Snowflake 用户等了多年;GA 版本仍需大量定制才能融入 DevOps 流程。
Version control is table stakes for data platforms; Snowflake users waited years; the GA version still needs heavy customization to fit DevOps flows.
- 生产验证
来源 60:DAS42 Summit 2023 回顾(2023-07),"commonly-requested feature"原话出处;
来源 61:disqr DevOps 横评(2026),"requires significant customization and development effort"原话出处;
来源 64:Snowflake 官方通稿转述(bizwire,2025-11-04),GA 时间锚点(仅作时间核验用)。
—
- 证据等级
`多方印证(3 个独立来源)`,咨询公司回顾 + 第三方横评 + 官方 GA 时间锚点。
`Corroborated (3 independent sources)`, consultancy recap + third-party bake-off + official GA time anchor.
- 备注
缺口状态:已补上于 2025-11(GA);但按 2026 年第三方横评仍需大量定制,算部分补上,故保留收录。本卡主题可能与本站 [避坑] 卡重叠。
gap status: fixed Nov 2025 (GA); per 2026 third-party review still needs heavy customization — partially fixed, kept on record. This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
Dynamic PIVOT:ANY 选项补上了,但在存储过程/UDF 里依然不行
多方印证
生态与信任
- 一句话
很长一段时间 Snowflake 的 PIVOT 要求把转置列一个个手写进 `IN (...)`——2023 年 Retool 论坛用户:"So I've searched several times for a solution to be able to pivot data in Retool. It's fairly complicated to do it in SQL (Snowflake) and have it be dynamic";后来 Snowflake 给 PIVOT 加了 `IN (ANY ...)`,新增列自动变新列,40 行代码变 6 行;但 2026 年 1 月实测指出两个至今未补的洞:"Dynamic pivot is not supported in the code of stored procedure or UDF"(编译时就要定死列);view 上用动态透视,新数据一来列数变了还可能直接报错——adhoc 查询爽了,工程化场景(UDF/SP/view)依然绕行。
For a long time Snowflake's PIVOT required hand-listing every transposed column in `IN (...)` — a 2023 Retool forum user: "So I've searched several times for a solution to be able to pivot data in Retool. It's fairly complicated to do it in SQL (Snowflake) and have it be dynamic"; Snowflake later added `IN (ANY ...)`, new columns automatically become new columns, 40 lines → 6; but a Jan 2026 measurement found two holes still open: "Dynamic pivot is not supported in the code of stored procedure or UDF" (columns must be fixed at compile time); on views, arriving new data can change column counts and error out — adhoc queries are happy, engineered surfaces (UDF/SP/view) still work around.
- 窄场景
需要动态透视的报表/UDF/存储过程;列数随数据变化的宽表。
Reports/UDFs/procedures needing dynamic pivots; wide tables whose columns shift with data.
- 机制
`IN (ANY)` 解决了 adhoc 场景;但存储过程/UDF 编译时定列、view 上列数变化可能报错——动态透视补了一半。
`IN (ANY)` fixed the adhoc case; stored procedures/UDFs fix columns at compile time, views risk errors on column-count changes — dynamic pivot is half-fixed.
- 生产验证
来源 62:Retool 社区用户实战帖,2023——动态透视缺失期的 workaround 实录;
来源 63:独立从业者 Fabien Monnery 个人博客,2026-01——ANY 原理+限制实测。
—
- 证据等级
`多方印证(2 个独立来源)`,社区用户实战 + 独立从业者实测。
`Corroborated (2 independent sources)`, community user field post + independent practitioner measurement.
- 备注
缺口状态:部分补上(`PIVOT ... IN (ANY)` 已可用;存储过程/UDF 中仍不支持,view 上有报错风险)。本卡主题可能与本站 [避坑] 卡重叠。
gap status: partially fixed (`PIVOT ... IN (ANY)` available; still unsupported in procedures/UDFs, error risk on views). This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
"MySQL 兼容"的代价:分布式版的功能阉割清单
多方印证
升级迁移生态与信任
- 一句话
TDSQL for MySQL(分布式版)顶着"MySQL 兼容"招牌,但触发器、存储过程、外键、全文索引、自建分区、视图、自定义函数一律不支持——存量 MySQL 应用迁过去等于重写。
TDSQL for MySQL (distributed edition) flies the "MySQL compatible" flag, yet triggers, stored procedures, foreign keys, full-text indexes, self-built partitioning, views, and user-defined functions are all unsupported — migrating an existing MySQL application means rewriting it.
- 窄场景
存量 MySQL 应用(含存储过程/触发器/外键/全文索引)评估迁移到 TDSQL 分布式版。
Teams evaluating a move of an existing MySQL application (one that uses stored procedures/triggers/foreign keys/full-text indexes) to the TDSQL distributed edition.
- 机制
Proxy + 多 SET 的 Shared Nothing 架构下,跨分片语义无法廉价实现的能力被直接砍掉。官方《使用限制》(2024-01-06)列明:不支持自定义函数、事件、表空间、视图、存储过程、触发器、游标、外键、自建分区、复合语句(BEGIN END/LOOP);DML 层面不支持不带 WHERE 的 UPDATE/DELETE、LOAD DATA/XML、用户变量引用;管理 SQL 连 KILL、ANALYZE TABLE 都不支持(需走透传语法)。
Under the Proxy-plus-multiple-SET shared-nothing architecture, features whose cross-shard semantics cannot be implemented cheaply were cut outright. The official Usage Restrictions (Jan 6 2024) state: no user-defined functions, events, tablespaces, views, stored procedures, triggers, cursors, foreign keys, self-built partitioning, or compound statements (BEGIN END / LOOP); at the DML level, no WHERE-less UPDATE/DELETE, no LOAD DATA/XML, no user-variable references; even admin SQL like KILL and ANALYZE TABLE is unsupported (must use pass-through syntax).
- 生产验证
来源 1:独立开发者 colask 2026-05 选型调研(GitHub,逐条引用官方文档)——"需要保留触发器/存储过程/外键/全文索引 → 不能选 TDSQL for MySQL";"TDSQL for MySQL 主要痛点:兼容性损失大、需 shardkey 规划";"如果你现有 MySQL 应用代码里出现了上述任何一项,迁移到 TDSQL for MySQL 都意味着重写";
来源 2:社区 SQL 方言参考 wenshao/sql-dialects(2026)"已知不足"——"与 MySQL 的兼容性在分布式特性(如全局 AUTO_INCREMENT、跨分片外键)上存在限制"、"部分 MySQL 存储过程和触发器在分布式场景下行为可能不同"。
Source 1: independent developer colask, May 2026 selection study (GitHub, citing official docs item by item) — "Need to keep triggers/stored procedures/foreign keys/full-text indexes → do not choose TDSQL for MySQL"; "TDSQL for MySQL's main pain points: heavy compatibility loss, shard-key planning required"; "if any of the above appears in your existing MySQL application code, migrating to TDSQL for MySQL means a rewrite";
Source 2: community SQL-dialect reference wenshao/sql-dialects (2026), "Known limitations" — "MySQL compatibility has restrictions in distributed features (e.g., global AUTO_INCREMENT, cross-shard foreign keys)" and "some MySQL stored procedures and triggers may behave differently in distributed scenarios."
- 证据等级
`多方印证(2 个独立来源)`,独立选型评估 ×1 + 社区方言参考 ×1(官方《使用限制》文档仅作机制佐证,不计入来源数)。
`Corroborated (2 independent sources)`, independent selection evaluation x1 + community dialect reference x1 (the official Usage Restrictions doc is mechanism corroboration only and not counted as a source).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site on this site.
腾讯云 TDSQL 年份:2026
升级不能跳版本:1.24 直升 1.26,150 万条数据"凭空消失"
多方印证
升级迁移
- 一句话
Weaviate 官方只支持逐个 minor 版本升级——跳过任何一个 minor 版本,migration 脚本不跑,数据静默不可见,降级也救不回来。
Weaviate only supports upgrading one minor version at a time — skip any minor in between and migration scripts never run: data goes silently invisible, and downgrading won't bring it back.
- 窄场景
自部署 Weaviate(Helm/K8s/Docker Compose),版本滞后多个 minor 后想一次升到最新;多副本集群风险更高(单副本 staging 升级正常不代表生产集群没事)。
Self-hosted Weaviate (Helm/K8s/Docker Compose) that has fallen several minors behind and tries to jump straight to latest; multi-replica clusters are riskier (a clean single-replica staging upgrade says nothing about production).
- 机制
Weaviate 每次 minor 发布都可能带存储层/schema 迁移(如 1.25 引入 Raft、1.22 的 fs 层级迁移),升级流程默认只执行"上一版本→当前版本"的迁移链。跳过版本后,旧数据目录的新版本二进制仍然能启动(日志无报错),但部分 shard/class 的元数据迁移没做 → /v1/nodes 显示 0 对象、或查 class 报 "shard not found"。数据文件还在磁盘上,只是读不出来;直接降级同样读不出来(schema 格式已部分改写),唯一出路是恢复升级前的备份后逐级重升。
Every Weaviate minor may ship storage/schema migrations (1.25 introduced Raft, 1.22 had filesystem-level migrations), and the upgrade path only executes the "previous version → current version" migration chain. After skipping versions, the old data directory still boots under the new binary (no errors in the logs), but metadata migrations for some shards/classes never ran → /v1/nodes shows 0 objects, or querying a class returns "shard not found". The data files are still on disk — just unreadable; downgrading directly doesn't help either (the schema format was partially rewritten), so the only way out is restoring the pre-upgrade backup and climbing version by version.
- 生产验证
来源 1:2024-10,用户从 1.24.23 直接升到 1.26.4——150 万 chunks 变 0,无任何报错日志;官方客服确认"不允许跳版本",推荐路径 1.24→1.25.latest→1.26.latest;降级回 1.24.23 后依然 0,必须从备份恢复;
来源 2:2025-10,EKS 3 副本集群从 1.22.6 升 1.25.0——Raft schema 迁移不完整,19 个 class 里 7 个报 "shard not found" 不可访问,pod 全部正常启动无报错;staging 单副本升级正常,生产多副本才翻车,最终回滚。官方客服回复升级链要求 1.22.latest→1.23.latest→1.24.latest,并建议"数据量大的话直接搭新集群导数据更省事";
来源 3:2026-10 Dify 官方 troubleshooting 文档——Dify 1.17.1 内置 Weaviate 从 1.27.0 升到 1.39.2(跨 12 个 minor)被明确警告"必须逐个 minor 升级,直接重启会让向量检索静默永久损坏";文档还附带"修复孤儿 LSM 数据"的 shell 脚本——升级后部分 LSM 数据目录会被孤立,需要停机后手工 `cp -a` 搬回原位才能读到。
Source 1: Oct 2024, a user jumped from 1.24.23 straight to 1.26.4 — 1.5M chunks became 0, with zero error logs; official support confirmed "skipping versions is not allowed" and recommended 1.24 → 1.25.latest → 1.26.latest; downgrading back to 1.24.23 still showed 0, recovery from backup was the only option;
Source 2: Oct 2025, an EKS 3-replica cluster went 1.22.6 → 1.25.0 — incomplete Raft schema migration left 7 of 19 classes "shard not found" and unreachable, all pods started cleanly with no errors; the single-replica staging upgrade was fine, only production multi-replica broke, ending in rollback. Official support replied the upgrade chain must be 1.22.latest → 1.23.latest → 1.24.latest, and suggested "for large datasets it's easier to just stand up a new cluster and import";
Source 3: Oct 2026 Dify official troubleshooting doc — Dify 1.17.1's bundled Weaviate going 1.27.0 → 1.39.2 (12 minors) is explicitly warned "must upgrade minor by minor; restarting directly will silently and permanently corrupt vector search"; the doc even ships a shell script for "repairing orphaned LSM data" — after upgrading, some LSM data directories get orphaned and need downtime plus manual `cp -a` to be readable again.
- 证据等级
`多方印证(3 个独立来源)`,用户踩坑帖 ×2 + 第三方项目官方迁移文档。
`Corroborated (3 independent sources)`, user incident threads ×2 + third-party project's official migration doc.
Weaviate 年份:2026
Tablet 数量爆炸:索引点查被拖到 13-15 秒
多方印证
性能问题运维复杂度
- 一句话
tablet 太多时,连"按索引查一行"这种点查都会被拖到十几秒——分片元数据与后台开销吃掉了整台机器。
With too many tablets, even a "fetch one row by index" point lookup degrades to tens of seconds — sharding metadata and background overhead eat the whole machine.
- 窄场景
表数量多(2000 张非 colocated 表)、分区表多(单表 100+ 分区)、或建表时 SPLIT INTO 过度分片,导致单节点 tablet 副本数过万的大集群。
Large clusters with many tables (2,000 non-colocated tables), heavily partitioned tables (100+ partitions per table), or over-sharding via SPLIT INTO at creation — pushing per-node tablet replica counts past ten thousand.
- 机制
每个 tablet 是一个独立 Raft 组,附带内存中的 tablet 映射、后台 compaction 线程与心跳;tablet 数膨胀 → TServer 内存/线程压力、compaction 排队、tablet 路由查找变慢。用户实测:tablet 少时同一查询 Execution Time 仅 2.178ms,tablet 过万后多数点查 13-15 秒,且空载 CPU 达 15-20%。
Each tablet is an independent Raft group carrying an in-memory tablet map, background compaction threads, and heartbeats; tablet count inflation means TServer memory/thread pressure, compaction queuing, and slower tablet routing lookups. The user measured the same query at 2.178ms execution time with few tablets, but 13-15 seconds for most point lookups once tablets exceeded ten thousand per node, with 15-20% CPU burn at zero workload.
- 生产验证
来源 1:用户 Meteo 2024-12 发帖——43 节点(64C/256GB/6×NVMe),2000 张表,单节点 10,000+ tablet,"select one row by index"类查询多数 13-15 秒;官方人员回复 "You're most likely hitting tablet replica overhead",建议降低 tablet 数,但用户表示 tablet 数与表数成正比,"无法减少";
来源 2:letsbuildsolutions.com 2026-06 独立深挖——单 TServer 2,000-5,000 tablet 可管理,超过 10,000 会出现 compaction 延迟升高与内存压力,建议监控 tserver_tablets_num 并按节点内存封顶。
Source 1: user Meteo, Dec 2024 — 43 nodes (64C/256GB/6xNVMe), 2,000 tables, 10,000+ tablets per node; "select one row by index"-style queries took 13-15 seconds for the majority of queries. Yugabyte staff replied "You're most likely hitting tablet replica overhead" and advised lowering tablet counts, but the user said tablet count scales with table count and "we can not decrease the number of tablets";
Source 2: letsbuildsolutions.com, Jun 2026, independent deep dive — 2,000-5,000 tablets per TServer is manageable; above 10,000 per node expect elevated compaction latency and memory pressure; advises monitoring tserver_tablets_num and capping tablets against node memory.
- 证据等级
`多方印证(2 个独立来源)`,厂商论坛客户发帖 + 独立技术博客(注:论坛为厂商运营,执笔人为客户)。
`Corroborated (2 independent sources)`, customer thread on the vendor forum + independent tech blog (note: the forum is vendor-operated; the posts are customer-authored).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site on this site.
YugabyteDB 年份:2026
连接模型:YSQL 继承 PG 进程模型,连接数按节点算
多方印证
运维复杂度
- 一句话
YSQL 是"真 PG 进程",max_connections 默认 300/节点且按节点算——想撑 10 万连接,官方答案是"部署 350 个节点"或外挂连接池。
YSQL is a genuine Postgres process: max_connections defaults to 300 per node and is counted per node — for 100k connections the official answer is "deploy 350 nodes" or add an external pooler.
- 窄场景
物联网/网关类长连接、大并发短连接的应用;想直连、不想引入池化层的团队。
IoT/gateway-style long-lived connections or massive short-lived connection fan-out; teams wanting direct connections without a pooling layer.
- 机制
YSQL 复用 PG 的 process-per-connection 模型,每个物理后端是独立进程、占常驻内存;max_connections 按节点限制总内存(调大吃内存,挤占文件系统缓存)。长连接持续忙碌时池化层也无法复用(用户原话:pgBouncer 只是复用空闲连接,长期忙碌的连接池化帮不上忙)。
YSQL reuses Postgres's process-per-connection model — every physical backend is a separate process holding resident memory; max_connections caps total memory per node (raising it eats memory needed for filesystem cache). With long-lived busy connections, poolers can't help either (user's own words: pgBouncer only reuses idle connections; long-term busy connections can't be multiplexed away).
- 生产验证
来源 6:用户 Meteo 2024-05 售前咨询帖——100+ 地域组、1000+ Go 服务每 5 秒上报,问 10 万+ 并发连接怎么撑(文档 max_connections 默认 300/节点);官方回复:max_connections 是按节点算的,"with the default setting, you can have 100000+ connections if you deploy to 350 nodes";当时内置 Connection Manager 还是 tech preview,建议 PgBouncer/Odyssey;
来源 2:letsbuildsolutions.com 2026-06——"500+ raw connections saturate the TServer processes quickly",建议 PgBouncer 或内置 YSQL Connection Manager(recent versions 已内置)。
Source 6: user Meteo, May 2024 pre-sales thread — 100+ region groups, 1,000+ Go services reporting every 5 seconds, asking how to sustain 100k+ concurrent connections (docs say max_connections defaults to 300/node). Official reply: max_connections is per database node — "with the default setting, you can have 100000+ connections if you deploy to 350 nodes"; the built-in Connection Manager was still tech preview then, so PgBouncer/Odyssey was recommended;
Source 2: letsbuildsolutions.com, Jun 2026 — "500+ raw connections saturate the TServer processes quickly"; recommends PgBouncer or the built-in YSQL Connection Manager (built into recent versions).
- 证据等级
`多方印证(2 个独立来源)`,厂商论坛客户咨询帖 + 独立技术博客(如实标注:来源 6 为售前咨询帖,非生产事故)。
`Corroborated (2 independent sources)`, customer question thread on the vendor forum + independent tech blog (stated as-is: source 6 is a pre-sales question, not a production incident).
- 备注
后续版本内置 YSQL Connection Manager 已 GA(据独立博客,recent versions 已内置),本卡反映 2024-2025 年情况;本卡主题可能与本站 [避坑] 卡重叠。
the built-in YSQL Connection Manager has since gone GA (per the independent blog, built into recent versions); this card reflects the 2024-2025 situation; this card's topic may overlap an existing [Pitfall] card on this site on this site.
YugabyteDB 年份:2026
v1 停服强制迁 v2:迁移路这么绕,版本还可能被顺手升级
多方印证
升级迁移
- 一句话
v1 停服,AWS 给的"自动升级"发生在维护窗口里——如果你的引擎版本 v2 不支持,它会顺手给你升大版本。
For the v1 shutdown, AWS's "automatic upgrade" happens inside your maintenance window — and if your engine version isn't supported on v2, it upgrades your major engine version too.
- 窄场景
2024 年底前仍跑 Aurora Serverless v1 的集群(MySQL 5.7 / PG 13 及更早)。
Clusters still on Aurora Serverless v1 before end of 2024 (MySQL 5.7 / PG 13 and earlier).
- 机制
v1 与 v2 是两套架构(v1 搬迁式扩容,v2 常驻式细粒度扩容),没有原地一键切换;官方迁移指南承认迁移"operationally complex"且"需要不可忽视的停机";官方通知写明:到期未迁的集群将在维护窗口被自动升级到 v2,若引擎版本不被 v2 支持则一并升级引擎大版本。
v1 and v2 are two architectures (v1's relocate-to-scale vs v2's always-on fine-grained scaling) with no in-place one-click switch; AWS's own migration guide concedes the migration is "operationally complex" with "non-negligible downtime"; the official notice states unmigrated clusters will be auto-upgraded to v2 during the maintenance window, including an engine major-version bump when the running version isn't available on v2.
- 生产验证
来源 6:Datanami 2024-01-08 独立报道——AWS 邮件通知 v1 于 2024-12-31 停止支持(后延至 2025-03-31);JPMorgan 云架构负责人 Ganesh Swaminathan 具名吐槽"告别能自己缩到零的关系型数据库,你好,双倍账单(或更多)";
来源 7:re:Post 用户——"迁移路径极其绕"(migration path is so convoluted),质疑"既然 AWS 能自动迁,为什么现在不给我这个选项"。
Source 6: Datanami, Jan 8 2024, independent reporting — AWS emailed that v1 support ends Dec 31 2024 (later extended to Mar 31 2025); JPMorgan cloud architecture head Ganesh Swaminathan, named: "Goodbye to a relational database that could idle down to zero on its own. Hello to double the bill (or more)";
Source 7: re:Post user — "the migration path is so convoluted," asking "if Amazon could automatically migrate it, why not provide that option today?"
- 证据等级
`多方印证(2 个独立来源)`,独立媒体报道 + re:Post 用户。
`Corroborated (2 independent sources)`, independent media reporting + re:Post user.
- 备注
v2 自 2024 年底支持缩容到 0(自动暂停),"v2 不能缩零"的批评已部分缓解;但强制迁移本身的运维复杂度不受影响。
v2 gained scale-to-zero (auto-pause) in late 2024, partially answering the "v2 can't scale to zero" criticism; the operational complexity of the forced migration itself is unaffected.
Amazon Aurora 年份:2025
Serverless v2:0.5 ACU 常开,"serverless"但账单不睡觉
多方印证
成本账单
- 一句话
v2 上线时最低 0.5 ACU 全天在线——一个几乎没流量的库,一个月也要交约 45 美元的"占位费"。
At launch, v2 kept a minimum 0.5 ACU online around the clock — a nearly idle database still paid roughly $45/month just for existing.
- 窄场景
低频/间歇负载、dev/QA 环境;2024 年 11 月之前(尚无自动暂停)的 v2。
Low-frequency/intermittent workloads and dev/QA environments; v2 before Nov 2024 (pre auto-pause).
- 机制
v2 为消除 v1 的冷启动改为常驻式细粒度扩容,代价是最低 0.5 ACU(约 2GB 内存及对应 CPU/网络)永远在线、按秒计费;HN 用户测算约 $45/月保底,Reddit 用户约 $50/月。
To kill v1's cold starts, v2 moved to always-on fine-grained scaling — at the price of a permanent 0.5 ACU floor (~2 GiB memory plus corresponding CPU/network), billed by the second; HN users estimated ~$45/month minimum, a Reddit user ~$50/month.
- 生产验证
来源 4:HN 用户——"v2 价格翻倍","1 个开发者付不起 3 个 $45/月的 dev/QA 环境";
来源 6:Datanami 报道——JPMorgan 的 Ganesh Swaminathan:"你好,双倍账单(或更多)";Reddit 用户 zmose:"v2 似乎不能缩到 0,这可是 serverless 本该有的样子?最低也要 $50/月";
来源 8:独立工程师 HoangMNguyen 的 ADR(2025-01)——给 Aurora Serverless v2 配了 min 0 ACU 后实测:暂停后首次连接约 15 秒、单实例单 AZ,最终把库迁到了 Neon。
Source 4: HN user — "V2 has DOUBLED in price," "I'm 1 developer, I can't pay 3x$45/mo" for dev/QA environments;
Source 6: Datanami — JPMorgan's Ganesh Swaminathan: "Hello to double the bill (or more)"; Reddit user zmose: "v2 cannot seem to scale to 0 ACU — ya know, what 'serverless' is supposed to mean? …around like $50/month at minimum";
Source 8: independent engineer HoangMNguyen's ADR (Jan 2025) — after configuring min 0 ACU on Aurora Serverless v2: first connection after a pause takes ~15 s, single instance single AZ, and the database was ultimately moved to Neon.
- 证据等级
`多方印证(3 个独立来源)`,HN 用户 + 独立媒体引述 + 独立工程师 ADR。
`Corroborated (3 independent sources)`, HN user + independent media quotes + independent engineer ADR.
- 备注
2024 年底起 v2 支持 0 ACU 自动暂停,"不能缩零"已部分缓解;但唤醒约 15 秒、且有预置实例/在 Global Database 中/挂了 RDS Proxy 的集群不会暂停,"serverless"体验仍打折。
since late 2024 v2 supports 0-ACU auto-pause, partially answering "can't scale to zero"; but resume takes ~15 s, and clusters with provisioned instances, Global Database membership, or RDS Proxy attached never pause — the "serverless" experience is still discounted.
Amazon Aurora 年份:2025
GC STW 暂停:整节点停顿打爆尾延迟
多方印证
性能问题
- 一句话
JVM 的 Stop-The-World 暂停在 Cassandra 上不是"慢一点",而是整节点在集群视角短暂失联——病态读写模式触发的 GC 压力曾逼得一线团队在前面自研 Rust 查询层做限流,结果还是切到了 ScyllaDB。
JVM stop-the-world pauses on Cassandra are not "a little slower" — the node briefly goes dark from the cluster's perspective; one frontline team built a Rust query layer in front just to throttle pathological patterns, and eventually moved to ScyllaDB anyway.
- 窄场景
大分区(数百 MB ~ 1GB+)、高频访问热 key、墓碑堆积叠加的集群;CMS/G1 调优未跟上的老版本。
Large partitions (hundreds of MB to 1GB+), hot keys hit at high frequency, clusters where tombstone buildup compounds the above; older versions without tuned CMS/G1.
- 机制
查询命中分区时,Cassandra 把整分区(或大切片)从 SSTable 读出并反序列化为堆内 Java 对象(行、列、时间戳、墓碑全进堆);大分区 + 高并发 = young-gen 迅速撑爆、对象晋升老年代 → 频繁 Minor GC → 堆扛不住时 Full GC 全停顿,停顿期间该节点对集群表现为不可用。病态的租户读写模式会不成比例地放大这一效应(GC 压力 → 整节点 STW → 尾延迟雪崩)。
On a partition hit, Cassandra reads the whole partition (or large slices) from SSTables and deserializes it into on-heap Java objects — rows, columns, timestamps, tombstones all land on the heap; large partitions under high concurrency fill the young generation fast, objects promote to the old generation, frequent minor GCs escalate into full-GC pauses during which the node looks unavailable to the cluster. Pathological tenant read/write patterns amplify this disproportionately (GC pressure → whole-node STW → tail-latency collapse).
- 生产验证
来源 2:2022-12,HN 一线运维者——"pathological read/write patterns…triggering garbage collection pressure that would cause whole node GC STW pauses and severe tail latency";缓解手段包括自研 Rust 查询层(read coalescing + 限流/降载)以及最终"switching to ScyllaDB, a C++ rewrite of Cassandra which is of course void of any garbage collection issues";
来源 3:2025-05,客户生产实录——超 1GB 的大分区"contributing to long GC pauses";Full GC 期间 "Cassandra pauses all activity (Stop-The-World), and shows as node not available"。
Source 2: 2022-12, HN frontline operator — "pathological read/write patterns…triggering garbage collection pressure that would cause whole node GC STW pauses and severe tail latency"; mitigations included a self-built Rust query layer (read coalescing + throttling/load shedding) and ultimately "switching to ScyllaDB, a C++ rewrite of Cassandra which is of course void of any garbage collection issues";
Source 3: 2025-05, client production account — 1GB+ partitions "contributing to long GC pauses"; during full GC, "Cassandra pauses all activity (Stop-The-World), and shows as node not available."
- 证据等级
`多方印证(2 个独立来源)`,一线运维者评论 + 客户生产实录。
`Corroborated (2 independent sources)`, frontline operator comment + client production account.
Apache Cassandra / ScyllaDB 年份:2025
数据建模反直觉:关系型思维建模,必翻车
多方印证
运维复杂度
- 一句话
Cassandra 的建模铁律是"按查询建模"而非"按实体建模"——带着关系型范式思维进来的人,几乎注定要经历一次 schema 推倒重来;连文档都要"逐行审计",默认配置几乎没有能直接用的。
Cassandra's modeling law is "model for your queries," not "model your entities" — teams arriving with relational normalization habits are almost guaranteed a schema tear-down and rebuild; even the docs deserve a line-by-line audit, and almost none of the defaults are usable as-is.
- 窄场景
从关系型转过来的团队;按实体直觉建表、滥用二级索引、集合类型嵌套 UDT 的 schema。
Teams migrating from relational databases; schemas built on entity intuition, abused secondary indexes, or UDT-nested collection types.
- 机制
无 join、无子查询,查询必须命中完整分区键;"范式化"在这里是反模式,正确姿势是为每种查询模式建宽表、主动反范式化。二级索引是各节点本地索引,全局查询要 fan-out 到所有节点;物化视图有 read-before-write 开销与一致性边缘情况;`frozen<list<>>` 等集合类型误用会行为怪异。默认配置(分区大小、压缩策略、一致性级别)几乎都要按 workload 重调。
No joins, no subqueries — queries must hit the full partition key; "normalization" is an anti-pattern here, the correct posture is one wide denormalized table per query pattern. Secondary indexes are per-node local indexes, so global queries fan out to every node; materialized views carry read-before-write cost and consistency edge cases; misused collection types like `frozen<list<>>` behave oddly. Defaults (partition sizing, compaction strategy, consistency levels) almost all need retuning per workload.
- 生产验证
来源 6:2015-01,HN 用户——"you can't run Cassandra with almost any of its default settings…It's also almost mandatory to read the internal design docs of cassandra even if you're not the admin…And modelling data is a lot less trivial than it looks - and almost always not what you assume";另一用户:"Cassandra…is very sensitive to how you model and store your database. The whole tombstone saga is never a fun one";
来源 7:2025-06,why amit 迁移实录——"Review the Data Model Like You're the Auditor…review your schema line by line";`frozen<list<>>` "can get weird if misused, especially when combined with nested UDTs";二级索引"just… don't",全部改写为显式反范式化表才"Way more predictable"。
Source 6: 2015-01, HN users — "you can't run Cassandra with almost any of its default settings…It's also almost mandatory to read the internal design docs of cassandra even if you're not the admin…And modelling data is a lot less trivial than it looks - and almost always not what you assume"; another user: "Cassandra…is very sensitive to how you model and store your database. The whole tombstone saga is never a fun one";
Source 7: 2025-06, why amit's migration account — "Review the Data Model Like You're the Auditor…review your schema line by line"; `frozen<list<>>` "can get weird if misused, especially when combined with nested UDTs"; secondary indexes: "just… don't", everything rewritten as explicitly denormalized tables became "Way more predictable."
- 证据等级
`多方印证(2 个独立来源)`,社区用户 + 2025 年迁移实录(2015 年的建模痛点在 2025 年实录中依然成立)。
`Corroborated (2 independent sources)`, community users + a 2025 migration account (the 2015 modeling pain still holds in the 2025 account).
- 备注
与现有 [避坑] 卡主题重叠(CQL 无 join、查询必须事先建模 query-first),双方保留。
Topic overlaps with existing [Pitfall-avoidance] cards (no joins in CQL, query-first modeling is mandatory); both are kept.
Apache Cassandra / ScyllaDB 年份:2025
没有内置查询路由/负载均衡:自建集群必须在前面架 chproxy/NGINX
多方印证
运维复杂度
- 一句话
`max_concurrent_queries` 只在单节点生效——"集群层面没有任何办法限制并发查询数",用户不得不在集群前面维护 HTTP 代理,把 INSERT 打散、把 SELECT 发往可限流节点。
`max_concurrent_queries` is per-node only — "there is no way to limit the number of concurrent queries at the cluster level." Users maintain an HTTP proxy in front: scattering INSERTs, sending SELECTs to throttlable nodes.
- 窄场景
自建多节点集群;读写分离、故障转移需求。
Self-hosted multi-node clusters; read-write splitting and failover needs.
- 机制
chproxy 作者第一人称:"我们不得不在 ClickHouse 集群前面维护两个不同的 HTTP 代理……这很脆弱、很难管,于是就有了 chproxy。" `Distributed` 表只解决分片内数据定位,不解决跨节点并发限流/读写分离/故障转移。Tinybird 2025 年生产博客仍明示 load balancer 是架构关键。
chproxy's author, first-hand: "we had to maintain two different HTTP proxies in front of the ClickHouse cluster… it was fragile and hard to manage, so chproxy was born." The `Distributed` table only solves intra-shard data location — not cross-node concurrency limits, read-write splitting, or failover. Tinybird's 2025 production blog still calls the load balancer architecturally key.
- 生产验证
—
Source 29: Tinybird CTO Javi Santana, Apr 2025 — "The load balancer is the key to all of this.";
Source 10: same series Part II — "At Tinybird, we handle it with a combination of app logic and a load balancer.";
chproxy official docs, History (its author a ClickHouse production user).
- 证据等级
`多方印证(3 个独立来源)`,工具作者 + CTO 生产博客。
`Corroborated (3 independent sources)`, tool author + CTO production blogs.
ClickHouse 年份:2025
算错"计算类型":同一个集群,DBU 单价差 2–4 倍
多方印证
成本账单
- 一句话
把定时任务跑在交互式集群上,DBU 单价直接翻 2–4 倍——账单爆炸时,代码和集群规模看起来都"没问题"。
Run a scheduled job on an interactive cluster and the DBU rate jumps 2–4x — when the bill explodes, the code and the cluster size both "look fine."
- 窄场景
交互式(All-Purpose)集群被复用于定时 pipeline;从 notebook 开发环境直接搬上生产的团队。
Interactive (All-Purpose) clusters reused for scheduled pipelines; teams that lift a notebook dev environment straight into production.
- 机制
Databricks 按计算类型分类定价,All-Purpose Compute 的 DBU 费率显著高于 Jobs Compute(约 2–4 倍);Classic/Pro 仓库另有第二张云厂商 VM 账单。Auto-termination 只管"闲置"不管"忙而贵",选错计算类型在成本审计里看起来一切正常。
Databricks prices by compute classification — All-Purpose Compute carries a much higher DBU rate than Jobs Compute (roughly 2–4x); Classic/Pro warehouses add a second, separate cloud-VM bill on top. Auto-termination only catches idle clusters, not clusters that are busy being expensive, so the wrong compute type survives a cost review looking perfectly correct.
- 生产验证
—
Source 1: AT&T Israel tech blog, Dec 2025 — the team tracked DBU spend in Power BI, halved its Databricks cost in weeks, and listed "wrong compute type" among the systemic mistakes;
Source 2: Ahmed Youssef, Oct 2025 production postmortem — one job accounted for 70%+ of all job-related spend; moving automated workloads from All-Purpose to Jobs Compute was "the easiest win," cutting total spend by 80%+.
- 证据等级
`多方印证(2 个独立来源)`,具名团队博客 + 个人生产复盘。
`Corroborated (2 independent sources)`, named team blog + personal production postmortem.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
topic may overlap the existing [Pitfall] card.
Databricks 年份:2025
Delta 小文件:OPTIMIZE 是个隐形运维岗位
多方印证
来源存疑
运维复杂度
- 一句话
Delta 的小文件不会自己消失——OPTIMIZE/VACUUM/ZORDER 是每个团队都要自建的"隐形运维岗",排重了还会互相踩死。
Delta's small files never clean themselves up — OPTIMIZE/VACUUM/ZORDER is an "invisible ops role" every team must staff, and overlapping schedules can kill each other.
- 窄场景
流式/高频增量写入的 Delta 表(Silver 层);多团队共享表。
Streaming/high-frequency incremental writes into Delta tables (Silver layer); tables shared across teams.
- 机制
高频小批量写入天然产生大量小 Parquet 文件,查询变成元数据开销主导;OPTIMIZE 是重写数据的昂贵操作,需要调度、错峰、监控;VACUUM 负责清掉 OPTIMIZE 标记作废的旧文件,不跑则存储一直涨。维护任务本身无中心协调时,不同团队的 OPTIMIZE 会撞车争抢同一张表。
Frequent micro-batch writes naturally produce masses of tiny Parquet files, turning queries into metadata-overhead-bound scans; OPTIMIZE is an expensive full-rewrite operation needing scheduling, off-peak windows, and monitoring; VACUUM clears the stale files OPTIMIZE invalidates — skip it and storage grows forever. With no central coordination, different teams' OPTIMIZE jobs collide on the same table.
- 生产验证
—
Source 1: AT&T Israel, Dec 2025 — skipping OPTIMIZE/ZORDER/VACUUM leads to bloated tables and sluggish queries; the "small files problem" is a common bill driver, and poor storage layout inflates compute cost too;
Source 9: Rohit Pradhan, Dec 2025 — hourly incremental writes accumulated 100K+ files over 90 days, query latency up 5–8x; the article also contains a STAR-format "overlapping OPTIMIZE incident" narrative.
- 证据等级
`多方印证(2 个独立来源)[来源存疑:来源 9 的事故叙述发表于面试问答格式文章,其独立生产复盘身份无法确认]`,具名团队博客 + 个人技术文章。
`Corroborated (2 independent sources) [Questionable source: source 9's incident narrative appears in an interview-Q&A-format article; its status as an independent production postmortem cannot be confirmed]`, named team blog + personal technical article.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
topic may overlap the existing [Pitfall] card.
Databricks 年份:2025
自带 BI 太弱,分析被迫外挂 Power BI/Tableau(vs BigQuery + Looker)
多方印证
生态与信任
- 一句话
多名具名用户评价 Databricks 自带可视化 "average"、"not comparable to Power BI/Tableau"——分析工作流必须依赖外部 BI,而 BigQuery 有 Looker 原生集成。
Multiple named users rate Databricks' built-in visualization "average," "not comparable to Power BI/Tableau" — the analytics workflow depends on external BI, while BigQuery ships native Looker integration.
- 窄场景
想"一个平台走完数仓+分析"的团队;语义层/指标治理要原生落点的组织。
Teams wanting "one platform for warehouse + analytics"; orgs needing a native home for the semantic layer and metric governance.
- 机制
Databricks Dashboards(原 Lakeview)仍在追赶;语义层无原生落点,指标定义散落在各 BI 工具里。BigQuery + LookML 是"模型即治理"的对照组。
Databricks Dashboards (ex-Lakeview) are still catching up; the semantic layer has no native home, so metric definitions scatter across BI tools. BigQuery + LookML is the "model as governance" reference.
- 生产验证
—
Source 7: PeerSpot named reviews — Jithin James (Juniper Networks, Mar 2024): "It has an average dashboarding tool… They can bring advanced features so we don't depend on other BI tools"; Roberto Messora (Jakala, Oct 2022): "It's not comparable to a solution like Power BI, Luca, or Tableau."; Sandesh Nagaraj (Dec 2020): "It would be better to do it right in Databricks.";
Source 47: Data Mediator, Nov 2025 migration reflection — "Databricks gives you data engineering power, but analytics often depend on external tools like Power BI, Tableau, or Looker Studio. BigQuery, however, integrates natively with Looker."
- 证据等级
`多方印证(4 个独立来源)`,3 具名复评(2020→2024 口径一致)+ 1 迁移博客。
`Corroborated (4 independent sources)`, 3 named reviews (2020→2024, consistent) + 1 migration blog.
Databricks 年份:2025
"号称 drop-in replacement":MySQL 8 ↔ MariaDB 早就不是一个东西
多方印证
升级迁移生态与信任
- 一句话
从 MySQL 8 迁到 MariaDB(或反向),dump 不再是直通车——认证、JSON、GTID、字符集四处硬分化,"无缝替换"只剩历史叙事。
Migrating MySQL 8 ↔ MariaDB in either direction, dumps are no longer a straight shot — authentication, JSON, GTID, and collations all hard-diverge, and the "drop-in replacement" story is now history.
- 窄场景
MySQL 8.x ↔ MariaDB 10.5+ 之间任何方向的迁移;习惯用 Navicat/mysqldump 直导直灌的团队。
Any-direction migration between MySQL 8.x and MariaDB 10.5+; teams used to exporting from one and importing into the other with Navicat/mysqldump.
- 机制
分叉十几年后的硬分化点:①JSON——MySQL 是二进制原生类型,MariaDB 是 LONGTEXT 别名,不给 JSON 列补 `CHECK (JSON_VALID(col))` 后续非 JSON 写入会静默污染;②认证——MySQL 8 默认 caching_sha2_password,MariaDB 11.4 之前没有兼容插件,用户密码要轮换重建;③GTID——MySQL 的 `uuid:seqno` vs MariaDB 的 `domain-server-sequence`,格式不互通,跨库复制无平滑路径;④字符集——`utf8mb4_0900_ai_ci` 在 MariaDB 11.4.5 之前根本不存在;⑤sqldump 互通已死——来源 1 原话:"drifted apart drastically and there is no longer a direct pathway from one to the other without using enterprise solutions"。
Hard divergence points after a decade-plus fork: 1) JSON — a native binary type in MySQL, a LONGTEXT alias in MariaDB, where non-JSON writes silently pollute a column unless you add `CHECK (JSON_VALID(col))`; 2) authentication — MySQL 8 defaults to caching_sha2_password, which had no compatible plugin in MariaDB before 11.4, so user passwords need rotation and rebuild; 3) GTID — MySQL's `uuid:seqno` vs MariaDB's `domain-server-sequence` are incompatible formats with no smooth cross-server replication path; 4) collations — `utf8mb4_0900_ai_ci` simply did not exist in MariaDB before 11.4.5; 5) sqldump interop is dead — Source 1's own words: "drifted apart drastically and there is no longer a direct pathway from one to the other without using enterprise solutions".
- 生产验证
来源 1:infophreak(SH3LL)MariaDB 11→MySQL 8 迁移指南,开篇即承认两者 sqldump 曾互通、如今"漂移严重、无直接路径",只能走应用层导出/导入;
来源 2:CSDN 2025-11,MySQL 8.0→MariaDB 10.5 签到系统迁移——导表第一条就报 `1273 - Unknown collation: 'utf8mb4_0900_ai_ci'`,随后 ALTER 权限不足、外键约束拦截、`Illegal mix of collations` 连环坑,全程手工修脚本。
Source 1: infophreak (SH3LL) MariaDB 11 → MySQL 8 migration guide opens by conceding the two used to exchange sqldumps and now "drifted apart drastically", forcing app-level export/import;
Source 2: CSDN 2025-11, a MySQL 8.0 → MariaDB 10.5 migration of a sign-in system — the first import statement failed with `1273 - Unknown collation: 'utf8mb4_0900_ai_ci'`, followed by ALTER-privilege errors, foreign-key blocks, and `Illegal mix of collations`, all patched by hand-editing scripts.
- 证据等级
`多方印证(2 个独立来源)`,个人博客/迁移指南 ×2。
`Multi-source corroborated (2 independent sources)`, personal blogs / migration guides ×2.
- 备注
与现有 [避坑] 卡主题重叠(档案吐槽清单已含 GTID/JSON/认证三行),本卡补上真实迁移现场的连环坑细节。
Overlaps the existing profile's "pitfall" list (GTID/JSON/authentication rows); this card adds the real-world chain-reaction detail from migration war stories.
MariaDB 年份:2025
SkySQL/Xpand 说砍就砍:云业务是战略,客户是成本
多方印证
生态与信任
- 一句话
2023 年 10 月 MariaDB plc 砍掉 SkySQL 与 Xpand、裁员 28%,把"战略级产品"的用户(包括跑着 50 个 Xpand 节点的三星)晾在迁移计划上。
In October 2023 MariaDB plc discontinued SkySQL and Xpand and laid off 28% of staff, leaving "strategic product" users — including Samsung with 50 Xpand nodes — stranded on a migration plan.
- 窄场景
2020–2023 年间采纳 MariaDB SkySQL / Xpand 的企业客户。
Enterprise customers who adopted MariaDB SkySQL / Xpand between 2020 and 2023.
- 机制
SPAC 上市后财务崩盘(市值从 4.45 亿美元跌到约 1000 万美元,NYSE 发出退市警告),公司为降本砍掉"未来增长引擎":SkySQL(2020 年发布,对标云厂商 RDS)与 Xpand(2021 年加入的分布式后端,5 个月前刚推出 PG 兼容前端)。SEC 文件措辞是"帮助现有客户迁出"——即断供。
Post-SPAC financial collapse (market cap fell from USD 445M to roughly USD 10M, NYSE delisting warning) forced the company to cut its "future growth engines": SkySQL (launched 2020 as the answer to cloud-vendor RDS) and Xpand (the distributed backend added in 2021, whose PostgreSQL-compatible frontend had launched just five months earlier). The SEC filing's euphemism was "help existing customers migrate off" — i.e. end of supply.
- 生产验证
来源 6:The Register 2023-10,引 SEC 文件——停售 SkySQL+Xpand,裁员 84 人(28%),靠 2600 万美元、年息 10% 的贷款续命;三星 50 个 Xpand 节点、日均数百亿事务受影响;
来源 7:InfoWorld 2024-02,分析师 Tony Baer(dbInsight):"客户被迫对公司失去信心……会担心自己的投入会不会是下一个被砍的";Matt Aslett(Ventana/ISG)称影响"深远";同期微软宣布 Azure Database for MariaDB 于 2025-09-19 退役。
Source 6: The Register 2023-10, citing the SEC filing — SkySQL+Xpand discontinued, 84 people (28%) laid off, survival on a USD 26.5M loan at 10% interest; Samsung's 50 Xpand nodes processing hundreds of billions of transactions daily affected;
Source 7: InfoWorld 2024-02, analyst quotes — Tony Baer (dbInsight): the discontinuation "forced enterprise customers to lose confidence… they're going to be insecure about whether their investments will be next" (author quote); Matt Aslett (Ventana/ISG) called the impact "far-reaching"; in the same period Microsoft announced Azure Database for MariaDB would retire on 2025-09-19.
- 证据等级
`多方印证(2 个独立来源)`,独立媒体 ×2(含具名分析师引述)。
`Multi-source corroborated (2 independent sources)`, independent media ×2 (including named analyst quotes).
- 备注
事件已属历史(公司 2024 年私有化重组),但"战略产品说砍就砍"的信任折价仍在;与档案判决中"MariaDB Cloud 2025-08 买回重建、连续性风险待验证"主题相关。
The event is now historical (the company was taken private and restructured in 2024), but the "strategic products can be cut overnight" trust discount persists; related to the profile verdict's "MariaDB Cloud was repurchased and rebuilt in 2025-08, continuity risk TBD" theme.
MariaDB 年份:2025
CVE-2025-64513:一个 HTTP header 就绕过全部鉴权
多方印证
已修复于 2.4.24
生态与信任
- 一句话
Milvus Proxy 把"内部组件流量"和"外部客户端流量"的区分押在一个客户端可控的 header 上——伪造该 header,即可无密码拿到整个集群的管理员权限。
Milvus Proxy staked the distinction between "internal component traffic" and "external client traffic" on a client-controlled header — forge that header and you get full administrator access to the whole cluster with no password.
- 窄场景
Proxy/gRPC 端口暴露在不可信网络、且版本在 2.4.24 / 2.5.21 / 2.6.5 之前的任何部署;开了鉴权也挡不住这次绕过。
Any deployment with the Proxy/gRPC port exposed to an untrusted network and running a version before 2.4.24 / 2.5.21 / 2.6.5; enabling authentication did not stop this bypass.
- 机制
Proxy 的认证拦截器读取请求元数据中的 `sourceID`,base64 解码后与硬编码常量 `@@milvus-member@@` 比对;命中则判定为内部组件流量,直接跳过标准鉴权。该校验只认 header 值、不验证连接的真实身份,因此任何能向 Proxy 发请求的人都能伪造。官方公告明确:攻击者可借此读写删数据、管理数据库与集合;来源 2 的复现在 v2.4.23 上用伪造 header 成功建库,v2.4.24 上同样的请求被拒绝。
The Proxy's authentication interceptor reads the `sourceID` field from request metadata, base64-decodes it, and compares it against the hardcoded constant `@@milvus-member@@`; a match is treated as internal component traffic and skips standard authentication entirely. The check trusts the header value without verifying the connection's real identity, so anyone able to send requests to the Proxy can forge it. Per the official advisory, an attacker could read, modify, or delete data and manage databases and collections; Source 2's reproduction created a database with a forged header on v2.4.23, and the identical request was rejected on v2.4.24.
- 生产验证
来源 1:官方安全公告——未认证攻击者可绕过 Proxy 全部鉴权机制,获得集群完全管理权限;影响所有受影响版本用户,强烈建议立即升级;
来源 2:2025-11-13 由字节跳动 Volcengine 团队负责任披露,CVSS 9.3;披露与补丁同步发布;
来源 3:独立漏洞库交叉印证影响版本(2.4.24 / 2.5.21 / 2.6.5 之前)与修复版本一致,NVD 收录日期 2025-11-10。
Source 1: official security advisory — an unauthenticated attacker could bypass all authentication mechanisms in the Milvus Proxy component, gaining full administrative access; all users on affected versions strongly advised to upgrade immediately;
Source 2: responsibly disclosed on Nov 13, 2025 by ByteDance's Volcengine team, CVSS 9.3; disclosure and patches shipped simultaneously;
Source 3: independent vulnerability database corroborates the affected versions (before 2.4.24 / 2.5.21 / 2.6.5) and the fixed versions, NVD publication date Nov 10, 2025.
- 证据等级
`多方印证(3 个独立来源)`,官方安全公告 ×1 + 独立安全分析 ×2。
`Corroborated (3 independent sources)`, official security advisory x1 + independent security analyses x2.
- 备注
已修复于 2.4.24 / 2.5.21 / 2.6.5(2025-11),按规则不删除、作"已修复"标注。来源 2 系竞品厂商(CyborgDB)博客,技术细节与 GHSA 一致,其营销性结论("应用内加密才是正解")已剔除未引用;来源 3 页面自带 AI 生成声明,仅作版本与日期交叉印证。本卡主题可能与本站 [避坑] 卡重叠。
fixed in 2.4.24 / 2.5.21 / 2.6.5 (Nov 2025); retained with a "fixed in" label per the rules. Source 2 is a competitor vendor (CyborgDB) blog; its technical details match the GHSA, and its marketing conclusions ("encryption-in-use is the answer") were excluded. Source 3 carries an AI-generated-content notice and was used only to cross-check versions and dates. May overlap an existing [Pitfall] card on this site.
Milvus 年份:2025
改个分词器 = 重建整个集合:Milvus 没有 ALTER,只有"复制-交换"
多方印证
升级迁移
- 一句话
在 Milvus 里调个 BM25 analyzer 参数、改 schema 属性,没有原地 ALTER——标准做法是全量复制到新集合、校验行数、改名交换、旧集合留作备份,迁移期间应用还得停写。
Tune a BM25 analyzer parameter or change a schema attribute in Milvus and there is no in-place ALTER — the standard procedure is a full copy into a new collection, row-count verification, rename-and-swap with the old collection kept as backup, and the application must stop writing during the migration.
- 窄场景
用了全文/BM25 稀疏索引需要调 analyzer 参数,或任何需要变更 schema 属性(非简单加减字段)的集合。
Collections using full-text/BM25 sparse indexes that need analyzer tuning, or any collection needing a schema-attribute change (beyond simply adding fields).
- 机制
Milvus 的 `analyzer_params` 等 schema 属性不支持原地变更(开着 text match 时连关都关不掉)。独立项目 OpenRAG 的迁移文档记录的标准流程是"重建集合":把行复制到 `<collection>_v2_rebuild`、校验行数、改名交换、旧集合保留为备份;复制读的是快照,迁移中途写入的行会丢失(行数对不上则中止重试),所以必须先停应用;两份拷贝同时存在,需预留约 2 倍磁盘。
Milvus does not support in-place changes to schema attributes such as `analyzer_params` (it cannot even be turned off while text match is enabled). The independent OpenRAG project's migration doc records the standard procedure as a collection rebuild: copy rows into `<collection>_v2_rebuild`, verify row counts, rename-and-swap, keep the old collection as a backup; the copy reads a snapshot, so rows written mid-migration would be lost (the migration aborts and retries if counts move) — hence the application must be stopped first; both copies coexist, so roughly 2x disk is required.
- 生产验证
来源 5:2025-11 Software Finder 认证客户评价(中型企业用户 Issa M.)原话:"Changing collection schemas requires going through a migration process and that can be fairly complicated and pretty time-consuming to handle."(引用客户原话);
来源 6:Linagora OpenRAG 迁移文档——版本 2 的迁移 "rebuilds the collection rather than altering it",并列出停机、双倍磁盘、停应用、先备份等注意事项。
Source 5: Nov 2025 Software Finder verified customer review (Issa M., mid-market), quoted verbatim: "Changing collection schemas requires going through a migration process and that can be fairly complicated and pretty time-consuming to handle.";
Source 6: Linagora OpenRAG migration doc — the version-2 migration "rebuilds the collection rather than altering it", listing downtime, double disk, stop-the-app, and back-up-first caveats.
- 证据等级
`多方印证(2 个独立来源)`,认证客户评价 ×1 + 独立开源项目迁移文档 ×1。
`Corroborated (2 independent sources)`, verified customer review x1 + independent open-source project migration doc x1.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
may overlap an existing [Pitfall] card on this site.
Milvus 年份:2025
执行计划说翻就翻:复杂 SQL 计划突变,只能靠 DBA 人工干预
多方印证
性能问题
- 一句话
复杂语句的执行计划容易突然变化导致性能问题,稳住计划要靠管理员手工绑定。
Plans for complex statements are prone to sudden changes that cause performance incidents, and the only way to pin them down is manual plan binding by an administrator.
- 窄场景
多表关联、分区表、带复制表的复杂查询;统计信息/参数变化频繁的业务库。
Multi-table joins, partitioned tables, queries touching replicated tables; business databases with frequently changing statistics or parameters.
- 机制
优化器代价评估依赖分区采样(partition_index_dive_limit)与统计信息,采样分区小或数据为空时易误判;同一 SQL 模板在不同参数下复用同一计划(计划缓存),长尾参数直接翻车;含复制表的 SQL 在特定版本下无法复用执行计划;local rescan 规则曾错误裁剪掉低代价的 NLJ 计划。官方给出的解法是执行计划绑定(plan binding)把计划"钉死",等于把优化器的活转嫁给 DBA。
The optimizer's cost model depends on partition sampling (partition_index_dive_limit) and statistics — small sampled partitions or empty data skew the estimates; one cached plan is reused for the same SQL template under different parameters, so long-tail parameters blow up; SQL touching replicated tables could not reuse plans in certain versions; a local-rescan rule once wrongly pruned the cheaper nested-loop plan. The official remedy is plan binding, which moves the optimizer's job onto the DBA.
- 生产验证
来源 1:Gartner Peer Insights 验证用户原话——"Execution Plan: Complex statement execution plans are prone to sudden changes, leading to performance issues and requiring administrator intervention.";
来源 2:百丽/卢文豪 2025-11-25——执行计划问题涉及 partition_index_dive_limit 采样误判、同一计划用于不同参数(用 /*+USE_PLAN_CACHE(NONE)*/ 规避但引入硬解析 CPU 开销)、IN 参数过多导致硬解析慢(调 _inlist_rewrite_threshold);另有三个版本 bug:local rescan 计划不优(已修复于 V4.2.5.3)、tablegroup 调整遇空分区导致均衡任务卡住(已修复于 V4.2.5.4)、含复制表 SQL 无法复用执行计划(已修复于 V4.2.5.5)。
Source 1: Gartner Peer Insights validated user, verbatim — "Execution Plan: Complex statement execution plans are prone to sudden changes, leading to performance issues and requiring administrator intervention.";
Source 2: Belle/Lu Wenhao, Nov 25 2025 — plan problems included partition_index_dive_limit sampling misestimates, one plan reused across different parameters (worked around with /*+USE_PLAN_CACHE(NONE)*/ at the cost of hard-parse CPU), slow hard parsing with many IN-list parameters (tune _inlist_rewrite_threshold); plus three version bugs: suboptimal local-rescan plans (fixed in V4.2.5.3), rebalance jobs stuck when a tablegroup adjustment met empty partitions (fixed in V4.2.5.4), and plans not reusable for SQL over replicated tables (fixed in V4.2.5.5).
- 证据等级
`多方印证(2 个独立来源)`,Gartner 验证用户评价 + 具名客户生产复盘。
`Corroborated (2 independent sources)`, Gartner validated user review + named customer production postmortem.
OceanBase 年份:2025
Terraform provider:社区野路子扛了多年,官方 v1 直到 2024 年底才来
多方印证
已修复于 v1.0.0
运维复杂度
- 一句话
2022 年实战文:"the Snowflake provider for Terraform is documented, but you will come to find out that not all the resources are updated or that some contain copy-paste bugs sometimes. Just don't take everything for granted";背景是 Snowflake 官方长期没有自己的 provider,全靠社区(chanzuckerberg)的野路子实现;Snowflake-Labs 接手后 0.x 一路 breaking change,有用户在 GitHub issue 直言:"It also degrades our overall confidence in using the Snowflake terraform provider in general because it shows that Snowflake is not maintaining a stable interface."——直到 2024 年底 v1.0.0 发布、2025 年 4 月 GA,官方支持才算落地。
A 2022 field piece: "the Snowflake provider for Terraform is documented, but you will come to find out that not all the resources are updated or that some contain copy-paste bugs sometimes. Just don't take everything for granted"; the backstory: Snowflake had no official provider for years — the community (chanzuckerberg) wildcat implementation carried it; after Snowflake-Labs took over, 0.x shipped breaking change after breaking change, and a user wrote in a GitHub issue: "It also degrades our overall confidence in using the Snowflake terraform provider in general because it shows that Snowflake is not maintaining a stable interface." Official support only landed with v1.0.0 (end of 2024) and GA (Apr 2025).
- 窄场景
用 Terraform 管 Snowflake 的平台团队;2024 年底之前上 IaC 的用户。
Platform teams managing Snowflake with Terraform; anyone who adopted IaC before end-2024.
- 机制
IaC 这条线用户自己凑合了好几年:社区实现→Snowflake-Labs 接手→0.x breaking change不断→v1.0.0(2024 年底)→GA(2025-04)。
Users patched it together for years on the IaC front: community implementation → Snowflake-Labs takeover → endless 0.x breaking changes → v1.0.0 (end 2024) → GA (2025-04).
- 生产验证
来源 57:dataroots dev.to 实战文,2022-05——最早一批系统记录 provider 坑的独立文章;
来源 58:Snowflake-Labs 官方仓库 issue(约 2025-02),用户原话见上;
来源 59:官方 ROADMAP,2025-04-10 GA 公告。
—
- 证据等级
`多方印证(3 个独立来源)`,咨询公司实战文 + 用户 issue 原话 + 官方 ROADMAP 时间线。
`Corroborated (3 independent sources)`, consultancy field piece + user issue verbatim + official ROADMAP timeline.
- 备注
缺口状态:已补上于 v1.0.0(2024 年底发布)、GA 于 2025-04;补上之前用户凑合多年,故保留收录。本卡主题可能与本站 [避坑] 卡重叠。
gap status: fixed in v1.0.0 (end 2024), GA Apr 2025; kept to record the years users patched it together. This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2025
延迟税:ACID 永远在线,低延迟永远不在线
多方印证
性能问题
- 一句话
TrueTime 的外部一致性不是免费的——每次提交都要"等时钟不确定性过去",跨区写再叠加共识往返,写延迟天然比单机库高一个数量级。
TrueTime's external consistency is not free — every commit waits out the clock uncertainty, and cross-region writes stack consensus round trips on top, so write latency is inherently an order of magnitude above single-node databases.
- 窄场景
写密集、对写入延迟敏感、跨区部署的 OLTP;指望"全球一致还和本地 PG 一样快"的团队。
Write-heavy, latency-sensitive OLTP across regions; teams expecting "global consistency at local-Postgres speed."
- 机制
TrueTime 返回的是时间区间 [earliest, latest](ε 通常 1–7ms);提交时取 commit timestamp s = latest,然后 commit-wait 到全系统时钟都越过 s 才向客户端确认——每次写入固定支付约 2ε 的等待;跨区还要再付 Paxos 多数派复制的 WAN 往返。读默认是强读,跨区强读可能还要先找 leader 确认最新位点。
TrueTime returns a time interval [earliest, latest] (epsilon typically 1–7 ms); on commit, Spanner picks a commit timestamp s = latest, then commit-waits until every clock in the system has passed s before acknowledging — a fixed ~2-epsilon wait on every write; cross-region writes additionally pay Paxos quorum WAN round trips. Reads default to strong reads, which may need to check with the leader for the latest position across regions.
- 生产验证
来源 1(HN 2023):"It's the best, but it's far from perfect. Default mode is non-ACID, and going serializable mode makes it very slow. Spanner is always ACID... but always slow.";
来源 4(Medium 2025-04,独立从业者):"Spanner's TrueTime API enables global consistency... but it also introduces some latency, especially for global writes",并给出设计建议 "Be mindful of cross-region writes — writes that span regions will always incur higher latency due to consensus and time synchronization"。
Source 1 (HN, 2023): "It's the best, but it's far from perfect. Default mode is non-ACID, and going serializable mode makes it very slow. Spanner is always ACID... but always slow.";
Source 4 (Medium, Apr 2025, independent practitioner): "Spanner's TrueTime API enables global consistency... but it also introduces some latency, especially for global writes", with the design advice "Be mindful of cross-region writes — writes that span regions will always incur higher latency due to consensus and time synchronization".
- 证据等级
`多方印证(2 个独立来源)`,HN 用户 + 独立从业者博客。
`Corroborated (2 independent sources)`, one HN user plus one independent practitioner blog.
Google Spanner 年份:2025
按核许可:想读扩展或加硬件,先给微软写支票
多方印证
来源存疑
成本账单
- 一句话
2012 年从按 socket 改按核计费后,Always On 读扩展的数学就崩了——5 节点 AG 光许可就要 41 万美元,客户直接把 AG 从项目清单上划掉。
After the 2012 switch from per-socket to per-core licensing, the math for scaling out with Always On read replicas broke — a 5-node AG costs $412,440 in licenses alone, so customers crossed AGs off their project lists.
- 窄场景
想用 AG 可读副本做读扩展、或靠堆硬件掩盖性能问题的团队;ISV 客户(最终用户要为每核买单,堆内存不再便宜)。
Teams wanting AG readable secondaries for read scale-out, or masking performance problems by throwing hardware at them; ISV customers (end users pay per core, so adding memory is no longer cheap).
- 机制
Enterprise 按核许可(当前约 $7k/核),AG 的每个可读副本都要全额许可(仅被动故障转移副本可免)。Brent Ozar 算过账:2-socket×6 核的服务器组 5 节点 AG,许可费 $412,440——"简直不可想象"。Standard 版又有 16 核/128GB 上限(后放宽到 24 核),硬件堆到头就只能跳 Enterprise。于是 DBA 的经典动作——"加内存掩盖问题"——变成了先写支票的财务决策。
Enterprise is licensed per core (roughly $7k/core today); every readable AG replica needs a full license (only purely passive failover replicas are exempt). Brent Ozar did the math: a 5-node AG on 2-socket x 6-core servers costs $412,440 in licensing — "simply unthinkable." Standard Edition caps out at 16 cores/128GB (later relaxed to 24 cores), so once hardware tops out you must jump to Enterprise. The classic DBA move — "hide the sins with memory" — became a finance decision first.
- 生产验证
来源 1:Brent Ozar 2011-12 与客户访谈——客户两类反应,一类直接把 Availability Groups 从 2012 年项目清单划掉("pricing just makes this a no-go"),一类 ISV 抱怨客户"没法花 $1000 加内存就获得好性能";
来源 2:2025-10 个人转库实录——"SQL Server's licensing model quickly became restrictive… Scaling up meant moving to costly editions",免费 Express 版又有库大小/内存/CPU 三重上限,规模一涨就被迫跳付费版。
Source 1: Brent Ozar, Dec 2011, after interviewing clients — two reactions: some crossed Availability Groups off their 2012 project lists entirely ("pricing just makes this a no-go"), and ISV clients complained their customers "can't just buy $1,000 worth of memory to get awesome performance";
Source 2: Oct 2025 personal migration account — "SQL Server's licensing model quickly became restrictive… Scaling up meant moving to costly editions," while the free Express edition caps database size, memory, and CPU, forcing an expensive edition jump as workloads grow.
- 证据等级
`多方印证(2 个独立来源)`,具名顾问客户现场证据 ×1 + [来源存疑]个人博客 ×1
`Corroborated (2 independent sources)`, named consultant client field evidence x1 + [Questionable source] personal blog x1.
- 备注
与现有 [避坑] 卡「授权复杂度 —— 微软式"税"」主题重叠;该卡曾诚实标注"独立生产账单故事缺失",本卡补充了具名顾问的客户场景与报价(5 节点 AG $412,440),可视为该证据缺口的一次填补。另:来源 1 发表于 2011 年,但按核计费模型自 2012 年起未变(2025 年定价页仍为 Enterprise $13,748/2 核包),属结构性问题而非已修复缺陷,故保留收录。
overlaps the existing [Pitfall] card "Licensing complexity — the Microsoft tax," which honestly flagged "no named production billing stories"; this card fills that gap with a named consultant's client scenario and quote ($412,440 for a 5-node AG). Also: Source 1 is from 2011, but the per-core model has been unchanged since 2012 (the 2025 pricing page still lists Enterprise at $13,748 per 2-core pack) — a structural issue, not a fixed defect, so it is retained.
Microsoft SQL Server 年份:2025
Linux 版"二等公民":首发功能缺席,性能/兼容弱于 Windows
多方印证
来源存疑
性能问题生态与信任
- 一句话
SQL Server on Linux 跑在 SQLPAL 兼容层上——2017 年首发就没有 SSRS/SSAS/复制/Stretch DB,八年后用户仍在抱怨 Linux 版性能和兼容性"明显弱于 Windows"。
SQL Server on Linux runs on the SQLPAL compatibility layer — the 2017 launch shipped without SSRS/SSAS, replication, or Stretch DB, and eight years later users still complain the Linux edition's performance and compatibility are "noticeably weaker" than Windows.
- 窄场景
Linux 标准化企业、容器化部署团队;从 Windows 迁移到 Linux 指望"同一份许可换个 OS"的用户。
Linux-standardized enterprises and containerized deployments; anyone migrating from Windows to Linux expecting "the same license on a different OS."
- 机制
Linux 版通过 SQLPAL(Platform Abstraction Layer)把 Windows 库调用翻译到 Linux,SOS 内存/线程管理直接调 Linux API。架构上能跑,但周边生态长期缺席:2017 年首发无 SSRS/SSAS/机器学习服务、无复制(HA 场景除外)、无 Stretch DB、FileTable 不工作、管理工具基本只有 Windows 版。微软当时给 Linux 版定的性能目标只是"达到 Windows 的 75% 以上"。
The Linux port translates Windows library calls through SQLPAL (Platform Abstraction Layer), with the SOS memory/thread manager calling native Linux APIs. It runs, but the surrounding ecosystem was long absent: at the 2017 launch there were no Reporting/Analysis Services, no Machine Learning Services, no replication (outside HA), no Stretch DB, no FileTable, and management tools were mostly Windows-only. Microsoft's stated performance target for the Linux port was merely "75% or better of Windows."
- 生产验证
来源 3:The Register 2017-09 首发报道——缺席功能清单如上;微软总经理 Rohan Kumar 对 SSAS/SSRS 的回应是"问题是,有需求吗?"("What's the demand?"),暗示短期不会补;
来源 2:2025-10 个人转库实录——"SQL Server does support Linux, but performance and compatibility were noticeably weaker compared to Windows",这是其转向 PG 的原因之一。
Source 3: The Register, Sep 2017 launch coverage — the missing-feature list above; when asked about SSAS/SSRS on Linux, GM Rohan Kumar answered "What's the demand?", signaling no near-term port;
Source 2: Oct 2025 personal migration account — "SQL Server does support Linux, but performance and compatibility were noticeably weaker compared to Windows," one reason for switching to Postgres.
- 证据等级
`多方印证(2 个独立来源)`,独立媒体报道 ×1 + [来源存疑]个人博客 ×1
`Corroborated (2 independent sources)`, independent media coverage x1 + [Questionable source] personal blog x1.
- 备注
2017 年的具体缺席清单随版本已部分收窄(如复制、SQL Agent 后续补上),但"Linux 版弱于 Windows"的抱怨从 2017 持续到 2025,故保留收录;本卡仅针对 SQL Server on Linux,不涉及 Windows 版。
the specific 2017 gap list has narrowed across versions (replication, SQL Agent were added later), but the "Linux weaker than Windows" complaint persisted from 2017 to 2025, so it is retained. This card covers SQL Server on Linux only, not the Windows edition.
Microsoft SQL Server 年份:2025
Azure SQL MI:"永远最新版"的营销,现实是 2019 的功能 2025 年还没有
多方印证
成本账单生态与信任
- 一句话
微软官网说 Managed Instance"永远运行最新版 SQL,不用担心升级"——Kendra Little 逐条审计发现,2022 的旗舰功能 MI Link 上线一年还在排队,2019 的 tempdb 元数据优化到 2025 年 11 月都没上,错误日志故障转移就丢。
Microsoft's site says Managed Instance "always operates on the latest version of SQL, stop worrying about upgrades" — Kendra Little's line-by-line audit found the 2022 flagship MI Link feature still queueing a year after release, 2019's tempdb metadata optimization still absent as of Nov 2025, and error logs that vanish on failover.
- 窄场景
指望"lift & shift 上云就不用管版本"的团队;用到 Filestream/复制/事件通知/CHECKDB 修复选项等边缘功能的库。
Teams expecting "lift and shift to the cloud and forget about versions"; databases using edge features like Filestream, replication, event notifications, or CHECKDB repair options.
- 机制
MI 是"近全功能"的 PaaS,但"近"字很贵:tempdb 加文件可以、调文件大小不行(G1 时代还要靠建超大文件换 IOPS,多花的存储照单全收);错误日志不持久化,故障转移就可能丢;DBCC CHECKDB 的 REPAIR 选项、单用户模式、数据库快照、修改时区全都不支持;实例关不掉——没有 Serverless 档,闲置也按 24×7 计费。Kendra 的原话(引):"Isn't it always a little creepy when it's easier to get into something than it is to get out?"(进去容易出来难,不觉得瘆人吗?)
MI is a "near-full-feature" PaaS, but the "near" is expensive: you can add tempdb files but not size them (in the GPv1 era you had to create artificially large files to get IOPS — and pay for the storage); error logs aren't persisted and can be erased on failover; DBCC CHECKDB REPAIR options, single-user mode, database snapshots, and changing the time zone are all unsupported; and an instance can't be shut down — no serverless tier, so idle time bills 24/7. Kendra's own words (quoted): "Isn't it always a little creepy when it's easier to get into something than it is to get out?"
- 生产验证
来源 11:Kendra Little 2023-12(2025-11 更新)——逐项列出缺席功能:2022 的 IQP 特性缺席(2025 年 11 月大部分已补,但 Query Store 优化计划强制仍无)、2019 的内存优化 tempdb 元数据缺席、错误日志不持久化、100 库上限、无最小日志恢复模式;并点名营销话术与现实的落差;
来源 12:Red9 咨询公司 2024——"写这篇文章时,我们在 Red9 没有一个客户在用 SQL Managed Instance";"限制让现有应用迁移变得困难(对大型复杂负载,大多数客户根本不会考虑)"。
Source 11: Kendra Little, Dec 2023 (updated Nov 2025) — itemized missing features: 2022 Intelligent Query Processing features absent at publication (most added by Nov 2025, but Query Store optimized plan forcing still missing), 2019's memory-optimized tempdb metadata still missing, error logs not persisted, the 100-database cap, no minimal-logging recovery models; calls out the gap between the marketing pitch and reality;
Source 12: Red9 consulting firm, 2024 — "As of this writing, we don't have a single client here at Red9 where we are using SQL Managed Instances"; "limitations make migrating existing apps difficult (for large, complex workloads, most clients just won't go here)."
- 证据等级
`多方印证(2 个独立来源)`,独立顾问功能审计 ×1 + 咨询公司客户侧观察 ×1
`Corroborated (2 independent sources)`, independent consultant feature audit x1 + consulting-firm client-side observation x1.
Microsoft SQL Server 年份:2025
升级后执行计划回归:有合理索引却走全表扫描
多方印证
性能问题升级迁移
- 一句话
升级到 v7.5.x 后,优化器对部分 SQL 倾向选全表扫描或错误索引,215 亿行表的查询 60 秒被 kill,强制绑索引只要 0.4 秒。
After upgrading to v7.5.x, the optimizer prefers full table scans or wrong indexes for some queries — a 21.5-billion-row query was killed at 60 seconds, while forcing the index took 0.4 seconds.
- 窄场景
从 v4/v5 老版本升级到 v7.5.x 的生产集群;超大表点查、按时间范围聚合的查询。
Production clusters upgraded from v4/v5 to v7.5.x; point lookups on very large tables and time-range aggregate queries.
- 机制
优化器成本模型与统计信息解读随大版本变化(v7.5.x 的行数估计/代价偏好与 v4/v5 不同),老版本"恰好正确"的计划在新版本翻转。TiDB 没有 Oracle SPM 式开箱即用的计划基线冻结,得物团队的解法是手动绑定执行计划(SPM)或调系统变量(`tidb_opt_prefer_range_scan`、`tidb_opt_objective=determinate`),等于每个回归 SQL 都要人工兜底。
The optimizer's cost model and statistics interpretation changed across major versions — plans that were "accidentally correct" on v4/v5 flipped on v7.5.x. TiDB has no out-of-the-box plan-baseline freezing like Oracle SPM, so the Dewu team's remedies were manual plan binding (SPM) or session variables (`tidb_opt_prefer_range_scan`, `tidb_opt_objective=determinate`) — every regressed query needs a human safety net.
- 生产验证
来源 1:得物 2025 年升级实录——§5.1,215 亿行表、有合理索引却走全表扫描,60 秒被 kill,绑索引仅 0.4 秒;§5.2,聚合查询 v4.0.11 跑 12 秒→v7.5.6 跑 2 分 32 秒;
来源 2:asktug 1047626(2025)——生产 v7.5.6,215 亿行表合理索引下优化器倾向全表扫描,重收统计信息无效;
来源 3:asktug 1046610(2025-09)——生产 v7.5.6 聚合查询计划回归,评论区多方印证("线上已经多次 SQL 走错索引导致应用抖动"、"7.5.6 下执行计划确实有问题")。
Source 1: Dewu's 2025 upgrade account — section 5.1, a 21.5-billion-row table with a suitable index went full-scan, killed at 60s, 0.4s with a forced index; section 5.2, an aggregate query went from 12s on v4.0.11 to 2m32s on v7.5.6;
Source 2: asktug 1047626 (2025) — production v7.5.6, optimizer prefers full scan on a 21.5-billion-row table despite a suitable index; re-collecting statistics did not help;
Source 3: asktug 1046610 (Sep 2025) — production v7.5.6 aggregate-query plan regression, corroborated in comments ("production has repeatedly seen wrong indexes chosen, causing app jitter", "execution plans on 7.5.6 do have problems").
- 证据等级
`多方印证(3 个独立来源)`,具名生产复盘 ×1 + 社区生产问题帖 ×2(含评论区多方印证)。
`Corroborated (3 independent sources)`, named production postmortem x1 + community production threads x2 (with corroborating comments).
- 备注
评论区有用户反馈 v8.5.3 下同类问题已正常,属版本回归而非架构缺陷,保留收录并如实标注。
commenters report the same class of issue is gone on v8.5.3 — a version regression, not an architectural defect; retained with the note.
TiDB 年份:2025
自建起步 8 节点:分布式税还没跑业务就先交了
多方印证
运维复杂度成本账单
- 一句话
自建 TiDB 起步最少 3 PD + 3 TiKV + 2 TiFlash(8 节点),比同规模的 MySQL/MongoDB/PostgreSQL 贵得多——规模还没上来,账单先上来了。
Self-hosted TiDB starts at a minimum of 3 PD + 3 TiKV + 2 TiFlash (8 nodes) — far pricier than MySQL/MongoDB/PostgreSQL at the same scale; the bill arrives before the scale does.
- 窄场景
自建部署、中小规模业务;用"MySQL 替代"的预期做预算的团队。
Self-hosted deployments, small-to-medium workloads; teams budgeting with "MySQL replacement" expectations.
- 机制
TiDB 是多组件分布式架构,最小生产拓扑天然就是多节点:PD 管元数据与调度、TiKV 存数据(三副本)、TiFlash 列存、再加监控备份全套。节点数、SSD、网络都是硬成本,且三副本存储冗余让"每 GB 有效数据"的硬件成本天然高于单机库。
TiDB is a multi-component distributed architecture whose minimum production topology is multi-node by nature: PD for metadata and scheduling, TiKV for storage (three replicas), TiFlash for columnar, plus the full monitoring/backup stack. Node count, SSDs, and networking are hard costs, and three-replica storage redundancy makes per-GB-of-effective-data hardware cost structurally higher than single-node databases.
- 生产验证
来源 4:PeerSpot 具名评价,Mafiree DBA——"自建最少 3 PD + 3 TiKV + 2 TiFlash,比 MySQL/MongoDB/PostgreSQL 更贵"(引自原文大意);
来源 1:得物 DBA 团队 2025 年升级实录引言——"能用分库分表能解决的问题尽量选择 MySQL,毕竟运维成本相对较低、数据库版本更加稳定、单点查询速度更快、单机 QPS 性能更高这些特性是分布式数据库无法满足的"(引自原文)。
Source 4: PeerSpot named review, Mafiree DBA — self-hosting needs at minimum 3 PD + 3 TiKV + 2 TiFlash, "more expensive than MySQL/MongoDB/PostgreSQL" (paraphrased);
Source 1: Dewu DBA team's 2025 upgrade account, in their own words — "problems solvable with sharding should stay on MySQL: lower ops cost, more stable versions, faster point queries, higher single-node QPS — a distributed database cannot deliver these" (quoted in translation).
- 证据等级
`多方印证(2 个独立来源)`,具名用户评价 ×1 + 具名生产复盘引言 ×1。
`Corroborated (2 independent sources)`, named user review x1 + named production postmortem x1.
- 备注
与现有 [避坑] 卡「小数据量强行上 TiDB 是过度设计」主题部分重叠(该卡已含最小生产拓扑的运维成本论述)。
partially overlaps the existing [Pitfall] card "Small Data Forced onto TiDB Is Over-Engineering" (that card already covers the minimum-topology ops cost).
TiDB 年份:2025
版本 EOL 二选一:停机升级,或者自动开始交 Extended Support 费
多方印证
升级迁移成本账单
- 一句话
Aurora 的大版本 EOL 不给你"无痛"选项——PG 版是"关机、升级、重启,可能重启好几次";MySQL 版是"不升就自动 enroll 进按 vCPU 计费的 Extended Support"。
Aurora major-version EOLs offer no painless option — the Postgres path is "shut down, upgrade, restart, possibly restart several times"; the MySQL path is "don't upgrade and you're auto-enrolled into per-vCPU-billed Extended Support."
- 窄场景
跑在 PG 11(2024-01-31 EOL)或 Aurora MySQL 2.x(MySQL 5.7,2024-10-31 EOL)上的生产集群。
Production clusters on PG 11 (EOL Jan 31 2024) or Aurora MySQL 2.x (MySQL 5.7, EOL Oct 31 2024).
- 机制
Aurora 的大版本升级是停机式原地升级(时长取决于对象数量);AWS 对 MySQL 5.7 采用"到期自动 enroll Extended Support、下个月自动开始按 vCPU 每小时收费",最长收到 2027-02-28——不主动操作就等于默认开始付费。
Aurora major upgrades are downtime in-place upgrades (duration depends on object count); for MySQL 5.7 AWS chose "auto-enroll into Extended Support at expiry, billing starts the next month per vCPU-hour," running through Feb 28 2027 — doing nothing means silently starting to pay.
- 生产验证
来源 10:The Register 2023-02-14 独立报道——AWS 致 Aurora PG 用户邮件原文:"升级过程将关闭数据库实例、执行升级、重启实例……可能重启多次";Percona 的 Charly Batista 评价"把迁移的成本和负担推给用户很方便";
来源 11:AWS re:Post 官方公告——2024-10-31 后未升级的 Aurora MySQL 2.x 自动进入 RDS Extended Support,2024-12-01 起自动收费、按 vCPU 每小时计价,直到用户升级或删库。
Source 10: The Register, Feb 14 2023, independent reporting — AWS's email to Aurora PG customers verbatim: "The upgrade process will shut down the database instance, perform the upgrade, and restart the database instance… may be restarted multiple times"; Percona's Charly Batista: "it is convenient for AWS to push this migration process onto their users, so they carry the expense and burden";
Source 11: AWS re:Post official announcement — Aurora MySQL 2.x clusters not upgraded by Oct 31 2024 are auto-enrolled in RDS Extended Support with automatic billing from Dec 1 2024, priced per vCPU-hour, until the user upgrades or deletes the cluster.
- 证据等级
`多方印证(2 个独立来源)`,独立媒体报道(PG 线)+ AWS 官方公告(MySQL 线)。
`Corroborated (2 independent sources)`, independent media reporting (Postgres track) + AWS official announcement (MySQL track).
- 备注
来源 10 中 PlanetScale CEO Sam Lambert 为直接竞争者,其"embarrassingly low bar"等评价为竞争者口径,仅采信邮件原文与事实部分。
in Source 10, PlanetScale CEO Sam Lambert is a direct competitor; his "embarrassingly low bar" style judgments are competitor rhetoric — only the email's verbatim text and facts are relied upon.
Amazon Aurora 年份:2024
"免费数据集"的 $14,000 账单:2 小时烧掉 2.5PB
多方印证
成本账单
- 一句话
HTTP Archive 数据集免费开放,但查询费算在查询者头上——一个 Python 脚本 2 小时跑出 $14,000,零预警。
The HTTP Archive dataset is free to access, but queries are billed to the querier — one Python script ran up $14,000 in 2 hours with zero warning.
- 窄场景
用 Python/官方客户端库(而非网页控制台)跑 BigQuery 的新手、学生、研究人员;成本控制默认关闭的项目。
Newcomers, students, and researchers running BigQuery through Python/the official client libraries instead of the web console; projects with cost controls left at their defaults (off).
- 机制
BigQuery 按量计费 $6.25/TiB,费用记在发起查询的项目上,与数据集是否"公开免费"无关;网页 UI 会在运行前显示"本次查询将处理 X 数据"的预估,但 Python 客户端没有这个机制;cost controls(自定义配额/熔断)默认不开启。用户脚本 2 小时处理了约 2.5PB,按单价折合约 $14,000。
BigQuery on-demand billing is $6.25/TiB, charged to the project that initiates the query — independent of whether the dataset is "public and free." The web UI shows a "this query will process X" estimate before running; the Python client has no such mechanism; cost controls (custom quotas / circuit breakers) are off by default. The user's script processed ~2.5PB in 2 hours, which at list price is about $14,000.
- 生产验证
来源 1:用户 Tim 在 HTTP Archive 论坛的原帖(2024-02-20)——"被收 $14,000,zero warning whatsoever,Google 不给退";他指出用官方 Python 库跑查询时"unlike the web ui there's no mechanism to show costs";提议默认加一个 $5k 熔断器,手动确认才继续;
来源 2:The Register 2024-02-22 独立报道——印证了原帖内容与金额(2.5PB × $6.25/TiB),并引述维护者回应(已在 FAQ 加显式警告、99% 用户只看免费月报)。
Source 1: user Tim's original HTTP Archive forum post (Feb 20 2024) — "billed $14,000, zero warning whatsoever, and they won't remove the fee"; he notes that with the official Python libraries "unlike the web ui there's no mechanism to show costs"; he proposed a default $5k circuit breaker requiring manual confirmation to continue;
Source 2: The Register, Feb 22 2024, independent coverage — corroborates the post and the amount (2.5PB x $6.25/TiB), quoting the maintainer's response (explicit warning added to the FAQ; 99% of users only read the free monthly reports).
- 证据等级
`多方印证(2 个独立来源)`,具名论坛原帖 + 独立媒体报道。
`Corroborated (2 independent sources)`, named forum post + independent media coverage.
Google BigQuery 年份:2024
没有原生 Upsert:ReplacingMergeTree 是"最终一致"的去重,不是主键更新
多方印证
生态与信任
- 一句话
ClickHouse 没有 `INSERT ON CONFLICT`/原生主键 upsert——官方推荐的 ReplacingMergeTree 只是"写多版本、后台异步合并",不加 FINAL 会读到重复行;Kafka 重试导致重复消息,"clickhouse will create a new row for duplicated message"。
ClickHouse has no `INSERT ON CONFLICT` or native primary-key upsert — the officially recommended ReplacingMergeTree is "write multiple versions, merge in the background"; skip FINAL and you read duplicates. Kafka retries producing duplicate messages? "Clickhouse will create a new row for duplicated message."
- 窄场景
CDC、Kafka 重试场景;从 PG/MySQL/Doris(Unique Key)迁来的用户。
CDC, Kafka-retry scenarios; users arriving from PG/MySQL/Doris (Unique Key).
- 机制
ReplacingMergeTree 的去重发生在后台 merge 时,查询时若不加 FINAL 新旧行并存;加了 FINAL 则查询期付去重税。没有开箱即用的幂等写入语义——用户被迫在写入前先查库判重。
ReplacingMergeTree dedups at background-merge time; pre-merge, old and new rows coexist and queries need FINAL; FINAL pays a dedup tax at query time. No out-of-the-box idempotent write semantics — users check-then-write by hand.
- 生产验证
—
Source 26: Cloudera community user (NiFi+Kafka→ClickHouse, circa 2024) — "Unfortunately clickhouse will create a new row for duplicated message";
Source 15: Hacker News discussion (Sep 2024) — multiple practitioners confirm the awkward "overlapping inserts + ReplacingMergeTree background dedup, FINAL on read" pattern.
- 证据等级
`多方印证(2 个独立来源)`。
`Corroborated (2 independent sources)`.
ClickHouse 年份:2024
两次换证:Apache 2.0 → BSL → 专有 Enterprise
多方印证
生态与信任
- 一句话
2019 年刚把 Apache 2.0 换成 BSL,2024 年连 BSL 的"开源遮羞布"都不要了——自托管只剩专有 Enterprise 许可,免费与否看你年营收过没过 1000 万美元。
Barely five years after swapping Apache 2.0 for the BSL, CockroachDB dropped even the "fauxpen source" fig leaf in 2024 — self-hosting is now proprietary Enterprise only, free or not depending on whether your revenue clears $10 million.
- 窄场景
自托管部署的团队;把 CockroachDB 当"开源 Postgres 替代品"选型的公司;年营收接近或超过 1000 万美元门槛的企业。
Self-hosting teams; companies that picked CockroachDB as an "open Postgres alternative"; businesses near or above the $10M revenue threshold.
- 机制
许可证是信任锚。BSL 1.1 虽非 OSI 开源许可,但至少有"3 年后转 Apache 2.0"的滚动回退条款(来源 5);2024-11 的 24.3 起 Core 退场,自托管统一为 Enterprise 许可:年营收 <1000 万美元免费(须每年自证),之上按 CPU 核数付费(来源 1、4)。代码仍"可见"但不可自由使用——从"fauxpen source"变成纯专有 + 免费试吃(来源 2)。
The license is the trust anchor. BSL 1.1 was never OSI-approved, but it had a rolling fallback: three years after each release the code reverted to Apache 2.0 (Source 5). With 24.3 (Nov 2024), Core was retired and self-hosting consolidated under a single Enterprise license: free below $10M annual revenue (self-attested yearly), per-CPU-core paid above it (Sources 1, 4). The code stays "visible" but not freely usable — from fauxpen source to pure proprietary with a free taste (Source 2).
- 生产验证
来源 1:Percona 联创 Peter Zaitsev 称"CockroachDB 完成了从开源的彻底转向……又一个 Oracle";OpenUK CEO Amanda Brock 称"Cockroach 一直是开源许可的抱怨者,对其自有产品建不起开源商业模式不意外"(2024-08);
来源 2:HN 用户原话被引用——"Enterprise 版的问题是它很贵、销售导向,咬一口可能就是跟未来的 Oracle/房东型厂商签卖身契";另一条——"VC 支持的开源项目迟早都这样,开源只是诱饵,等投资人要增长就扯地毯"(2024-08);
来源 3:OSI 执行董事 Stefano Maffulli 称中途换成 BUSL 这类限制性许可是"从用户社区脚下扯地毯"、"打破开源社区信任的 switcheroo";Yugabyte 创始人 Karthik Ranganathan 称这是"short term thinking","开发者现在会犹豫选 CockroachDB,因为他们知道一旦长大撞上营收门槛,用法就要变"(2024)。
Source 1: Percona co-founder Peter Zaitsev — "CockroachDB has completed its transition away from open source... yet another Oracle"; OpenUK CEO Amanda Brock — "Cockroach has been a longstanding open-source-license-grumbler" that failed to build a business on its open-source product (Aug 2024);
Source 2: HN users quoted verbatim — "the problem with the Enterprise edition is that it's quite expensive, 'contact us' salesy, and it feels like taking a bite of this edition is possibly getting into bed with a future Oracle/landlord type of relationship"; and "whenever an open source project is run by a VC-backed company, it sooner or later ends up like this... the rug gets pulled" (Aug 2024);
Source 3: OSI executive director Stefano Maffulli called midstream switches to BUSL-class licenses "pulling the rug from beneath the user community's feet," a "switcheroo that breaks the trust of the open source community"; Yugabyte founder Karthik Ranganathan called it "short term thinking" — "developers and small organizations will likely be hesitant to adopt CockroachDB now because they know that if they grow and hit that revenue amount, there will be implications in how they use the database" (2024).
- 证据等级
`多方印证(5 个独立来源)`,独立媒体 ×4(含具名第三方评论)+ 2019 年换证历史报道。
`Corroborated (5 independent sources)`, independent media x4 (with named third-party commentary) + 2019 relicense history.
CockroachDB 年份:2024
免费版长期无增量备份与 PITR:以弹性著称的库,丢数据类故障零方案
多方印证
已修复于 24.3
运维复杂度
- 一句话
节点挂了能自愈,数据写错了/删错了呢?免费版长期没有增量备份和 PITR——HN 生产用户的原话:"有一类错误/故障,(免费版至少)是 ZERO solution。"
A dead node heals itself — but a wrong write or an accidental delete? The free tier long shipped without incremental backups or PITR. In the HN production user's words: for that class of errors and failures, the free offering had "ZERO solution."
- 窄场景
用免费 Core 跑生产、指望"弹性=备份"的团队;需要延迟只读副本做误操作兜底的场景。
Teams running production on free Core expecting "resilience = backups"; anyone wanting a delayed read replica as a fat-finger safety net.
- 机制
Raft 解决的是节点故障,不是逻辑错误。增量备份、延迟副本、PITR 这些"防人祸"能力长期是 Enterprise 付费功能;免费 Core 用户只有全量备份可用,误删/误写只能认。讽刺的是:一个以 resilience 为招牌的数据库,免费档对最常见的生产事故(人为误操作)没有恢复手段。
Raft solves node failure, not logical error. Incremental backups, delayed replicas, and PITR — the defenses against human error — were Enterprise-only paid features for years; free Core users got full backups at best, so a bad UPDATE/DELETE meant accepting the loss. The irony: a database marketed on resilience had no recovery story in its free tier for the most common production incident (operator error).
- 生产验证
来源 11:2022-11 HN 生产用户——"免费版刚开始时备份基本不可用;现在改善了,但对一个以弹性为卖点的数据库,免费版仍有严重短板:没有只读副本/延迟只读副本,没有增量备份,做不了 PITR……有一类错误/故障是 ZERO solution";
来源 4:2024-08 SD Times——24.3 起所有用户(含免费)可用此前付费才有的 cluster optimization、disaster recovery、backup、streaming 等高级功能。
Source 11: Nov 2022 HN production user — "When we started using CockroachDB, the free version barely had working backups. This has improved, but for a database built around resilience, the free offering still has serious shortcomings... you can't do incremental backups or use other approaches that can facilitate PITR... there's a class of errors/failures where (the free version at least), has ZERO solution";
Source 4: Aug 2024 SD Times — from 24.3, all users (including free) get previously paid-only capabilities: cluster optimization, disaster recovery, backup, streaming, and advanced security.
- 证据等级
`多方印证(2 个独立来源)`,HN 生产用户 + 独立媒体(前后对照)。
`Corroborated (2 independent sources)`, HN production user + independent media (before/after).
- 备注
已修复于 24.3(2024-11,Enterprise Free 全量开放备份/容灾能力),按规则保留收录并标注修复版本。
fixed in 24.3 (Nov 2024 — Enterprise Free opened up backup/disaster-recovery); retained per the rule with the fixed-in version labeled.
CockroachDB 年份:2024
集群启动:5 分钟是常态,35 分钟"每周几次"
多方印证
性能问题运维复杂度
- 一句话
交互式/任务集群冷启动 5–6 分钟是常态,随机飙到 20–35 分钟"每周几次"——跑 5 分钟的任务要付 11 分钟甚至 35 分钟的账。
Cold starts of 5–6 minutes are normal for classic clusters, with random spikes to 20–35 minutes "several times a week" — a 5-minute job can bill 11 or even 35 minutes.
- 窄场景
Azure 上的 job compute;短时高频任务;被 Spark Connect 破坏性变更卡住、切不到 shared/serverless 的团队。
Job compute on Azure; short, high-frequency tasks; teams blocked from shared/serverless compute by Spark Connect breaking changes.
- 机制
经典集群每次从零拉 VM、装库、起 Spark;启动阶段 DBU 不计费但云厂商 VM 照收钱;启动时长受云厂商容量、网络、init 脚本多重因素影响,方差大。Serverless 靠 warm pool 把启动压到秒级,但那是另一套计费。
Classic clusters pull VMs, install libraries, and start Spark from scratch on every run; DBUs are not billed during startup but the cloud provider's VMs are; startup time varies with cloud capacity, networking, and init scripts. Serverless compresses startup to seconds via warm pools — at a different price.
- 生产验证
—
Source 8: Databricks official community thread (circa 2024) — a 1-driver + 1-worker cluster normally takes 5–6 minutes to start, several times a week 20–35 minutes; "Paying for 35 minutes of compute to run a 5 minute job is hard to justify" (direct quote); stuck on single-user compute because of Spark Connect breaking changes;
Source 3: TDS 2024 measurement — Serverless starts in 5–10 seconds, "much less than starting up a cluster from scratch" (reverse confirmation of how slow classic startup is).
- 证据等级
`多方印证(2 个独立来源)`,官方社区用户实测 + 独立基准实验。
`Corroborated (2 independent sources)`, official-community user measurement + independent benchmark experiment.
Databricks 年份:2024
存储过程缺失多年,2025 年中才补上——且仅限 Unity Catalog
多方印证
升级迁移生态与信任
- 一句话
2022 年 SQL Server 迁移用户问 DECLARE/存储过程等价物,被告知"用 Python 拼 SQL 字符串";具名 `CREATE PROCEDURE` 2025-06/07(DBR 17.0)才来,且仅 UC 可用。
In 2022 a SQL Server migrant asking for a DECLARE/stored-procedure equivalent was told to "build up a SQL string in Python"; named `CREATE PROCEDURE` arrived in Jun/Jul 2025 (DBR 17.0) — UC-only.
- 窄场景
从 SQL Server/Oracle 迁移、重度依赖存储过程的团队;想给 Power BI 暴露可调存储过程的场景。
Teams migrating from SQL Server/Oracle with heavy stored-procedure reliance; exposing callable procedures to Power BI.
- 机制
Databricks SQL 长期没有过程化 SQL 与具名存储过程,2024 年仍有人问"Power BI 可调存储过程等价物"被建议用 view 凑合。补上后仍有前提:仅 Unity Catalog,游标等完整能力要 DBR 18.1+。
Databricks SQL had no procedural SQL or named procedures for years; as late as 2024 users asking for a "Power BI–callable stored procedure equivalent" were told to make do with views. The fill-in carries prerequisites: Unity Catalog only; full capabilities like cursors need DBR 18.1+.
- 生产验证
—
Source 34: community thread, circa 2022–2023 (41k+ views) — SQL Server-background user asking for a DECLARE equivalent; the answer: "Use python and build up a sql string which can be executed.";
Source 35: community thread, circa Aug 2024 — "stored procedures don't seem to be an option," advised to use a view instead.
- 证据等级
`多方印证(2 个独立来源)`,修复版本有官方 release notes 证据。
`Corroborated (2 independent sources)`, with the fix version evidenced by official release notes.
Databricks 年份:2024
Atlas Data API / Device Sync 说砍就砍:生产系统一年内被迫迁移
多方印证
升级迁移生态与信任
- 一句话
MongoDB 于 2024-09 宣布 Atlas Data API、Custom HTTPS Endpoints、Device Sync(Realm 同步)停服(2025-09 EOL),大批已上线生产系统的客户被迫在一年内重写,社区出现信任危机。
In September 2024 MongoDB announced end-of-life for Atlas Data API, Custom HTTPS Endpoints, and Device Sync (Realm sync) with EOL in September 2025, forcing many production systems into rewrites within a year and triggering a trust crisis in the community.
- 窄场景
基于 Atlas App Services(Data API/HTTPS endpoints/Functions/Realm 同步)构建生产系统的团队,尤其是移动端与独立开发者。
Teams that built production systems on Atlas App Services (Data API / HTTPS endpoints / Functions / Realm sync), especially mobile developers and independents.
- 机制
App Services 是 MongoDB 曾主推的 serverless 配套层,客户按官方引导把鉴权、函数、同步建在上面;停服后无官方迁移路径(密码无法导出、用户需重建),替代方案(如 Ditto)在当时尚未就绪。厂商战略转向的成本全部由客户承担。
App Services was the serverless companion layer MongoDB itself had promoted; customers built auth, functions, and sync on it per official guidance. After the shutdown announcement there was no official migration path (passwords cannot be exported, users must be recreated), and suggested replacements (e.g. Ditto) were not ready at the time. The cost of a vendor strategy pivot landed entirely on customers.
- 生产验证
MongoDB 官方论坛停服公告帖(2024-09-11)下多名独立客户留言:"You lost trust"(14 赞)、全站基于 HTTPS endpoints + functions 的电商站(9 赞)、靠订阅收入的 iOS 开发者担忧迁移期用户流失、4 年 3 万用户的认证体系无迁移方案。
MongoDB's official forums deprecation announcement thread (2024-09-11) with multiple independent customer replies: "You lost trust" (14 likes), an e-commerce site built entirely on HTTPS endpoints + functions (9 likes), a subscription-revenue iOS developer worried about churn during migration, and a 4-year-old 30k-user auth system with no migration plan.
- 证据等级
`多方印证(4+ 独立客户)`(官方论坛公告帖内多名独立客户留言)
—
MongoDB 年份:2024
审计是"销售赋能工具":合规团队的 KPI 连着销售提成
多方印证
成本账单生态与信任
- 一句话
四名前 Oracle 许可证合规(LMS)高管具名作证:审计的产出直接交给销售谈判,销售部门在审计流程中的权力远大于审计团队。
Four former Oracle License Management Services (LMS) executives testified on the record: audit findings go straight to the sales team for negotiation leverage, and sales holds far more power than audit inside Oracle.
- 窄场景
任何规模的 Oracle EE 客户;合同续签前、架构变更(虚拟化/上云/并购)后是高发期。
Oracle EE customers of any size; audits cluster around contract renewals and infrastructure changes (virtualization, cloud, M&A).
- 机制
LMS 审计发现"不合规项"后并不止于补款,而是作为销售谈判杠杆——客户常被引导用"买更多许可/转 Oracle 云"来"解决"审计问题。2020 年 City of Sunrise 消防员养老基金诉 Oracle 的投资者诉讼即指控 Oracle 以审计威胁迫使客户上 Oracle 云(Oracle 否认)。授权咨询顾问另指出,Oracle 销售与审计团队会以"健康检查"等名义从客户会议中收集许可缺口情报(来源 13)。
An LMS audit finding "non-compliance" rarely ends with a true-up payment — it becomes sales negotiation leverage, with customers steered toward "buy more licenses / move to Oracle Cloud" to "resolve" the audit. A 2020 investor lawsuit (City of Sunrise Firefighters' Pension Fund v. Oracle) alleged Oracle used audit threats to force customers onto Oracle Cloud (Oracle denied). Licensing consultants further note that Oracle sales and audit teams gather license-gap intelligence from customer meetings under the guise of "health checks" (source 13).
- 生产验证
—
Source 2: former Oracle LMS manager Adi Ahuja told a Palisade Compliance customer webinar that audits are "a sales enablement tool"; "Sales has far more power within Oracle than the audit team"; testifying alongside him were former LMS VP Craig Guarente, former senior manager Ryan Bendana, and former executive Max Shlopak (combined ~50 years at Oracle);
Source 3: a 2024 HN thread where multiple customers shared audit experiences (six-figure bills, helplessness under toolchain lock-in), echoing the "audits manufacture fear" narrative.
- 证据等级
`多方印证(2 个独立来源)`,具名前 Oracle 高管 ×4 + 社区客户多方呼应。
`Multi-source corroboration (2 independent sources)`, 4 named former Oracle executives + multiple community customers echoing.
- 备注
主题可能与现有 [避坑] 卡(许可证审计类)重叠。
topic may overlap with existing [Pitfall] cards (license-audit theme).
Oracle Database(甲骨文) 年份:2024
Database@Azure:上了 Exadata 就被强制 RAC,算力需求翻倍
多方印证
成本账单
- 一句话
Oracle Database@Azure 跑在 Exadata 上,而 Exadata 必须开 RAC——从单实例迁过去,算力需求直接翻倍,被许可专家称为"运行 Oracle 数据库最贵的方式"。
Oracle Database@Azure runs on Exadata, and Exadata requires RAC — migrating from a single-instance setup doubles the compute capacity needed; licensing experts call it "the most expensive way to run an Oracle database".
- 窄场景
考虑把 Oracle 迁到 Azure 的 Database@Azure 服务的客户;从单实例/非 RAC 架构迁移时。
Customers considering Oracle Database@Azure; especially migrations from single-instance/non-RAC architectures.
- 机制
Database@Azure 的底层是 Oracle Exadata,而 Exadata 平台要求启用 RAC;RAC 要求多节点,客户从单实例迁过去需要约两倍的计算容量;BYOL 许可按 2 vCPU 折 1 个 processor 计算。独立许可专家 Eric Guyer(Remend)原话:"Any customer moving from another Oracle database would require twice the compute capacity, 'if not more'... That is the most expensive way to run an Oracle database"。
Database@Azure sits on Oracle Exadata, which requires RAC; RAC means multiple nodes, so customers moving from single-instance need roughly twice the compute; BYOL licenses convert at 2 vCPUs = 1 processor. Independent licensing expert Eric Guyer (Remend): "Any customer moving from another Oracle database would require twice the compute capacity, 'if not more'... That is the most expensive way to run an Oracle database."
- 生产验证
来源 9:The Register 2024-02-05 报道——前 Oracle LMS VP Craig Guarente 指出 Oracle 云许可文件"none of this is contractual; it's just stuff Oracle makes up"(都不是合同条款,是 Oracle 自己编的);Eric Guyer 给出"两倍算力"的量化判断。Oracle 官方回应称 Database@Azure 按 OCI 定价。
Source 9: The Register 2024-02-05 — former Oracle LMS VP Craig Guarente on Oracle cloud licensing documents: "none of this is contractual; it's just stuff Oracle makes up"; Eric Guyer gave the "twice the compute" quantification. Oracle responded that Database@Azure is priced on OCI terms.
- 证据等级
`多方印证(2 个独立来源)`,独立许可专家 ×2(均曾长期从事 Oracle 许可工作)。
`Multi-source corroboration (2 independent sources)`, 2 independent licensing experts (both long-time Oracle licensing practitioners).
Oracle Database(甲骨文) 年份:2024
主键只是个"建议":声明了也不强制,重复数据自己查
多方印证
运维复杂度生态与信任
- 一句话
Redshift 接受 PRIMARY KEY/FOREIGN KEY/UNIQUE 语法,但一律不强制——ETL 写重复了不会报错;更坑的是优化器会"相信"这些约束,声明错了查询结果直接静默算错。
Redshift accepts PRIMARY KEY/FOREIGN KEY/UNIQUE syntax but enforces none of it — an ETL that writes duplicates gets no error; worse, the optimizer "trusts" those constraints, so a wrong declaration silently computes wrong results.
- 窄场景
从传统 RDBMS 迁移过来、习惯靠数据库保唯一性的团队;CDC 重放、幂等没做好的管道。
Teams migrating from traditional RDBMS who expect the database to guarantee uniqueness; pipelines with CDC replays or weak idempotency.
- 机制
约束纯信息性(informational),只给优化器做计划用;唯一性必须由 ETL/应用层保证,需自建去重校验。AWS 文档原话:"Do not define primary key and foreign key constraints unless your application enforces the constraints."(来源 8,已打开核实)
Constraints are informational only — hints for the query planner; uniqueness must be guaranteed by the ETL/application layer with self-built dedup checks. AWS docs, verbatim: "Do not define primary key and foreign key constraints unless your application enforces the constraints." (Source 8, opened and verified.)
- 生产验证
来源 4:PeerSpot 2024-05-09,FNU AKSHANSH(Senior Data Engineer,平台标注 Real User)——"Redshift does not have primary-key tools. The vendor must consider adding them.";且元数据难取,"too many layers to get simple information";
来源 5:TrustRadius 认证用户评论,Cons 明写 "Missing option to restrict duplicate records";
来源 3:同一 HN 讨论串(2022-05)匿名评论——"Our most shocking discovery on Redshift was that primary key constraints are not honored."
Source 4: PeerSpot, May 9 2024, FNU AKSHANSH (Senior Data Engineer, site-labeled Real User) — "Redshift does not have primary-key tools. The vendor must consider adding them."; also "too many layers to get simple information" from metadata;
Source 5: TrustRadius verified-user review, Cons: "Missing option to restrict duplicate records";
Source 3: anonymous comment in the same HN thread (May 2022) — "Our most shocking discovery on Redshift was that primary key constraints are not honored."
- 证据等级
`多方印证(3 个独立来源)`,具名点评用户 + 点评平台认证用户 + HN 匿名评论。
`Corroborated (3 independent sources)`, named reviewer + platform-verified reviewer + anonymous HN comment.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap existing [Pitfall] cards on this site.
Amazon Redshift 年份:2024
2024 失窃事件:Snowflake 怪客户没开 MFA,但当时管理员根本没有强制开关
多方印证
生态与信任
- 一句话
2024 年 4–6 月 UNC5537 用窃取的凭证登录约 165 家 Snowflake 客户租户,Snowflake 的公开口径是"客户没开 MFA"——但当时租户管理员根本没有"一键全员强制 MFA"的开关,"it cannot be enforced on users by the admin of the tenant"(租户管理员无法对用户强制执行),等于把开不开 MFA 交给每个用户自己决定。
In Apr–Jun 2024 the UNC5537 crew logged into ~165 Snowflake customer tenants with stolen credentials; Snowflake's public line was "customers didn't enable MFA" — but tenant admins had no "force MFA for everyone" switch at the time: "it cannot be enforced on users by the admin of the tenant," leaving MFA enrollment up to each individual user.
- 窄场景
2024 年中的所有 Snowflake 租户;用用户名+密码(未接 SSO)的用户。
Every Snowflake tenant in mid-2024; users on username+password without SSO.
- 机制
平台侧缺失租户级 MFA 强制能力,安全责任被推给终端用户;讽刺的是 Snowflake 自己承认攻击者用"前员工的个人凭证"攻破了该员工名下的 demo 账号——同样因为"not behind Okta or MFA",而 Snowflake 对记者的十几个问题至少六次拒绝回答。
The platform lacked tenant-level MFA enforcement, pushing security responsibility onto end users; ironically Snowflake admitted the attackers breached a demo account under a former employee's name with "personal credentials" — also "not behind Okta or MFA" — while declining to answer at least six of a reporter's dozen questions.
- 生产验证
来源 26:Hacker News 2024-06-02 社区讨论(49 条评论)——有租户安全实施经验的从业者:"I generally dislike Snowflake in this case and think they're attempting to deflect the blame off themselves for the architecture you get forced into as a tenant";
来源 27:InformationWeek 2024-06 引述 Mitiga CTO Ofer Maor、Stern Security、IP Architects、Silverfort CISO 四家独立信源,一致确认管理员无法强制 MFA;
来源 28:TechCrunch 2024-06-07 调查报道(仅作背景)。
—
- 证据等级
`多方印证(2 个独立来源 + 4 家专家信源)`,社区一线批评 + 媒体引述的多家独立安全专家。
`Corroborated (2 independent sources + 4 expert voices)`, frontline community criticism + multiple independent security experts quoted by media.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2024
升级优化器回归:2019 的 UDF 内联翻车,2022 的 CE 估计仍不准
多方印证
已修复于 2019
性能问题升级迁移
- 一句话
每次大版本升级都动优化器——2019 的标量 UDF 内联让电商客户升级后性能暴跌只能全局关闭;2022 的基数估计换个数据类型长度就能给出"垃圾"估计,排序直接溢出到磁盘。
Every major version moves the optimizer — 2019's scalar UDF inlining tanked an e-commerce customer's performance so badly they disabled it globally; in 2022, changing a function parameter's data-type length alone produces "hot garbage" estimates and sorts spilling to disk.
- 窄场景
升级 2019+ 并启用新兼容级别的库;大量使用标量 UDF、或有数据倾斜/长尾参数的查询。
Databases upgraded to 2019+ with the compatibility level raised; heavy scalar-UDF usage or skewed/long-tail query parameters.
- 机制
2019 引入标量 UDF 内联(把逐行执行的函数展平进计划),但内联后的基数估计并不稳定;2022 又在 CE 周边持续加东西(CE 反馈等)。Brent Ozar 用 Stack Overflow 数据库实测:NVARCHAR(40) 的函数版本 2022 估计精准、1 秒跑完;把函数参数改成 NVARCHAR(MAX),2022 直接忽略索引,"估计只有 1 行出来",排序内存不足溢写磁盘——"连'2019 内联后估计更准了'都说不出口"。
SQL Server 2019 introduced scalar UDF inlining (flattening row-by-row functions into the plan), but post-inlining cardinality estimates proved unstable; 2022 kept adding things around the CE (CE feedback, etc.). Brent Ozar tested on the Stack Overflow database: the NVARCHAR(40) function version got bang-on estimates and finished in 1 second on 2022; changing the parameter to NVARCHAR(MAX) made 2022 ignore the index entirely — "how many rows did it actually think were going to come out after the function's filter? Just one" — and the sort spilled to disk from under-allocated memory.
- 生产验证
来源 4:Pinal Dave 2020-08——某大型电商客户刚升级 2019 就遭遇大面积性能问题,排查到标量 UDF 内联,执行 `ALTER DATABASE SCOPED CONFIGURATION SET TSQL_SCALAR_UDF_INLINING = OFF` 后"性能大幅改善"(improved big time),客户决定保持关闭;
来源 5:Brent Ozar 2024-09——2022 下仅改函数参数的数据类型长度就导致估计崩坏、排序溢盘;并指出"即使你一直在用 2014+ 的新 CE,2022 的 CE 反馈也能把估计改回老版本"。
Source 4: Pinal Dave, Aug 2020 — a large e-commerce client hit widespread performance problems immediately after upgrading to 2019; root-caused to scalar UDF inlining; setting `ALTER DATABASE SCOPED CONFIGURATION SET TSQL_SCALAR_UDF_INLINING = OFF` "improved performance big time," and the client left it off;
Source 5: Brent Ozar, Sep 2024 — on 2022, merely changing a function parameter's data-type length collapsed the estimate and spilled the sort; he also notes 2022's CE feedback can move estimates back to older versions even for workloads long on the post-2014 CE.
- 证据等级
`多方印证(2 个独立来源)`,具名顾问客户案例 ×1 + 带复现的机制分析 ×1
`Corroborated (2 independent sources)`, named consultant client case x1 + mechanism analysis with reproduction x1.
- 备注
标量 UDF 内联的已知问题已修复于 2019 CU31 / 2022 CU17(KB4538581),该部分标注"已修复";但 CE 随版本持续变化、估计不稳定是 2024 年实测现状,未修复。
the known scalar UDF inlining issues were fixed in 2019 CU31 / 2022 CU17 (KB4538581) — that part is marked fixed. But CE behavior changing with each version, with unstable estimates, is the measured state as of 2024 — not fixed.
Microsoft SQL Server 年份:2024
非 shardkey 建不了唯一索引:PG 迁 TDSQL,迁移评估第一天就被挡下
多方印证
运维复杂度升级迁移
- 一句话
PostgreSQL 里一条普通的 `create unique index`,在 TDSQL PostgreSQL 版上报 `Unique index of partitioned table must contain the hash/modulo distribution column`——非分布键列建不了唯一索引,业务强依赖的唯一约束在迁移评估阶段就被挡下。
A plain `create unique index` that works fine in PostgreSQL fails on TDSQL PostgreSQL Edition with `Unique index of partitioned table must contain the hash/modulo distribution column` — you cannot build a unique index on a non-distribution-key column, so business-critical uniqueness constraints get blocked at the migration-evaluation stage.
- 窄场景
PostgreSQL → TDSQL PostgreSQL 版(分布式)迁移评估;表上有非分布键唯一索引、组合唯一索引的业务。
Evaluating a move from PostgreSQL to TDSQL PostgreSQL Edition (distributed); workloads with unique or composite-unique indexes on non-distribution-key columns.
- 机制
分布式下唯一约束必须在单分片内可判定;TDSQL PG 要求分区表的唯一索引必须包含 hash 分布列(shardkey),否则跨分片无法保证全局唯一。绕法都有代价:把 shardkey 塞进索引会失去原列的唯一语义;换 shardkey 但分布键只能选一个字段,多张唯一索引无解;触发器+影子表方案发帖人自己都未实践。
In a distributed layout, uniqueness must be decidable within a single shard; TDSQL PG requires every unique index on a partitioned table to include the hash distribution column (shard key), otherwise global uniqueness cannot be enforced across shards. Every workaround costs something: stuffing the shard key into the index destroys the original column's uniqueness semantics; switching the shard key is impossible when a table needs several unique indexes (only one distribution column is allowed); the trigger-plus-shadow-table approach was never even tried by the reporter.
- 生产验证
来源 1:V2EX 用户 2024-10 迁移实录——users 表(id 主键、user_name 唯一)在原生 PG 建唯一索引正常,TDSQL PG 直接报错;尝试三种绕法均不理想。引用原话:"我去咨询了官方客服,她们给我的反馈是分布式数据库是有挺多的限制,建议我使用她们的 postgresql"(即腾讯云客服建议改用普通 PostgreSQL);
来源 2:CSDN 用户 usoa 2024-09-29 实操笔记——建表 `distribute by shard(ct)` 而主键为 id 时,亲手复现同一报错 `Unique index of partitioned table must contain the hash distribution column`,并总结"有主键的时候,指定的 shardkey 必须为主键中的一个"。
Source 1: V2EX user, Oct 2024 migration account — a `users` table (`id` primary key, `user_name` unique) built its unique index fine in native PG but errored on TDSQL PG; three workaround attempts all fell short. Quoted: "I asked official support, and their feedback was that distributed databases have quite a few limitations — they suggested I use their [plain] PostgreSQL" (i.e., Tencent Cloud's own support advised against TDSQL PG);
Source 2: CSDN user usoa, Sep 29 2024 hands-on notes — reproduced the identical error (`Unique index of partitioned table must contain the hash distribution column`) when creating a table `distribute by shard(ct)` with primary key `id`, concluding "when there is a primary key, the specified shard key must be one of the primary key columns."
- 证据等级
`多方印证(2 个独立来源)`,V2EX 迁移实录 ×1 + CSDN 实操笔记 ×1(后者为中性技术笔记,仅作机制印证)。
`Corroborated (2 independent sources)`, V2EX migration account x1 + CSDN hands-on notes x1 (the latter is a neutral technical note, cited for mechanism corroboration only).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site on this site.
腾讯云 TDSQL 年份:2024
"Postgres 兼容"不等于搬过来就快:跨分片延迟税
多方印证
性能问题
- 一句话
表和索引不一定存在同一个节点上,join、外键检查、索引查找都可能变成多次跨节点 RPC;厂商 CEO 自己承认,开发者要"花额外周期为分布式特性优化 PG 应用"。
Tables and indexes don't necessarily live on the same node, so joins, foreign-key checks, and index lookups can turn into multiple cross-node RPCs; the vendor's own CEO admits developers must "spend extra cycles optimizing their existing Postgres apps" for the distributed nature.
- 窄场景
从单机 PG 迁移过来的 OLTP 应用,未按分布键重做数据建模,直接 lift-and-shift 上 YugabyteDB。
OLTP applications migrated from single-node Postgres without remodeling around distribution keys — straight lift-and-shift onto YugabyteDB.
- 机制
YSQL 复用 PG 查询层,但存储是 DocDB 分布式 KV:相关表/索引按 tablet 散列到不同节点,一次 join 或外键检查可能触发多次内部网络往返;分布式事务要走两阶段提交与 tablet 间协调。单分片事务快,跨分片事务付延迟税。厂商的缓解手段(colocated 表、批量、pushdown)都要求改建模,不是零成本。
YSQL reuses the Postgres query layer but stores data in DocDB, a distributed KV store: related tables/indexes hash across nodes per tablet, so one join or foreign-key check can trigger many internal network round trips; distributed transactions go through two-phase commit and cross-tablet coordination. Single-shard transactions are fast; cross-shard transactions pay the latency tax. The vendor's mitigations (colocated tables, batching, pushdowns) all require remodeling — not free.
- 生产验证
来源 4:HN 独立用户 2024-02 评论——"I think this con is very real":相关表与索引不一定存放在一起,join、外键评估、索引查找都可能产生过多内部网络 hops,强事务保证的额外锁与协调也会拖慢性能;"把整张表存在单节点"的缓解手段"违背了分片数据库的大部分收益";
来源 5:InfoWorld 2024-09 报道(zicos 转存)——Yugabyte 联合创始人兼 co-CEO Karthik Ranganathan 承认 "some developers faced challenges with query performance when porting applications from Postgres to YugabyteDB","Developers had to spend extra cycles optimizing their existing Postgres apps for YugabyteDB's distributed nature";2.19 为此加入 CBO、双模执行等特性。
Source 4: independent HN user, Feb 2024 — "I think this con is very real": related tables and indexes are not necessarily stored together, so joins, foreign-key evaluation, and even simple index lookups can incur excessive internal network hops, and the extra locks/coordination of strong transactional guarantees drag performance; the "store the whole table on a single node" workaround "defeats many of the benefits of these sharded SQL databases";
Source 5: InfoWorld, Sep 2024 (via zicos mirror) — Yugabyte co-founder and co-CEO Karthik Ranganathan admitted "some developers faced challenges with query performance when porting applications from Postgres to YugabyteDB" and that "Developers had to spend extra cycles optimizing their existing Postgres apps for YugabyteDB's distributed nature"; 2.19 added a CBO and bimodal execution in response.
- 证据等级
`多方印证(2 个独立来源)`,独立社区评论 + 独立媒体转述的厂商承认(非生产复盘,已如实标注)。
`Corroborated (2 independent sources)`, independent community comment + vendor admission relayed by independent media (not production postmortems — stated as-is).
- 备注
2.19 已加入 CBO/双模执行等缓解手段;本卡主题可能与本站 [避坑] 卡重叠。
2.19 added mitigations (CBO, bimodal execution); this card's topic may overlap an existing [Pitfall] card on this site on this site.
YugabyteDB 年份:2024
运维税:每个慢查询都要走一遍"为什么这么慢"排查手册
多方印证
性能问题运维复杂度
- 一句话
存算耦合、WLM 队列靠手调、每张表都要设计分布键和排序键——Redshift 的慢查询排查是一套固定手册,团队为此长期投入专人。
Coupled compute/storage, hand-tuned WLM queues, per-table distribution and sort key design — slow-query diagnosis on Redshift is a fixed runbook, and teams staff for it permanently.
- 窄场景
DC2/DS2 时代及未做精细调优的集群;ETL、BI 报表、ad-hoc 查询混跑,查询模式多样的团队。
DC2/DS2-era clusters or any cluster without meticulous tuning; teams mixing ETL, BI reports, and ad-hoc queries with diverse query patterns.
- 机制
Redshift 是 shared-nothing MPP,数据按分布键切分到各节点。查询性能同时取决于:分布键是否避免广播与倾斜、排序键是否让 zone map 生效、WLM 队列的并发槽位与内存分配、表是否及时 VACUUM/ANALYZE。任一环节失配就慢,而"慢"的原因分散在 STL_WLM_QUERY、SVL_QUERY_REPORT 等系统表里,只能按手册逐项排查,没有一键归因。
Redshift is a shared-nothing MPP; data is sliced across nodes by distribution key. Query performance depends simultaneously on whether the distribution key avoids broadcast/skew, whether the sort key lets zone maps prune, the WLM queue's concurrency slots and memory allocation, and whether tables got timely VACUUM/ANALYZE. Any one of these being off means slow, and the "why" is scattered across system tables (STL_WLM_QUERY, SVL_QUERY_REPORT) — diagnosis is a manual checklist with no one-click attribution.
- 生产验证
来源 1:Faire 2023-02 具名复盘——Redshift "very high maintenance cost and needed consistent efforts to keep it fast";非 RA3 架构存算耦合,无法隔离工作负载;为回答 "why is my query so slow?" 有一整套 on-call 排查手册流程;100+ Airflow ETL、Mode 报表、SageMaker/Jupyter、AWS Batch 混跑时,WLM 队列调优与 SLA 难以兼顾;
来源 2:Robin 2023-09 迁移复盘——迁往 Snowflake 的理由之一是"不用再担心扩缩容时的维护窗口与停机",并期待 zero-copy cloning、time travel、integrated monitoring 等"quality of life"改进,反衬日常运维负担;
来源 3:HN 讨论串 2022-05——Redshift 性能调优公司创始人直言 "Too many knobs to turn, and it's just not something customers wanted to do";同串另一 5 年用户称 Redshift 经验是"constantly keep giving it TLC",从建用户/schema、WLM 调优到 Spectrum 的 Parquet 文件组织"everything was a chore",切到 BigQuery 后"mostly self driving"。
Source 1: Faire, Feb 2023, named retrospective — Redshift had "very high maintenance cost and needed consistent efforts to keep it fast"; the non-RA3 coupled architecture made workload isolation impossible; answering "why is my query so slow?" required a full on-call triage runbook; with 100+ Airflow ETLs, Mode reports, SageMaker/Jupyter, and AWS Batch sharing one cluster, WLM tuning and SLAs were hard to reconcile;
Source 2: Robin, Sep 2023, migration retrospective — one reason for moving to Snowflake was "we wouldn't have to worry about maintenance windows or downtime when scaling up or down," plus hoped-for quality-of-life gains (zero-copy cloning, time travel, integrated monitoring) that underline the day-to-day burden;
Source 3: HN thread, May 2022 — the founder of a Redshift performance-tuning company: "Too many knobs to turn, and it's just not something customers wanted to do"; another 5-year user in the same thread called Redshift "constantly keep giving it TLC," from creating users/schemas to WLM tuning to organizing Parquet files for Spectrum access — "everything was a chore" — and "mostly self driving" after switching to BigQuery.
- 证据等级
`多方印证(3 个独立来源)`,具名公司工程博客 ×2 + HN 业内人士评论。
`Corroborated (3 independent sources)`, named company engineering blogs x2 + HN industry comment.
- 备注
本卡主题可能与本站 [避坑] 卡重叠(WLM/运维税)。
this card's topic may overlap existing [Pitfall] cards on this site (WLM/ops tax).
Amazon Redshift 年份:2023
扩缩容就是一次停服:换节点类型/规格要整集群停机
多方印证
运维复杂度升级迁移
- 一句话
在 Redshift 上扩容或换节点类型不是在线操作——旧流程要停整个集群再拉起来,终端用户会感知到明显中断。
Scaling a Redshift cluster or switching node types is not an online operation — the classic path stops the entire cluster and brings it back up, and end users feel a noticeable interruption.
- 窄场景
DC1/DC2 时代及需要换节点家族(DC2→RA3)或纵向扩缩的集群。
DC1/DC2-era clusters, or any cluster needing a node-family switch (DC2 to RA3) or vertical rescaling.
- 机制
classic resize 会终止所有执行中的查询,期间集群只读甚至不可用;elastic resize 仍需短暂停机;classic resize 只能按 2 倍乘子调整节点数。存算耦合架构下"加存储"必须"连计算一起买"。
Classic resize terminates all running queries and leaves the cluster read-only or unavailable during the operation; elastic resize still needs brief downtime; classic resize only allows 2x node-count multipliers. Under the coupled architecture, "add storage" means "buy compute with it."
- 生产验证
来源 4:PeerSpot 2023-07-27,Barathwaj Ramamoorthy(Senior Director Data Architecture,平台标注 Real User)——因扩展性不足切到 Snowflake:"if I need to switch from DC1 to DC2, or from one compute/storage optimization to another, I have to bring down the entire cluster and then bring it back up. That's a pain point.";
来源 2:Robin 2023-09——迁移动因之一:"We wouldn't have to worry about maintenance windows or downtime when scaling up or down.";
来源 4:PeerSpot 同页匿名评论(2023-05)——对比 Snowflake 可即时调整虚拟仓库大小,Redshift 做不到。
Source 4: PeerSpot, Jul 27 2023, Barathwaj Ramamoorthy (Senior Director Data Architecture, site-labeled Real User) — switched to Snowflake over scalability: "if I need to switch from DC1 to DC2, or from one compute/storage optimization to another, I have to bring down the entire cluster and then bring it back up. That's a pain point.";
Source 2: Robin, Sep 2023 — one migration driver: "We wouldn't have to worry about maintenance windows or downtime when scaling up or down.";
Source 4: another PeerSpot comment on the same page (May 2023, anonymous) — contrasts Snowflake's instant virtual-warehouse resizing, which Redshift cannot do.
- 证据等级
`多方印证(3 个独立来源)`,具名点评用户 + 匿名点评用户 + 具名公司工程博客(点评页身份经平台标注,未经独立核实)。
`Corroborated (3 independent sources)`, named reviewer + anonymous reviewer + named company engineering blog (reviewer identities site-labeled, not independently verified).
Amazon Redshift 年份:2023
Robin:Redshift→Snowflake 迁移中最磨人的不是数据搬运,而是 SQL 翻译
多方印证
升级迁移
- 一句话
5 个月迁移里最痛的是 "SQL Translation"——一串静默行为差异:Redshift 里 `greatest(1, 10, 1000, null)` 返回 1000,Snowflake 遇到 null 直接返回 null;时间戳默认精度不同导致 `md5(created_at::varchar)` 两仓算出不同哈希;`type`、`start` 这类列名在 Snowflake 是保留关键字——"查得出来、修得回去,但机器翻译工具扫不出来"。
In a 5-month migration the most painful part was "SQL Translation" — a string of silent behavior differences: `greatest(1, 10, 1000, null)` returns 1000 in Redshift but NULL in Snowflake; different default timestamp precision made `md5(created_at::varchar)` hash differently in each warehouse; column names like `type` and `start` are reserved keywords in Snowflake — "findable, fixable, but invisible to machine translation tools."
- 窄场景
从 Redshift(或其他数仓)迁入 Snowflake;历史 SQL 包袱重的团队。
Migrating into Snowflake from Redshift (or another warehouse); teams with heavy legacy SQL.
- 机制
函数语义、类型精度、保留字等静默差异不在语法层面,SnowConvert 类官方工具最容易漏掉;团队总结的 QA 原则:"It's not reasonable to expect that data will be 100% the same between the two warehouses… the purpose of QA is to be able to explain differences"(别指望两仓数据 100% 一致,QA 的目标是能解释差异)。
Function semantics, type precision and reserved words differ silently below the syntax level — exactly what official tools like SnowConvert miss most; the team's QA principle: "It's not reasonable to expect that data will be 100% the same between the two warehouses… the purpose of QA is to be able to explain differences between Redshift and Snowflake."
- 生产验证
来源 40:Robin 数据工程师 Eric Wurtzbacher 2023-09-21 第一人称复盘;
来源 40 内引述:Instacart 数据团队早前在类似迁移中记录过同类 SQL 函数翻译坑("The SQL Challenge")。
—
- 证据等级
`多方印证(2 个独立来源)`,Robin 一线复盘 + Instacart 同类记录(经 Robin 引述)。
`Corroborated (2 independent sources)`, Robin frontline postmortem + Instacart's analogous record (via Robin's citation).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2023
没有查询工作负载管理器:重要查询和跑批只能"分仓",没有 WLM 式的资源调度
多方印证
运维复杂度
- 一句话
"No Snowflake query workload manager: Unlike other data warehouses, Snowflake lacks a resource manager to assign resources to queries based on their importance and overall system workload."(不像其他数仓,它缺少按查询重要性和系统整体负载分配资源的资源管理器。)媒体批评汇总:"Workload management. Or lack thereof. If you want to isolate higher-performance workloads with Snowflake, you just spin up a separate virtual warehouse. It works generally but will add expense."(想隔离高优先级负载,办法就是另开一台虚拟仓库。管用,但要加钱。)
"No Snowflake query workload manager: Unlike other data warehouses, Snowflake lacks a resource manager to assign resources to queries based on their importance and overall system workload." Media criticism roundup: "Workload management. Or lack thereof. If you want to isolate higher-performance workloads with Snowflake, you just spin up a separate virtual warehouse. It works generally but will add expense."
- 窄场景
同一 warehouse 里跑批 ETL 和交互式 BI 互相挤占的团队;需要"这个查询优先"语义的组织。
Teams where batch ETL and interactive BI contend in the same warehouse; orgs needing "this query goes first" semantics.
- 机制
Snowflake 的隔离单位是 warehouse 整台,没有查询级优先级调度;隔离=拆 warehouse=每个 warehouse 独立计费="隔离"直接等于"加钱"。Redshift 的 WLM 队列(查询优先级、并发槽位、内存配比)是参照物。
Snowflake's isolation unit is the whole warehouse — no query-level priority scheduling; isolation means splitting warehouses, each a separate billing unit, so "isolation" literally equals "more spend." Redshift's WLM queues (query priority, concurrency slots, memory ratios) are the reference point.
- 生产验证
来源 14:Slim Baltagi §2.17(2023-11-08,与本清单
来源 14 为同一篇文章的不同章节);
来源 53:SiliconANGLE 2020-11-14 媒体批评汇总(汇总"dozens and dozens"客户/从业者观点,仅作印证)。
—
- 证据等级
`多方印证(2 个独立来源)`,独立从业者长文 + 媒体批评汇总。
`Corroborated (2 independent sources)`, independent practitioner long-form + media criticism roundup.
- 备注
缺口状态:至今缺失(截至 2026-10,隔离手段仍是拆 warehouse)。本卡主题可能与本站 [避坑] 卡重叠。
gap status: still missing (as of 2026-10, isolation still means splitting warehouses). This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2023
节点小时计费:低吞吐 workload 在为"不存在的流量"买单
多方印证
成本账单
- 一句话
Spanner 按预置节点小时收费,存储配额又和节点数绑定——吞吐低、数据量大的 workload 会被迫为闲置算力买单。
Spanner bills by provisioned node-hours, and storage quotas are tied to node counts — workloads with low throughput but large data footprints end up paying for idle compute.
- 窄场景
低 QPS 但数据量大的库(被迫加节点只为拿存储配额)、流量波动大(只能按峰值预置)的业务。
Low-QPS but data-heavy databases (forced to add nodes just for storage headroom), and spiky workloads (forced to provision for the peak).
- 机制
Spanner 的计算容量(节点/PU)同时决定吞吐上限和存储上限:小于 1 节点时每 100 PU 约 1TiB(1024.0 GiB),1 节点及以上每节点 10TiB;官方文档明确"Spanner doesn't have a suspend mode"——算力是专用资源,闲置也在计费。于是出现两种交税姿势:低吞吐大存储 → 加节点只为存储配额,闲置算力照单全收;流量波动大 → 按峰值预置,平均利用率低时大量容量空转。
Spanner's compute capacity (nodes/PUs) determines both the throughput ceiling and the storage ceiling: ~1 TiB (1024.0 GiB) per 100 PUs under 1 node, 10 TiB per node at 1 node and above; the docs state explicitly "Spanner doesn't have a suspend mode" — compute is a dedicated resource, billed even when idle. That yields two ways to pay the tax: low-throughput, large-data workloads add nodes purely for storage quota and absorb the idle compute; spiky workloads provision for the peak and watch most capacity sit idle on average.
- 生产验证
来源 1(HN 2023):"We had a huge spanner db with low throughput so had to add idle nodes just for storage which also ballooned costs";
来源 1(HN 2023,另一用户):"I have to provision for peak throughput on Spanner. Average throughput is much lower than the peak throughput, so I'm doubtful of seeing savings";
来源 1(HN 2023,另两位用户):"a production ready instance starts at $65/mo. DynamoDB can run for ~$0.00/month with per-request pricing" / "But you have to start by paying like $1k/mo, it's not serverless"。
Source 1 (HN, 2023): "We had a huge spanner db with low throughput so had to add idle nodes just for storage which also ballooned costs";
Source 1 (HN, 2023, a different user): "I have to provision for peak throughput on Spanner. Average throughput is much lower than the peak throughput, so I'm doubtful of seeing savings";
Source 1 (HN, 2023, two more users): "a production ready instance starts at $65/mo. DynamoDB can run for ~$0.00/month with per-request pricing" / "But you have to start by paying like $1k/mo, it's not serverless".
- 证据等级
`多方印证(4 个独立声音)`,同一 HN 讨论帖内的 4 位不同用户;计费与存储耦合机制经官方文档核实。
`Corroborated (4 independent voices)`, four different users in the same HN thread; the billing/storage coupling verified against official documentation.
- 备注
2023 年帖子里的 "$1k/mo 起步"吐槽基于当时的 3 节点起步口径;2024–2026 年 Google 改为 Standard/Enterprise/Enterprise Plus 版本体系并支持 100 PU 起步,绝对数字已变,但"按节点小时预置计费"的模型未变,本卡保留。
the "$1k/mo to start" complaints in the 2023 thread reflect the 3-node minimum of that era; Google's 2024–2026 Standard/Enterprise/Enterprise Plus editions allow starting at 100 PUs, so the absolute numbers have changed — but the provisioned node-hour billing model has not, so the card stands.
Google Spanner 年份:2023
热分区:你付了 1000 RPS 的钱,热点键却只能吃到其中一小份
多方印证
性能问题成本账单
- 一句话
Spanner 把键空间切成多个 split 并把配额均分——写入集中在少数键范围时,只有"热点分片"那一份配额能用,要么整体超配,要么在应用层自建缓存削峰。
Spanner splits the key space and divides quotas evenly across splits — when writes concentrate on a few key ranges, only the hot split's share of the quota is usable, forcing you to overprovision the whole instance or build a cache layer yourself.
- 窄场景
写分布不均的 workload(少数热键/热租户)、读多写少但写集中的业务。
Skewed-write workloads (a few hot keys or hot tenants), read-heavy workloads with concentrated writes.
- 机制
数据按主键字典序切分为 splits 分散到各服务器;单个键范围的写入吞吐受单个 split 的服务能力上限约束,加节点并不能提升热点范围的写入能力。热点出现后只有两条路:为热点超额预置整体算力(大部分闲置),或在应用层做 write-through 缓存把热点削平——而后者等于把本该数据库做的事搬回应用层。
Data is split lexicographically by primary key across servers; write throughput for a single key range is capped by what one split can serve, and adding nodes does not raise a hot range's write ceiling. Once a hotspot appears there are two escapes: overprovision the whole instance for the hotspot (most of it idle), or build a write-through cache in the application to flatten the hotspot — which moves work the database should do back into the application.
- 生产验证
来源 1(HN 2023,细节充分的生产叙述):"we ended up creating a very complicated write through cache system in front of spanner that dynamically added memory/CPU capacity as needed to prevent hot shards",Spanner 每月数万美元 + 前置缓存计算每月数万美元;写只有几百 RPS、读是写的 1000 倍,"Postgres behind this cache system would have handled the load just as well and cost less than half as much";并抱怨 "Google's docs are incomplete (as usual); there are lots of performance gotchas like this that exist throughout the entire service, and they aren't clearly documented";
来源 1(HN 2023,另一用户):"I have seen two applications in my career... we paid an absolute arm and a leg for the privilege and ended up implementing a distributed cache in front of Spanner to reduce costs"。
Source 1 (HN, 2023, detailed first-hand account): "we ended up creating a very complicated write through cache system in front of spanner that dynamically added memory/CPU capacity as needed to prevent hot shards" — tens of thousands of dollars a month for Spanner plus tens of thousands a month for the caching compute; writes were only a few hundred RPS with reads 1000x that: "Postgres behind this cache system would have handled the load just as well and cost less than half as much"; plus "Google's docs are incomplete (as usual); there are lots of performance gotchas like this that exist throughout the entire service, and they aren't clearly documented";
Source 1 (HN, 2023, a different user): "I have seen two applications in my career... we paid an absolute arm and a leg for the privilege and ended up implementing a distributed cache in front of Spanner to reduce costs".
- 证据等级
`多方印证(2 个独立来源)`,同一 HN 帖内 2 位不同用户。
`Corroborated (2 independent sources)`, two different users in the same HN thread.
Google Spanner 年份:2023
DeWitt 条款:你不可以公开 benchmark 它
多方印证
生态与信任
- 一句话
Google Cloud 服务条款第 7 条要求——公开 Spanner 的 benchmark 结果必须先拿到 Google 书面同意,还得允许 Google 反过来测你的产品。
Google Cloud's service terms, section 7, require written Google consent before publishing Spanner benchmark results — and grant Google the right to benchmark your products in return.
- 窄场景
想做独立第三方性能对比、发表选型 benchmark 的团队与媒体。
Teams and media wanting independent third-party performance comparisons or publishing selection benchmarks.
- 机制
Google Cloud 通用服务条款 §7 "Benchmarking":Customer may only publicly disclose the results of such Tests if (a) obtains Google's prior written consent, (b) provides Google all necessary information to replicate the Tests, (c) allows Google to conduct benchmark tests of Customer's publicly available products or services and publicly disclose the results;且"on behalf of a hyperscale public cloud provider"时连"测"本身都要事先书面同意。效果:独立第三方不愿碰,公开可复现的对比数据长期缺席——这也是 Spanner 深度吐槽稀少的原因之一。
Google Cloud General Service Terms, section 7 "Benchmarking": "Customer may only publicly disclose the results of such Tests if (a) obtains Google's prior written consent, (b) provides Google all necessary information to replicate the Tests, (c) allows Google to conduct benchmark tests of Customer's publicly available products or services and publicly disclose the results" — and tests "on behalf of a hyperscale public cloud provider" need prior written consent even to run. The effect: independent third parties stay away, and reproducible public comparisons remain absent — one reason deep Spanner rants are so rare.
- 生产验证
来源 2(Cube 博客 2022-06,逐字引述条款原文):把 GCP(含 Spanner 适用的通用条款)列入"有 DeWitt 条款"名单;
来源 3(Google 开发者论坛 2022-08):用户核实 "customers cannot publicly disclose Cloud Spanner benchmarking results without getting written consent from Google";
来源 1(HN 2023):"If only someone could actually run and publish comparison benchmarks, but DeWitt clause by Spanner makes it impossible."(同帖另一用户贴出条款原文逐条解读)。
Source 2 (Cube blog, Jun 2022, quoting the terms verbatim): lists GCP (whose general terms cover Spanner) among vendors with a DeWitt clause;
Source 3 (Google developer forum, Aug 2022): users verify "customers cannot publicly disclose Cloud Spanner benchmarking results without getting written consent from Google";
Source 1 (HN, 2023): "If only someone could actually run and publish comparison benchmarks, but DeWitt clause by Spanner makes it impossible." (another commenter in the same thread posted and parsed the clause text line by line).
- 证据等级
`多方印证(3 个独立来源)`,独立厂商博客 + 官方论坛用户核实 + HN 用户。
`Corroborated (3 independent sources)`, independent vendor blog + official-forum user verification + HN user.
Google Spanner 年份:2023
Google 式定价阴影:今天半价,明天呢
多方印证
成本账单生态与信任
- 一句话
Spanner 只能跑在 GCP 上,而 Google 有过 Maps 10–20 倍、App Engine 10 倍式涨价的前科——把核心数据库押在 GCP 上,等于把定价权交给一家"说变就变"的公司。
Spanner only runs on GCP, and Google has a track record of Maps-style 10–20x and App Engine 10x price hikes — betting your core database on GCP means handing pricing power to a company that changes its mind.
- 窄场景
把 Spanner 作为核心系统数据库、计划用 5–10 年的团队;对多云/可迁移性有要求的组织。
Teams adopting Spanner as the core system of record for 5–10 years; organizations that need multi-cloud or exit options.
- 机制
Spanner 是 GCP 专属服务(无自托管、无其他云),迁出意味着重写数据层;Google 历史上对 Maps API(2018 年约 10–20 倍涨价)、App Engine(定价上调约 10 倍)都做过大幅调价。用户担心的不是"会不会涨",而是"涨了你毫无议价能力"。
Spanner is GCP-exclusive (no self-hosting, no other cloud) — leaving means rewriting the data layer; Google previously imposed ~10–20x price increases on Maps APIs (2018) and ~10x on App Engine. What users fear is not "whether" but "zero leverage when it happens."
- 生产验证
来源 1(HN 2023,多位用户):"Google Maps is the key lesson here. 10 to 20 times price increase, just because someone had a meeting." / "After they introduced Google Cloud they got bored of App Engine and 10x'd the price." / "Google also has a history of massively spiking the cost of its services. Vendor lockin is a dangerous thing." / "At the end of the day... I trust AWS to be a stable, long term foundation to build a product on, I don't trust GCP to be the same."
Source 1 (HN, 2023, multiple users): "Google Maps is the key lesson here. 10 to 20 times price increase, just because someone had a meeting." / "After they introduced Google Cloud they got bored of App Engine and 10x'd the price." / "Google also has a history of massively spiking the cost of its services. Vendor lockin is a dangerous thing." / "At the end of the day... I trust AWS to be a stable, long term foundation to build a product on, I don't trust GCP to be the same."
- 证据等级
`多方印证(4 个独立声音)`,同一 HN 帖内 4 位不同用户。
`Corroborated (4 independent voices)`, four different users in the same HN thread.
- 备注
涨价案例(Maps/App Engine)是 Google 其他产品线的历史,非 Spanner 本身的涨价记录;本卡收录的是"绑定 GCP 的定价信任风险"这一用户真实顾虑,不代表 Spanner 涨过价。
the price-hike cases (Maps, App Engine) are other Google product lines' history, not Spanner's own record; the card captures users' genuine "GCP-bound pricing trust" concern, not a claim that Spanner itself raised prices.
Google Spanner 年份:2023
为不存在的规模提前买单:Spanner 是"火车",你只想要"货车"
多方印证
性能问题运维复杂度
- 一句话
Spanner 的复杂度与成本是为 planet-scale 设计的——团队在只有个位数用户时就为"未来规模"选它,等于提前支付分布式协调税,却迟迟拿不到收益。
Spanner's complexity and cost are engineered for planet scale — teams that pick it with a handful of users "just in case" pay the distributed-coordination tax up front while the payoff stays perpetually out of reach.
- 窄场景
初创/中小团队、流量远没到单机瓶颈、"以防万一要全球扩展"而选 Spanner。
Startups and small teams far from any single-node bottleneck, choosing Spanner "in case we need to go global."
- 机制
分布式事务的 2PC/Paxos/commit-wait 成本是每笔操作都要付的固定税,而收益(突破单机上限)只在超过单机容量后才兑现。规模没到之前,你同时支付了"分布式复杂度"和"单机本可免费拿到的低延迟"。
Distributed-transaction costs (2PC/Paxos/commit-wait) are a fixed tax on every operation, while the benefit (breaking past single-node capacity) only materializes after you outgrow one machine. Before that point, you pay both "distributed complexity" and "the low latency a single node would have given you for free."
- 生产验证
来源 1(HN 2023,多位用户):"It seems like a lot of people like to plan for massive scale while they have a handful of actual users." / "Spanner is hard to recommend from an engineering perspective at anything except the absolute most massive of scales, and even then it will create nearly as many problems as it solves." / "My team bought the 'scale down' thing and got bit." / "The more I deal with scalable relational DBMSes (Spanner in particular), the more I doubt their usefulness even at large scale."
Source 1 (HN, 2023, multiple users): "It seems like a lot of people like to plan for massive scale while they have a handful of actual users." / "Spanner is hard to recommend from an engineering perspective at anything except the absolute most massive of scales, and even then it will create nearly as many problems as it solves." / "My team bought the 'scale down' thing and got bit." / "The more I deal with scalable relational DBMSes (Spanner in particular), the more I doubt their usefulness even at large scale."
- 证据等级
`多方印证(4 个独立声音)`,同一 HN 帖内 4 位不同用户。
`Corroborated (4 independent voices)`, four different users in the same HN thread.
Google Spanner 年份:2023
Serverless v1:扩容要先找"事务间隙",找不到就卡你 5–50 秒
多方印证
性能问题运维复杂度
- 一句话
v1 扩容的办法是"冻结数据库、搬到另一台 VM、再启动"——前提是找到一个没有重叠事务的时间窗口;业务一忙,这个窗口根本不存在。
v1 scaled by "freezing the database, moving to another VM, and restarting" — which required finding a window with no overlapping transactions; on a busy database that window simply does not exist.
- 窄场景
持续有并发事务的 v1 集群("chatty" 数据库);依赖缩容到 0 省钱的 dev/staging 环境。
v1 clusters with sustained concurrent transactions ("chatty" databases); dev/staging environments relying on scale-to-zero savings.
- 机制
v1 按 1 ACU 起步、翻倍扩容;扩容需暂停数据库并跨 VM 迁移内存状态(buffer pool 跨网络拷贝),大内存下拷贝本身就很慢;AWS 自述:很忙的库找不到事务间隙时,扩容耗时 5–50 秒,偶尔还会直接打断数据库;且 v1 的计算层没有多可用区高可用。
v1 started at 1 ACU and doubled on scale-up; scaling paused the database and migrated in-memory state (buffer pool copied across the network) to a new VM — slow by itself at large memory sizes; AWS's own account: when no transaction gap could be found on a busy database, scaling took 5-50 seconds and occasionally disrupted the database outright; v1 compute had no multi-AZ high availability either.
- 生产验证
来源 3:The Register 2022-04-29(AWS 赞助内容,此处仅引用 AWS 产品经理 Chayan Biswas 的自述)——亲口承认上述 5–50 秒机制与"有时打断数据库",并称这限制了 v1 只适合零星负载;
来源 4:HN 用户 2022 年详述——v1 扩容"慢且经常失败",高峰前只能手动先扩到超大;dev/QA 环境从休眠唤醒"差不多要一分钟";Data API"是个错误决定";
来源 5:re:Post 用户——开了缩容到 0 的 dev/staging 环境出现随机超时,连 `SELECT * FROM users WHERE user_id = 1` 都会超时,CPU 与慢日志一切正常,被判定为 v1 冷启动/扩容问题。
Source 3: The Register, Apr 29 2022 (AWS-sponsored piece; only AWS PM Chayan Biswas's self-description is used) — he openly described the 5-50 second mechanism and the "sometimes disrupts the database" behavior, saying it confined v1 to sporadic workloads;
Source 4: HN user, 2022, in detail — v1 scaling was "slow and often failed," forcing manual pre-scaling before events; dev/QA environments took "almost a minute" to wake from pause; the Data API was "a mistake";
Source 5: re:Post user — scale-to-zero dev/staging environments saw random timeouts, even on `SELECT * FROM users WHERE user_id = 1`, with CPU and slow logs all clean; diagnosed as a v1 cold-start/scaling issue.
- 证据等级
`多方印证(3 个独立来源)`,AWS 官方自述(经赞助媒体转述)+ HN 用户实测 + re:Post 生产求助。
`Corroborated (3 independent sources)`, AWS's own account (via sponsored media) + HN user field report + re:Post production help request.
- 备注
v1 已于 2025-03-31 停止支持(见下张卡片),本卡为历史记录;继任者 v2 的问题见"Serverless v2:0.5 ACU 常开"卡。
v1 reached end of support on Mar 31 2025 (see next card); this card is a historical record. v2's issues are covered in the "Serverless v2: 0.5 ACU always on" card.
Amazon Aurora 年份:2022
CQL 的坑位清单:长得像 SQL,处处是例外
多方印证
运维复杂度
- 一句话
CQL 的语法在"骗"你以为它是 SQL——LWT 只是单分区、quorum 写失败后数据在不在要靠猜、物化视图加了又弃、二级索引只能查单分区、insert 和 select 的时间戳格式还不一样;用下来感觉像"例外之上的例外"。
CQL's syntax tricks you into thinking it is SQL — LWTs are single-partition only, after a failed quorum write you have to guess whether the data is there, materialized views were added then deprecated, secondary indexes only query single partitions, and insert and select do not even share a timestamp format; it feels like "exceptions on top of exceptions."
- 窄场景
把 CQL 当 SQL 用的团队;依赖 LWT 做条件写、依赖 MV/二级索引做查询的 schema。
Teams treating CQL as SQL; schemas relying on LWTs for conditional writes or on MVs/secondary indexes for queries.
- 机制
LWT(`IF NOT EXISTS`/`IF`)底层是 Paxos,但只保证单分区线性化,且延迟数倍于普通写、热点竞争下更糟;quorum 写"失败"不等于没写入(超时 vs 拒绝语义模糊),错误处理要按"maybe"写;BATCH 保证原子性但不保证隔离性;物化视图经历"加入—标记实验性—长期边缘"的反复;二级索引本质是各节点本地索引。
LWTs (`IF NOT EXISTS`/`IF`) are Paxos underneath but guarantee only single-partition linearizability, cost multiples of a normal write, and degrade further under hot-partition contention; a "failed" quorum write does not mean the write did not happen (timeout vs. rejection semantics are fuzzy), so error handling must be written for "maybe"; BATCH is atomic but not isolated; materialized views went through an add—experimental—long-edge-case cycle; secondary indexes are fundamentally per-node local indexes.
- 生产验证
来源 2:2022-12,HN 用户逐条列举——"LWT are single row only…error handling is asinine. E.g: A quorum insert fails. is the data there or not? Maybe!…material views. Added, removed, readded, deprecated…BATCH is atomic, but not isolated…CQL feels like a hack…It's just exception on top of exception";
来源 2:2022-12,同一讨论串另一用户——"Cassandra is good at ingesting data, bad at deleting, really really bad at anything remotely relational. Errors are almost pointless…Don't trust what the CQL language says you can do."
Source 2: 2022-12, HN user enumerating — "LWT are single row only…error handling is asinine. E.g: A quorum insert fails. is the data there or not? Maybe!…material views. Added, removed, readded, deprecated…BATCH is atomic, but not isolated…CQL feels like a hack…It's just exception on top of exception";
Source 2: 2022-12, another user in the same thread — "Cassandra is good at ingesting data, bad at deleting, really really bad at anything remotely relational. Errors are almost pointless…Don't trust what the CQL language says you can do."
- 证据等级
`多方印证(2 个独立作者,同一 HN 讨论串)`,两条评论相互独立、指向同一组语义陷阱。
`Corroborated (2 independent authors in one HN thread)`, two independent comments pointing at the same set of semantic traps.
- 备注
与现有 [避坑] 卡"兼容坑"(CQL 无 join/子查询)角度不同——现有卡讲"缺能力",此卡讲"语义陷阱",双方保留。
Different angle from existing [Pitfall-avoidance] "compatibility pitfall" cards (no joins/subqueries in CQL) — those document missing capabilities, this one documents semantic traps; both are kept.
Apache Cassandra / ScyllaDB 年份:2022
最终一致性的心智负担:应用层替数据库操心
多方印证
运维复杂度
- 一句话
读写一致性级别选什么、每档的延迟与可见性代价是多少、读到旧数据时应用层怎么兜——这些本该是数据库的事,在 Cassandra 里全是应用工程师的必修课;quorum 读"太慢没人用",ONE 读"读到旧数据出 bug",两头为难。
Which read/write consistency level to pick, what each level costs in latency and visibility, how the application copes when it reads stale data — things that should be the database's job are mandatory homework for every Cassandra application engineer; quorum reads are "too slow, nobody uses them," and ONE reads produce "why isn't the data I just wrote showing up" bugs.
- 窄场景
多副本、多 DC 部署;读写一致性级别混用的应用;对"读到自己刚写的数据"有预期的产品场景。
Multi-replica, multi-DC deployments; applications mixing consistency levels; product scenarios that expect read-your-writes.
- 机制
无主复制 + 可调一致性:`W+R>N` 只保证读写集合相交,不保证语义收敛;低级别读可能命中尚未收到最新写的副本;quorum 读写要多等副本,延迟翻倍且继承最慢副本的尾延迟;想跨分区强一致只有 LWT(Paxos,单分区、昂贵)。"最终一致"收敛靠 read repair 与 anti-entropy repair,而 repair 又是另一个运维科目(见本页 repair 卡)。
Leaderless replication + tunable consistency: `W+R>N` only guarantees the read and write sets intersect, not semantic convergence; low-level reads can hit replicas that have not received the latest write; quorum reads wait on more replicas, multiplying latency and inheriting the slowest replica's tail; the only cross-partition strong consistency is LWTs (Paxos, single-partition, expensive). "Eventual" convergence relies on read repair and anti-entropy repair — and repair is itself another operational discipline (see the repair card on this page).
- 生产验证
来源 5:HN 用户生产经历——"We didn't have a good handle on the exact perf implications of different values of read/write replication. Writing product code to handle a range of eventual consistency scenarios is challenging";
来源 2:2022-12,HN 讨论串——"nobody does reads at quorum, they're slow"(quorum 读太慢没人用);"Unless they want 'read your writes', otherwise bugs start appearing 'why isn't the data i just put showing'?"(不要 read-your-writes,就等着"刚写的数据怎么查不到"的 bug)。
Source 5: HN user production experience — "We didn't have a good handle on the exact perf implications of different values of read/write replication. Writing product code to handle a range of eventual consistency scenarios is challenging";
Source 2: 2022-12, HN thread — "nobody does reads at quorum, they're slow"; "Unless they want 'read your writes', otherwise bugs start appearing 'why isn't the data i just put showing'?"
- 证据等级
`多方印证(2 个独立来源)`,HN 生产经历分享 + 2022 年 HN 讨论串印证。
`Corroborated (2 independent sources)`, HN production-experience share + a 2022 HN thread corroborating.
- 备注
主贴经历约 2016 年前后,评论者自述"情况可能已有变化";2022 年讨论串显示一致性级别选型困境仍在被一线讨论,故保留收录,未作"已修复"标注。与现有 [避坑] 卡主题重叠(深水区一:可调一致性),双方保留。
The main experience dates to circa 2016, and the commenter themselves noted things may have changed; the 2022 thread shows the consistency-level dilemma still being debated by frontline users, so the card is retained without a "fixed" marker. Topic overlaps with existing [Pitfall-avoidance] cards (deep-dive on tunable consistency); both are kept.
Apache Cassandra / ScyllaDB 年份:2022
Ghost 5 只认 MySQL 8:应用生态用脚投票
多方印证
生态与信任
- 一句话
Ghost 从 5.0 起官方只支持 MySQL 8,MariaDB 用户要么锁死旧版、要么被迫迁库——生态漂移的账最终由用户买单。
Ghost has officially supported only MySQL 8 since 5.0 — MariaDB users either pin to an old release or migrate the database, and the ecosystem drift lands on the user.
- 窄场景
用 MariaDB 跑 Ghost(及类似只在 MySQL 上做 CI 的上游应用)的自托管用户。
Self-hosters running Ghost (and other upstream apps whose CI only covers MySQL) on MariaDB.
- 机制
应用厂商的测试矩阵只覆盖 MySQL 8;"协议兼容"不等于"行为兼容"——Ghost 4.46 在 MariaDB 上的更新失败并回滚,根因是"MariaDB 与 MySQL 之间的某些差异";厂商随后把支持矩阵收紧为 MySQL 8 only,用户失去选择权:不迁库就等于放弃安全更新。
App vendors' test matrices cover only MySQL 8; "protocol compatible" is not "behavior compatible" — Ghost 4.46 failed to update on MariaDB and rolled back due to "some discrepancy between MariaDB and MySQL"; the vendor then narrowed its support matrix to MySQL 8 only. Users lose the choice: don't migrate the database and you forfeit security updates.
- 生产验证
来源 3:zblesk 2022-06,多年 Ghost+MariaDB 用户——4.46 更新失败被迫回滚;为升 Ghost 5 把库迁到 Docker 里的 MySQL 8,第一次 dump→导入直接失败、推倒重来,戏称"grit my teeth and make the switch";
来源 1:infophreak 专门给"被困在 MariaDB 上的 Ghost 用户"写了 14 步迁移指南(MariaDB 11→MySQL 8),开篇点名 Ghost 官方已宣布 MySQL 8 要求。
Source 3: zblesk 2022-06, a multi-year Ghost+MariaDB user — the 4.46 update failed and was rolled back; to get Ghost 5 the author moved to MySQL 8 in Docker, the first dump→import failed outright and had to be redone from scratch ("grit my teeth and make the switch", author quote);
Source 1: infophreak wrote a dedicated 14-step guide for "Ghost users stuck on MariaDB" (MariaDB 11 → MySQL 8), naming Ghost's official MySQL 8 requirement.
- 证据等级
`多方印证(2 个独立来源)`,个人博客 ×2。
`Multi-source corroborated (2 independent sources)`, personal blogs ×2.
MariaDB 年份:2022
只读节点的慢查询,能把主库的 DDL 一票否决
多方印证
运维复杂度升级迁移
- 一句话
PolarDB 的 DDL 要等所有只读节点同步 MDL——只读节点上一个没结束的大查询/长事务,主库加索引就直接超时失败回滚。
PolarDB DDL must synchronize MDL locks across every read replica — one unfinished long query or long transaction on a replica, and the primary's index build times out and rolls back.
- 窄场景
一写多读集群、只读节点同时跑报表/慢查询;做分区表变更、加索引等 DDL 时。
One-write-many-read clusters where read replicas also serve reporting/slow queries; partition changes, index additions, or other DDL.
- 机制
共享存储架构下,DDL 变更表结构前必须保证所有节点的读操作结束:主节点自己获取 MDL 锁后写一条 redo 日志,只读节点解析到后尝试获取同一表的 MDL,失败则反馈主节点;主节点等待全部只读节点同步(loose_innodb_primary_abort_ddl_wait_replica_timeout 默认 1 小时),只读节点加锁超时由 replica_lock_wait_timeout 控制(默认 50 秒)。于是 MDL 同步的"选民"从主库本地扩大到了每一个只读节点。
Under the shared-storage architecture, a DDL that changes table structure must wait for all read activity on every node to finish: the primary takes the MDL lock itself, writes a redo log entry, and each read replica parses it and tries to take the MDL on the same table, reporting failure back to the primary; the primary waits for all replicas to sync (loose_innodb_primary_abort_ddl_wait_replica_timeout, default 1 hour) while each replica's lock attempt is bounded by replica_lock_wait_timeout (default 50s). The MDL-sync "electorate" thus grows from the local primary to every read replica.
- 生产验证
来源 1:2022-01 事故记录——1 亿多行表改成分区表后,作者第二天执行 DDL 改回非分区表,跑了半小时后报 `ERROR 8007 (HY000): Fail to get MDL on replica during DDL synchronize`;查 `show full processlist` 发现只读节点上有几个 select 持续了一天没结束,正是它们挡住了 MDL 获取,kill 掉才继续;
来源 2:2022-12 社区问答——另一用户提问"polardb添加索引报错Fail to get MDL on replica during DDL synchronize"(已解决;该提问正文无细节,仅报错标题,故作为第二独立声音但偏薄)。
Source 1: Jan 2022 incident record — after converting a 100M+ row table to a partitioned table, the author ran DDL to convert it back the next day; it failed after 30 minutes with `ERROR 8007 (HY000): Fail to get MDL on replica during DDL synchronize`; `show full processlist` showed several SELECTs stuck on a read replica for a full day, and they were what blocked MDL acquisition — the DDL only proceeded after they were killed;
Source 2: Dec 2022 community Q&A — a different user asking about the exact same error when adding an index (marked resolved; the question body is thin — just the error title, no detail — so it counts as a second independent voice, but a thin one).
- 证据等级
`多方印证(2 个独立来源)`,具名事故记录 ×1 + 社区用户提问 ×1(第二个声音仅报错标题、无细节)。
`Corroborated (2 independent sources)`, named incident record x1 + community user question x1 (the second voice is the error title only, no detail).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。另外:两起客户声音均早于 2023 年,但该故障模式源于共享存储 MDL 同步的架构约束,官方文档当前版本仍在描述同一报错与手动缓解参数(非默认修复),故保留收录,未作"已修复"标注。来源 1 文中引用的机制解释转述自淘宝数据库内核月报(厂商侧材料),仅作机制背景参考。
this card's topic may overlap an existing [Pitfall] card on this site on this site. Also: both customer voices predate 2023, but the failure mode stems from the shared-storage MDL-sync architectural constraint, and the vendor's current documentation still describes the same error and the same manual mitigation parameters (not a default fix), so it is retained without a "fixed in" label. The mechanism explanation quoted in source 1 is relayed from the Taobao database kernel monthly report (vendor-side material), used only as mechanism background.
PolarDB 年份:2022
Python Connector 里 fetch_pandas_all() 跑个建表语句就抛"未知错误"
多方印证
生态与信任
- 一句话
`cur.execute("create temp table ...")` 后调 `cur.fetch_pandas_all()`,直接 `NotSupportedError: Unknown error`——cursor.py 里硬编码检查 `if self._query_result_format != "arrow": raise NotSupportedError`;只有 SELECT 返回 Arrow 结果集,DDL/DML 返回 JSON,四个 pandas/arrow 快捷 fetch 方法全部拒绝工作。
Run `cur.execute("create temp table ...")` then `cur.fetch_pandas_all()` and you get `NotSupportedError: Unknown error` — a hardcoded check in cursor.py (`if self._query_result_format != "arrow": raise NotSupportedError`); only SELECT returns Arrow result sets, DDL/DML return JSON, and all four pandas/arrow shortcut fetch methods refuse to work.
- 窄场景
用 Python connector 跑 DDL/DML 后想拿 DataFrame 的脚本;把 connector 当通用执行层用的工具。
Scripts using the Python connector for DDL/DML that then want a DataFrame; tools treating the connector as a generic execution layer.
- 机制
Snowflake 维护者承认这是已知取舍:为保"这四个函数永远比取 JSON 快"的保证,只愿意在用户显式 opt-in(如 `slow_ok=True`)时才做 JSON→DataFrame 转换;Snowflake 自家新 universal driver 的行为差异文档把"cursor 级 pandas/arrow fetch 方法拒绝 JSON 结果集"列为**破坏性行为差异**(`is_breaking_change: true`),且 Snowpark 的 `DataFrame.to_pandas()` 内部恰好依赖这个 fallback 分支——这个坑连 Snowflake 自己的上层工具都踩。
A known tradeoff admitted by Snowflake maintainers: keeping "these 4 fetch functions are always guaranteed to be faster than fetching Json results" means JSON→DataFrame conversion only happens with explicit opt-in (e.g. `slow_ok=True`); Snowflake's own new universal driver's behavior-diff doc lists "cursor-level pandas/arrow fetch methods rejecting JSON result sets" as a **breaking change** (`is_breaking_change: true`) — and Snowpark's `DataFrame.to_pandas()` internally depends on exactly that fallback branch. Snowflake's own upper-layer tooling steps in this hole too.
- 生产验证
来源 45:snowflake-connector-python 官方仓库用户 issue #1202(2022,复现堆栈完整);
Snowflake 自家 `snowflakedb/drivers` 仓库 SNOW-4072353 记录(官方文档,作机制印证)。
—
- 证据等级
`多方印证(2 个独立来源)`,真实用户 issue + Snowflake 自家新 driver 文档承认该行为差异。
`Corroborated (2 independent sources)`, real user issue + Snowflake's own new-driver doc acknowledging the behavior difference.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2022
dataroots 实测:Snowpark Python 做 ML,"报错没有堆栈、包只能用 Anaconda 子集、HuggingFace 模型直接 OOM"
多方印证
来源存疑
生态与信任
- 一句话
只用 Snowflake 做完整 ML 方案的第一手实测:UDF/存储过程只能用 Snowflake 提供的 Anaconda 包子集("not always the latest"),pip 独有包得打 wheel 走 imports 上传——"feels like a hack rather than a feature";"Error handling is maybe the biggest of the 'issues' I have"——UDF 在 Snowflake 侧调用失败时 "I was not able to find tracebacks in the Snowflake IDE";HuggingFace transformer 权重不能直接下载、加载直接 OOM,"even tried increasing the warehouse size with no success",且 warehouse 内存/CPU 规格不透明,"hard to know what will work and what won't",也没有 GPU;UDF 限 1 分钟、存储过程限 1 小时且只返回标量。
A first-hand attempt at "full ML on Snowflake only": UDFs/stored procedures can use only Snowflake's Anaconda package subset ("not always the latest") — pip-only packages must be wheeled up via imports, which "feels like a hack rather than a feature"; "Error handling is maybe the biggest of the 'issues' I have" — when a UDF failed server-side, "I was not able to find tracebacks in the Snowflake IDE"; HuggingFace transformer weights couldn't be downloaded directly, loading OOM'd, "even tried increasing the warehouse size with no success"; warehouse memory/CPU specs are opaque — "hard to know what will work and what won't" — and there's no GPU; UDFs cap at 1 minute, procedures at 1 hour and return scalars only.
- 窄场景
想在 Snowflake 里做 ML 训练/推理的团队;被"Snowpark ML"宣传吸引的选型。
Teams wanting to train/run ML inside Snowflake; anyone attracted by "Snowpark ML" marketing.
- 机制
Snowpark 执行环境是受限沙箱:包白名单、执行时长上限、无 GPU、硬件黑盒;报错链路为 SQL 执行环境设计,Python 调试体验差。
Snowpark's execution environment is a restricted sandbox: package allowlist, execution time caps, no GPU, opaque hardware; the error path was designed for a SQL execution environment, so Python debugging is poor.
- 生产验证
来源 47:dataroots 工程师 dev.to 实测复盘 2022(Snowpark Python public preview 时期);
来源 49:Empower 数据科学家横评([来源存疑],约 2022)——在 UDF/存储过程限制、包子集、无 GPU、内存黑盒四点上交叉印证,另补:单节点训练只有 UDF/存储过程两条路、真正的分布式训练在 Snowflake 执行环境内 "impossible"。
—
- 证据等级
`多方印证(2 个独立来源,含 1 个 [来源存疑])`,第三方咨询公司实测 + 从业者横评,四点交叉印证。
`Corroborated (2 independent sources, including 1 [source questionable])`, third-party consultancy hands-on + practitioner bake-off, four points cross-confirmed.
- 备注
2022 年 public preview 时期的实测,Snowpark 此后持续迭代,但"包白名单、无 GPU、执行时长上限"的架构约束是平台设计使然。本卡主题可能与本站 [避坑] 卡重叠。
measured in the 2022 public-preview era; Snowpark has iterated since, but the allowlist/no-GPU/time-cap constraints are architectural. This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2022
半同步复制的 ACK 等待:gh-ost 迁移丢数据、管理命令堆积打满连接数
多方印证
稳定与故障运维复杂度
- 一句话
半同步复制下主库提交要等从库 ACK;ACK 迟迟不来时,等待中的事务对别的线程"不可见",能让 gh-ost 在线 DDL 丢数据,也能让 FLUSH BINARY LOGS / SHOW MASTER STATUS 堆积到打满 max_connections。
—
- 窄场景
开启半同步复制(rpl_semi_sync_master_enabled=1)、从库延迟或宕机、timeout 设得很大的主库;用 gh-ost 做在线 DDL 的团队。
Primaries with semi-sync enabled (rpl_semi_sync_master_enabled=1), a lagging or dead replica, and a generously large timeout; teams running online DDL with gh-ost.
- 机制
AFTER_SYNC 模式下事务先等 ACK 才返回提交;等待期间该事务已写入 binlog 并发往从库,但对新起的读事务不可见——gh-ost 读到的 max id 偏小,该行既没被 DML 监听捕获、也没被行拷贝复制,直接丢失。另一方面,FLUSH BINARY LOGS / SHOW MASTER STATUS 会被"Waiting for semi-sync ACK" 的会话间接阻塞;kill 掉的只是客户端,会话不释放,堆积可打满 max_connections,届时连"关闭半同步"这条救命命令都执行不了。
In AFTER_SYNC mode a transaction waits for the ACK before commit returns; during the wait it is already in the binlog and shipped to the replica, but invisible to newly started reads — gh-ost read a stale max id, so the row was neither captured by the DML listener nor copied by the row-copy phase: lost. Separately, FLUSH BINARY LOGS / SHOW MASTER STATUS get indirectly blocked behind a session stuck in "Waiting for semi-sync ACK"; killing them only kills the client, not the session, so the pile-up can reach max_connections — at which point even the life-saving "disable semi-sync" command cannot run.
- 生产验证
github/gh-ost#1039(2021-10,MySQL 5.7.26,AFTER_SYNC,timeout=50000ms;id=3297 的行在迁移中丢失,附 gh-ost debug 日志与 processlist "Waiting for semi-sync ACK from slave" 实录);bugs.mysql.com#104012(Jean-François Gagné 提交,8.0.25/5.7.34 可复现:FLUSH BINARY LOGS 与 SHOW MASTER STATUS 被阻塞,kill 不释放会话,S2 级别,唯一 workaround 是禁用半同步)。
github/gh-ost#1039 (Oct 2021, MySQL 5.7.26, AFTER_SYNC, timeout=50000ms; the row with id=3297 went missing mid-migration, with gh-ost debug logs and a processlist showing "Waiting for semi-sync ACK from slave"); bugs.mysql.com#104012 (filed by Jean-François Gagné, reproducible on 8.0.25/5.7.34: FLUSH BINARY LOGS and SHOW MASTER STATUS block, kill does not free the session, rated S2, only workaround is disabling semi-sync).
- 证据等级
`多方印证(2 个独立来源)`,来源性质:GitHub 生产 issue 复盘 + MySQL bug 库具名提交。
—
MySQL 年份:2021
Receipt Bank:迁完一年才敢说——Snowflake 比 Redshift 贵得多
多方印证
升级迁移
- 一句话
"It's important to say we pay much more for Snowflake than what we have been paying for Redshift, but it works well for our use-case."(必须说:我们为 Snowflake 付的钱比 Redshift 多得多,但它确实适合我们的场景。)——迁移前 "back-of-the-napkin calculations" 测算的降本,4 个月后才验证成立;"迁入 Snowflake 后账单是否真的下降"本身就是迁移中的不确定项。
"It's important to say we pay much more for Snowflake than what we have been paying for Redshift, but it works well for our use-case." The pre-migration back-of-the-napkin savings math took 4 months to validate — "whether the bill actually drops after moving to Snowflake" is itself an uncertainty in the migration.
- 窄场景
从 Redshift(预留实例/固定成本模式)迁入 Snowflake(弹性计费)的团队;按"弹性=省钱"做立项测算的迁移。
Teams moving from Redshift (reserved/fixed-cost) to Snowflake (elastic billing); migrations justified on "elastic = cheaper" math.
- 机制
Redshift DDL 在 Snowflake 下直接不兼容,团队自写爬虫式脚本枚举全部 schema/table/view 重建;全球多时区用户"缺一小时数据就抓瞎",并行 dump 速度与保留算力跑 SELECT 之间要找平衡——迁移工程本身就有成本,账单对比要在有监控(Looker 仪表盘看平均查询耗时和"快查询"占比)的前提下依然可能超预期。
Redshift DDL is directly incompatible with Snowflake, so the team hand-wrote crawler scripts to enumerate every schema/table/view; global multi-timezone users meant "every hour without this data makes them blind," forcing a tradeoff between parallel dump speed and keeping enough compute for SELECTs — the migration itself costs, and the bill comparison can surprise even with monitoring (Looker dashboards on avg query time and "fast query" share) in place.
- 生产验证
来源 41:Receipt Bank 数据工程师 Yordan Ivanov 约 2020-02 第一人称复盘;
来源 40:Robin 同样提到迁移前测算降本、4 个月后才验证成立。
—
- 证据等级
`多方印证(2 个独立来源)`,两家 Redshift→Snowflake 迁移复盘在"账单不确定性"上互相印证。
`Corroborated (2 independent sources)`, two Redshift→Snowflake postmortems corroborating the "bill uncertainty" phenomenon.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2020
XID 回绕保护停机:写库被"正确"地拒绝写入
多方印证
稳定与故障运维复杂度
- 一句话
事务号耗尽前,PostgreSQL 宁可停写也不让你继续——autovacuum 没跟上 freeze,就等于计划外停机。
Before transaction IDs run out, PostgreSQL would rather stop accepting writes than keep going — if autovacuum falls behind on freezing, that means an unplanned outage.
- 窄场景
高写入 OLTP、存在超大表、autovacuum 被调得保守或被长事务/复制槽挡住 freeze 的集群;任何大版本(机制为架构性设计)。
Write-heavy OLTP clusters with very large tables where autovacuum is tuned conservatively or freezing is blocked by long transactions/replication slots; any major version (the mechanism is architectural).
- 机制
PG 用 32 位 XID(约 21 亿可用),MVCC 靠 XID 判断行可见性;vacuum 必须定期 freeze 旧元组以推进 frozen xid。freeze 跟不上 → datfrozenxid 老化 → 剩余不足约 100 万时触发保护模式拒绝一切写入,只能单用户模式 vacuum 或删表自救。Sentry 还发现当时 vacuum 的 maintenance_work_mem 有 1GB 硬上限,加内存也救不了。
Postgres uses 32-bit XIDs (~2.1B usable); MVCC visibility depends on them, so vacuum must periodically freeze old tuples to advance the frozen XID. Fall behind → datfrozenxid ages → with roughly 1M XIDs left, protection mode refuses all writes; recovery requires single-user-mode vacuum or deleting tables. Sentry additionally discovered vacuum's maintenance_work_mem had a 1GB hard limit at the time — throwing memory at it did not help.
- 生产验证
来源 1:Sentry 2015-07-20 停机大半个美国工作日,autovacuum 几小时跑完仍无效果,最终 truncate 一张事件映射大表,5 分钟后恢复;
来源 2:Mailchimp/Mandrill 2019-02-04,shard4(K/V 分片中最热)触发回绕保护停机约 40 小时,全量 vacuum 预估最长 40 天,最终 truncate TB 级 Search/Url 表恢复。两家此前都"知道风险但没排上修"。
Source 1: Sentry, Jul 20 2015 — down most of the US working day; hours of autovacuum accomplished nothing; they truncated one large event-mapping table and recovered 5 minutes later;
Source 2: Mailchimp/Mandrill, Feb 4 2019 — the hottest shard (shard4 of a sharded K/V store) hit wraparound protection and stayed degraded ~40 hours; a full vacuum was estimated at up to 40 days, so they truncated TB-scale Search/Url tables. Both teams had known about the risk but never got the fix scheduled.
- 证据等级
`多方印证(2 个独立来源)`,具名生产复盘 ×2。
`Corroborated (2 independent sources)`, named production postmortems x2.
- 备注
与现有 [避坑] 卡主题部分重叠(该卡已含 XID 回绕)。另:两起复盘分别发生于 2015 与 2019 年;该故障模式源于 32 位 XID 架构约束(需定期 freeze),非已修复缺陷,故保留收录,未作"已修复"标注。
partially overlaps the existing [Pitfall] card (which already covers XID wraparound). Also: both postmortems predate 2023 (2015 and 2019); the failure mode stems from the 32-bit XID architectural constraint (periodic freezing required), not a since-fixed defect, so it is retained without a "fixed in" label.
PostgreSQL(社区版) 年份:2019
迁移工具链:导个 CLOB 表,逼得 DBA 用 Python 生成 INSERT 再粘贴
多方印证
运维复杂度升级迁移
- 一句话
一个只有几百行、带 CLOB 字段的简单表迁移,变成耗时一个月的工程:SQL Developer 导出随机截断 CLOB、Data Pump 因小版本差异拒绝工作、SQLLDR 报 ORA-42321,最后靠 Python 生成 INSERT 语句、粘贴进 GUI 分批执行——200 小时的 Oracle 顾问账单,客户多花约 100 万美元。
Migrating a simple few-hundred-row table with a CLOB column turned into a month-long project: SQL Developer exports randomly truncated CLOBs, Data Pump refused to work over a minor version difference, SQL*Loader threw ORA-42321 — the final fix was generating INSERT statements with Python and pasting them into the GUI in chunks. 200 billable Oracle-consultant hours; the customer paid about $1M extra.
- 窄场景
跨实例/跨版本的数据迁移;表含 CLOB/BLOB 大对象字段时。
Cross-instance/cross-version data migrations; tables with CLOB/BLOB columns.
- 机制
Oracle 有十几种数据导入导出路径(exp/imp、Data Pump、SQL*Loader、SQL Developer 导出等),各工具行为"微妙而不兼容":SQL Developer 导出 CLOB 随机截断、部分格式导出时不转义导致无法解析;Data Pump 对源/目标小版本差异零容忍;SQL*Loader 报索引/约束错,而同样的 INSERT 粘贴进 GUI 却能成功——工具链把简单任务变成了需要"Oracle 黑魔法"知识的专家活。
Oracle offers a dozen-plus subtly incompatible data import/export paths (exp/imp, Data Pump, SQL*Loader, SQL Developer export...): SQL Developer exports truncate CLOBs at random and skip escaping in some formats so output cannot be parsed; Data Pump is zero-tolerant of source/target minor-version differences; SQL*Loader threw index/constraint errors while the same INSERTs pasted into the GUI succeeded — the toolchain turns simple tasks into expert-only work requiring "Oracle lore".
- 生产验证
来源 12:HN 讨论串(2017)——原帖作者完整记录了上述全过程(200 billable hours 的 Oracle 顾问、$400/小时费率、客户多支出约 100 万美元);跟帖中至少两位独立从业者呼应:"I wanted to move schema and data... It took 6 days to move it"(有 3 名全职 Oracle DBA 的客户,迁个 schema 花了 6 天);"My experience is that the tools and the 'experts' leave a lot to be desired"。
Source 12: HN thread (2017) — the original poster documented the full saga (200 billable hours of Oracle consultants, $400/hr rates, ~$1M extra customer cost); at least two independent practitioners echoed in-thread: "I wanted to move schema and data... It took 6 days to move it" (a customer with 3 full-time Oracle DBAs, moving one schema took 6 days); "My experience is that the tools and the 'experts' leave a lot to be desired."
- 证据等级
`多方印证(同一讨论串内 3 个独立声音)`。
`Multi-source corroboration (3 independent voices within one thread)`.
- 备注
2017 年旧帖,涉及的工具为 10g/11g 时代版本;按时效规则降级保留——抱怨的是"简单任务需要专家级内部知识"的结构性问题,而非某个已修复的缺陷。英文原帖含发泄性措辞,本卡仅转述技术事实。
2017 thread; tools discussed are the 10g/11g-era versions. Retained under a downgrade per the recency rule — the complaint is structural ("trivial tasks require expert-level internals knowledge"), not a since-fixed defect. The original thread contains venting language; this card reports only the technical facts.
Oracle Database(甲骨文) 年份:2017
Reader endpoint:连接池一建,负载均衡就"钉死"在一台 reader 上
多方印证
性能问题运维复杂度
- 一句话
Reader endpoint 的"负载均衡"只发生在 DNS 解析那一刻——连接池建连一次,之后所有查询都钉在同一台 reader 上,加再多 reader 也分不到流量。
The Reader endpoint's "load balancing" happens only at DNS resolution time — a pool connects once and every subsequent query sticks to the same reader; adding more readers sends them no traffic.
- 窄场景
PgBouncer/应用连接池 + 多 reader 的 Aurora PG 集群;读放大的报表/OLTP 混合负载。
PgBouncer/app connection pools against multi-reader Aurora PG clusters; mixed OLTP/reporting read-heavy workloads.
- 机制
Reader endpoint 是 DNS 轮询的连接级分发,不是查询级负载均衡;长连接一旦建立就固定在某台 reader;想扩 reader 分流,必须让池子重建连接或应用重连;JVM 等运行时的 DNS 缓存还会让"重建"本身不生效。
The Reader endpoint is DNS round-robin distribution at connection level, not query-level load balancing; a long-lived connection is pinned to whichever reader it landed on; spreading load across readers requires the pool to cycle connections or the app to reconnect; JVM-style DNS caching can additionally defeat the "reconnect" itself.
- 生产验证
来源 12:独立开源项目 pgbouncer-aurora-operator 的 README——作者为解决"Reader endpoint 无法分散连接池、长连接钉死导致读负载 skew"专门写了一个 operator:逐实例 1:1 配 PgBouncer、用 K8s Service 做成员管理;
来源 2:The Build 独立分析——"这是所有 DNS 负载均衡的通病:长连接会粘在第一次分到的 reader 上;要分流就得让连接池轮换连接或应用重连"。
Source 12: independent open-source project pgbouncer-aurora-operator's README — the author wrote an entire operator to fix "the Reader endpoint cannot spread a connection pool; long-lived connections pin and skew read load": one PgBouncer per instance (1:1) with K8s Service membership management;
Source 2: The Build's independent analysis — "the same limitation every DNS-based load balancer has: long-lived pooled connections stick to whichever reader they were first handed; spreading load requires the connection pool to cycle connections or the application to reconnect."
- 证据等级
`多方印证(2 个独立来源)`,独立开源项目文档 + 独立技术分析。
`Corroborated (2 independent sources)`, independent open-source project docs + independent technical analysis.
Amazon Aurora 年份:—
Failover"30 秒"是集群事件的时间,不是你的应用恢复的时间
多方印证
稳定与故障
- 一句话
AWS 的 30 秒是从"集群事件"口径量的——你的应用还要过 DNS 缓存、连接池重建、冷缓存三关,每一关都能把中断拉长。
AWS's 30 seconds is measured at the cluster-event level — your application still has to clear DNS caching, pool rebuild, and a cold cache, each of which can stretch the outage.
- 窄场景
直连集群 endpoint、无 RDS Proxy、无拓扑感知驱动的应用;JVM 等有 DNS 缓存的运行时;p99 敏感的业务。
Apps connecting directly to the cluster endpoint with no RDS Proxy and no topology-aware driver; runtimes with DNS caching (JVM); p99-sensitive services.
- 机制
Failover 本体(reader 晋升并绑定共享存储卷)确实通常小于 30 秒;但应用侧恢复取决于:DNS TTL(默认 5 秒,JVM/容器层层缓存可放大到分钟级)、连接池重建、被晋升 reader 的 buffer cache 是按"读负载"预热的、写负载上来后有 p99 毛刺。单实例集群(无 reader)则要现起计算节点,按分钟计。
The failover itself (a reader promoted and bound to the shared storage volume) is indeed typically under 30 seconds; but application-side recovery depends on DNS TTL (5 s default, amplified to minutes by JVM/container-layer caching), connection-pool rebuild, and the promoted reader's buffer cache being warmed for its read workload — not the writer's — producing p99 spikes. Single-instance clusters (no reader) must launch a fresh compute node first: minutes.
- 生产验证
来源 12:pgbouncer-aurora-operator README——"集群 endpoint 强烈受 DNS 缓存/TTL 影响,failover 后即使 endpoint 已指向新 writer,陈旧的解析结果仍会被继续使用";
来源 2:The Build——failover 是"新 writer 快速绑定存储卷",但"不是缓存热的";p99/p999 敏感的应用会在首次生产 failover 时感受到延迟毛刺;无 reader 的集群 failover 按分钟计;
来源 13:Harshith——"单实例 Aurora 集群得不到 Aurora 的 failover 故事;没有 reader 时要先起一台新计算实例,以分钟计"。
Source 12: pgbouncer-aurora-operator README — "the cluster endpoint is strongly affected by DNS cache, TTL, and refresh behavior; after failover, even if the endpoint points at the new writer, stale resolution results keep being used";
Source 2: The Build — failover is "the new writer binds to the storage volume quickly" but "not with a warm cache"; p99/p999-sensitive apps feel latency spikes on their first production failover; reader-less clusters fail over on the order of minutes;
Source 13: Harshith — "a single-instance Aurora cluster does not get you Aurora's failover story; with no reader, Aurora has to launch a fresh compute instance first, and you're into minutes."
- 证据等级
`多方印证(3 个独立来源)`,独立开源项目 + 独立技术分析 ×2(AWS 官方手册亦承认"集群事件 30 秒、应用中断更久"的现象,作为机制佐证)。
`Corroborated (3 independent sources)`, independent open-source project + independent technical analyses x2 (AWS's own handbook also acknowledges the "30-second cluster event, longer app outage" phenomenon, used here as mechanism reference).
Amazon Aurora 年份:—
没有真 superuser:扩展白名单 + 参数子集 + 摸不到文件系统
多方印证
生态与信任
- 一句话
`rds_superuser` 不是 superuser——装不了白名单外的扩展(含自定义 C 扩展),调不了的参数只能认,`COPY ... FROM PROGRAM` 这类操作想都别想。
`rds_superuser` is not superuser — no installing extensions outside the allow-list (custom C extensions included), no tuning the unavailable knobs, no `COPY ... FROM PROGRAM`.
- 窄场景
依赖特定扩展(PostGIS、TimescaleDB、完整版 pg_cron、新版 pgvector、自研 C 扩展)或深度调参的工作负载。
Workloads depending on specific extensions (PostGIS, TimescaleDB, full pg_cron, recent pgvector, home-grown C extensions) or deep parameter tuning.
- 机制
Aurora 沿用 RDS 的权限模型:最高只有 rds_superuser;扩展来自 AWS 维护的 allow-list(与 RDS 的清单还不完全一致,且随小版本变化);可调的 postgresql.conf 参数是子集,静态参数改完要重启;没有文件系统访问。
Aurora inherits the RDS permission model: the ceiling is rds_superuser; extensions come from an AWS-maintained allow-list (not identical to RDS's, and shifting with minor versions); the tunable postgresql.conf subset is smaller than stock, static parameters need a reboot; no filesystem access.
- 生产验证
来源 2:The Build——"没有真正的 SUPERUSER";扩展 allow-list 短于社区生态,TimescaleDB/Citus 这类直接出局;pg_tle 只允许过程语言表达的扩展,C 扩展不在 allow-list 就没门;
来源 13:Harshith——"你交出了 superuser:拿到的是 rds_superuser、精选扩展清单、没有文件系统;清单外的扩展就是不可用"。
Source 2: The Build — "no true SUPERUSER"; the extension allow-list is shorter than the community universe, ruling out the likes of TimescaleDB/Citus; pg_tle only permits extensions expressible in a procedural language — C extensions outside the allow-list are simply unavailable;
Source 13: Harshith — "You give up superuser. You get rds_superuser, a curated extension list, and no filesystem access. Extensions outside the supported set are simply unavailable."
- 证据等级
`多方印证(2 个独立来源)`,独立技术分析 ×2。
`Corroborated (2 independent sources)`, independent technical analyses x2.
Amazon Aurora 年份:—
Global Database 写转发:备区域写一次,跨洋跑一趟
多方印证
性能问题
- 一句话
Write forwarding 让备区域"看起来"能写——但每一笔写都要先转发到主区域提交,备区域的写延迟是主区域的十几倍;这叫 active-passive,不叫 active-active。
Write forwarding makes secondaries look writable — but every write is forwarded to the primary region before commit, so secondary-region write latency is an order of magnitude worse; that's active-passive, not active-active.
- 窄场景
多区域部署、备区域有就近写入需求的应用;开了 write forwarding 却没读一致性模式文档的团队。
Multi-region deployments with near-write needs in secondary regions; teams enabling write forwarding without reading the consistency-mode docs.
- 机制
Global Database 只有一个写区域;备区域 endpoint 接受写后经存储层转发到主区域 writer,跨区域 RTT 直接加在每次写上;区域间复制延迟按"秒、通常小于 1 秒"计;一致性模式(eventual/session/global)在延迟与可见性之间做交易,不读文档直接开会有"惊喜"。
Global Database has exactly one writer region; a secondary endpoint accepts a write then forwards it through the storage layer to the primary writer — the cross-region RTT lands on every write; inter-region replication lag is measured in "seconds, usually under one"; the consistency modes (eventual/session/global) trade latency against visibility, and enabling forwarding without reading them produces "surprises."
- 生产验证
来源 2:The Build——"备区域秒级延迟对很多应用够用,对另一些不够";write forwarding"有一致性模式的微妙语义,不读文档就开的应用会遇到惊喜";
来源 17:独立工程师的本地 PoC(swa-roopa/event_ticketing_platform)——演示模型下备区域转发写约 80–120ms、主区域本地写约 5ms,结论"这是 active-passive,不是 active-active"。
Source 2: The Build — "seconds, usually under one" of lag is fine for many apps and not for others; write forwarding has "consistency-mode subtleties" that become "the source of production bugs in applications that do not read the documentation carefully";
Source 17: an independent engineer's local PoC (swa-roopa/event_ticketing_platform) — in the demo model, forwarded secondary writes measured ~80-120 ms vs ~5 ms local primary writes; conclusion: "this is active-passive, not active-active."
- 证据等级
`多方印证(2 个独立来源)`,独立技术分析 + 独立工程师的概念验证。
`Corroborated (2 independent sources)`, independent technical analysis + independent engineer's proof of concept.
- 备注
来源 17 为本地 PoC(MySQL 模拟 Aurora、dynamodb-local 模拟 Global Tables),80–120ms 为演示模型值、非生产实测,仅用于说明行为差异,不采信为精确数字。
Source 17 is a local PoC (MySQL standing in for Aurora, dynamodb-local for Global Tables); the 80-120 ms figures are demo-model values, not production measurements — used only to illustrate the behavioral difference, not as precise numbers.
Amazon Aurora 年份:—
换了 Aurora,vacuum 还是你的
多方印证
运维复杂度
- 一句话
很多人以为上了"云原生"就不用管 vacuum 了——但 Aurora 的 vacuum 跑在 writer 进程里,受 writer 的 CPU/内存限制;writer 配小了,膨胀和 XID 回绕一个都少不了。
Many assume "cloud-native" means never thinking about vacuum again — but Aurora's vacuum runs inside the writer process, bounded by the writer's CPU and memory; undersize the writer and bloat plus XID wraparound are all still on the menu.
- 窄场景
高频 UPDATE/DELETE 的 OLTP 跑在偏小 writer 上;"托管=不用管 vacuum"心智的团队。
High-churn UPDATE/DELETE OLTP on an undersized writer; teams with a "managed means no vacuum duty" mindset.
- 机制
Aurora 的存储层只管 I/O,找死元组、清索引项、更新可见性图这些活还是 writer 进程干;vacuum 重的负载下瓶颈是 writer 的 CPU/内存,不是存储吞吐;MVCC/XID 回绕的架构约束原样继承。
Aurora's storage layer only handles I/O; finding dead tuples, clearing index entries, and updating visibility maps still happens in the writer process; under vacuum-heavy load the bottleneck is writer CPU/memory, not storage throughput; the MVCC/XID-wraparound architectural constraints are inherited unchanged.
- 生产验证
来源 2:The Build——"分布式存储不是 vacuum 溶剂,它只是存储";writer 配小会在重更新负载下复现社区 PG 的膨胀问题;
来源 13:Harshith——"Vacuum 还是完全你的问题……如果你没想过 autovacuum_vacuum_cost_limit 和 oldest XID,你的 outage 已经排期了,只是日期未知"。
Source 2: The Build — "distributed storage is not a vacuum solvent; it is storage"; an undersized writer reproduces community Postgres bloat problems under heavy update load;
Source 13: Harshith — "Vacuum is still entirely your problem… If you're running Aurora and you haven't thought about autovacuum_vacuum_cost_limit and your oldest XID, you have an outage scheduled for a date you don't know yet."
- 证据等级
`多方印证(2 个独立来源)`,独立技术分析 ×2。
`Corroborated (2 independent sources)`, independent technical analyses x2.
- 备注
主题可能与现有 [避坑] 卡及 PG「用户吐槽」autovacuum 卡重叠。
this card's topic may overlap existing [Pitfall] cards on this site and the Postgres "User Rants" autovacuum card.
Amazon Aurora 年份:—
Too Many Parts 墙:小批量写入能把 MergeTree 写死
多方印证
来源存疑
性能问题运维复杂度
- 一句话
每次 INSERT 至少生成一个不可变 part,小批量高频写入让 part 数超过后台合并能力,最终 ClickHouse 先限流、再直接拒绝写入。
Every INSERT creates at least one immutable part; small, frequent batches let part counts outrun background merges until ClickHouse first throttles writes, then refuses them outright.
- 窄场景
Kafka/队列直连的流式写入、单行/小批次 INSERT、物化视图目标表(每个 INSERT 在每张 MV 表上都至少多一个 part)、高基数分区键把合并队列拆散的表。
Streaming writes wired straight from Kafka/queues, single-row or small-batch INSERTs, materialized-view target tables (each INSERT adds at least one part per MV table), and high-cardinality partition keys that fragment the merge queues.
- 机制
MergeTree 把写入先落成不可变 part,后台 merge 线程再把小 part 折叠成大 part;合并只发生在同一分区内、受 CPU/磁盘吞吐上限约束,而 part 的产生速度只受客户端并发限制。part 堆积触发两道护栏:先 `parts_to_delay_insert` 人为拖慢 INSERT,再超 `parts_to_throw_insert` 直接抛 `Too many parts` 拒绝写入。`async_insert`、Buffer 表只是把"攒批"搬到服务端,并没有改变 MergeTree 批式合并的本质;`wait_for_async_insert=0`(fire-and-forget)下,进程崩溃会直接丢掉缓冲中尚未落盘的行。
MergeTree lands writes as immutable parts; background merge threads fold small parts into bigger ones. Merging happens only within one partition and is bounded by CPU/disk throughput, while part creation is bounded only by client concurrency. Accumulation trips two guardrails: `parts_to_delay_insert` artificially slows INSERTs first, then `parts_to_throw_insert` throws `Too many parts` and rejects writes. `async_insert` and Buffer tables only move batching server-side — they do not change MergeTree's batch-oriented nature; with `wait_for_async_insert=0` (fire-and-forget), a process crash loses whatever rows were still buffered.
- 生产验证
来源 1:OpenPanel 创始人总结——"ClickHouse loves big inserts. It hates small ones…Do that one row at a time and you'll drown it",他们被迫在应用层用队列先攒批;
来源 2:Tinybird 指出读性能"无论 schema 多好"都会被 part 数拖累("If you are inserting data often, you'll get penalized while reading no matter what schema your table has"),回填时必须"increase the rows you push per block to avoid generating thousands of parts";
来源 6:长文拆解了阈值演进史与高基数分区键如何把合并队列拆散([来源存疑],见上);
来源 8:Altinity 记录"见过团队因为不懂 `wait_for_async_insert` 的含义丢了数百万行",且 async insert 默认关闭去重。
Source 1: OpenPanel's founder — "ClickHouse loves big inserts. It hates small ones…Do that one row at a time and you'll drown it"; they had to add application-side queue batching;
Source 2: Tinybird — read performance is dragged down by part counts "no matter what schema your table has"; backfills must "increase the rows you push per block to avoid generating thousands of parts";
Source 6: long-form walkthrough of the threshold history and how high-cardinality partition keys fragment merge queues ([Questionable source], see above);
Source 8: Altinity reports teams that "lost millions of records" without understanding `wait_for_async_insert`, and notes async inserts disable deduplication by default.
- 证据等级
`多方印证(3 个独立来源,含 1 个 [来源存疑])`,具名客户 ×2(OpenPanel 创始人、Tinybird)+ 机制长文 ×1。
`Corroborated (3 independent sources, including 1 [Questionable source])`, named customers x2 (OpenPanel founder, Tinybird) + mechanism deep-dive x1.
- 备注
与现有深水区一(part 机制)与吐槽清单"高频小 INSERT"行主题重叠,互为佐证。
overlaps the existing Deep-Dive 1 (part mechanics) and the rant-list row on high-frequency small INSERTs — kept as corroboration.
ClickHouse 年份:—
ZooKeeper 运维税:副本一致性外包给外部共识服务
多方印证
稳定与故障运维复杂度
- 一句话
ClickHouse 自身不实现共识算法,复制全靠 ZooKeeper/ClickHouse Keeper——ZK 一抖,所有副本表降级只读,写入直接被拒。
ClickHouse implements no consensus algorithm of its own — replication depends entirely on ZooKeeper/ClickHouse Keeper, and when ZK wobbles, every replicated table degrades to read-only and writes are refused.
- 窄场景
自建 ReplicatedMergeTree 集群;ZK 与数据节点混部、ZK 节点内存不足或磁盘打满的部署。
Self-hosted ReplicatedMergeTree clusters; deployments where ZK shares machines with data nodes or ZK nodes run out of memory/disk.
- 机制
ZK 存 part 元数据、副本选举、mutation 队列、分布式 DDL 队列;会话断开后副本表为保一致性降级为只读。part 总数膨胀会直接放大 ZK 的元数据压力(每个 part 都是一批 znode)。
ZK holds part metadata, replica elections, the mutation queue, and the distributed-DDL queue; when sessions drop, replicas go read-only to protect consistency. Part-count growth directly multiplies ZK metadata pressure (every part is a batch of znodes).
- 生产验证
来源 4:TechWolf 自建 K8s 集群初期用 ZooKeeper,"很快遇到节点 OOM、写入失败",被迫迁到 ClickHouse Keeper(C++ 重写,解决 ZK 扩展性上限);
来源 3:Tinybird CTO 明确要求至少 3 个 ZK 副本、且必须与数据节点物理隔离——"ZK 和数据库同机,机器一过载就全慢、最终全失败,表变成只读",这是纯粹的额外硬件与运维成本;
来源 5:Cloudflare 在 part 数涨到 16 万/副本后,ZK 同样出问题,原文预告"也许哪天讲讲我们的 100GB ZooKeeper 集群的故事"。
Source 4: TechWolf's self-hosted K8s cluster initially used ZooKeeper and "quickly ran into stability issues with nodes going out-of-memory and inserts failing," forcing a move to ClickHouse Keeper (the C++ rewrite addressing ZK's scalability limits);
Source 3: Tinybird's CTO requires at least 3 ZK replicas physically isolated from data nodes — co-locating means "everything will be slow and eventually fail, leaving your tables in read-only" — pure extra hardware and ops cost;
Source 5: Cloudflare hit ZK problems too once parts reached 160k/replica, teasing "perhaps one day we'll tell the story of the 100 gigabyte ZooKeeper cluster."
- 证据等级
`多方印证(3 个独立来源)`,具名客户 ×3(TechWolf、Tinybird、Cloudflare)。
`Corroborated (3 independent sources)`, named customers x3 (TechWolf, Tinybird, Cloudflare).
- 备注
与现有深水区三(复制与一致性)部分重叠,但"ZK 运维税"角度在深水区中未展开。
partially overlaps Deep-Dive 3 (replication and consistency), but the "ZK operations tax" angle is not developed there.
ClickHouse 年份:—
Mutation:ALTER UPDATE 是整 part 重写的异步作业
多方印证
性能问题
- 一句话
`ALTER TABLE ... UPDATE/DELETE` 不原地改行,而是把受影响的 part 整个重写一遍——改几行逻辑行,代价按 part 大小算。
`ALTER TABLE ... UPDATE/DELETE` never edits rows in place — it rewrites every affected part in full, so the cost follows part size, not the number of logical rows changed.
- 窄场景
把 ClickHouse 当 OLTP 用、需要频繁单行/小批量 UPDATE/DELETE 的业务;大表上跑大范围 mutation 的运维窗口。
Treating ClickHouse as an OLTP store with frequent single-row/small-batch UPDATE/DELETE; running wide-range mutations on large tables inside an ops window.
- 机制
mutation 是异步后台任务:为每个受影响 part 生成重写后的新 part,旧 part 标记待删;写放大与受影响 part 的总大小成正比,与逻辑修改行数无关;默认异步、提交后不可普通回滚(只能 `KILL MUTATION`),查询可能同时读到新旧 part;大 mutation 与后台 merge 抢 IO/CPU,拖慢整表。轻量 delete/update 按版本逐步改善,但"重写 part"的本质未变。
A mutation is an async background job: it produces rewritten copies of each affected part and marks the old ones for deletion; write amplification scales with total affected part size, not logical rows touched. Mutations are async by default, not ordinarily rollbackable (only `KILL MUTATION`), and queries can see old and new parts simultaneously; big mutations fight background merges for IO/CPU and slow the whole table. Lightweight deletes/updates improve this version by version, but the "rewrite the part" essence is unchanged.
- 生产验证
来源 1:OpenPanel 创始人——"ClickHouse was never designed for frequent updates or deletes. It's a write-once, append-forever kind of database…painful if you expect relational behavior. I still run into these kind of issues today",并表示指望 beta 中的 lightweight updates 救命;
来源 2:Tinybird 把"Mutations"列进必须监控的卡死清单——"Things get stuck sometimes, and killing the queries does not always work"。
Source 1: OpenPanel's founder — "ClickHouse was never designed for frequent updates or deletes. It's a write-once, append-forever kind of database…painful if you expect relational behavior. I still run into these kind of issues today," pinning hopes on lightweight updates then in beta;
Source 2: Tinybird lists "Mutations" among the things that get stuck and must be monitored — "things get stuck sometimes, and killing the queries does not always work."
- 证据等级
`多方印证(2 个独立来源)`,具名客户 ×2。
`Corroborated (2 independent sources)`, named customers x2.
- 备注
与现有深水区二(Mutation 重写税)与吐槽清单"ALTER UPDATE/DELETE"行主题重叠,互为佐证。
overlaps Deep-Dive 2 (mutation rewrite tax) and the rant-list row on `ALTER UPDATE/DELETE` — kept as corroboration.
ClickHouse 年份:—
没有优化器,你就是优化器:JOIN 能把内存吃光
多方印证
性能问题
- 一句话
ClickHouse 没有 Postgres/MySQL 那样的全功能查询优化器,大表 JOIN 默认 hash join 会把右表整个装进内存——不是变慢,是直接 OOM 被杀。
ClickHouse has no full-featured query optimizer like Postgres/MySQL; a big-table JOIN under the default hash join loads the entire right table into memory — not slower, but OOM-killed.
- 窄场景
BI/临时分析里的大表 JOIN、宽结果返回、高基数聚合状态列(如 `uniqExactState`)、`index_granularity` 设得过小的表。
Big-table JOINs in BI/ad-hoc analysis, wide result sets, high-cardinality aggregation states (e.g. `uniqExactState`), tables with `index_granularity` set too low.
- 机制
默认 hash join 先把右表完整建成内存 hash table 再扫左表;大右表直接触发 `MEMORY_LIMIT_EXCEEDED` 杀查询。排序键是物理布局合同,选错列查询退化 10-100 倍且无法低成本重排。聚合状态列(如高基数 `uniqExactState`)在合并时能把整台机器拖死;`index_granularity` 低于 128 同样可能拖垮服务器。`max_memory` 等参数是"熔断器"不是"优化器"。
Default hash join builds the complete right table as an in-memory hash table before scanning the left; a large right table trips `MEMORY_LIMIT_EXCEEDED` and the query is killed. The sort key is a physical-layout contract — pick the wrong columns and queries degrade 10–100x with no cheap way to re-sort later. Aggregation-state columns (high-cardinality `uniqExactState`) can take down a whole machine during merges; `index_granularity` below 128 can cripple a server too. `max_memory` and friends are circuit breakers, not an optimizer.
- 生产验证
来源 2:Tinybird——"ClickHouse does not have an optimizer; you are the optimizer",排序键设计不好"you are losing 10-100x on your queries","even ClickHouse has a max total memory config, and it'll OOM",`uniqExactState` 高基数列"can kill your cluster on merges",`index_granularity` 别低于 128;
来源 1:OpenPanel——"It doesn't plan your joins intelligently. If you join two big tables, it'll happily try and load everything in memory and die trying…You'll notice when you have a bad join since it will take ages or die trying"。
Source 2: Tinybird — "ClickHouse does not have an optimizer; you are the optimizer"; a bad sort-key design means "you are losing 10-100x on your queries"; "even ClickHouse has a max total memory config, and it'll OOM"; a high-cardinality `uniqExactState` column "can kill your cluster on merges"; keep `index_granularity` at 128 or above;
Source 1: OpenPanel — "It doesn't plan your joins intelligently. If you join two big tables, it'll happily try and load everything in memory and die trying…You'll notice when you have a bad join since it will take ages or die trying."
- 证据等级
`多方印证(2 个独立来源)`,具名客户 ×2。
`Corroborated (2 independent sources)`, named customers x2.
- 备注
与现有深水区五(JOIN 内存爆炸)、深水区六(排序键)与吐槽清单"默认 hash join"行主题重叠,互为佐证。
overlaps Deep-Dive 5 (JOIN memory explosion), Deep-Dive 6 (sort key), and the rant-list row on default hash joins — kept as corroboration.
ClickHouse 年份:—
自建的便宜是假象:副本数 × SSD × Keeper 都是钱
多方印证
成本账单
- 一句话
单机 ClickHouse 便宜是真的,但生产要的副本数、SSD、Keeper 节点、运维人力加起来,"便宜"只存在于 benchmark 里。
Single-node ClickHouse is genuinely cheap, but production needs replicas, SSDs, Keeper nodes, and ops headcount — "cheap" only exists in benchmarks.
- 窄场景
从单机 POC 放大到生产多副本集群;用 SSD 装全部数据的团队;被 ClickHouse Cloud 账单刺痛后转自建的团队。
Scaling a single-node POC to a production multi-replica cluster; teams putting all data on SSDs; teams moving off ClickHouse Cloud after bill shock.
- 机制
本地盘架构下每个副本都要存全量数据:300TB 表 × 10 副本(1000 QPS ÷ 单副本 100 QPS)= 3000TB SSD;另加至少 3 个独立的 Keeper/ZK 节点;读写分离、冷热分层都要自己搭。自建省的是云账单,花的是工程师时间——"Once you grow, it's a full-time job"。
With local-disk architecture every replica stores the full dataset: a 300TB table x 10 replicas (1000 QPS ÷ 100 QPS per replica) = 3000TB of SSD; plus at least 3 dedicated Keeper/ZK nodes; read/write separation and hot/cold tiering are DIY. Self-hosting saves the cloud bill and spends engineer time instead — "once you grow, it's a full-time job."
- 生产验证
来源 1:OpenPanel——云 vs 自建对照表:云"Cost at scale $$$…the bill can sting hard once you start pushing data around",自建"Running ClickHouse yourself works fine for smaller setups. Once you grow, it's a full-time job";
来源 3:Tinybird 给出副本成本算式(300TB × 10 副本 = 3000TB SSD,"you have a problem"),以及读写分离、按负载类型隔离副本的额外硬件开销。
Source 1: OpenPanel — cloud vs self-host table: cloud "cost at scale $$$…the bill can sting hard once you start pushing data around"; self-hosting "works fine for smaller setups. Once you grow, it's a full-time job";
Source 3: Tinybird's replica-cost arithmetic (300TB x 10 replicas = 3000TB of SSD, "you have a problem"), plus the extra hardware for read/write separation and per-workload replica isolation.
- 证据等级
`多方印证(2 个独立来源)`,具名客户 ×2。
`Corroborated (2 independent sources)`, named customers x2.
ClickHouse 年份:—
升级比社区版更复杂:数据库、工具链、扩展三者的版本要一起对齐
多方印证
运维复杂度升级迁移
- 一句话
EDB 不是单个数据库,而是一套互相咬合的发行版+工具链——升数据库大版本时,PEM、EFM、Barman/BART、迁移工具的版本兼容矩阵也要一起算,比社区 PG 麻烦一截。
EDB is not a single database but an interlocking distribution-plus-toolchain — a major-version upgrade means solving the version-compatibility matrix for PEM, EFM, Barman/BART, and migration tools together, a notch harder than community Postgres.
- 窄场景
跑 EDB 全家桶(EPAS + PEM 监控 + EFM 故障转移 + Barman/BART 备份)的生产环境;做 PG 大版本升级(如 PG13→16)或 EDB 订阅续期换版本的团队。
Production environments running the full EDB stack (EPAS + PEM monitoring + EFM failover + Barman/BART backup); teams doing a Postgres major-version upgrade (e.g. PG13 to 16) or switching EDB versions at subscription renewal.
- 机制
社区 PG 升级只需考虑 PG 本体与扩展;EDB 发行版在 PG 之上叠了自有分支补丁和一整套闭源工具链,每件工具有自己的版本线且与服务端版本绑定。升级变成"解多元一次方程":服务端、代理、监控、备份、故障转移组件必须落在兼容矩阵内,任一错位都可能导致监控断连或备份失败。订阅制的软件源又意味着版本获取节奏受厂商发布牵制。
A community Postgres upgrade only concerns Postgres itself plus extensions; the EDB distribution layers proprietary fork patches and a whole closed-source toolchain on top, each tool with its own release line bound to server versions. Upgrading becomes solving a multi-variable equation: server, agents, monitoring, backup, and failover components must all land inside the compatibility matrix, and any mismatch can break monitoring or backups. Subscription-gated software repositories add vendor-release pacing to the mix.
- 生产验证
来源 2:Kanishka R.(中型企业,4.0/5)原话:"Complexity of Upgrades: Managing version upgrades and ensuring compatibility across EDB's ecosystem of tools sometimes feels more complex than working with the community edition."(升级复杂度:在 EDB 工具生态里管理版本升级、保证兼容性,有时比社区版更复杂);
来源 3:匿名企业已验证用户(4.0/5)原话:"Cost, Time consuming Updates"(贵,更新耗时)。
Source 2: Kanishka R. (mid-market, 4.0/5): "Complexity of Upgrades: Managing version upgrades and ensuring compatibility across EDB's ecosystem of tools sometimes feels more complex than working with the community edition.";
Source 3: anonymous verified enterprise user (4.0/5): "Cost, Time consuming Updates".
- 证据等级
`多方印证(2 个独立来源)`,G2 评价 ×2(含 1 具名)。
`Corroborated (2 independent sources)`, G2 reviews x2 (1 named).
EDB Postgres / EDB Postgres Advanced Server(EPAS) 年份:—
内存三本账:dashboard 显示 16GB,内核按 32GB 杀你
多方印证
性能问题成本账单
- 一句话
dashboard 显示 used_memory 16GB,你按 16GB 配机器,凌晨 3 点内核按 RSS 31GB 把你 OOM 杀掉——Redis 的内存有三本账,只看一本必翻车。
The dashboard says used_memory is 16GB, you provision for 16GB, and at 3am the kernel OOM-kills you based on an RSS nearing 31GB — Redis keeps three sets of memory books, and reading only one guarantees a bad surprise.
- 窄场景
value 大小参差、高 churn 的实例;按 used_memory 做容量规划的团队;Redis 与 Valkey 同理(同一分配器、同一机制)。
Instances with uneven value sizes and high churn; teams doing capacity planning off used_memory; applies equally to Redis and Valkey (same allocator, same mechanism).
- 机制
`INFO memory` 里三个数是三层意思:used_memory(分配器交给 Redis 的逻辑占用)、used_memory_rss(OS 看到的常驻集,内核只认它)。jemalloc 按 size class 取整分配(33 字节占 48 字节槽)→ 天然内部碎片;删除 key 只把槽还到 size class 的 free list,一整页没有活对象才还给 OS → RSS 只增不减;ratio < 1.0 不是"负碎片",是部分内存被换到 swap(最坏状态)。maxmemory 按 used_memory(逻辑数)执行淘汰,于是"一边拼命淘汰丢缓存命中、一边 RSS 远超 maxmemory"可以同时发生。
`INFO memory` holds three different numbers: used_memory (logical, what the allocator handed Redis), used_memory_rss (resident set, the only number the kernel cares about). jemalloc rounds up to size classes (33 bytes occupies a 48-byte slot) — inherent internal fragmentation; deleting a key only returns the slot to the size class's free list, and a page is returned to the OS only when it has no live objects — so RSS grows and never shrinks; ratio < 1.0 is not "negative fragmentation", it means part of memory went to swap (the worst state). maxmemory enforces against used_memory (the logical number), so "aggressively evicting while losing cache hits, while RSS runs far past maxmemory" can happen simultaneously.
- 生产验证
来源 1:32GB 机器,used_memory 16GB,凌晨 3 点被 OOM killer 带走,RSS 逼近 31GB,"the graph wasn't lying, it was just answering a different question";
来源 2:排查清单明确要求同时看 used_memory_rss 是否远大于 used_memory(碎片信号),以及 mem_fragmentation_ratio 指标。
Source 1: a 32GB box with used_memory at 16GB, OOM-killed at 3am with RSS approaching 31GB — "the graph wasn't lying, it was just answering a different question";
Source 2: a troubleshooting checklist that explicitly requires watching whether used_memory_rss is far larger than used_memory (the fragmentation signal), plus the mem_fragmentation_ratio metric.
- 证据等级
`多方印证(2 个独立来源)`,个人博客 ×2。
`Corroborated (2 independent sources)`, personal blogs x2.
Redis / Valkey 年份:—
持久化 fork:备份窗口就是延迟毛刺与内存翻倍窗口
多方印证
性能问题稳定与故障
- 一句话
RDB 快照和 AOF 重写都要 fork;大内存实例上,fork 复制页表阻塞主线程、写负载下 CoW 让内存短时翻倍——备份时刻就是毛刺时刻。
RDB snapshots and AOF rewrites both fork; on a large instance, fork copies the page table while blocking the main thread, and under write load copy-on-write can briefly double memory — the backup moment is the spike moment.
- 窄场景
数十 GB 数据集、写负载高的生产实例;容器/VPS 上内存配得紧的部署;开了 RDB 或 AOF 重写的实例。
Production instances with tens of GB of data and heavy write load; tight-memory container/VPS deployments; instances with RDB or AOF rewrite enabled.
- 机制
fork() 复制页表是 O(RSS) 操作,大内存实例可阻塞主线程数十至上百毫秒(`latest_fork_usec` 可观测);子进程 CoW:快照期间父进程每写一页内核就复制一页,极端下子进程持有接近完整第二份数据 → 父子 RSS 之和趋近 2 倍数据集;vm.overcommit_memory=0 时内核按悲观会计直接拒绝 fork → 后台持久化静默失败;AOF 重写收尾"追加重写缓冲区并替换文件"同样阻塞主线程。
fork() copies the page table in O(RSS), which can block the main thread for tens to hundreds of milliseconds on large instances (observable via `latest_fork_usec`); with copy-on-write, every page the parent writes during the snapshot gets copied by the kernel — in the extreme, the child holds a near-complete second copy, so parent+child RSS approaches 2x the dataset; with vm.overcommit_memory=0 the kernel's pessimistic accounting refuses the fork outright — background persistence fails silently; the AOF rewrite's final "append the rewrite buffer and swap files" step also blocks the main thread.
- 生产验证
来源 1:16GB 数据集 + 写负载 + RDB save 进行中,父子进程 RSS 之和可逼近 32GB,"the save itself is what tips you into the killer";
来源 2:fork 阻塞原理、latest_fork_usec 监控指标、AOF 重写期额外阻塞;
来源 15(Valkey 实证,机制同源):2GB droplet 跑 1.2GB Valkey 数据集,BGSAVE 恰在流量 spike 时触发,CoW balloon,OOM killer 杀掉的正是 Valkey 自己。
Source 1: a 16GB dataset under write load with RDB save in progress, parent+child RSS approaching 32GB — "the save itself is what tips you into the killer";
Source 2: the fork-blocking mechanism, the latest_fork_usec metric, and the extra blocking during AOF rewrite;
Source 15 (Valkey evidence, same mechanism): a 2GB droplet running a 1.2GB Valkey dataset, BGSAVE triggered exactly during a traffic spike, CoW ballooning — and the OOM killer killed Valkey itself.
- 证据等级
`多方印证(3 个独立来源)`,含 1 个 Valkey 侧实证。
`Corroborated (3 independent sources)`, including 1 Valkey-side case.
Redis / Valkey 年份:—
noeviction 默认:内存一满,写请求直接被拒
多方印证
性能问题
- 一句话
默认 maxmemory-policy 是 noeviction——内存一满,Redis 不删数据、不崩溃,只是礼貌地拒绝你所有的写。
The default maxmemory-policy is noeviction — when memory fills, Redis does not delete data or crash, it just politely rejects all your writes.
- 窄场景
当纯缓存用、没设 maxmemory 或没改默认策略就上生产的实例;把写失败当致命错误处理的客户端。
Instances used as a pure cache that went to production with no maxmemory or with the default policy; clients that treat write failure as fatal.
- 机制
noeviction 下达到 maxmemory 后写命令返回 OOM 错误,读不受影响;应用若把写失败当致命错误,缓存层就在最需要它的流量 spike 时变成 fail-closed 的硬依赖。另一极端 allkeys-lru:压力下优先淘汰热 key → 命中率雪崩 → DB 被压垮;volatile-* 若 key 没设 TTL 则无 key 可淘汰,等于没配。
Under noeviction, writes past maxmemory return OOM errors while reads keep working; if the application treats write failure as fatal, the cache layer becomes a fail-closed hard dependency exactly during the traffic spike when it is needed most. The other extreme, allkeys-lru, evicts hot keys first under pressure — hit-rate avalanche, then the DB gets crushed; volatile-* keys with no TTL set simply have nothing eligible for eviction, which is the same as not configuring anything.
- 生产验证
来源 3:used_memory 3.98G / maxmemory 4.00G / noeviction,结算服务所有 SET 被 "OOM command not allowed" 打回,"Redis does not crash. Redis just stops saying yes.";
来源 4:"I have seen this take down checkout on a Tuesday that was supposed to be quiet";并指出重启不解决问题,20 分钟后照样回来。
Source 3: used_memory 3.98G / maxmemory 4.00G / noeviction, every SET in the checkout path rejected with "OOM command not allowed" — "Redis does not crash. Redis just stops saying yes.";
Source 4: "I have seen this take down checkout on a Tuesday that was supposed to be quiet"; a restart does not fix it — it comes back in 20 minutes.
- 证据等级
`多方印证(2 个独立来源)`,个人博客 ×2(其一为付费墙前半可见的生产叙述,已如实标注)。
`Corroborated (2 independent sources)`, personal blogs x2 (one is a partially paywalled production narrative, flagged as such).
Redis / Valkey 年份:—
单线程 + O(N):一条 KEYS 卡死整个实例 3.8 秒
多方印证
性能问题
- 一句话
Redis 的快建立在"所有操作都是 O(1)"的假设上;一条 O(N) 命令就是扔进单线程引擎的手榴弹——KEYS * 在 800 万 key 的库上卡了 3.8 秒。
Redis's speed rests on the assumption that every operation is O(1); a single O(N) command is a grenade in a single-threaded engine — KEYS * on an 8-million-key database blocked it for 3.8 seconds.
- 窄场景
key 数百万级的生产主库;后台脚本/数据分析同学直连生产库;大 Key(MB 级 String、百万成员集合、超长 Lua)。
Production primaries with millions of keys; background scripts or data analysts connected directly to production; big keys (MB-scale strings, million-member sets, long Lua scripts).
- 机制
命令执行单线程,一条 KEYS / FLUSHALL / 大 HGETALL / LRANGE 0 -1 / 长 Lua 执行期间,所有其他客户端命令(含 PING)在 TCP 缓冲区排队;DEL 大 Key 是同步递归释放内存(UNLINK 4.0+ 才把释放丢给后台线程),且 UNLINK 只解决"删"的阻塞,读大 Key 依然阻塞。
Command execution is single-threaded — while one KEYS / FLUSHALL / large HGETALL / LRANGE 0 -1 / long Lua script runs, every other client's commands (including PING) queue up in TCP buffers; DEL of a big key frees memory synchronously and recursively (only UNLINK, 4.0+, offloads the release to a background thread), and UNLINK only fixes the delete path — reading a big key still blocks.
- 生产验证
来源 5:数据分析同学在 800 万 key 生产主库执行 `KEYS user_tag_*`,SLOWLOG 显示 3.84 秒;Go 网关 P99 从 25ms 飙到 5000ms+,连接池耗尽级联雪崩,可用性跌至 20%;
来源 6:Mistake #6 "one slow command blocks every other client",50 万 key 的生产库上 KEYS * 是灾难;
来源 1:KEYS 是 O(n) 全程阻塞主线程,"排查事故时亲手制造事故"(用 KEYS 找大 key 的经典死循环)。
Source 5: a data analyst ran `KEYS user_tag_*` on an 8-million-key production primary, SLOWLOG showed 3.84 seconds; the Go gateway's P99 jumped from 25ms to 5000ms+, the connection pool drained and the outage cascaded, availability fell to 20%;
Source 6: Mistake #6 — "one slow command blocks every other client", KEYS * on a 500k-key production database is a disaster;
Source 1: KEYS is O(n) and blocks the main thread throughout — "causing the outage while troubleshooting the outage" (the classic loop of using KEYS to hunt big keys).
- 证据等级
`多方印证(3 个独立来源)`,含 1 个中文社区具名生产复盘。
`Corroborated (3 independent sources)`, including 1 named production postmortem from the Chinese community.
Redis / Valkey 年份:—
热 Key:一个 key 吃掉一个分片,加节点也救不了
多方印证
性能问题
- 一句话
一个 key 吃掉 40% 命令量,它所在的 slot、分片、CPU 核心就被吃满——集群其他节点再闲也帮不上忙,加节点也没用。
One key absorbing 40% of commands saturates its slot, its shard, its CPU core — idle nodes elsewhere in the cluster cannot help, and adding nodes does not either.
- 窄场景
排行榜/全局计数器/全员刷新的 session blob/配置 key;读写热点皆可。
Leaderboards, global counters, session blobs refreshed by everyone, config keys; read-heavy or write-heavy hotspots alike.
- 机制
单分片命令执行单线程,一个 key 固定属于一个 slot、一个分片;热点 key 的流量无法被分片数摊薄(扩容只切分 slot 空间,不切分 key);同分片的冷 key 也被拖慢(排队),p99 从 2ms 爬到 400ms 而命中率仍 99%,极具迷惑性;加副本只救读热点,写热点无解。
Single-shard command execution is single-threaded; one key always lives in one slot on one shard; hotspot traffic cannot be amortized by shard count (scaling only repartitions the slot space, never splits a key); cold keys on the same shard suffer too (queuing), so p99 climbs from 2ms to 400ms while the hit rate stays at 99% — maximally misleading; replicas only help read hotspots, write hotspots have no escape.
- 生产验证
来源 7:p99 2ms→400ms、命中率 99% 不变,单 key 占 40% 命令量导致分片 CPU 打满;"重启、垂直扩容、加副本(写热点下)都救不了";
来源 8:hot key 导致分片 CPU/内存热点,且 reshard 治标不治本——热点根因(坏哈希/热 key)不除,不均衡会回来。
Source 7: p99 2ms to 400ms with hit rate unchanged at 99%, one key at 40% of commands saturating a shard's CPU; "restarts, vertical scaling, and replicas (for write hotspots) cannot save you";
Source 8: hot keys create CPU/memory hotspots per shard, and resharding is only a temporary fix — if the root cause (bad hash, hot key) is not removed, the imbalance comes back.
- 证据等级
`多方印证(2 个独立来源)`,个人博客 ×2。
`Corroborated (2 independent sources)`, personal blogs x2.
Redis / Valkey 年份:—
缓存击穿:一个 key 过期,DB 被打穿
多方印证
性能问题
- 一句话
热门 key 一过期,成千上万个并发请求同时 miss、同时回源——缓存兢兢业业工作了几小时,死在它"正常过期"的那 200 毫秒。
A hot key expires, and thousands of concurrent requests miss at once and hit the origin — the cache worked diligently for hours, then died in the 200 milliseconds of its own normal expiration.
- 窄场景
高并发读热点 key 的 TTL 到期瞬间;部署/预热时批量写入相同 TTL 的 key 群(集体过期)。
High-concurrency reads on a hot key at the instant its TTL lapses; key populations bulk-written with identical TTLs during deployment or warmup (collective expiry).
- 机制
过期窗口内每个请求独立看到 miss、独立回源打 DB;连接池按稳态配,一波重复查询直接打满;更隐蔽的是"集体过期":同一秒写入的大批 key 五分钟后同一秒过期,单 key 防御拦不住。
Inside the expiry window, each request independently sees a miss and independently hits the DB; connection pools are sized for steady state, and one wave of duplicate queries fills them; the subtler variant is collective expiry: keys written in the same second all expire in the same second five minutes later, which per-key defenses cannot stop.
- 生产验证
来源 6:Mistake #3,key 过期 → 上千并发同时 miss → 上千并发打 DB;防御:TTL jitter、SETNX 分布式锁 singleflight、后台预刷新;
来源 14:200ms 窗口内上千请求同时 miss 的完整机制推演,"the busier and more successful your hot key is, the worse the stampede when it finally expires"。
Source 6: Mistake #3 — key expires, thousands of concurrent misses, thousands of concurrent DB hits; defenses: TTL jitter, SETNX distributed-lock singleflight, background pre-refresh;
Source 14: full mechanism walkthrough of thousands of requests missing inside a 200ms window — "the busier and more successful your hot key is, the worse the stampede when it finally expires".
- 证据等级
`多方印证(2 个独立来源)`,个人博客 + 公司技术博客。
`Corroborated (2 independent sources)`, personal blog + company engineering blog.
Redis / Valkey 年份:—
Cluster 跨槽:MULTI/EXEC 一到集群就报 CROSSSLOT
多方印证
运维复杂度
- 一句话
单机上一个 MULTI/EXEC 原子搞定的事,一上 Cluster 就报 CROSSSLOT——跨分片没有原子性,这是设计不是 bug,但教程从不提前说。
What one MULTI/EXEC handled atomically on a standalone server fails with CROSSSLOT on Cluster — there is no cross-shard atomicity; it is the design, not a bug, but tutorials never say so in advance.
- 窄场景
从单机/哨兵迁到 Cluster 的团队;库存扣减、转账类多 key 原子操作;Lua 脚本触多 key。
Teams migrating from standalone or Sentinel to Cluster; multi-key atomic operations like inventory deduction or transfers; Lua scripts touching multiple keys.
- 机制
key 按 CRC16(key)%16384 落 slot,一 slot 一主分片;多 key 命令/事务/Lua 要求所有 key 同 slot,否则拒绝;逃生舱是 hash tag({...} 内子串参与哈希),但 tag 选错(如 {listing} 通用词)等于把全部数据赶到一个分片,人造热 Key;已上线 key 加 tag 等于重写 key 名,retrofit 痛苦。
Keys land on slots by CRC16(key)%16384, one slot per primary shard; multi-key commands, transactions, and Lua require all keys in the same slot or they are rejected; the escape hatch is the hash tag (the substring in {...} participates in hashing), but a wrongly chosen tag (e.g. a generic {listing}) packs all data onto one shard — a hand-made hot key; adding tags to keys already in production means rewriting key names, a painful retrofit.
- 生产验证
来源 9:库存预订三 key 落三 slot 三节点,MULTI 直接 CROSSSLOT,"most tutorials never mention it until you hit it in production";
来源 10(具名):"If you want to run a command that touches multiple keys at once, those keys generally need to live in the same hash slot","reaching for Cluster before you have that problem trades a real, current cost for a hypothetical future one"。
Source 9: an inventory-reservation flow with three keys landing in three slots on three nodes, MULTI rejected with CROSSSLOT — "most tutorials never mention it until you hit it in production";
Source 10 (named): "If you want to run a command that touches multiple keys at once, those keys generally need to live in the same hash slot"; "reaching for Cluster before you have that problem trades a real, current cost for a hypothetical future one".
- 证据等级
`多方印证(2 个独立来源)`,个人博客 ×2。
`Corroborated (2 independent sources)`, personal blogs x2.
Redis / Valkey 年份:—
Reshard:在线扩容不是免费午餐
多方印证
运维复杂度
- 一句话
Cluster 在线 reshard 被当成"后台小事",实际上是带流量的运维事件——迁 slot 期间吃带宽吃 CPU,延迟毛刺可能超过 SLA。
Cluster resharding is treated as a "background detail", but it is really a traffic-carrying operational event — slot migration consumes bandwidth and CPU, and latency spikes can exceed the SLA.
- 窄场景
集群扩容/缩容、slot 迁移;把 --cluster reshard 当日常操作跑的团队。
Cluster scale-out/scale-in, slot migration; teams running --cluster reshard as routine maintenance.
- 机制
迁 slot 要在源/目标节点间搬 key,持续消耗带宽与 CPU;客户端 topology 缓存过期会在新旧节点间 ping-pong(MOVED/ASK 重定向风暴);更根本的是:若不均衡的根因是热 key/坏哈希,reshard 只是把歪移到别处,歪还会回来——治标不治本。
Migrating slots moves keys between source and target nodes, continuously consuming bandwidth and CPU; clients with stale topology caches ping-pong between old and new nodes (MOVED/ASK redirect storms); more fundamentally, if the imbalance's root cause is a hot key or bad hash, resharding just moves the skew elsewhere — it comes back.
- 生产验证
来源 8:reshard 期间 "latency spikes that may exceed acceptable SLAs";对坏哈希要极高的分片数才有效,等于拿大量闲置分片换均匀,浪费资源;
来源 10(具名):"moving slots between nodes as the cluster grows is an online operation but not a free one, and it adds load while it's happening"。
Source 8: resharding can cause "latency spikes that may exceed acceptable SLAs"; against bad hashes it only works with a very high shard count, which is wasted capacity buying evenness;
Source 10 (named): "moving slots between nodes as the cluster grows is an online operation but not a free one, and it adds load while it's happening".
- 证据等级
`多方印证(2 个独立来源)`,个人博客 ×2。
`Corroborated (2 independent sources)`, personal blogs x2.
- 备注
NextLevelDev 2026-07 有一篇 reshard 致 MOVED 风暴的生产叙述("Your Redis Cluster Reshard Has Been Running for 6 Hours"),因浏览器请求限流未能打开核实,未收录;如需可后续补充。
NextLevelDev's Jul 2026 piece on a reshard-caused MOVED storm ("Your Redis Cluster Reshard Has Been Running for 6 Hours") could not be opened for verification due to browser rate limiting and was not included; it can be added later if needed.
Redis / Valkey 年份:—
"全托管"但 WLM/VACUUM 还得自己调:几十亿行表的自动 vacuum sort 罢工
多方印证
性能问题运维复杂度
- 一句话
Redshift 自称托管数仓,但 WLM 队列、查询队列、VACUUM 策略仍是用户必修课;自动 vacuum sort 在几十亿行大表上直接不工作,大表维护还得人工排期。
Redshift calls itself a managed warehouse, yet WLM queues, query queues, and VACUUM strategy remain the user's homework; automatic vacuum sort simply doesn't work on multi-billion-row tables, so big-table maintenance still needs manual scheduling.
- 窄场景
TB/十亿行级大表、高频 UPDATE/DELETE 的集群。
TB/multi-billion-row tables with heavy UPDATE/DELETE.
- 机制
Redshift 列存追加写,UPDATE/DELETE 产生死行与乱序;AUTO VACUUM DELETE 与自动表排序在后台做,但大表的 VACUUM SORT 代价高、自动调度覆盖不到,只能人工在低峰期跑 VACUUM,且 VACUUM 与 ALTER DISTSTYLE 等操作互斥。
Redshift's columnar store is append-only; UPDATE/DELETE create dead rows and disorder. AUTO VACUUM DELETE and automatic table sort run in the background, but VACUUM SORT on huge tables is expensive and the automatic scheduler doesn't cover it — it must be run manually during off-peak hours, and VACUUM is mutually exclusive with operations like ALTER DISTSTYLE.
- 生产验证
来源 5:TrustRadius 认证用户评论(页面未标注日期)——"Amazon Redshift is a Managed Service. But it is Not a 100% managed service. We still need to configure it with WLM settings, and add Query Queues... 'Vacuum'... They recently started doing automated vacuuming. Prior to that we had to do that at regular intervals.";
来源 5:TrustRadius 同页 Cloudwalker 评论(5 年经验,认证评论)——"Automatic vacuum sort doesn't work for several billion rows tables"。
Source 5: TrustRadius verified-user review (no date shown) — "Amazon Redshift is a Managed Service. But it is Not a 100% managed service. We still need to configure it with WLM settings, and add Query Queues... 'Vacuum'... They recently started doing automated vacuuming. Prior to that we had to do that at regular intervals.";
Source 5: another TrustRadius review on the same page (Cloudwalker, 5 years experience, verified) — "Automatic vacuum sort doesn't work for several billion rows tables."
- 证据等级
`多方印证(2 个独立来源)`,点评平台认证用户 ×2。
`Corroborated (2 independent sources)`, platform-verified reviewers x2.
- 备注
本卡主题可能与本站 [避坑] 卡重叠(VACUUM 运维税)。
this card's topic may overlap existing [Pitfall] cards on this site (VACUUM ops tax).
Amazon Redshift 年份:—
2024-12-16 全球大故障:10 个区域 13 小时,CTAS 和 window function 批量报 internal error
多方印证
稳定与故障
- 一句话
一次软件更新引入向后不兼容的元数据 schema 变更,23 个区域中 10 个约 13 小时无法执行查询或 ingest——下游客户看到的是批量 `"SQL execution internal error"`,特定于 CTAS 和带 window function 的 INSERT。
A software update shipped a backwards-incompatible metadata schema change; 10 of 23 regions couldn't run queries or ingest data for ~13 hours — downstream customers saw batches of `"SQL execution internal error"`, specific to CTAS and INSERTs with window functions.
- 窄场景
2024-12-16 当天在受影响区域跑 ETL/ingest 的所有客户;依赖 Snowflake 做数据管道的下游 SaaS。
Anyone running ETL/ingest in affected regions on Dec 16 2024; downstream SaaS platforms built on Snowflake.
- 机制
服务端 rollout 的不兼容变更直接击穿查询引擎;客户端无任何自救手段,只能等 Snowflake 回滚。
A server-side rollout's incompatible change broke the query engine directly; clients had no self-service remedy and could only wait for Snowflake's rollback.
- 生产验证
来源 15:Keboola 官方状态页(Snowflake 下游数据平台)——其 Snowflake transformations 大量报错 `"SQL execution internal error ... incident 5370475"`,"specific to CTAS and INSERT statements with window functions",19:30 UTC 确认美国区恢复;
另有 Network World、InfoWorld、The Register 同日独立报道同一事件(10 区域、13 小时;Snowflake 归因于 "backwards-incompatible database schema update")。
—
- 证据等级
`多方印证(2 个独立来源)`,下游平台状态页 + 三家独立媒体报道同一事件。
`Corroborated (2 independent sources)`, downstream platform status page + three independent media reports of the same incident.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:—
自增 ID 跳号、不从 0 起,容灾还要手工复制序列
多方印证
运维复杂度生态与信任
- 一句话
不同表的自增 ID 跳号、不从 0 开始;做容灾时还得把自增序列从生产环境复制到容灾环境——MySQL 用户的直觉在这里处处碰壁。
Auto-increment IDs jump across tables and don't start from zero; for disaster recovery you must copy the sequences from production to the DR environment by hand — MySQL intuition breaks everywhere here.
- 窄场景
从 MySQL 迁移过来、依赖自增 ID 连续/可预测的业务;有异地容灾、需要序列在两边对齐的部署。
Workloads migrated from MySQL that rely on continuous/predictable auto-increment IDs; deployments with cross-site DR requiring sequence alignment on both sides.
- 机制
TiDB 的 AUTO_INCREMENT 是按 TiDB 节点批量申请、缓存分配的(`auto_increment_increment/offset` 语义与 MySQL 不同),各表独立计数,ID 不保证连续、也不从 0 起;序列(sequence)是独立对象,不在数据复制流里,容灾端不会自动对齐,切换后可能发号冲突或跳号。
TiDB's AUTO_INCREMENT is allocated in batches per TiDB node with caching (`auto_increment_increment/offset` semantics differ from MySQL); each table counts independently, so IDs are neither guaranteed continuous nor starting at zero; sequences are standalone objects outside the data replication stream, so the DR side never aligns automatically — after failover you can get ID collisions or jumps.
- 生产验证
来源 4:PeerSpot 具名评价,Casafari 高级工程师 Ivan Makarenko——"IDs in different tables jump and do not start from zero"(引自原文);
来源 4:PeerSpot 具名评价,Verinite 首席顾问 Shailesh Shandilya——需要把自增序列从生产环境复制到容灾环境。
Source 4: PeerSpot named review, Ivan Makarenko (Senior Engineer, Casafari) — "IDs in different tables jump and do not start from zero" (quoted verbatim);
Source 4: PeerSpot named review, Shailesh Shandilya (Principal Consultant, Verinite) — sequences must be copied from the production environment to the DR environment.
- 证据等级
`多方印证(2 个独立来源)`,具名用户评价 ×2。
`Corroborated (2 independent sources)`, named user reviews x2.
TiDB 年份:—
事务日志账单暗坑:向量重嵌一次生成约 23GB WAL
单方声音
成本账单
- 一句话
AlloyDB 的 WAL 既是复制机制也是备份机制——7 天之后的日志保留单独收费,向量场景一次全量重嵌就能烧出几十 GB 的 WAL 账单。
In AlloyDB, WAL is both the replication and the backup mechanism — log retention beyond 7 days is billed separately, and a single full re-embedding in a vector workload can burn tens of GB of WAL charges.
- 窄场景
在 AlloyDB 上跑 RAG/向量检索、定期全量重嵌 embedding 的团队;写放大大的向量工作负载。
Teams running RAG/vector search on AlloyDB with periodic full-corpus re-embedding; write-amplified vector workloads.
- 机制
AlloyDB 是日志型架构,事务日志支撑跨区复制与 PITR;前 7 天日志保留免费,超过 7 天按 $0.113/GB/月收费(备份同价)。向量重嵌本质是全表大 UPDATE,MVCC 下每行新版本都要写 WAL——10M 条 768 维向量一次全量重嵌约产生 23GB WAL,日志保留费用随重嵌频率线性累积。
AlloyDB's log-based architecture uses transaction logs for cross-region replication and PITR; the first 7 days of retention are free, beyond that $0.113/GB/month (same as backup storage). A re-embedding is a full-table mega-UPDATE — under MVCC every new row version writes WAL, so 10M 768-dim embeddings produce ~23GB of WAL per full pass, and log-retention charges accumulate linearly with re-embedding frequency.
- 生产验证
来源 1,2026-04 独立 field notes——明确点出 "The transaction log cost is a gotcha for write-heavy vector workloads",给出 10M embeddings / 768 维 / 全量重嵌 ≈ 23GB WAL 的测算,并提醒 "Plan accordingly"。
Source 1, Apr 2026 independent field notes — explicitly flags "The transaction log cost is a gotcha for write-heavy vector workloads," with the worked estimate of ≈ 23GB WAL for 10M embeddings at 768 dimensions, advising readers to "Plan accordingly."
- 证据等级
`单方声音`,独立博客 field notes(细节充分:具体单价、23GB 测算过程、重嵌场景)。
`Single voice`, independent blog field notes (detailed: exact unit prices, the 23GB worked estimate, the re-embedding scenario).
Google AlloyDB(AlloyDB for PostgreSQL) 年份:2026
托管黑盒:无真 SUPERUSER、扩展白名单、存储内部不透明
单方声音
生态与信任
- 一句话
"PostgreSQL-compatible, not PostgreSQL"——超管权限被阉割、扩展只能用白名单、存储层文档语焉不详,习惯按原理排障的团队会很难受。
"PostgreSQL-compatible, not PostgreSQL" — superuser privileges are curtailed, extensions are allow-list only, and storage-layer documentation is hand-wavy. Teams used to reasoning from first principles will chafe.
- 窄场景
依赖 PG 原生 superuser 操作、第三方/C 扩展(如全文检索、时序类扩展)的团队;习惯对照公开架构文档做性能归因的团队。
Teams relying on native superuser operations or third-party/C extensions (full-text search, time-series); teams used to attributing performance anomalies from published architecture internals.
- 机制
托管 AlloyDB 只给 `alloydbsuperuser`(与其他托管 PG 一样,无真 SUPERUSER);扩展走 allow-list,自定义 C 扩展不支持;存储层被描述为 "intelligent",但 Google 公开的复制协议、仲裁、写路径细节远少于 AWS 对 Aurora 的披露——客户对实例之下发生的事情可见性更低。
Managed AlloyDB grants only `alloydbsuperuser` (no true SUPERUSER, like other managed Postgres services); extensions go through an allow-list and custom C extensions are unsupported; the storage layer is marketed as "intelligent," but Google publishes far fewer specifics on replication protocol, quorum, and write path than AWS does for Aurora — customers get less visibility into what happens beneath the instance.
- 生产验证
来源 4,2026-05 独立评测 "Non-brochure concerns" 章节——逐条列出:存储文档 "thinner than Aurora's","the customer has less visibility into what is happening under the instance";"AlloyDB is PostgreSQL-compatible, not PostgreSQL. Behaviors differ at the edges, and teams that rely on stock internals for operational reasoning will be frustrated.";扩展白名单与无真 SUPERUSER 同列为 Negatives。
Source 4, May 2026 independent review, "Non-brochure concerns" section — itemizes each: storage documentation "thinner than Aurora's," "the customer has less visibility into what is happening under the instance"; "AlloyDB is PostgreSQL-compatible, not PostgreSQL. Behaviors differ at the edges, and teams that rely on stock internals for operational reasoning will be frustrated."; the extension allow-list and no-true-SUPERUSER both listed under Negatives.
- 证据等级
`单方声音`,独立技术评测(多条 concerns 出自同一篇系统性评测)。
`Single voice`, independent technical review (multiple concerns from one systematic review).
Google AlloyDB(AlloyDB for PostgreSQL) 年份:2026
扩缩容是中断性事件:改规格要重启,HA 下触发故障转移
单方声音
运维复杂度
- 一句话
AlloyDB 实例改规格不是在线的——要重启实例,HA 部署下还会触发一次故障转移,扩缩容得当成计划内中断来排。
Resizing an AlloyDB instance is not online — it restarts the instance, and on an HA deployment it triggers a failover. Treat every resize as a planned interruption.
- 窄场景
负载波动大、需要频繁纵向扩缩容的团队;把"弹性"当成理所当然的云原生用户。
Teams with spiky loads needing frequent vertical resizing; cloud-native users who take "elasticity" for granted.
- 机制
AlloyDB 计算实例仍是 VM 形态的规格绑定,改 vCPU/内存需要重启实例进程;在 HA 部署(主+备跨区)下,重启走故障转移路径完成。无论扩容还是缩容,都是 operationally significant 的事件——和"无缝弹性"不是一回事。
AlloyDB compute instances are VM-shaped and spec-bound; changing vCPU/memory restarts the instance process. On HA deployments (primary + cross-zone standby), the restart goes through the failover path. Scale-up and scale-down alike are operationally significant events — not "seamless elasticity."
- 生产验证
来源 4,2026-05 独立评测——"a resize involves an instance restart and, on an HA deployment, a failover. Both scale-up and scale-down are operationally significant events. Teams used to elastic-compute models tend to underweight how disruptive a resize is."
Source 4, May 2026 independent review — "a resize involves an instance restart and, on an HA deployment, a failover. Both scale-up and scale-down are operationally significant events. Teams used to elastic-compute models tend to underweight how disruptive a resize is."
- 证据等级
`单方声音`,独立技术评测。
`Single voice`, independent technical review.
Google AlloyDB(AlloyDB for PostgreSQL) 年份:2026
Resources exceeded:serverless 不等于无限,估错资源就拒绝执行
单方声音
性能问题
- 一句话
BigQuery 按"预估扫描量"给你分配执行资源——查询实际比预估重,它不加价接着跑,而是直接报错让你重写 SQL。
BigQuery allocates execution resources from the scan estimate — when a query turns out heavier than estimated, it doesn't scale up and charge more; it errors out and tells you to rewrite the SQL.
- 窄场景
内存密集型写法(大 GROUP BY、窗口函数、无 LIMIT 的 ARRAY_AGG);数据量增长后原本能跑的定时任务突然失败的数据管道。
Memory-heavy patterns (big GROUP BYs, window functions, unbounded ARRAY_AGG); scheduled pipelines that suddenly start failing as data grows.
- 机制
BigQuery 根据扫描量预估分配 slot/内存,单槽内存有硬上限;`ARRAY_REVERSE(ARRAY_AGG(…))` 这类写法会把每组的全量宽行数组物化在内存里再取一个,数据量一涨就触顶(Peak usage: 103% of limit);报错信息 "Resources exceeded during query execution: The query could not be executed in the allotted memory";解法是改写查询(如 `ARRAY_AGG(… LIMIT 1)`),而不是加钱扩容——serverless 的自动扩缩在此处不兜底。
BigQuery sizes slots/memory from the scan estimate, with a hard per-slot memory cap; patterns like `ARRAY_REVERSE(ARRAY_AGG(…))` materialize full wide-row arrays per group in memory before taking one element, and growing data pushes them over the cap (Peak usage: 103% of limit); the error reads "Resources exceeded during query execution: The query could not be executed in the allotted memory"; the fix is rewriting the query (e.g. `ARRAY_AGG(… LIMIT 1)`), not paying for more capacity — serverless autoscaling doesn't cover this case.
- 生产验证
来源 7,Mozilla bigquery-etl 公开 PR #9713(2026-07-22)——`search_terms_daily_v1` 自 2026-06-18 起在 Airflow 中持续失败,根因为 merino 流量增长后每组数组超内存上限;修复顺带暴露了更惨的后果:源表只保留 ~14 天,2026-06-18 → 2026-07-07 的数据**永久不可恢复**,形成数据缺口。具名生产事故,有 Bugzilla 编号(Bug 2048966)。
Source 7, Mozilla's bigquery-etl public PR #9713 (Jul 22 2026) — `search_terms_daily_v1` had been failing in Airflow since Jun 18 2026 because per-group arrays exceeded the memory cap as merino traffic grew; the fix exposed a harsher consequence: the source table retains only ~14 days, so Jun 18 → Jul 7 2026 data is **permanently unrecoverable** — a documented data gap. A named production incident with a Bugzilla ID (Bug 2048966).
- 证据等级
`单方声音`,具名生产事故修复记录(Mozilla 公开仓库 PR,含事故时间线与数据缺口文档)。
`Single voice`, named production-incident fix record (Mozilla public repo PR, with incident timeline and documented data gap).
- 备注
该机制早在 2019 年就有用户在 Google 开发者论坛记录过("Cartesian Blast"案例:估错资源后拒绝执行而非加价继续),Mozilla 2026 年事故证明该行为延续至今,故以新事故为主要证据。
the mechanism was reported as early as 2019 in a Google developer forum ("Cartesian Blast": underestimated jobs get rejected rather than upsized); Mozilla's 2026 incident shows the behavior persists, so the recent incident is the primary evidence.
Google BigQuery 年份:2026
分区键改出查询计划锁风暴:Cloudflare 账单管线的血泪
单方声音
性能问题
- 一句话
为了做按租户 retention 把分区键从 `(day)` 改成 `(namespace, day)`,结果查询计划阶段的全局 part 列表锁变成瓶颈,账单作业差点赶不上出账 deadline。
Changing the partition key from `(day)` to `(namespace, day)` for per-tenant retention turned the query planner's global part-list lock into the bottleneck, nearly missing the billing deadline.
- 窄场景
part 数以万计、并发查询数百的大表;分区键含高基数列(租户/namespace)且 part 总数持续增长的集群。
Tables with tens of thousands of parts and hundreds of concurrent queries; partition keys containing high-cardinality columns (tenant/namespace) with ever-growing total part counts.
- 机制
查询计划阶段每个线程都要拿 `MergeTreeData` 的**排他锁**、完整拷贝全表 part 列表、再按分区过滤。part 数涨到几万后,火焰图显示 45% 的叶查询 CPU 花在 `filterPartsByPartition`,超过一半的查询耗时是在等这一把互斥锁——"所有线程排成一队"。这不是读了更多 part,而是 part 的"存在本身"拖慢了计划。
During query planning, every thread took an **exclusive lock** on the `MergeTreeData` mutex, copied the entire table's part list, then filtered by partition. With tens of thousands of parts, flame graphs showed 45% of leaf-query CPU in `filterPartsByPartition` and over half of query time just waiting on that one mutex — "all standing in a single-file line." It was not reading more parts; the mere existence of the parts slowed planning.
- 生产验证
来源 5:Cloudflare 官方工程博客(2026-05)具名深挖——Ready-Analytics 表(2PiB,数百万行/秒写入),2025-01 开始迁移,3 月底账单聚合作业持续变慢;排查数天才发现 part 总数(3 万→一年后 16 万/副本)与查询耗时强相关。他们写了三个补丁:排他锁改共享锁(立竿见影)、延迟拷贝 part 向量(只拷贝过滤后的列表,已合入上游 PR #85535,随 25.11 发布)、按 namespace 二分查找裁剪 part(2026-03 上线,查询耗时再降 50%)。文末坦言"uneasy truce":优化只是买时间,分区方案长期是否正确仍是开放问题。
Source 5: Cloudflare's official engineering blog (May 2026), a named deep-dive — the Ready-Analytics table (2 PiB, millions of rows/sec ingested); migration began Jan 2025, billing aggregation jobs started slowing late March; days of investigation before correlating query duration with total part count (30k → 160k parts/replica a year later). Three patches: exclusive→shared lock (immediate relief), deferred part-vector copying (copy only the filtered list; merged upstream as PR #85535, shipped in 25.11), and namespace binary-search part pruning (deployed Mar 2026, another 50% off query durations). The post closes on an "uneasy truce": the patches buy time, but whether the partitioning scheme is right long-term stays open.
- 证据等级
`单方声音`,具名深度生产复盘(Cloudflare 官方工程博客,含火焰图数据与上游合入记录)。
`Single voice`, named deep production postmortem (Cloudflare official engineering blog, with flame-graph data and upstream merge records).
- 备注
与现有深水区一(part 机制)部分重叠,但"查询计划锁争用"这一机制在深水区中未覆盖。
partially overlaps Deep-Dive 1 (part mechanics), but the query-planning lock-contention mechanism is not covered there.
ClickHouse 年份:2026
后台合并 OOM:JSON 动态列撑爆 merge 内存
单方声音
稳定与故障
- 一句话
JSON 动态列在后台合并时要把所有被追踪的子列解压进内存,600GB 的日分区直接把 merge 干到 `MEMORY_LIMIT_EXCEEDED`,失败的合并反复重试还拖垮了写入。
JSON dynamic columns force every tracked sub-column to be decompressed into memory during background merges; ~600GB daily partitions drove merges straight into `MEMORY_LIMIT_EXCEEDED`, and the endlessly retrying merges then starved writes.
- 窄场景
用 JSON/Object 动态类型存半结构化日志、分区粒度过大(日分区数百 GB)、merge 内存上限收紧的 ClickHouse Cloud/自建集群。
Semi-structured logs in JSON/Object dynamic types, oversized partitions (hundreds of GB per day), and tightened merge memory caps on ClickHouse Cloud or self-hosted clusters.
- 机制
MergeTree 后台合并需要把参与合并的 part 的列数据读入内存;JSON 列的每个动态子路径都是独立子列,`max_dynamic_paths=64` 意味着合并时要同时解压 64+ 个子列,内存需求随分区大小线性放大,超过 `max_memory` 上限即被杀;失败的合并会不断重试,重试风暴反过来挤占写入的内存/IO/CPU。
MergeTree background merges must read the merging parts' column data into memory; each dynamic sub-path of a JSON column is an independent sub-column, so `max_dynamic_paths=64` means decompressing 64+ sub-columns at once — memory demand scales linearly with partition size and dies past the `max_memory` cap; failed merges retry forever, and the retry storm squeezes the memory/IO/CPU that inserts need.
- 生产验证
来源 7:icanbwell/fhir-server 公开 PR(2026-06)——生产事故,ClickHouse 官方支持工单 Case #00046381(2026-05-31):`fhir.AccessLog` 的 `details JSON(max_dynamic_paths=64)` 列在约 600GB 日分区上合并时 OOM(Code 241),超过 14.4 GiB 上限,"Continuously retrying-and-failing merges starved inserts of memory, IO, and CPU",造成写入延迟尖刺。修复:`max_dynamic_paths` 64→16(实际只用 11 个 key,merge 内存降约 4 倍)、分区改按月、`insert_deduplicate=0`。
Source 7: icanbwell/fhir-server public PR (Jun 2026) — production incident, ClickHouse official support Case #00046381 (May 31, 2026): the `details JSON(max_dynamic_paths=64)` column on `fhir.AccessLog` OOMed merges (Code 241) on ~600GB daily partitions, exceeding the 14.4 GiB cap; "continuously retrying-and-failing merges starved inserts of memory, IO, and CPU," spiking insert latency. Fix: `max_dynamic_paths` 64→16 (only 11 keys actually used, ~4x less merge memory), monthly instead of daily partitions, `insert_deduplicate=0`.
- 证据等级
`单方声音`,具名开源项目生产事故(公开 PR + 官方支持工单号;PR 正文为 AI 辅助起草,已如实标注)。
`Single voice`, named open-source project's production incident (public PR + official support case number; PR body was AI-assisted — stated as-is).
ClickHouse 年份:2026
Distributed 表"写入成功"但数据静默堆积:队列无上限,最后 inode 耗尽整机宕机
单方声音
稳定与故障运维复杂度
- 一句话
写 Distributed 表只落本地磁盘队列、后台线程异步刷往分片——刷失败客户端无感知,队列默认无上限,文件涨到数百万耗尽 inode,`df -h` 还有空间但整机无法建文件。
Writing to a Distributed table lands only in a local disk queue; a background thread flushes to shards asynchronously — on failure the client never notices. The queue is unbounded by default; at millions of files the ext4 inodes run out — `df -h` shows free space but the box can't create files.
- 窄场景
多分片 Distributed 表写入;Keeper 抖动、超大 block、并发槽位配错的集群。
Writes to multi-shard Distributed tables; clusters with Keeper flapping, giant blocks, or misconfigured concurrency slots.
- 机制
三根因:Keeper 挂了→底层表只读、队列持续堆积;单个超大 block(如 5 亿行)触发 `max_execution_time` 被杀,又因队列按分片**串行**刷而成为"瓶塞"卡住后面所有正常 block;`max_concurrent_queries_for_user` 设太低,SELECT 占满槽位致后台 INSERT 拿不到执行位。缓解靠 `bytes_to_throw_insert` 当熔断器 + 告警 `DistributedFilesToInsert`——但默认全都没开。
Three root causes: Keeper down → underlying tables read-only, queue keeps growing; one giant block (e.g., 500M rows) trips `max_execution_time` and gets killed, then acts as a "cork" because the queue flushes per shard **serially**, blocking every healthy block behind it; `max_concurrent_queries_for_user` set too low lets SELECTs hog all slots so background INSERTs never get one. Mitigations (`bytes_to_throw_insert` as a circuit breaker, alerting on `DistributedFilesToInsert`) exist — but none are on by default.
- 生产验证
—
Source 9: Pranav Mehta, Medium Mar 2026 — multi-environment reproduction with all three root causes and the mitigation chain.
- 证据等级
`单方声音`,具名作者多环境复现、机制完整。
`Single voice`, named author with multi-environment reproduction and a complete mechanism.
ClickHouse 年份:2026
备份在超多表下慢到不可用:FREEZE 与哈希查询抢锁,9.5 小时 vs 26 分钟
单方声音
性能问题运维复杂度
- 一句话
约 5 万张 MergeTree 表做一次全量备份要 9.5 小时——`FREEZE` 与取 `hash_of_all_files` 的查询在抢同一把锁,单表 FREEZE 从 100ms 膨胀到 5 秒;跳过该查询后备份降到 26 分钟。
A full backup over ~50K MergeTree tables takes 9.5 hours — `FREEZE` and the `hash_of_all_files` lookup fight over the same mutex, inflating per-table FREEZE from ~100ms to ~5s; skipping that query drops the backup to 26 minutes.
- 窄场景
表数量极多(万级)的集群;用 clickhouse-backup 做全量备份。
Clusters with tens of thousands of tables; full backups with clickhouse-backup.
- 机制
备份工具为校验完整性从 `system.parts` 取 `hash_of_all_files`,与 FREEZE 存在互斥锁争用;表越多争用越烈。用户自己提 PR 加 opt-in 开关绕过。
The backup tool reads `hash_of_all_files` from `system.parts` for integrity checking, which contends mutex-wise with FREEZE; the more tables, the worse the contention. The user filed a PR adding an opt-in switch to bypass it.
- 生产验证
—
Source 11: public GitHub issue (clickhouse-backup #1465), user ruslanen, Jul 2026 — 954 GiB / 50K MergeTree + 500K Join tables, with concrete before/after timings.
- 证据等级
`单方声音`,生产数据具体的公开 issue。
`Single voice`, public issue with concrete production numbers.
ClickHouse 年份:2026
没有多语句事务:一次 sync 的新旧数据同时可见,逼得用户自建"类事务层"
单方声音
运维复杂度生态与信任
- 一句话
ClickHouse 不支持多语句事务——新一轮 sync 写入时用户同时看到新旧两批数据,CloudQuery 被迫在应用层造了个 high-level transaction layer 让每次 sync 对用户呈现原子性。
ClickHouse has no multi-statement transactions — during a fresh sync, users see old and new batches at the same time. CloudQuery built a high-level transaction layer in the application so each sync appears atomic to users.
- 窄场景
ETL sync 场景;需要"一次同步原子可见"的数据产品。
ETL sync workloads; data products needing "each sync atomically visible."
- 机制
"technically wasn't a race condition, but it created a confusing experience"——没有跨语句原子提交,应用层只能自己实现"事务层"。官方有实验性事务,但限制极多(见"缺口"卡)。
"Technically wasn't a race condition, but it created a confusing experience" — with no cross-statement atomic commit, the application implements the "transaction layer" itself. The official experimental transactions carry heavy restrictions (see the gap card).
- 生产验证
—
Source 12: CloudQuery company engineering blog, circa Apr 2026 — "If we'd put a queuing or buffering layer in place earlier, we might've avoided some of the mess."
- 证据等级
`单方声音`,具名公司生产复盘。
`Single voice`, named-company production retrospective.
ClickHouse 年份:2026
物化视图是异步黑盒:失败了不知道、何时算完不知道,CloudQuery 直接弃用
单方声音
运维复杂度
- 一句话
MV 相对创建/insert 语句是异步的——"Lack of visibility into failures…Unpredictable completion timing",换表迁移时无法确定 MV 何时追完,且无内置历史重算;CloudQuery 弃用 MV,改用定时任务显式刷新的 snapshot 表。
MVs are async relative to their CREATE/INSERT — "Lack of visibility into failures…Unpredictable completion timing." During table migrations you can't tell when an MV has caught up, and there's no built-in historical recompute; CloudQuery dropped MVs for explicitly-refreshed snapshot tables via scheduled jobs.
- 窄场景
用 MV 做预聚合/换表迁移的团队。
Teams using MVs for pre-aggregation or table migrations.
- 机制
MV 无失败可见性、无完成时间可预测性、无内置回填机制。"we know exactly what ran, when it ran, and what the output was—no guessing, no race conditions, no half-finished state."——这是弃用 MV 改手工快照表的理由。
No failure visibility, no predictable completion, no built-in backfill. "We know exactly what ran, when it ran, and what the output was—no guessing, no race conditions, no half-finished state." — the case for abandoning MVs in favor of hand-rolled snapshot tables.
- 生产验证
—
Source 12: CloudQuery official engineering blog, circa Apr 2026 — a strong statement: abandoning the entire feature.
- 证据等级
`单方声音`,具名公司。
`Single voice`, named company.
ClickHouse 年份:2026
ClickHouse Cloud 迁回自建是个坑:备份恢复不干净、查询行为有细微差异
单方声音
升级迁移
- 一句话
"We assumed we could just back up our managed cluster and restore it locally. Turns out, not so simple."——备份/恢复、查询行为、配置三处都有细微差异,"干净迁移"变成 patching/testing/head-scratching。
"We assumed we could just back up our managed cluster and restore it locally. Turns out, not so simple." Backup/restore, query behavior, and configs all differ subtly — a "clean migration" becomes patching/testing/head-scratching.
- 窄场景
从 ClickHouse Cloud 迁回自建(成本/合规原因)。
Moving off ClickHouse Cloud back to self-hosted (cost or compliance reasons).
- 机制
Cloud 与自建在配置、查询行为上存在细微差异,备份恢复过去不保证行为一致。Cloud 起步快,迁出难——"上云容易下云难"的 ClickHouse 版本。
Cloud and self-hosted differ subtly in configs and query behavior; backups don't guarantee behavioral parity on restore. Fast to start on Cloud, hard to leave — ClickHouse's version of "easy on, hard off."
- 生产验证
—
Source 12: CloudQuery official engineering blog, circa Apr 2026.
- 证据等级
`单方声音`,具名公司。
`Single voice`, named company.
ClickHouse 年份:2026
可靠写入没有官方答案:不用 Kafka,很多人根本不知道怎么把数据"可靠地"写进去[来源存疑]
单方声音
来源存疑
运维复杂度生态与信任
- 一句话
"every time I try to use it, I get stuck on 'how do I get data into it reliably'…inevitably end up with 'by combining clickhouse and Kafka', at which point my desire to keep going drops to zero."——方案全靠自拼:Vector 缓冲、攒批到 S3、Python connector 攒 100K 批。
"Every time I try to use it, I get stuck on 'how do I get data into it reliably'…inevitably end up with 'by combining clickhouse and Kafka', at which point my desire to keep going drops to zero." Every workable setup is hand-assembled: Vector as buffer, batching to S3, a Python connector batching 100K rows.
- 窄场景
第一次用 ClickHouse 做 ingestion 的团队;不想引入 Kafka 的轻量场景。
Teams ingesting into ClickHouse for the first time; lightweight setups that don't want Kafka.
- 机制
没有官方"标准写入姿势":小批次写触发 Too Many Parts,大批次要自己攒,不用 Kafka/ZK 就没有开箱即用的可靠写入语义。refreshable MV 当时还是 experimental。
No official "standard write path": small batches trigger Too Many Parts, big batches need hand-rolled buffering, and without Kafka/ZooKeeper there's no out-of-the-box reliable-write semantic. Refreshable MVs were still experimental at the time.
- 生产验证
—
Source 15: Hacker News thread (2026) — the OP plus sympathetic replies ("That's the same stage I get stuck every time") [Questionable source: anonymous].
- 证据等级
`单方声音[来源存疑]`,多人附和但全匿名。
`Single voice [Questionable source]`, multiple sympathetic replies but all anonymous.
ClickHouse 年份:2026
分布式 JOIN 忘写 GLOBAL:结果静默变错,没有任何报错[来源存疑]
单方声音
来源存疑
稳定与故障
- 一句话
分片集群上对 Distributed 表做普通 JOIN/IN,每分片只跟自己本地的数据 join——"returns results that are silently smaller or different than expected…with no error at all…nobody investigates a query that returns a plausible-looking, wrong number."
On a sharded cluster, a plain JOIN/IN against a Distributed table only joins each shard's local data — "returns results that are silently smaller or different than expected…with no error at all…nobody investigates a query that returns a plausible-looking, wrong number."
- 窄场景
分片集群上的 JOIN/IN 查询;从单机迁到分布式的团队。
JOIN/IN queries on sharded clusters; teams moving from single-node to distributed.
- 机制
普通 JOIN 只在分片本地执行,不加 GLOBAL 就得不到全局正确结果——但查询"正常"返回,看起来合理、实际是错的。修法是 GLOBAL JOIN/IN,但"记得加"是人的责任。
Plain JOINs execute shard-locally; without GLOBAL you don't get globally correct results — but the query "succeeds," looking plausible while being wrong. The fix is GLOBAL JOIN/IN, and "remembering to add it" is a human responsibility.
- 生产验证
—
Source 17: GitHub personal course notes, circa Sep 2026 [Questionable source: course-marketing notes, though the mechanism is real].
- 证据等级
`单方声音[来源存疑]`。
`Single voice [Questionable source]`.
ClickHouse 年份:2026
轻量更新的生产 bug:patch parts 致 MERGE_PARTS 卡死、重试 4000+ 次[来源存疑]
单方声音
来源存疑
稳定与故障
- 一句话
`fhir.dim_practitioners` 表出现 stuck MERGE_PARTS、>4000 次重试——"Likely cause: a ClickHouse bug in the lightweight updates/deletes ('patch parts') implementation for 25.10+",关联上游 open issue #89836、#89472。
A `fhir.dim_practitioners` table hit stuck MERGE_PARTS with 4,000+ retries — "Likely cause: a ClickHouse bug in the lightweight updates/deletes ('patch parts') implementation for 25.10+," linked to upstream open issues #89836 and #89472.
- 窄场景
用轻量更新(25.10+ patch parts)的表。
Tables using lightweight updates (25.10+ patch parts).
- 机制
patch parts 实现的 bug 致合并队列卡死、反复重试拖垮集群。与同一团队的 JSON 合并 OOM 事故形成"连环踩坑"。
A bug in the patch-parts implementation stalls the merge queue; retries pile up and drag the cluster down. Pairs with the same team's JSON-merge OOM incident as a "serial stumbling" record.
- 生产验证
—
Source 18: Altinity Jan 2026 slide deck (retelling the icanbwell team's incident) [Questionable source: vendor-retold].
- 证据等级
`单方声音[来源存疑]`。
`Single voice [Questionable source]`.
ClickHouse 年份:2026
Mutation 报成功但数据是坏的:静默损坏 + 重启崩溃循环 + 静默跳行[来源存疑]
单方声音
来源存疑
稳定与故障
- 一句话
UInt32→IPv6 的 ALTER UPDATE 被接受但写下 4 字节的"UInt32 形"数据到 16 字节列——`mutations_sync=2` 返回成功、`is_done=1`、无日志,坏数据只有读时才爆 Code 48;随后复合 RENAME 撞上坏 part 进入无限重试,重启后继续崩溃→systemd 重启循环。
An ALTER UPDATE from UInt32 to IPv6 was accepted yet wrote 4-byte "UInt32-shaped" data into a 16-byte column — `mutations_sync=2` returned success, `is_done=1`, no logs; the corruption only exploded as Code 48 on read. A compound RENAME then hit the bad part and entered infinite retries; restarts kept crashing → systemd restart loop.
- 窄场景
大表 ALTER UPDATE/类型变更;带 DEFAULT 列的 UPDATE。
ALTER UPDATE/type changes on huge tables; UPDATEs on columns with DEFAULTs.
- 机制
三宗"报成功但结果错":类型变更写坏数据(成功+无日志)、RENAME 撞坏 part 无限重试(KILL MUTATION 在启动期无效)、带 DEFAULT 的列 UPDATE 约一半行被静默跳过且重跑写不进去。共同 UX 失败:"the mutation subsystem reports success while the underlying write either didn't happen or happened incorrectly."
Three flavors of "reported success but wrong": type-change writes corrupt data (success + no logs), RENAME hits the bad part with infinite retries (KILL MUTATION ineffective during startup), and UPDATEs on DEFAULT columns silently skip ~half the rows with rewrites unable to fix them. The shared UX failure: "the mutation subsystem reports success while the underlying write either didn't happen or happened incorrectly."
- 生产验证
—
Source 19: ClickHouse official repo public issue, Apr 2026 (billion-row CollapsingMergeTree table, minimal repro included) [Questionable source: anonymous + version 23.7.5 + unverified on newer versions].
- 证据等级
`单方声音[来源存疑]`。
`Single voice [Questionable source]`.
ClickHouse 年份:2026
字典不是免费午餐:逐行 dictGet 成瓶颈,外部源 credential 得写进配置
单方声音
性能问题运维复杂度
- 一句话
用字典替代 JOIN 后内存从 50GB+ 降到 3.5GB,但"if you perform a dictionary lookup per row (e.g. using dictGet in the SELECT for millions of rows), it can bottleneck";从集群内其他表建字典还得把 credential 写进 SOURCE 配置。
Replacing JOINs with dictionaries cut memory from 50GB+ to 3.5GB — but "if you perform a dictionary lookup per row (e.g. using dictGet in the SELECT for millions of rows), it can bottleneck." Building a dictionary from another in-cluster table or custom SQL also requires writing credentials into the SOURCE config.
- 窄场景
用字典替代大表 JOIN;从外部源/自定义 SQL 建字典。
Replacing big-table JOINs with dictionaries; dictionaries from external sources or custom SQL.
- 机制
逐行 dictGet 在数百万行上成瓶颈;字典源引用集群内表/自定义 SQL 时 credential 必须写进配置(后被 named collections 缓解),带来配置复杂度和安全顾虑。
Per-row dictGet bottlenecks at millions of rows; referencing an in-cluster table or custom SQL as a dictionary source requires credentials in the config (later eased by named collections), adding config complexity and security concerns.
- 生产验证
—
Source 12: CloudQuery official engineering blog, circa Apr 2026.
- 证据等级
`单方声音`,具名公司次要抱怨。
`Single voice`, named company's secondary complaint.
ClickHouse 年份:2026
分布式语义全靠手写:加分片不搬历史数据,"no safe built-in online rebalance"
单方声音
运维复杂度生态与信任
- 一句话
ClickHouse 没有"透明分布式"——每分片建本地表 + 外层 Distributed 表 + 手设分片键 + DDL 全加 ON CLUSTER;更狠的是加分片不搬历史数据,新分片只接新写入,老分片持续 hot。
ClickHouse has no "transparent distribution" — per-shard local tables + an outer Distributed table + hand-set sharding keys + ON CLUSTER on every DDL. Worse: adding a shard doesn't move historical data; new shards only take new writes while old shards stay hot.
- 窄场景
集群扩容;从 Snowflake/BigQuery/Doris(透明分布式)迁来的团队。
Cluster scale-out; teams arriving from Snowflake/BigQuery/Doris (transparent distribution).
- 机制
OneUptime 运维手册称为扩容前必读的 "biggest gotcha":"Adding a shard does not move historical rows... ClickHouse has no safe built-in online rebalance, and naively re-inserting live rows through the Distributed table duplicates them." 想搬历史数据只能停写手动导。
OneUptime's ops manual calls it the must-read "biggest gotcha" before scaling: "Adding a shard does not move historical rows... ClickHouse has no safe built-in online rebalance, and naively re-inserting live rows through the Distributed table duplicates them." Moving history means stopping writes and hand-loading.
- 生产验证
—
Source 28: OneUptime (open-source observability platform, heavy ClickHouse user) official ops manual (updated 2026);
Source 15's HN thread: newcomers' first impression is "shards must be managed by hand."
- 证据等级
`单方声音`,重度用户运维手册为主。
`Single voice`, led by a heavy user's ops manual.
ClickHouse 年份:2026
强制遥测:免费版关不掉
单方声音
生态与信任
- 一句话
Enterprise Free 不要钱,但遥测必须开——"免费"的代价是你的集群使用数据持续回传,且没有 opt-out。
Enterprise Free costs nothing, but telemetry is mandatory — the price of "free" is your cluster's usage data phoning home continuously, with no opt-out.
- 窄场景
对数据出境/合规敏感(金融、医疗、政务)的自托管用户;内网隔离集群;把"免费"当纯开源用的团队。
Self-hosters sensitive to data egress or compliance (finance, healthcare, government); air-gapped clusters; teams treating "free" as pure open source.
- 机制
24.3 起自托管统一 Enterprise 许可,年营收 <1000 万美元免费;免费档的条款里遥测是强制的、不可关闭(来源 6)。这不是技术缺陷,是商业模式写进许可条款:用数据回传换免费。
Since 24.3, self-hosting is unified under the Enterprise license, free below $10M annual revenue — and telemetry is compulsory on the free tier with no way to disable it (Source 6). This is not a technical defect; the business model is written into the license terms: usage data in exchange for free.
- 生产验证
来源 6,2026-09——"CockroachDB Core 于 2024-11-18 退场,被 Enterprise Free 许可取代(年营收 <1000 万美元企业);代价是强制遥测,免费许可下无法关闭,且集群还得自己运维打补丁。"
Source 6, Sep 2026 — "CockroachDB Core was retired on November 18, 2024 and replaced by an Enterprise Free license for businesses under $10 million in annual revenue. The catch is mandatory telemetry, which you cannot opt out of on the free license, and you still run and patch the cluster yourself."
- 证据等级
`单方声音`,独立评测博客(细节充分:退场日期、门槛、遥测不可关闭三要素齐全)。
`Single voice`, independent review blog (fully detailed: retirement date, threshold, and non-optional telemetry).
CockroachDB 年份:2026
迁出要重写:CRDB 特有构造没有 PG 对应物
单方声音
升级迁移
- 一句话
迁入时"PG 线协议兼容",迁出时才发现用了 `unique_rowid()` 默认值、hash-sharded 索引、多 region 表 locality——这些在真 PG 里都不存在,重写 schema 跑不掉。
Moving in was "Postgres wire-protocol compatible"; moving out reveals `unique_rowid()` defaults, hash-sharded indexes, and multi-region table localities — none of which exist in real Postgres, so the schema gets rewritten.
- 窄场景
从 CockroachDB 迁回 Postgres/Neon/Aurora 的团队;当初为压热点行按 CRDB 官方指南改了主键与索引设计的项目。
Teams migrating from CockroachDB back to Postgres/Neon/Aurora; projects that followed CRDB's official guidance and redesigned primary keys and indexes to tame hotspots.
- 机制
厂商锁定不一定靠许可证,靠方言。CRDB 为分布式做的官方最佳实践(UUID/hash-sharded 主键打散热点、`unique_rowid()`、REGIONAL BY ROW 等 locality 语义)都是 PG 没有的概念;迁出等于一次反向 schema 改造。来源 10 也承认:"迁离分布式 SQL 主库去单区方案是重新架构,不是改配置。"
Lock-in is not always the license — sometimes it is the dialect. CRDB's official distributed best practices (UUID/hash-sharded primary keys to scatter hotspots, `unique_rowid()`, REGIONAL BY ROW locality semantics) are concepts Postgres simply does not have; migrating out is a reverse schema transformation. Source 10 concedes the point: "migrating off a distributed SQL primary to a sharded single-region one is a re-architecture, not a config change."
- 生产验证
来源 6,2026-09——从 CockroachDB 迁到 Neon 的检查项:"注意 CockroachDB 特有功能:unique_rowid() 默认值、hash-sharded 索引、多 region 表 locality 在 Postgres 里没有直接对应物,需要重写。"
Source 6, Sep 2026 — migration checklist for CockroachDB to Neon: "Watch for CockroachDB-specific features: unique_rowid() defaults, hash-sharded indexes and multi-region table localities have no direct Postgres equivalent and need rewriting."
- 证据等级
`单方声音`,独立评测博客(迁移检查清单,细节充分)。
`Single voice`, independent review blog (migration checklist, fully detailed).
CockroachDB 年份:2026
auto-stop 的两道裂缝:BI 心跳"幽灵"让它永不触发,5 分钟下限还是 UI 定的
单方声音
成本账单
- 一句话
Tableau 每 8 分钟一次的心跳被判定为"活跃查询",10 分钟 auto-stop 计时器无限重置——一个周末烧掉 $14k;而 5 分钟的下限只是 UI 的假限制,API 早就能设 1 分钟。
A Tableau heartbeat every 8 minutes counts as an "active query," resetting the 10-minute auto-stop timer forever — one weekend burned $14k; meanwhile the 5-minute minimum was never a platform limit, only a console UI constraint (the API has long accepted 1 minute).
- 窄场景
Serverless SQL warehouse 对接 BI 工具长连接;用默认 10 分钟 auto-stop 的团队。
Serverless SQL warehouses with long-lived BI connections; teams on the default 10-minute auto-stop.
- 机制
Serverless 仓库的"空闲"按是否有查询判定,BI 工具对 `system.information_schema` 的周期性心跳足以让计时器永远归零;Large 规格整周末为空查询全速烧 DBU。另一道裂缝:5 分钟最小值并非平台限制,只是控制台 UI 的约束,API 可设 1 分钟——差值是静默烧钱。
A Serverless warehouse judges "idle" by query presence, so a BI tool's periodic heartbeat against `system.information_schema` zeroes the timer indefinitely; a Large warehouse burns DBUs at full speed all weekend for empty queries. The second crack: the 5-minute floor is not a platform limitation, just a UI constraint — the API accepts 1 minute, and the difference is silent burn.
- 生产验证
—
Source 10: DEV community, aniketsoni, circa Sep 2026 — a Tableau heartbeat reset the 10-minute auto-stop every 8 minutes; a Large warehouse burned 80% of the monthly cloud budget in 48 hours ("The warehouse wasn't idling; it was being kept alive by a ghost."); fixed by isolating the heartbeat on a small warehouse and tightening permissions;
Source 14: Pravin Ghavare, Jan 2026, tested in production — "The 5-minute minimum is not a Databricks platform limitation. It is a UI constraint… DBUs burn silently while nobody is running queries." A 1-minute auto-stop works on real workloads.
- 证据等级
`单方声音`,两篇独立生产复盘(机制不同、主题同属 suspend 失效)。
`Single voice`, two independent production postmortems (different mechanisms, same suspend-failure theme).
Databricks 年份:2026
账单根本拆不清:billing 记录里 custom_tags 全空,分账靠事后"刑侦"
单方声音
运维复杂度成本账单
- 一句话
计费记录的 custom_tags 是空的、service principal 是一串不透明 UUID——钱花完之后,只能从 query history 反向"刑侦"这是谁花的。
Billing records carry empty `custom_tags`, service principals show up as opaque UUIDs — after the money is spent, the only way to know who spent it is to reverse-engineer it from query history.
- 窄场景
多团队共享 warehouse;FinOps 要做成本分摊/chargeback 的平台团队。
Multiple teams sharing warehouses; platform teams doing FinOps chargeback.
- 机制
`system.billing.usage` 的粒度是 compute-hour,没有 query_id/statement_id,join 不上 `system.query.history`;budget policy 打的标签不追溯历史;共享仓库的标签只能归到"shared"。归因基础设施缺失,账单精确到分却不知道该递给谁。
`system.billing.usage` is metered at compute-hour granularity with no query_id/statement_id, so it cannot join `system.query.history`; budget-policy tags don't apply retroactively; shared-warehouse tags collapse to "shared." The attribution infrastructure is missing: the bill is precise to the cent, but nobody knows whose desk to send it to.
- 生产验证
—
Source 11: Ke Zhu, Mar 2026 production postmortem — "The billing records had `custom_tags: {}`. The warehouses had `tags: {}`. Service principals showed up as opaque UUIDs. The only way to attribute cost was to reverse-engineer it from query history after the money was already spent.";
Source 12: official community thread — a large UK bank platform team in the same bind: precise totals, no owner [Questionable source: anonymous].
- 证据等级
`单方声音`,具名工程师生产复盘为主。
`Single voice`, led by a named engineer's production postmortem.
Databricks 年份:2026
Auto Loader 默认"目录扫描":S3 LIST 请求在规模下变成烧钱项
单方声音
性能问题成本账单
- 一句话
Auto Loader 默认用目录 LIST 轮询发现新文件,规模一大,S3 LIST API 本身变成账单上的一项——切事件通知模式后 ingestion 从 15 分钟降到 3 分钟。
Auto Loader discovers new files by polling with directory LIST calls by default — at scale the S3 LIST API itself becomes a billable line item; switching to event-notification mode cut ingestion from 15 minutes to 3.
- 窄场景
S3/ADLS 上文件多、到达频繁的 ingestion;直接用默认配置的 Auto Loader。
Ingestion from S3/ADLS with many files arriving frequently; teams running Auto Loader on defaults.
- 机制
directory listing 模式周期性对存储路径发 LIST 请求做全量扫描;文件数上来后单次扫描可达 15 分钟,LIST 按次计费且拖慢 ingestion。事件通知模式(SNS+SQS)只处理增量,成本与延迟双降。
Directory-listing mode periodically LIST-scans the whole storage path; with enough files a single scan takes ~15 minutes, LIST calls are billed per request and slow ingestion. Event-notification mode (SNS+SQS) processes only deltas — cheaper and faster on both axes.
- 生产验证
—
Source 13: Seemplicity (named Databricks customer) engineering blog, May 2026 — "The default, directory listing, periodically scans your storage path with LIST API calls… S3 LIST operations aren't free, and at scale those API calls add up. Our scans were taking around 15 minutes every run." Part of a $2,000→$500/day (75% cut) postmortem.
- 证据等级
`单方声音`,具名客户生产复盘。
`Single voice`, named-customer production postmortem.
Databricks 年份:2026
实例族 DBU 费率迷雾:内存翻倍的机器反而便宜一半
单方声音
成本账单
- 一句话
DBU 费率不是按"算力"统一定价的——换到内存翻倍(64GB)的内存优化型实例,账单反而便宜了一半。
DBU rates aren't uniform per unit of compute — moving to memory-optimized instances with twice the RAM (64GB) halved the bill.
- 窄场景
按默认推荐选通用型实例的团队;做实例选型的成本优化。
Teams picking general-purpose instances by default; anyone doing instance-level cost optimization.
- 机制
"the DBU rate is not the same across all instance families"——不同实例族的 DBU 费率不同,内存优化型实例的费率低到足以抵消规格上涨。按"配置越高越贵"的直觉选型,会系统性选贵。
"The DBU rate is not the same across all instance families" — memory-optimized families carry rates low enough to offset the bigger box. Sizing by the "bigger costs more" intuition systematically overspends.
- 生产验证
—
Source 13: Seemplicity engineering blog, May 2026 — "The memory-optimized instances we switched to had twice the RAM (64 GB) and cost roughly half as much as the general-purpose instances we were running."
- 证据等级
`单方声音`,具名客户实测。
`Single voice`, named-customer measurement.
Databricks 年份:2026
Auto Loader 的运维坑二连:schema 推断只看先到文件,checkpoint 还能自我摄取
单方声音
运维复杂度
- 一句话
schema 推断只采样"先到的文件",回填历史数据时类型漂移全进 _rescued_data;更绝的是 checkpoint 若嵌套在源路径里,Auto Loader 会把自己的 RocksDB 文件当数据源吃掉。
Schema inference samples only "the files that landed first," so backfilled history with type drift lands in _rescued_data; worse, nest the checkpoint inside the source path and Auto Loader ingests its own RocksDB files as data.
- 窄场景
历史数据回填(backfill);checkpointLocation 随手放在源路径子目录的流任务。
Historical backfills; streaming jobs whose checkpointLocation sits under the monitored source path.
- 机制
首次运行只采样最多 50GB/1000 文件推断 schema 并写入 schema location 成"合同"——"schema inference is only as good as your sample, and your sample is not random — it's whatever landed first." 回填 EDI JSON 时早期 string 后期 double 的列全被 rescue。另一坑:checkpoint 目录若在被监控路径内,Auto Loader 把 checkpoint 自己的 .zip/.log 反序列化(FAILED_READ_FILE.NO_HINT),RocksDB 损坏后关选项重跑也无法恢复,只能重置 checkpoint。
The first run samples at most 50GB/1,000 files, writes the schema to the schema location as a contract — "schema inference is only as good as your sample, and your sample is not random — it's whatever landed first." Backfilling EDI JSON where early files had strings and later files had doubles rescues whole columns. Separately: a checkpoint directory inside the watched path gets deserialized as source data (FAILED_READ_FILE.NO_HINT); the RocksDB corruption survives option toggles — only a checkpoint reset recovers.
- 生产验证
—
Source 17: Divyansh Goyal, Jun 2026 — backfilling 2,000+ history files; StructType-vs-ArrayType drift all landed in _rescued_data;
Source 18: official community thread, circa Sep 2026 — checkpoint self-ingestion killed the stream; the documented fix didn't work for the reporter, who reset the checkpoint.
- 证据等级
`单方声音`,两起独立运维事故(机制不同、同属 Auto Loader 运维坑)。
`Single voice`, two independent ops incidents (different mechanisms, same Auto Loader ops-pit theme).
Databricks 年份:2026
UC External Location 陷阱:报错与 IAM 真实配置脱节,200 个 job 的迁移教训
单方声音
运维复杂度升级迁移
- 一句话
S3 路径没显式注册为 UC External Location,作业报 cryptic 的 PERMISSION_DENIED——即使底层云 IAM 角色配得完全正确,报错也指不到点子上。
Skip registering an S3 path as a UC External Location and the job fails with a cryptic PERMISSION_DENIED — even when the underlying cloud IAM role is perfectly configured; the error never points at the missing registration.
- 窄场景
200+ 生产 job 的 Hive → UC 迁移;跨云存储路径多的老工作区。
Hive → UC migrations with 200+ production jobs; older workspaces with many cloud storage paths.
- 机制
UC 把权限链条拆成两段:云 IAM 管存储、UC 管 External Location 注册。两段脱节时报错只说权限不足,不告诉你是"没注册";另有 metadata sync 窗口期双写损坏分区元数据、Hive 视图需重写 DDL。
UC splits the permission chain in two: cloud IAM owns storage, UC owns External Location registration. When the two diverge, the error just says permission denied — never "you forgot to register." A metadata-sync window can also corrupt partition metadata with dual writes, and Hive views need DDL rewrites.
- 生产验证
—
Source 23: Aniket Soni, DEV circa Sep 2026 — field notes from migrating 200 production jobs: "don't let UC migration become your resume-generating event."
- 证据等级
`单方声音`,独立工程师具名生产复盘。
`Single voice`, named independent engineer's production postmortem.
Databricks 年份:2026
Airflow 侧点了 skip,Databricks 侧照跑不误:"幽灵执行"写重数据
单方声音
稳定与故障运维复杂度
- 一句话
`DatabricksWorkflowTaskGroup` 把整个 workflow 打包成一次 Jobs API 调用,Airflow 的 skip 信号传不进去——Airflow UI 显示 Skipped,Databricks 侧任务照常运行、写数据、完成。
`DatabricksWorkflowTaskGroup` packs the whole workflow into one Jobs API call, so Airflow's skip signal never reaches Databricks — Airflow UI shows Skipped while the Databricks side runs, writes data, and completes.
- 窄场景
Airflow 编排 Databricks workflow 且依赖条件跳过(回填、分支逻辑)。
Airflow-orchestrated Databricks workflows relying on conditional skips (backfills, branching logic).
- 机制
skip 只在 Airflow 侧生效(AirflowSkipException),Databricks 侧收到的是一次性整包调用,感知不到上游的跳过语义。结果:静默数据重复写入 + 审计轨迹失真(两边 UI 显示的状态互相矛盾)。
The skip only takes effect on the Airflow side (AirflowSkipException); Databricks receives a single bundled call and never sees the upstream skip semantics. Result: silent duplicate writes plus contradictory audit trails (the two UIs disagree about what happened).
- 生产验证
—
Source 24: Anil Reddaboina, Medium Jun 2026, in-depth postmortem;
Source 25: Apache Airflow issue #47024 (kind:bug), confirmed as a real problem by the community.
- 证据等级
`单方声音`,具名工程师复盘 + 开源 issue 印证。
`Single voice`, named engineer postmortem + open-source issue corroboration.
Databricks 年份:2026
LakeFlow 文件到达触发器:同名覆盖不触发、删过的路径"永远不报错也不触发"
单方声音
稳定与故障运维复杂度
- 一句话
从 Airflow 切到 LakeFlow 拿文件到达触发器做数据感知调度,结果同名文件覆盖不触发、指向已删路径的触发器"never errors and never fires"——静默不跑比报错更贵。
Moving from Airflow to LakeFlow for data-aware scheduling, file-arrival triggers silently miss: same-name overwrites don't fire, and triggers pointing at deleted paths "never error and never fire" — silent no-runs cost more than errors.
- 窄场景
用 LakeFlow file arrival triggers 替代 cron/Airflow 的团队。
Teams replacing cron/Airflow with LakeFlow file-arrival triggers.
- 机制
触发器多个 gotcha:同名覆盖写不触发、旧文件元数据过期后被修改会误触发、S3/GCS 上指向已删除路径的触发器既不报错也不触发;且切过去后失去 Airflow 的全局 DAG 视图,排查链路变长。
Multiple trigger gotchas: same-name overwrites don't fire; stale file metadata modified later misfires; triggers on S3/GCS pointing at deleted paths neither error nor fire. And you lose Airflow's global DAG view, lengthening the debugging chain.
- 生产验证
—
Source 28: official community thread, Sep 2026 user reply with a concrete gotcha list.
- 证据等级
`单方声音`,社区用户回帖。
`Single voice`, community user reply.
Databricks 年份:2026
notebook 习惯污染定时任务:一行 restartPython() 让 Job 随机失败且报错误导
单方声音
运维复杂度
- 一句话
notebook 里顺手写的 `dbutils.library.restartPython()` 被带进定时任务,在 Job 模式下杀死 Python 进程被判定为 Cancelled——报错却是 `AZURE_QUOTA_EXCEEDED_EXCEPTION`,排查方向全错。
A `dbutils.library.restartPython()` casually written in a notebook, carried into a scheduled job, kills the Python process in Job mode and gets judged Cancelled — but the error surfaces as `AZURE_QUOTA_EXCEEDED_EXCEPTION`, sending debugging in entirely the wrong direction.
- 窄场景
notebook 开发完直接转成 workflow 定时任务的团队;上云迁移后 job 随机失败。
Teams converting notebooks straight into workflow jobs; post-cloud-migration jobs failing randomly.
- 机制
交互式 notebook 的"重启 Python 进程换依赖"习惯,在 Job 运行时语义完全不同:进程被杀 → runner 判 Cancelled → 冒出来的却是 Azure 配额超限的报错。误导性报错让一次"代码习惯问题"被当成"云配额问题"查了半天。
The interactive habit of "restart the Python process to pick up dependencies" means something completely different under Job runtime semantics: the killed process → runner marks Cancelled → what bubbles up is an Azure quota-exceeded error. A "code habit" problem gets investigated as a "cloud quota" problem.
- 生产验证
—
Source 29: Tushar Sable, Medium Jun 2026 — post-migration random workflow failures, root-caused to the notebook line.
- 证据等级
`单方声音`,具名从业者生产复盘。
`Single voice`, named practitioner production postmortem.
Databricks 年份:2026
Shallow clone 的足枷:源表 VACUUM 会打坏 clone,UC 下还不能覆盖重建
单方声音
运维复杂度
- 一句话
Snowflake zero-copy clone 是元数据级 CoW;Databricks shallow clone 一串限制:源表一跑 VACUUM 就 FileNotFoundException、UC 下不能 CREATE OR REPLACE 覆盖、streaming 表不能做 clone 源——还有用户把 clone drop 掉后源表 VACUUM 永久损坏。
Snowflake's zero-copy clone is metadata-level CoW; Databricks shallow clone ships a list of caveats: VACUUM on the source raises FileNotFoundException, UC forbids CREATE OR REPLACE over a clone, streaming tables can't be clone sources — and one user permanently broke VACUUM on the source table by dropping the clone.
- 窄场景
用 clone 做测试环境、蓝绿发布的团队;UC 托管表。
Teams using clones for test environments and blue-green releases; UC managed tables.
- 机制
shallow clone 共享源表文件,VACUUM 清掉源文件即打坏 clone;UC 下覆盖重建 clone 不被允许;2023 年有用户经 SQL warehouse drop 掉 shallow clone 后源表 VACUUM 永久损坏(是否修复未见公开说明)。
Shallow clones share the source's files — VACUUM the source and the clone breaks; overwriting a clone via CREATE OR REPLACE is disallowed in UC; in 2023 a user dropping a shallow clone through a SQL warehouse left the source table's VACUUM permanently broken (no public word on whether the bug was fixed).
- 生产验证
—
Source 41: official community thread, Nov 2023 (repro notebook attached) — "the source table is broken beyond repair. Data reads and writes still work, but vacuum will remain forever broken.";
Official docs, "Delta clone" current limitations (verified Oct 2026).
- 证据等级
`单方声音`,用户 bug 报告 + 官方文档。
`Single voice`, user bug report + official docs.
Databricks 年份:2026
Delta Sharing 传 1.3TB 数据,在 60 分钟 token 过期处翻车
单方声音
性能问题运维复杂度
- 一句话
Delta Sharing recipient token 默认寿命 3600 秒——TB 级传输超 1 小时,整点处以 400 Bad Request 崩掉;修要改 workspace 参数且只对新 share 生效,老 recipient 得逐个手动 rotate。
Delta Sharing recipient tokens default to a 3,600-second lifetime — a transfer past the one-hour mark dies with a 400 Bad Request; fixing it means changing a workspace parameter that only applies to new shares, and manually rotating every existing recipient.
- 窄场景
经 Delta Sharing 做 TB 级跨平台迁移/分享(如 Databricks→BigQuery)。
TB-scale cross-platform migration/sharing over Delta Sharing (e.g., Databricks→BigQuery).
- 机制
token 寿命是 workspace 级参数 `delta_sharing_recipient_token_lifetime_in_seconds`,默认 3600;传输约 1 小时 40 分钟的任务在 60 分钟处被拒,整体失败。分享"开箱即用",上规模后续航要人肉调参。
Token lifetime is the workspace-level `delta_sharing_recipient_token_lifetime_in_seconds`, default 3600; a ~100-minute transfer gets rejected at minute 60 and fails wholesale. Sharing is "it just works" — at scale the keepalive is manual tuning.
- 生产验证
—
Source 44: Arjit Shukla, Jun 2026 load-test diary — a 1.3TB dataset's Direct Method failed wholesale at token expiry; the fix was raising the parameter to 14400 plus per-recipient rotation.
- 证据等级
`单方声音`,细节充分的压测复盘。
`Single voice`, detailed load-test postmortem.
Databricks 年份:2026
UC 治理边界=region:同云同 region 的两个库,分享还得走"对外协议"
单方声音
运维复杂度
- 一句话
UC metastore 是 region 级的,一个 workspace 只能挂一个——同 region 建了俩 metastore,跨库分享数据还得走 Delta Sharing,"对外分享协议"用在了"自家后院"。
A UC metastore is region-scoped; one workspace mounts exactly one — build two metastores in the same region and cross-metastore sharing still goes through Delta Sharing, the "external sharing protocol" used in your own backyard.
- 窄场景
prod/non-prod 分 metastore 的大企业;同 region 多业务线。
Large enterprises with prod/non-prod metastores; multiple business lines in one region.
- 机制
架构设计:单 region 单 metastore。Valcon 顾问实录明确建议"不要在同 region 建两个 metastore"——Client B 用双 metastore 后,workspace 只能挂其一,跨 metastore 分享必须走 Delta Sharing,治理复杂度显著上升。
By architecture: one metastore per region. Valcon's field notes explicitly advise "don't have two metastores in the same region" — their Client B did, and each workspace could only mount one, forcing all cross-metastore sharing through Delta Sharing and multiplying governance complexity.
- 生产验证
—
Source 46: Valcon consulting, Karlo Kotarac, Jan 2026 two-client field notes — "Don't have two metastores in the same region";
Official docs, "Unity Catalog best practices": "You can have only one metastore per region."
- 证据等级
`单方声音`,顾问客户实录 + 官方文档。
`Single voice`, consultant field notes + official docs.
Databricks 年份:2026
memory_limit 只是"建议":有一块缓存根本不归它管
单方声音
性能问题
- 一句话
`memory_limit` 设了 12GB,进程 RSS 照样涨到 14.8 GiB——`enable_external_file_cache` 这块内存绕过 buffer manager,内存上限管不着它。
With `memory_limit` set to 12GB, process RSS still climbed to 14.8 GiB — the `enable_external_file_cache` bypasses the buffer manager, and the memory limit cannot touch it.
- 窄场景
在容器/K8s 里给 DuckDB 设内存上限、跑"外部 Parquet/Arrow 文件只读一次"的 ETL/压缩/采集管道。
Teams running DuckDB inside containers/K8s with a memory cap, on "read each external Parquet/Arrow file exactly once" ETL/compaction/ingest pipelines.
- 机制
`enable_external_file_cache` 是外部文件数据的内存 LRU 缓存,它不走 buffer manager,因此不受 `memory_limit` 约束。对"每个文件只读一次就丢弃"的管道(压缩、合并、采集),这个缓存永远命中不了,只会单调累积内存。关掉它之后吞吐不变,内存占用回到正常水平。
`enable_external_file_cache` is an in-memory LRU over external file data that bypasses the buffer manager, so `memory_limit` does not bound it. On pipelines that read each file once and discard it (compaction, merge, ingest), the cache can never hit and only accumulates memory. Disabling it changed nothing about throughput and returned memory usage to normal.
- 生产验证
来源 3,dazzleduck-sql-server 项目 2026-09 提交记录——压缩器在 12GB `memory_limit` 下约 44 分钟涨到 14.8 GiB RSS;关闭该缓存后稳定在 0.7 GiB,周期时长不变。项目方把 `SET enable_external_file_cache = false` 写进默认配置并在注释里说明了原因。
Source 3, dazzleduck-sql-server project commit, Sep 2026 — the compactor grew to 14.8 GiB RSS in ~44 minutes under a 12GB `memory_limit`; with the cache disabled it stayed flat at ~0.7 GiB with unchanged cycle duration. The project made `SET enable_external_file_cache = false` a shipped default with a comment explaining why.
- 证据等级
`单方声音`,具名项目提交记录(带实测数据与已验证的修复方案)。
`Single voice`, named project commit record (measured data, verified fix).
DuckDB 年份:2026
二级 ART 索引检查点损坏:最危险的是静默错结果,不是崩溃
单方声音
稳定与故障
- 一句话
1.4.0 引入的回归——同一进程内两个独立引擎共开一个文件做 checkpoint,二级 ART 索引会被持久化写坏;之后按索引查永远返回 0 行,而全表扫描结果正常。
A regression since 1.4.0 — when two independent engines in one process co-open a file and one checkpoints it, a secondary ART index is persisted corrupt; index lookups then return 0 rows forever while full scans stay correct.
- 窄场景
同一进程内加载了两份 libduckdb(混合 C API/Python 或插件场景常见)、且存在 `CREATE INDEX` 二级索引的库;DuckDB ≥ 1.4.0。
Libraries loading two copies of libduckdb in one process (common in mixed C API/Python or plugin setups) with `CREATE INDEX` secondary indexes; DuckDB >= 1.4.0.
- 机制
DuckDB 的文件锁按 PID 粒度,同一进程内两份独立引擎可以绕过锁共开同一文件;第二个引擎对未 checkpoint 的 WAL 做 checkpoint 时,二级 ART 索引被错误序列化。损坏写入文件后永久生效:索引扫描返回错误结果、全表扫描正常——是"静默错结果"而非崩溃。若写者之后干净关闭会用正确的内存索引重写文件"自愈",掩盖问题;写者异常退出则损坏永久留存。
DuckDB's file lock is per-PID, so two independent engines loaded in the same process can co-open one file and bypass the lock; when the second engine checkpoints a pending WAL, the secondary ART index is serialized incorrectly. The corruption is persisted: index scans return wrong results, full scans return correct ones — silent wrong results, not a crash. If the writer later closes cleanly it rewrites the file from its correct in-memory index and "heals" it, masking the problem; if the writer dies abnormally the corruption is permanent.
- 生产验证
来源 4,GitHub issue duckdb/duckdb#23788(2026)——Paradigm4 工程师 Rares Vernica 具名报告,附完整复现脚本与版本矩阵:1.1.3/1.3.2 正常,1.4.1/1.5.1/1.5.4 均损坏。
Source 4, GitHub issue duckdb/duckdb#23788 (2026) — named report by Paradigm4 engineer Rares Vernica with a complete repro script and version matrix: 1.1.3/1.3.2 unaffected, 1.4.1/1.5.1/1.5.4 all corrupt.
- 证据等级
`单方声音`,具名工程师报告(复现脚本完整、版本边界清晰)。
`Single voice`, named engineer report (complete repro script, clear version boundary).
DuckDB 年份:2026
存储格式单向升级:新版打开旧文件,旧版就永远打不开了
单方声音
升级迁移
- 一句话
新版 DuckDB 打开旧文件会不可逆地升级存储格式——v1.5 写的库 v2.0 能读,v2.0 写的库 v1.5 只能报错,没有降级路径。
Opening an old database file with a newer DuckDB irreversibly upgrades the storage format — a v2.0-written file is unreadable by v1.5, with no downgrade path.
- 窄场景
多组件/多版本共存的环境(CI 与生产 DuckDB 版本不一致、桌面工具与服务端混用);把 `.duckdb` 文件当作可分发产物的团队。
Environments with mixed versions across components (CI vs. production DuckDB versions, desktop tools mixed with servers); teams treating `.duckdb` files as distributable artifacts.
- 机制
存储格式版本号随版本演进(v1.5 写 header version 64、读 64–68;v2.0 写 69、读 64–69)。升级是单向的:新版打开即升级,官方唯一的跨版本迁移手段是 EXPORT/IMPORT 全量导出导入。v1.4 起虽有 LTS 线,但 LTS 只保证同一大版本内的格式稳定。
The storage format version advances with releases (v1.5 writes header version 64, reads 64–68; v2.0 writes 69, reads 64–69). The upgrade is one-way: the new version upgrades on open, and the only supported cross-version migration is full EXPORT/IMPORT. The v1.4 LTS line only guarantees format stability within the same major line.
- 生产验证
来源 5,DataverseDuck 项目(2026-09)实测记录——明确把"缓存单向升级"列为已知问题:v2 链接的客户端打开旧缓存后,该文件即超出 v1.5 客户端可读范围,打开报错 `Trying to read a database file with version number 69, but we can only read versions between 64 and 68`。
Source 5, DataverseDuck project (Sep 2026) measured records — lists "cache upgrades one way" as a known issue: once a v2-linked client opens an old cache, the file is beyond what a v1.5 client can read, failing with `Trying to read a database file with version number 69, but we can only read versions between 64 and 68`.
- 证据等级
`单方声音`,具名项目实测记录(版本边界精确)。
`Single voice`, named project measurement record (precise version boundary).
DuckDB 年份:2026
DuckLabs 被 AWS 收购:许可证没变,但写代码的人换了老板
单方声音
生态与信任
- 一句话
DuckDB 本体仍是 MIT + 独立基金会,但写代码、定路线的人现在归 AWS——社区担心的是路线图,不是许可证。
DuckDB itself stays MIT under an independent foundation, but the people who write the code and set the roadmap now report to AWS — the community's worry is the roadmap, not the license.
- 窄场景
把 DuckDB 嵌入自家产品/管线、赌的是"轻量本地优先"这条路线的团队;与 AWS 有竞争关系的厂商生态用户。
Teams embedding DuckDB into their own products/pipelines, betting on the "lightweight, local-first" trajectory; users in ecosystems competing with AWS.
- 机制
DuckLabs(约 30 人,DuckDB 原作者创办)2026-08-26 签署被 AWS 收购的最终协议,9 月初交割;IP 与商标留在独立的 DuckDB Foundation,MIT 许可证不变。但历史上绝大多数贡献与路线决策来自 DuckLabs,外部贡献者提个简单修复 CI 就要跑五小时,社区实质影响力本就有限。收购前九天发布的 DuckDB 2.0 预览主打 server 模式(Quack)与 S3 对象存储加速——正是 AWS 想要的方向。
DuckLabs (~30 people, founded by DuckDB's creators) signed a definitive acquisition agreement with AWS on Aug 26, 2026, closing early September; IP and trademarks stay with the independent DuckDB Foundation under the MIT license. But the vast majority of contributions and roadmap decisions have historically come from DuckLabs, and external contributors already faced friction (a five-hour CI run for simple fixes). The DuckDB 2.0 preview published nine days before the announcement led with server mode (Quack) and S3 object-storage acceleration — exactly the direction AWS wants.
- 生产验证
来源 6,独立数据教育者 Walter Shields(LinkedIn Learning 讲师,2026-08-31)梳理 HN(1058 分、300+ 评论)与 Reddit 社区反应:一方庆祝创始人五年 bootstrapping 的成功,另一方担忧核心团队加入"靠云计算与存储赚钱"的公司后,路线图从"笔记本上的分析师"转向"AWS 的企业客户"。MotherDuck CEO Jordan Tigani 亦公开表示"他们收购 DuckLabs 不是因为热爱开源"。
Source 6, independent data educator Walter Shields (LinkedIn Learning instructor, Aug 31 2026) summarizing the HN (1,058 points, 300+ comments) and Reddit reaction: one side celebrated the founders' five-year bootstrapped success; the other worried that a core team joining a company "whose business model depends on cloud compute and storage" shifts the roadmap from "the analyst on a laptop" to "AWS's enterprise customers." MotherDuck CEO Jordan Tigani publicly said "they're not acquiring DuckLabs just because they love open source."
- 证据等级
`单方声音`,独立评论者对 HN/Reddit 社区反应的梳理(引注了 GeekWire/The Register/SiliconANGLE 报道与 HN 高赞讨论)。
`Single voice`, an independent commentator's synthesis of the HN/Reddit community reaction (citing GeekWire/The Register/SiliconANGLE coverage and the top HN thread).
- 备注
交割时基金会技术咨询委员会的实际治理结构尚未确定,后续值得跟进;若治理落地且社区确有话语权,本卡应更新。
at closing, the Foundation's technical advisory board governance was still undecided — worth following up; if it lands with real community influence, this card should be updated.
DuckDB 年份:2026
拆掉 DAX:缓存成了架构里唯一的单点故障
单方声音
稳定与故障运维复杂度
- 一句话
DAX 缓存成了架构里最脆弱的一环:扩容要手工、打满后恢复要 30 分钟;拆掉后系统反而更稳、更便宜。
DAX became the most fragile link in the architecture: manual scaling, ~30-minute recovery after saturation; removing it made the system more reliable and cheaper.
- 窄场景
读密集但流量尖刺、DynamoDB 本体其实扛得住的业务;把 DAX 当"无脑加速器"引入的团队。
Read-heavy but spiky workloads where DynamoDB itself could handle the load; teams that adopted DAX as a "drop-in accelerator".
- 机制
DynamoDB 表级自动扩缩容近乎无缝,DAX 却是按节点小时计费的固定集群——扩缩容全手工,重建/rebalance 期间缓存失效;DAX 只支持最终一致读,不支持 TransactWriteItems、PartiQL、Streams;流量上涨时 DAX 先到容量上限,延迟反升,成为比数据库更脆弱的一环。
DynamoDB table auto-scaling is near-seamless, but DAX is a fixed cluster billed per node-hour — scaling is fully manual, and rebuild/rebalance windows invalidate the cache; DAX serves eventually consistent reads only and supports neither TransactWriteItems, PartiQL, nor Streams; under rising traffic DAX hits its capacity ceiling first, increasing latency instead of reducing it.
- 生产验证
来源 5,Muhammad Ali 2026-04——为降读延迟引入 DAX(Client → DAX → DynamoDB),流量上涨后 DAX 到容量上限、缓存性能退化、延迟不降反升;故障后约 30 分钟才恢复健康(期间 API 延迟飙升、错误率上升);评估发现性能收益边际、DynamoDB 本体可扛住流量,拆除后省 ~$600,架构简化为 Client → DynamoDB。
Source 5, Muhammad Ali, Apr 2026 — introduced DAX to cut read latency (Client → DAX → DynamoDB); as traffic grew the DAX cluster hit capacity limits, cache performance degraded, latency rose instead of falling; after an incident it took ~30 minutes to return to healthy (API latency spiked, error rates rose); evaluation found marginal gains and that DynamoDB alone handled traffic, so they removed DAX, saved ~$600, and simplified to Client → DynamoDB.
- 证据等级
`单方声音`,具名生产复盘(恢复时长、金额、决策过程完整)。
`Single voice`, named production postmortem (recovery time, dollar amount, decision process complete).
Amazon DynamoDB 年份:2026
Global Tables 双活:LWW 静默吃掉 847 条客户资料更新
单方声音
稳定与故障
- 一句话
Global Tables 的冲突解决是"最后写入胜出":一次 DNS 抖动让双区域同时写了 4 分钟,847 条客户资料更新静默消失,日志里没有任何错误。
Global Tables resolves conflicts with last-writer-wins: one DNS flap caused 4 minutes of dual-region writes, 847 customer profile updates vanished silently, and the logs showed zero errors.
- 窄场景
Global Tables 做多活/温备 + Route53 故障转移的团队;把 Global Tables 当"开箱即用灾备"的团队。
Teams using Global Tables for active-active/warm-standby with Route53 failover; teams treating Global Tables as out-of-the-box disaster recovery.
- 机制
默认 MREC 模式为异步复制,且复制延迟没有 SLA;两区域同时写同一 item 即冲突,按写入时间戳 last-writer-wins——没有版本向量、没有冲突记录(CloudWatch/CloudTrail 均无);时钟漂移下"时间戳更晚"未必是"实际更晚"的写入;MRSC 强一致模式代价是更高的写延迟且不支持事务。
The default MREC mode replicates asynchronously with no SLA on replication lag; concurrent writes to the same item in two regions conflict and resolve by write-timestamp last-writer-wins — no vector clocks, no conflict records (nothing in CloudWatch or CloudTrail); under clock drift, the "later" timestamp is not necessarily the actually-later write; the MRSC strong-consistency mode costs higher write latency and does not support transactions.
- 生产验证
来源 9,Illya Yalovoy 2026-06——温备架构下局部区域降级导致 Route53 健康检查抖动,约 4 分钟内两区域同时写入同一批用户记录;复制收敛后 847 条更新被静默覆盖("overwritten by whichever region happened to have the later timestamp"),应用日志零错误、DynamoDB 零异常;排查花了两天,原话 "every observability signal said the system was healthy"(部分付费墙,仅前半可见)。
Source 9, Illya Yalovoy, Jun 2026 — a partial regional degradation flapped Route53 health checks; for ~4 minutes both regions wrote the same user records; after replication converged, 847 updates had been silently overwritten ("overwritten by whichever region happened to have the later timestamp"); zero application-log errors, zero DynamoDB exceptions; investigation took two days: "every observability signal said the system was healthy" (partly paywalled; visible portion verified).
- 证据等级
`单方声音`,个人博客(细节充分:847 条、4 分钟、两天排查;部分付费墙,已如实标注)。
`Single voice`, personal blog (detailed: 847 items, 4 minutes, two-day investigation; partial paywall disclosed as-is).
Amazon DynamoDB 年份:2026
S1 级支持也要先交"作业":要的是快速解决,给的是"请提供日志"
单方声音
稳定与故障
- 一句话
连最高优先级的 S1 工单,支持流程的第一步也是让用户收集上传诊断包,而不是先给止血方案。
Even for top-priority S1 tickets, the support process's first step is having the customer collect and upload a diagnostics bundle — not a containment plan.
- 窄场景
买了 EDB 24x7 支持、指望"出事有人兜底"的生产团队;真正遇到 S1 停机、按合同期待快速响应的场合。
Production teams paying for EDB 24x7 support who expect "someone has our back"; real S1 outages where the contract promises rapid response.
- 机制
厂商支持的标准分诊流程要求"证据先行":Lasso 等诊断工具收集系统元数据、日志、配置,打包上传后支持工程师才开始分析。这套流程为降低支持成本而设计,但在 S1 场景下把"收集证据"放在了"恢复服务"之前——用户感知到的就是"要日志而不是要方案"。另一位用户在评价中也侧面印证了这套流程的存在(onboarding 即要求安装 Lasso 诊断工具)。
Vendor support triage is built "evidence-first": Lasso-style diagnostic tools collect system metadata, logs, and configs into an upload bundle before engineers start analysis. That workflow is designed to lower support cost, but in an S1 scenario it puts evidence collection ahead of service recovery — what the user perceives is "give us logs, not a plan." Another reviewer independently confirms the workflow exists (onboarding requires installing the Lasso diagnostic tool).
- 生产验证
来源 3:Rakesh kumar b.(Technical Authority expert,企业用户,2026-07-31,5/5 好评中仍提)原话:"Sometimes in S1 case they asking for logs and all insted of quick resolution"(S1 工单里他们有时要日志这日志那,而不是快速解决)。
Source 3: Rakesh kumar b. (technical authority expert, enterprise user, Jul 31 2026, inside an otherwise 5/5 review): "Sometimes in S1 case they asking for logs and all insted of quick resolution".
- 证据等级
`单方声音`,具名评价(企业用户,Technical Authority expert,细节具体:S1 场景+日志要求)。
`Single voice`, named review (enterprise user, technical authority expert; specific: S1 scenario + log demands).
EDB Postgres / EDB Postgres Advanced Server(EPAS) 年份:2026
Galera:官方文档推荐的配置,丢已提交事务
单方声音
稳定与故障
- 一句话
Jepsen 按官方"更安全、推荐"的配置(`innodb_flush_log_at_trx_commit=0`)测试 Galera,节点相继崩溃时已确认提交的事务照丢不误。
Jepsen tested Galera with the officially "safer, recommended" setting (`innodb_flush_log_at_trx_commit=0`) — when nodes crash one after another, confirmed-committed transactions still get lost.
- 窄场景
多节点同时/相继故障(断电、洪水、网络 bug 等关联故障);以及无故障健康集群上的普通读写事务。
Simultaneous or sequential multi-node failures (power, flooding, correlated network bugs); and ordinary read-write traffic on a healthy cluster.
- 机制
官方文档称 `innodb_flush_log_at_trx_commit=0` "是 Galera 下更安全、推荐的选项,因为不一致总能从其他节点恢复"——这个"总能"的前提是故障不关联;关联故障下未刷盘的提交在所有节点上一起消失(MDEV-38974)。即使设为 1,进程崩溃+网络分区下仍偶发丢失已提交事务(MDEV-38976)。更糟的是健康集群也测出 P4 Lost Update(MDEV-38977)与常态 Stale Read(MDEV-38999)——而官方口径是"无丢失事务"、"隔离级别介于 Serializable 与 Repeatable Read 之间"。
The official docs call `innodb_flush_log_at_trx_commit=0` "a safer, recommended option with Galera Cluster, since inconsistency can always be recovered from another node" — that "always" assumes uncorrelated failures; under correlated failures, unflushed commits vanish on all nodes together (MDEV-38974). Even with `=1`, process crash plus network partition still occasionally loses committed transactions (MDEV-38976). Worse, a healthy cluster also produced P4 Lost Update (MDEV-38977) and routine Stale Reads (MDEV-38999) — while the official line is "no lost transactions" and "isolation between Serializable and Repeatable Read".
- 生产验证
来源 4:Jepsen 2026,MariaDB Galera Cluster 12.1.2(Galera 26.4.13–26.4.25),声明独立无偿测试——关联崩溃下 1 分钟测试丢 9 个已确认提交的值;`=1` 配置下数小时出现一次、一次丢约 19 秒写入;健康集群 P4 与 Stale Read 可复现;已提交 MDEV-38974/38976/38977/38999。
Source 4: Jepsen 2026, MariaDB Galera Cluster 12.1.2 (Galera 26.4.13–26.4.25), declared independent and uncompensated testing — correlated crashes lost 9 acknowledged committed values in a one-minute test; the `=1` configuration lost about 19 seconds of writes once every several hours; P4 and Stale Reads reproduced on a healthy cluster; MDEV-38974/38976/38977/38999 filed.
- 证据等级
`单方声音`,独立第三方实验室测试(非用户生产投诉,但细节充分、可复现;已如实标注)。
`Single-source`, independent third-party lab test (not a user production incident, but fully detailed and reproducible; labeled honestly).
- 备注
与现有 [避坑] 卡部分重叠(档案吐槽清单已有 Galera 认证冲突/流控行),本卡新增"已提交事务丢失"的正确性质疑角度。
Partially overlaps the existing profile's "pitfall" list (Galera certification-conflict/flow-control row); this card adds the committed-transaction-loss correctness angle.
MariaDB 年份:2026
压实把去重结果反转了:数据静默"回到"第一次写入
单方声音
稳定与故障
- 一句话
同一批 insert 里对同一个主键写了 5 次,查出来是对的(最后一次写入胜出);等自动压实跑完一遍,再查就"回到"了第一次写入——全程无报错、无日志。
Insert the same primary key 5 times in one batch and queries correctly return the last write; after automatic compaction runs once, queries return the first write instead — with zero errors and zero log output.
- 窄场景
流式管道 at-least-once 投递、客户端批量重试、ETL 分块未去重的写入模式;单 batch 内出现重复主键的集合;报告在 Milvus v2.6.20 / v2.6.21(standalone)上复现。
At-least-once streaming pipelines, client batch retries, or ETL jobs that chunk data without dedup; collections where duplicate PKs land in a single batch; reproduced on Milvus v2.6.20 / v2.6.21 (standalone).
- 机制
同一 growing segment 内的行共享同一时间戳。查询层对同 segment 内重复主键按"最后一次出现"去重(last-write-wins),而压实在合并 segment、遇到等时间戳重复主键时取了"第一次出现"。两层去重语义不一致 → 压实后可观察到的"当前值"静默回退到最早版本。报告者另指出:重复主键会留下多个"幽灵"向量条目——每个历史向量都可被检索命中、`count(*)` 膨胀,压实不清理;只有改用 upsert(delete+insert)能同时避开回退与幽灵条目。
Rows inside one growing segment share the same timestamp. The query layer dedups duplicate PKs within a segment keeping the *last* occurrence (last-write-wins), but compaction, when merging a segment and meeting equal-timestamp duplicate PKs, keeps the *first* occurrence. The two layers disagree on dedup semantics, so the observable "current value" silently reverts to the earliest version after compaction. The reporter additionally notes duplicate PKs leave multiple "ghost" vector entries — every historical vector stays searchable, `count(*)` stays inflated, and compaction never cleans them; only switching to upsert (delete+insert) avoids both the revert and the ghosts.
- 生产验证
来源 4,2026-07 GitHub issue——REST API 完整复现:建集合 → 单 batch 插入 5 条 pk=1(order 1..5)→ flush → compact → 再查。压实前查到 `order=5, label="fifth_LAST"`(正确),压实后查到 `order=1, label="first"`(数据回退),`count(*)=5`(幽灵条目未清理);全程返回 code 0,无服务端错误日志。报告者判定:"静默、自动触发(auto-compaction 定期跑)、压实删掉原 segment 后不可恢复"。
Source 4, Jul 2026 GitHub issue — full REST API reproduction: create collection, insert 5 rows with pk=1 in one batch (order 1..5), flush, compact, query again. Before compaction the query returned `order=5, label="fifth_LAST"` (correct); after compaction it returned `order=1, label="first"` (reverted), with `count(*)=5` (ghosts not cleaned); every request returned code 0 with no server-side error logs. The reporter classifies it as "silent, automatically triggered (auto-compaction runs periodically), and non-recoverable once the original segment is deleted."
- 证据等级
`单方声音`,GitHub 社区 bug 报告(细节充分:完整 curl 复现、压实前后输出对比、根因定位到时间戳相同的去重分支)。
`Single voice`, community bug report on GitHub (detailed: complete curl reproduction, before/after output comparison, root-cause traced to the equal-timestamp dedup branch).
- 备注
报告者在 2026-07-24 于 v2.6.21 上二次确认可复现;是否已在后续版本修复,未找到公开记录,未作"已修复"标注。本卡主题可能与本站 [避坑] 卡重叠。
the reporter re-confirmed the reproduction on v2.6.21 on 2026-07-24; no public record of a fix in later versions was found, so no "fixed in" label. May overlap an existing [Pitfall] card on this site.
Milvus 年份:2026
WiredTiger 逐出线程被内核软中断饿死:CPU 看似空闲,写延迟却飙升
单方声音
性能问题运维复杂度
- 一句话
WiredTiger 后台逐出线程被内核网络软中断(ksoftirqd)抢占 CPU,导致缓存压力上升、读写延迟飙升,而常规监控指标(总 CPU、磁盘、内存)全部正常,极难定位。
WiredTiger's background eviction threads were preempted by kernel network softIRQs (ksoftirqd), driving cache pressure up and read/write latency with it — while every conventional metric (aggregate CPU, disk, memory) looked healthy, making it extremely hard to diagnose.
- 窄场景
写密集、分片集群、高网络吞吐的自建 MongoDB;逐出/检查点线程被调度到软中断集中的 CPU 核上时。
Write-heavy, sharded, self-hosted MongoDB with high network throughput; when eviction/checkpoint threads get scheduled onto CPU cores saturated with softIRQs.
- 机制
WiredTiger 依靠后台逐出线程持续回收缓存;当逐出线程长期得不到调度,缓存占用缓慢爬升,应用线程被迫亲自执行逐出,读写延迟被逐出开销污染。MongoDB 不会根据内核级争用动态重排后台线程,且对此类争用几乎没有可观测性——聚合 CPU 指标会系统性撒谎(个别核跑满、其余核空闲)。
WiredTiger depends on background eviction threads to continuously reclaim cache; when they stop getting scheduled, cache occupancy creeps up until application threads are forced to perform eviction themselves, polluting read/write latency with eviction overhead. MongoDB never rebalances background threads based on kernel-level contention, and offers almost no observability into it — aggregate CPU metrics systematically lie (a few cores pegged, the rest idle).
- 生产验证
Gojek 工程博客 Yuvaraj A《When MongoDB Isn't Slow — It's Starved: A Performance Mystery》(2026-01-12);用 /proc/interrupts 与 /proc/softirqs 实证网络软中断在少数 CPU 核上堆积、逐出线程被抢占。
Gojek engineering blog, Yuvaraj A, "When MongoDB Isn't Slow — It's Starved: A Performance Mystery" (2026-01-12); proven with /proc/interrupts and /proc/softirqs showing network softIRQs piling onto a few CPU cores and eviction threads being preempted.
MongoDB 年份:2026
跨大版本升级没有直达路:老系统只能逐版爬升或逻辑迁移
单方声音
运维复杂度升级迁移
- 一句话
MongoDB 不支持跨大版本直接升级,老版本只能逐版爬升(每个大版本一次停机窗口)或走 dump/restore 逻辑迁移;配套的 Mongoose/驱动升级还带来大量应用层 breaking。
MongoDB does not support direct upgrades across major versions — old deployments must either climb every major version one at a time (one downtime window each) or do a dump/restore logical migration; the accompanying Mongoose/driver upgrades add a pile of application-level breakages on top.
- 窄场景
多年未升级的老系统(3.x/4.x)、单节点部署、无预发布环境。
Legacy systems many versions behind (3.x/4.x), single-node deployments, no staging environment.
- 机制
featureCompatibilityVersion 链式约束要求 3.4→3.6→…→8.0 逐版升级;4.2 起强制 WiredTiger(MMAPv1 被移除);6.0 起 mongo shell 被 mongosh 取代;Mongoose 7 把 ObjectId 改为 class(必须 new 调用)、强制 model 注册顺序。官方升级路径与老旧 OS/单节点的现实脱节。
The featureCompatibilityVersion chain forces 3.4 → 3.6 → … → 8.0 step by step; WiredTiger became mandatory in 4.2 (MMAPv1 removed); the mongo shell was replaced by mongosh in 6.0; Mongoose 7 turned ObjectId into a class (must be called with new) and enforces model registration order. The official upgrade path simply does not match the reality of old OSes and single nodes.
- 生产验证
Blen Redwan《How I Upgraded MongoDB 3.4 to 6 on a Legacy Production System》(2026-07);14 万行 Express 生产单体,MongoDB 3.4.23 → 6,Ubuntu 16.04,逐条列出 breaking 清单与取舍(最终选逻辑迁移而非 7 次停机窗口的逐版爬升)。
Blen Redwan, "How I Upgraded MongoDB 3.4 to 6 on a Legacy Production System" (2026-07); a 140k-line Express production monolith, MongoDB 3.4.23 → 6 on Ubuntu 16.04, with an itemized breakage list and the final call (logical migration instead of 7 separate downtime windows).
- 证据等级
`单方声音`(具名个人生产复盘,版本/规模/报错俱全)
—
MongoDB 年份:2026
mongorestore --nsInclude 对 gzip 归档静默失效,归档内容还无法查看
单方声音
运维复杂度
- 一句话
mongorestore 的 --nsInclude/--nsFrom/--nsTo 对 `--gzip --archive` 归档静默失效——会恢复归档内全部集合;且归档是不透明二进制流,没有任何内置命令可查看其中包含哪些集合。
mongorestore's --nsInclude/--nsFrom/--nsTo are silently ignored when restoring from a `--gzip --archive` — it restores every collection in the archive; and since an archive is an opaque binary stream, no built-in command can list which collections it contains.
- 窄场景
用 `--archive --gzip` 做生产备份流水线、需要按集合选择性恢复的自建用户。
Self-hosted users whose production backup pipeline uses `--archive --gzip` and who need selective per-collection restores.
- 机制
目录式 dump 每个集合是独立 .bson 文件,恢复时可跳过;archive 是多路复用的单一二进制流,mongorestore 只能顺序读取、无法 seek,命名空间过滤在该路径下不可靠。JIRA TOOLS-2023(选择性逻辑改进)已 open 六年以上。雪上加霜的是没有任何 --list/--inspect 语义能枚举归档内容,作者被迫在每次备份时额外保存集合清单文件。
A directory dump keeps one .bson file per collection, so restore can skip files; an archive is a single multiplexed binary stream that mongorestore can only read sequentially — it cannot seek, so namespace filtering is unreliable on that path. JIRA TOOLS-2023 (selectivity-logic improvement) has been open for over six years. Worse, there is no --list/--inspect semantic to enumerate an archive's contents, so the author now saves a collection-inventory file alongside every backup.
- 生产验证
thedecipherist《MongoDB Backups》(2026-02-25);作者自建 MongoDB 生产十年(34 个电商站、12GB 库、约 3650 次日备零丢失),实录一次只想恢复 products/orders 却恢复了 130+ 全集合的经历。
thedecipherist, "MongoDB Backups" (2026-02-25); the author has run self-hosted MongoDB in production for a decade (34 e-commerce sites, 12 GB database, ~3,650 daily backups with zero loss) and recounts asking for just products/orders and getting all 130+ collections restored instead.
MongoDB 年份:2026
社区版 Search(mongot):索引卡 PENDING 无报错,状态字段"撒谎",查不存在的索引名静默返回空
单方声音
性能问题运维复杂度
- 一句话
社区版 Search(mongot)问题频出:磁盘超 ~89% 时索引永久 PENDING 且客户端无任何报错;$listSearchIndexes 汇总状态取"历史最差"(含已死主机);查询不存在的索引名静默返回空结果;$search 简单词查询比老 $text 索引慢 10-14 倍;mongot 空闲常驻内存是 mongod 的 4 倍多。
Community-edition Search (mongot) is full of sharp edges: indexes stay PENDING forever past ~89% disk usage with zero client-side error; $listSearchIndexes reports the worst historical status (including dead hosts); querying a nonexistent index name silently returns empty results; $search on plain term queries is 10–14x slower than the legacy $text index; and mongot's idle memory footprint is 4x+ mongod's.
- 窄场景
自建 Community Server 9.0.2 + mongot-community 1.70.5,磁盘配额受限环境。
Self-managed Community Server 9.0.2 + mongot-community 1.70.5 in disk-quota-constrained environments.
- 机制
mongot 的磁盘保护阈值(~90% 停、~85% 恢复)按底层设备全容量计算,与容器内 df 读数可差 60 个百分点;状态聚合逻辑取所有历史主机的最差值(含已不存在的主机);不存在的索引名只在 mongot 日志里 WARN,客户端无法区分"拼写错误"与"真无结果";$search 走 gRPC 跨进程,简单词项查询多一跳。
mongot's disk-protection thresholds (~90% stop, ~85% resume) are computed against the underlying device's full capacity, which can differ from in-container df readings by 60 percentage points; the aggregated status takes the worst value across all hosts it has ever seen, including ones that no longer exist; a nonexistent index name only produces a WARN in mongot's own log, so clients cannot distinguish "typo" from "genuinely empty"; $search crosses a gRPC hop to a separate process, adding a round trip that plain term lookups don't need.
- 生产验证
dev.to alexgeorgiev17《MongoDB Community Edition's new search index stayed pending above 89% disk use》(2026-09);5 万文档实测,min/median/p95/max 对比数据俱全($search 中位 9.36ms vs $text 0.67ms)。
dev.to, alexgeorgiev17, "MongoDB Community Edition's new search index stayed pending above 89% disk use" (2026-09); 50k-document benchmark with full min/median/p95/max numbers ($search median 9.36 ms vs $text 0.67 ms).
- 证据等级
`单方声音`(具名个人实测复盘,版本/数据俱全)
—
MongoDB 年份:2026
utf8mb3 遗毒:emoji 被静默截断,迁 utf8mb4 又撞三连坑
单方声音
运维复杂度升级迁移
- 一句话
MySQL 的 `utf8` 从来不是真正的 UTF-8(只是 utf8mb3 别名);存 emoji 等四字节字符在非严格模式下被静默截断,而迁往 utf8mb4 又会撞上 767 字节索引上限、连接字符集、CONVERT TO 锁表三连坑。
—
- 窄场景
2010 年代建库、字符集为 utf8(即 utf8mb3)的老业务;用户昵称、商品目录等含 emoji/国际文本的场景。
Databases created in the 2010s with charset utf8 (i.e. utf8mb3); user nicknames, catalogs, and other emoji/international-text-heavy workloads.
- 机制
utf8mb3 每字符最多 3 字节,四字节字符在非严格模式下静默截断(无报错,数周后才在客服工单里发现);迁移时索引按"最宽字符"计字节,VARCHAR(255) 从 765 字节(255×3)涨到 1020 字节(255×4),老实例/旧行格式下触发 ERROR 1071;字符集是连接级协商的,只改表不改 SET NAMES 照样乱码;CONVERT TO 是全表重写。
utf8mb3 allows at most 3 bytes per character, so four-byte characters are silently truncated in non-strict mode (no error; discovered weeks later in support tickets); on migration, indexes are sized by widest possible character, so VARCHAR(255) jumps from 765 bytes (255×3) to 1020 bytes (255×4), tripping ERROR 1071 on older instances/row formats; charset is negotiated per connection — converting tables without fixing SET NAMES still corrupts data; and CONVERT TO CHARACTER SET is a full table rewrite.
- 生产验证
GitHub 个人博客(2026-06-13):Magento 2 真实店铺迁移,客户 display name 带 emoji 被静默截断且日志无报错;ERROR 1071(Specified key was too long; max key length is 767 bytes)复现;三处必须联动(表、my.cnf server 默认、应用连接字符集),缺一则"mysql CLI 看正常、店面继续乱码";CONVERT TO 全表重写需维护窗口或 pt-osc/gh-ost;半库转换导致 Illegal mix of collations。
Personal blog (Jun 13, 2026): a real Magento 2 store migration where a customer's emoji-bearing display name was silently truncated with nothing in the logs; reproduced ERROR 1071 (Specified key was too long; max key length is 767 bytes); three changes must move together (tables, my.cnf server defaults, application connection charset) — miss one and "the mysql CLI looks fine while the storefront keeps corrupting data"; CONVERT TO needs a maintenance window or pt-osc/gh-ost; half-converted schemas throw Illegal mix of collations.
- 证据等级
`单方声音`,来源性质:个人博客(细节充分,含报错原文与复现路径)。
—
- 备注
与本站现有 [避坑] 卡提及的"utf8mb3 包袱"部分重叠;本卡角度为具体迁移事故(静默截断 + 迁移三连坑),而非性能。
Partially overlaps the existing site card's mention of the "utf8mb3 burden"; this card's angle is a concrete migration incident (silent truncation + the three migration traps), not performance.
MySQL 年份:2026
异步复制主从延迟:从库跑个报表,读到"刚下的订单不存在"
单方声音
性能问题稳定与故障
- 一句话
从库默认单线程回放又被拿去跑分析查询,一次主库 DDL 引发的写入突发就能让 Seconds_Behind_Master 冲到 1800 秒以上,读从库的业务看到"用户刚提交的订单不存在"。
—
- 窄场景
读写分离(写主读从)、从库被复用跑报表/分析查询、未开并行复制的 8.0 集群。
Read-write splitting (writes to primary, reads from replica), replicas reused for reporting/analytics queries, 8.0 clusters without parallel replication.
- 机制
主库 ALTER 拿 MDL 排他锁 → 写被阻塞排队 → DDL 完成后突发提交,binlog 密集;从库 I/O 线程照单拉取(relay log 涨到 4.7GB),但 SQL 线程默认单线程(replica_parallel_workers=0)且与分析查询争 CPU,追赶速度跟不上,形成"越慢越积、越积越慢";Seconds_Behind_Master 只反映 SQL 线程正在执行的事件时间戳、不反映 relay log 积压量,故障早期具有欺骗性。
ALTER on the primary takes an exclusive MDL → writes queue behind it → burst-commit after the DDL, producing dense binlog; the replica's I/O thread dutifully pulls everything (relay log grew to 4.7 GB), but the SQL thread is single-threaded by default (replica_parallel_workers=0) and fights the analytics query for CPU, so catch-up never keeps up — "the slower it gets, the more it accumulates". Seconds_Behind_Master only reflects the timestamp of the event the SQL thread is executing, not relay-log backlog, so it is deceptive early in the incident.
- 生产验证
个人博客故障复盘(约 2026-06):MySQL 8.0.32,1 主 1 从,orders 表约 1000 万行;14:00 主库 DDL → 14:25 Seconds_Behind_Master>1800、Relay_Log_Space 4.7GB → 定位到从库上一个跑了 812 秒的 GROUP BY 分析查询 → KILL 后 15 分钟恢复;附 SHOW REPLICA STATUS / PROCESSLIST 实录。
Personal blog incident postmortem (circa Jun 2026): MySQL 8.0.32, one primary + one replica, ~10M-row orders table; 14:00 DDL on primary → 14:25 Seconds_Behind_Master > 1800, Relay_Log_Space 4.7 GB → root cause: an 812-second GROUP BY analytics query running on the replica → KILLed, full recovery in 15 minutes; includes SHOW REPLICA STATUS / PROCESSLIST transcripts.
- 证据等级
`单方声音`,来源性质:匿名个人博客故障复盘(时间线与命令实录细节充分,但作者身份不明,页面含 SEO 式关键词表、有 AI 辅助写作痕迹,采信时打折)。
—
MySQL 年份:2026
普通账号两道权限墙:装不了扩展,public 里建不了表
单方声音
运维复杂度
- 一句话
控制台建的"普通账号"既装不了扩展,又在 public 里建不了表——而库的 owner 是你碰不到的内建账号,连自建 schema 绕开都不行。
The console-created "ordinary account" can neither install extensions nor create tables in public — and the database owner is a built-in account you can never touch, so creating your own schema to work around it is blocked too.
- 窄场景
PolarDB PostgreSQL 版;用控制台"普通账号"做初始化 DDL 的团队。
PolarDB for PostgreSQL; teams running initial DDL with the console "ordinary account".
- 机制
两道独立的墙:① `CREATE EXTENSION` 要求 `polar_superuser`,普通账号 `rolsuper=f` 且不属于任何角色,直接 `permission denied to create extension "vector"`;② PG15 起 public schema 默认 ACL 只给普通用户 USAGE,而该库 owner 是 PolarDB 内建的 `aurora`(另有 `polardb_admin`、`replicator`,口令都不在用户手里),普通账号对库也没有 CREATE——于是连"自建一个 schema 绕开"这条路也被堵死。必须先建"高权限账号"(属 `pg_polar_superuser` 角色)执行 `CREATE EXTENSION` + `GRANT CREATE ON SCHEMA public TO <普通账号>`,之后才能全程用普通账号。
Two independent walls: (1) `CREATE EXTENSION` requires `polar_superuser`; the ordinary account has `rolsuper=f` and no role membership, so it gets `permission denied to create extension "vector"`; (2) since PG15 the public schema's default ACL grants ordinary users only USAGE, and the database owner is PolarDB's built-in `aurora` (plus `polardb_admin` and `replicator`, whose passwords users never hold) — the ordinary account has no CREATE on the database either, closing the "create my own schema" escape hatch. You must first create a "high-privilege account" (member of `pg_polar_superuser`) and run `CREATE EXTENSION` + `GRANT CREATE ON SCHEMA public TO <ordinary account>` before the ordinary account becomes usable.
- 生产验证
来源 3,2026-08-30 实测记录——作者的五步冒烟脚本第一次跑出 2/5,第 2 步(装 vector 扩展)、第 3 步(public 建表)当场失败,完整报错与修复步骤均有记录;
来源 4 为同一作者同一实例的配套记录(非独立来源)。
Source 3, field test Aug 30 2026 — the author's five-step smoke script first came back 2/5, with step 2 (installing the vector extension) and step 3 (creating tables in public) failing outright; full error messages and remediation steps recorded; source 4 is the same author's same-instance companion record (not an independent source).
- 证据等级
`单方声音`,开源项目实测文档(细节充分:完整报错、pg_roles 盘点、修复命令)。
`Single voice`, open-source project field-test doc (detailed: full errors, pg_roles inventory, remediation commands).
PolarDB 年份:2026
白名单没放行 = TCP 静默超时:连不上时先怀疑人生
单方声音
运维复杂度
- 一句话
白名单没放行 IP 时,PolarDB 公网地址的表现是 TCP 静默超时而不是拒绝——你会先去查网络,而不是查白名单。
When your IP isn't whitelisted, PolarDB's public endpoint goes TCP-silent-timeout instead of refusing — so you investigate the network, not the whitelist.
- 窄场景
走公网地址连接 PolarDB PostgreSQL 版;白名单配错或出口 IP 变化时。
Connecting over the public endpoint; misconfigured whitelist or changed egress IP.
- 机制
白名单拦截发生在网络层静默丢包,客户端看到的不是连接被拒绝(RST),而是 SYN 发出去石沉大海直到超时。超时症状与"网络不通/DNS 故障/实例挂了"无法区分,排查方向天然先偏。
The whitelist drops packets silently at the network layer; the client never gets a refusal (RST) — its SYNs vanish until timeout. A timeout is indistinguishable from "network down / DNS broken / instance dead," so triage naturally points the wrong way first.
- 生产验证
来源 4,2026-08-30 实测——作者明确记录"白名单不放行时的症状是 TCP 静默超时,不是拒绝,容易误判成网络故障";且他第一次连接失败还叠加了 DSN 占位符未替换的自检盲区,分层诊断(DNS→TCP→鉴权)才定位。
Source 4, field test Aug 30 2026 — the author explicitly notes "the symptom of a non-whitelisted IP is a TCP silent timeout, not a refusal — easy to misread as a network failure"; his first connection failure was further compounded by an unreplaced DSN placeholder in his own self-check, and only layered diagnosis (DNS -> TCP -> auth) located it.
- 证据等级
`单方声音`,开源项目实测文档(细节充分:症状描述、排查分层过程)。
`Single voice`, open-source project field-test doc (detailed: symptom description, layered diagnosis process).
PolarDB 年份:2026
SSL 默认关闭:公网连接串开局明文传口令
单方声音
生态与信任
- 一句话
PolarDB PostgreSQL 版实例默认 `SHOW ssl = off`,`sslmode=require` 直接连不上——公网连接串默认走明文,加密要自己去控制台手动开。
A PolarDB for PostgreSQL instance ships with `SHOW ssl = off`, and `sslmode=require` simply fails to connect — passwords travel in plaintext over the public internet unless you enable encryption yourself in the console.
- 窄场景
走公网地址连接的实例;安全合规要求加密传输的团队。
Instances accessed over the public endpoint; teams with encryption-in-transit compliance requirements.
- 机制
实例出厂默认不启用 SSL,客户端配 `sslmode=require` 会连接失败;不配则口令与数据明文过公网。作者 2026-08-31 在控制台手动开启 SSL 后 `SHOW ssl = on` 恢复正常,但明文期用过的口令是否轮换成了残留风险。
SSL is not enabled on new instances; a client configured with `sslmode=require` fails outright, and without it credentials and data cross the public network in cleartext. The author enabled SSL in the console on Aug 31 2026 (`SHOW ssl = on`) and recovered — but whether the password used during the plaintext window was rotated remains an open residual risk.
- 生产验证
来源 3(§3.5)与
来源 4,2026-08-30/31 实测——同一实例开启前后的 `SHOW ssl` 取值、`sslmode=require` 连不上的现象均有记录。
Sources 3 (section 3.5) and 4, field tests Aug 30-31 2026 — `SHOW ssl` values before and after enabling, and the `sslmode=require` failure, are both recorded on the same instance.
- 证据等级
`单方声音`,开源项目实测文档(细节充分:开启前后对照)。
`Single voice`, open-source project field-test doc (detailed: before/after comparison).
PolarDB 年份:2026
扩展版本被平台锁定:pgvector 落后上游两个小版本
单方声音
生态与信任
- 一句话
PolarDB PG 实例上的 pgvector 是 0.8.3.1,而同期社区已到 0.8.6——版本由平台定,用户自己升不了。
The pgvector on a PolarDB for PostgreSQL instance was 0.8.3.1 while upstream was already 0.8.6 — versions are set by the platform, and users cannot upgrade them themselves.
- 窄场景
PolarDB PostgreSQL 版上用 pgvector 等扩展做 AI/向量检索的团队。
Teams running pgvector or other extensions for AI/vector search on PolarDB for PostgreSQL.
- 机制
托管实例的可用扩展版本由平台统一提供:作者实测该实例 `pg_available_extensions` 共 189 个,vector 可装版本为 0.8.3.1;用户无权自行升级扩展或内核版本,只能等平台发布新修订版本。作者实测检索行为与本机 0.8.6 逐字节一致(本层只用 vector 类型与 `<=>`),但版本差意味着上游的安全修复与新特性滞后。
Available extension versions on managed instances are fixed by the platform: the tested instance listed 189 available extensions with vector installable at 0.8.3.1; users have no way to upgrade extensions or kernel versions themselves and must wait for the platform's next revision. The author verified retrieval behavior byte-identical to local 0.8.6 (this layer only uses the vector type and `<=>`), but the version gap means upstream security fixes and new features lag.
- 生产验证
来源 3(§3.2),2026-08-30 实测——版本号对照表(本机 0.8.6 vs PolarDB 0.8.3.1)、`SELECT version()` 输出 `PolarDB 16.14.20.0 build 1f03f15d` 均有记录。
Source 3 (section 3.2), field test Aug 30 2026 — version comparison table (local 0.8.6 vs PolarDB 0.8.3.1) and the `SELECT version()` output `PolarDB 16.14.20.0 build 1f03f15d` both recorded.
- 证据等级
`单方声音`,开源项目实测文档(版本号实测)。
`Single voice`, open-source project field-test doc (measured version numbers).
PolarDB 年份:2026
hot_standby_feedback:保读库查询还是保主库清理,只能二选一
单方声音
运维复杂度
- 一句话
开了 feedback,读库一个 4 小时的慢查询就能让主库 vacuum 四小时删不掉一行;关了,读库查询超 30 秒就被 cancel。
With feedback on, a single 4-hour slow query on the replica can stop the primary's vacuum from removing a single row for four hours; with it off, replica queries get cancelled after ~30 seconds.
- 窄场景
物理从库同时承担 HA 和 BI/报表查询的架构(读库"身兼两职")。
A physical replica doubling as both HA standby and BI/reporting workhorse.
- 机制
feedback 把从库 backend_xmin 回传主库,vacuum 不敢删除该快照仍可见的行版本;走复制槽连接时 xmin 钉在 pg_replication_slots.xmin 上,从库断开后依然钉着。旧的折中参数 vacuum_defer_cleanup_age 在 PG16 已被移除,中间档没了。
Feedback ships the replica's backend_xmin back to the primary, so vacuum won't remove row versions that snapshot might still need; via a replication slot the xmin stays pinned in pg_replication_slots.xmin even after the standby disconnects. The old middle-ground knob vacuum_defer_cleanup_age was removed in PG16 — no middle setting anymore.
- 生产验证
来源 12,2026-09 实战复盘——给 BI 用的从库开了 feedback,一个写坏 join 的 dashboard 查询跑了 4 小时才被发现;这 4 小时里主库最热表的 vacuum"按时成功但一行没删",n_dead_tup 一路上涨,膨胀只能事后慢慢消化。
Source 12, Sep 2026 field postmortem — feedback enabled on a BI-serving replica; a dashboard query with a bad join plan ran 4+ hours before anyone noticed; during those 4 hours the primary's hottest tables vacuumed "successfully on schedule but removed nothing," n_dead_tup climbing all the while, bloat only digestible afterward.
- 证据等级
`单方声音`,个人博客(细节充分:4 小时查询、排查过程、PG16 版本注记)。
`Single voice`, personal blog (detailed: 4-hour query, investigation steps, PG16 version note).
PostgreSQL(社区版) 年份:2026
GIN 索引锁雪崩:1.5 万个进程排队等一把锁
单方声音
性能问题稳定与故障
- 一句话
GIN 索引 pending list 批量合并要拿扩展锁,写入突发时所有 writer 串行化;死锁检测再把 16 个 LWLock 全抓一遍,雪上加霜。
GIN's pending-list batch merge takes an extension lock; under a write burst every writer serializes on it, and deadlock detection then grabs all 16 LWLocks, piling on.
- 窄场景
GIN 索引(jsonb/数组/全文检索)+ 高并发写入突发的库。
GIN indexes (jsonb/arrays/full-text) under bursty high-concurrency writes.
- 机制
GIN pending list 攒到约 512 页/4MB 才合并,每次合并拿 extension lock;突发写入让所有后端排队等这一把锁;锁释放瞬间数千进程同时惊醒,各自触发 CheckDeadlock——而死锁检测要排他抓取全部 16 个 lock-manager 分区 LWLock,形成第二个 convoy。若还开了 log_lock_waits,每个后端再遍历等待队列写 10KB+ 日志,日志子系统跟着添乱。
The GIN pending list accumulates ~512 pages/4MB before merging, each merge taking an extension lock; a synchronized write burst queues every backend on that one lock; when it releases, thousands wake at once and each fires CheckDeadlock — which exclusively grabs all 16 lock-manager partition LWLocks, forming a second convoy. With log_lock_waits on, every backend additionally walks the wait queue to write 10KB+ log messages, and the logging subsystem adds its own pressure.
- 生产验证
来源 4,WebProNews 2026-07 对 Recall.ai 具名复盘(Brendan Lockhart)的详细转述——AWS RDS 一夜之间 100% CPU(全是 system 态),1.5 万个进程等 GIN 扩展锁,等待队列超 4500,重启实例 60 分钟没救回来(负载一回来就重现)。
Source 4, WebProNews Jul 2026 detailed recount of the named Recall.ai postmortem (Brendan Lockhart) — an AWS RDS instance pinned at 100% CPU (all system time) overnight; 15,000 processes waited on one GIN extension lock, wait queue past 4,500; a reboot didn't help for 60 minutes (the load returned and re-formed the convoy).
- 证据等级
`单方声音`,独立媒体对具名复盘的详细转述(原始复盘页未直接打开核实,已如实标注)。
`Single voice`, independent media's detailed recount of a named postmortem (original postmortem page not directly opened and verified — stated as-is).
PostgreSQL(社区版) 年份:2026
逻辑复制三连坑:DDL 不复制、序列不同步、初始全量拷贝拖垮源库
单方声音
运维复杂度升级迁移
- 一句话
逻辑复制只传 DML——加列要在订阅端手工同步,序列值根本不传(切主后主键冲突),400GB 表的初始 COPY 是源库上的一个长事务。
Logical replication ships DML only — added columns need manual sync on the subscriber, sequence values are never sent (duplicate keys after failover), and the initial COPY of a 400GB table is one long transaction on the source.
- 窄场景
用逻辑复制做跨版本升级、报表从库、CDC 的团队;照着"两条命令就跑起来"的教程上生产的。
Teams using logical replication for cross-version upgrades, reporting replicas, or CDC, set up from "two commands and you're replicating" tutorials.
- 机制
publication 只解码 DML,DDL 变更订阅端无感知(要么静默漏数,要么复制 worker 直接报错停掉);序列是独立对象不在 WAL 数据流里,订阅端序列停在初始值;初始同步是单事务 COPY 全表 + 快照,期间抢 I/O、拉长 vacuum 相关膨胀。
Publications decode DML only; DDL changes go unnoticed by the subscriber (silent data gaps, or the apply worker errors out and stops); sequences are standalone objects outside the WAL data stream, so the subscriber's sequence sits at its initial value; initial sync is a single-transaction full-table COPY plus snapshot, competing for I/O and stretching vacuum-related bloat.
- 生产验证
来源 6,2026-09 报表从库实战——主库加列后订阅端出现静默的数据缺口(dashboard 空了才发现);"只读"从库转故障转移后 nextval 发出已存在的 id,几周后报 duplicate key;400GB 事实表的初始同步放在业务时段做,被 pager 叫醒。
Source 6, Sep 2026 reporting-replica field notes — a column added on the primary produced silent gaps (only noticed when a dashboard went empty); a "read-only" replica promoted in failover handed out already-used ids from nextval, surfacing as duplicate-key errors weeks later; a 400GB fact table's initial sync scheduled during business hours paged the author.
- 证据等级
`单方声音`,个人博客(细节充分:三个坑各有现象与处置)。
`Single voice`, personal blog (detailed: each pitfall with symptoms and handling).
PostgreSQL(社区版) 年份:2026
RF=1 丢一个节点:p99 飙 31 倍,还"半死不活"
单方声音
稳定与故障
- 一句话
三节点集群 RF=1 时 kill 掉一个节点,p99 从 12.7ms 飙到 392.5ms,5% 查询直接报错——实测作者说这种"半死"比彻底挂掉更糟,因为客户端收不到明确的故障信号。
Kill one node of a 3-node RF=1 cluster and p99 jumps from 12.7ms to 392.5ms with 5% of queries erroring outright — the test author calls this "half-dead" state worse than a clean outage, because clients get no clear failure signal.
- 窄场景
3 节点集群、RF=1(每分片仅一份副本)的生产部署;实测版本 v1.13.6。
3-node clusters running RF=1 (a single copy of each shard) in production; tested on v1.13.6.
- 机制
RF=1 下某分片只有一份数据,所属节点被 kill -9 后,查询协调器仍会尝试向死亡节点的分片扇出,内部超时后才返回部分结果或报错——超时等待直接打进 p99;约 1/3 查询命中死亡分片(实测 5% 直接报错、其余变慢),形成"部分可用 + 严重长尾"的混合降级,而非干净的快速失败。
At RF=1 each shard exists on exactly one node; after that node is kill -9'd, the query coordinator still fans out to the dead node's shard, waits out an internal timeout, then returns partial results or an error — the timeout wait lands directly in p99; roughly a third of queries hit the dead shard (measured: 5% errored outright, the rest slowed), producing a mixed "partially available plus severe tail latency" degradation instead of a clean fast failure.
- 生产验证
来源 2:独立故障注入实测(2026-04,Qdrant v1.13.6,100 万 SIFT 向量,50 QPS)——RF=1 kill -9 后 60 秒窗口内 p99 最高 392.5ms(基线 12.7ms,31 倍),149/3000 查询报错;成功查询的召回率不受影响(0.9989 vs 基线 0.9987)。作者结论引文:"produces a mixed degradation mode — partial availability with severe latency spikes — that is arguably worse than a clean outage, because clients receive slow, incomplete results without a clear signal that the system is degraded."(混合降级模式——部分可用叠加严重长尾延迟,可以说比干净的宕机更糟,因为客户端收到的是缓慢、不完整的结果,却没有明确的系统降级信号。)
Source 2: independent fault-injection test (Apr 2026, Qdrant v1.13.6, 1M SIFT vectors, 50 QPS) — after an RF=1 kill -9, p99 peaked at 392.5ms inside the 60s fault window (baseline 12.7ms, a 31x spike), 149 of 3,000 queries errored; recall of successful queries was unaffected (0.9989 vs 0.9987 baseline). Author's verdict, quoted: "produces a mixed degradation mode — partial availability with severe latency spikes — that is arguably worse than a clean outage, because clients receive slow, incomplete results without a clear signal that the system is degraded."
- 证据等级
`单方声音`,独立可复现实验(脚本与数据开源在仓库中),细节充分。
`Single voice`, independent reproducible experiment (scripts and data open-sourced in the repo), fully detailed.
Qdrant 年份:2026
混合检索的 BM25 分支:增量写入让 IDF 权重慢慢变质
单方声音
性能问题
- 一句话
Qdrant/bm25 稀疏模型的 IDF 统计量在写入时按当时语料计算,语料持续增量更新而不刷新,稀有词的加权优势就慢慢丢了——关键词检索质量静默下降。
The Qdrant/bm25 sparse model's IDF statistics are computed over the corpus as of write time; as the corpus keeps growing incrementally without a refresh, rare terms quietly lose their weighting advantage — keyword retrieval quality degrades silently.
- 窄场景
持续增量写入语料的混合检索(dense+sparse RRF)管线;FastEmbed 本地推理模式。
Hybrid retrieval (dense+sparse RRF) pipelines with continuously incrementally-written corpora; local-inference FastEmbed mode.
- 机制
BM25 的核心是 IDF(逆文档频率)——词越稀有加权越高。FastEmbed 的 Qdrant/bm25 模型在 upsert 时按当前语料计算 IDF 表;后续增量写入改变词频分布,但已存点的稀疏向量权重不自动更新,IDF 表变陈旧→稀有词失去加权→关键词分支排名退化。Qdrant 侧没有"增量刷新 IDF"的机制,只能靠定期重建/重嵌。
BM25 hinges on IDF (inverse document frequency) — rarer terms get higher weight. FastEmbed's Qdrant/bm25 model computes the IDF table over the corpus at upsert time; later incremental writes shift term-frequency distributions, but already-stored points' sparse weights are never refreshed, so the IDF table goes stale, rare terms lose their boost, and the keyword branch's rankings decay. Qdrant has no "incrementally refresh IDF" mechanism — only periodic rebuilds or re-embedding.
- 生产验证
来源 1:Effloow 独立实测(2026-09)——稀疏单检 sanity check"返回的全是垃圾排名,含常见词的文档都排在前面";定位为 IDF 陈旧;workaround 是几百篇一批批量写入、在语料大变化时重建索引,"对实时更新的语料,要么换固定权重的稀疏模型,要么接受定期重嵌"。
Source 1: Effloow independent test (Sep 2026) — a sparse-only sanity check "returned garbage rankings; every document with a common word scored high"; root-caused to stale IDF; the workaround was batching upserts in groups of a few hundred and re-indexing on major corpus changes: "for a live-updating corpus, we'd switch to a sparse model with fixed weights or accept periodic re-embedding."
- 证据等级
`单方声音`,独立实测(含现象、定位、workaround)。
`Single voice`, independent test (symptoms, root cause, workaround all documented).
Qdrant 年份:2026
HNSW 构建期内存尖峰:一次 5 万批量写入差点把容器 OOM
单方声音
性能问题运维复杂度
- 一句话
bulk upsert 时 HNSW 图在后台并发构建,容器内存直线爬升——实测团队靠"先调高 indexing_threshold、导完再降"才躲过一次 OOM 重启。
During bulk upserts the HNSW graph builds concurrently in the background and container memory climbs steeply — the test team only dodged an OOM restart by raising indexing_threshold first and lowering it after the backfill.
- 窄场景
大批量初始导入/回填(backfill),默认索引配置直接灌数。
Large initial imports/backfills with default index settings ingesting straight in.
- 机制
Qdrant 写入时先攒 segment,达到 indexing_threshold 后触发 HNSW 构建;构建期既要持有原始向量又要建图,内存是平稳期的数倍;单批 5 万点 + 并发构建直接把内存顶穿。默认的 indexing_threshold 对回填场景过于激进。
Qdrant accumulates segments on write and triggers HNSW construction once indexing_threshold is reached; during construction it holds both the raw vectors and the graph being built, multiplying memory usage several-fold over steady state; a single 50k-point batch plus concurrent construction blew straight through the container's memory. The default indexing_threshold is too aggressive for backfill scenarios.
- 生产验证
来源 1:Effloow 独立实测(2026-09)——"kicked off a bulk upsert of 50k points in a single batch while the HNSW graph was building concurrently, and the container's memory usage climbed sharply";采用的 pattern:hnsw_config m=16/ef_construct=100、每批约 1000 点、回填期预留 2 倍索引大小的内存空闲、"build with indexing_threshold high, backfill, then lower it"——"saved us a full out-of-memory crash and restart on the second attempt"(第二次尝试靠这套配置才躲过 OOM 崩溃重启)。
Source 1: Effloow independent test (Sep 2026) — "kicked off a bulk upsert of 50k points in a single batch while the HNSW graph was building concurrently, and the container's memory usage climbed sharply"; the adopted pattern: hnsw_config m=16/ef_construct=100, batches of ~1,000, at least 2x the index size kept free in RAM during backfills, "build with indexing_threshold high, backfill, then lower it" — which "saved us a full out-of-memory crash and restart on the second attempt."
- 证据等级
`单方声音`,独立实测(含具体配置与避险 pattern)。
`Single voice`, independent test (exact config and the avoidance pattern documented).
Qdrant 年份:2026
RRF 融合掩盖单分支故障:稀疏分支挂了你都看不出来
单方声音
运维复杂度
- 一句话
dense+sparse 双分支做 RRF 融合时,如果稀疏分支静默返回 0 命中,融合照样吐出 dense 结果——管线"看起来健康",只是质量可测地变差了。
With dense+sparse dual-branch RRF fusion, if the sparse branch silently returns zero hits, fusion still emits dense results — the pipeline "looks healthy" while quality measurably degrades.
- 窄场景
生产环境混合检索(prefetch + FusionQuery RRF)。
Production hybrid retrieval (prefetch + FusionQuery RRF).
- 机制
RRF 按排名位置融合,不看原始分数;某分支返回空/错(如 IDF 陈旧、using 名字写错、模型不匹配),融合结果只是少了一个排序信号,不会报错。故障被"优雅降级"吞掉,没有告警面。
RRF fuses by rank position, not raw scores; when one branch returns empty or wrong results (stale IDF, a mistyped `using` name, model mismatch), the fused result simply loses one ranking signal without raising any error. The failure is swallowed by "graceful degradation" with no alerting surface.
- 生产验证
来源 1:Effloow 独立实测(2026-09)——"If your sparse prefetch silently returns zero hits (stale IDF, wrong using name, model mismatch), fusion still returns dense results and your pipeline looks healthy; just measurably worse."(如果稀疏预取静默返回零命中——IDF 陈旧、using 名写错、模型不匹配——融合照样返回 dense 结果,管线看起来健康,只是可测地变差了。)团队加了一个"融合前对比双分支命中数"的两行健康检查,"would have caught gotcha #1 a day earlier"(早一天就能抓到上一条的 IDF 问题)。
Source 1: Effloow independent test (Sep 2026) — "If your sparse prefetch silently returns zero hits (stale IDF, wrong using name, model mismatch), fusion still returns dense results and your pipeline looks healthy; just measurably worse." The team added a two-line health check comparing per-branch hit counts before fusion, which "would have caught gotcha #1 a day earlier."
- 证据等级
`单方声音`,独立实测。
`Single voice`, independent test.
- 最后核验
2026-10-03
---
2026-10-03
---
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
these cards' topics may overlap existing [Pitfall] cards on this site (no access to site sources; no card titles verified).
Qdrant 年份:2026
Concurrency Scaling:WLM 路由配错,弹性变成每天 18 小时的加价
单方声音
来源存疑
性能问题成本账单
- 一句话
Concurrency Scaling 本是应对突发峰值的,但只要混杂负载都涌进开了弹性伸缩的默认 WLM 队列,它就会从"峰值保险"变成"基线容量",账单悄悄变大。
Concurrency Scaling is meant for burst spikes, but once mixed workloads all land in the default WLM queue that has scaling enabled, it stops being "peak insurance" and becomes baseline capacity — and the bill quietly grows.
- 窄场景
BI 报表、ETL、ML、应用查询混跑在同一集群、WLM 路由没做精细隔离的团队。
Teams running BI dashboards, ETL, ML, and application queries on one cluster without fine-grained WLM routing.
- 机制
Auto WLM 下查询按用户组/query group 路由进队列;只有开了 concurrency scaling 的队列能溢出到弹性集群。若未分类的混合负载都落在默认队列(通常开了弹性),CPU 压力稍一上来就触发弹性集群扩容,且按弹性计算秒数单独计费。案例中弹性集群一天在线 18–20 小时、多弹性集群并行。
Under Auto WLM, queries route into queues by user group/query group; only queues with concurrency scaling enabled can spill onto scaling clusters. If unclassified mixed workloads all land in the default queue (which typically has scaling on), any CPU pressure triggers paid scaling-cluster seconds. In the investigated case, scaling clusters were online 18–20 hours a day, several in parallel.
- 生产验证
来源 7,DoiT 2026 年对某匿名客户生产集群的账单调查——RA3 集群,账单波动几乎全来自混合用途的 analytics 集群;CloudWatch 显示 00:00–09:00 UTC 批处理与报表撞车、弹性集群每天 18–20 小时在线;根因是"太多混合负载流经唯一开了扩缩容的默认队列",WLM 路由成了成本放大器。
Source 7, DoiT, 2026, bill investigation of an anonymous customer's production RA3 cluster — nearly all spend volatility came from the mixed-use analytics cluster; CloudWatch showed batch and dashboard workloads colliding 00:00–09:00 UTC with scaling active most of the day; root cause: "too much mixed workload flowing through the one queue that had access to scaling capacity" — WLM routing as cost amplifier.
- 证据等级
`单方声音` `[来源存疑]`,云成本咨询公司博客的匿名客户案例分析(方法透明:账单数据 + CloudWatch 指标 + SYS_QUERY_HISTORY,但作者有商业利益,独立性无法完全确认)。
`Single voice` `[Questionable source]`, a cloud-cost consultancy's anonymous customer case analysis (transparent method: billing data + CloudWatch metrics + SYS_QUERY_HISTORY, but the author has commercial interests and independence cannot be fully confirmed).
- 备注
本卡主题可能与本站 [避坑] 卡重叠(并发扩展/成本)。
this card's topic may overlap existing [Pitfall] cards on this site (concurrency scaling/cost).
Amazon Redshift 年份:2026
Time Travel + Fail-safe:存储账单的"撤销税"乘数
单方声音
成本账单
- 一句话
Time Travel 按"每次变更保留一份"计费——高频变更的表开 90 天保留,存储账单能翻上百倍;Fail-safe 7 天还关不掉。
Time Travel bills "one retained copy per change" — a high-churn table with 90-day retention can multiply the storage bill a hundredfold; and the 7-day Fail-safe cannot be turned off.
- 窄场景
高 churn 表(ETL staging、频繁 UPDATE/DELETE)+ Enterprise 版长保留期。
High-churn tables (ETL staging, frequent UPDATE/DELETE) combined with long Enterprise-edition retention.
- 机制
微分区不可变,UPDATE/DELETE 写新分区、旧分区保留至保留期结束;Time Travel 存储 ≈ 保留天数 × 数据变更量;Fail-safe 在 Time Travel 到期后再加 7 天,不可配置、不可关闭、不可查询,但照样计费。
Micro-partitions are immutable — UPDATE/DELETE writes new partitions and retains the old ones until retention expires; Time Travel storage scales as retention days x data churn; Fail-safe adds 7 more days after Time Travel expires — non-configurable, non-disablable, non-queryable, but still billed.
- 生产验证
来源 6,Nazeer Syed 2026-01——1TB staging 表每小时 truncate+reload、7 天保留,"turn a 1TB bill into a 168TB bill overnight";90 天 Time Travel + 7 天 Fail-safe = 删除数据要在付费存储里躺 97 天;并给出 `TABLE_STORAGE_METRICS` 审计膨胀的 SQL。
Source 6, Nazeer Syed, Jan 2026 — a 1TB staging table truncated and reloaded hourly with 7-day retention: "turn a 1TB bill into a 168TB bill overnight"; 90 days of Time Travel + 7 days of Fail-safe means deleted data sits in paid storage for 97 days; includes a `TABLE_STORAGE_METRICS` query for auditing the bloat.
- 证据等级
`单方声音`,具名工程师博客(细节充分:微分区机理、168TB 测算、审计 SQL)。
`Single voice`, named engineer blog (detailed: micro-partition mechanics, the 168TB calculation, audit SQL).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
SQL API 高并发直接 429/503:瓶颈在云服务层,加大 warehouse 没用
单方声音
性能问题
- 一句话
10 rps 持续压力下 SQL API 撞上 429/503,而 warehouse 远未到瓶颈——直觉是"加 warehouse",作者实测证明"it doesn't help - the bottleneck is upstream of it"。
Under sustained 10 rps the SQL API hits 429/503 while the warehouse is nowhere near its limit — the instinct is "scale up the warehouse," and the author measured that "it doesn't help - the bottleneck is upstream of it."
- 窄场景
用 Snowflake SQL API 做高并发在线 serving;轻量集成测试通过、上生产才发现限流的团队。
High-concurrency online serving on the Snowflake SQL API; teams whose light integration tests pass but production throttles.
- 机制
SQL API 无状态,每个请求都要走 auth→编译→路由→执行;Cloud Services 层的限流独立于 warehouse size,瓶颈在 warehouse 上游。
The SQL API is stateless; every request goes through auth → compile → route → execute; Cloud Services throttling is independent of warehouse size, so the bottleneck sits upstream of the warehouse.
- 生产验证
来源 12:Mechanical Rock 2026-10 生产复盘——"At 10 requests per second sustained, the Cloud Services layer starts to feel it before the warehouse does. We hit 429 (TooManyRequests) and 503 (ServiceUnavailable) errors under peak load.";只能靠指数退避重试、批量写或换有连接池的 connector 绕行;"If you're expecting sustained high concurrency, test this early. It's not a problem you'll see in light integration testing."
—
- 证据等级
`单方声音`,独立咨询公司一手生产实测(数据具体:10 rps 阈值、429/503 错误码)。
`Single voice`, independent consultancy first-hand production measurement (concrete: 10 rps threshold, 429/503 codes).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
CLONE DATABASE 挂 22 分钟:元数据操作是单线程串行 [来源存疑]
单方声音
来源存疑
性能问题
- 一句话
对 1000 张表的 schema 执行标准 `CLONE DATABASE` 挂了 22 分 34 秒——database 级 clone 是 Cloud Services 层的串行元数据操作,"You aren't compute-bound; you're dispatch-bound."(你卡的不是算力,是调度。)
A standard `CLONE DATABASE` on a 1,000-table schema hung for 22 minutes 34 seconds — database-level clone is a serial metadata operation in the Cloud Services layer: "You aren't compute-bound; you're dispatch-bound."
- 窄场景
大 schema 的 database 级 clone;CI/CD 里用 clone 做环境复制的 pipeline。
Database-level clones of large schemas; CI/CD pipelines using clone for environment provisioning.
- 机制
database 级 clone 不搬运数据,只处理元数据指针,但 dispatch 是单线程队列;改用 Python API 对单表并行 clone(10 线程)可把总耗时压到 22 秒,约 60 倍加速——证明瓶颈在调度而非数据量。
Database-level clone moves no data, only metadata pointers, but dispatch is a single-threaded queue; parallel per-table clone via the Python API (10 threads) cut total time to 22 seconds (~60×) — proving the bottleneck is scheduling, not data volume.
- 生产验证
来源 11:Alexandra Sampietro 2026-02-20([来源存疑])——"cloning a database with 1,000 tables took 22 minutes and 34 seconds… across 10 parallel threads, the total time dropped to just 22 seconds";"a database-level clone is a serial metadata operation handled by the Cloud Services layer. It's a single-threaded queue."
—
- 证据等级
`单方声音`,[来源存疑]:作者身份背景未能独立核实,但数字具体、机制(Cloud Services 串行瓶颈)与本页其他卡片的根因一致。
`Single voice`, [source questionable]: author background could not be independently verified, but the numbers are concrete and the mechanism (serial Cloud Services bottleneck) matches the root cause in other cards on this page.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
Masking 策略静默破坏 hybrid 表索引:查询退化为全列扫描
单方声音
性能问题
- 一句话
masking policy 一落到 hybrid 表的索引列上,Snowflake 就把扫描模式从 ROW_BASED 静默切换为 COLUMN_BASED——索引被完全绕过,查询退化为全列扫描;"a massive performance regression you won't see coming until you deploy to an environment where masking policies are actually active (which, in our case, was production)"(直到部署到真正启用 masking 的环境——也就是生产——你才看得到)。
Apply a masking policy to a hybrid-table index column and Snowflake silently switches scan mode from ROW_BASED to COLUMN_BASED — the index is bypassed entirely and the query degrades to a full column scan; "a massive performance regression you won't see coming until you deploy to an environment where masking policies are actually active (which, in our case, was production)."
- 窄场景
hybrid 表 + 列级 masking 的组合;dev/staging 没配 masking、生产才配的团队。
Hybrid tables combined with column-level masking; teams whose dev/staging lack masking but production has it.
- 机制
masking 与 hybrid 表行存储索引路径不兼容,引擎静默降级为列扫;workaround 是给索引列打 `CLASSIFICATION_EXEMPT` 剥离 masking,代价是这些列不再脱敏,只能靠更严格的 RBAC 补偿——安全与性能二选一。
Masking is incompatible with the hybrid-table row-store index path, so the engine silently downgrades to column scans; the workaround is tagging index columns `CLASSIFICATION_EXEMPT` to strip masking — at the cost of leaving those columns unmasked and compensating with stricter RBAC. Security or performance, pick one.
- 生产验证
来源 12:Mechanical Rock 2026-10 生产复盘——"When a masking policy is applied to a column that's part of a hybrid table index, Snowflake switches from ROW_BASED scan mode to COLUMN_BASED. That bypasses the index entirely and falls back to a full column scan."
—
- 证据等级
`单方声音`,独立咨询公司一手生产实测(hybrid 表较新,未见第二家公开复盘)。
`Single voice`, independent consultancy first-hand production measurement (hybrid tables are new; no second public postmortem found).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
Result cache 按 query text 缓存、忽略 bind 参数:用户 A 可能拿到用户 B 的数据
单方声音
生态与信任
- 一句话
Snowflake 按查询文本做结果缓存、完全忽略 bind 参数——两个用户用同一个参数化 SQL 查不同 customer ID,会拿到同一份缓存结果;这是正确性问题,不是性能问题。
Snowflake caches results by query text and ignores bind parameters entirely — two users running the same parameterized SQL for different customer IDs get the same cached result set. That's a correctness bug, not a performance quirk.
- 窄场景
多租户应用共用参数化 SQL 模板;依赖 result cache 提速的在线查询。
Multi-tenant apps sharing parameterized SQL templates; online queries relying on the result cache for speed.
- 机制
结果缓存的 key 是查询文本,bind 参数值不参与 key;参数化查询在缓存命中时直接返回别人的结果集。
The result-cache key is the query text; bind parameter values don't participate, so a parameterized query can return another user's result set on a cache hit.
- 生产验证
来源 12:Mechanical Rock 2026-10 生产复盘——"Snowflake caches query results based on query text only - bind parameters are ignored entirely. Two users querying with different customer IDs but the same parameterised SQL will get the same cached result."
—
- 证据等级
`单方声音`,独立咨询公司生产测试观察;行为描述具体可复核。
`Single voice`, independent consultancy production-test observation; behavior described concretely and reproducibly.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
Gen2 warehouse 上 bulk ALTER TABLE 稳定报引擎内部错误,生产每天 60+ 失败
单方声音
稳定与故障
- 一句话
迁到 Gen2 warehouse 后,dbt 生成的 bulk `ALTER TABLE ... ALTER`(宽表一次改几十上百列 comment)稳定报错 `"000603 (XX000): SQL execution internal error: Processing aborted due to error 300002"`;Gen1 上 500+ 列的表才偶发,Gen2 上 120 列左右即可靠复现——而 Snowflake 文档声称 Gen2 "SQL semantics are unchanged"。
After migrating to Gen2 warehouses, dbt-generated bulk `ALTER TABLE ... ALTER` (dozens to hundreds of column comments on wide tables) reliably fails with `"000603 (XX000): SQL execution internal error: Processing aborted due to error 300002"`; on Gen1 only 500+-column tables hit it occasionally, on Gen2 ~120 columns reproduces it — while Snowflake docs claim Gen2 "SQL semantics are unchanged."
- 窄场景
已迁移或被默认迁移到 Gen2 warehouse 的 dbt 用户;宽表 comment 批量维护。
dbt users migrated (or default-migrated) to Gen2 warehouses; bulk comment maintenance on wide tables.
- 机制
引擎侧 bug(Snowflake 文档口径与实测行为矛盾);dbt-adapters 只能改 emission shape 绕行;Gen2 自 2026-07 起成为新 org 的默认 warehouse 类型,影响面随默认切换扩大。
An engine-side bug contradicting Snowflake's own docs; dbt-adapters can only work around it by changing emission shape; Gen2 became the default warehouse type for new orgs in Jul 2026, widening the blast radius.
- 生产验证
来源 16:dbt-adapters GitHub issue #1919(dbt Labs 维护者代 dbt Platform 客户上报,约 2026-05)——生产环境每天 60+ 失败,issue 仍 open。
—
- 证据等级
`单方声音`,dbt Labs 代客户上报的生产故障(错误原文、复现条件、绕行方案俱全);dbt-snowflake#842 曾报告同类 300002/XX000 错误模式,可作侧面参考。
`Single voice`, dbt Labs filing a customer's production failure (error text, repro conditions and workaround all documented); dbt-snowflake#842 reported a similar 300002/XX000 pattern as a side reference.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
Catalog-linked Iceberg 表刷新静默失败,下游查询报 Parquet file inaccessible
单方声音
稳定与故障
- 一句话
catalog-linked 的 Iceberg 表 refresh 看似成功、实际没完全更新——某表 delete 文件过多触及内部元数据处理上限,refresh 静默损坏,下游查询间歇性报 `"Parquet file inaccessible"`。
Refresh on catalog-linked Iceberg tables looks successful but doesn't fully apply — when a table accumulates too many delete files it hits an internal metadata-processing ceiling, the refresh silently corrupts, and downstream queries intermittently fail with `"Parquet file inaccessible"`.
- 窄场景
Snowflake catalog-linked Iceberg 表 + 高 churn(delete 文件多)的表。
Snowflake catalog-linked Iceberg tables with high churn (many delete files).
- 机制
refresh 的元数据处理有内部上限,超限后 refresh 不报错但元数据停留在旧版本;读到的元数据是旧的,查询时才暴露为文件不可访问。
Refresh metadata processing has an internal cap; exceeding it leaves refresh unerrored but metadata stale; reads then surface as file-inaccessible errors.
- 生产验证
来源 17:独立工程师 Soumil Shah 2026-07 复盘——第一反应怀疑存储层 compaction 与读请求的竞态,双边开 support case 排查后排除;诊断方法 `SYSTEM$CATALOG_LINK_STATUS`,workaround 逐表手动处理;文章开篇声明为个人技术复盘,不代表雇主或厂商。
—
- 证据等级
`单方声音`,单源客户复盘,但故障机制描述具体可核(函数名、错误原文、复现条件俱全),且已向 Snowflake 开 support case。
`Single voice`, single-customer postmortem, but the failure mechanism is concretely verifiable (function names, exact errors, repro conditions) and a Snowflake support case was opened.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
Zero-copy clone:免费午餐的三张账单
单方声音
运维复杂度
- 一句话
clone 是元数据操作、看似免费,但有三张账单:删源表即断 clone("把源表扣在元数据层当人质")、放几个月的 clone 在源表高 churn 下会"拥有"旧分区的完整副本(亲历:账户存储账单翻三倍)、有 SELECT 就能 clone 导致 masking 没配好就把未脱敏生产数据端给 dev(亲历:合规审计挂掉)。
Clone is a metadata operation and looks free, but there are three bills: drop the source table and the clone loses its micro-partition references ("holding the source table hostage in the metadata layer"); a months-old clone on a high-churn source ends up owning full copies of old partitions (first-hand: the account's storage bill tripled); SELECT privilege usually suffices to clone, so a misconfigured masking policy hands unmasked prod data to dev (first-hand: a failed compliance audit).
- 窄场景
用 clone 做 dev/test 环境、CI 临时环境的团队;DBA 有 drop 重建 staging 表习惯的组织。
Teams using clone for dev/test environments and CI; orgs whose DBAs habitually drop and rebuild staging tables.
- 机制
clone 与源表共享 micro-partition,任何修改(RECLUSTER/DELETE)都让分区偏离;clone 断裂于源表删除;clone 权限通常随 SELECT 走,masking 策略若未覆盖 clone 路径即形同虚设。
Clones share micro-partitions with the source; any modification (RECLUSTER/DELETE) diverges partitions; clones break when the source is dropped; clone privilege typically follows SELECT, so masking policies that don't cover the clone path are decorative.
- 生产验证
来源 22:aniketsoni 2026-09 个人工程博客——"you are essentially holding the source table hostage in the metadata layer";junior dev clone 大事实表六周未删,"The storage bill for that account tripled because we were essentially versioning the entire fact table history twice.";"I've seen compliance audits fail specifically because a developer cloned a production table to debug a join issue, inadvertently exposing PII to the development environment."
—
- 证据等级
`单方声音`,细节充分(三倍账单实例、合规审计失败实例、机制解释完整),未找到第二独立来源,未硬凑。
`Single voice`, richly detailed (tripled-bill instance, failed-audit instance, complete mechanism); no second independent source found, none invented.
- 备注
存储账单翻三倍也沾"成本账单",但根因是 clone 生命周期无人管理,归运维复杂度。本卡主题可能与本站 [避坑] 卡重叠。
the tripled storage bill also touches "cost billing," but the root cause is unmanaged clone lifecycle, so operational complexity is the better tag. This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
MFA 强制上线打碎了官方文档的告警方案:邮件发不到 distribution list
单方声音
运维复杂度
- 一句话
Snowflake 邮件告警要求 `ALLOWED_RECIPIENTS` 里的每个地址都属于"有已验证邮箱的 Snowflake 用户",官方 KB 的变通方案是建"桥接用户"验证邮箱;但 BCR 2086/2097 上线后,所有用密码登录 Snowsight 的 PERSON 用户被强制注册 MFA——而绑定邮件组的桥接用户没有手机、没有 Duo、没有 passkey,"You're stuck."(你卡死了。)
Snowflake email alerts require every `ALLOWED_RECIPIENTS` address to belong to "a Snowflake user with a verified email"; the official KB workaround is a "bridge user" whose email is the mailing-list address. Then BCR 2086/2097 forced MFA enrollment for all PERSON users signing into Snowsight with a password — and a bridge user tied to a mailing list has no phone, no Duo app, no passkey: "You're stuck."
- 窄场景
想把 Snowflake 告警发到团队邮件组(data-team@)的运维;按官方 KB 建了桥接用户的团队。
Ops teams wanting Snowflake alerts sent to a team mailing list; teams that built the KB's bridge user.
- 机制
安全加固(MFA 强制)与官方文档的变通方案互相打架,文档没提这堵墙;绕行方案是 `MINS_TO_BYPASS_MFA=30` 开 30 分钟豁免窗口完成验证、再立刻封死用户——等于为了发个告警邮件,得走一套"临时开后门再焊死"的流程。
The security hardening (mandatory MFA) conflicts with the official workaround, and the docs never mention the wall; the escape hatch is a 30-minute `MINS_TO_BYPASS_MFA` exemption to complete verification, then immediately sealing the user — a "temporarily open a backdoor, then weld it shut" procedure just to send an alert email.
- 生产验证
来源 29:Stéphane Bizard 2026-09(生产账号实测)——"Since the BCR 2086/2097 rollout, Snowflake enforces MFA enrollment for all PERSON users who sign in to Snowsight with a password. The moment you log in with your bridge user to validate the email, Snowsight forces you to enroll in MFA — and a bridge user tied to a mailing list has no phone, no Duo app, no passkey. You're stuck."
—
- 证据等级
`单方声音`,一线工程师生产账号实测;BCR 2086/2097 强制 MFA 本身是公开变更,"桥接用户被 MFA 卡住"这一具体冲突未找到第二独立来源。
`Single voice`, frontline engineer tested on a production account; BCR 2086/2097 mandatory MFA is public, but this specific conflict has no second independent source.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
列级 PII 脱敏是 Enterprise 版专属:Standard 版里 SELECT 到的就是明文
单方声音
生态与信任
- 一句话
"CREATE MASKING POLICY doesn't exist on Standard Edition"——没有它,"every reader with SELECT access sees raw values, no matter how carefully everything else in this account is scoped"(任何有 SELECT 权限的读者看到的都是明文,不管账号里其他权限管得多细)。
"CREATE MASKING POLICY doesn't exist on Standard Edition" — without it, "every reader with SELECT access sees raw values, no matter how carefully everything else in this account is scoped."
- 窄场景
Standard 版上存 PII 的团队;以为"权限管细了就安全"的账号。
Teams storing PII on Standard edition; accounts that assumed "tight permissions = safe."
- 机制
基础安全能力按版本收费;作者的应对是提前把分类角色的授权图谱建好、等升级 Enterprise 那天再"一键打开"脱敏——等于承认加钱之前 PII 保护是裸奔的。
A basic security capability sold by edition; the author's approach was to pre-build the classified-role grant graph and flip masking on the day of the Enterprise upgrade — an admission that PII protection runs naked until you pay up.
- 生产验证
来源 31:krish0502 2026-09(55 角色实建脚本的多环境账号架构实录,非纸上谈兵)。
—
- 证据等级
`单方声音`,"masking policy 需 Enterprise 版"是产品事实(官方文档可核),"Standard 版裸奔"的批评口径未找到第二独立客户信源。
`Single voice`; "masking requires Enterprise" is a product fact (verifiable in official docs), the "naked on Standard" criticism has no second independent customer voice.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
R1 RCM:12 周把 847 个 dbt 模型从 Snowflake 搬到 Databricks,Snowflake 专有函数是最大坑 [来源存疑]
单方声音
来源存疑
升级迁移
- 一句话
847 个 dbt 模型、5 层、35 张信息集市表、每天跑多次覆盖数千亿行,12 周迁完、单次运行成本降约 77%——但 Snowflake 的 `HASH()` 不可移植(Databricks 无等价实现,surrogate key 注定两侧不同)、Snowflake MERGE 容忍源端重复键而 Delta MERGE 要求目标键唯一,被迫在数百个模型上游统一加 `QUALIFY + ROW_NUMBER()`。
847 dbt models, 5 layers, 35 marts, multiple daily runs over hundreds of billions of rows — migrated in 12 weeks with ~77% lower per-run cost; but Snowflake's `HASH()` isn't portable (no Databricks equivalent, surrogate keys guaranteed to differ on both sides) and Snowflake MERGE tolerates duplicate source keys while Delta MERGE requires unique target keys, forcing `QUALIFY + ROW_NUMBER()` across hundreds of upstream models.
- 窄场景
重度使用 Snowflake 专有函数/语义的 dbt 项目;Data Vault 建模。
dbt projects heavy on Snowflake-proprietary functions/semantics; Data Vault modeling.
- 机制
专有函数与宽松语义是迁移时的隐性锁定——不是"数据搬不走",是"语义对不上",测试策略被迫围绕"key 一定不同"来设计;NULL 排序默认、隐式类型转换、timestamp/timezone、窗口函数 frame 语义等执行引擎差异是工程量的大头。
Proprietary functions and lenient semantics are hidden lock-in — not "data can't move" but "semantics don't line up"; the test strategy had to be designed around "keys will definitely differ"; NULL sort defaults, implicit casts, timestamp/timezone and window-frame semantics made up the bulk of the engineering.
- 生产验证
来源 36:Databricks 社区 technical blog 2026-05([来源存疑]:咨询公司 Lovelytics 与 R1 的 Zheng Zhu 联合撰写,厂商社区合作稿,非客户独立执笔)——客户与人物具名、数字具体;技术细节(HASH 不可移植、MERGE 语义差异)与已知的 Snowflake→Databricks 迁移通用知识一致。
—
- 证据等级
`单方声音`,[来源存疑]:未见 R1 自家博客或第三方独立复述。
`Single voice`, [source questionable]: no R1-owned blog or third-party independent retelling found.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
Appcues:告别 Snowflake+Airflow 迁往 ClickHouse,P95 查询 20 秒→2 秒 [来源存疑]
单方声音
来源存疑
升级迁移
- 一句话
管理 1.31 PB 数据、4100 亿事件的 SaaS 平台,多年 Snowflake+Airflow 分析栈:Airflow 每 5–10 分钟 rollup 预聚合再写回,带来额外成本和庞大 job 网,客户只能看到 10 分钟窗口的数据,查询涨到 20 秒;迁 ClickHouse 后 P95 降到 2 秒以内,新事件可查延迟从 10 分钟降到约 5 秒,整体分析栈支出仍降 23%——"No matter how we tuned our previous cloud data warehouse, ClickHouse was faster."(无论我们怎么调优之前的云数仓,ClickHouse 就是更快。)
The SaaS onboarding platform (1.31 PB, 410B events) ran Snowflake + Airflow for years: Airflow rollup pre-aggregations every 5–10 minutes added cost and a sprawling job mesh, customers saw 10-minute-old data, queries hit 20 seconds; on ClickHouse P95 fell under 2 seconds, event freshness went from 10 minutes to ~5 seconds, and total analytics spend still fell 23% — "No matter how we tuned our previous cloud data warehouse, ClickHouse was faster."
- 窄场景
高基数事件分析、需要秒级新鲜度的用户行为分析;Airflow rollup 越堆越多的团队。
High-cardinality event analytics needing second-level freshness; teams drowning in Airflow rollups.
- 机制
Snowflake 行列混存对高并发点查/事件分析不是最优解;预聚合链路(Airflow rollup)是架构税;ClickHouse 物化视图在写入时预聚合,退役 Airflow。
Snowflake's row-column hybrid isn't optimal for high-concurrency point/event queries; the pre-aggregation pipeline (Airflow rollups) is an architecture tax; ClickHouse materialized views pre-aggregate at write time and Airflow was retired.
- 生产验证
来源 37:ClickHouse 官方博客客户访谈 2026([来源存疑]:厂商邀请的客户访谈,非客户自发撰写)——受访人 Appcues 工程副总裁 Chris Brookins 具名;迁移方式:Snowflake 全量导出为 S3 Parquet 再导入(PoC 快 73%),随后一年期 "in-flight migration"("We were doing this while the plane was flying")。
—
- 证据等级
`单方声音`,[来源存疑]:未见 Appcues 自家工程博客的独立版本。
`Single voice`, [source questionable]: no independent Appcues engineering-blog version found.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
"Snowflake 原谅低效查询,账单最后才让你疼"
单方声音
升级迁移
- 一句话
一线从业者十年三次迁移(Teradata→Snowflake→BigQuery)的毒观察:on-prem 时代"efficiency wasn't a virtue, it was a constraint"(效率不是美德,是硬约束);迁入 Snowflake 后,"the platform that was supposed to free us from infrastructure concerns has instead freed us from the consequences of bad engineering"(本该把你从基础设施中解放的平台,反而把你从烂工程的后果中解放了)——弹性算力让低效查询、脏数据悄无声息堆积,直到有人看账单那一刻才爆:"Snowflake forgives (inefficient queries). BigQuery punishes."
A practitioner's venomous observation from a decade of three migrations (Teradata→Snowflake→BigQuery): on-prem, "efficiency wasn't a virtue, it was a constraint"; after moving to Snowflake, "the platform that was supposed to free us from infrastructure concerns has instead freed us from the consequences of bad engineering" — elastic compute lets inefficient queries and dirty data pile up silently until someone looks at the bill: "Snowflake forgives (inefficient queries). BigQuery punishes."
- 窄场景
从 on-prem/固定容量迁入 Snowflake 的团队;SQL 质量无人看管的组织。
Teams moving from on-prem/fixed capacity to Snowflake; orgs where nobody reviews SQL quality.
- 机制
弹性计费移除了"写烂 SQL 立刻疼"的反馈回路;数据质量侧同样:"Neither is a true relational database and neither enforces referential integrity. So errors that a proper RDBMS would have rejected at the gate instead quietly propagate through your data."(两家都不强制引用完整性;正规 RDBMS 会在入口拒绝的错误,在这里悄悄向下游传播。)
Elastic billing removes the "bad SQL hurts immediately" feedback loop; data-quality side too: "Neither is a true relational database and neither enforces referential integrity. So errors that a proper RDBMS would have rejected at the gate instead quietly propagate through your data."
- 生产验证
来源 42:Steven Jeffrey Feldman 2026-03(短文回应体,细节较少,作为观点性记录);
来源 41:Receipt Bank"账单反而更高"的实例与此机制同调。
—
- 证据等级
`单方声音`,独立从业者亲历观察(观点性,机制与来源 41 的实例互相印证)。
`Single voice`, independent practitioner's lived observation (opinionated; mechanism corroborated by source 41's instance).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
JDBC 把"密码错误"报成"连不上":SQLState 08001,连接池会无限重试错误密码
单方声音
生态与信任
- 一句话
错误密码登录返回 `errorCode=390100 sqlState=08001 msg=Incorrect username or password was specified`——但 08001 的含义是"SQL client 无法建立连接"(网络类故障),认证被拒按 SQL 标准应为 28000;所有按 SQLState 做错误分类的通用工具(jOOQ、连接池、重试框架)都会把"密码错了"这种永久性失败当成瞬时网络抖动去重试,jOOQ 还会映射成 503 而不是 401。
Logging in with a wrong password returns `errorCode=390100 sqlState=08001 msg=Incorrect username or password was specified` — but 08001 means "SQL client unable to establish connection" (a network-class failure); auth rejection should be 28000 per the SQL standard. Every generic tool classifying errors by SQLState (jOOQ, connection pools, retry frameworks) treats "wrong password" — a permanent failure — as transient network flakiness and retries it; jOOQ even maps it to a 503 instead of a 401.
- 窄场景
JDBC + 连接池/重试框架/jOOQ 的 Java 服务;靠 SQLState 做错误分类的中间件。
Java services on JDBC + connection pools/retry frameworks/jOOQ; middleware classifying errors by SQLState.
- 机制
`SessionUtil.java` 登录失败抛错处硬编码了 `SQLCLIENT_UNABLE_TO_ESTABLISH_SQLCONNECTION`,无视服务端返回的 error code;讽刺的是驱动对本地检测到的缺用户名/缺密码反而正确用了 28000。
`SessionUtil.java` hardcodes `SQLCLIENT_UNABLE_TO_ESTABLISH_SQLCONNECTION` at the login-failure throw site, ignoring the server's error code; ironically the driver correctly uses 28000 for locally-detected missing username/password.
- 生产验证
来源 46:snowflake-jdbc 官方仓库用户 issue #2681(2026-06-26,作者 andriivaliukh,数据库工具方向工程师)——附完整复现代码和修复方案(含 8 个认证拒绝码的映射表),"reproduced firsthand"(亲手复现)。
—
- 证据等级
`单方声音`,用户亲手复现+修复方案;同仓库同期多个连接/认证类 issue 说明该管线是投诉集中区,但 SQLState 误标这一点是该 issue 首次系统性指出。
`Single voice`, firsthand reproduction plus fix proposal; sibling connection/auth issues in the same repo show the pipeline is a complaint hotspot, but this is the first systematic writeup of the SQLState mislabeling.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
没有 schema-on-read:COPY INTO 要你预知未来,Databricks Auto Loader 不用
单方声音
生态与信任
- 一句话
同一套欺诈分析 workload(含 6 种"脏数据"注入)双平台实跑:Databricks Auto Loader "infers schemas on first sight, evolves them as files arrive, and rescues unknown fields"(初见即推断、文件到了就演进、未知字段自动兜底),schema v1→v2 演进零代码改动;Snowflake `COPY INTO` 要求目标列预先声明,"Schema mismatches are hard errors unless you wrote the COPY transform defensively"——"both produced the same downstream silver table, but Snowflake required me to know the future schema in advance and Databricks didn't."
The same fraud-analytics workload (with 6 kinds of "dirty data" injected) run on both platforms: Databricks Auto Loader "infers schemas on first sight, evolves them as files arrive, and rescues unknown fields" — schema v1→v2 evolution with zero code changes; Snowflake `COPY INTO` requires pre-declared target columns — "Schema mismatches are hard errors unless you wrote the COPY transform defensively": "both produced the same downstream silver table, but Snowflake required me to know the future schema in advance and Databricks didn't."
- 窄场景
半结构化/ schema 演进的源数据入仓;"脏数据"常态化的 pipeline。
Ingesting semi-structured / schema-evolving source data; pipelines where "dirty data" is the norm.
- 机制
Snowflake 的 COPY INTO 是 schema-on-write:目标结构必须预先声明,schema 对不上即硬报错;Databricks Auto Loader 是 schema-on-read:推断+演进+兜底。Snowflake 的 Schema Inference 功能长期 private preview,未见 GA 公告。
Snowflake's COPY INTO is schema-on-write: target structure must be pre-declared, mismatches are hard errors; Databricks Auto Loader is schema-on-read: infer + evolve + rescue. Snowflake's Schema Inference feature has sat in private preview with no GA announcement.
- 生产验证
来源 50:独立从业者 btriani 公开实测 repo(2026-05,同一 workload 双平台实跑 + 4 项 parity 校验,可复跑代码)。
—
- 证据等级
`单方声音`,一人实测 repo,但为可复跑代码 + parity 校验,非空口对比。
`Single voice`, one-person measurement repo, but rerunnable code + parity checks, not armchair comparison.
- 备注
缺口状态:至今缺失。本卡主题可能与本站 [避坑] 卡重叠。
gap status: still missing. This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
流批一体缺失:Databricks 切流批改个参数,Snowflake 得换架构
单方声音
生态与信任
- 一句话
Databricks 侧 Structured Streaming 是 first-class,"same code works batch and stream",从批切到流是 "Flip one option (`trigger=continuous`)";Snowflake 侧 `COPY INTO` "inherently batch (one file/execution)",想上流式得换 Snowpipe auto-ingest 架构——而 Snowpipe auto-ingest 需要 S3+SNS 的外部事件链路,"out of scope for trial"(连试用都搭不起来)——"on Databricks the path from batch → streaming is a flag flip. On Snowflake it's an architecture change."
On Databricks, Structured Streaming is first-class — "same code works batch and stream," batch→stream is "Flip one option (`trigger=continuous`)"; on Snowflake `COPY INTO` is "inherently batch (one file/execution)" — going streaming means switching to the Snowpipe auto-ingest architecture, which needs an S3+SNS external event chain, "out of scope for trial": "on Databricks the path from batch → streaming is a flag flip. On Snowflake it's an architecture change."
- 窄场景
从批处理向流处理演进的 pipeline;PoC 阶段想快速验证流式链路的团队。
Pipelines evolving from batch to streaming; teams wanting to validate a streaming path quickly in PoC.
- 机制
Databricks 的批流是同一套 API 的两种触发模式;Snowflake 的批(COPY INTO)与流(Snowpipe)是两套架构、两套 API,切换=重搭链路。
Databricks batch/stream are two trigger modes of one API; Snowflake batch (COPY INTO) and streaming (Snowpipe) are two architectures and two APIs — switching means rebuilding the pipeline.
- 生产验证
来源 50:btriani 双平台实测 repo 第 6 问(2026-05)。
—
- 证据等级
`单方声音`,同一实测 repo 的另一问结论(可复跑代码 + parity 校验)。
`Single voice`, another question from the same measurement repo (rerunnable code + parity checks).
- 备注
缺口状态:至今缺失(Snowpipe Streaming 是另一套 API/架构,无"同一份代码批流通用"的声明式流处理)。本卡主题可能与本站 [避坑] 卡重叠。
gap status: still missing (Snowpipe Streaming is a separate API/architecture; no declarative "same code for batch and stream"). This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
Unity Catalog 治理的是表+模型+端点+notebook,Horizon 只看得见表
单方声音
生态与信任
- 一句话
Databricks Unity Catalog 的治理对象是"Tables + files + ML features + model versions + serving endpoints + vector indexes + notebooks + agents","UC's surface area is broader because Databricks ships ML and notebooks as first-class governable objects";Snowflake Horizon 是"Tables + views + business metrics + row policies + column masks + dashboards"——"Narrower scope but cleaner primitives"(范围窄,但原语干净);当治理对象超出"仓里的表"(模型、端点、notebook 协作)时,Horizon 没有对应物。
Databricks Unity Catalog governs "Tables + files + ML features + model versions + serving endpoints + vector indexes + notebooks + agents" — "UC's surface area is broader because Databricks ships ML and notebooks as first-class governable objects"; Snowflake Horizon governs "Tables + views + business metrics + row policies + column masks + dashboards" — "Narrower scope but cleaner primitives"; once governed objects go beyond "tables in the warehouse" (models, endpoints, notebook collaboration), Horizon has no counterpart.
- 窄场景
ML 与数据平台共用一套治理的组织;需要治理 notebook 协作与模型端点的团队。
Organizations governing ML and data platforms under one model; teams needing to govern notebook collaboration and model endpoints.
- 机制
Horizon 的设计边界是"仓内对象";Databricks 把 ML/ notebook 当一等治理对象。作者对 Horizon 的 masking/row-access-policy 语法亦有公允评价,非一边倒。
Horizon's design boundary is "in-warehouse objects"; Databricks treats ML/notebooks as first-class governable objects. The author fairly credits Horizon's more elegant masking/row-access-policy syntax — not one-sided.
- 生产验证
来源 50:btriani 实测 repo 第 3 问(2026-05)。
—
- 证据等级
`单方声音`,独立实测 repo(作者对双方均有公允评价)。
`Single voice`, independent measurement repo (fair to both sides).
- 备注
缺口状态:至今缺失。本卡主题可能与本站 [避坑] 卡重叠。
gap status: still missing. This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
地理空间负载:BigQuery GIS 又快又原生,Snowflake 复杂空间 join 落后 [来源存疑]
单方声音
来源存疑
性能问题
- 一句话
30 天双平台实测(同一 workload 双跑:pipeline + adhoc、三个 BI 工具、每天约 8TB 扫描)专列 "Native GIS Performance":"BigQuery handled our geospatial queries significantly faster. Its native GIS functions and columnar storage for geography types are a genuine advantage for logistics, delivery, or any location-aware workloads. Snowflake supports geospatial queries but lags on performance for complex spatial joins at scale.";选型建议直接写:"Choose BigQuery if… Geospatial or GIS workloads are part of your stack"。
A 30-day dual-platform measurement (same workload on both: pipelines + adhoc, three BI tools, ~8TB scanned/day) dedicated a "Native GIS Performance" section: "BigQuery handled our geospatial queries significantly faster. Its native GIS functions and columnar storage for geography types are a genuine advantage for logistics, delivery, or any location-aware workloads. Snowflake supports geospatial queries but lags on performance for complex spatial joins at scale." The buying advice: "Choose BigQuery if… Geospatial or GIS workloads are part of your stack."
- 窄场景
物流、配送、位置相关的地理空间负载。
Logistics, delivery, and location-aware geospatial workloads.
- 机制
Snowflake 有 GEOGRAPHY 类型与 ST_ 函数,但在规模化复杂空间 join 的性能与函数完备度上不及 BigQuery GIS 的原生实现。
Snowflake has GEOGRAPHY types and ST_ functions, but trails BigQuery GIS's native implementation in at-scale complex spatial join performance and function completeness.
- 生产验证
来源 54:@aidelearning 30 天双平台生产 workload 实测,2026-03-31([来源存疑]:课程引流性质的内容营销,但测试方法与数字具体:8TB/天、$4,820 vs $3,940)。
—
- 证据等级
`单方声音`,[来源存疑]:主来源为课程引流文章;另有一条 2020 年 HN 评论侧面印证"地理空间长期是短板印象",时间较早未编号。
`Single voice`, [source questionable]: course-funnel article; a 2020 HN comment side-confirms the long-standing "geospatial is a weak spot" impression (old, unnumbered).
- 备注
缺口状态:至今缺失。本卡主题可能与本站 [避坑] 卡重叠。
gap status: still missing. This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
成本可观测性:BigQuery 按用户/标签查账单零配置,Snowflake 视图延迟 3 小时还要多表 join [来源存疑]
单方声音
来源存疑
运维复杂度
- 一句话
同一篇双平台实测:"BigQuery's INFORMATION_SCHEMA is significantly richer for cost attribution. The JOBS_BY_PROJECT view lets you attribute costs to users, labels, and time windows with zero configuration. Snowflake has ACCOUNT_USAGE views, but they're delayed by up to 3 hours and require more joins to get useful breakdowns."——"feeding directly into cost dashboards… often the first thing that changes analyst behavior when they can see their own tab"(直接喂给成本看板;当分析师能看到自己的账单时,行为往往最先改变)——Snowflake 侧要做到同样效果得多花几小时搭 pipeline。
Same dual-platform measurement: "BigQuery's INFORMATION_SCHEMA is significantly richer for cost attribution. The JOBS_BY_PROJECT view lets you attribute costs to users, labels, and time windows with zero configuration. Snowflake has ACCOUNT_USAGE views, but they're delayed by up to 3 hours and require more joins to get useful breakdowns." — "feeding directly into cost dashboards… often the first thing that changes analyst behavior when they can see their own tab": on Snowflake, the same costs hours of pipeline building.
- 窄场景
想做"按用户/标签"成本归因、推动分析师为自己账单负责的 FinOps 团队。
FinOps teams wanting per-user/per-label cost attribution that makes analysts own their spend.
- 机制
这不是"不能做",是"原生能力缺口":BigQuery 原生就有、Snowflake 得自己造;ACCOUNT_USAGE 延迟(文档口径最长 3 小时)至今仍是现状。
Not "impossible" but a native-capability gap: BigQuery ships it, Snowflake makes you build it; ACCOUNT_USAGE latency (up to 3 hours per docs) remains current.
- 生产验证
来源 54:@aidelearning 30 天双平台实测,2026-03-31([来源存疑]:课程引流性质,但测试方法与数字具体)。
—
- 证据等级
`单方声音`,[来源存疑]。
`Single voice`, [source questionable].
- 备注
缺口状态:至今缺失。本卡主题可能与本站 [避坑] 卡重叠。
gap status: still missing. This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
DDL 版本控制至今没有:ALTER 一下,旧结构就永远没了
单方声音
运维复杂度
- 一句话
"Snowflake does not natively version your object definitions."(Snowflake 不对你的对象定义做原生版本管理。)"If you `ALTER TABLE` something, the old structure is gone. If a stored procedure gets replaced at 2am, there is no rollback."——作者团队亲历:生产表上季度丢了一个列,"Nobody knew when. Nobody knew who. It took three engineers half a day to piece together what happened."(没人知道什么时候、没人知道谁干的,三个工程师花了半天才拼凑出真相。)
"Snowflake does not natively version your object definitions." "If you `ALTER TABLE` something, the old structure is gone. If a stored procedure gets replaced at 2am, there is no rollback." The author's team lived it: a column vanished from a prod table last quarter — "Nobody knew when. Nobody knew who. It took three engineers half a day to piece together what happened."
- 窄场景
多人协作、多环境、需要审计"谁改了表结构"的团队。
Multi-person, multi-environment teams needing to audit "who changed the table structure."
- 机制
讽刺的是 2025 年 11 月 Snowflake 的 Git 集成才 GA,而 Git 集成只管代码文件同步——对象定义的 DDL 历史至今仍要靠自建 task 定时扫 `ACCOUNT_USAGE` 再 push 到 GitHub 来补,作者为此写了一整套 nightly 捕获+AI 摘要的管线。
Ironic: Snowflake's Git integration only went GA in Nov 2025, and it syncs code files only — object DDL history still requires a homegrown task sweeping `ACCOUNT_USAGE` nightly and pushing to GitHub; the author built an entire nightly-capture + AI-summary pipeline for it.
- 生产验证
来源 55:独立从业者 Nimish Nagpal 个人博客,2026-08——一线数据工程师,为自家 Snowflake 环境自建 DDL 版本捕获管线。
—
- 证据等级
`单方声音`,一线工程师实战记录(自建管线的存在本身就是缺口的证据)。
`Single voice`, frontline engineer field record (the homemade pipeline is itself evidence of the gap).
- 备注
缺口状态:至今缺失(Git 集成 GA 只解决代码文件同步,不解决对象 DDL 历史版本管理)。本卡主题可能与本站 [避坑] 卡重叠。
gap status: still missing (Git integration GA covers code-file sync, not object DDL history). This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
元数据接口残缺:想导出一份完整 DDL,得自己写 600 行存储过程
单方声音
运维复杂度
- 一句话
数据架构师为客户做 Snowflake 环境现代化,被迫手写 610 行 Python 存储过程导出整库 DDL,因为原生能力处处是洞:"Snowflake's metadata surface is inconsistent across object types, SHOW commands have silent truncation limits that can break downstream logic"——`GET_DDL('DATABASE', ...)` 导出的视图没有依赖排序、直接跑就报错,还漏掉 stage、task、pipe、database role、grant 以及 Notebook、Semantic View 等新对象;`GET_DDL` 根本不支持 STAGE;`SHOW PROCEDURES` 的参数列静默截断长参数列表,拿到的签名是残的——被迫改从 `INFORMATION_SCHEMA` 取 16MB TEXT 完整签名再手工清洗。
A data architect modernizing a client's Snowflake environment hand-wrote a 610-line Python stored procedure to export full database DDL, because native capabilities are riddled with holes: "Snowflake's metadata surface is inconsistent across object types, SHOW commands have silent truncation limits that can break downstream logic" — `GET_DDL('DATABASE', ...)` emits views with no dependency ordering (runs fail), and omits stages, tasks, pipes, database roles, grants, plus newer objects like Notebooks and Semantic Views; `GET_DDL` doesn't support STAGE at all; `SHOW PROCEDURES` silently truncates long parameter lists, so captured signatures are broken — forcing a fallback to 16MB TEXT signatures from `INFORMATION_SCHEMA` with manual cleanup.
- 窄场景
想把 Snowflake 纳入版本控制/IaC 的团队;做环境迁移、审计的架构师。
Teams bringing Snowflake under version control/IaC; architects doing environment migrations and audits.
- 机制
元数据接口在对象类型间不一致 + SHOW 命令静默截断 + 无依赖排序——三者叠加让"导出完整 DDL"这件本该原生的事变成 600 行手工作业。
Inconsistent metadata surface across object types + silent SHOW truncation + no dependency ordering — together they turn "export complete DDL," which should be native, into 600 lines of hand-rolled work.
- 生产验证
来源 56:数据架构师 Andy Brown 个人博客,2026-03-31——客户现场实战长文(22 分钟),细节极充分。
—
- 证据等级
`单方声音`,客户现场实战记录(610 行存储过程的存在本身就是缺口的证据)。
`Single voice`, client-site field record (the 610-line procedure is itself evidence of the gap).
- 备注
缺口状态:至今缺失(2026-03 实测仍成立)。本卡主题可能与本站 [避坑] 卡重叠。
gap status: still missing (measured Mar 2026). This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2026
升级工程税:三分之一的升级会出意外,兼容级别"手刹"常年不松
单方声音
升级迁移
- 一句话
SQL Server 升级本身只要 20 分钟,周围的工程要几周——第三方软件不认新版本、兼容级别升完不敢提、"手刹"一拉就是几个月。
The SQL Server upgrade itself takes 20 minutes; the surrounding project takes weeks — third-party vendors won't certify the new version, and the compatibility level stays un-raised ("handbrake on") for months.
- 窄场景
有 ISV/ERP/医疗系统等第三方应用的生产库;in-place 升级后直接把兼容级别提到最新的团队。
Production databases with ISV/ERP/hospital-system applications; teams that upgrade in place and raise the compatibility level on day one.
- 机制
升级真正的杀手是组织性的:ISV 没给新版本做认证(有的 HIS 系统甚至硬编码了 SQL Server 版本检查,2016 以上直接报错);升级后数据库保留旧兼容级别——新优化器特性(CE、PSP)全不生效,等于"拉着手刹开车";而一旦提兼容级别,计划回归又要靠 Query Store 盯一周。作者的经验数字:约三分之一的升级会出点意外("something unexpected happens on roughly one in three upgrades"),所以他默认推荐 side-by-side 而非 in-place——旧服务器留着才能秒回滚。
The real upgrade killers are organizational: ISVs slow to certify new versions (some hospital information systems hard-code a SQL Server version check and error out on anything newer than 2016); after upgrade, databases keep their old compatibility level — new optimizer features (CE, PSP) stay off, i.e. "driving with the handbrake on"; raising it later means watching Query Store for plan regressions for a week. The author's field number: something unexpected happens on roughly one in three upgrades, which is why he defaults to side-by-side migration over in-place — the old server must still exist for instant rollback.
- 生产验证
来源 6,2026-04 实战指南——作者列出升级失败的常见真因:缺厂商书面确认、没 staging 环境(第一次测试就是生产)、忘了 SSIS 包/Reporting Services/CLR 依赖、兼容级别升完"手刹"数月;并给出完整 pre/during/post 检查清单。
Source 6, Apr 2026 field guide — lists the real-world reasons upgrades fail: missing vendor sign-off ("don't rely on 'it should work'"), no staging environment (first test is production), forgotten dependencies (SSIS packages, Reporting Services, CLR assemblies needing recompilation), compatibility level left at the old value for months; plus a full pre/during/post checklist.
- 证据等级
`单方声音`,个人博客(细节充分:三升级路径对比、检查清单、1/3 意外率的现场经验)。
`Single voice`, personal blog (detailed: three upgrade methods compared, checklist, the one-in-three field estimate).
Microsoft SQL Server 年份:2026
Always On 只读路由:配了一半,读流量全砸在主库上
单方声音
运维复杂度
- 一句话
把副本设成可读、连接串加上 ApplicationIntent=ReadOnly,还差一步——主库上的只读路由列表没建,所有读请求默默回落到主库,直到报表超时才有人发现。
You marked the secondaries readable and added ApplicationIntent=ReadOnly to the connection string — but missed the read-only routing list on the primary, so every read silently fell back to the primary until the reporting dashboards timed out.
- 窄场景
用 AG 可读副本做读扩展(read-scale)的团队;照着向导点完"允许读连接"就以为完事的。
Teams using AG readable secondaries for read scale-out; anyone who clicked "allow read connections" in the wizard and assumed the job was done.
- 机制
只读路由是两个独立配置对象:副本侧的"可读 + 路由 URL",和主库侧的"只读路由列表"。SSMS 向导不会强制你走完第二步;没有路由列表时,带读意图的连接直接落在主库上按读写执行——静默失败,无任何报错。更坑的是路由列表定义在"当前主库"上,故障转移后新主库也得有一份,每副本都要配一遍;负载均衡还是按连接轮询,不看副本延迟,副本 redo 落后 40 秒它也照样把查询扔过去。
Read-only routing is two independent configuration objects: the secondary's "readable + routing URL," and the primary's "read-only routing list." The SSMS wizard doesn't force you through the second step; without a routing list, read-intent connections land on the primary and run read-write — silently, with no error. Worse, the routing list lives on whichever replica is currently primary, so after a failover the new primary needs its own copy — every replica must be configured. And load balancing is per-connection round-robin with no regard for replica lag: a secondary 40 seconds behind on redo still gets queries sent to it.
- 生产验证
来源 7,2026-09 顾问复盘——客户凌晨 6 点被叫醒:报表 dashboard 超时,主库被所有读查询打满,两个健康的只读副本"除了 redo 啥也没干"。排查发现数月前配的 read-scale 漏了主库路由列表;作者称这是"AG 读扩展配置里最常见的错误"。
Source 7, Sep 2026 consultant postmortem — a client paged at 6am: reporting dashboards timing out, the primary hammered by every read query while two healthy readable secondaries "did nothing but redo." Investigation found the read-scale setup from months earlier had never built the routing list on the primary; the author calls it the single most common mistake in AG read-scale setups.
- 证据等级
`单方声音`,个人博客(细节充分:6 点叫醒、排查过程、DMV 验证查询)。
`Single voice`, personal blog (detailed: the 6am page, investigation, DMV verification query).
- 备注
该文还提醒——可读副本不是免费的,除被动故障转移配置外都要按正常 SQL Server 许可;Standard 版的 Basic AG 根本不支持可读副本。与 [吐槽] 按核许可卡交叉引用。
the post also reminds that readable secondaries are not free — outside a passive failover-only configuration they need full SQL Server licenses; Standard Edition's Basic AG doesn't support readable secondaries at all. Cross-reference with the per-core licensing rant card.
Microsoft SQL Server 年份:2026
Always On + 加密:自动种子跳过主密钥密码,故障转移后解密失效
单方声音
稳定与故障运维复杂度
- 一句话
用自动种子(automatic seeding)把加密库加进 AG,向导会跳过数据库主密钥密码校验——故障转移后新主库打不开 DMK,应用的加解密直接报错。
Add an encrypted database to an AG with automatic seeding and the wizard skips the Database Master Key password validation — after failover the new primary can't open the DMK and the application's encryption/decryption fails.
- 窄场景
用对称密钥/证书做列级加密、且跑在 AG 上的库;用自动种子而非"完整备份+日志备份"方式加副本的团队。
Databases using symmetric keys/certificates for column-level encryption inside an AG; teams joining replicas via automatic seeding instead of "full backup + log backup."
- 机制
SQL Server 加密链是 SMK(服务主密钥,实例级)→ DMK(数据库主密钥)→ 证书 → 对称密钥。DMK 默认同时被本地 SMK 加密保护;故障转移后新主库的 SMK 变了,打不开 DMK——除非每个副本上都用 `sp_control_dbmasterkey_password` 存一份 DMK 密码凭据。坑在于:AG 向导选"完整备份+日志备份"做初始同步时会自动建这个凭据,选自动种子时却跳过密码校验这一步,结果就是"配的时候一切正常,第一次故障转移才爆"。
SQL Server's encryption hierarchy is SMK (instance-level Service Master Key) → DMK (Database Master Key) → certificate → symmetric key. The DMK is by default also encrypted by the local SMK; after failover the new primary has a different SMK and can't open the DMK — unless every replica stores the DMK password via `sp_control_dbmasterkey_password`. The trap: the AG wizard's "Full Database and Log Backup" seeding path creates that credential automatically, while automatic seeding skips the password-validation step — so everything looks fine until the first failover.
- 生产验证
来源 8,2026-06 咨询公司客户案例复盘——作者完整复现:自动种子加库后故障转移,应用无法加解密;逐副本执行 `sp_control_dbmasterkey_password … @action='add'` 后恢复;并追问"为什么自动种子不做备份方式会自动做的事",表示自己也无法解释。
Source 8, Jun 2026 consulting-firm client case — the author fully reproduced it: after automatic seeding and a failover, the application could no longer encrypt/decrypt; running `sp_control_dbmasterkey_password … @action='add'` on each replica restored it; he closes by asking why automatic seeding doesn't do what the backup-based path does automatically — "I have no clue yet."
- 证据等级
`单方声音`,咨询公司博客(细节充分:客户案例 + 完整复现步骤 + 两种种子方式的对照实验)。
`Single voice`, consulting-firm blog (detailed: client case + full reproduction steps + A/B comparison of the two seeding methods).
Microsoft SQL Server 年份:2026
Always On 滚动升级八坑:先 drain 后挂起,代价是主库停写
单方声音
运维复杂度升级迁移
- 一句话
升级 AG 集群时自动化脚本先 drain 节点再挂起数据移动——顺序反了,同步提交的主库等一个正在被抽走的副本确认,写请求在主库上卡了两三分钟,而你动的明明是备库。
The automation drained the cluster node before suspending data movement — the wrong order — so a synchronous-commit primary stalled writes for two to three minutes waiting on a replica being pulled out from under it, on a node nobody touched.
- 窄场景
同步提交 + 自动故障转移的 AG;做 OS/集群滚动升级、按直觉写"先排空节点再停复制"的自动化。
Synchronous-commit + automatic-failover AGs; rolling OS/cluster upgrades with automation written in the intuitive "drain the node, then stop replication" order.
- 机制
同步提交 AG 里主库提交要等备库硬化日志。drain 节点时数据移动还是 active + synchronous,主库的提交就卡在"等一个正在被拔走的副本"上。正确顺序是先 `SUSPEND` 数据移动(把主库从该副本解耦)再 drain——但前提是 `required_synchronized_secondaries_to_commit=0`,否则挂掉唯一的同步备库反而会堵死主库提交,错上加错。同一次升级还踩了:统计作业在故障转移中途被 kill、回滚卡在 0%,四个库 NOT SYNCHRONIZING;两节点 CU 差一个版本,备库 build 低于主库不支持 redo;26GB 的 Windows CU 差点撑爆系统盘;Microsoft Update 通道会"顺手"装上没计划的 SQL CU。
In a synchronous-commit AG the primary doesn't acknowledge a commit until the secondary hardens the log. Draining the node while data movement is still active and synchronous leaves the primary waiting on a replica being removed — commits stall until the suspend finally lands. The correct order is SUSPEND data movement first, then drain — but only if `required_synchronized_secondaries_to_commit=0`; otherwise suspending your only synchronous secondary blocks the primary's commits instead, the opposite mistake. The same upgrade also hit: a stats job mid-run during failover got killed with rollback stuck at 0%, leaving four databases NOT SYNCHRONIZING; replicas one CU apart, where a lower-build secondary can't redo a higher-build primary's log; a 26GB Windows cumulative update nearly filling the OS drive; and the Microsoft Update channel silently installing an unplanned SQL Server CU.
- 生产验证
来源 9,2026-09 生产实录——两节点 SQL Server 2022 Enterprise AG(AWS+GCP 混合),WS2022→WS2025 原地滚动升级,8 个坑逐个记录命令与修复;作者强调"两步都在脚本里、两步都'成功'了,只有顺序是错的,而伤害发生在根本没碰过的主库上"。
Source 9, Sep 2026 production account — a two-node SQL Server 2022 Enterprise AG (mixed AWS+GCP fleet), WS2022-to-WS2025 in-place rolling upgrade, all eight surprises documented with commands and fixes; the author stresses "both steps were in the script, both 'worked' — only the order was wrong, and the damage happened on the primary, a node we hadn't touched at all."
- 证据等级
`单方声音`,个人博客(细节充分:8 个坑各有现象、命令与修复,生产环境实录)。
`Single voice`, personal blog (detailed: each of the eight surprises with symptoms, commands, and fixes; production account).
Microsoft SQL Server 年份:2026
Iceberg 分区缓存:查 10 个分区,先把 660 万个分区全 load 进堆内存
单方声音
性能问题稳定与故障
- 一句话
StarRocks 的 Iceberg Catalog 查分区是"全量加载再过滤"——`getPartitionsByNames()` 内部调 `getPartitions()`,查 10 个分区也要先把 660 万个分区对象塞进 FE 堆,24GB 堆直接 OOM,FE pod 崩溃重启。
StarRocks's Iceberg catalog loads partitions greedily — `getPartitionsByNames()` calls `getPartitions()` internally, so querying 10 partitions first stuffs all 6.6M partition objects into the FE heap; a 24GB heap OOMs and the FE pod crash-restarts.
- 窄场景
Iceberg 外表分区数达到百万级(用户实测 660 万+ 分区)、FE 堆 24GB、K8s 部署;任何带分区谓词的查询在规划期即触发。
Iceberg external tables with millions of partitions (user measured 6.6M+), 24GB FE heap, Kubernetes deployment; any query with a partition predicate triggers it during planning.
- 机制
三处叠加:① `IcebergCatalog.getPartitions()` 贪婪加载全部分区;② `partitionCache` 用 maximumSize(按表数限)而非按权重限,管不住单个大表的分区数,分区级内存无界增长;③ `getPartitionsByNames()` 内部直接调 `getPartitions()`,只想要 10 个分区也得先全量加载。发生在查询规划期(memo 阶段 37 秒超时),不是后台任务——后台 refresh 只是让情况更糟。用户还附了堆 dump 证据。
Three compounding flaws: (1) `IcebergCatalog.getPartitions()` eagerly loads every partition; (2) `partitionCache` is bounded by maximumSize (table count), not by weight, so per-partition memory grows unbounded for a single huge table; (3) `getPartitionsByNames()` calls `getPartitions()` internally — even 10 wanted partitions require a full load first. It happens during query planning (37s timeout in the memo phase), not in a background task — background refresh only makes it worse. The reporter attached heap-dump evidence.
- 生产验证
来源 2:GitHub issue #67760,2026-01,用户在 GKE 上 40GB pod(FE 堆 24GB)跑 3.5.9,对 660 万分区的 Iceberg 表做分区裁剪查询,堆涨到 20GB+ 后 `OutOfMemoryError: Java heap space`,FE pod 崩溃重启;用户给出三处代码级根因定位与修复建议。
Source 2: GitHub issue #67760, Jan 2026 — user on GKE, 40GB pod (24GB FE heap), StarRocks 3.5.9, partition-pruning queries against a 6.6M-partition Iceberg table grew the heap past 20GB into `OutOfMemoryError: Java heap space`, crashing and restarting the FE pod; the report includes code-level root-cause analysis and fix proposals.
- 证据等级
`单方声音`,GitHub 用户 issue(细节充分:版本号、堆大小、分区数、planner 超时日志、堆 dump 附件、代码级根因)。
`Single voice`, GitHub user issue (detailed: version, heap size, partition count, planner timeout logs, heap-dump attachment, code-level root cause).
StarRocks 年份:2026
嵌套物化视图刷新:内层 MV 新鲜度一未知,外层直接 NPE
单方声音
运维复杂度
- 一句话
异步物化视图套娃(外层 MV 依赖内层 MV)时,内层 MV 的分区新鲜度一旦是 FULL/UNKNOWN,外层刷新不降级、不报错,直接空指针异常失败。
With stacked async MVs (an outer MV depending on an inner MV), the moment the inner MV's partition freshness is FULL/UNKNOWN, the outer refresh doesn't degrade or error cleanly — it dies with a NullPointerException.
- 窄场景
多层嵌套的异步分区 MV(`REFRESH ASYNC` + `auto_refresh_partitions_limit` 等分区刷新参数),基表为内部表(Duplicate/PK 表)即可复现,无需外部 catalog。
Multi-level nested async partitioned MVs (`REFRESH ASYNC` with partitioned-refresh properties like `auto_refresh_partitions_limit`); reproducible with internal base tables (Duplicate/PK) alone, no external catalog needed.
- 机制
外层 MV 刷新前检查基表新鲜度,递归走到内层 MV 时,`MvRefreshArbiter.getMvBaseTableUpdateInfo()` 在 timeliness 为 FULL/UNKNOWN 时返回 null,而 `MVTimelinessArbiter.collectBaseTableUpdatePartitionNames()` 不做空检查直接解引用 → `NullPointerException`,刷新任务失败。保守做法应是降级为全量刷新或给明确错误,但当时是直接 NPE。临时规避:先手工刷内层 MV,再刷外层——依赖链的拓扑顺序要人肉保证。
Before refreshing, the outer MV checks base-table freshness and recurses into the inner MV; `MvRefreshArbiter.getMvBaseTableUpdateInfo()` returns null when timeliness is FULL/UNKNOWN, and `MVTimelinessArbiter.collectBaseTableUpdatePartitionNames()` dereferences it without a null check → `NullPointerException`, failing the refresh task. The conservative behavior would be degrading to a full refresh or a readable error; instead it is a raw NPE. Workaround: manually refresh inner MVs first, then the outer — the topological order of the dependency chain becomes the human's job.
- 生产验证
来源 3:GitHub issue #73716,2026-05,用户在 3.5.15 上报嵌套分区 MV 刷新 NPE,附完整堆栈、复现的 MV 定义与规避方法;关联 issue #71975 为查询改写路径的同类空指针问题。
Source 3: GitHub issue #73716, May 2026 — user on 3.5.15 reports nested partitioned MV refresh NPE with full stack trace, reproducing MV definition, and workaround; related issue #71975 is the same null-handling class of bug in the query-rewrite path.
- 证据等级
`单方声音`,GitHub 用户 issue(细节充分:版本、堆栈、复现配置、规避方案)。
`Single voice`, GitHub user issue (detailed: version, stack trace, reproducing config, workaround).
- 备注
修复 PR #73644 已提出,合并状态本次未核验,未作"已修复"标注;本卡主题可能与本站 [避坑] 卡重叠。
fix PR #73644 has been proposed; merge status not verified in this pass, so no "fixed in" label applied; this card's topic may overlap an existing [Pitfall] card on this site on this site.
StarRocks 年份:2026
"TDSQL" 到底指哪个:产品线命名与版本线混乱
单方声音
生态与信任
- 一句话
"TDSQL" 不是一款产品——同名之下有分库分表型(for MySQL,原 DCDB)、云原生型(TDSQL-C,原 CynosDB)、PG 版(TBase 系),连 PG 版内部都有 V2/V5.06/V5.21 三条内核线并行停售;选型第一步就可能买错。
"TDSQL" is not one product — under the same name sit a sharded type (for MySQL, formerly DCDB), a cloud-native type (TDSQL-C, formerly CynosDB), and a PG edition (TBase lineage); even the PG edition alone runs three kernel lines (V2/V5.06/V5.21) being retired in parallel. You can buy the wrong one at step one of selection.
- 窄场景
初次接触腾讯云数据库产品线的选型者;按"TDSQL 评测/对比"关键词做功课的团队。
First-time evaluators of Tencent Cloud's database line; teams doing homework via "TDSQL review/comparison" search keywords.
- 机制
品牌统一为 TDSQL 后,原 DCDB、原 CynosDB、TBase 共用名前缀,但架构(Shared Nothing 分片 vs 存算分离 vs PGXC 系)与兼容性完全不同。PG 版三条内核线:V2(纯 PG 兼容)2023.12.30 停售、V5.06(Oracle 兼容)2025.12.30 停售、V5.21(2024.06 起售)为当前主推;V2 老用户需走 DTS 迁到 V5.21。
After the brand consolidation, the former DCDB, former CynosDB, and TBase share the TDSQL prefix but differ completely in architecture (shared-nothing sharding vs. compute-storage separation vs. PGXC lineage) and compatibility. The PG edition's three kernel lines: V2 (pure PG compatibility) stopped selling Dec 30 2023, V5.06 (Oracle compatibility) stops selling Dec 30 2025, V5.21 (on sale since Jun 2024) is the current recommendation; V2 users must migrate to V5.21 via DTS.
- 生产验证
来源 1:colask 2026-05 选型调研——"本章解决一个常见误解:TDSQL 不是单一产品。市面上很多对比文章把'TDSQL'当作一个东西来写,得出'TDSQL 不支持触发器/存储过程'等结论。这只对 TDSQL for MySQL(分布式版)成立,对 TDSQL-C MySQL 版(云原生版)完全错误。混淆这点会导致选型彻底跑偏。"
Source 1: colask, May 2026 selection study — "This chapter addresses a common misconception: TDSQL is not a single product. Many comparison articles on the market treat 'TDSQL' as one thing and conclude 'TDSQL does not support triggers/stored procedures.' That holds only for TDSQL for MySQL (distributed edition) and is flat-out wrong for TDSQL-C MySQL Edition (cloud-native). Confusing the two derails selection entirely."
- 证据等级
`单方声音`,独立选型评估(细节充分:产品矩阵图 + 版本线梳理;官方《PG 版产品简介》仅作版本生命周期机制佐证)。
`Single voice`, independent selection evaluation (detailed: product-matrix diagram + version-line mapping; the official PG Edition product brief is mechanism corroboration for the version lifecycle only).
腾讯云 TDSQL 年份:2026
shardkey 是建表时的一次性赌注:选错无后悔药
单方声音
性能问题运维复杂度
- 一句话
shardkey 在建表时一次定终身——选错导致数据倾斜或跨分片查询,没有在线改键的后悔药,只能重建表、迁数据、切路由。
The shard key is fixed for life at CREATE TABLE time — pick wrong and get data skew or cross-shard queries, with no online key-change remedy; the only fix is rebuilding the table, migrating data, and switching routing.
- 窄场景
业务增长后发现分片键选择不当(热点商户打爆单分片、时间列做键导致热分片)的 TDSQL for MySQL 实例。
TDSQL for MySQL instances where growth exposed a bad shard-key choice (a hot merchant saturating one shard, a time column as key creating a hot shard).
- 机制
shardkey 决定 hash 路由与数据分布。官方明确:不支持 ALTER 对分表键改名;不要 UPDATE shardkey 列的值(会改变数据分布,需先 DELETE 再 INSERT);shardkey 值不应含中文(网关不转换字符集,不同字符集可能路由到不同分区);查询不带 shardkey 即全分片扫描+网关聚合。选错后的标准解法只有重建表迁数据。社区方言参考亦提醒:"shardkey 选择直接决定查询性能和数据分布均匀度"。
The shard key determines hash routing and data distribution. Official docs are explicit: ALTER cannot rename the shard key; do not UPDATE the shard-key column (it changes data distribution — DELETE then INSERT instead); shard-key values should not contain Chinese characters (the gateway does no charset conversion, so different charsets may route to different partitions); any query without the shard key fans out to all shards for a full scan plus gateway-side aggregation. The standard remedy for a bad choice is rebuilding the table and migrating data. The community dialect reference likewise warns that "shard-key choice directly determines query performance and data-distribution evenness."
- 生产验证
来源 1:colask 2026-05 选型调研——"需要在建表时规划 shardkey,否则会遇到数据倾斜与跨分片查询问题"。
Source 1: colask, May 2026 selection study — "you must plan the shard key at table-creation time, otherwise you will run into data skew and cross-shard query problems."
- 证据等级
`单方声音`,独立选型评估(官方文档仅作机制佐证;未找到具名的"选错后返工"生产复盘,机制层面风险明确,如实标注缺口)。
`Single voice`, independent selection evaluation (official docs are mechanism corroboration only; no named "picked-wrong-and-reworked" production postmortem was found — the gap is stated as-is while the mechanism-level risk is explicit).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site on this site.
腾讯云 TDSQL 年份:2026
流式 UPDATE:更新一行扫 31GB,每次作业还有 10MB 起步价
单方声音
成本账单
- 一句话
BigQuery 不是为行级更新设计的——改一行一列要重写整个列块;雪上加霜的是,每次作业还有最低计费,流式小更新等于按"起步价"论次收费。
BigQuery isn't built for row-level updates — changing one row rewrites whole column blocks; worse, every job has minimum billing, so streaming micro-updates are charged per-job "starting fares."
- 窄场景
把 BigQuery 当 OLTP 用、逐行流式更新维度表/状态表的 CDC 式写入;按量计费下高频小 UPDATE 的管道。
Treating BigQuery like an OLTP store with row-by-row streaming updates to dimension/status tables; high-frequency small UPDATE pipelines on on-demand pricing.
- 机制
列式存储以大压缩块存放数据,单行 UPDATE 触发所在列块的解压-修改-重写,扫描量远超同条件 SELECT(作者实测:取一行一列 SELECT 扫 667MB,同条件 UPDATE 扫 31.71GB);按量计费下每个查询作业有 10MB 最低计费(实际扫 6.1KB 也按 10MB 收);容量计费(slot)下按 slot-second 计费但有 1 分钟最低消费(实际跑 21 毫秒按 1 分钟收)。UPDATE 还不能用索引搜索,只能靠分区/聚簇减扫描。
Columnar storage keeps data in large compressed blocks, so a single-row UPDATE triggers decompress-modify-rewrite of the affected column blocks — scanning far more than an equivalent SELECT (the author measured: fetching one row/one column scanned 667MB; updating it scanned 31.71GB); on-demand pricing bills a 10MB minimum per query job (6.1KB actually scanned, billed as 10MB); capacity (slot) pricing bills per slot-second with a 1-minute minimum (a 21ms job billed as a full minute). UPDATEs also can't use index search — only partitioning/clustering reduce the scan.
- 生产验证
来源 5,Karanrat Rattanawichai 2025-09-27 实测记录——带截图的计费明细:6.1KB 实际扫描被收 10MB、21 毫秒作业被收 1 分钟 slot;结论是"不要流式 UPDATE 到 BigQuery",改用追加写入 + 窗口函数去重或批量 MERGE。
Source 5, Karanrat Rattanawichai, Sep 27 2025 — measured billing records with screenshots: 6.1KB of actual scan billed as 10MB, a 21ms job billed as 1 slot-minute; conclusion: "don't streaming-update BigQuery" — append instead and dedupe with window functions or batch MERGE.
- 证据等级
`单方声音`,个人博客(细节充分:带截图的实测计费数据、SELECT/UPDATE 对比扫描量)。
`Single voice`, personal blog (detailed: screenshot-backed billing measurements, SELECT vs UPDATE scan comparison).
Google BigQuery 年份:2025
动态分区过滤被优化器"不信任":2.5GB vs 85MB,30 倍账单差
单方声音
成本账单
- 一句话
同样的分区表,写死日期只扫 85MB,换成 `MAX(DATE())` 子查询动态取最新一天就扫 2.5GB——优化器不信任你算出来的"常量",分区裁剪直接失效。
Same partitioned table: a hardcoded date scans 85MB, but fetching "the latest day" via a `MAX(DATE())` subquery scans 2.5GB — the optimizer doesn't trust your computed "constant," so partition pruning silently fails.
- 窄场景
按小时/天分区的表 + 看板类"取最新一天数据"的动态过滤查询;把日期逻辑写成子查询/CTE 的 BI 查询。
Hour/day-partitioned tables with dashboard-style "latest available day" dynamic filters; BI queries that express date logic as subqueries/CTEs.
- 机制
分区裁剪只对静态字面量可靠;`WHERE DATE(ts) = (SELECT MAX(DATE(ts)) …)` 这类动态谓词下,优化器无法确信子查询结果可作为裁剪常量(作者实验:连从 INFORMATION_SCHEMA.PARTITIONS 取 MAX(partition_id) 再经 CTE 传入的"常量"都不被信任),退化为大范围扫描;更进一步,把这类子查询写进 JOIN ON 会直接报错 "Unsupported subquery with table in join predicate"。同一张表、同一语义,硬编码日期 85MB,动态写法 2.5GB。
Partition pruning is only reliable for static literals; with dynamic predicates like `WHERE DATE(ts) = (SELECT MAX(DATE(ts)) …)`, the optimizer can't treat the subquery result as a pruning-safe constant (the author's experiments: even a "constant" derived from INFORMATION_SCHEMA.PARTITIONS via CTE wasn't trusted), degrading to a broad scan; pushing such a subquery into a JOIN ON fails outright with "Unsupported subquery with table in join predicate." Same table, same semantics: 85MB hardcoded, 2.5GB dynamic.
- 生产验证
来源 6,Nitheesh Varma 2025-07-01 团队复盘——为 Looker Studio 看板做"最新可用一天"查询,逐轮实验记录扫描量:硬编码日期 ~85MB;`MAX(DATE())` 子查询稳定 2.5GB(约 30 倍);INFORMATION_SCHEMA + CTE 方案依然 2.5GB;最终靠"先 CTE 预过滤再 JOIN"的改写绕过。
Source 6, Nitheesh Varma's Jul 1 2025 team postmortem — building a "latest available day" query for a Looker Studio dashboard, with per-round scan measurements: hardcoded date ~85MB; the `MAX(DATE())` subquery consistently 2.5GB (~30x); the INFORMATION_SCHEMA + CTE route still 2.5GB; resolved by restructuring to pre-filter in a CTE before joining.
- 证据等级
`单方声音`,个人博客(细节充分:多轮对照实验的扫描量数据、报错原文)。
`Single voice`, personal blog (detailed: multi-round controlled experiments with scan numbers, verbatim error text).
Google BigQuery 年份:2025
ON CLUSTER DDL 在副本故障时长时间挂起:一个副本挂了,全集群 DDL 卡住
单方声音
稳定与故障运维复杂度
- 一句话
某个副本挂了,ON CLUSTER 操作"will take a lot of time",超时没配好应用跟着遭殃——表删不掉、副本"无理由"只读是家常便饭,修法是删表重建让复制自己恢复。
With a replica down, ON CLUSTER operations "will take a lot of time" — and if app timeouts aren't set right ( "trust me, they're not"), the app suffers too. Tables that won't drop and "unreasonably" read-only replicas are routine; the fix is to drop the table and let replication heal itself.
- 窄场景
多副本集群做 schema 变更;有副本故障未及时发现的集群。
Schema changes on multi-replica clusters; clusters with an unnoticed dead replica.
- 机制
ON CLUSTER 要等所有副本 ack,挂掉的副本让 DDL 长时间挂起;应用侧超时若没配好会被拖死。运维日常包括"表删不掉"和"副本无理由只读",标准修法简单粗暴:删表重建。
ON CLUSTER waits for every replica to ack; a dead replica stalls DDL for a long time, dragging misconfigured app timeouts with it. Daily ops include "can't drop this table" and "replica read-only for no reason"; the standard fix is blunt: drop and rebuild.
- 生产验证
—
Source 10: Tinybird official account series (CTO Javi Santana's team ops retrospective), Apr 2025 — "Be careful handling `ON CLUSTER` operations… If a replica is down, `ON CLUSTER` operations will take a lot of time (depending on the config) and may generate issues in your app if timeouts are not properly set (and trust me, they're not)."
- 证据等级
`单方声音`,具名大厂 CTO 级复盘。
`Single voice`, named senior-CTO-level retrospective.
ClickHouse 年份:2025
FINAL 是"代码坏味道":去重税从写入转嫁到每次查询
单方声音
性能问题
- 一句话
ReplacingMergeTree 用户曾给所有查询加 FINAL 当"安全网",结果"forces an extra merge for every query…On large datasets, it slowed queries down drastically"——FINAL 正确但慢,且并行度差。
A ReplacingMergeTree user once added FINAL to every query as a "safety net" — "forces an extra merge for every query…On large datasets, it slowed queries down drastically." FINAL is correct but slow and parallelizes poorly.
- 窄场景
ReplacingMergeTree 去重表;"查询不加 FINAL 就不对"的团队。
ReplacingMergeTree dedup tables; teams where "queries are wrong without FINAL."
- 机制
FINAL 在查询期做去重合并,无法很好并行;大数据集上每个查询都付一次合并税。结论是把 FINAL 当调试工具——"If your queries always need FINAL, the problem is your data model"("FINAL is a code smell")。
FINAL performs dedup merging at query time with poor parallelism; on large datasets every query pays a merge tax. The conclusion: treat FINAL as a debugging tool — "If your queries always need FINAL, the problem is your data model" ("FINAL is a code smell").
- 生产验证
—
Source 16: dev.to personal author, circa Jul 2025 — production lesson.
- 证据等级
`单方声音`。
`Single voice`.
ClickHouse 年份:2025
没有 time travel:原生表误删,只能去翻备份
单方声音
运维复杂度
- 一句话
Tinybird CTO:"有两种数据工程师:一种是曾经误删过表的,另一种是将来会误删的。如果你误删了一张表,就只能去翻备份了,那真是个 pain in the ass."
Tinybird's CTO: "There are two kinds of data engineers: those who have dropped a table by accident, and those who will. If you drop a table by accident, you're going to backups — and that's a pain in the ass."
- 窄场景
误删表/误删分区后的恢复;想回看历史快照。
Recovery after mistaken drops; wanting historical snapshots.
- 机制
Time travel 仅 Iceberg 外表支持(25.4+);原生 MergeTree 表截至 2026-10 无 time travel。`UNDROP TABLE`(23.3 起实验性)只能救"误删表"这一种场景,不是通用时间回溯;行级误删、历史快照回看仍只能靠备份。
Time travel exists only for Iceberg external tables (25.4+); native MergeTree tables have none as of Oct 2026. `UNDROP TABLE` (experimental since 23.3) rescues exactly one scenario — the dropped table — not general time-travel; row-level mistakes and snapshot lookbacks still mean backups.
- 生产验证
—
Source 10: Tinybird CTO, Apr 2025 production blog, "Table deletion" section.
- 证据等级
`单方声音`。
`Single voice`.
ClickHouse 年份:2025
DBR 版本同步:升级本身是一个项目
单方声音
升级迁移
- 一句话
DBR 大版本一跳,集群配置、依赖库、Spark Connect 行为全要重测——"跟上版本"本身就是一个常年项目。
Jump a DBR major version and cluster configs, libraries, and Spark Connect behavior all need retesting — "staying current" is a permanent project.
- 窄场景
长期运行的老工作区;从旧 LTS 向新 LTS 迁移的团队。
Long-lived workspaces; teams migrating from old LTS releases to new ones.
- 机制
Databricks Runtime 捆绑 Spark、Delta、Python、Java 大版本,LTS 之间跳跃常伴随破坏性变更(Python 小版本、Spark Connect 行为、UC 强制要求);集群与代码深度耦合 runtime,升级等于全量回归。
Databricks Runtime bundles major versions of Spark, Delta, Python, and Java; LTS-to-LTS jumps routinely carry breaking changes (Python minor versions, Spark Connect behavior, UC enforcement); clusters and code are deeply coupled to the runtime, so an upgrade means full regression.
- 生产验证
来源 7,PeerSpot 具名复评——Parag Bhosale(物流公司 Senior Data Engineer,2025-01):"We need to stay in sync with the DVR versions, and migrations can pose challenges. For example, issues arose when we moved a cluster from a previous version to the latest one."(引用原话)
Source 7, PeerSpot named review — Parag Bhosale (Senior Data Engineer at a logistics company, Jan 2025): "We need to stay in sync with the DVR versions, and migrations can pose challenges. For example, issues arose when we moved a cluster from a previous version to the latest one." (direct quote)
- 证据等级
`单方声音`,具名用户复评(细节充分:版本同步→迁移挑战→具体升级出问题)。
`Single voice`, named user review (detailed: version sync → migration challenges → specific upgrade incident).
Databricks 年份:2025
真实故障问 RCA,答复是"请提工单":事故透明度不足
单方声音
稳定与故障生态与信任
- 一句话
AWS us-east-1 上 Classic/Serverless/DLT/Jobs 曾全量不可用约 50 分钟,客户在社区公开索要 RCA,官方回复"请向 Support 提工单"——没有公开事故复盘。
Classic/Serverless/DLT/Jobs once went fully unavailable in AWS us-east-1 for ~50 minutes; a customer publicly asked for an RCA in the community and the official reply was "please file a support ticket" — no public postmortem.
- 窄场景
把 Databricks 当核心生产依赖、需要事故复盘做灾备设计的团队。
Teams running production on Databricks that need incident reviews for DR planning.
- 机制
控制面集中化意味着区域性故障是"全量不可用"而非降级;第三方状态页镜像记录 2026-09 三起事故(ES-2231583 us-east-1 50 分钟全不可用、Azure East US 降级、AWS DLT 降级 2.5 小时),但官方无公开 RCA 流程。
A centralized control plane means a regional incident is "fully unavailable," not degraded; a third-party status-page mirror logged three September 2026 incidents (ES-2231583: 50 min full outage in us-east-1; Azure East US degradation; 2.5h AWS DLT degradation) with no public RCA process from the vendor.
- 生产验证
—
Source 19: official community thread, May 2025 — a customer publicly requested an RCA after the May 15 outage; the official response was "file a ticket."
- 证据等级
`单方声音`,客户原声帖(事故事实有第三方状态页交叉记录)。
`Single voice`, customer-voiced thread (incident facts cross-checked against a third-party status mirror).
Databricks 年份:2025
SQL warehouse 并发排队:单集群约 10 并发,扩缩还要先等 5 分钟
单方声音
性能问题
- 一句话
一个 SQL warehouse 同时只能跑约 10 个查询,第 11 个进排队;Classic/Pro 的自动扩缩是"慢规则"——排队的查询要先等 5 分钟扩缩才开始。
One SQL warehouse runs about 10 concurrent queries; the 11th queues. Classic/Pro autoscaling runs on "slow rules" — a queued query waits 5 minutes before scale-out even begins.
- 窄场景
Power BI 同一身份多报表并发;突发流量打到 Classic/Pro 仓库。
Power BI with many concurrent reports on one identity; bursty traffic hitting Classic/Pro warehouses.
- 机制
单集群并发上限约 10,超了进 queuing state,官方建议"每 10 个并发查询配 1 个集群"。Classic 的自动扩缩:排队查询等 5 分钟才开始扩,缩容要连续 15 分钟低负载——"没有什么东西会在配置与负载错配的那一刻提醒你"。
Per-cluster concurrency caps around 10; beyond that queries enter queuing state ("one cluster per 10 concurrent queries" is the official guidance). Classic autoscaling: queued queries wait 5 minutes before scale-out starts; scale-in needs 15 continuous minutes of low load — "nothing alerts you the moment config and load mismatch."
- 生产验证
—
Source 21: official community thread, circa 2022 — Power BI concurrency test: "the first 10 queries run concurrently, the 11th must wait";
Source 22: Keebo blog, Dec 2025 — "On a classic warehouse, a queued query can wait five minutes before scaling even begins" [Questionable source: vendor selling optimization tooling, but the scaling rules are confirmed by official docs].
- 证据等级
`单方声音`,社区实测 + 厂商博客(机制有官方文档交叉证实)。
`Single voice`, community measurement + vendor blog (mechanism cross-confirmed by official docs).
Databricks 年份:2025
递归 CTE 缺失到 2025-08:求助帖的 workaround 是"用 Scala 手写循环"
单方声音
生态与信任
- 一句话
2025 年 7 月还有用户在社区求递归 CTE,唯一的 workaround 是 Scala 手写迭代循环(迭代次数 hard-code 10)——次月官方才 GA。
In Jul 2025 users were still asking the community for recursive CTEs; the only workaround was a hand-written Scala iteration loop (iteration count hard-coded to 10) — GA came the next month.
- 窄场景
组织层级、物料 BOM、图遍历等递归查询;从 Postgres/Snowflake 迁移的 SQL。
Org hierarchies, BOM explosions, graph traversals; SQL ported from Postgres/Snowflake.
- 机制
Databricks SQL 直到 Databricks SQL 2025.25(2025-08-20 rollout)/ DBR 17.2+ 才 GA 递归 CTE。在此之前 SQL 层面无解,只能换语言手写循环。
Databricks SQL had no recursive CTEs until Databricks SQL 2025.25 (rollout from Aug 20, 2025) / DBR 17.2+. Before that there was simply no SQL-side answer — you switched languages and hand-rolled the loop.
- 生产验证
—
Source 36: community thread, circa Jul 8, 2025 — plea → Scala workaround ("If there are more levels, you need to adjust") → GA the next month; a complete evidence chain.
- 证据等级
`单方声音`,单帖但细节充分。
`Single voice`, one thread but a complete story.
Databricks 年份:2025
原生 GIS 缺失多年,2025-07 才补上:之前只能装 Sedona 或手写 UDF
单方声音
生态与信任
- 一句话
DBR 17.1(2025-07)之前 Databricks 无原生地理空间类型与 ST_* 函数,空间分析只能装 Apache Sedona/Mosaic 或手写 UDF——" significantly slower" 且有运维负担。
Before DBR 17.1 (Jul 2025) Databricks had no native geospatial types or ST_* functions — spatial analytics meant installing Apache Sedona/Mosaic or hand-writing UDFs: "significantly slower" with maintenance overhead.
- 窄场景
LBS、风控反欺诈、物流等空间分析场景。
Location-based services, fraud detection, logistics spatial analytics.
- 机制
2025-07 前无 GEOMETRY/GEOGRAPHY 原生类型;第三方库要装 cluster library(维护负担),UDF 方案慢。DBR 17.1 引入原生类型,2025.25(2025-08)新增 80+ spatial SQL 表达式。
No native GEOMETRY/GEOGRAPHY types before Jul 2025; third-party libraries meant cluster-library maintenance, UDFs meant slow. DBR 17.1 added native types; 2025.25 (Aug 2025) added 80+ spatial SQL expressions.
- 生产验证
—
Source 43: Marian Reuss, Nov 2025 retrospective — "Prior to DBR 17.1, users had to resort to installing third-party tooling… These approaches come with additional maintenance overhead and are, in the case of UDFs, significantly slower."
- 证据等级
`单方声音`,事后回顾性描述(缺口期直接抱怨帖未找到)。
`Single voice`, retrospective account (no contemporary complaint thread found from the gap period).
Databricks 年份:2025
Time travel 的版本"保不保得住"要自己算:默认 VACUUM 7 天静默清历史
单方声音
运维复杂度
- 一句话
Delta time travel 能回看几个版本不是平台保证的——默认 VACUUM 7 天清理历史版本,"用了几年的代码突然报 `Cannot time travel…Available versions: [3, 23]`"。
How many Delta versions you can time-travel isn't a platform guarantee — the default 7-day VACUUM cleans history silently: "code that worked for years suddenly throws `Cannot time travel…Available versions: [3, 23]`."
- 窄场景
依赖 time travel 做审计/回滚的 pipeline;长期运行的老代码。
Pipelines relying on time travel for audit/rollback; long-running legacy code.
- 机制
历史版本保留取决于三个表属性(`delta.logRetentionDuration` / `delta.deletedFileRetentionDuration` 等),默认 7 天 VACUUM 静默清理。想保长期历史得手动调参——"可配置但默认坑人"。
History retention depends on three table properties (`delta.logRetentionDuration` / `delta.deletedFileRetentionDuration`, …); the 7-day default VACUUM silently clears old versions. Keeping long history means hand-tuning — "configurable but the default bites."
- 生产验证
—
Source 45: MS Q&A user Stephen01, Feb 2025 — "This behavior is unexpected, as the code has worked reliably until now." A moderator confirmed the 7-day VACUUM root cause with the manual tuning recipe.
- 证据等级
`单方声音`。
`Single voice`.
Databricks 年份:2025
Flink 写 Doris 超时:FE 给的是 BE 私网地址,报错只说 timeout
单方声音
来源存疑
运维复杂度
- 一句话
Flink SQL 写入 Doris 时,FE 返回的 BE 地址是内网 IP,Flink 侧直连超时;报错只有一句 `Connection timed out`,不告诉你真正连的是谁。
When Flink SQL writes to Doris, the FE returns BE addresses as private IPs; Flink's direct connections time out, and the error message is a bare `Connection timed out` that never tells you who it actually tried to reach.
- 窄场景
Flink 与 Doris 不在同一内网(跨网段/VPC),用 flink-doris-connector 做 Kafka→Doris 实时写入。
Flink and Doris on different networks (cross-subnet/VPC), using the flink-doris-connector for Kafka→Doris streaming writes.
- 机制
Flink connector 先问 FE 要 BE 节点列表,FE 按自身配置解析出 BE 的私网 IP 返回;connector 再直连 BE 的 8040 端口做 Stream Load。跨网段时私网 IP 不可达 → `HttpHostConnectException: Connection timed out`。报错只给超时,不暴露"FE 返回的地址不可达"这一层信息,排查要靠读 connector 源码/官方文档才知道要在 with 参数里配 `benodes` 显式指定可达地址。
The Flink connector first asks the FE for the BE node list; the FE resolves BE addresses from its own config and returns private IPs. The connector then connects directly to BE port 8040 for Stream Load. Across network segments the private IPs are unreachable → `HttpHostConnectException: Connection timed out`. The error carries no hint of the address-resolution layer — finding out you must set `benodes` in the connector's WITH options requires reading connector source or docs.
- 生产验证
来源 5:CSDN 个人博客(2025-10-10)——Kafka 经 Flink SQL 入 Doris,BE 明明存活且监听正常,始终 `Connect to ***:8040 failed: Connection timed out`;按官方文档解释在 with 参数加 `'benodes'` 后解决。
Source 5: CSDN personal blog (Oct 10 2025) — Kafka data written to Doris via Flink SQL; the BE was alive and listening, yet every write failed with `Connect to ***:8040 failed: Connection timed out`; resolved by adding `'benodes'` to the connector WITH options per the official docs.
- 证据等级
`单方声音 [来源存疑]`,CSDN 匿名个人博客(内容具体:完整报错栈、日期、解法;但作者归属与独立性无法确认)。
`Single voice [Questionable source]`, anonymous CSDN personal blog (concrete: full error stack, date, fix; but author attribution and independence cannot be confirmed).
Apache Doris 年份:2025
Flink CDC 同步 MySQL 到 Doris:类型映射处处是坑
单方声音
来源存疑
运维复杂度
- 一句话
MySQL 的 JSON/ENUM/TIMESTAMP(6)/BIT 在 Doris 里没有直接对应,Flink CDC 同步要么写入失败、要么小数被截断、要么时间整体偏移 8 小时。
MySQL's JSON/ENUM/TIMESTAMP(6)/BIT have no direct Doris counterparts — Flink CDC syncs either fail to write, silently truncate decimals, or shift every timestamp by 8 hours.
- 窄场景
用 Flink CDC + Debezium 把 MySQL 同步到 Doris 做实时数仓。
Using Flink CDC + Debezium to sync MySQL into Doris for a real-time warehouse.
- 机制
Doris 是严格关系模型,类型系统与 MySQL 不对齐:JSON/ENUM 无原生对应(需转 VARCHAR/STRING);Doris TIMESTAMP 默认秒级精度,MySQL TIMESTAMP(6) 微秒写入会失败或丢精度;DECIMAL 精度/标度不一致会被截断;BIT(1) 需显式转换;Flink 默认 UTC 而 MySQL 按 Asia/Shanghai 存 TIMESTAMP,不统一时区整列偏移 8 小时。Debezium 的 snapshot.mode、inconsistent.schema.handling.mode 配错还会导致 schema 变更不同步、类型解析失败。
Doris is a strict relational model whose type system does not align with MySQL's: JSON/ENUM have no native equivalents (must be mapped to VARCHAR/STRING); Doris TIMESTAMP defaults to second precision, so MySQL TIMESTAMP(6) microsecond writes fail or lose precision; mismatched DECIMAL precision/scale gets truncated; BIT(1) needs explicit conversion; Flink defaults to UTC while MySQL stores TIMESTAMP in Asia/Shanghai — without a unified timezone the whole column shifts 8 hours. Misconfigured Debezium snapshot.mode / inconsistent.schema.handling.mode additionally breaks schema-change propagation and type resolution.
- 生产验证
来源 6:CSDN 个人博客(2025-11-13)——"在Flink CDC同步MySQL到Doris的过程中,数据类型不兼容是高频错误点",逐项列出 JSON/ENUM、TIMESTAMP(6)、DECIMAL、BIT、时区 8 小时偏移的映射冲突与配置解法。
Source 6: CSDN personal blog (Nov 13 2025) — "data type incompatibility is a high-frequency failure point when syncing MySQL to Doris with Flink CDC," enumerating the JSON/ENUM, TIMESTAMP(6), DECIMAL, BIT, and 8-hour timezone-shift mapping conflicts with configuration fixes.
- 证据等级
`单方声音 [来源存疑]`,CSDN 匿名个人博客(细节充分:逐项类型冲突 + 配置解法;作者归属与独立性无法确认)。
`Single voice [Questionable source]`, anonymous CSDN personal blog (detailed: per-type conflicts + configuration fixes; author attribution and independence cannot be confirmed).
Apache Doris 年份:2025
"随手加个索引":一个 GSI 让账单涨 $800/月、写入慢 40%
单方声音
性能问题成本账单
- 一句话
为修慢查询随手加的 GSI,因为投影配置不当,每月多花 $800,写入还慢了 40%。
A hastily added GSI to fix a slow query, misconfigured on projection, cost an extra $800/month and slowed writes by 40%.
- 窄场景
为修复慢查询临时加 GSI、投影选 ALL 的高写入表。
Adding a GSI ad hoc to fix a slow query on a write-heavy table, with ALL projection.
- 机制
每个 GSI 是一份独立计费的副本:基表每次写入 → 每个 GSI 产生一次写入(1 个 GSI = 2 倍 WCU,N 个 GSI = N+1 倍);投影选 ALL 会把整行复制进索引,存储与写入成本双涨;GSI 有自己独立的容量配置,配不足会反压基表写入(写入变慢)。
Every GSI is an independently billed copy: each base-table write produces one write per GSI (1 GSI = 2x WCU, N GSIs = N+1x); ALL projection replicates whole items into the index, inflating both storage and write cost; each GSI has its own capacity settings, and an under-provisioned GSI back-pressures base-table writes.
- 生产验证
来源 4,Rob Abbott 2025-10——新版本上线后慢查询,团队"just add an index",查询从秒级降到毫秒级;账单日发现正是该 GSI 的投影配置导致每月 +$800 且写入慢 40%(原文 "cost us an extra $800 per month and slowed down our writes by 40%";部分付费墙,仅前半可见)。
Source 4, Rob Abbott, Oct 2025 — after a release, a slow query was fixed with "just add an index" (seconds to milliseconds); the bill then revealed that GSI's projection choice cost "an extra $800 per month and slowed down our writes by 40%" (partly paywalled; visible portion verified).
- 证据等级
`单方声音`,个人博客(细节充分:金额、写入 slowdown 比例;部分付费墙,已如实标注)。
`Single voice`, personal blog (concrete: dollar amount, write-slowdown percentage; partial paywall disclosed as-is).
Amazon DynamoDB 年份:2025
MaxScale 2025 纯商业化:BSL 的隐性契约被撕毁
单方声音
成本账单生态与信任
- 一句话
MaxScale 走完 GPLv2→BSL→纯商业三步,2025 年起新版源码不再公开——当年 BSL 承诺的"最终会开源"只兑现到 21.06。
MaxScale completed its GPLv2 → BSL → fully-commercial journey; since 2025 new source code is no longer published — the BSL's "eventually open source" promise only held through 21.06.
- 窄场景
生产架构依赖 MaxScale 做读写分离/查询路由/Galera 监控的用户。
Production architectures depending on MaxScale for read/write splitting, query routing, and Galera monitoring.
- 机制
BSL 的隐性契约是"生产付费、代码最终转 GPL";2025 年 MaxScale 25.01 改为闭源商业许可,源码不再可得,契约破裂。最后免费版 21.06 不再收安全更新——用户三选一:付费、守旧版裸奔、迁到 ProxySQL(缺 MaxGUI、galeramon 级 Galera 监控、数据脱敏、binlog 路由等)。
The BSL's implicit contract was "pay for production use, code turns GPL eventually"; in 2025 MaxScale 25.01 switched to a closed commercial license and the source stopped being available. The last free version, 21.06, receives no security updates — users face three options: pay, run an old version unpatched, or migrate to ProxySQL (which lacks MaxGUI, galeramon-class Galera monitoring, data masking, binlog routing, etc.).
- 生产验证
来源 8:PmaControl 2025-06-16(Sylvain Arbaudie 署名)——"By moving to pure commercial, MariaDB plc breaks this contract. MaxScale 25.01 code will never be free. Transparency disappears. And with it, the trust of part of the community."(引文为作者原话)
Source 8: PmaControl 2025-06-16 (bylined Sylvain Arbaudie) — "By moving to pure commercial, MariaDB plc breaks this contract. MaxScale 25.01 code will never be free. Transparency disappears. And with it, the trust of part of the community." (author quote)
- 证据等级
`单方声音`,具名独立顾问(细节充分:三段许可史、版本号、影响矩阵)。
`Single-source`, named independent consultant (fully detailed: three-stage licensing history, version numbers, impact matrix).
- 备注
与现有 [避坑] 卡主题重叠(档案吐槽清单"授权坑"已记 MaxScale BSL 生产使用受限),本卡是 2025 年事态升级:BSL→纯商业。
Overlaps the existing profile's "pitfall" list (the licensing-pitfall row already flags MaxScale BSL production-use restrictions); this card is the 2025 escalation — BSL → fully commercial.
MariaDB 年份:2025
小版本滚动更新翻车:10.11.9 改了系统表定义,错误日志 1GB/天
单方声音
运维复杂度升级迁移
- 一句话
zypper 从 10.11.x 升到 10.11.9 后,`mysql.column_stats` 表定义与新版预期不一致,错误日志灌到每天 1GB、服务器几小时后无响应。
After zypper upgraded MariaDB from 10.11.x to 10.11.9, the `mysql.column_stats` table definition no longer matched the new server's expectations — the error log ballooned to 1 GB/day and the server went unresponsive within hours.
- 窄场景
用发行版包(openSUSE/Leap)滚动更新 MariaDB 小版本、且未手动跑 mariadb-upgrade 的实例。
Instances rolling minor-version upgrades through distro packages (openSUSE/Leap) without manually running mariadb-upgrade.
- 机制
小版本包升级改变了服务端对系统表的期望定义(`mysql.column_stats`:`hist_type` 加 `JSON_HB` 枚举值、`histogram` 改 longblob),但包管理器不会自动执行 mariadb-upgrade;每次统计信息查询都刷两条 `[ERROR]`,日志爆炸拖慢磁盘 IO;`mariadb-check` 报 OK 也修不好,只能手工 ALTER mysql 系统表——普通用户不敢动系统表,陷入两难。
The minor-version package changed the server's expected system-table definition (`mysql.column_stats`: `hist_type` gained a `JSON_HB` enum value, `histogram` changed to longblob), but the package manager never runs mariadb-upgrade automatically; every statistics query then logged two `[ERROR]` lines, and the exploding log dragged disk I/O down; `mariadb-check` reported OK and fixed nothing, leaving hand-editing a mysql system table — something ordinary users are terrified to touch — as the only cure.
- 生产验证
来源 9:openSUSE 论坛 2025-01(用户 Dnk1287 发帖,hui 回复指路)——完整报错行、表结构、"error log grows to 1GB per day and the server gets unresponsive after some hours"(引文为用户原话);最终备份后手工 ALTER 系统表解决。
Source 9: openSUSE forum 2025-01 (posted by user Dnk1287) — full error lines, table structure, "error log grows to 1GB per day and the server gets unresponsive after some hours" (user quote); resolved by backing up and hand-altering the system table.
- 证据等级
`单方声音`,具名论坛用户(细节充分:完整报错、表结构、处置过程)。
`Single-source`, named forum user (fully detailed: complete errors, table structure, resolution steps).
- 备注
与现有 [避坑] 卡部分重叠(档案吐槽清单"大版本升级无官方回退"属同类升级运维税),本卡是小版本维度的实例。
Partially overlaps the existing profile's "pitfall" list ("no official rollback for major upgrades" is the same class of upgrade tax); this card is the minor-version instance.
MariaDB 年份:2025
Zilliz Cloud:量到 10 亿向量,账单"有点贵"
单方声音
成本账单
- 一句话
Zilliz Cloud 跑到约 10 亿向量量级时,账单开始"有点贵"——按 CU 小时 / GB 月计费的托管省心,是有标价的。
At roughly a billion vectors on Zilliz Cloud, pricing starts to feel "a bit expensive" — the operational peace of mind of per-CU-hour / per-GB-month billing has a price tag.
- 窄场景
向量规模冲向 10 亿量级、或 Dedicated 集群利用率不高的 Zilliz Cloud 用户。
Zilliz Cloud users whose vector volume approaches the billion-vector scale, or whose Dedicated clusters run at low utilization.
- 机制
Zilliz Cloud Dedicated 按计算单元(CU)× 小时计费,集群存在即计费、与查询量无关;Serverless 按存储 GB/月计费。向量规模上一个数量级,CU 与存储同步放大;用户原话指向的正是这个拐点。
Zilliz Cloud Dedicated bills per compute unit (CU) per hour — the cluster bills for existing, regardless of query volume; Serverless bills per GB/month of storage. Grow vectors by an order of magnitude and CUs plus storage scale with it; the quoted user points at exactly that inflection point.
- 生产验证
来源 5,2025-11 Software Finder 认证客户评价(小企业匿名用户,自述存了约 5000 万向量,价值评分 8/10)原话:"Pricing can get a bit expensive when the volume grows to around 1 billion vectors."(引用客户原话)
Source 5, Nov 2025 Software Finder verified customer review (anonymous small-business user, self-reported ~50M vectors, value-for-money score 8/10), quoted verbatim: "Pricing can get a bit expensive when the volume grows to around 1 billion vectors."
- 证据等级
`单方声音`,认证客户评价(匿名;细节:自述 5000 万向量规模)。
`Single voice`, verified customer review (anonymous; detail: self-reported 50M-vector scale).
Milvus 年份:2025
TTL 索引静默罢工:一条 2038 年后的时间戳让整集合停止过期(已修复于 8.0.12)
单方声音
已修复于 8.0.12
稳定与故障
- 一句话
MongoDB 8.0.4 上,时序集合中只要有一条文档的时间戳超过 2038-01-19,整个集合的 TTL 过期删除静默停摆——无报错、无指标异常,磁盘持续上涨。
On MongoDB 8.0.4, a single document with a timestamp past 2038-01-19 in a time-series collection silently halts TTL expiry for the entire collection — no errors, no anomalous metrics, disk usage keeps climbing.
- 窄场景
8.0.4 时序集合 + TTL 自动过期;脏数据/未来时间戳写入。
8.0.4 time-series collections with TTL auto-expiry; dirty or future-dated timestamp writes.
- 机制
SERVER-97368:TTL 评估遇到超出 32 位 Unix epoch 窗口的时间戳时直接跳过整个集合,而非跳过单条问题文档;TTL monitor 的 passes 计数器看起来正常,属于静默失败。根因是 2038 年问题(32 位秒级时间戳上限)。
SERVER-97368 — when TTL evaluation hits a timestamp outside the 32-bit Unix epoch window, it skips the whole collection instead of skipping the single offending document; the TTL monitor's pass counters look normal, so the failure is silent. Root cause is the year-2038 problem (32-bit seconds-since-epoch ceiling).
- 生产验证
Mydbops 客户实录《MongoDB TTL Failure in 8.0.4》(2025);定位到一条时间戳为 2038-08 的 rider 数据文档,升级到 8.0.12 后 TTL 立即恢复、磁盘企稳。
Mydbops client write-up, "MongoDB TTL Failure in 8.0.4" (2025); traced to one rider-data document dated August 2038; TTL resumed and disk stabilized immediately after upgrading to 8.0.12.
- 证据等级
`单方声音`(第三方运维商具名客户实录,具名 JIRA);备注:已修复于 8.0.12
—
MongoDB 年份:2025
5.7 升 8.0:多表 JOIN 查询卡死在"优化器思考人生"阶段
单方声音
性能问题升级迁移
- 一句话
5.7 升 8.0 后,多表 JOIN(尤其 ORM 生成的 10+ 表查询)的优化器 planning 时间从毫秒级膨胀到秒级,"编译 SQL"本身成了延迟主体。
—
- 窄场景
5.7→8.0 升级;高 JOIN 数、ORM 生成 SQL 的 OLTP 业务。
5.7→8.0 upgrades; OLTP workloads with high JOIN counts and ORM-generated SQL.
- 机制
optimizer_search_depth 默认 62,8.0 优化器分支更多(hash join、直方图、新代价模型),深度搜索的组合爆炸让 planning 阶段 CPU 开销大涨;设为 0(自动按查询规模限界)后,多表 JOIN 查询的编译时间下降一个数量级,部分端点总延迟下降 50–80%。
optimizer_search_depth defaults to 62, and the 8.0 optimizer explores more branches (hash join, histograms, new cost model), so the combinatorial explosion of deep search sharply raised planning-phase CPU; setting it to 0 (automatic per-query bounding) cut compile time by an order of magnitude and total latency on some endpoints by 50–80%.
- 生产验证
Medium 个人工程博客(2025-10-16):用 EXPLAIN ANALYZE 定位到查询卡在 STATISTICS/optimizing 阶段;staging 对比不同 depth 取值,生产灰度验证正确性(结果集不变)后全量。
Personal engineering blog (Oct 16, 2025): EXPLAIN ANALYZE pinned the stall to the STATISTICS/optimizing phase; depth values compared in staging, then rolled out in production during a low-traffic window with result-set correctness validated.
- 证据等级
`单方声音`,来源性质:个人工程博客(细节充分,但作者身份信息单薄,原文对 5.7/8.0 配置差异的叙述前后略有含混,采信时以"planning 阶段耗时异常、可复现、可缓解"为核心事实)。
—
- 备注
与本站现有 [避坑] 卡"8.x 简单 workload 性能倒退 20–40%"主题部分重叠(该卡已提 optimizer_search_depth regression);本卡角度为具体生产事故——编译期爆炸而非执行期吞吐倒退。
Partially overlaps the existing site card "8.x simple-workload performance regression of 20–40%" (which already mentions the optimizer_search_depth regression); this card's angle is a concrete production incident — compile-time explosion, not execution-time throughput regression.
MySQL 年份:2025
OMS 复制大字段丢数据:4K 以上 LOB 的前镜像没吐出来,下游被置空
单方声音
稳定与故障升级迁移
- 一句话
OMS 4.2.5.2 之前版本,行外存储的超 4K LOB 字段在 DML 时 OBCDC 不吐前镜像,下游 Kafka 消费者拿到的是空字段。
Before OMS 4.2.5.2, out-of-row LOB columns over 4K had no before-image emitted by OBCDC on DML, so downstream Kafka consumers received empty fields.
- 窄场景
用 OMS 把 OB 变更下发到 Kafka/数仓/Oracle 下游,且表里有大文本/大对象字段的链路。
Pipelines using OMS to ship OB changes to Kafka/data warehouse/Oracle downstream, with large text/object columns in the tables.
- 机制
LOB 行外存储时,DML 变更事件的前镜像(before image)需要 OBCDC 从存储层回填;老版本 OBCDC 在字段未被修改时不吐 LOB 列的前镜像,OMS 下发到 Kafka 时该字段被置空。下游若按"没变化的字段保持原值"假设消费,就会静默丢数据——只改更新时间也能把大字段"改没"。
For out-of-row LOB storage, the DML change event's before image must be backfilled by OBCDC from the storage layer; older OBCDC did not emit the LOB column's before image when the column itself was unmodified, so OMS published the field as empty to Kafka. Downstream consumers assuming "unchanged columns keep their values" silently lost data — touching only an update timestamp could "delete" a large field.
- 生产验证
来源 2:百丽/卢文豪 2025-11-25——"OMS V4.2.5.2 之前的版本需要关注字段超过 4K 的复制情况……执行 DML 时 OBCDC 可能不吐出 LOB 列的前镜像,导致下游数据不一致……建议使用 4.2.5.2 以后的版本(OMS 4.2.5.2 版本已解决该问题)"。
Source 2: Belle/Lu Wenhao, Nov 25 2025 — "versions before OMS 4.2.5.2 need attention when replicating fields over 4K... OBCDC may not emit the LOB column's before image on DML, causing downstream data inconsistency... use versions after 4.2.5.2 (fixed in OMS 4.2.5.2)."
- 证据等级
`单方声音`,具名客户生产复盘(细节充分:版本号、触发条件、修复版本)。
`Single voice`, named customer production postmortem (detailed: version, trigger, fixed-in version).
- 备注
已修复于 OMS V4.2.5.2,按规则保留收录并标注。
fixed in OMS V4.2.5.2; retained with the fix labeled per the rules.
OceanBase 年份:2025
DDL 在线还是离线,官方不告诉你:离线 DDL 直接锁表,只能自研检测工具
单方声音
运维复杂度
- 一句话
OceanBase 的 DDL 分在线/离线两种,离线 DDL 会锁表,但判断一次变更是哪种只能靠"肉眼"——百丽被迫自研了 Table ID 变化检测。
OceanBase DDL comes in online and offline flavors, and offline DDL locks the table — but there is no authoritative way to know which kind a given statement is, so Belle built a Table-ID-change detector.
- 窄场景
用 Archery/工单平台做 DDL 发布的团队;大表结构变更频繁的业务。
Teams running DDL through Archery/ticketing platforms; businesses with frequent large-table schema changes.
- 机制
OB 的 DDL 按是否重建表分为在线(只改元数据)与离线(重建表、期间锁表);但官方没有在事前明确告知某条 DDL 属于哪一类。百丽的解法是:在测试环境先跑一遍,看 Table ID 是否变化——变了就是离线、会锁表,再在工单上打告警。
OB DDL is online (metadata-only) or offline (rebuilds the table, locking it) depending on whether the table is rebuilt; but nothing tells you in advance which category a statement falls into. Belle's workaround: run it once in a test environment and check whether the Table ID changed — changed means offline and will lock.
- 生产验证
来源 2:百丽/卢文豪 2025-11-25——"在使用 Archery 平台管理 MySQL 时,我们通常通过 ptosc 或 gos 等方式实现 Online DDL 操作。然而,这种方式对于 OceanBase 的离线操作,可能会直接锁表……在提交工单时,我们会在测试环境中的 OceanBase 中运行一遍,检查 Table ID 是否发生变化。如果 Table ID 发生变化,则表明操作是 Offline 的"。
Source 2: Belle/Lu Wenhao, Nov 25 2025 — "When managing MySQL with Archery we usually get online DDL via pt-osc or gh-ost. With OceanBase, an offline operation may lock the table directly... at ticket submission time we run it once in the test OceanBase and check whether the Table ID changed. If it changed, the operation is offline."
- 证据等级
`单方声音`,具名客户生产复盘。
`Single voice`, named customer production postmortem.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site on this site.
OceanBase 年份:2025
Binlog Service 单线程:租户级生产消费,高峰期就是瓶颈
单方声音
性能问题
- 一句话
OceanBase Binlog Service 以租户为维度生产与消费 Binlog,只能单线程,高峰期下游同步跟不上。
OceanBase Binlog Service produces and consumes binlog per tenant on a single thread — downstream sync cannot keep up at peak.
- 窄场景
从 OB 向下游(数仓/Oracle)同步数据、依赖 Binlog 消费链路的架构。
Architectures syncing OB data downstream (warehouse/Oracle) over a binlog consumption chain.
- 机制
MyCat 架构下 8 个 MySQL 分片是 8 个线程并行采集;换成 OB 后,Binlog Service 按租户单线程生产、单线程消费,并发能力从"分片数"跌到"1"。业务高峰期变更量一大,单线程消费就成为整条同步链路的瓶颈。
Under MyCat, eight MySQL shards were collected by eight parallel threads; after the move, Binlog Service produces and consumes per tenant on exactly one thread, so concurrency collapses from "number of shards" to "1". At business peaks the single-threaded consumer becomes the bottleneck of the whole sync chain.
- 生产验证
来源 2:百丽/卢文豪 2025-11-25——"OceanBase Binlog Service 是以租户为维度,无论是生产 Binlog 还是消费 Binlog,都只能有 1 个线程去处理。在业务高峰期间,可能存在性能瓶颈。"百丽最终被迫改用 OMS→Kafka 链路并自研下游同步工具。
Source 2: Belle/Lu Wenhao, Nov 25 2025 — "OceanBase Binlog Service is tenant-scoped; whether producing or consuming binlog, only 1 thread can do the work. At business peaks this can be a performance bottleneck." Belle was forced onto an OMS-to-Kafka chain plus home-built downstream sync tooling.
- 证据等级
`单方声音`,具名客户生产复盘。
`Single voice`, named customer production postmortem.
OceanBase 年份:2025
批量写入慢一个数量级:Oracle 3 分钟的活,OB 4.2.5.6 跑了 29 分钟
单方声音
性能问题
- 一句话
同一份批量写入,Oracle 3 分钟,换成 OceanBase 4.2.5.6 后 29 分钟——默认参数下写入性能差近 10 倍。
The same bulk load took 3 minutes on Oracle and 29 minutes after switching to OceanBase 4.2.5.6 — nearly 10x slower with default parameters.
- 窄场景
数据同步/ETL 类批量写入场景;按默认参数部署、未调转储与并发参数的集群。
Bulk-write workloads like data sync/ETL; clusters deployed with defaults, memstore/compaction and concurrency parameters untuned.
- 机制
OB 的写入先进入 MemStore(内存),MemStore 满阈值触发转储(minor compaction)落盘;转储参数过小会导致批量写入中频繁转储,CPU/IO 被转储抢走。同时日志归档若开启,海量 Redo 同步归档存储成为瓶颈;租户工作线程数与连接池并发不足则并行度上不去。三者叠加,批量写入吞吐远低于单机 Oracle 的直接路径写入。
OB writes go to the MemStore first; when it hits its threshold it triggers minor compactions to disk. With a small compaction trigger, bulk writes compact constantly and CPU/IO is stolen by the compaction. If log archiving is on, the flood of redo synced to archive storage becomes the bottleneck; too few tenant worker threads and too-small client connection pools cap parallelism. Together, bulk throughput lands far below a single-node Oracle's direct-path writes.
- 生产验证
来源 3:「WEL测试」2025-10-31——"早期 B 数据库采用 Oracle 时,全量批量写入耗时稳定在 3 分钟。为适配业务架构升级,将 B 数据库替换为 OceanBase 4.2.5.6 版本后,相同数据量的批量写入耗时骤增至 29 分钟,性能下降近 10 倍。"作者锁定转储参数、日志归档、并发配置三个排查方向。
Source 3: "WEL Test," Oct 31 2025 — "With Oracle as database B, the full bulk write held steady at 3 minutes. After switching B to OceanBase 4.2.5.6 with no business-logic changes, the same data volume took 29 minutes — nearly a 10x drop." The author narrowed it to compaction parameters, log archiving, and concurrency configuration.
- 证据等级
`单方声音`,个人博客实战记录(数字具体:3 分钟 vs 29 分钟,版本具体)。
`Single voice`, personal blog field notes (concrete numbers: 3 vs 29 minutes, concrete version).
- 备注
该文为 CSDN 推荐文章(平台推广位),作者为独立技术博主;文中只给出排查方向、未记录最终调优结果,如实标注。
the piece ran in a CSDN promoted slot; the author is an independent blogger; it lists investigation directions but no final tuning outcome — stated as-is.
OceanBase 年份:2025
手动部署的隐形门槛:OBProxy 密码要传 sha1,bootstrap 被 5G 的 unit 最小内存拦下
单方声音
运维复杂度
- 一句话
不用 OBD、手动装 OB 集群时,OBProxy 初始化密码必须传 sha1 哈希、明文不行;bootstrap 还会被默认 5G 的 unit 最小内存挡住报错。
Deploying an OB cluster by hand (no OBD/OCP), OBProxy's init password must be passed as a sha1 hash — plaintext fails; bootstrap is blocked by the default 5G unit minimum memory.
- 窄场景
手动 RPM 部署(不用 OBD/OCP)的团队;小规格测试机。
Manual RPM deployments (no OBD/OCP); small test machines.
- 机制
OBProxy 初始化参数 observer_sys_password 要求传入密码的 sha1 值,传明文直接报 ERROR 2013 连接失败,报错信息与密码问题毫无关联;集群 bootstrap 时 __min_full_resource_pool_memory 默认 5G,小内存机器直接报 ERROR 1235,必须手动把隐藏参数调小才能继续。两个都是"报错信息不说人话"的经典坑。
OBProxy's observer_sys_password init parameter requires the sha1 of the password; passing plaintext fails with ERROR 2013, whose message says nothing about passwords. Cluster bootstrap fails on small-memory machines with ERROR 1235 because __min_full_resource_pool_memory defaults to 5G and must be lowered by hand. Both are classic "the error message doesn't speak human" traps.
- 生产验证
来源 4:loxehate 个人 GitHub 博客 2025-05-08「问题总结」——"部署 OBProxy 连接数据库失败,使用 obclient 可以正常连接。ERROR 2013 (HY000): Lost connection to MySQL server at 'reading authorization packet'……proxyro 的密码必须为 sha1 后的密码,明文密码不行";"集群 bootstrap 操作失败 ERROR 1235 (0A000): unit min memory less than __min_full_resource_pool_memory not supported……__min_full_resource_pool_memory 默认值是 5G"。
Source 4: loxehate personal GitHub blog, May 8 2025, "problem summary" — "Deploying OBProxy failed to connect to the database while obclient connected fine. ERROR 2013 (HY000): Lost connection to MySQL server at 'reading authorization packet'... the proxyro password must be the sha1-hashed password, plaintext does not work"; "cluster bootstrap failed with ERROR 1235 (0A000): unit min memory less than __min_full_resource_pool_memory not supported... __min_full_resource_pool_memory defaults to 5G."
- 证据等级
`单方声音`,个人博客部署实战记录(命令与报错原文俱全)。
`Single voice`, personal blog deployment notes (commands and verbatim errors included).
OceanBase 年份:2025
分区裁剪要人肉补条件:少写一个分区键,优化器就"看不见"分区
单方声音
性能问题运维复杂度
- 一句话
从 MyCat 迁过来后,原来靠物理分片"天然隔离"的查询,在 OB 里必须手动补上分区条件,否则分区裁剪失效。
After moving off MyCat, queries that used to be "naturally isolated" by physical sharding need the partition predicate added by hand in OB, or pruning silently fails.
- 窄场景
从分库分表中间件(MyCat/ShardingSphere)迁移到 OB 的团队;按大区/租户分区的表。
Teams migrating from sharding middleware (MyCat/ShardingSphere) to OB; tables partitioned by region/tenant.
- 机制
MyCat 下不同大区是不同物理库,查询天然只扫一个分片;OB 里这些分片变成同一张分区表的不同分区,优化器做分区裁剪需要查询条件中显式带有分区键。业务 SQL 里少写分区条件,优化器无法推导裁剪范围,只能全分区扫描。等于把原来中间件在物理层免费做掉的事,变成了每个开发都要记得的心智负担。
Under MyCat, regions were separate physical databases so queries only ever hit one shard; in OB those shards become partitions of one table, and the optimizer needs the partition key explicitly in the predicate to prune. Leave the partition condition out of the SQL and the optimizer cannot derive the range — full-partition scan. What the middleware did for free at the physical layer becomes a mental tax on every developer.
- 生产验证
来源 2:百丽/卢文豪 2025-11-25——"在涉及 DTL 订单明细表的 order by 关联查询中,指定某个大区条件,在 MyCat 中没有……这是因为 MyCat 的分区规则相同……在物理层面进行了隔绝……而在 OceanBase 中,如果将这个条件去掉,会引发 om 表正常进行分区裁剪,但 od 表不知道需要在这个大区内进行……OceanBase 需要指定这个条件以确保分区裁剪的正确性。这是一个非常典型的分区裁剪问题。"
Source 2: Belle/Lu Wenhao, Nov 25 2025 — "In an order-by join over the DTL order-detail table with a region condition, MyCat needed no such condition because its sharding rules physically isolated regions... in OceanBase, dropping the condition leaves the om table pruning fine but the od table not knowing it must stay within the region... OceanBase needs the condition specified to guarantee correct partition pruning. A very typical partition-pruning problem."
- 证据等级
`单方声音`,具名客户生产复盘(有具体表与查询场景)。
`Single voice`, named customer production postmortem (concrete tables and query shape).
OceanBase 年份:2025
19.26 RU 补丁翻车:打完补丁,PDB 里大量对象变 invalid
单方声音
稳定与故障升级迁移
- 一句话
19.26 Release Update 在 3 台 Oracle Linux 9 服务器上把 PDB 里大量对象打成 invalid,datapatch 跑不完,官方支持让 DBA 重复已经试过的步骤,最后靠手动跑 catupgrd.sql 才救回来。
The 19.26 Release Update left large numbers of objects invalid in PDBs across 3 Oracle Linux 9 servers; datapatch could not complete, official support had the DBA repeat already-tried steps, and only a manual catupgrd.sql run saved it.
- 窄场景
19c 打 Release Update 补丁的多租户(PDB)环境。
Multi-tenant (PDB) environments patching 19c Release Updates.
- 机制
RU 补丁需要 datapatch 完成数据字典/组件升级;此案例中 PDB 内大量对象 invalid,datapatch 无法完成,utlrp.sql、catalog.sql、catproc.sql 均无效,手动重编译报 ORA-04020(SYS.AQ$_REG_INFO 死锁);一线 SR 先让 DBA 重复已试步骤无果,升级后的支持才建议在 PDB 上跑 catupgrd.sql 解决——根因至今不明。
RU patches require datapatch to finish data-dictionary/component upgrades; here masses of objects in PDBs went invalid, datapatch could not complete, and utlrp.sql, catalog.sql, and catproc.sql all failed — manual recompilation hit ORA-04020 (deadlock on SYS.AQ$_REG_INFO); front-line SR asked the DBA to repeat already-tried steps to no effect, and only escalated support suggested running catupgrd.sql in the PDB — root cause remains unknown.
- 生产验证
来源 6:Tim Hall(ORACLE-BASE 站长,26 年 Oracle 社区作者)2025-03-10 完整记录——3 台 OL9 服务器逐台复现、完整的报错与排查链条、最终解决方式与"根因不明"的诚实结论。
Source 6: Tim Hall (ORACLE-BASE webmaster, 26-year Oracle community author) documented it fully on 2025-03-10 — reproduced server by server across 3 OL9 machines, complete error and troubleshooting chain, the final fix, and the honest "root cause unknown" conclusion.
- 证据等级
`单方声音`,具名资深 DBA(细节充分:版本号、平台、报错码、排查全过程)。
`Single voice`, named veteran DBA (rich details: versions, platform, error codes, full troubleshooting process).
Oracle Database(甲骨文) 年份:2025
Reddit 实测:写入一忙,查询就慢——同构节点的代价
单方声音
性能问题
- 一句话
Qdrant 的写入和查询跑在同一批节点上,Reddit 3.4 亿向量实测发现写入负载对查询延迟的干扰远大于 Milvus。
Qdrant serves writes and queries from the same nodes; Reddit's 340M-vector test found write load interfered with query latency far more on Qdrant than on Milvus.
- 窄场景
持续写入 + 在线查询的混合负载(RAG 增量索引、实时推荐写入);评估基于 v1.12,结论为架构性。
Mixed workloads with continuous writes plus live queries (incremental RAG indexing, real-time recommendation writes); evaluated on v1.12; the conclusion is architectural.
- 机制
Qdrant 是同构节点架构——segment 构建、HNSW 索引、优化器合并与查询服务争抢同一节点的 CPU 和内存带宽;批量写入期的索引构建是 CPU/内存密集型,直接挤占查询。Milvus 把写入/索引/查询拆到异构节点类型,天然隔离。
Qdrant uses a homogeneous node architecture — segment building, HNSW indexing, optimizer merges, and query serving all compete for the same nodes' CPU and memory bandwidth; index construction during bulk writes is CPU/memory-intensive and directly crowds out queries. Milvus splits writes, indexing, and queries across heterogeneous node types, isolating them by design.
- 生产验证
来源 3:Reddit 工程评估(Chris Fournie 牵头,2025-11)——恒定吞吐(100 QPS)下对比,"Qdrant 上写入与查询负载的相互干扰远大于 Milvus";作者归因于架构差异:"Milvus 把大部分写入分散到与查询服务不同的节点类型上,而 Qdrant 的写入和查询走同一批节点"。
Source 3: Reddit engineering evaluation (led by Chris Fournie, Nov 2025) — at constant throughput (100 QPS), "there was far more of an interaction between ingestion and query load on Qdrant than on Milvus"; attributed to architecture: "Milvus splits much of its ingestion over separate node types from those that serve query traffic, whereas Qdrant serves both ingestion and query traffic from the same nodes."
- 证据等级
`单方声音`,具名工程评估(Reddit Staff SWE,含多名工程师致谢名单),3.4 亿向量实测数据。
`Single voice`, named engineering evaluation (Reddit Staff SWE, multiple engineers acknowledged), 340M-vector measured data.
- 备注
评估基于 Qdrant v1.12;后续版本未获独立复测,不作"已修复"标注。
evaluated on Qdrant v1.12; no independent re-test on later versions, so no "fixed in" label.
Qdrant 年份:2025
强制 Duo 推送:本地跑一次多线程 dbt,手机震八次
单方声音
运维复杂度
- 一句话
"Snowflake's mandatory multi-factor authentication is great for security, but it's terrible for developer flow"——本地跑一次多线程 dbt,手机要震八次批准;作者最终给个人账号换上 key-pair 认证,"再也不用 Duo 批准或输 OTP"——MFA 的安全收益在开发场景被自己的交互设计抵消了一半。
"Snowflake's mandatory multi-factor authentication is great for security, but it's terrible for developer flow" — one local multi-threaded dbt run meant eight approval buzzes; the author ended up switching his personal account to key-pair auth, "never a Duo approval or OTP again" — MFA's security gains half-undone by its own interaction design in dev workflows.
- 窄场景
本地高频跑 dbt/CLI 调试的数据工程师;多线程并行建连接的开发流程。
Data engineers running dbt/CLI locally at high frequency; dev flows that open many parallel connections.
- 机制
Snowflake 当时的 MFA 实现仅支持 Duo App 推送一种方式;每个新连接都要人工批准,多线程=多次打扰。
Snowflake's MFA implementation at the time supported only Duo App push; every new connection demanded a manual approval; multi-threading means multi-interruption.
- 生产验证
来源 30:Mamad Kajbaf 2025-09-12——"Every time I ran a multi-threaded dbt command locally, my phone buzzed for approval EIGHT TIMES."(Medium 正文后半被付费墙挡住,核心吐槽段落在公开部分,已核实。)
—
- 证据等级
`单方声音`,一线开发者具名实录;Duo 推送疲劳是通用痛点,针对 Snowflake 的具名实录只找到这一篇。
`Single voice`, named frontline developer record; Duo fatigue is a generic pain, but this is the only named Snowflake-specific record found.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2025
GetYourGuide:把 Looker 从 Snowflake 搬到 Databricks SQL,降本 20%、砍掉整套数据复制链路
单方声音
升级迁移
- 一句话
柏林旅游平台复审 BI 展示层发现:数据湖源头本就躺在 AWS S3 + Databricks,Snowflake 只是个"展示层"——要把 Databricks 产出的数据再复制一份进 Snowflake 供 Looker 查询;迁移后运营成本降约 20%,还砍掉了整套复制链路。
The Berlin travel platform's BI review found the lakehouse source already lived on AWS S3 + Databricks, with Snowflake as just a "presentation layer" — Databricks-produced data had to be copied into Snowflake for Looker; after migrating, operating costs fell ~20% and the whole copy pipeline was deleted.
- 窄场景
lakehouse 已建成、Snowflake 只剩"BI 展示层"定位的团队;为同一份数据付两份存储+两套计算的组织。
Teams whose lakehouse is built but Snowflake remains only as the "BI presentation layer"; orgs paying double storage + double compute for the same data.
- 机制
架构冗余——数据在 Databricks 产出后复制进 Snowflake 只为 serving BI,多一层复制就多一份存储成本、一份同步延迟、一套故障面。
Architectural redundancy — data produced in Databricks copied into Snowflake just to serve BI; one more copy means one more storage bill, one more sync lag, one more failure surface.
- 生产验证
来源 34:GetYourGuide 数据平台团队官方工程博客 2025 年初——迁移范围:2 套 Looker(内部 300 日活+外部 1500 日活)、25 个模型 430+ explores、2024 上半年 6.5M 内部+16M 外部查询、20K+ 唯一 SQL 语法校验、750 张 Snowflake 表迁为 Delta 表;PoC 用 Looker SDK 把测量过程代码化;痛点:Looker 按供应商生成不同 SQL 方言,Spark 下 file/partition pruning 偶发不生效需逐一绕过;Databricks 客户团队资助了迁移期间的外部咨询支持(已注明)。
—
- 证据等级
`单方声音`,客户官方工程博客(数字具体:20% 降本、750 张表、22.5M 查询)。
`Single voice`, customer engineering blog (concrete numbers: 20% savings, 750 tables, 22.5M queries).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2025
SmarterX:80+ 库迁 BigQuery,"Snowflake 像嫁接到云上的传统数仓" [来源存疑]
单方声音
来源存疑
升级迁移
- 一句话
80 多个数据库、数千张表、21 个数据源,不到一个月从 Snowflake 搬进 BigQuery;客户原话:"Snowflake felt like a traditional enterprise data warehouse grafted onto the cloud, completely uninfluenced by the AI revolution, with a database that forced us to work in a specific, predetermined way."(Snowflake 像一座嫁接到云上的传统企业数仓,完全没受 AI 革命影响,逼着人按它预设的方式干活。)
80+ databases, thousands of tables, 21 data sources moved from Snowflake to BigQuery in under a month; the customer's sharpest line: "Snowflake felt like a traditional enterprise data warehouse grafted onto the cloud, completely uninfluenced by the AI revolution, with a database that forced us to work in a specific, predetermined way."
- 窄场景
AI 公司需要弹性扩展、快速 onboarding 新客户的数据平台。
AI companies needing elastic scale and fast customer onboarding on their data platform.
- 机制
迁移采用 external table 快速接入 + native table 优化性能的双轨策略;业务收益:新产品发布快 10 倍、新客户 onboarding 从 6 个月缩到不到 1 周、pipeline 数据量是之前的 100 倍。
Dual-track migration — external tables for fast onboarding, native tables for performance; business outcomes: 10× faster product launches, new-customer onboarding from 6 months to under a week, 100× pipeline data volume.
- 生产验证
来源 38:Google Cloud 官方博客客户案例 2025-07([来源存疑]:Google 执笔的厂商案例稿)——官方口径成本 "cut costs in half";迁移由 Google Technical Onboarding Center 深度参与。
—
- 证据等级
`单方声音`,[来源存疑]:未见 SmarterX 自家工程博客或第三方独立报道。
`Single voice`, [source questionable]: no SmarterX-owned engineering blog or third-party independent coverage found.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2025
Dynamic Tables 对 DLT:要指定 warehouse、只会 SQL、lag 不保证
单方声音
运维复杂度
- 一句话
咨询公司受客户委托落地 Dynamic Tables,几个月后放弃退回 Tasks+存储过程:"Snowflake forces you to specify which warehouse you will use… moving back to a hardcoded warehouse will be the same as moving to the Stone Age"(回到硬编码 warehouse 跟回到石器时代一样);target lag 只是 best-effort——"A refresh 5 minutes early? Sure, why not? A delay of 4 minutes? It happens.";底层表无 DML 时"it will refuse to refresh until a change occurs (even a manual refresh)"(连手动刷新都拒绝);增量模式禁非确定函数、禁 UDF,pivot/读外部表/读安全共享视图还得回退存储过程。对比 Databricks DLT:serverless 计算、Python/SQL 双语言、数据质量 expectations 原生。
A consultancy engaged to land Dynamic Tables gave up after months and rolled back to Tasks + stored procedures: "Snowflake forces you to specify which warehouse you will use… moving back to a hardcoded warehouse will be the same as moving to the Stone Age"; target lag is best-effort — "A refresh 5 minutes early? Sure, why not? A delay of 4 minutes? It happens."; with no DML on base tables "it will refuse to refresh until a change occurs (even a manual refresh)"; incremental mode bans nondeterministic functions and UDFs; pivoting, reading external tables, or secure shared views still fall back to procedures. Compare Databricks DLT: serverless compute, Python/SQL, native data-quality expectations.
- 窄场景
想用声明式管道替代 Task+存储过程的团队;对数据新鲜度有 SLA 要求的 pipeline。
Teams wanting declarative pipelines instead of Tasks + procedures; pipelines with freshness SLAs.
- 机制
Dynamic Tables 的刷新是"尽力而为",lag 是目标不是保证;无变更不刷新;能力子集(SQL-only、无 UDF/非确定函数);且必须绑定 warehouse,无 serverless 选项。
Dynamic Table refresh is best-effort — lag is a target, not a guarantee; no changes means no refresh; capability subset (SQL-only, no UDFs/nondeterministic functions); and a warehouse must be named — no serverless option.
- 生产验证
来源 52:Ulpia Tech(保加利亚独立数据咨询公司)客户实战复盘,2025-04——第一人称"we",5 大类限制逐条可核。
—
- 证据等级
`单方声音`,咨询公司客户实战复盘(细节充分)。
`Single voice`, consultancy customer field postmortem (rich detail).
- 备注
缺口状态:部分至今缺失(SQL-only、lag best-effort、无 DML 不刷新等限制在官方文档中延续;serverless DT 截至 2026-10-07 未见官方 GA 公告)。本卡主题可能与本站 [避坑] 卡重叠。
gap status: partially still missing (SQL-only, best-effort lag, no-refresh-without-DML persist in official docs; no serverless DT GA announced as of 2026-10-07). This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2025
SSIS:2019→2022 升级后,包直接跑不起来
单方声音
升级迁移
- 一句话
SQL Server 从 2019 升到 2022,SSIS 包部署完一个都跑不起来——ISServerExec.exe 抛未处理异常(error 27203),operations_messages 里还查不到任何记录。
After upgrading SQL Server from 2019 to 2022, deployed SSIS packages refused to run at all — ISServerExec.exe died with an unhandled exception (error 27203), and the SSISDB operations_messages table showed nothing.
- 窄场景
用 SSISDB 目录(project deployment model)跑 ETL 的团队;随数据库引擎一起升级 SSIS 的。
Teams running ETL on the SSISDB catalog (project deployment model); upgrading SSIS alongside the database engine.
- 机制
SSIS 目录(SSISDB)版本要与引擎匹配,升级时目录不一定自动跟着升;ISServerExec.exe 报 `System.IO.FileNotFoundException` 式未处理异常,错误信息只给 error 27203,SSISDB 的 operations_messages 表里空空如也——排错只能靠 SSMS 日志文件查看器翻堆栈。社区给出的偏方包括:重建 SSIS 目录、核对 SSIS 服务账号权限、检查第三方组件兼容性。SSIS 的升级体验本质上是"引擎升了,ETL 运行时是另一个产品"。
The SSIS catalog (SSISDB) version must match the engine, and the catalog doesn't always upgrade itself during a server upgrade; ISServerExec.exe throws `System.IO.FileNotFoundException`-style unhandled exceptions, the only error text is error 27203, and SSISDB's operations_messages table is empty — troubleshooting means digging through the SSMS Log File Viewer for stack traces. Community workarounds include rebuilding the SSIS catalog from scratch, fixing SSIS service-account permissions, and checking third-party component compatibility. The SSIS upgrade experience is essentially "the engine upgraded, but the ETL runtime is a different product."
- 生产验证
来源 10,2025-03 社区论坛——用户详述已试过三种部署方式(升级部署项目、导出重部署、按目标版本构建),全部报 error 27203;回帖者分享了自己历次升级踩坑的排查清单(目录版本、服务账号、重建目录)。
Source 10, Mar 2025 community forum — the user documents three attempted deployment approaches (upgrade-and-deploy the project, export and redeploy, build packages per target version), all failing with error 27203; responders share their own upgrade-battle checklists (catalog version, service account, catalog rebuild).
- 证据等级
`单方声音`,社区论坛(细节充分:具体错误号、堆栈、已尝试的三种部署方式)。
`Single voice`, community forum (detailed: exact error number, stack trace, three attempted deployment approaches).
- 备注
SSIS 在 Linux 上长期残血(无目录、无 Agent 调度),且微软官方文档已确认 SSIS 2025 不再支持 Linux、legacy Integration Services Service 在 2025 被废弃——"买 license 附赠的 BI 套件"正在被慢慢晾干。
SSIS on Linux has long been crippled (no catalog, no Agent scheduling), and Microsoft's own documentation confirms SSIS is unavailable on Linux for SQL Server 2025 and the legacy Integration Services Service is deprecated in 2025 — the "BI suite thrown in with the license" is being quietly left to dry.
Microsoft SQL Server 年份:2025
用户名里的下划线能被任意字符替换:资源组配额形同虚设
单方声音
生态与信任
- 一句话
`test_1` 能用 `test=1`、`test.1` 甚至 `test1` 登录——用户名特殊字符在鉴权时被归一化替换,按用户名绑定的资源组规则直接被绕过,配额限制名存实亡。
`test_1` can log in as `test=1`, `test.1`, even `test1` — special characters in usernames get normalized away at authentication, so per-username resource-group rules are trivially bypassed and quota enforcement is fiction.
- 窄场景
用 `_` 等特殊字符做用户名分隔规范、且依赖 resource group 做多租户资源隔离的集群;3.3.3/3.3.9/3.5.0 均受影响,存算一体/分离架构都中招。
Clusters using `_`-style username conventions with resource groups for multi-tenant isolation; reproduced on 3.3.3/3.3.9/3.5.0, on both shared-nothing and shared-data architectures.
- 机制
登录时用户名中的特殊字符可被 `.`、`=`、`*` 等任意替换后通过鉴权(`test_1` ≡ `test=1` ≡ `test1`),而资源组绑定规则只匹配含 `_` 的用户名。用户用 `test*1` 这类变体登录即落到 `default_wg` 默认资源组,预设的资源配额完全失效。报告者明确指出业务影响:大量用户挤进默认资源组 → 集群负载激增 → 有资源耗尽导致宕机的风险,且合规用户的资源被挤占。
At login, special characters in a username can be substituted with `.`, `=`, `*`, etc. and still authenticate (`test_1` ≡ `test=1` ≡ `test1`), while resource-group binding rules only match usernames containing `_`. Logging in with a variant like `test*1` lands the session in the `default_wg` default group, voiding all preset quotas. The reporter spells out the business impact: users pile into the default group → cluster load surges → risk of resource-exhaustion outage, crowding out compliant users.
- 生产验证
来源 4:GitHub issue #61525,2025-08,企业用户报告("Our company's username naming convention uses _ as a separator"),附复现步骤与业务影响分析,横跨三个版本复现。
Source 4: GitHub issue #61525, Aug 2025 — enterprise user report ("Our company's username naming convention uses _ as a separator") with reproduction steps and business-impact analysis, reproduced across three versions.
- 证据等级
`单方声音`,GitHub 企业用户 issue(细节充分:复现步骤、三版本验证、业务影响)。
`Single voice`, GitHub enterprise-user issue (detailed: repro steps, three-version verification, business impact).
StarRocks 年份:2025
Iceberg 联邦查询:高频写入下小文件+Compaction 压力,分钟级时延保不住
单方声音
性能问题
- 一句话
Fresha 最初想把热点链路直接跑在 Iceberg 上——功能都通,但高频写入产生的大量小文件和 Compaction 压力让"分钟级时延"根本稳不住,最后只能把热点链路迁回 StarRocks 内部表,Iceberg 退回去只当历史存档。
Fresha first tried running its hot path directly on Iceberg — everything functioned, but the flood of small files and compaction pressure from high-frequency writes made "minute-level freshness" impossible to sustain; the hot path had to move back to StarRocks internal tables, with Iceberg demoted to historical archive.
- 窄场景
以 Iceberg/Paimon 为唯一事实源、同时想在其上跑秒级~分钟级准实时查询的湖仓架构;CDC 高频写入场景。
Lakehouse architectures treating Iceberg/Paimon as the single source of truth while also wanting second-to-minute quasi-real-time queries on top; CDC high-frequency write workloads.
- 机制
开放表格式的高频写入必然产生大量小文件,查询侧要付出文件数膨胀的代价(list/scan 开销、Compaction 压力)。Fresha 的实践结论是:StarRocks 直查 Iceberg"功能上没问题但运行层面不稳定",分钟级时延无法持续保证。最终架构变成 Hot(秒级)走内部表、Warm/Deep history 才走 Iceberg 联邦——"统一湖仓入口"的宣传与高频写入现实之间有一道小文件鸿沟,填坑要靠 Spark 在外部做 compaction/backfill。
High-frequency writes to open table formats inevitably produce masses of small files, and the query side pays in file-count inflation (list/scan overhead, compaction pressure). Fresha's conclusion: querying Iceberg directly through StarRocks "worked functionally but was unstable operationally" — minute-level latency could not be guaranteed. The resulting architecture is Hot (seconds) on internal tables, Warm/Deep history on Iceberg federation — a small-file chasm sits between the "unified lakehouse entry point" marketing and high-frequency-write reality, bridged only by Spark-side compaction/backfill outside StarRocks.
- 生产验证
来源 5:Fresha 工程师 Anton Borisov 执笔的案例(2025-12,StarRocks 官方博客"全球用户精选案例"栏目,厂商邀请、客户执笔)——"我们最初尝试使用 Iceberg,功能上没问题但运行层面不稳定——高频写入产生的大量小文件和 Compaction 压力,使分钟级时延难以持续保证。于是,热点链路切换至 StarRocks 内部表";自 2025 年春季生产上线。
Source 5: case authored by Fresha engineer Anton Borisov (Dec 2025, StarRocks official blog "global user spotlight" column — vendor-invited, customer-penned): "We first tried Iceberg; functionally fine but operationally unstable — the large number of small files and compaction pressure from high-frequency writes made minute-level latency impossible to guarantee. So the hot path moved to StarRocks internal tables"; in production since spring 2025.
- 证据等级
`单方声音`,具名客户工程师执笔(发表于厂商官方博客栏目,渠道已如实标注)。
`Single voice`, named customer engineer authorship (published in the vendor's official blog column; channel disclosed as-is).
StarRocks 年份:2025
DDL 是异步的:Schema 变更没有同步语义,得自己造迁移工具轮询
单方声音
运维复杂度
- 一句话
StarRocks 的许多 DDL 是异步执行的——发完 DDL 不代表变更生效,Fresha 被迫自研了一套 ActiveRecord 风格的迁移工具,靠轮询后台任务到 FINISHED 状态来保证 Schema 演进可逆、安全。
Many StarRocks DDL operations execute asynchronously — a returned DDL says nothing about the change being effective, so Fresha built an ActiveRecord-style migration tool that polls background tasks to FINISHED state to keep schema evolution reversible and safe.
- 窄场景
CI/CD 管道里做 Schema 变更、多人协作演进表结构的团队;依赖"DDL 成功即生效"假设的 MySQL 思维迁移者。
Teams running schema changes through CI/CD pipelines or evolving tables collaboratively; MySQL-brained migrants assuming "DDL success means effective".
- 机制
StarRocks 的 DDL(如加列、建 MV 等)提交后转后台任务异步执行,客户端收到成功不代表变更完成,更不保证顺序与原子性。Fresha 的解法是自建迁移工具:层级命名规范、每项变更配显式 up/down SQL、在 StarRocks 内维护声明式 Schema 版本号(单一事实源)、轮询后台任务直到 FINISHED 才更新版本号、失败用配对 down SQL 回滚——相当于把 Flyway/Liquibase 对 MySQL/PG 免费提供的能力,自己重新实现了一遍。
StarRocks DDL (adding columns, building MVs, etc.) is handed to background tasks after submission — client success neither means completion nor guarantees ordering/atomicity. Fresha's answer was a homegrown migration tool: hierarchical naming conventions, explicit up/down SQL per change, a declarative schema version number kept inside StarRocks as the single source of truth, polling background tasks until FINISHED before bumping the version, rolling back with the paired down SQL on failure — reimplementing from scratch what Flyway/Liquibase give MySQL/Postgres users for free.
- 生产验证
来源 5:Fresha 工程师 Anton Borisov 执笔的案例(2025-12)——"由于 StarRocks 的许多 DDL 操作是异步执行的",团队构建 ActiveRecord 风格迁移工具,"持续轮询变更状态,直到所有后台任务达到最终的 FINISHED 状态后才会更新版本号"。
Source 5: case authored by Fresha engineer Anton Borisov (Dec 2025): "since many of StarRocks's DDL operations execute asynchronously," the team built an ActiveRecord-style migration tool that "keeps polling the change status until all background tasks reach the terminal FINISHED state before updating the version number."
- 证据等级
`单方声音`,具名客户工程师执笔(发表于厂商官方博客栏目,渠道已如实标注)。
`Single voice`, named customer engineer authorship (published in the vendor's official blog column; channel disclosed as-is).
StarRocks 年份:2025
Iceberg REST Catalog:库建得,删不得
单方声音
生态与信任
- 一句话
4.0.0 的 Iceberg REST Catalog 上,`CREATE DATABASE` 成功,紧接着 `DROP DATABASE` 就报"库不存在"——建库时没写 location 属性,删库时按 location 找不到,用户自己读源码才定位到。
On the 4.0.0 Iceberg REST catalog, `CREATE DATABASE` succeeds and the immediate `DROP DATABASE` fails with "database doesn't exist" — the create path never wrote the location property the drop path needs; the user had to read the source to find it.
- 窄场景
存算分离集群 + Iceberg REST Catalog(AWS S3 Tables),对 catalog 做库表 DDL 管理的用户;4.0.0。
Shared-data clusters with an Iceberg REST catalog (AWS S3 Tables) doing database DDL against the catalog; 4.0.0.
- 机制
`createDb` 实现未在 database properties 里写入 `location`,而 drop 路径依赖该属性定位库,导致"建完即删不掉"。用户做了源码走查 + 远程调试,定位到 `IcebergRESTCatalog.java` 的 createDb(L190)缺 location、drop 路径(L233)因缺 location 报错。FE 日志里只有一条 info,没有任何 error——排障全靠用户自己啃代码。
The `createDb` implementation omits `location` from the database properties while the drop path depends on that property to locate the database — "created but undroppable". The user did a source walkthrough plus a remote debug session, pinpointing `IcebergRESTCatalog.java`: createDb (L190) missing location, drop path (L233) erroring on its absence. FE logs show only an info line, no error at all — debugging meant reading the code yourself.
- 生产验证
来源 6:GitHub issue #65023,2025-11,用户在 4.0.0 存算分离集群(AWS S3 Tables + Iceberg REST)复现,附完整复现 SQL、FE 日志与源码级根因定位。
Source 6: GitHub issue #65023, Nov 2025 — user on a 4.0.0 shared-data cluster (AWS S3 Tables + Iceberg REST) reproduces it, with full repro SQL, FE logs, and source-level root-cause analysis.
- 证据等级
`单方声音`,GitHub 用户 issue(细节充分:版本、复现步骤、源码级定位)。
`Single voice`, GitHub user issue (detailed: version, repro steps, source-level diagnosis).
StarRocks 年份:2025
大版本原地升级不支持回退,跨版本只能迁移升级
单方声音
升级迁移成本账单
- 一句话
原地升级不支持回退、长抖动,v4/v5 到 v7 这种大跨度还得递增升级——得物最终选了迁移升级,代价是双倍硬件 + TiCDC 增量同步。
In-place upgrades support no rollback and cause long jitter; jumping v4/v5 to v7 also requires stepping through intermediate versions — Dewu chose migration upgrades instead, at the cost of double hardware plus TiCDC incremental sync.
- 窄场景
运行 v4/v5 老版本、想上 v7/v8 的生产集群;不能接受长时间性能抖动的业务。
Production clusters on v4/v5 aiming for v7/v8; workloads that cannot tolerate extended performance jitter.
- 机制
原地升级是滚动替换二进制,升级过程集群持续对外服务但有长时间性能抖动,且一旦开始无法回退;大版本跨度大时还需逐个中间版本递增升级,抖动时间翻倍。迁移升级(搭新集群 + TiCDC 双写/增量 + 流量切换)可灰度可回滚,但要多搭一套集群、原集群还得部署 TiCDC 做增量同步——升级成本变成实打实的硬件账单。
In-place upgrade is a rolling binary swap — the cluster keeps serving but with prolonged performance jitter, and once started there is no way back; large version jumps additionally require sequential intermediate upgrades, doubling the jitter window. Migration upgrade (build a new cluster + TiCDC incremental sync + traffic cutover) is gradual and rollback-safe, but means provisioning a whole second cluster and deploying TiCDC on the old one — the upgrade cost becomes a literal hardware bill.
- 生产验证
来源 1,得物 2025 年升级实录——原地升级"不支持回退、并且升级过程会有长时间的性能抖动",所有升级方案均采用迁移升级,"搭建新集群将产生额外的成本支出,同时,原集群还需要部署 TiCDC 组件用于增量同步"。
Source 1, Dewu's 2025 upgrade account — in-place upgrade "does not support rollback, and the upgrade process causes prolonged performance jitter"; all their upgrades used the migration approach: "building a new cluster incurs extra cost, and the old cluster also needs a TiCDC component deployed for incremental sync."
- 证据等级
`单方声音`,具名生产复盘(细节充分:两种升级方式的优劣对比与成本测算)。
`Single voice`, named production postmortem (detailed: side-by-side comparison of both approaches with cost analysis).
TiDB 年份:2025
TiCDC 在 v6.5.0 前经常延迟/OOM
单方声音
已修复于 v6.5.0
稳定与故障升级迁移
- 一句话
v6.5.0 之前的 TiCDC 运行稳定性差,经常出现数据同步延迟或 OOM——而迁移升级恰恰依赖它做增量同步。
Before v6.5.0, TiCDC was unstable in operation — sync lag and OOM were common — yet migration upgrades depend on it for incremental sync.
- 窄场景
v6.5.0 之前版本的 TiCDC changefeed;用 TiCDC 做跨集群迁移、容灾同步的场景。
TiCDC changefeeds on pre-v6.5.0 versions; cross-cluster migrations and DR sync built on TiCDC.
- 机制
TiCDC 的 checkpoint 推进是全局木桶——任一 Region 的 resolvedTs 卡住,全局同步延迟就涨;早期版本在内存管理与背压处理上有缺陷,大流量下易 OOM。对迁移升级而言,TiCDC 是增量同步的唯一官方通道,它的稳定性直接决定割接窗口。
TiCDC's checkpoint advance is a global weakest-link — if any Region's resolvedTs stalls, global sync latency rises; early versions had memory-management and backpressure defects that OOMed under heavy traffic. For migration upgrades, TiCDC is the only official incremental-sync channel, so its stability directly determines the cutover window.
- 生产验证
来源 1,得物 2025 年升级实录——"TiCDC 作为增量数据同步工具,在 v6.5.0 版本以前在运行稳定性方面存在一定问题,经常出现数据同步延迟问题或者 OOM 问题";得物称新版同步性能提升数十倍。
Source 1, Dewu's 2025 upgrade account — "as an incremental data sync tool, TiCDC had operational stability problems before v6.5.0, frequently showing data sync lag or OOM"; Dewu reports sync performance improved dozens of times in newer versions.
- 证据等级
`单方声音`,具名生产复盘。已修复于 v6.5.0(得物称新版同步性能提升数十倍),保留收录并标注修复版本。
`Single voice`, named production postmortem. Fixed in v6.5.0 (Dewu reports dozens-of-times sync improvement in newer versions); retained with the fixed-in label.
TiDB 年份:2025
BR 备份又慢又吃资源:每天备份 >8 小时,负载 +30%
单方声音
性能问题运维复杂度
- 一句话
集群每天备份耗时超过 8 小时,备份期间集群负载上升超 30%,撞上业务高峰直接让应用 RT 上升。
Daily backups took over 8 hours, pushing cluster load up more than 30% — colliding with peak business hours and raising application response times.
- 窄场景
数据量大的生产集群、备份窗口与业务高峰重叠的团队;需要定期全量备份的合规场景。
Large production clusters; teams whose backup windows overlap business peaks; regulated environments requiring periodic full backups.
- 机制
BR 做的是分布式快照备份,要协调所有 TiKV 节点扫描 SST 数据并上传外部存储;备份流量与业务流量争抢同一批 TiKV 的 CPU/磁盘/网络。分布式快照没有单机库"停写瞬间 cp 数据文件"那种轻量路径,备份即全集群参与的重活。
BR performs distributed snapshot backups, coordinating every TiKV node to scan SST data and upload it to external storage; backup traffic contends with business traffic for the same TiKV CPU/disk/network. A distributed snapshot has no lightweight equivalent of a single-node "quiesce and copy the data files" — backup is whole-cluster heavy labor.
- 生产验证
来源 1,得物 2025 年升级实录——"集群每天备份时间大于 8 小时,在此期间,数据库备份会导致集群负载上升超过 30%,当备份时间赶上业务高峰期,会导致应用 RT 上升"。
Source 1, Dewu's 2025 upgrade account — "daily backup time exceeds 8 hours; during backups the database raises cluster load by more than 30%, and when backups collide with business peaks, application RT rises."
- 证据等级
`单方声音`,具名生产复盘(细节充分:时长、负载涨幅、业务影响链条)。
`Single voice`, named production postmortem (detailed: duration, load increase, business-impact chain).
TiDB 年份:2025
单点查询与小规模场景:MySQL 更快更便宜
单方声音
性能问题成本账单
- 一句话
单点查询速度、单机 QPS——这些 MySQL 的基本盘,分布式数据库给不了;数据量没到量级时上 TiDB 是纯负担。
Point-query latency, single-node QPS — MySQL's home turf, which a distributed database cannot take; running TiDB before data reaches scale is pure overhead.
- 窄场景
点查为主、数据量单机能扛住的业务;被"MySQL 兼容"吸引、以為它是"更好的 MySQL"的团队。
Point-query-heavy workloads whose data fits on one machine; teams drawn by "MySQL compatibility" assuming it means "a better MySQL".
- 机制
分布式 SQL 的固定成本:每次查询都要经过 TiDB Server 解析优化、PD 取 TSO、走网络到 TiKV 取数(多跳 RPC + 两阶段提交)。单点查询下这些固定开销无法摊薄,延迟与 QPS 天花板天然低于单机 MySQL;小规模下还没有分片收益,等于只付税、不享受。
Distributed SQL has fixed costs: every query goes through TiDB Server parsing/optimization, fetches a TSO from PD, and travels the network to TiKV (multi-hop RPC + two-phase commit). On point queries these fixed costs cannot be amortized, so latency and QPS ceilings are structurally below single-node MySQL; at small scale there is no sharding benefit either — all tax, no reward.
- 生产验证
来源 1,得物 DBA 团队 2025 年升级实录引言——"能用分库分表能解决的问题尽量选择 MySQL,毕竟运维成本相对较低、数据库版本更加稳定、单点查询速度更快、单机 QPS 性能更高这些特性是分布式数据库无法满足的"(引自原文)。
Source 1, Dewu DBA team's 2025 upgrade account, in their own words — "problems solvable with sharding should stay on MySQL: lower ops cost, more stable versions, faster point queries, higher single-node QPS — a distributed database cannot deliver these" (quoted in translation).
- 证据等级
`单方声音`,具名生产复盘引言(大厂 DBA 团队的选型结论)。
`Single voice`, named production postmortem (a large-company DBA team's selection conclusion).
- 备注
与现有 [避坑] 卡「小数据量强行上 TiDB 是过度设计」主题重叠(该卡已含小数据量场景的分布式架构负担论述)。
overlaps the existing [Pitfall] card "Small Data Forced onto TiDB Is Over-Engineering" (that card already covers the distributed-architecture burden at small data scale).
TiDB 年份:2025
每个 class 一张 HNSW 图:几十个近乎空的 class 吃掉 6GB 内存
单方声音
性能问题运维复杂度
- 一句话
Weaviate 按 class 建独立的 HNSW 索引并全部加载进内存——内存占用正比于 class 数量而非向量数量,schema 膨胀的代价直接转嫁到内存账单。
Weaviate builds a separate HNSW index per class and keeps them all in memory — RAM usage scales with class count, not vector count, so schema bloat bills you directly in memory.
- 窄场景
按 KB/文档/租户动态建 class(或被上层应用如 Dify 自动建 class)的 schema 设计;class 数量达到几十上百但每个 class 数据量很小的部署。
Schema designs that create classes dynamically per knowledge base, document, or tenant (or upper-layer apps like Dify that auto-create classes); deployments with dozens to hundreds of classes, each holding little data.
- 机制
Weaviate 的存储模型是"一 class 一索引":每个 class 有自己独立的 HNSW 图和 LSM 结构,默认全部常驻内存。即使单个 class 只有几十条向量,其索引结构、commitlog、shard 元数据也要占一份常驻内存;class 数量一多,Go GC 压力随之上升(用户观察到 CPU 100–200% 的 GC 抖动),最终触发 OOM。官方文档的默认 `vectorCacheMaxObjects` 是 1e12(1 万亿),即默认不设限。
Weaviate's storage model is "one index per class": each class has its own HNSW graph and LSM structures, all resident in memory by default. Even a class with a few dozen vectors carries its own index structures, commit logs, and shard metadata as permanent residents; as class count grows, Go GC pressure climbs (the user observed 100–200% CPU in GC thrashing) until OOM. The official default for `vectorCacheMaxObjects` is 1e12 (one trillion) — effectively unlimited.
- 生产验证
来源 4,2025-12 Dify GitHub issue(用户自部署 Dify 1.8.1 + Weaviate)——Dify 每次上传文件自动建一个 `Vector_index_<uuid>_Node` class,schema 里攒了 50+ 个 class;`docker stats` 显示少量文档的 KB 照样占 3–6GB 内存、CPU 持续 100–200%、搜索变慢、启动变长,作者结论是"无法投入生产"。
Source 4, Dec 2025 Dify GitHub issue (self-hosted Dify 1.8.1 + Weaviate) — Dify auto-created one `Vector_index_<uuid>_Node` class per uploaded file, accumulating 50+ classes in the schema; `docker stats` showed 3–6GB RAM for a small knowledge base, sustained 100–200% CPU, slower search, longer startups; the author's conclusion was "cannot go to production".
- 证据等级
`单方声音`,GitHub issue(细节充分:class 命名、内存/CPU 数据、复现步骤)。
`Single voice`, GitHub issue (detailed: class naming, memory/CPU numbers, reproduction steps).
- 备注
该 issue 的表层根因是 Dify 的 class 设计,但暴露的是 Weaviate 的机制特性(per-class 常驻内存),故以机制为卡片主题收录。
the surface root cause was Dify's class design, but what it exposes is a Weaviate mechanism (per-class resident memory), so the card is themed on the mechanism.
Weaviate 年份:2025
单节点试用写入慢:第一次接触就留下错误第一印象
单方声音
性能问题
- 一句话
在单节点上试 YugabyteDB,万行级批量插入慢得离谱——而这恰恰是大多数人第一次接触它的方式。
Try YugabyteDB on a single node and batch inserts are shockingly slow — which is exactly how most people first meet it.
- 窄场景
单节点 yugabyted / Docker 本地试用,YSQL,万行以上批量插入。
Single-node yugabyted / Docker local trial, YSQL, batch inserts of 10k+ rows.
- 机制
YSQL 写入路径按"分片 + 同步复制"的生产拓扑设计:Raft quorum、事务状态 tablet、分布式提交都有固定成本,单节点下这些成本无法摊薄。另有 sequences 默认 cache=100,批量插入时对 sequences 表单行高频递增形成写入热点。官方回复直言:"We always deploy in production with sharding & synchronous replication, so it doesn't make sense to add optimizations that don't apply in this context (single node with no replication setups)"——单节点场景不值得做优化。
YSQL's write path is designed for the "sharded + synchronously replicated" production topology: Raft quorum, transaction-status tablets, and distributed commit all carry fixed costs that cannot be amortized on one node. Sequences also default to cache=100, so a batch insert hammers a single row of the sequences table. The official reply was blunt: "We always deploy in production with sharding & synchronous replication, so it doesn't make sense to add optimizations that don't apply in this context (single node with no replication setups)" — single-node setups are not worth optimizing.
- 生产验证
来源 3,2025-05 用户 NicoleSanders14 发帖——单节点、SSD、"decent hardware",10k+ 行批量插入 "pretty slow","I expected better";官方回复列出 5 点解释,核心是单节点无复制场景不在优化范围内。
Source 3, May 2025, user NicoleSanders14 — single node, SSD, "decent hardware", yet "pretty slow" insert speeds on 10k+ row batches: "I expected better." The official reply listed five explanations, boiling down to single-node-without-replication being out of optimization scope.
- 证据等级
`单方声音`,厂商论坛客户发帖(细节充分:硬件配置、批量规模、官方回复原文;注:为试用场景,非生产事故)。
`Single voice`, customer thread on the vendor forum (detailed: hardware, batch size, verbatim official reply; note: a trial scenario, not a production incident).
YugabyteDB 年份:2025
DMS 迁移作业生产环境崩溃:失败就从头再来
单方声音
升级迁移
- 一句话
Database Migration Service 迁 AlloyDB,生产环境作业崩了就是重来——Bayer 的工程师建议直接按"预估时间的两倍"做计划。
Migrating to AlloyDB with Database Migration Service crashed in production — the Bayer engineer who lived through it recommends budgeting double the time you think you need.
- 窄场景
从 Cloud SQL/自建 PostgreSQL 经 DMS 迁往 AlloyDB 的生产迁移;数据量大、迁移窗口紧张的团队。
Production migrations from Cloud SQL or self-managed PostgreSQL to AlloyDB via DMS; large datasets with tight cutover windows.
- 机制
DMS 迁移基于逻辑复制:先做初始快照全量同步,再持续复制增量。迁移作业在早期阶段有多种崩溃方式,且崩溃后是整个作业重来而非断点续传——全量快照阶段的时间成本直接翻倍。
DMS migration runs on logical replication: an initial full snapshot sync followed by continuous change replication. In its early stages the migration job has multiple ways to crash, and a crash restarts the entire job rather than resuming from a checkpoint — the full-snapshot time cost doubles outright.
- 生产验证
来源 6,Bayer 工程师 Aaron Joyce 在 Google Cloud Next 2024 分享——非生产环境迁移"really well",但"we get to production it didn't go quite as well";原话建议 "leave about twice the length of time that you think that the migration is going to take","that job especially in the beginning has a lot of ways it can crash and if it crashes you're starting it over again"。
Source 6, Bayer engineer Aaron Joyce at Google Cloud Next 2024 — the non-production migration went "really well," but "we get to production it didn't go quite as well"; verbatim advice to "leave about twice the length of time that you think that the migration is going to take," because "that job especially in the beginning has a lot of ways it can crash and if it crashes you're starting it over again."
- 证据等级
`单方声音`,具名客户工程师会议分享(厂商大会上的客户演讲,细节充分:非生产 vs 生产对比、两倍时间建议)。
`Single voice`, named customer engineer conference talk (customer speaking on a vendor stage; detailed: non-prod vs prod contrast, the 2x-time recommendation).
Google AlloyDB(AlloyDB for PostgreSQL) 年份:2024
读 4000 个分区并发要 4000 个 slot:slot 模型像个黑盒
单方声音
运维复杂度
- 一句话
BigQuery 的 slot 是"读一个分区占一个 slot"这种粒度的资源——想并发读 4000 个分区,你得有 4000 个 slot,而 slot 的分配过程对你完全不透明。
A BigQuery slot reads roughly one partition at a time — to read 4,000 partitions concurrently you'd need 4,000 slots, and the slot allocation process is completely opaque to you.
- 窄场景
分区数上千的大表、重度并发查询的项目;在按量计费与容量预留(Editions)之间做选择的中大型团队。
Large tables with thousands of partitions; projects with heavy query concurrency; mid-to-large teams choosing between on-demand and capacity (Editions) pricing.
- 机制
slot 是 BigQuery 内部计算单元(CPU+内存),按量模式下用的是共享池 + 突发容量,容量预留模式下按 slot-hour 付费;据作者与 Google 团队面谈确认,slot 一次只能读一个分区,并发读 4000 个分区理论上需要 4000 个 slot——这也是"单表 4000 分区上限"这个配额存在的保护性原因。用户侧看不到 slot 分配明细,查询忽快忽慢只能归因于"slot 天气"(社区语:BQ weather / slot contention)。
Slots are BigQuery's internal compute units (CPU+memory); on-demand uses a shared pool with burst capacity, reservations bill per slot-hour; per the author's session with the Google team, a slot can only read one partition at a time, so concurrently reading 4,000 partitions would theoretically require 4,000 slots — which is also the protective rationale behind the 4,000-partitions-per-table quota. Users see no slot-allocation breakdown; queries that randomly run slow get attributed to "slot weather" (community term: BQ weather / slot contention).
- 生产验证
来源 8,Christophe Oudar 2024-01-29 跟进文章——此前发表"14 个 BigQuery 短板"后被 Google 团队约谈 1 小时+,逐条跟进;slot/分区并发关系是 Google 团队亲口确认的机制;另有 HN 同期评论印证:"really big queries in BQ can take 10x the slots on some runs just because of 'BQ weather' and slot contention"。
Source 8, Christophe Oudar's Jan 29 2024 follow-up — after his "14 BigQuery shortfalls" post, Google's team invited him for an hour-plus session and walked through each point; the slot/partition-concurrency relationship was confirmed by the Google team directly; a contemporaneous HN comment corroborates the opacity: "really big queries in BQ can take 10x the slots on some runs just because of 'BQ weather' and slot contention."
- 证据等级
`单方声音`,个人博客(与 Google 团队面谈后的一手跟进记录)。
`Single voice`, personal blog (first-hand follow-up record of a session with the Google team).
Google BigQuery 年份:2024
DBFS 存 init script 被废弃:存量脚本要么搬家,要么集群起不来
单方声音
运维复杂度升级迁移
- 一句话
存在 DBFS 里的集群初始化脚本被官方废弃,不搬家集群就起不来——"It will be a re-work to move all?",还被告知 CLI 导不了 shell 脚本只能走网页。
Cluster init scripts stored in DBFS were deprecated — don't move them and the cluster won't start. "It will be a re-work to move all?" — and the CLI can't even import shell scripts, web UI only.
- 窄场景
用 DBFS 存了大量 init script 的老工作区;靠脚本装系统包/配代理的集群。
Older workspaces with many DBFS init scripts; clusters relying on scripts to install system packages or configure proxies.
- 机制
legacy init script 存储位置被废弃,官方要求搬到云存储/workspace 文件/UC volumes——是强制迁移,不是可选优化。迁移工具链还不完整:CLI 不支持导入 shell 脚本。
Legacy init-script storage was deprecated — a forced migration, not an optional optimization — with the official path being cloud storage, workspace files, or UC volumes. The migration tooling is incomplete: no CLI import for shell scripts.
- 生产验证
—
Source 20: official community thread, circa 2024 — a user asks what happens to existing scripts; the deprecation and the move paths are confirmed, plus complaints about web-only import.
- 证据等级
`单方声音`,社区帖(年代 2024,但废弃是强制性的)。
`Single voice`, community thread (2024, but the deprecation is mandatory).
Databricks 年份:2024
Query history 只留 30 天:Snowflake 给 365 天,查旧账得自建管线
单方声音
运维复杂度
- 一句话
Databricks 查询历史(UI/API/`system.query.history`)30 天自动删除,想查 30 天前的记录?自己建增量持久化管线。可配置保留期 2026-09 才进 Beta。
Databricks query history (UI/API/`system.query.history`) auto-deletes after 30 days — need something older? Build your own incremental persistence pipeline. Configurable retention only entered Beta in Sep 2026.
- 窄场景
做季度成本复盘、合规审计要查历史查询的团队。
Quarterly cost reviews; compliance audits needing historical queries.
- 机制
保留期默认 30 天,过期自动删;2024 年用户问 API 能否返回 30 天前的数据,官方答复"只能等 system tables roadmap"。可配置保留 2026-09 进 Beta,截至 2026-10 仍非 GA。
30-day default retention, auto-deleted; in 2024 a user asking the API for data older than 30 days was told to wait for the system-tables roadmap. Configurable retention entered Beta in Sep 2026 — still not GA as of Oct 2026.
- 生产验证
—
Source 42: official community thread, Apr 2024 — user plea.
- 证据等级
`单方声音`。
`Single voice`.
Databricks 年份:2024
无模式 + 事务门槛:Infisical 举家迁往 PostgreSQL
单方声音
运维复杂度升级迁移
- 一句话
Infisical 因 MongoDB 的事务配置门槛、无模式导致的数据不一致、缺失关系型能力(CASCADE),历时 3-4 个月把全站迁往 PostgreSQL,迁移后数据库账单下降 50%。
Infisical abandoned MongoDB for PostgreSQL after a 3–4 month migration because of the transaction setup bar, schemaless data inconsistency, and missing relational capabilities (CASCADE); the database bill dropped 50% after the move.
- 窄场景
数据实际是关系型的产品、需要支持自托管/客户 POC 的 SaaS。
Products whose data is genuinely relational; SaaS vendors that need self-hosting or customer POCs.
- 机制
多文档事务要求副本集/集群模式,单节点跑不起来,客户连 POC 都跑不通;无 CASCADE,应用层手写级联删除删不干净,留下悬空数据;无模式 + 校验只在 Mongoose 层实现,任何绕过 ODM 的写入直接产生数据不一致;关系型查询只能靠多段 $lookup 模拟,低效且要靠扩容数据库+应用实例来扛。
Multi-document transactions require replica-set/cluster mode, so customers could not even run a POC on a single node; no CASCADE meant hand-written cascading deletes that never fully worked, leaving dangling resources; schema validation lived only in the Mongoose layer, so any write bypassing the ODM produced inconsistent data; relational queries had to be simulated with chains of $lookup stages, which was inefficient enough to require scaling up both database and application tiers.
- 生产验证
Infisical 官方博客《The Great Migration from MongoDB to PostgreSQL》(2024);密钥管理 SaaS,日均处理 5000 万+ secrets;迁移窗口 6 小时只读、零数据丢失。
Infisical official blog, "The Great Migration from MongoDB to PostgreSQL" (2024); a secrets-management SaaS processing 50M+ secrets daily; six-hour read-only migration window with zero data loss.
MongoDB 年份:2024
小版本升级引入复制停滞:官方标"已修复",生产每周仍复现 2-3 次
单方声音
稳定与故障生态与信任
- 一句话
从 4.2 升级到 4.4.28 后,生产分片集群的 secondary 复制随机停滞(CursorNotFound),CPU/磁盘归零,只能重启恢复;按官方 JIRA(SERVER-70155,标记为已修复)升到 4.4.29 后仍每周复现 2-3 次。
After upgrading from 4.2 to 4.4.28, a production sharded cluster's secondaries randomly stall replication (CursorNotFound) with CPU/disk dropping to zero — only a restart recovers; upgrading to 4.4.29 per the official JIRA (SERVER-70155, marked fixed) did not help, and stalls still recur 2–3 times a week.
- 窄场景
4.4.28/4.4.29 分片集群 secondary,7×24 高写入。
4.4.28/4.4.29 sharded-cluster secondaries under 24/7 heavy writes.
- 机制
oplog fetcher 游标丢失后复制停滞且不自愈;oplog 窗口充足(22 小时)、重启即恢复,指向版本回归而非配置问题。官方标记修复但生产仍复现,修复的可信度存疑。
Once the oplog fetcher loses its cursor, replication stalls and never self-heals; the oplog window was ample (22 hours) and restarts recover instantly, pointing at a version regression rather than misconfiguration. The vendor marked it fixed while production still reproduces it, which puts the fix's credibility in doubt.
- 生产验证
MongoDB 官方论坛帖《Mongo replication stalls》(2024-03~05),用户 Fory Horio;生产分片环境,6-7 年老集群,4.2 稳定运行 1-2 年无此问题,含完整报错日志与时间线(5 月仍在复现)。
MongoDB official forums thread "Mongo replication stalls" (2024-03 to 2024-05), user Fory Horio; production sharded environment, a 6–7 year old cluster that ran 4.2 stably for 1–2 years, with full error logs and a timeline showing recurrences into May.
- 证据等级
`单方声音`(具名论坛用户生产实录,版本/报错/时间线俱全)
—
MongoDB 年份:2024
租户端点永远对外可解析:IP 白名单只做在 L7,SSO 登录页还能枚举用户名 [来源存疑]
单方声音
来源存疑
生态与信任
- 一句话
Snowflake 的"IP 过滤"发生在共享的 Layer7 ALB 上、返回 HTTP 403 而非网络层丢包;即使只想走 PrivateLink,也关不掉对外可解析的地址;SSO 登录页返回 200 还能做用户名枚举——"Tenant's should not have externally resolvable endpoints to worry about IP filtering"(租户本不该有对外可解析的端点、还得自己操心 IP 过滤)。
Snowflake's "IP filtering" happens on a shared Layer-7 ALB returning HTTP 403, not network-layer drops; even PrivateLink-only deployments can't turn off the publicly resolvable address; the SSO login page returns 200 and enables username enumeration — "Tenant's should not have externally resolvable endpoints to worry about IP filtering."
- 窄场景
对公网暴露面零容忍的金融/政务租户;做 PrivateLink 私有化部署的团队。
Financial/government tenants with zero tolerance for public exposure; teams deploying PrivateLink.
- 机制
多租户共享架构下租户没有真正的私有托管能力;IP 白名单维护变成"永远补不完的 SaaS 防火墙名单维护闭环"。
The multi-tenant shared architecture offers no true private hosting; IP allowlist maintenance becomes "an open loop in Ops for maintaining these SAAS software firewall lists."
- 生产验证
来源 26:Hacker News 2024-06-02 匿名单条长评论(自述有租户端安全实施经验)——实测随手挑一个租户子域名即对外可解析并返回 IP 过滤的 HTTP 403;"you can't go private exposure via private-link only, without also having externally resolvable addresses existing for your instance";"Snowflake's fault here for never building true dedicated org capabilities with private hosting so far as I can tell"。
—
- 证据等级
`单方声音`,[来源存疑]:匿名单方声音,技术细节具体、可对照官方文档核验,但未找到第二独立客户信源复述。
`Single voice`, [source questionable]: anonymous single voice, technically concrete and checkable against official docs, but no second independent customer source.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2024
Definite:整仓从 Snowflake 搬到自建 DuckDB,降本超 70%,代价是自研读写分离
单方声音
升级迁移
- 一句话
数据产品公司把内部全部分析数据与查询从 Snowflake 迁到自建 DuckDB,成本 "radical reduction in cost (>70%)",动机是成本与 "vendor lock-in concerns";代价是 DuckDB 单写者设计逼他们自研双实例读写分离,"required a good bit of work"(花了不少功夫)。
The data-product company moved all internal analytics data and queries from Snowflake to self-hosted DuckDB — "radical reduction in cost (>70%)", motivated by cost and "vendor lock-in concerns"; the price was DuckDB's single-writer design forcing a DIY dual-instance read/write split, which "required a good bit of work."
- 窄场景
中小规模分析 workload、对 Snowflake 最小 warehouse 规格都嫌贵的团队。
Small-to-mid analytics workloads for which even Snowflake's smallest warehouse feels oversized.
- 机制
Snowflake 按 warehouse 规格×时间计费,小 workload 也要为整台"仓库"付费;DuckDB 嵌入式按实际资源付费,但单写者锁需要自研同步机制。
Snowflake bills per warehouse size×time — small workloads still pay for a whole "warehouse"; DuckDB bills actual resources, but the single-writer lock needs a homemade sync mechanism.
- 生产验证
来源 35:Definite 公司博客 2024——1TB 数据、16 vCPU/64GB 假设下,每天 12 小时使用自建 DuckDB 比最小 Snowflake 仓库便宜约 55%、比 Small 便宜约 77%;DuckDB SQL 方言接近 Postgres,大部分查询可直接翻译,数据存 Parquet 搬走无锁定。
—
- 证据等级
`单方声音`,客户自家迁移复盘;70% 数字被多篇第三方文章转述,但均为转述同一篇博客,无第二个独立实测。
`Single voice`, customer's own migration postmortem; the 70% figure is widely requoted but all trace to the same blog — no second independent measurement.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2024
dbt-snowflake 把已建好的 Dynamic Table 当成新的,触发全量重建并在裸 schema 名上报错
单方声音
生态与信任
- 一句话
一个早已物化好的 dynamic table 突然被 dbt 当成"首次创建",走完备份-重命名-重建全套流程,并在 `Cannot perform CREATE TABLE. This session does not have a current schema`(090106)上失败;用户确认 dbt Cloud 与本地 Docker 均复现,问题次日自行恢复(怀疑 Snowflake 侧变化,未能复现)。
A long-materialized dynamic table was suddenly treated by dbt as "first creation," running the full backup-rename-rebuild flow and failing on `Cannot perform CREATE TABLE. This session does not have a current schema` (090106); the user reproduced it on dbt Cloud and local Docker; it self-recovered the next day (suspected Snowflake-side change, never reproduced).
- 窄场景
dbt-snowflake + Dynamic Table 的增量管线;依赖 session 默认 schema 推断的部署。
dbt-snowflake + Dynamic Table incremental pipelines; deployments relying on session-default schema inference.
- 机制
dbt 判断 dynamic table 是否已物化的逻辑可能因"检查了未限定的位置"而看不见已存在的表;重建流程里 `rename to MONTHLY__dbt_backup` 用不限定库名的裸名,依赖 Snowflake 按 session 默认 schema 推断——"Since that behavior is subject to change, dbt might want to be more explicit";Snowflake 官方文档至今保留相关警告:Snowflake 计划 2026 年 9 月扩大 string/binary 默认列宽,低版本 dbt-snowflake 在特定增量模型上会构建失败——Snowflake 侧行为变更反复击穿 dbt-snowflake 的增量物化逻辑。
dbt's "does this dynamic table exist" check may look in an unqualified location and miss the existing table; the rebuild's `rename to MONTHLY__dbt_backup` uses a bare unqualified name, depending on Snowflake's session-default schema inference — "Since that behavior is subject to change, dbt might want to be more explicit"; Snowflake's own docs still carry a related warning: a planned Sep 2026 widening of default string/binary column widths will break older dbt-snowflake versions on certain incremental models — Snowflake-side behavior changes keep punching through dbt-snowflake's incremental materialization logic.
- 生产验证
来源 48:dbt-snowflake 官方仓库用户 issue #1017(2024-05-02 事件,附完整 run artifacts SQL 与 dbt debug 环境信息)。
—
- 证据等级
`单方声音`,真实用户 issue(模板规范、复现环境完整);同 issue 提到 dbt 工程师在 Slack 同期讨论关联 bug,可作侧面参考。
`Single voice`, real user issue (proper template, complete repro environment); the thread notes a dbt engineer discussing a related bug in Slack as a side reference.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2024
TiKV 磁盘空间失控:日志堆积、Titan 拿空间换性能
单方声音
稳定与故障运维复杂度
- 一句话
TiKV 磁盘被吃满的原因有四个:老版本日志不 rotation、混部下占位文件重复占用、GC 过慢、Titan 引擎默认拿空间换性能——单个 titandb 实例曾占 937G/1.1T。
TiKV disks filled for four reasons: no log rotation in old versions, duplicated placeholder files under mixed deployment, slow GC, and the Titan engine trading disk space for write performance by default — one titandb instance hit 937G of 1.1T.
- 窄场景
TiKV 独立部署用作 KV 存储(非 TiDB SQL 场景)、混部、老版本的集群;磁盘水位监控不完善的团队。
TiKV deployed standalone as a KV store (not the TiDB SQL stack), mixed deployments, older versions; teams with weak disk-waterline monitoring.
- 机制
rocksdb.info/raftdb.info 在老版本无日志轮转会无限堆积;混部时 space_placeholder_file 在每块盘重复占位;默认关闭的 gc.enable-compaction-filter 让过期数据回收过慢;Titan 把 value 存 blob 文件、默认 discardable-ratio=0.5,等于用磁盘空间换写入性能,空间放大显著。
rocksdb.info/raftdb.info grew unbounded without log rotation in old versions; space_placeholder_file was duplicated per disk under mixed deployment; the default-off gc.enable-compaction-filter left expired data reclaimed too slowly; Titan stores values in blob files with a default discardable-ratio of 0.5 — disk space traded for write performance, with significant space amplification.
- 生产验证
来源 5,vivo 互联网技术团队(袁建伟)约 2024-11 排查实录——逐项定位上述四因,给出各因的处置(开 rotation、调占位、开 compaction-filter、调 Titan 参数)。
Source 5, vivo internet technology team (Yuan Jianwei), circa Nov 2024 investigation — all four causes pinned down one by one, with each fix documented (enable rotation, adjust placeholders, enable compaction-filter, tune Titan parameters).
- 证据等级
`单方声音`,具名生产复盘(细节充分:四因逐项定位与处置)。
`Single voice`, named production postmortem (detailed: four causes isolated with fixes).
- 备注
该复盘的场景是 TiKV 独立用作 KV 存储(Redis 协议兼容层),非 TiDB SQL 集群,机制结论对 TiKV 组件通用,但场景口径已如实标注。
the postmortem's scenario is TiKV standalone as a KV store (Redis-protocol compatibility layer), not a TiDB SQL cluster; the mechanism conclusions generalize to the TiKV component, and the scope is labeled as observed.
TiDB 年份:2024
还原状态接口谎报 SUCCESS:恢复还在跑,API 先说成功了
单方声音
运维复杂度
- 一句话
filesystem 备份的 restore 是异步的,但 GET 备份状态接口在 restore 刚触发时就返回 SUCCESS——运维以为恢复完了,实际数据还在写入中。
Filesystem backup restores are asynchronous, but the GET backup-status endpoint returns SUCCESS moments after the restore is triggered — operators think the restore finished while data is still being written.
- 窄场景
用 backup-filesystem 模块做跨机器恢复/灾备演练的自部署用户;用脚本轮询备份状态接口做自动化恢复的流程。
Self-hosted users running cross-machine restores or disaster-recovery drills with the backup-filesystem module; scripted recovery flows that poll the status endpoint.
- 机制
POST /v1/backups/filesystem/{id}/restore 只登记恢复任务并立即返回 STARTED;随后 GET 同一接口返回的 status 反映的是"任务已登记"而非"恢复已完成",于是立刻显示 SUCCESS。真正的进度只能通过再发一次 restore(返回 "already in progress" 错误)或等待一段时间后查数据来确认——状态接口与真实进度脱节。
POST /v1/backups/filesystem/{id}/restore only registers the restore job and immediately returns STARTED; the subsequent GET on the same endpoint reports the job as registered rather than actually completed, so it shows SUCCESS right away. True progress can only be detected by issuing another restore (which errors with "already in progress") or by waiting and then checking the data — the status API is decoupled from actual progress.
- 生产验证
来源 5,2024-07 论坛帖——用户把备份从 nodeA 恢复到另一台机器,POST 恢复返回 STARTED 后立刻 GET 状态得到 SUCCESS,误以为完成;再次 POST 触发恢复才暴露 "restoration ... already in progress";等待一段时间后数据才真正出现。用户原话评价该接口"misleading"。
Source 5, Jul 2024 forum thread — the user restored a backup from nodeA onto another machine; the POST returned STARTED and an immediate GET returned SUCCESS, suggesting completion; re-issuing the POST exposed "restoration ... already in progress"; the data only appeared after waiting. The user called the endpoint "misleading" in their own words.
- 证据等级
`单方声音`,论坛帖(细节充分:curl 命令、返回体原文、排查过程)。
`Single voice`, forum thread (detailed: curl commands, verbatim response bodies, investigation steps).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's this card's topic may overlap an existing [Pitfall] card on this site on this site.
Weaviate 年份:2024
InnoDB 间隙锁:删了 0 行的 DELETE,也能锁住别人的 INSERT
单方声音
性能问题运维复杂度
- 一句话
REPEATABLE READ 下,一个"删了 0 行"的 DELETE 也会拿走间隙锁,让毫不相干的 INSERT 排队甚至死锁;死锁日志只保留最近一次,排查靠人肉拼 performance_schema.data_locks。
—
- 窄场景
默认隔离级别 REPEATABLE READ;外键/非唯一索引上的范围删除与并发插入混跑;Rails/ORM 长事务。
Default isolation level REPEATABLE READ; range deletes on foreign keys/non-unique indexes mixed with concurrent inserts; Rails/ORM long transactions.
- 机制
InnoDB 在 RR 下对非唯一索引扫描加 next-key lock(record + gap);DELETE ... WHERE table1_id=75 即使匹配 0 行,也会锁住 (74,80] 区间;另一事务往该区间 INSERT 拿 insert intention lock 等待;双向"删除+插入"即成环。SHOW ENGINE INNODB STATUS 只保留 LATEST DETECTED DEADLOCK,高频死锁下现场稍纵即逝,需 innodb_print_all_deadlocks 或 pt-deadlock-logger 常开。
InnoDB takes next-key locks (record + gap) on non-unique index scans under RR; DELETE ... WHERE table1_id=75 matching zero rows still locks the (74,80] range; another transaction's INSERT into that range waits on an insert intention lock; "delete + insert" in both directions closes the cycle. SHOW ENGINE INNODB STATUS keeps only LATEST DETECTED DEADLOCK — under frequent deadlocks the evidence vanishes instantly; you need innodb_print_all_deadlocks or pt-deadlock-logger always on.
- 生产验证
个人博客(2023-06-12,作者 nikhil):Sentry 捕获死锁,用 performance_schema.data_locks 实录 gap/insert-intention 锁等待链,定位到 AASM 状态机把多次回调塞进同一长事务;修复=拆小事务 + 删掉"删不存在行"的多余 DELETE。
Personal blog (Jun 12, 2023, author nikhil): deadlocks caught by Sentry, gap/insert-intention wait chains reconstructed from performance_schema.data_locks, root-caused to an AASM state machine stuffing multiple callbacks into one long transaction; fix = split transactions + remove the redundant "delete rows that don't exist" DELETE.
- 证据等级
`单方声音`,来源性质:具名个人博客(含锁表实录与修复验证;在测试环境由 Sentry 捕获,非严格生产事故,但机制与生产一致)。
—
MySQL 年份:2023
串行化隔离 + FIFO 锁队列:ETL 的表切换被一张报表的长查询卡住
单方声音
性能问题稳定与故障
- 一句话
长查询拿着 AccessShareLock,ETL 的 ALTER TABLE … RENAME(AccessExclusiveLock)只能等;队列是 FIFO,排在写锁后面的读查询也要一起等——一次日常的表切换能拖成连锁等待。
A long query holding an AccessShareLock blocks the ETL's ALTER TABLE … RENAME (AccessExclusiveLock); the queue is FIFO, so read queries arriving after the write lock wait too — one routine table swap cascades into chained waits.
- 窄场景
ETL 用"建临时表 + RENAME 切换"模式发布、同时有 BI 长查询的集群。
Clusters where ETL publishes via "build temp table + RENAME swap" while long BI queries run.
- 机制
Redshift 有 3 种表锁:SELECT/UNLOAD 拿 AccessShareLock,会阻塞 DDL 的 AccessExclusiveLock(如 ALTER TABLE RENAME);查询按到达顺序进队列,排在写锁后面的读也要等它。Faire 踩坑时 Redshift 只支持 serializable 隔离(snapshot isolation 2022-05 才引入),锁冲突面更大。Faire 的两种 workaround:给 RENAME 步骤设极短超时 + 大量重试做成 best-effort,或改用 DELETE+INSERT 模式(但 ETL 任务数翻倍,Airflow SLA 承压);且"不能指望每个用数据的人都遵守最佳实践"。
Redshift has three table lock types: SELECT/UNLOAD take AccessShareLock, which blocks the AccessExclusiveLock needed by DDL like ALTER TABLE RENAME; queries enter the queue in arrival order, so reads queued behind a write lock also wait. At the time of Faire's incident Redshift supported only serializable isolation (snapshot isolation arrived May 2022), widening the lock-conflict surface. Faire's two workarounds: give the RENAME step a very short timeout with many retries as a best-effort process, or switch to a DELETE+INSERT pattern (which doubles Airflow task count and strains DAG SLAs); and "you can't expect everyone touching data to follow best practices religiously."
- 生产验证
来源 1,Faire 2023-02 具名复盘——早 9 点 Mode 报表长查询 + 9:15 ETL 的 RENAME 表切换真实踩坑,附锁类型与队列顺序的完整分析。
Source 1, Faire, Feb 2023, named retrospective — a real collision between a 9 AM Mode report's long query and a 9:15 AM ETL RENAME swap, with full analysis of lock types and queue ordering.
- 证据等级
`单方声音`,具名公司工程博客(细节充分:锁类型、复现场景、两种 workaround 的代价分析)。
`Single voice`, named company engineering blog (detailed: lock types, reproduction scenario, cost analysis of both workarounds).
- 备注
Faire 发文时已注明 snapshot isolation 于 2022-05 引入,串行化误杀部分缓解;但锁类型与 FIFO 队列语义未变,故保留收录,未作"已修复"标注。
Faire noted snapshot isolation was introduced in May 2022, partially relieving serializable false conflicts; but lock types and FIFO queue semantics are unchanged, so this is retained without a "fixed in" label.
Amazon Redshift 年份:2023
Serverless 按 60 秒起收:2–3GB 的小仓库一个月 260–520 美元
单方声音
成本账单
- 一句话
Redshift Serverless 每个查询按 60 秒最低时长 × RPU 数计费——亚秒级查询也一样,一个每天跑十几条 ETL 的小仓库月账单轻松上几百美元。
Redshift Serverless charges every query for at least 60 seconds × RPU count — sub-second queries included — so a tiny warehouse running a dozen daily ETL queries easily bills hundreds of dollars a month.
- 窄场景
数据量小(GB 级)、查询短平快、每天只有少量查询的轻量场景。注意区分 Redshift provisioned(常驻集群)与 Redshift Serverless。
Small data volumes (GB scale), short fast queries, only a handful of queries per day. Note the distinction between provisioned Redshift (always-on clusters) and Redshift Serverless.
- 机制
计费 = RPU 小时 × 每秒单价,但每次查询至少按 60 秒计;默认最小 32 RPU 下,单次查询最低约 0.2 美元(60×32×0.375/3600)。作者实测:几乎每次查询都被收了 1920 秒(60s×32RPU),SYS_SERVERLESS_USAGE 表可查。
Billing = RPU-hours × per-second price, but each query is charged a 60-second minimum; at the 32-RPU default floor, one query costs at least ~$0.20 (60×32×0.375/3600). The author measured nearly every query billed 1,920 seconds (60s×32 RPU), auditable in SYS_SERVERLESS_USAGE.
- 生产验证
来源 6,HackerNoon 2023 年工程师为客户做的实测——预期月 60 美元,首日试用金就掉了约 25 美元;按 10–20 条日 ETL + 100–200 条手动查询估算,月 260–520 美元,还没算 BI 工具连接的费用;结论是"在 AWS 拿掉 60 秒起收之前,BigQuery 仍是小仓库首选"。
Source 6, HackerNoon, 2023, a field engineer's measured analysis for a client — expected $60/month, burned ~$25 of trial credit on day one; estimated $260–520/month for 10–20 daily ETL queries plus 100–200 manual queries, before BI-tool connection costs; conclusion: "Until now, BigQuery has been the top 1 for small data warehouses. When AWS removes a minimal 60-second period, it can change."
- 证据等级
`单方声音`,一线工程师账单实测(含 SYS_SERVERLESS_USAGE 查询与计算过程)。
`Single voice`, field engineer's measured bill analysis (includes SYS_SERVERLESS_USAGE queries and the arithmetic).
- 备注
AWS 后来推出 4 RPU 档(2025),小负载可降本;但 60 秒最低计费规则未变。本卡主题可能与本站 [避坑] 卡重叠(Serverless 成本)。
AWS later added a 4-RPU tier (2025) that lowers cost for small workloads; the 60-second minimum rule is unchanged. Topic may overlap existing [Pitfall] cards (Serverless cost).
Amazon Redshift 年份:2023
Spectrum:有些算子只能回集群算,S3 外表查询说慢就慢
单方声音
来源存疑
性能问题
- 一句话
Redshift Spectrum 查 S3 外表时,部分 SQL 算子下推不了,只能回 Redshift 集群本地执行——"透明加速"的期望落空,慢的原因还不好定位。
When Redshift Spectrum queries S3 external tables, some SQL operators cannot be pushed down and fall back to executing on the Redshift cluster — the "transparent acceleration" promise breaks, and the slowness is hard to attribute.
- 窄场景
冷热分层(热数据在集群、冷数据在 S3)、大量用 Spectrum 做联邦查询的团队。
Hot/cold tiering (hot data on cluster, cold data on S3) with heavy Spectrum federated querying.
- 机制
Spectrum 把部分操作下推到 S3 层的独立 fleet 执行,但并非所有算子都支持下推;不支持的回退到集群计算,跨层数据搬运与集群资源争抢让延迟不可预测。
Spectrum pushes part of the work to its S3-side fleet, but not all operators support pushdown; unsupported ones fall back to cluster compute, and cross-tier data movement plus cluster resource contention make latency unpredictable.
- 生产验证
来源 4,PeerSpot 2023-05-04 匿名 Data Scientist(平台标注 Real User)——"some SQL commands have to run on Redshift instead of Redshift Spectrum, which slows down a few things";"有些命令走 Redshift、有些走 Spectrum,which is problematic"。
Source 4, PeerSpot, May 4 2023, anonymous Data Scientist (site-labeled Real User) — "some SQL commands have to run on Redshift instead of Redshift Spectrum, which slows down a few things"; "some of the commands are run on Redshift itself, and some of the commands are completed by Spectrum, which is problematic."
- 证据等级
`单方声音` `[来源存疑]`,匿名点评用户的一句话评论(细节有限,独立性无法确认)。
`Single voice` `[Questionable source]`, one anonymous reviewer's short comment (limited detail, independence cannot be confirmed).
Amazon Redshift 年份:2023
并发写上限 20:文档里没写的 DML 天花板
单方声音
性能问题
- 一句话
同一张表最多 20 条并发的 COPY/INSERT/MERGE/UPDATE/DELETE——"Snowflake users might not be aware of such a concurrent write limit as this is not in Snowflake documentation"(文档里没有,用户无从得知);"Snowflake is designed for high-volume high-concurrency reading and not writing."
At most 20 concurrent COPY/INSERT/MERGE/UPDATE/DELETE statements against the same table — "Snowflake users might not be aware of such a concurrent write limit as this is not in Snowflake documentation"; "Snowflake is designed for high-volume high-concurrency reading and not writing."
- 窄场景
把 Snowflake 当近实时写入目标的团队(streaming ingestion、高频 MERGE、hybrid 表写入)。
Teams using Snowflake as a near-real-time write target (streaming ingestion, high-frequency MERGE, hybrid-table writes).
- 机制
限制指向 Global Service Layer 元数据仓库的约束;超过上限的并发写会被限流或失败,而官方文档未明确写出该数字。
The limit traces to Global Service Layer metadata-store constraints; concurrent writes beyond it get throttled or fail, and the number is not stated in official docs.
- 生产验证
来源 14:Slim Baltagi 2023-11-08——"Snowflake has a built-in limit of 20 DML statements that target the same table concurrently… this is not in Snowflake documentation."
—
- 证据等级
`单方声音`,独立工程师长文实测;该限制未见官方文档记载亦未见官方否认。
`Single voice`, independent engineer's long-form measurement; the limit appears in no official documentation and has not been officially denied.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2023
神秘报错码官方文档查不到,只能去 GitHub 翻驱动源码,官方"暂无计划"补文档
单方声音
生态与信任
- 一句话
"Snowflake might display cryptic error messages that are not available in Snowflake documentation. You need to contact Snowflake technical support or dig into source code of connectors and drivers in GitHub repositories."——更尖锐的是后半句:"Although Snowflake is made aware of this issue, it does not have any plan on adding related documentation at this time!"(Snowflake 知道这个问题,但目前没有任何补文档的计划。)
"Snowflake might display cryptic error messages that are not available in Snowflake documentation. You need to contact Snowflake technical support or dig into source code of connectors and drivers in GitHub repositories." The sharper second half: "Although Snowflake is made aware of this issue, it does not have any plan on adding related documentation at this time!"
- 窄场景
半夜排障、搜报错码无果的 on-call;air-gapped/内网环境不能随手开 support case 的团队。
On-call engineers googling an error code at 3am; air-gapped teams that can't just open a support case.
- 机制
官方文档至今无系统性错误码手册(仅有零散 troubleshooting 页),与文中指控一致;排障路径=开工单或翻驱动源码。
Official docs still have no systematic error-code manual (only scattered troubleshooting pages) — consistent with the allegation; the debugging path is "open a ticket or read driver source."
- 生产验证
来源 14:Slim Baltagi 2023-11-08 长文第 9 节专讲文档缺失。
—
- 证据等级
`单方声音`,独立工程师长文;"官方文档无系统性错误码手册"可对照 docs.snowflake.com 现状核验,与指控一致。
`Single voice`, independent engineer's long-form piece; "no systematic error-code manual in official docs" is checkable against docs.snowflake.com today and matches the allegation.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2023
GROUP BY ALL:一个"盼了好多年"的语法糖,2023 年才补上
单方声音
已修复于 2023
生态与信任
- 一句话
2023 年 Summit 回顾里,DAS42 的顾问写道 GROUP BY ALL "a feature that we at DAS42 have anticipated for many years"(我们盼了好多年);在此之前分析师写聚合要手工数 `GROUP BY 1, 2, 3, ... 35`——列一多就数错、列一动就得重数,是个持续多年的日常摩擦;Snowflake 2023 年才把这个别家早有的便利语法补上。
In a Summit 2023 recap, a DAS42 consultant wrote that GROUP BY ALL was "a feature that we at DAS42 have anticipated for many years"; before it, analysts hand-counted `GROUP BY 1, 2, 3, ... 35` — miscount with many columns, recount when columns move — a daily friction for years; Snowflake only added this long-standard convenience syntax in 2023.
- 窄场景
写宽表聚合的分析师;2023 年之前的 Snowflake 用户。
Analysts writing wide-table aggregations; Snowflake users before 2023.
- 机制
SQL DX 的"小事"最见响应速度:一个语法糖让用户盼了好多年。
The "small things" of SQL DX reveal response speed: a syntactic sugar kept users waiting for years.
- 生产验证
来源 60:DAS42 Summit 2023 回顾,"anticipated for many years"原话出处。
—
- 证据等级
`单方声音`,第三方咨询公司回顾(原话明确)。
`Single voice`, third-party consultancy recap (verbatim explicit).
- 备注
缺口状态:已补上于 2023(Summit 2023 宣布);保留收录以记录"补上之前盼了多年"的事实。本卡主题可能与本站 [避坑] 卡重叠。
gap status: fixed in 2023 (announced at Summit 2023); kept to record the "waited years" fact. This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:2023
从 TDSQL 迁回 MySQL:存在,但没人细说
单方声音
升级迁移
- 一句话
有用户把 TDSQL 用过之后又迁回了 MySQL——原话只有一句 "TDSQL -> Mysql",动因、规模、时间线一概没说。
Some user ran TDSQL and then moved back to MySQL — the entire account is one line, "TDSQL -> Mysql", with no motive, scale, or timeline disclosed.
- 窄场景
不详(原帖为信创选型讨论下的跟帖,未披露业务背景)。
Unknown (a follow-up comment in a domestic-substitution selection thread; no business background disclosed).
- 机制
动因未披露,无法归因。注:分布式带来的运维与兼容性代价超过收益时,回迁单机 MySQL 是常见的止损路径,但这只是一般性推断,不是对该用户的断言。
Motive undisclosed, so no attribution is possible. Note: when a distributed database's operational and compatibility costs outweigh its benefits, moving back to single-node MySQL is a common damage-control path — but that is a general observation, not a claim about this user.
- 生产验证
来源 1:V2EX 用户 dddd1919 2023-05-06 在信创选型帖跟帖,原话:"用过的:达梦 -> Oracle TiDB -> Mysql TDSQL -> Mysql"。
Source 1: V2EX user dddd1919, May 6 2023, commenting on a domestic-substitution selection thread, verbatim: "used: Dameng -> Oracle, TiDB -> Mysql, TDSQL -> Mysql".
- 证据等级
`单方声音`,匿名社区跟帖(细节极少,仅作"存在从 TDSQL 迁出"方向的弱信号;未找到具名迁出复盘,如实标注缺口)。
`Single voice`, anonymous community comment (extremely thin on detail; kept only as a weak directional signal that migrations off TDSQL exist; no named migration-off postmortem was found — gap stated as-is).
腾讯云 TDSQL 年份:2023
分区表 DDL 之后,执行计划走错,只读节点 CPU 从 20% 涨到 50%+ 持续
单方声音
性能问题
- 一句话
一次分区表 DDL 让统计信息重建,优化器选了全量扫索引的坏计划——只读节点 CPU 翻倍还多,几个查询卡了一整天。
One partition-table DDL rebuilt statistics, the optimizer picked a full-index-scan plan — the read replica's CPU more than doubled, and several queries stayed stuck for a whole day.
- 窄场景
大表做分区表 DDL 变更;只读节点承担业务查询的集群。
Large tables converted to partitioned tables; clusters where read replicas serve live query traffic.
- 机制
DDL 触发统计信息重新维护后,优化器对子查询选择了非最优索引:异常时段 CPU 主要消耗在 IndexScanIterator(对 `idx_created_on_ent_id(created_on, ent_id)` 全量扫),正常时走 IndexRangeScanIterator(用 ent_id+agent_id 定范围再拿 created_on 过滤)。计划走错后查询不报错、只是慢,几个 select 在只读节点上持续了一天(单独拿出来执行都在 0.1 秒内)。
After the DDL rebuilt statistics, the optimizer chose a non-optimal index for a subquery: during the incident CPU was dominated by IndexScanIterator (full scan of `idx_created_on_ent_id(created_on, ent_id)`), while the healthy plan used IndexRangeScanIterator (range on ent_id+agent_id, filter on created_on). A wrong plan doesn't error — it just runs slow, and several SELECTs sat on the read replica for a day (each ran in under 0.1s when executed standalone).
- 生产验证
来源 1,2022-01 事故记录——1 亿多行表改成 分区表后,只读节点 CPU 从长期稳定 20% 以下升到 50% 多并持续;作者对阿里云服务群给出的"统计信息重建导致计划走错"解释表示怀疑,原文引用:"但是,一亿多行的数据也花不了一天吧",怀疑是分区表 DDL 成功瞬间的某种竞态条件。
Source 1, Jan 2022 incident record — after converting a 100M+ row table to partitioned, the read replica's CPU rose from a long-stable <20% to 50%+ and stayed; the author doubted Alibaba's support-group explanation ("statistics rebuilt, plan went wrong"), quoted: "but 100 million rows wouldn't take a whole day to scan" — suspecting a race condition at the instant the partition DDL succeeded.
- 证据等级
`单方声音`,具名事故记录(细节充分:CPU 曲线、执行计划对比、排查过程)。
`Single voice`, named incident record (detailed: CPU curves, plan comparison, investigation steps).
PolarDB 年份:2022
Uber 把金融账本搬出 DynamoDB:官方口径一年省 600 万美元
单方声音
成本账单
- 一句话
Uber 把支付账本 LedgerStore 从 DynamoDB 迁到自研 Docstore,官方复盘称每年节省约 600 万美元。
Uber migrated its LedgerStore payment ledger from DynamoDB to in-house Docstore; its official postmortem cites ~$6M/year in savings.
- 窄场景
写密集、数据量大(2500 亿条记录/~300TB)、需要强一致索引与审计可验证性的账本类负载。
Write-heavy, very large (250B records / ~300TB) ledger workloads needing strongly consistent indexes and audit verifiability.
- 机制
DynamoDB 按操作次数 + 存储量计费,热数据单价高;账本数据的典型生命周期是"写入后几周内高频读、之后转冷",而 DynamoDB 没有原生的冷热分层,长期把全量数据放在热库 = 持续支付热存储溢价;记录占其存储足迹与成本的 75%,是主要矛盾。
DynamoDB bills per operation + storage, with hot data priced at a premium; ledger data is typically read heavily for weeks after writing, then goes cold — DynamoDB has no native hot/cold tiering, so keeping everything in the hot store means paying the hot-storage premium indefinitely; records were 75% of its storage footprint and cost.
- 生产验证
来源 3,Uber 官方工程博客 2020——LedgerStore(Gulfstream 支付平台的不可变账本,带 sealing 签名审计)在 DynamoDB 上跑近两年后"becoming expensive";回填 2500 亿条唯一记录(约 300TB)零生产事故;原文 "The estimated yearly savings are $6 million per year";同时实现技术栈整合、减少外部依赖。
Source 3, Uber official engineering blog, 2020 — LedgerStore (the immutable, sealing-signed ledger behind the Gulfstream payments platform) ran on DynamoDB for almost 2 years before "becoming expensive"; backfilled 250 billion unique records (~300TB) with zero production incidents; quote: "The estimated yearly savings are $6 million per year"; plus technology consolidation and fewer external dependencies.
- 证据等级
`单方声音`,具名公司官方工程复盘(Uber 自撰,非厂商邀请稿)。
`Single voice`, named company's official engineering postmortem (written by Uber itself, not a vendor-invited piece).
- 备注
该复盘发表于 2020 年;按请求计费的成本结构为架构性设计、至今未变,故保留收录。与本站 Cognito Forms 弃选案例角度不同(彼为选型阶段弃用,此为生产运行后的迁出 + 量化节省),不重复。
published 2020; the per-request pricing structure is architectural and unchanged, so it is retained. Different angle from the site's Cognito Forms case (evaluation-stage rejection vs. post-production migration with quantified savings).
Amazon DynamoDB 年份:2020
defrag:你必须定期执行的"停机式"维护
单方声音
运维复杂度
- 一句话
compaction 不缩文件,defrag 才缩——但 defrag 是阻塞操作,执行期间该成员**读写全停**,DB 越大停越久,连 HA 集群都救不了你,只能逐个成员错峰做。
Compaction doesn't shrink the file, defrag does — but defrag is a blocking operation: while it runs, that member serves **neither reads nor writes**, longer for bigger DBs, and HA doesn't save you — every member must take its turn, staggered.
- 窄场景
DB 逼近配额上限、长期只 compact 从不 defrag 的集群;defrag 跑在 leader 上且耗时超过选主超时时会顺带触发一次 leader 选举。
Clusters near the quota ceiling that compact but never defrag; defrag running on the leader and exceeding the election timeout also triggers a leader election on the side.
- 机制
defrag 重建整个 bbolt 文件以消除内部空洞,重建期间成员不响应任何读写;耗时随 DB 大小线性增长(AWS EKS 实测约每 GB 10 秒);`etcdctl defrag --cluster` 也只是逐个成员串行,不是并行无损。Gojek 的原话是"there is always going to be an unavoidable pause"——HA 架构对此无能为力,因为每个成员都必须单独经历一次。
Defrag rebuilds the entire bbolt file to eliminate internal free-space holes; during the rebuild the member answers nothing; duration grows linearly with DB size (~10 seconds per GB measured by AWS EKS); `etcdctl defrag --cluster` still just walks members serially, not a lossless parallel op. Gojek's words: "there is always going to be an unavoidable pause" — HA is no help because each member must individually endure it.
- 生产验证
来源 4,2020-04,Gojek 生产 etcd 运维实录——"Defragmentation to a live member blocks the system from reading and writing data while rebuilding its state",因此 defrag 必须低频执行、DB 保持小巧(默认 2GB 别贪大);该阻塞行为在现行版本官方维护文档中依然存在,未被修复或优化掉。
Source 4, Apr 2020, Gojek production etcd operations notes — "Defragmentation to a live member blocks the system from reading and writing data while rebuilding its state," so defrag must run infrequently and the DB kept small (don't get greedy past the 2GB default); the blocking behavior still exists in current-version official maintenance docs — not fixed or optimized away.
- 证据等级
`单方声音`,公司工程博客生产运维实录(具名作者;阻塞机制经现行官方文档确认仍存在,故保留收录,未作"已修复"标注)。
`Single voice`, company engineering blog production operations record (named author; the blocking mechanism is confirmed to still exist in current official docs, so it is retained without a "fixed in" label).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site on this site.
etcd 年份:2020
Streams 免费变 GoldenGate 天价:用了十年的免费复制功能被砍
单方声音
成本账单生态与信任
- 一句话
随 EE 免费的 Streams/Advanced Replication 在 12c 被弃用、19c 起不受支持,替代品 GoldenGate 是独立许可、天价——客户原话:"Oracle 拿走了一个免费产品,换成了一个贵得离谱的!"
Streams/Advanced Replication, free with EE, was deprecated in 12c and desupported from 19c; the replacement GoldenGate is separately licensed and eye-wateringly expensive — in a customer's words: "Oracle took away a 'free' product and replaced it with a very expensive one!!!"
- 窄场景
用 Streams/Advanced Replication 做复制/数据同步的老 EE 客户;升级到 19c 时被迫面对。
Long-time EE customers using Streams/Advanced Replication for replication/sync; forced to face it when upgrading to 19c.
- 机制
Streams 曾是 EE 自带("免费")的复制方案;12c 起 deprecated,19c 起 desupported;官方迁移路径是 GoldenGate(独立产品、独立许可)或 XStream API(需自行开发)。功能没有消失,但免费到天价的落差全部由客户承担。
Streams shipped with EE ("free"); deprecated in 12c, desupported in 19c; the official migration path is GoldenGate (separate product, separate license) or the XStream API (build it yourself). The functionality did not disappear — but the entire free-to-sky-high delta lands on the customer.
- 生产验证
来源 8:Experts Exchange 问答(约 2020 年,围绕 19c 许可)——客户 slightwv 原话:"Oracle took away a 'free' product and replaced it with a very expensive one!!!";另一位客户 Alex 跟帖:"GoldenGate is horrobly expensive"(拼写为原文);讨论串确认 RAC 亦是 EE 之外的额外付费选项。
Source 8: Experts Exchange Q&A (circa 2020, around 19c licensing) — customer slightwv verbatim: "Oracle took away a 'free' product and replaced it with a very expensive one!!!"; fellow customer Alex: "GoldenGate is horrobly expensive" (spelling as in the original); the thread also confirms RAC is an extra paid option on top of EE.
- 证据等级
`单方声音`,同一问答串内多名客户呼应(细节充分:产品名、版本线、价格体感)。
`Single voice`, multiple customers echoing within one Q&A thread (solid details: product names, version lines, price sentiment).
- 备注
问答发生在 2020 年前后,但 Streams 在 19c 起 desupported 的机制未变,故保留。
the Q&A dates to circa 2020, but Streams being desupported from 19c is unchanged, hence retained.
Oracle Database(甲骨文) 年份:2020
"我们得教育 Oracle 我们的合同":审计按现行政策、不按签约条款算
单方声音
成本账单生态与信任
- 一句话
JM Huber 前 CIO 作证:2019 年审计时,Oracle 按 2019 年的许可政策、而非 1999 年签约时的合同条款衡量合规,客户反过来给 Oracle 上课。
JM Huber's former CIO testified: during a 2019 audit, Oracle measured compliance against 2019 licensing policies rather than the contract terms signed in 1999 — the customer ended up teaching Oracle how to read its own contract.
- 窄场景
与 Oracle 签约历史悠久、合同条款与现行政策存在代际差异的老客户。
Long-tenured Oracle customers whose contract terms predate several generations of policy revisions.
- 机制
Oracle 许可政策随版本与年份不断修订(虚拟化认定、选项包定义、云折算系数等),而审计团队倾向于用"当前政策"解读历史合同;客户若无专人逐条核对签约文本,就可能为按新政策算出的"缺口"买单。
Oracle licensing policies are revised constantly (virtualization recognition, option-pack definitions, cloud conversion factors), while audit teams tend to interpret historical contracts through "current policy"; without someone on the customer side checking every clause of the signed text, customers can end up paying for "gaps" computed under the new rules.
- 生产验证
来源 1:Michael Cahoon(时任 JM Huber 公司 CIO)2019 年亲历审计——合同可追溯至 1999 年,审计方却按 2019 年条款执行;Cahoon 原话:"It felt like we were having to educate Oracle on our contract"(感觉是我们在给 Oracle 上课,教他们读我们自己的合同)。Palisade Compliance 在报道中称 Oracle 审计年收入约 30 亿美元。
Source 1: Michael Cahoon (then CIO of JM Huber) on his 2019 audit — the contract dated back to 1999, but auditors applied 2019 terms; his words: "It felt like we were having to educate Oracle on our contract." Palisade Compliance claims in the report that Oracle audits generate ~$3B/year.
- 证据等级
`单方声音`,具名 CIO 经独立媒体报道(细节充分:年份、公司、合同年限)。
`Single voice`, named CIO via independent press (solid details: year, company, contract vintage).
- 备注
主题可能与现有 [避坑] 卡(许可证审计类)重叠。
topic may overlap with existing [Pitfall] cards (license-audit theme).
Oracle Database(甲骨文) 年份:2019
RDS Proxy 的隐藏保底:8 ACU 起步,比数据库本身还贵
单方声音
成本账单
- 一句话
给 Serverless v2 配 RDS Proxy 做连接池,Proxy 的最低收费是 8 ACU——0.5 ACU 的库,配了个 16 倍大的"门卫"。
Put RDS Proxy in front of Serverless v2 for connection pooling and the proxy's minimum charge is 8 ACUs — a 0.5-ACU database with a 16x "doorman."
- 窄场景
Serverless v2 + RDS Proxy(常见于 Lambda 架构)的组合;东京区域 0.5 ACU 常驻的小库。
Serverless v2 + RDS Proxy combos (typical for Lambda architectures); small 0.5-ACU-always-on databases in the Tokyo region.
- 机制
RDS Proxy 对 Serverless v2 按 ACU 计费且有 8 ACU 最低消费;口径藏在定价页细则里,控制台切换实例类型时不会提醒你 Proxy 的计费口径变了(从按 vCPU 小时变为按 ACU)。
RDS Proxy for Serverless v2 bills in ACUs with an 8-ACU minimum, buried in pricing-page fine print; switching instance types in the console never warns you the proxy's billing unit changed (from per-vCPU-hour to per-ACU).
- 生产验证
来源 9,dev.to 独立实践者——开发环境从 db.t3.medium 切到 Serverless v2(0.5 ACU 常驻约 $2.21/天),Cost Explorer 却每天多出约 $7;按使用类型分组发现 `Proxy-ASv2-Usage` 一项约 $10/天,根因是之前 Lambda 架构留下的 RDS Proxy 没关。
Source 9, dev.to independent practitioner — dev environment moved from db.t3.medium to Serverless v2 (0.5 ACU steady at ~$2.21/day), yet Cost Explorer showed ~$7/day extra; grouping by usage type revealed `Proxy-ASv2-Usage` at ~$10/day — root cause was a leftover RDS Proxy from the previous Lambda architecture that never got turned off.
- 证据等级
`单方声音`,dev.to 个人实践记录(细节充分:Cost Explorer 分组、东京区单价、根因定位过程)。
`Single voice`, dev.to personal field record (detailed: Cost Explorer grouping, Tokyo pricing, root-cause walkthrough).
Amazon Aurora 年份:—
Aurora Limitless:分片键要进主键,跨分片还可能"分布式死锁"
单方声音
运维复杂度
- 一句话
Limitless 不是透明的水平扩展——主键必须包含分片键、序列有空洞、Read Committed 下会出现原生 PG 没有的"分布式死锁"可重试错误,一批 Aurora 功能直接不支持。
Limitless is not transparent horizontal scaling — primary keys must include the shard key, sequences have gaps, Read Committed can throw "distributed deadlock" retryable errors unknown to native Postgres, and a long list of Aurora features is unsupported.
- 窄场景
想用 Limitless 做写扩展、且带着原生 PG 心智(主键、序列、事务语义)迁过来的团队。
Teams adopting Limitless for write scale-out while carrying native-Postgres mental models (primary keys, sequences, transaction semantics).
- 机制
Limitless 是"路由器+分片"的分片集群:主键必须含分片键(否则建不了);序列由路由器预留区间实现(空洞是常态);跨分片写走 2PC,Read Committed 下会出现 `aborting transaction participating in a distributed deadlock` 这类原生 PG 没有的可重试错误;Serializable 直接不支持;官方文档另有一长串不支持的功能(RDS Proxy、Global Database、克隆、自定义 endpoint、zero-ETL 等)。
Limitless is a router-plus-shards sharded cluster: primary keys must contain the shard key or creation is rejected; sequences are implemented via router-reserved ranges (gaps are normal); cross-shard writes go through 2PC and can raise `aborting transaction participating in a distributed deadlock` — a retryable error class native Postgres never produces; Serializable is unsupported outright; the docs list 40+ unsupported features (RDS Proxy, Global Database, cloning, custom endpoints, zero-ETL, and more).
- 生产验证
来源 16,AWS Hero Franck Pachot(独立数据库顾问,非 AWS 员工)的 Limitless 实测系列——pgbench 跨分片转账测出分布式死锁错误,只能先锁表绕过;`ALTER TABLE ... ADD PRIMARY KEY(id)` 因主键不含分片键被拒;序列 chunk 机制导致取值不连续。
Source 16, AWS Hero Franck Pachot (independent database consultant, not an AWS employee) and his Limitless hands-on series — pgbench cross-shard transfers produced the distributed-deadlock error, worked around only by locking the table first; `ALTER TABLE ... ADD PRIMARY KEY(id)` was rejected because the key lacked the shard key; sequence chunking produced discontinuous values.
- 证据等级
`单方声音`,独立实测系列(同一作者多篇,细节充分:完整报错原文、绕过方法、版本行为)。
`Single voice`, independent hands-on series (multiple posts by one author, detailed: verbatim errors, workarounds, version behavior).
- 备注
作者为 AWS Hero(社区身份),非 AWS 员工;内容为第一手实测,非厂商软文。
the author is an AWS Hero (community title), not an AWS employee; the content is first-hand testing, not vendor collateral.
Amazon Aurora 年份:—
升级:读源码、等一个月、混版本 CI 才敢升
单方声音
升级迁移
- 一句话
ClickHouse 每月发版、行为变化快,升级前要 diff 源码里的 `*Settings.h`、跑混版本集群 CI,大版本数据格式不兼容真实发生过。
Monthly releases and fast-moving behavior mean upgrades require diffing `*Settings.h` in the source, running mixed-version cluster CI, and real data-format incompatibilities have happened.
- 窄场景
自建多副本集群的版本升级;用了实验性功能、依赖特定 SQL 行为或默认参数的团队。
Version upgrades of self-hosted multi-replica clusters; teams relying on experimental features, specific SQL behaviors, or default settings.
- 机制
发版节奏快带来四类升级风险:数据存储格式不兼容变更(降级即丢数据)、bugfix 顺带改 SQL 行为(查询结果变化)、性能回退、默认参数/开关变化。复制协议向后兼容,但数据格式与行为不保证。
Fast release cadence brings four upgrade risks: data-storage-format incompatible changes (downgrade = data loss), bug fixes that change SQL behavior (query results change), performance regressions, and changed defaults/flags. The replication protocol is backward compatible; data formats and behaviors are not guaranteed.
- 生产验证
来源 3:Tinybird CTO 具名复盘——第一次升级花了 3 小时 + 2 周准备,用了 4 年才做到 CI/CD 无停机升级;"过去 2 年里遇到过 2-3 次数据格式不兼容"(多租户下各种类型组合全覆盖才撞上),"别在发布后立刻升级,至少等一个月"、"别用实验性功能";CI 跑生产版 + master + 目标版三版本、混版本集群、每天用下个版本重跑客户全部查询;"想有效运维 ClickHouse,你得读源码——我管过几百个 Postgres 集群,从没需要过这个"。
Source 3: Tinybird's CTO, named account — their first upgrade took 3 hours plus 2 weeks of preparation; it took 4 years to reach CI/CD zero-downtime upgrades; "2-3 data storage format incompatible changes in the last 2 years" (only hit because their multi-tenant base covers every type combination); "do not upgrade anything right away after the release; wait at least a month," "don't use any experimental feature"; CI runs prod version + master + target version, mixed-version clusters, and re-runs all customer queries against the next version daily; "if you want to operate ClickHouse effectively, you need to read the source code — this wasn't necessary with other databases I worked with (I was CTO of a company managing hundreds of Postgres clusters)."
- 证据等级
`单方声音`,具名 CTO 深度运维复盘(细节充分:升级流程 8 步、CI 矩阵、格式不兼容次数)。
`Single voice`, named CTO deep operations account (detailed: 8-step upgrade flow, CI matrix, incompatibility count).
- 备注
与现有吐槽清单"升级"行(默认值/SQL 行为/元数据变化影响降级)主题重叠,互为佐证。
overlaps the existing rant-list row on upgrades (defaults/SQL behavior/metadata changes affecting downgrade) — kept as corroboration.
ClickHouse 年份:—
写入后读不到:read-your-own-writes 默认不保证
单方声音
性能问题
- 一句话
刚写入的行,后续 SELECT 可能查不到——ClickHouse 默认不保证读己写一致性,打开 `select_sequential_consistency` 要拿吞吐换。
Rows you just wrote may not show up in the next SELECT — ClickHouse does not guarantee read-your-own-writes consistency out of the box, and `select_sequential_consistency` costs throughput.
- 窄场景
写入后立即查询的交互流程(表单提交后刷新、流水线"写完即查"断言)、对一致性有直觉预期的 OLTP 迁移团队。
Write-then-immediately-read flows (refresh after form submit, "write then assert" pipeline checks), and teams migrating from OLTP with intuitive consistency expectations.
- 机制
写入落盘与查询可见性之间是异步链路(part 落盘、副本同步都是后台的);`select_sequential_consistency` 能强制保证,但"at a measurable performance cost"。这是反复出现的主题:更强的保证都有,只是都要拿吞吐/延迟换。
The path from write landing to query visibility is async (part persistence and replica sync both happen in the background); `select_sequential_consistency` enforces the guarantee "at a measurable performance cost." A recurring theme: stronger guarantees exist, but each one trades throughput/latency.
- 生产验证
来源 4:TechWolf 工程博客——"ClickHouse does not guarantee read-your-own-writes consistency out of the box. The records you insert may not immediately be visible in subsequent SELECT queries",并把"想要关系型数据库的行为?可以,但学习曲线陡峭、处处是 tradeoff"列为采用 ClickHouse 的最大挑战之一。
Source 4: TechWolf engineering blog — "ClickHouse does not guarantee read-your-own-writes consistency out of the box. The records you insert may not immediately be visible in subsequent SELECT queries," listing "want relational-database behavior? possible, but steep learning curve and trade-offs everywhere" as a top adoption challenge.
- 证据等级
`单方声音`,具名公司工程博客(细节充分:参数名、性能代价说明)。
`Single voice`, named company engineering blog (detailed: setting name, performance-cost statement).
ClickHouse 年份:—
物化视图:回填没有安全路,POPULATE 有竞态
单方声音
运维复杂度
- 一句话
给存量表加物化视图想回填历史数据?官方 `POPULATE` 有并发竞态会丢数/重数,安全做法是全套手工作业。
Adding a materialized view to an existing table and backfilling history? The official `POPULATE` has concurrency races that lose or duplicate data — the safe route is an entirely manual job.
- 窄场景
给已有大表新增物化视图并回填历史;MV 越加越多、写入越来越慢的集群。
Adding a new materialized view to a large existing table and backfilling history; clusters where MVs keep piling up and writes keep slowing.
- 机制
MV 本质是"触发器 + 另一张 MergeTree 表",每次写入主表都会同步执行 MV 的聚合写入——MV 越多写入越慢、part 越多、内存压力越大。`POPULATE` 在视图创建期间到达的写入不会被正确处理(官方文档明确不推荐);安全的回填是:先建 Null 表过渡、手动分批回填、控制每批行数避免产生数千个 part。
An MV is essentially "trigger + another MergeTree table" — every main-table write synchronously executes the MV's aggregation write, so more MVs mean slower ingestion, more parts, more memory pressure. `POPULATE` mishandles writes arriving during view creation (official docs explicitly discourage it); safe backfill means staging through a Null table, manually chunking the backfill, and sizing each chunk to avoid spawning thousands of parts.
- 生产验证
来源 2:Tinybird——"MVs are a killer feature, but they are hard to manage, they lead to memory issues, and they generate a lot of parts…every new MV will make your ingestion slower";回填章节标题即"Backfills…it's so painful","You might be tempted to use POPULATE…Don't. It's broken because data can be duplicated",并给出 Null 表 + 手动分批的标准作业流程。
Source 2: Tinybird — "MVs are a killer feature, but they are hard to manage, they lead to memory issues, and they generate a lot of parts…every new MV will make your ingestion slower"; the backfill chapter opens with "it's so painful"; "You might be tempted to use POPULATE…Don't. It's broken because data can be duplicated," followed by the standard Null-table + manual-chunking runbook.
- 证据等级
`单方声音`,具名 CTO 运维复盘(细节充分:POPULATE 竞态、回填作业步骤)。
`Single voice`, named CTO operations account (detailed: POPULATE race, backfill runbook steps).
ClickHouse 年份:—
开源与商业的信任裂缝:零拷贝复制"没人爱"
单方声音
生态与信任
- 一句话
社区贡献的零拷贝复制(S3 存算分离的关键特性)buggy、会丢数据、还在 S3 留垃圾——而 ClickHouse Inc 自己的 SharedMergeTree 只给 Cloud 用,还一度计划删掉社区版特性。
The community-contributed zero-copy replication — the key feature for S3 compute-storage separation — is buggy, can lose data, and leaves garbage in S3, while ClickHouse Inc's own SharedMergeTree stays Cloud-only and the company once planned to delete the community feature.
- 窄场景
想在开源版上做 S3 存算分离降成本的团队;评估"开源版会不会被釜底抽薪"的选型者。
Teams wanting S3 compute-storage separation on the open-source build to cut costs; anyone evaluating whether the open-source edition could be hollowed out.
- 机制
开源版 S3 零拷贝复制由外部贡献者实现,ClickHouse Inc 态度冷淡(buggy、丢数据风险、S3 垃圾回收问题);ClickHouse Cloud 用的 SharedMergeTree 存储引擎不开源。开源的是代码,但最好的存储架构只存在于商业版。
Open-source S3 zero-copy replication was implemented by an outside contributor and gets a cold shoulder from ClickHouse Inc (buggy, data-loss risk, S3 garbage-collection problems); ClickHouse Cloud's SharedMergeTree storage engine is not open source. The code is open, but the best storage architecture lives only in the commercial edition.
- 生产验证
来源 3:Tinybird CTO 具名陈述——零拷贝复制"was contributed by someone outside ClickHouse, Inc., and it looks like they don't like it…it's buggy, you can lose data, it leaves garbage in S3";"ClickHouse Cloud has its own storage, but it's not open source";"ClickHouse, Inc. could remove the feature any time (they had actually planned to do it)";Tinybird 为此维护了私有 fork。同文开头的律师免责声明("mainly for ClickHouse, Inc lawyers: we have nothing to do with ClickHouse, Inc")本身也是商标关系紧张的信号。
Source 3: Tinybird's CTO, named account — zero-copy replication "was contributed by someone outside ClickHouse, Inc., and it looks like they don't like it…it's buggy, you can lose data, it leaves garbage in S3"; "ClickHouse Cloud has its own storage, but it's not open source"; "ClickHouse, Inc. could remove the feature any time (they had actually planned to do it)"; Tinybird maintains a private fork for it. The piece's opening lawyer disclaimer ("mainly for ClickHouse, Inc lawyers: we have nothing to do with ClickHouse, Inc") is itself a signal of trademark tension.
- 证据等级
`单方声音`,具名 CTO 陈述(细节充分:特性归属、删除计划、私有 fork)。
`Single voice`, named CTO account (detailed: feature provenance, removal plan, private fork).
ClickHouse 年份:—
列级脱敏(masking)仅 Cloud 可用:自建版执行 DDL 直接被拒
单方声音
运维复杂度生态与信任
- 一句话
Snowflake 有 Dynamic Data Masking;ClickHouse 的 `MASKING POLICY` 是 ClickHouse Cloud 专属(25.12+),开源/自建执行 DDL 直接被拒(`SUPPORT_IS_DISABLED`)——合规敏感的自建用户只能手写 view 变通。
Snowflake has Dynamic Data Masking; ClickHouse's `MASKING POLICY` is ClickHouse Cloud–exclusive (25.12+) — the open-source/self-hosted edition rejects the DDL with `SUPPORT_IS_DISABLED`. Compliance-sensitive self-hosters hand-roll views.
- 窄场景
自建 ClickHouse + 数据合规/脱敏需求。
Self-hosted ClickHouse with data-compliance/masking requirements.
- 机制
行级 `ROW POLICY`(OSS 可用)与列级 `MASKING POLICY`(Cloud 专属)被切分到两个版本线。自建版教脱敏的课程只能教"建 `v_employees_masked` view + `replaceRegexpAll` 手工脱敏"。
Row-level `ROW POLICY` (OSS-available) and column-level `MASKING POLICY` (Cloud-exclusive) are split across edition lines. Courses teaching masking on self-hosted Docker can only teach "build a `v_employees_masked` view + hand-rolled `replaceRegexpAll`."
- 生产验证
—
Source 32: ClickHouse official Terraform provider docs — "masking policies are only available on ClickHouse Cloud (version 25.12+)"; the open-source edition rejects the DDL.
- 证据等级
`单方声音`,Cloud-only 事实确凿但无具名用户抱怨。
`Single voice`, Cloud-only fact confirmed, no named user complaint found.
ClickHouse 年份:—
大版本之间没有升级通道:1.x→2.x、3.x→4.x 按"迁移"处理,1.x 迁移工具还要付费
单方声音
升级迁移成本账单
- 一句话
OceanBase 的大版本之间不支持原地升级,跨大版本等于重新做一次数据迁移;早期版本的官方迁移工具还是付费的。
OceanBase offers no in-place upgrade across major versions — moving across a major boundary means redoing a full data migration; the early official migration tooling was commercial and paid.
- 窄场景
运行 1.x/2.x/3.x 老版本、希望跟进 4.x 新特性的团队。
Teams running 1.x/2.x/3.x and wanting to follow 4.x features.
- 机制
OceanBase 内核在 1.x→2.x、3.x→4.x 之间存在不兼容的存储/元数据变更,官方升级路径只覆盖小版本与相邻版本序列内的"经停"升级(如 4.2.5→4.3.5→4.4.x 需按升级依赖序列逐段经停);跨代只能走 OMS/逻辑迁移重建集群。早期 1.x 时代的官方迁移工具是商业收费的,而同期开源迁移工具普遍免费或随服务提供。
Incompatible storage/metadata changes sit between 1.x to 2.x and 3.x to 4.x, so the official upgrade path only covers patch versions and staged hops within a release line (e.g. 4.2.5 to 4.3.5 to 4.4.x must stop at each intermediate per the upgrade dependency sequence); crossing a generation means rebuilding the cluster via OMS/logical migration. The official 1.x-era migration tools were commercially licensed, while comparable open-source migration tooling of the time was free or bundled with the service.
- 生产验证
来源 1:Gartner Peer Insights 验证用户原话——"their major version is sometimes not upgradable, e.g. 1.x to 2.x, 3.x to 4.x are considered migrations not upgrades. Data migration tools for version 1.x are paid tools, while out there data migration tools are free or as part of the services"。
Source 1: Gartner Peer Insights validated user, verbatim — "their major version is sometimes not upgradable, e.g. 1.x to 2.x, 3.x to 4.x are considered migrations not upgrades. Data migration tools for version 1.x are paid tools, while out there data migration tools are free or as part of the services."
- 证据等级
`单方声音`,Gartner Peer Insights 验证用户评价(匿名,平台验证身份)。
`Single voice`, Gartner Peer Insights validated user review (anonymous, platform-verified identity).
OceanBase 年份:—
CLOG 能吃掉近一半数据盘:高频写入场景的日志税
单方声音
成本账单
- 一句话
高频写入下 CLOG 日志空间占比极高,接近数据盘容量的一半,磁盘规划不留足就等着扩容。
Under high-frequency writes, CLOG log space takes a very large share — close to half of data-disk capacity — so disk planning without headroom ends in an expansion project.
- 窄场景
高频写入 OLTP;按数据量 1:1 规划磁盘、没给日志留余量的部署。
Write-heavy OLTP; deployments sized 1:1 on data volume with no log headroom.
- 机制
OceanBase 是 Paxos 多副本共识写入,每笔事务的 Redo(CLOG)要在多数派落盘才算提交;为保证 RPO=0 与故障恢复,CLOG 保留窗口通常远大于单机库的 WAL/binlog 保留。高频写入时日志产生速度与数据写入速度同量级,磁盘占用里日志与数据几乎对半。
OceanBase is a Paxos multi-replica consensus writer — every transaction's redo (CLOG) must land on a quorum before commit; the CLOG retention window is typically far longer than a single-node database's WAL/binlog retention to guarantee RPO=0 and crash recovery. Under high-frequency writes, log generation runs at the same order of magnitude as data writes, so disk splits roughly half-and-half between log and data.
- 生产验证
来源 1:Gartner Peer Insights 验证用户原话——"CLOG Log: In high-frequency write scenarios, the CLOG log space accounts for a high proportion, nearly half of the data disk capacity."
Source 1: Gartner Peer Insights validated user, verbatim — "CLOG Log: In high-frequency write scenarios, the CLOG log space accounts for a high proportion, nearly half of the data disk capacity."
- 证据等级
`单方声音`,Gartner Peer Insights 验证用户评价(匿名)。
`Single voice`, Gartner Peer Insights validated user review (anonymous).
OceanBase 年份:—
老版本改个配置要做节点迁移:弹性这事,新版本才有
单方声音
运维复杂度
- 一句话
老版本 OceanBase 调整资源配置需要做节点(Unit)迁移,又慢又重;在线改配是新版本才补上的能力。
On older OceanBase versions, resizing resources requires a Unit/node migration — slow and heavy; online resizing only arrived in newer versions.
- 窄场景
仍在跑 2.x/3.x 老版本的集群;需要频繁调整租户 CPU/内存规格的业务。
Clusters still on 2.x/3.x; businesses that resize tenant CPU/memory often.
- 机制
早期版本租户资源与 Unit/资源池强绑定,改规格要走"新建 Unit → 分区复制 → 切换"的重分布流程,本质是一次小规模数据迁移;4.x 才把"改 Unit Config 即时生效"做成在线操作。版本越老,弹性越接近"纸面"。
Early versions bound tenant resources tightly to Units/resource pools, so a resize went through "create Unit, copy partitions, switch over" — essentially a small-scale data migration; only 4.x made "change the Unit config and it takes effect immediately" an online operation. The older the version, the more "elasticity" is a paper claim.
- 生产验证
来源 1:Gartner Peer Insights 验证用户原话——"Elasticity: Configuration changes in older Oceanbase versions require node migration, which is time-consuming."
Source 1: Gartner Peer Insights validated user, verbatim — "Elasticity: Configuration changes in older Oceanbase versions require node migration, which is time-consuming."
- 证据等级
`单方声音`,Gartner Peer Insights 验证用户评价(匿名)。
`Single voice`, Gartner Peer Insights validated user review (anonymous).
OceanBase 年份:—
监控粒度粗、第三方集成弱:想接自己的可观测体系得自己造轮子
单方声音
生态与信任
- 一句话
集群监控粒度不够细、第三方监控集成能力弱,ODC 的可定制功能也被点名不足,运维想深度定制只能自己动手。
Cluster monitoring granularity is too coarse and third-party monitoring integration too weak; ODC's customization was also called out, so deep customization falls on the ops team.
- 窄场景
已有 Prometheus/Grafana/自研运维平台、希望把 OB 纳入统一可观测体系的团队。
Teams with existing Prometheus/Grafana/in-house ops platforms that want OB inside one observability system.
- 机制
OB 的可观测主要围绕自家 OCP 与内部视图(GV$ 系列)建设,对外暴露的指标粒度和第三方生态对接(exporter、标准告警集成)成熟度不如 MySQL/PG 生态;ODC 作为开发者工具,可定制能力被用户点名不足。
OB's observability is built around its own OCP and internal views (the GV$ family); externally exposed metric granularity and third-party ecosystem hooks (exporters, standard alert integrations) lag the MySQL/Postgres ecosystems; ODC as a developer tool was flagged for insufficient customization.
- 生产验证
来源 1:Gartner Peer Insights 验证用户原话——"1. Cluster monitoring granularity and third-party integration capabilities 2. Suggested real-world use cases for AI models 3. Customizable features and functionalities, such as those available on the ODC (Optical Distribution Center)."
Source 1: Gartner Peer Insights validated user, verbatim — "1. Cluster monitoring granularity and third-party integration capabilities 2. Suggested real-world use cases for AI models 3. Customizable features and functionalities, such as those available on the ODC (Optical Distribution Center)."
- 证据等级
`单方声音`,Gartner Peer Insights 验证用户评价(匿名)。
`Single voice`, Gartner Peer Insights validated user review (anonymous).
OceanBase 年份:—
12c 自适应优化器:升级后性能倒退,关掉才恢复正常
单方声音
来源存疑
性能问题升级迁移
- 一句话
两家 PeopleSoft 客户从 10g/11g 升到 12.1.0.2 后出现升级前不存在的性能问题,罪魁祸首是默认开启的自适应查询优化,关掉 OPTIMIZER_ADAPTIVE_FEATURES 后"性能极佳、再无用户投诉"。
Two PeopleSoft customers upgrading from 10g/11g to 12.1.0.2 hit performance problems that did not exist before; the culprit was the default-on adaptive query optimization — after setting OPTIMIZER_ADAPTIVE_FEATURES=FALSE, "performance was excellent and we had no more end user complaints."
- 窄场景
升级到 12c(12.1/12.2)的 OLTP/ERP 类应用(PeopleSoft、E-Business Suite 等)。
OLTP/ERP-style applications (PeopleSoft, E-Business Suite) upgrading to 12c (12.1/12.2).
- 机制
Adaptive Query Optimization 允许 Oracle 在 SQL 执行过程中根据运行时收集的"自适应统计信息"调整执行计划;该机制本身有可观开销,且在某些负载下导致计划选择恶化。Oracle Support 上有大量相关性能 note(1335892.1、1320300.1、2182373.1 等)。12.2 起该参数被拆分为 OPTIMIZER_ADAPTIVE_PLANS 与 OPTIMIZER_ADAPTIVE_STATISTICS 两个参数。
Adaptive Query Optimization lets Oracle adjust execution plans mid-execution using runtime-gathered "adaptive statistics"; the mechanism itself carries significant overhead and worsens plan choices under some workloads. Oracle Support hosts numerous related performance notes (1335892.1, 1320300.1, 2182373.1, etc.). From 12.2 the parameter was split into OPTIMIZER_ADAPTIVE_PLANS and OPTIMIZER_ADAPTIVE_STATISTICS.
- 生产验证
来源 11:Solvaria 咨询公司博客——两家 PeopleSoft 客户升级后性能问题、常规统计信息收集无效、关闭自适应特性后恢复的全过程。
Source 11: Solvaria consultancy blog — two PeopleSoft clients' post-upgrade performance issues, conventional statistics collection failing to help, and full recovery after disabling the adaptive features.
- 证据等级
`单方声音`[来源存疑],咨询公司博客(未具名客户,营销文体)。
`Single voice` [Questionable source], consultancy blog (unnamed clients, marketing tone).
- 备注
内容年代较早(12c 时代机制),但自适应参数机制在后续版本中仍然存在,未标注"已修复"。主题可能与现有 [避坑] 卡(升级类)重叠。
the content is from the 12c era, but the adaptive-parameter mechanism persists in later versions — not marked "fixed". This card's topic may overlap with existing [Pitfall] cards on this site (upgrade theme).
Oracle Database(甲骨文) 年份:—
加副本提吞吐:Qdrant 在 400 QPS 测崩了
单方声音
性能问题
- 一句话
副本从 1 加到 2,Qdrant p99 有改善,但吞吐上不去——Reddit 的 400 QPS 测试因高延迟和错误根本没跑完。
Raising the replication factor from 1 to 2 improved Qdrant's p99, but throughput hit a wall — Reddit's 400 QPS run died of high latency and errors before completing.
- 窄场景
高吞吐在线查询、靠加副本扩读吞吐的集群;评估基于 v1.12。
High-throughput live query clusters scaling read throughput via replicas; evaluated on v1.12.
- 机制
加副本后查询要向更多副本扇出并合并结果,协调开销上升;同构节点上每个节点既服务查询又参与复制协调,扇出开销随副本数放大,吞吐先撞墙。RF=1 时 Qdrant 单机吞吐占优,恰恰说明它的扩展瓶颈在"多副本协调"而非单机性能。
More replicas mean each query fans out to more replicas and merges results, raising coordination overhead; on homogeneous nodes every node both serves queries and participates in replication coordination, so fan-out cost grows with replica count and throughput hits the wall first. That Qdrant won at RF=1 single-node throughput shows the bottleneck is multi-replica coordination, not single-node speed.
- 生产验证
来源 3:RF=1 时 Qdrant 在更高吞吐下仍能给出满意延迟;RF=2 后 Qdrant p99 改善,但 Milvus 能以可接受延迟维持更高吞吐——"Qdrant 400 QPS 未展示,因为测试因高延迟和错误未能完成"。
Source 3: at RF=1 Qdrant delivered satisfactory latency at higher throughput than Milvus; at RF=2 Qdrant's p99 improved, but "Milvus was able to sustain higher throughput than Qdrant was, with acceptable latency (Qdrant 400 QPS not shown because the test did not complete due to high latency and errors)."
- 证据等级
`单方声音`,具名工程评估(量化对比数据)。
`Single voice`, named engineering evaluation (quantitative comparison data).
Qdrant 年份:—
开源版加副本要手工建/删分片:没有自动再平衡
单方声音
运维复杂度
- 一句话
想提高副本数,开源版 Qdrant 得手工创建/删除分片——Reddit 团队说这功能"要么自己造,要么用非开源版"。
To increase the replication factor on open-source Qdrant you create and drop shards by hand — Reddit's team called it a feature they'd "have to build ourselves or use the non-open-source version" for.
- 窄场景
需要在线调整副本数/分片数的自运维集群;评估基于 v1.12。
Self-operated clusters that need to change replica or shard counts online; evaluated on v1.12.
- 机制
Qdrant 的分片放置与副本拓扑变更走共识操作,但开源版缺少"一键改 RF 自动搬数据"的编排能力;运维需手工发起分片转移并盯进度,步骤多、易出错。Milvus 在提高集合副本数时有自动再平衡。
Shard placement and replica-topology changes go through consensus operations, but the open-source edition lacks one-click "change RF and move the data" orchestration; operators must manually trigger shard transfers and watch them, with many steps and room for error. Milvus rebalances automatically when a collection's replication factor grows.
- 生产验证
来源 3:评估团队原话——"Milvus 在提高集合副本数时有自动再平衡,而开源版 Qdrant 需要手工创建或删除分片来提高副本数(这个功能我们得自己造,或者用非开源版本)"。
Source 3: the evaluation team's own words — "Milvus also has automatic rebalancing when increasing the replication factor of a collection, whereas in open-source Qdrant, manual creation or dropping of shards is required to increase the replication factor (a feature we would have had to build ourselves or use the non-open-source version)."
- 证据等级
`单方声音`,具名工程评估。
`Single voice`, named engineering evaluation.
- 备注
评估基于 v1.12;Qdrant 后续版本加入了 resharding 相关能力(以官方 release notes 为准),但无独立验证确认该缺口已补齐,故不作"已修复"标注。
evaluated on v1.12; Qdrant has since added resharding-related capabilities (per official release notes), but no independent verification confirms the gap is closed, so no "fixed in" label.
Qdrant 年份:—
进坏状态后,Qdrant 比 Milvus 难调试
单方声音
生态与信任
- 一句话
Reddit 团队的原话——"任一系统进入坏状态时,调试和修复 Milvus 比 Qdrant 更容易"。
Reddit's team's own words — "we had an easier time debugging and fixing Milvus than Qdrant when either solution entered a bad state."
- 窄场景
生产异常排查(节点失速、查询变慢但不报错);评估基于 v1.12。
Production incident triage (a node going slow, queries degrading without errors); evaluated on v1.12.
- 机制
Milvus 异构组件边界清晰(proxy、coordinator、query、data 节点各司其职),日志/指标按组件划分,故障域可定位;Qdrant 同构节点内 segment 优化、HNSW 构建、Raft 共识交织在同一进程里,坏状态时难判断是哪个子系统在拖慢。加上社区可借鉴的排障案例更少,on-call 更多只能靠自己啃。
Milvus's heterogeneous components have clean boundaries (proxy, coordinator, query, data nodes each with a job), so logs and metrics are partitioned per component and the fault domain is localizable; inside Qdrant's homogeneous node, segment optimization, HNSW construction, and Raft consensus interleave in one process, making it hard to tell which subsystem is dragging. Fewer community triage stories to borrow from means on-call engineers mostly debug alone.
- 生产验证
来源 3:评估团队在"运维"维度上的直接定性结论(引用原话)。
Source 3: the evaluation team's direct qualitative verdict on the "operations" dimension (quoted verbatim).
- 证据等级
`单方声音`,具名工程评估(定性结论,非量化数据)。
`Single voice`, named engineering evaluation (qualitative verdict, not quantified).
Qdrant 年份:—
半结构化数据二等公民:ARRAY 变字符串,JSON 支持残缺
单方声音
生态与信任
- 一句话
在 Redshift 里 ARRAY 列默认被转成字符串,JSON 支持"非常有限"——数据科学家直言跟其他数仓比"支持很差",半结构化分析处处碰壁。
In Redshift, ARRAY columns are silently converted to strings and JSON support is "very limited" — a data scientist calls it plain poor versus other warehouses, and semi-structured analysis hits walls everywhere.
- 窄场景
事件流、点击流等半结构化数据为主的分析;从 Snowflake/BigQuery 迁移过来的团队。
Event/clickstream-heavy semi-structured analytics; teams migrating from Snowflake/BigQuery.
- 机制
Redshift 基于 PG 8 系老内核,SUPER 类型 2021 年才引入;数组/嵌套的表达与函数覆盖远不如 Snowflake VARIANT,JDBC/工具链常把 ARRAY 退化成字符串,下游解析成本转嫁给用户。
Redshift descends from the PostgreSQL 8 lineage; the SUPER type only arrived in 2021. Expression and function coverage for arrays/nesting lags far behind Snowflake's VARIANT, and JDBC/tooling often degrades ARRAY to strings, pushing parse costs onto the user.
- 生产验证
来源 5,TrustRadius 认证评论(Manager in Engineering,页面未标注日期)——"Json support in sql is very limited... Array type columns are missing. They are by default converted to strings... For a data scientist too, the SQL is a bit limited when it comes to unstructured columns in the tables. Arrays, jsons, etc have very poor support compared to other warehouses."
Source 5, TrustRadius verified review (Manager in Engineering, no date shown) — "Json support in sql is very limited... Array type columns are missing. They are by default converted to strings... For a data scientist too, the SQL is a bit limited when it comes to unstructured columns in the tables. Arrays, jsons, etc have very poor support compared to other warehouses."
- 证据等级
`单方声音`,点评平台认证用户(细节充分:ARRAY 转字符串、JSON 支持弱、与其他数仓对比)。
`Single voice`, platform-verified reviewer (detailed: ARRAY-to-string, weak JSON, comparison with other warehouses).
Amazon Redshift 年份:—
Snowflake 侧 incident 4 小时 45 分,dbt 任务间歇性报 000603 内部错误
单方声音
稳定与故障
- 一句话
2025-03-19 15:00 UTC 起 Snowflake 一次 incident,dbt Platform 上所有 Snowflake 用户的 IDE 调用与定时任务间歇性失败,报错 `"000603 (XX000): SQL execution internal error ... incident 7507128"`,影响窗口约 4 小时 45 分。
Starting 15:00 UTC Mar 19 2025, a Snowflake incident made IDE calls and scheduled jobs fail intermittently for all Snowflake users on the dbt Platform, erroring `"000603 (XX000): SQL execution internal error ... incident 7507128"` for about 4 hours 45 minutes.
- 窄场景
2025-03-19 当天用 dbt Cloud 跑 Snowflake 任务的团队(dbt US Cell 1 AWS 的 IDE 与 Scheduled Jobs 两个组件)。
Teams running Snowflake jobs on dbt Cloud that day (dbt US Cell 1 AWS, IDE and Scheduled Jobs components).
- 机制
Snowflake 服务端 incident;下游平台只能记录、等待,无自救手段。
A Snowflake server-side incident; downstream platforms could only record and wait.
- 生产验证
来源 18:dbt Labs 官方状态页——当日 04:50 UTC(美东)开始调查、05:12/06:44 UTC 两次确认是 Snowflake 侧 incident、次日 01:48 UTC 宣布恢复。
—
- 证据等级
`单方声音`,下游平台官方状态页的时间线与错误原文记录;与本页 2024-12-16 事件同属 000603/XX000 引擎内部错误报错族。
`Single voice`, downstream platform's official status-page timeline with exact error text; same 000603/XX000 engine-internal-error family as this page's Dec 16 2024 incident.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:—
服务端一次 rollout,Sigma 客户查询错误率 100% 长达 5 小时
单方声音
稳定与故障
- 一句话
2023-10-03 Snowflake 当天早晨发布的 server 7.35 合并了登录请求的 account 解析逻辑,导致 connection string 用 legacy 格式 account identifier 的 Go/JDBC driver 登录失败——受影响客户"所有 Sigma 相关 warehouse 操作错误率 100%",直到 Snowflake 全局回滚到 7.34 才恢复,约 5 小时。
Snowflake's server 7.35 release on the morning of Oct 3 2023 merged account-resolution logic for login requests, breaking Go/JDBC driver logins that used legacy-format account identifiers — affected customers saw "100% error rate on all Sigma-related warehouse operations" until Snowflake rolled back globally to 7.34, about 5 hours.
- 窄场景
通过 Sigma(及同类 BI 工具)连接 Snowflake、用 legacy account identifier 的组织。
Organizations connecting to Snowflake via Sigma (and similar BI tools) with legacy account identifiers.
- 机制
服务端升级的向后不兼容变更 + 客户端 identifier 格式旧——两边都没错,但组合起来全挂;Sigma 的预防措施是"内部测试账号永远先于客户吃到最新 server 更新",等于承认对 Snowflake 的 rollout 节奏没有可见性。
A backwards-incompatible server upgrade combined with an old client identifier format — neither side wrong alone, everything broken together; Sigma's preventive action was "internal test accounts always eat the newest server release before customers," an admission of zero visibility into Snowflake's rollout cadence.
- 生产验证
来源 19:Sigma Computing 工程团队 postmortem——13:00 UTC 起大量组织无法执行查询、报错 `"Bad request; operation not supported"`;排查确认是 server 7.35 的 account 解析逻辑变更;约 18:03 UTC 恢复。
—
- 证据等级
`单方声音`,下游 BI 平台以客户视角撰写的事故复盘(时间线、根因、恢复动作俱全)。
`Single voice`, downstream BI platform's customer-perspective incident postmortem (timeline, root cause and recovery all documented).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:—
SCIM 离职同步可能静默失败:Okta 侧显示解绑成功,Snowflake 里账号还活着
单方声音
生态与信任
- 一句话
Okta 给 Snowflake 做用户注销时,"Snowflake app un-assignment succeeds, but SCIM deactivation fails"——管理员在 Okta 里看到的是"已解绑",离职员工的 Snowflake 账号实际还活着,访问回收出现空档,且没有任何告警。
During Okta→Snowflake user deprovisioning, "Snowflake app un-assignment succeeds, but SCIM deactivation fails" — the admin sees "deprovisioned" in Okta while the departed employee's Snowflake account is still alive: a gap in access revocation with zero alerting.
- 窄场景
Okta + Snowflake SCIM 对接的团队;有离职员工账号回收合规要求的组织。
Teams with Okta + Snowflake SCIM integration; orgs with compliance requirements on leaver account cleanup.
- 机制
Snowflake SCIM 访问令牌过期是可能原因,而令牌过期本身没有任何提前通知机制;修复要去 Snowflake 里重新生成 token、再到 Okta 更新配置。
An expired Snowflake SCIM access token is the likely cause, and token expiry has no advance notification; the fix is regenerating the token in Snowflake and updating the Okta config.
- 生产验证
来源 32:加州政府数据团队(cagov)GitHub issue #586(团队工程师 Kevin 记录的一线运维记录)——"The likely cause is that Snowflake SCIM access token expired"。
—
- 证据等级
`单方声音`,客户一线运维记录;SCIM 配置繁琐在社区有多篇讨论,但"解绑成功、注销失败"的静默失败模式只在此 issue 见到具名实录。
`Single voice`, customer frontline ops record; SCIM setup pain is widely discussed, but this silent "deprovisioned-yet-alive" mode has only this named record.
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:—
Empower 横评:Snowpark 做 AI 训练,Databricks 全方位胜出——"it is not" [来源存疑]
单方声音
来源存疑
生态与信任
- 一句话
从业者横评 Snowpark vs Databricks 做 AI 模型训练,TLDR 直接写 "**it is not**"(Snowpark 不是强力竞争者):工作区只是"SQL-only worksheet",无原生 Git 集成、多人不能协作编辑 notebook;第三方 Python 包"limited list, not all library versions supported";UDF 限 1 分钟且不能返回对象(模型训完必须在 UDF 内推理完再扔掉),存储过程限 1 小时只支持标量返回;实测不支持 OpenCV、音频库、强化学习库,"removes quite a few branches of AI from your toolkit";无 GPU("critical for big-model AI training");warehouse 内存上限撞到要找 support 逐个仓库提;真正的分布式训练在 Snowflake 执行环境内 "impossible"——"are you truly doing AI **with** Snowflake or with a Snowflake partner?"(你到底是在用 Snowflake 做 AI,还是在用 Snowflake 的合作伙伴做 AI?)
A practitioner bake-off of Snowpark vs Databricks for AI model training, TL;DR: "**it is not**" (Snowpark is not a strong competitor): the workspace is a "SQL-only worksheet" with no native Git and no collaborative notebook editing; third-party Python packages are a "limited list, not all library versions supported"; single-node training offers only UDFs (1-minute cap, can't return objects — train the model, run inference, throw it away inside the UDF) or stored procedures (1-hour cap, scalars only); OpenCV, audio libs and RL libraries unsupported in testing — "removes quite a few branches of AI from your toolkit"; no GPU ("critical for big-model AI training"); warehouse memory caps require per-warehouse support tickets to raise; true distributed training "impossible" inside Snowflake's execution environment — "are you truly doing AI **with** Snowflake or with a Snowflake partner?"
- 窄场景
评估"Snowflake All-in-One AI"的团队;需要 GPU/分布式训练的 ML 负载。
Teams evaluating "all-in-one AI on Snowflake"; ML workloads needing GPUs or distributed training.
- 机制
Snowpark 的执行环境为 SQL 数仓设计,不是为 ML 训练设计;HN 社区同期争论中 Snowpark 被批 "walled garden… only 'approved' python libraries are allowed",Snowflake 员工反驳可用 IMPORTS 上传纯 Python 包,随即被怼 "Good luck with trying to install any non-trivial python library this way"。
Snowpark's execution environment was designed for a SQL warehouse, not ML training; a parallel HN debate called Snowpark a "walled garden… only 'approved' python libraries are allowed" — a Snowflake employee countered that IMPORTS allows arbitrary pure-Python packages, and was answered with "Good luck with trying to install any non-trivial python library this way."
- 生产验证
来源 49:Empower 公司博客([来源存疑]:数据平台咨询商,有立场倾向,结尾导向 Databricks Lakehouse 方案)——所列技术事实具体可核,且与
来源 47 在四点上交叉印证。
—
- 证据等级
`单方声音`,[来源存疑]:作者立场倾向明显,但技术事实具体、与独立实测交叉印证。
`Single voice`, [source questionable]: the author's slant is visible, but the technical facts are concrete and cross-confirmed with an independent hands-on review.
- 备注
约 2022 年 Snowpark 早期横评,Snowpark 此后持续迭代。本卡主题可能与本站 [避坑] 卡重叠。
an early-2022-era Snowpark bake-off; Snowpark has iterated since. This card's topic may overlap an existing [Pitfall] card on this site.
Snowflake 年份:—
主键表写入先把主键索引"吃"进内存:没分区的表直接写挂
单方声音
性能问题运维复杂度
- 一句话
主键模型写入时要把目标分区的主键索引加载进 BE 内存——表没做分区,就等于把整张表的主键索引一次性塞进内存,Update 内存一超,导入直接报错。
Primary Key tables load the target partition's primary index into BE memory on write — with no partitioning, that means stuffing the entire table's primary index into memory at once; once the update-memory budget is blown, loads fail outright.
- 窄场景
主键模型(Primary Key)表、未按时间分区(或无冷热特征)的表、高频 Flink/Stream Load 写入;BE 的 update 内存默认 27GB 上限。
Primary Key model tables without time partitioning (or without hot/cold data characteristics), under frequent Flink/Stream Load writes; BE update memory defaults to a 27GB cap.
- 机制
StarRocks 主键表为支持行级实时更新,在写入时把目标分区的主键索引(persistent index)加载到 BE 内存做去重/更新定位,主键索引的加载粒度是分区。没有分区键的表,最小加载单元就是整张表。微盟案例中一张设计不合理的主键表在写入时把 BE 的 update 内存(默认 27GB)打爆,导入报 `close index channel failed`。另有一重硬约束:主键编码后总长度上限 127 字节且不可调,主键列多/宽的表建表即埋雷,写入时才暴露。
To support row-level real-time updates, StarRocks Primary Key tables load the target partition's primary key index (persistent index) into BE memory for dedup/update positioning, at partition granularity. A table with no partition key has the whole table as its minimum load unit. In Weimob's case, a poorly designed PK table blew through the BE's 27GB update-memory budget during writes and loads failed with `close index channel failed`. A second hard constraint: the encoded primary key is capped at 127 bytes total and the cap is not tunable — wide multi-column PKs are landmines that only detonate at write time.
- 生产验证
来源 1:微盟技术中心生产案例一——多个主键表导入报错 `close index channel failed`,排查发现某主键表无分区,写入时整表主键索引进内存,BE update 内存(27GB)超限;最终把该表从主键模型改回更新模型解决;
来源 1:同团队案例三——50 张主键表的集群,8 台 BE 平均每台被主键索引常驻吃掉 7.5GB 内存(BE 总共只给了 50GB),治理(改更新模型/加分区键/清生命周期)后释放约 48GB 常驻内存;
来源 1:同团队案例二——7 列宽主键(多 VARCHAR)建表后写入失败,主键编码超 127 字节硬上限,去掉一列主键后才写入成功。
Source 1: Weimob tech center, case 1 — multiple PK tables failing loads with `close index channel failed`; root cause was a PK table with no partitions loading its whole primary index into memory and exceeding the 27GB BE update budget; fixed by switching the table back to the Update model;
Source 1: same team, case 3 — a cluster with ~50 PK tables had 7.5GB per BE (of 50GB allocated) permanently resident in primary indexes across 8 BEs; governance (switching to Update model, adding partition keys, lifecycle cleanup) freed ~48GB of resident memory;
Source 1: same team, case 2 — a 7-column wide PK (multiple VARCHARs) failed writes because the encoded key exceeded the 127-byte hard limit; writes succeeded only after dropping one PK column.
- 证据等级
`单方声音`,独立公司工程团队博客(同一来源内 3 个具名生产案例,细节充分:报错原文、内存数值、治理前后对比)。
`Single voice`, independent company engineering blog (3 detailed production cases in one source: verbatim errors, memory figures, before/after governance numbers).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site on this site.
StarRocks 年份:—
Tablet 数量失控:3T 数据、140 万个 tablet,FE 内存先顶不住
单方声音
运维复杂度
- 一句话
StarRocks 的 tablet 数 = 分区数 × 分桶数 × 副本数,随手建表就能造出百万级 tablet——每个 tablet 都在 FE 常驻约 5KB 元数据,FE 内存先被"小文件式"的元数据淹没。
StarRocks tablet count = partitions x buckets x replicas, so careless table design mints millions of tablets — each one resident in FE metadata at ~5KB, and the FE drowns in metadata long before the data gets big.
- 窄场景
分区粒度过细(天级分区配小数据量)、分桶数拍脑袋设大、历史表从未治理的集群;分桶数建表后不可修改。
Overly fine partitioning (daily partitions on small data), oversized bucket counts, never-governed legacy tables; bucket count cannot be changed after table creation.
- 机制
Tablet 是 StarRocks 最小的数据管理单元,FE 要为每个 tablet 维护元数据(微盟压测:约 5KB/tablet,随总量线性增长)。微盟集群 3T 数据攒出 140 万个 tablet,平均每个 tablet 仅 2.6MB(官方推荐 100MB–1GB),FE 内存峰值被推到 18.6GB。治理(三批:天改月分区、砍分桶数)后 tablet 降到 13 万,FE 内存峰值回落到 4.5GB。雪上加霜的是分桶数建表后不能直接改,前期设错只能重建表。
Tablets are StarRocks's smallest data-management unit and the FE keeps metadata for every tablet (Weimob measured ~5KB/tablet, growing linearly with total count). Weimob's cluster accumulated 1.4M tablets over just 3TB of data — 2.6MB per tablet on average (official guidance: 100MB–1GB) — pushing FE memory to an 18.6GB peak. Governance (three rounds: daily→monthly partitions, cutting bucket counts) brought tablets down to 130K and FE peak memory back to 4.5GB. Worse, bucket count is immutable after creation — a bad early guess can only be repaid by rebuilding the table.
- 生产验证
来源 1:微盟技术中心生产案例四——Tablet 数 120 万+(优化前 140 万),3TB 数据平均 tablet 2.6MB,大量 tablet 仅几 KB;治理后降至 13 万,FE 内存峰值 18.6GB → 4.5GB。
Source 1: Weimob tech center, case 4 — 1.2M+ tablets (1.4M before governance) over 3TB, average tablet 2.6MB with many at just a few KB; after governance 130K tablets, FE memory peak 18.6GB → 4.5GB.
- 证据等级
`单方声音`,独立公司工程团队博客(细节充分:tablet 数、平均大小、FE 内存治理前后数值)。
`Single voice`, independent company engineering blog (detailed: tablet counts, average sizes, FE memory before/after).
- 备注
本卡主题可能与本站 [避坑] 卡重叠。
this card's topic may overlap an existing [Pitfall] card on this site on this site.
StarRocks 年份:—