内核 进程内 OLAP:pip install 即拥有的分析引擎
- 一句话
- 以库的形式链进你的进程:无 server、无端口、无守护进程,直接对 Parquet/CSV/S3 做向量化 OLAP——笔记本、CI、边缘设备、Lambda 里都能跑。Linked into your process as a library: no server, no ports, no daemons — vectorized OLAP directly against Parquet/CSV/S3. It runs in notebooks, CI, edge devices, and Lambda.
- 窄场景
- 数据科学家笔记本里 GB~百 GB 级分析(SELECT 直接查 Parquet 文件,无需建库导数);ETL 管道里的变换算子(替代 pandas 的内存效率、Spark 的部署重量);Serverless 函数(单 binary、无依赖、冷启动快)。GB-to-hundreds-of-GB analytics in a data scientist's notebook (SELECT straight off Parquet files — no database setup, no data loading); transform operators inside ETL pipelines (replacing pandas' memory inefficiency or Spark's deployment weight); serverless functions (single binary, no dependencies, fast cold start).
- 机制
- 列存 + 向量化执行(借鉴 MonetDB/X100 研究),但关键差异在部署形态:进程内意味着零 IPC/网络开销、零运维。对比 Spark(同样做聚合要起 JVM 集群)、pandas(行式内存模型在 group-by/join 上慢一个量级且内存爆炸)。读 Parquet in-place(文件夹即表)是另一机制要点。MotherDuck 的总结:"Simplify. Replace Apache Spark with DuckDB where possible."Columnar + vectorized execution (borrowing from MonetDB/X100 research) — but the real differentiator is deployment form: in-process means zero IPC/network overhead, zero ops. Contrast Spark (needs a JVM cluster for the same aggregation) and pandas (row-wise memory model is an order of magnitude slower on group-by/join and explodes memory). In-place Parquet reads ("the folder is the table") are the other key mechanism. MotherDuck's summary: "Simplify. Replace Apache Spark with DuckDB where possible."
- 生产验证
- FinQore(原 SaaSWorks,金融):财务 ETL 管道 PostgreSQL→DuckDB,8 小时→8 分钟(https://github.com/cyyeh/skills-playground/blob/HEAD/examples/system-explorer/duckdb/06-use-cases.md);Stockly:数据工程师用 DuckDB 做 TB 级数据迁移引擎替代 pandas("We're talking terabytes, and Pandas is not really fitted for this",Medium 具名访谈)(https://medium.com/@admin_31497/duckdb-in-practice-migrating-terabytes-with-sql-first-analytics-3ee735149f22);MotherDuck 梳理 15+ 生产用例(Rill、Evidence、Mode、Hex、Mosaic、Count 等产品内嵌)(https://MotherDuck.com/blog/15-companies-duckdb-in-prod/)。诚实备注:FinQore/Stockly 为社区转述,已降档。FinQore (ex-SaaSWorks, finance): financial ETL pipeline PostgreSQL→DuckDB, 8 hours→8 minutes (https://github.com/cyyeh/skills-playground/blob/HEAD/examples/system-explorer/duckdb/06-use-cases.md); Stockly: a data engineer used DuckDB as a terabyte-scale migration engine replacing pandas ("We're talking terabytes, and Pandas is not really fitted for this" — named Medium interview) (https://medium.com/@admin_31497/duckdb-in-practice-migrating-terabytes-with-sql-first-analytics-3ee735149f22); MotherDuck documents 15+ production use cases (Rill, Evidence, Mode, Hex, Mosaic, Count embedding it) (https://MotherDuck.com/blog/15-companies-duckdb-in-prod/). Honest note: FinQore/Stockly are community-reported — downgraded.
- 竞品差距
- ClickHouse/StarRocks/Doris 都是 server 形态,要部署、要运维、要网络;10GB 以下单机场景 DuckDB 通常更快社区共识。SQLite 同为嵌入式单文件但行存 OLTP——"same niche, opposite workload"。20GB 日志分析社区对比:DuckDB ~15 秒 vs PostgreSQL 3-5 分钟。31 款中"进程内 + 列存向量化 + 直接读 Parquet"的产品没有第二家。ClickHouse/StarRocks/Doris are all server-shaped — they need deployment, ops, and a network; below ~10GB single-node DuckDB usually wins 社区共识. SQLite is equally embedded but row-store OLTP — "same niche, opposite workload". A community 20GB log-analysis comparison: DuckDB ~15s vs PostgreSQL 3-5 minutes. No second product among the 31 is "in-process + columnar vectorized + direct Parquet reads".
- 证据等级
- 社区共识 + 具名生产案例(部分为社区转述,已标注)Community consensus + named production cases (some community-reported, as noted)
- 最后核验
- 2026-10-012026-10-01