内核 Photon 向量化引擎
- 一句话
- 用 C++ 重写的向量化执行引擎替换 JVM 上的 Spark 执行层,SQL/DataFrame workload 快 3-8 倍——且因为跑得快,总账单反而下降。A from-scratch C++ vectorized execution engine replacing the JVM Spark layer — SQL/DataFrame workloads run 3-8x faster, and because they finish faster, the total bill goes down.
- 窄场景
- 已有 Spark/Delta Lake 资产、不想重写 pipeline 但要降本提速的团队;Databricks SQL Warehouse 上的 BI/ETL(Photon 默认开启)。边界:UDF-heavy 的 Python workload 收益小(社区实测结论)。Teams with existing Spark/Delta Lake assets that want lower cost and higher speed without rewriting pipelines; BI/ETL on Databricks SQL Warehouse (Photon on by default). Boundary: UDF-heavy Python workloads gain little (community-measured).
- 机制
- 经典 Spark 跑在 JVM 上,行式处理 + 虚函数调用 + GC 开销;Photon 是从零写的 C++ 向量化引擎,直接操作列式批量数据、SIMD 指令,绕过 JVM。关键洞见是计费悖论:Photon 的 DBU 单价贵约 2 倍,但查询快 3-8 倍,总 DBU-hours 下降——"为性能付费结果省了钱"的反直觉案例。Classic Spark on the JVM means row-wise processing + virtual-function calls + GC overhead; Photon is a C++ vectorized engine operating directly on columnar batches with SIMD, bypassing the JVM. The key insight is the billing paradox: Photon DBUs cost ~2x more per unit, but queries run 3-8x faster, so total DBU-hours fall — "paying for performance and saving money", the counter-intuitive case.
- 生产验证
- 独立实测(Reliable Data Engineering,2026-01):500M 行复杂聚合不用 Photon 4 分 23 秒 / $0.87,用 Photon 1 分 48 秒 / $0.65——快 59%,便宜 25%(https://medium.com/@reliabledataengineering/15-databricks-features-youre-paying-for-but-not-using-09b12d7dea45);社区 FinOps 实践已是生产调优 checklist 标准条目:"Turn on Photon for SQL and DataFrame-heavy ETL. It usually pays for its DBU premium in reduced runtime."(https://github.com/surfalytics/rockyourdata/blob/HEAD/src/content/blog/databricks-lakehouse-architecture.md);Intel 白皮书 TPC-DS 快 65%/省 35% 为厂商口径,仅作对照(https://cdrdv2-public.intel.com/755539/Photon-Databricks-on-Azure-Edsv4.pdf)。Independent benchmark (Reliable Data Engineering, Jan-2026): a 500M-row complex aggregation took 4m23s / $0.87 without Photon vs 1m48s / $0.65 with it — 59% faster, 25% cheaper (https://medium.com/@reliabledataengineering/15-databricks-features-youre-paying-for-but-not-using-09b12d7dea45); community FinOps practice lists it as a standard production-tuning checklist item: "Turn on Photon for SQL and DataFrame-heavy ETL. It usually pays for its DBU premium in reduced runtime." (https://github.com/surfalytics/rockyourdata/blob/HEAD/src/content/blog/databricks-lakehouse-architecture.md); Intel's TPC-DS whitepaper (65% faster / 35% cheaper) is vendor-sourced, reference only (https://cdrdv2-public.intel.com/755539/Photon-Databricks-on-Azure-Edsv4.pdf).
- 竞品差距
- Spark 原生/EMR 无等价物(开源 Comet、Gluten+Velox 在追赶,但成熟度/集成度不及);Snowflake/BigQuery 各有自己的向量化引擎,但那是"买数仓送引擎",Photon 是 Spark 生态的加速器——存量 Spark 代码不用改;StarRocks/Doris 本身就是 C++ 向量化 MPP 但它们不是 Spark。31 款中"让存量 Spark workload 不动代码变快 3-8 倍"的方案没有第二家。Native Spark/EMR have no equivalent (open-source Comet and Gluten+Velox are chasing the same idea but trail in maturity/integration); Snowflake/BigQuery each have their own vectorized engines, but those are "a warehouse that ships an engine" — Photon is an accelerator for the Spark ecosystem, and your existing Spark code doesn't change; StarRocks/Doris are C++ vectorized MPPs but they aren't Spark. No second product among the 31 makes existing Spark workloads 3-8x faster without code changes.
- 证据等级
- 社区共识 + 独立实测;Intel 白皮书数字为厂商口径,已标注Community consensus + independent benchmarks; Intel figures are vendor-sourced, as noted
- 最后核验
- 2026-10-012026-10-01