内核 1M–100M 区间的性价比之王——Rust 单机的高 QPS 与低延迟
- 一句话
- 同样的硬件,Qdrant 的单机向量检索 QPS/延迟显著优于 pgvector,比 Milvus 轻一个数量级的运维——"千万到亿级、要快、要便宜"场景的社区默认答案。Qdrant 的招牌不是"分布式",而是"单机效率 + 过滤工程"。On identical hardware, Qdrant's single-node vector retrieval QPS/latency clearly beats pgvector, with an order of magnitude less ops than Milvus — the community default for "tens of millions to a hundred million vectors, needs to be fast and cheap". Qdrant's signature isn't "distribution", it's "single-node efficiency + filtering engineering".
- 窄场景
- RAG 检索、Agent 记忆检索、实时推荐/搜索联想(高并发读、延迟敏感);1M–100M 向量(单机或小集群);5–20 人工程团队,有基础 Docker/K8s 能力但无专职 DBA/平台团队。典型画像:AI SaaS 的检索层、Agent infra。RAG retrieval, agent memory retrieval, real-time recommendation / search-as-you-type (high-concurrency reads, latency-sensitive); 1M–100M vectors (single node or small cluster); 5–20 person engineering teams with basic Docker/K8s skills but no dedicated DBA/platform team. Typical profiles: retrieval layers for AI SaaS, agent infra.
- 机制
- Rust 单二进制:无 GC 停顿(对比 Weaviate 的 Go GC 与 pgvector 的 Postgres MVCC 开销),内存用 jemalloc 按 size class 分配,HNSW 索引可 mmap 落盘,内存不够时退化而非 OOM。量化是工程而非噱头:标量量化(SQ,float32→uint8,4x 压缩)、二值量化(BQ,32x 内存压缩)、乘积量化(PQ)都是一等公民,可在查询时动态决定用原始向量 rescore,保证"压缩省内存、查询保召回"的可调 trade-off。架构选择的结果:Qdrant 故意只做 HNSW + 量化,把单条路走深,而不是像 Milvus 那样铺索引矩阵。A Rust single binary: no GC pauses (vs Weaviate's Go GC and pgvector's Postgres MVCC overhead), jemalloc size-class memory allocation, HNSW indexes that can mmap to disk — degrades instead of OOMing when memory runs short. Quantization is engineering, not gimmick: scalar quantization (SQ, float32→uint8, 4x compression), binary quantization (BQ, 32x memory compression), product quantization (PQ) are all first-class, with query-time rescoring against original vectors — a tunable "compress to save memory, rescore to protect recall" trade-off. An architectural choice with consequences: Qdrant deliberately does only HNSW + quantization, going deep on one path, instead of spreading an index matrix like Milvus.
- 生产验证
- HubSpot 的 200 亿+ 向量 VaaS(Vector-as-a-Service)平台:工程师在 Vector Space Day SF 2026(2026-06-11)公开演讲《Building the Infra Behind 20 Billion+ Vectors》,从手动 Helm 部署迁移到自研 K8s Operator(rolling upgrade、自动扩缩、自愈),服务 38+ 团队、200+ 索引、140+ 集群、5 地域,写峰值 10 万 QPS;选型理由明确写了"on-prem 部署 + named vectors + hybrid search + 多阶段查询 + 加权 rerank + 量化/on-disk 的成本控制"(独立媒体报道 https://www.snappr.com/news/story/qdrant-vector-space-day-sf-2026,数字来自演讲整理 https://briefly.co/anchor/DevOps/story/how-hubspot-scaled-semantic-search-to-20-billion-vectors)。Pinecone→Qdrant 独立迁移复盘:500 万法律文档向量,Pinecone Serverless $210/月 → 自托管 Qdrant $6/月服务器,同 1 万条测试查询 P50 23ms→4ms、P99 89ms→12ms、Recall@10 都是 0.97,迁移一个下午完成(独立博客、单人实测;硬件不对等——Pinecone 是 serverless 远端、Qdrant 是本地内存,延迟对比有方法学水分,成本对比可信度高,https://dev.to/chnby/i-run-5m-vectors-on-a-6mo-server-pinecone-would-charge-me-210-41lm)。独立基准(dev.to,2026-09,open harness 1M×1536):同等精度下 Qdrant p99 8.7ms,Milvus 576.7ms(https://dev.to/devrudals/enterprise-vector-database-2026-qdrant-vs-milvus-vs-pgvector-vs-pinecone-16a0,非学术基准,谨慎引用)。诚实备注:"xAI Grok 由 Qdrant 驱动"信源是 Qdrant 官方 2023 年推文经 TechCrunch 转述(https://techcrunch.com/2024/01/23/qdrant-open-source-vector-database/),xAI 从未官方确认技术栈细节——HubSpot 案例才是硬证据。HubSpot's 20-billion+ vector VaaS (Vector-as-a-Service) platform: engineers gave a public talk at Vector Space Day SF 2026 (2026-06-11), "Building the Infra Behind 20 Billion+ Vectors", on migrating from manual Helm deploys to a self-built K8s Operator (rolling upgrades, autoscaling, self-healing), serving 38+ teams, 200+ indexes, 140+ clusters, 5 regions, with peak write load of 100K QPS; selection reasons explicitly listed "on-prem deployment + named vectors + hybrid search + multi-stage queries + weighted rerank + quantization/on-disk cost control" (independent media coverage https://www.snappr.com/news/story/qdrant-vector-space-day-sf-2026, figures compiled from the talk https://briefly.co/anchor/DevOps/story/how-hubspot-scaled-semantic-search-to-20-billion-vectors). Independent Pinecone→Qdrant migration write-up: 5M legal-document vectors, Pinecone Serverless $210/mo → self-hosted Qdrant on a $6/mo server, same 10K test queries P50 23ms→4ms, P99 89ms→12ms, Recall@10 both 0.97, migrated in an afternoon (independent blog, single-person measurement; hardware not comparable — Pinecone is remote serverless, Qdrant is local memory, so the latency comparison has methodological caveats while the cost comparison is credible, https://dev.to/chnby/i-run-5m-vectors-on-a-6mo-server-pinecone-would-charge-me-210-41lm). Independent benchmark (dev.to, 2026-09, open harness 1M×1536): at equal precision Qdrant p99 8.7ms vs Milvus 576.7ms (https://dev.to/devrudals/enterprise-vector-database-2026-qdrant-vs-milvus-vs-pgvector-vs-pinecone-16a0, non-academic benchmark, cite with care). Honest note: the "xAI Grok runs on Qdrant" claim comes from a 2023 Qdrant tweet quoted by TechCrunch (https://techcrunch.com/2024/01/23/qdrant-open-source-vector-database/) and xAI has never confirmed its stack details — the HubSpot case is the hard evidence.
- 竞品差距
- vs pgvector:同规模 QPS 差一个数量级是社区反复验证的结论(10M×1536 下 pgvector ~30–100 QPS vs Qdrant 150–300;1M 向量 768 维某实测 Qdrant ~15K QPS vs pgvector ~1.5K QPS),pgvector 的 HNSW 与 Postgres 缓冲/MVCC 抢内存是机制性差距;vs milvus:1M–100M 区间 Milvus 是"杀鸡用牛刀"——运维复杂度 ★★★★ vs ★★,小规模下分布式开销反而拖累延迟,但 10 亿级以上反转;vs weaviate:同级延迟下 Qdrant 内存效率更高(Rust vs Go,量化成熟度),Weaviate 50M+ 后内存需求陡增是其官方生态都承认的。vs pgvector — same-scale QPS differing by an order of magnitude is a repeatedly validated community conclusion (~30–100 QPS for pgvector vs 150–300 for Qdrant at 10M×1536; a 1M-vector 768-dim test showed Qdrant ~15K QPS vs pgvector ~1.5K). pgvector's HNSW competing with the Postgres buffer pool/MVCC for memory is a structural gap; vs milvus — at 1M–100M Milvus is overkill, ops complexity ★★★★ vs ★★, and distribution overhead drags latency at small scale — reversed above a billion; vs weaviate — Qdrant is more memory-efficient at equal latency (Rust vs Go, more mature quantization); Weaviate's memory needs spiking past 50M+ vectors is admitted by its own ecosystem.
- 证据等级
- 社区共识(r/LocalLLaMA、HN、Discord 的一致画像)+ 独立实测(dev.to harness、个人迁移复盘)+ 具名生产案例(HubSpot 演讲);"二值量化 32x 内存、40x 检索速度"出自融资期 CEO 口径,未见独立复现,不单独成招牌Community consensus (consistent picture across r/LocalLLaMA, HN, Discord) + independent measurements (dev.to harness, personal migration write-up) + named production case (HubSpot talk); the "binary quantization 32x memory, 40x retrieval speed" figure came from a funding-era CEO quote and has no independent reproduction — not a standalone credential
- 最后核验
- 2026-10-012026-10-01