Submit projectSubmit

TOPIC GUIDE

Vector databases and RAG infrastructure

Pinecone, Qdrant, Weaviate, Milvus, Chroma and turbopuffer, plus tools that prepare data for retrieval.

products
19
things to compare
3
updated
2026-10-02

Retrieval-augmented generation needs three pieces: getting clean text out of documents and web pages, storing embeddings, and retrieving the right passages. Start with data preparation, since retrieval quality is usually limited by messy inputs before the database matters.

What to compare

  1. Decide between a managed service and self-hosting before comparing features.
  2. Test retrieval quality on your own documents and questions, not benchmark datasets.
  3. Estimate cost at your expected data size and query volume, including minimum monthly charges.

This is a focused selection, not an exhaustive ranking. See each profile for evidence and review scope.

Projects to explore

Pinecone

A managed vector database and knowledge platform for semantic search and retrieval in AI applications.

Read the project profile

Qdrant

An open-source vector search engine and database, available self-hosted or as Qdrant Cloud.

Read the project profile

Weaviate

An AI-native database with vector and hybrid search, available as source code to self-host or as Weaviate Cloud.

Read the project profile

Milvus

An open-source vector database for large-scale similarity search, also offered as managed Zilliz Cloud.

Read the project profile

Chroma

Open-source search infrastructure for AI with vector, full-text and metadata search, available locally or as Chroma Cloud.

Read the project profile

turbopuffer

A vector and full-text search engine built on object storage, designed to be cheaper and highly scalable.

Read the project profile

Unstructured

A platform that turns PDFs, documents and 64+ other file types into clean, structured inputs for RAG and AI applications.

Read the project profile

Firecrawl

A web data API for AI agents that searches, scrapes and crawls websites into clean markdown or structured data.

Read the project profile

LlamaIndex

AI agents for document parsing (LlamaParse) and workflows, alongside its framework for building on your data.

Read the project profile

Mem0

A memory layer for AI agents and apps that stores and retrieves what users said before, available as open source or a hosted platform.

Read the project profile

Jina AI

Embeddings, rerankers and a Reader API that turns web pages into Markdown for LLM search and grounding; now part of Elastic.

Read the project profile

Cohere

Enterprise AI models for generation, embeddings, reranking and speech, plus North, a secure workspace for agents and search.

Read the project profile

Docling

An open-source toolkit that converts PDFs, Office files, HTML, images, audio and more into structured data for AI, started at IBM Research.

Read the project profile

Crawl4AI

An open-source, LLM-friendly web crawler that turns any URL into clean Markdown or typed JSON, also available as a hosted API.

Read the project profile

RAGFlow

An open-source RAG engine with an integrated agent platform for giving AI agents reliable context from your documents.

Read the project profile

Graphiti

Zep's open-source framework for building temporal knowledge graphs that track how facts change over time, as memory for AI agents.

Read the project profile

Marker

Datalab's open-source tool for converting PDFs to Markdown and JSON quickly and accurately, including tables and layout.

Read the project profile

MarkItDown

Microsoft's lightweight Python tool for converting PDFs, Office documents and other files to Markdown for LLMs and text analysis.

Read the project profile

Onyx

An open-source enterprise search and AI chat platform that connects to company data and runs in your own cloud.

Read the project profile

Go a little deeper