qdrant-mcp
MCP server for document ingestion and semantic search on Qdrant.
Overview
qdrant-mcp provides tools to:
- ingest local documents into a Qdrant collection
- generate embeddings with OpenAI
- run vector search with optional metadata filters
Features
ingest_documents- converts files such as
docx,pptx, andpdfto Markdown via MarkItDown - splits content into chunks using
chunk_sizeandoverlap_ratio - embeds chunks with OpenAI Embeddings (
text-embedding-3-smallby default) - upserts chunk text and metadata into Qdrant
- converts files such as
search_documents- embeds query text with the same embeddings API
- retrieves top
kmatches from Qdrant - supports filtering by
categoryandpath
Requirements
- Python 3.11+
uv- Qdrant (for example,
http://localhost:6333) OPENAI_API_KEY
Setup
uv sync
Run inside the Codex CLI
[mcp_servers.qdrant-mcp]
command = "uv"
args = ["run", "qdrant-mcp"]
cwd = "/sandbox/qdrant-mcp"
env = {
OPENAI_API_KEY = "sk-...",
QDRANT_URL = "http://127.0.0.1:6333",
QDRANT_API_KEY = "QDRANT_API_KEY",
QDRANT_COLLECTION = "codex_collection",
CHUNK_HEADER_MODEL = "gpt-5.4-mini"
}
Testing
Set OPENAI_API_KEY, QDRANT_URL, and QDRANT_API_KEY in .env, then run:
uv run python -m unittest tests/integration/test_qdrant_integration.py
MCP Tools
ingest_documents
Parameters:
paths: list[str]category: strchunk_size: int = 1200overlap_ratio: float = 0.15embedding_model: str = "text-embedding-3-small"chunk_header_mode: Literal["enabled", "disabled"] = "enabled"
Returns:
collectionembedding_modelingested_filesingested_pointsfailed_files
search_documents
Parameters:
query: strtop_k: int = 5category: str | None = Nonepath: str | None = Noneembedding_model: str = "text-embedding-3-small"
Returns:
collectionembedding_modelquerycountresults(score,path,category,chunk_index,text)
delete_documents_by_path
Parameters:
path: strcategory: str | None = None
Returns:
collectionpathcategorystatusoperation_id
list_category
Parameters:
limit: int = 100
Returns:
collectioncountcategories
list_path
Parameters:
category: strlimit: int = 1000
Returns:
collectioncategorycountpaths
Notes
- If the target collection does not exist, it is created automatically on first ingestion.
- If payload indexes for
categoryandpathdo not exist, they are created during ingestion. - By default, ingestion prepends a generated
Chunk-Header(max 64 chars), derived from the first 4096 bytes, to every chunk. - The Chunk-Header model is read from
CHUNK_HEADER_MODELwhenchunk_header_modeisenabled(default:gpt-5.4-mini). - The collection name is configured only through
QDRANT_COLLECTION(not via MCP tool parameters). - With
text-embedding-3-small, the vector size is1536.
License
See LICENSE.