Qdrant and Vultr built Qdrant-FineWeb-10B, an open-source vector search benchmark of 10 billion embeddings derived from Hugging Face’s FineWeb corpus using the gte-multilingual-base model. To evaluate at this scale, the team developed Supernova, an open-source framework for embedding generation, brute-force ground-truth computation across dense and sparse vectors, and distributed benchmarking across cloud and HPC infrastructure. The release aims to address the limits of existing benchmarks, which typically top out at 10-100 million vectors, by providing exact ground truth and support for hybrid search patterns at billion-scale.
