Small. Fast. Reliable.
Choose any three.
Vec1 Tests on Publicly Available Data
Table Of Contents

1. Overview

This page contains some tests of the vec1 extension run on publicly available datasets distributed for testing ANN systems. The following datasets are used:

All four datasets are distributed with sample query vectors and results. "sift-128-euclidean" is distributed with dedicated training data (separate from the dataset itself), and the others are trained by sampling the main dataset. In all cases training and building the index are done with 16 threads. The "Training time" reported below is the wall-clock time the system spent running vec1_train() to create the model. The "Build time" is the wall-clock time spent running the 'rebuild' command to build the index.

Queries are run with a variety of values for parameters K and nprobe and the recall and throughput (queries/second) reported. The reported throughputs are for a single thread only.

Two flavors of recall are reported - recall@1 and recall@10. Recall@1 is the proportion of queries for the single nearest neighbor that do, in fact, return the true nearest neighbor. So a recall@1 of .916 means that if 1000 queries for the nearest neighbor are run, 916 of them return the true nearest neighbor.

Recall@10 is the average proportion of the true 10 nearest neighbors actually returned when the index is queried for the best 10 matches. i.e. if of the 10 results the query returns 7 of them are actually in the best 10 matches, the recall@10 is 0.7. The order of results returned does not matter - only the proportion that are part of the actual 10 best matches.

The SQL used for each query is:

SELECT *
FROM vec1tbl($query, '{K: $K, nprobe: $nprobe}')
ORDER BY vec1_l2_distance($query, vec1tbl.vector)
LIMIT $recall

Where $query is the query vector, $K is replaced by query parameter K, $nprobe by query parameter nprobe and $recall with the desired recall measure (1 or 10). For the "landmark-dino-768-cosine" index, vec1_cos_distance() is used instead of vec1_l2_distance(). In cases where $K==$recall the ORDER BY and LIMIT clauses are omitted from the query.

The results below were obtained on an AMD 5950X CPU based Linux workstation.

2. Dataset: sift-128-euclidean (1,000,000 128d vectors)

ParameterValue
Vec1 Versionversion 0.7 (AVX2, multi-threaded)
Index Parameters{nbucket:1024, quantizer:"pq", codesize:16}
Training time1.28s (100,000 samples, 16 threads)
Build time622ms (16 threads)

Recall@10 results:

K=10K=50K=100K=200K=300
nprobeRecallQPSRecallQPSRecallQPSRecallQPSRecallQPS
160.55679910.88252650.91940880.92728750.9292224
320.56546830.91635370.96229370.97522230.9771797
480.56733450.92426780.97323080.98818330.9901526
640.56826150.92721750.97719150.99215850.9951337

Recall@1 results:

K=1K=10K=50K=100K=200K=300
nprobeRecallQPSRecallQPSRecallQPSRecallQPSRecallQPSRecallQPS
160.46985720.88071660.94755520.95042530.95129090.9512250
320.47450550.90344400.98135860.98730040.98722520.9881820
480.47535930.90732210.98727160.99423410.99518540.9951541
640.47628000.90925360.99022010.99719540.99816160.9981351

3. Dataset: imagenet-clip-512-normalized (1,281,167 512d vectors)

ParameterValue
Vec1 Versionversion 0.7 (AVX2, multi-threaded)
Index Parameters{nbucket:1024, quantizer:"opq", codesize:32, residual:0}
Training time11.7s (100,000 samples, 16 threads)
Build time3.45s (16 threads)

Recall@10 results:

K=10K=50K=100K=200K=300
nprobeRecallQPSRecallQPSRecallQPSRecallQPSRecallQPS
160.52644170.89632610.95526840.97520050.9791611
320.52927110.90421810.96618870.98715150.9921270
480.52919800.90716590.97014760.99012160.9951059
640.52915680.90713460.97012150.99110390.996913

Recall@1 results:

K=1K=10K=50K=100K=200K=300
nprobeRecallQPSRecallQPSRecallQPSRecallQPSRecallQPSRecallQPS
160.41345210.87135630.97625970.98924370.99119330.9931610
320.41029350.87426150.98122020.99419010.99615230.9981278
480.41121350.87519310.98216760.99514840.99712310.9991064
640.41116850.87515370.98213570.99512260.99710450.999918

4. Dataset: landmark-dino-768-cosine (760,757 768d vectors)

ParameterValue
Vec1 Versionversion 0.7 (AVX2, multi-threaded)
Index Parameters{nbucket:1024, quantizer:"opq", codesize:48, residual:0}
Training time18.6s (100,000 samples, 16 threads)
Build time2.88s (16 threads)

Recall@10 results:

K=10K=50K=100K=200K=300
nprobeRecallQPSRecallQPSRecallQPSRecallQPSRecallQPS
160.58537470.90326360.93722260.94816830.9501370
320.59124510.92618560.96415790.97713120.9801111
480.59318690.93114720.97112710.98510480.988905
640.59315220.93312300.97410780.9889060.992792

Recall@1 results:

K=1K=10K=50K=100K=200K=300
nprobeRecallQPSRecallQPSRecallQPSRecallQPSRecallQPSRecallQPS
160.52039520.91430410.96821960.97020570.97215780.9721371
320.52526290.92622030.98618250.98815790.99012980.9901115
480.52520950.93018000.99114730.99312450.99510430.995902
640.52517070.93014750.99212360.99410840.9969170.996824

5. Dataset: agnews-mxbai-1024-euclidean (769,382 1024d vectors)

ParameterValue
Vec1 Versionversion 0.7 (AVX2, multi-threaded)
Index Parameters{nbucket:1024, quantizer:"opq", codesize:32, residual:0}
Training time25.9s (100,000 samples, 16 threads)
Build time4.75s (16 threads)

Recall@10 results:

K=10K=50K=100K=200K=300
nprobeRecallQPSRecallQPSRecallQPSRecallQPSRecallQPS
160.61839110.88927830.93022510.94416430.9491299
320.62328260.90421270.94917920.96513690.9701124
480.62521930.91117300.95814970.97411790.980985
640.62618040.91414660.96112900.97810440.983883

Recall@1 results:

K=1K=10K=50K=100K=200K=300
nprobeRecallQPSRecallQPSRecallQPSRecallQPSRecallQPSRecallQPS
160.61140250.92029220.96021290.96920880.96916260.9691291
320.60930620.92726370.97121480.98018030.98013800.9801131
480.61324140.93320810.98017400.99115050.99111840.991991
640.61219970.93417320.98214710.99312960.99310470.993887

This page was last updated on 2026-10-01 11:21:38Z