Pre-warm database before production for pgvector
It is useful to execute a few thousand warm-up queries before going into production. This helps with RAM utilization and can also help determine that you have selected the right compute size for your workload.
Fine-tune HNSW index parameters to improve performance
You can increase requests per second by increasing the `m` and `ef_construction` HNSW parameters. This allows you to use smaller values for `ef_search` and increase QPS. However, building the index takes longer with higher values for these parameters.
Vector workload scaling options for Supabase
There are two approaches to scaling vector workloads: (1) Increase the size of your database, or (2) Spread your workload across multiple databases by consulting the Engineering for Scale guide.
Embedding dimensionality is the most important factor for compute add-on choice
The number of dimensions in embeddings is the most important factor in choosing the right Compute Add-on. In general, lower dimensionality provides better performance.
Fine-tune IVFFlat index parameters to improve performance
You can increase requests per second by increasing the `lists` parameter for IVFFlat indexes. This also has the caveat that building the index takes longer with higher values for the lists parameter.
Benchmark methodology uses ANN Benchmarks techniques with vecs
Supabase follows techniques outlined in the ANN Benchmarks methodology. A Python test runner is responsible for uploading the data, creating the index, and running the queries using the vecs Python client for pgvector. Each test is run for a minimum of 30-40 minutes with experiments executed at different concurrency levels to measure engine performance under different load types. Results are then averaged. As a general recommendation, use a concurrency level of 5 or more for most workloads and 30 or more for high-load workloads.
HNSW benchmark results for 1536 dimensions (OpenAI embeddings)
HNSW benchmark using dbpedia-entities-openai-1M dataset (1,000,000 embeddings for compute sizes XL and above) and Wikipedia articles (224,482 embeddings for compute sizes Large and below) at 1536 dimensions created with OpenAI Embeddings API. Accuracy was 0.99 for all benchmarks. QPS can be improved by increasing m and ef_construction parameters. For example, increasing m to 32 and ef_construction to 80 for 4XL will increase QPS to 1280.
| Compute Size | Vectors | m | ef_construction | ef_search | QPS | Latency Mean | Latency p95 | RAM Usage | RAM |
| ------------ | --------- | --- | --------------- | --------- | ---- | ------------ | ----------- | ------------- | ------ |
| Micro | 15,000 | 16 | 40 | 40 | 480 | 0.011 sec | 0.016 sec | 1.2 GB (Swap) | 1 GB |
| Small | 50,000 | 32 | 64 | 100 | 175 | 0.031 sec | 0.051 sec | 2.2 GB (Swap) | 2 GB |
| Medium | 100,000 | 32 | 64 | 100 | 240 | 0.083 sec | 0.126 sec | 4 GB | 4 GB |
| Large | 224,482 | 32 | 64 | 100 | 280 | 0.017 sec | 0.028 sec | 8 GB | 8 GB |
| XL | 500,000 | 24 | 56 | 100 | 360 | 0.055 sec | 0.135 sec | 13 GB | 16 GB |
| 2XL | 1,000,000 | 24 | 56 | 250 | 560 | 0.036 sec | 0.058 sec | 32 GB | 32 GB |
| 4XL | 1,000,000 | 24 | 56 | 250 | 950 | 0.021 sec | 0.033 sec | 39 GB | 64 GB |
| 8XL | 1,000,000 | 24 | 56 | 250 | 1650 | 0.016 sec | 0.023 sec | 40 GB | 128 GB |
| 12XL | 1,000,000 | 24 | 56 | 250 | 1900 | 0.015 sec | 0.021 sec | 38 GB | 192 GB |
| 16XL | 1,000,000 | 24 | 56 | 250 | 2200 | 0.015 sec | 0.020 sec | 40 GB | 256 GB |
Performance impact of uploading more vectors than benchmark limits
It is possible to upload more vectors to a single table if memory allows (for example, 4XL plan and higher for OpenAI embeddings). However, this will affect query performance: QPS will be lower and latency will be higher. Scaling should be almost linear, but it is recommended to benchmark your workload to find the optimal number of vectors per table and per database instance.
IVFFlat benchmark results for 960 dimensions (gist-960 images, probes=10)
IVFFlat benchmark using gist-960 dataset with 1,000,000 embeddings of images at 960 dimensions with probe parameter set to 10.
| Compute Size | Vectors | Lists | QPS | Latency Mean | Latency p95 | RAM Usage | RAM |
| ------------ | --------- | ----- | ---- | ------------ | ----------- | ------------- | ------ |
| Micro | 30,000 | 30 | 75 | 0.065 sec | 0.088 sec | 1.1 GB (Swap) | 1 GB |
| Small | 100,000 | 100 | 78 | 0.064 sec | 0.092 sec | 1.8 GB | 2 GB |
| Medium | 250,000 | 250 | 58 | 0.085 sec | 0.129 sec | 3.2 GB | 4 GB |
| Large | 500,000 | 500 | 55 | 0.088 sec | 0.140 sec | 5 GB | 8 GB |
| XL | 1,000,000 | 1000 | 110 | 0.046 sec | 0.070 sec | 14 GB | 16 GB |
| 2XL | 1,000,000 | 1000 | 235 | 0.083 sec | 0.136 sec | 10 GB | 32 GB |
| 4XL | 1,000,000 | 1000 | 420 | 0.071 sec | 0.106 sec | 11 GB | 64 GB |
| 8XL | 1,000,000 | 1000 | 815 | 0.072 sec | 0.106 sec | 13 GB | 128 GB |
| 12XL | 1,000,000 | 1000 | 1150 | 0.052 sec | 0.078 sec | 15.5 GB | 192 GB |
| 16XL | 1,000,000 | 1000 | 1345 | 0.072 sec | 0.106 sec | 17.5 GB | 256 GB |
Production deployment checklist for vector search applications
Step 1: Decide if you will use indexes (can skip remaining steps if not). Step 2: Over-provision RAM during preparation; start with larger size (recommend at least 8XL for Supabase) to determine actual requirements, then scale down later. Step 3: Upload data to database; vecs library automatically generates index with default parameters. Step 4: Run benchmark with randomly generated queries using default index build parameters. Step 5: Monitor RAM usage and note it for future compute add-on selection. Step 6: Scale down compute add-on to one matching the observed RAM usage. Step 7: Reload data into RAM; QPS should increase on subsequent runs until it plateaus. Step 8: Run benchmark with real queries and tweak ef_search (HNSW) or probes (IVFFlat) until accuracy and QPS meet requirements. Step 9: For higher QPS, increase m and ef_construction (HNSW) or lists (IVFFlat) and rebuild index; repeat steps 6-7 to find optimal combination.
Tuning HNSW index build parameters for better performance
Fine-tune m and ef_construction for HNSW to accelerate queries at the expense of slower build times. For example, benchmarking 1,000,000 OpenAI embeddings with m=32 and ef_construction=80 resulted in 35% higher QPS compared to m=24 and ef_construction=56.
pgvector maximum dimensions for index creation
In pgvector versions 0.7.0 and above, the maximum dimensions supported for creating indexes are: vector up to 2,000 dimensions, halfvec up to 4,000 dimensions, and bit up to 64,000 dimensions.
Upgrade pgvector to access higher dimensional indexes
If you are on an earlier version of pgvector before 0.7.0, you should upgrade your project to access indexes supporting higher maximum dimensions.
Fewer dimensions perform better in pgvector
In general, embeddings with fewer dimensions perform best. Supabase has published analysis on the benefits of using fewer dimensions in pgvector.
IVFFlat benchmark results for 1536 dimensions (OpenAI embeddings)
IVFFlat benchmark using dbpedia-entities-openai-1M dataset with 1,000,000 embeddings at 1536 dimensions.
## Accuracy by Probe Parameter
- 10 probes: 0.91 accuracy for 1,000,000 vectors; 0.95-0.99 accuracy for 500,000 vectors and below
- 40 probes: 0.98 accuracy for 1,000,000 vectors
## Benchmark Results - 10 Probes
| Compute Size | Vectors | Lists | QPS | Latency Mean | Latency p95 | RAM Usage | RAM |
| ------------ | --------- | ----- | ---- | ------------ | ----------- | ------------- | ------ |
| Micro | 20,000 | 40 | 135 | 0.372 sec | 0.412 sec | 1.2 GB (Swap) | 1 GB |
| Small | 50,000 | 100 | 140 | 0.357 sec | 0.398 sec | 1.8 GB | 2 GB |
| Medium | 100,000 | 200 | 130 | 0.383 sec | 0.446 sec | 3.7 GB | 4 GB |
| Large | 250,000 | 500 | 130 | 0.378 sec | 0.434 sec | 7 GB | 8 GB |
| XL | 500,000 | 1000 | 235 | 0.213 sec | 0.271 sec | 13.5 GB | 16 GB |
| 2XL | 1,000,000 | 2000 | 380 | 0.133 sec | 0.236 sec | 30 GB | 32 GB |
| 4XL | 1,000,000 | 2000 | 720 | 0.068 sec | 0.120 sec | 35 GB | 64 GB |
| 8XL | 1,000,000 | 2000 | 1250 | 0.039 sec | 0.066 sec | 38 GB | 128 GB |
| 12XL | 1,000,000 | 2000 | 1600 | 0.030 sec | 0.052 sec | 41 GB | 192 GB |
| 16XL | 1,000,000 | 2000 | 1790 | 0.029 sec | 0.051 sec | 45 GB | 256 GB |
## Benchmark Results - 40 Probes
| Compute Size | Vectors | Lists | QPS | Latency Mean | Latency p95 | RAM Usage | RAM |
| ------------ | --------- | ----- | --- | ------------ | ----------- | --------- | ------ |
| 2XL | 1,000,000 | 2000 | 140 | 0.358 sec | 0.575 sec | 30 GB | 32 GB |
| 4XL | 1,000,000 | 2000 | 270 | 0.186 sec | 0.304 sec | 35 GB | 64 GB |
| 8XL | 1,000,000 | 2000 | 470 | 0.104 sec | 0.166 sec | 38 GB | 128 GB |
| 12XL | 1,000,000 | 2000 | 600 | 0.085 sec | 0.132 sec | 41 GB | 192 GB |
| 16XL | 1,000,000 | 2000 | 670 | 0.081 sec | 0.129 sec | 45 GB | 256 GB |
HNSW benchmark results: 384 dimensions (gte-small) and 960 dimensions (gist-960 images)
HNSW benchmarks testing vector scaling across two datasets with different dimensionalities. Accuracy was 0.99 for all benchmarks in both tests.
## Benchmark 1: 384 Dimensions (gte-small)
Dataset: dbpedia-entities-openai-1M with 1,000,000 embeddings at 384 dimensions using gte-small embeddings and inner-product distance measure.
| Compute Size | Vectors | m | ef_construction | ef_search | QPS | Latency Mean | Latency p95 | RAM Usage | RAM |
| ------------ | --------- | --- | --------------- | --------- | ---- | ------------ | ----------- | ---------- | ------ |
| Micro | 100,000 | 16 | 64 | 60 | 580 | 0.017 sec | 0.024 sec | 1.2 (Swap) | 1 GB |
| Small | 250,000 | 24 | 64 | 60 | 440 | 0.022 sec | 0.033 sec | 2 GB | 2 GB |
| Medium | 500,000 | 24 | 64 | 80 | 350 | 0.028 sec | 0.045 sec | 4 GB | 4 GB |
| Large | 1,000,000 | 32 | 80 | 100 | 270 | 0.073 sec | 0.108 sec | 7 GB | 8 GB |
| XL | 1,000,000 | 32 | 80 | 100 | 525 | 0.038 sec | 0.059 sec | 9 GB | 16 GB |
| 2XL | 1,000,000 | 32 | 80 | 100 | 790 | 0.025 sec | 0.037 sec | 9 GB | 32 GB |
| 4XL | 1,000,000 | 32 | 80 | 100 | 1650 | 0.015 sec | 0.018 sec | 11 GB | 64 GB |
| 8XL | 1,000,000 | 32 | 80 | 100 | 2690 | 0.015 sec | 0.016 sec | 13 GB | 128 GB |
| 12XL | 1,000,000 | 32 | 80 | 100 | 3900 | 0.014 sec | 0.016 sec | 13 GB | 192 GB |
| 16XL | 1,000,000 | 32 | 80 | 100 | 4200 | 0.014 sec | 0.016 sec | 20 GB | 256 GB |
## Benchmark 2: 960 Dimensions (gist-960 images)
Dataset: gist-960 with 1,000,000 image embeddings at 960 dimensions.
| Compute Size | Vectors | m | ef_construction | ef_search | QPS | Latency Mean | Latency p95 | RAM Usage | RAM |
| ------------ | --------- | --- | --------------- | --------- | ---- | ------------ | ----------- | ------------- | ------ |
| Micro | 30,000 | 16 | 64 | 65 | 430 | 0.024 sec | 0.034 sec | 1.2 GB (Swap) | 1 GB |
| Small | 100,000 | 32 | 80 | 60 | 260 | 0.040 sec | 0.054 sec | 2.2 GB (Swap) | 2 GB |
| Medium | 250,000 | 32 | 80 | 90 | 120 | 0.083 sec | 0.106 sec | 4 GB | 4 GB |
| Large | 500,000 | 32 | 80 | 120 | 160 | 0.063 sec | 0.087 sec | 7 GB | 8 GB |
| XL | 1,000,000 | 32 | 80 | 200 | 200 | 0.049 sec | 0.072 sec | 13 GB | 16 GB |
| 2XL | 1,000,000 | 32 | 80 | 200 | 340 | 0.025 sec | 0.029 sec | 17 GB | 32 GB |
| 4XL | 1,000,000 | 32 | 80 | 200 | 630 | 0.031 sec | 0.050 sec | 18 GB | 64 GB |
| 8XL | 1,000,000 | 32 | 80 | 200 | 1100 | 0.034 sec | 0.048 sec | 19 GB | 128 GB |
| 12XL | 1,000,000 | 32 | 80 | 200 | 1420 | 0.041 sec | 0.095 sec | 21 GB | 192 GB |
| 16XL | 1,000,000 | 32 | 80 | 200 | 1650 | 0.037 sec | 0.081 sec | 23 GB | 256 GB |
## Key Observations
- QPS can be improved by increasing m and ef_construction parameters, allowing for smaller ef_search values
- Higher dimensional embeddings (960 vs 384 dimensions) show higher RAM requirements and lower throughput at comparable compute sizes
- Both benchmarks demonstrate strong accuracy (0.99) across all scaling tiers