Vector Search Guide¶
Bring your own embeddings (OpenAI, HF, local models). This guide focuses on the recommended SQL-first way to build and query vector indexes (JVector) in embedded mode.
Quick Start (Embedded, Minimal)¶
import arcadedb_embedded as arcadedb
from arcadedb_embedded import to_java_float_array
texts = ["python database", "graph queries", "vector search"]
with arcadedb.create_database("./vector_demo") as db:
db.command("sql", "CREATE VERTEX TYPE Doc")
db.command("sql", "CREATE PROPERTY Doc.text STRING")
db.command("sql", "CREATE PROPERTY Doc.embedding ARRAY_OF_FLOATS")
db.command(
"sql",
"""
CREATE INDEX ON Doc (embedding)
LSM_VECTOR
METADATA {
"dimensions": 3,
"similarity": "COSINE"
}
""",
)
with db.transaction():
for i, t in enumerate(texts):
embedding = [float(i == j) for j in range(3)] # toy vectors
db.command(
"sql",
"INSERT INTO Doc SET text = ?, embedding = ?",
t,
to_java_float_array(embedding),
)
rows = db.query(
"sql",
"SELECT vectorNeighbors('Doc[embedding]', ?, 2) as res",
# bind the query vector; a bare Python list as the only argument would be
# read as the parameter array and bind ? to 0.9
to_java_float_array([0.9, 0.1, 0.0]),
).to_list()
for hit in rows[0].get("res", []):
record = hit.get("record")
if record is not None:
print(record.get("text"), hit.get("distance"))
API Essentials¶
Preferred split:
- Use SQL
CREATE INDEX ... LSM_VECTOR METADATA {...}for vector index creation. - Prefer SQL or Cypher for vector retrieval/search, because search composes naturally with filters, projections, and graph traversal.
-
Keep the secondary Python helper APIs in mind only for manual or maintenance cases; they are not the recommended application-facing workflow.
-
Vector property type is usually
ARRAY_OF_FLOATS. - Use
BINARYonly when you are storing pre-quantized INT8 bytes withencoding="INT8". CREATE INDEX ON Doc (embedding) LSM_VECTOR METADATA {...}is the preferred creation path.- SQL builds the vector graph immediately by default.
- Add
"buildGraphNow": falseonly if you intentionally want lazy preparation.
- Use
vectorNeighbors(..., k, ef_search)as the default SQL nearest-neighbor surface. get_metadata()remains available on the loaded vector index when you need to inspect index configuration from Python.get_stats()returns live counters (cache hits and misses, where vectors were read from, graph state) for sizing caches and diagnosing build cost. See Vector Caches.- After a database is opened, an index loads its graph on the first search, and that
search pays for it. A service that restarts can call
warm_up()on the loaded index right after opening (26.10.1 and later), or run one throwaway search on older engines. SeeVectorIndex.warm_up().
Distance Functions (scoring behavior)¶
cosine(default): returns cosine distance in [0,2]; lower is better.euclidean: returns squared Euclidean distance (d²); lower is better.dot_product: returns-(1 + A·B) / 2; lower is better. It expects unit-length vectors, and the engine logs a warning when sampled vectors are not.
Important:
- Vector-index search exposes cosine as distance: 1 - cos(θ).
- SQL
vectorCosineSimilarity(...)exposes raw cosine similarity, so its values follow cosine similarity semantics rather than vector-index distance semantics.
Tuning Knobs¶
dimensions: must match your embedding length.max_connections(Vamana per-layer degree; use 2*M to match an hnswlib M): higher → better recall, more memory/slower build (default: 32).beam_width(ef/efConstruction): higher → better recall, slower search/build (default: 100).ef_search(runtime, exact search only): higher → better recall, slower search.- Unset, the engine picks the beam: 100 below 10,000 vectors, and
max(2k, 20)from there up. On 20,000 random 32-dimension vectors that default found 80% of the true top 10, where a beam of 100 found 99.7%. Set it for the recall you need.
- Unset, the engine picks the beam: 100 below 10,000 vectors, and
The search beam in SQL¶
vectorNeighbors(index, vector, k) takes an optional fourth argument, the search
beam (efSearch): how many candidates the graph walk keeps before returning
k. Higher finds more of the true neighbours and costs time; the engine's
default is adaptive and small on larger graphs (see above). Pass it explicitly
when you compare recall across settings or engines, so the number you quote is
the number that ran:
Searches while the graph rebuilds¶
Deletes and updates collect as pending changes, and once they reach a fifth of
the graph (arcadedb.vectorIndex.rebuildGraphRatio, 0.2), and sometimes
earlier, the engine rebuilds the graph in the background while searches keep
running. On 26.9.1 and earlier, ArcadeDB
#8862 makes those searches
return noticeably worse neighbours for as long as the rebuild runs: on 20,000
vectors, recall@10 at a beam of 100 fell from 0.99 to between 0.14 and 0.31 for
the four seconds of the rebuild. Fixed in 26.10.1 (PR #8864, verified on its
merge): the same run keeps recall@10 between 0.985 and 1.000 through the
rebuild. On an older release, run bulk deletes and updates at a quiet time; a
rebuild takes about as long as building the index, and the engine logs
Built graph for index when the new graph is in place.
At about 10M vectors with sustained inserts while a rebuild runs, the next
rebuild can be deferred indefinitely, and every search then pays a growing scan
of the vectors not yet in the graph (ArcadeDB
#7260, open). A heap of
about twice the graph, a lower mutations_before_rebuild, or rebuilding in a
quiet window avoids it. Read-mostly use, and loading first and serving after,
never meet it.
Build-time cache: use the default¶
Building the graph for millions of vectors is dominated by reading the vectors
back, so the engine caches them while it builds. The default sizing is
automatic and is the right setting: it reads the heap the engine actually has
free and caches the whole corpus when it fits, inside a share of the heap
(arcadedb.vectorIndex.graphBuildCacheMaxHeapPercent, 25%). Give the JVM the
heap and leave graphBuildCacheSize alone; set an absolute count only to bound
a build on a deliberately small heap. Engines before 26.10 sized it from a
post-GC heap figure that included the page cache and could pick a fraction of
the corpus on a large heap (a 10M build took 7,000 s that way against 2,300 s
with the corpus cached); upstream fixed the sizing, so nothing in the bindings'
tests or examples sets it. See the
Memory & Heap section for the
heap side.
Memory & Heap Requirements (1024-dim vectors)¶
Vector index build is the most memory-hungry step. For 1024-dimensional vectors:
Build (heap):
- 1M vectors: at least 4G
- 2M vectors: at least 8G
- 4M vectors: at least 16G
- 8M vectors: at least 32G
Search (heap):
- 1M vectors: 1G works, at least 1G recommended
- 2M vectors: 1G works, at least 2G recommended
- 4M vectors: 1G works, at least 2G recommended
- 8M vectors: 1G OOM, 2G works, at least 4G recommended
If you reduce vector dimensions (e.g., 384-dim), you can substantially lower heap requirements.
Generating Embeddings (example)¶
Use any model you like; the bindings only need a Python list/NumPy array of floats. A typical text workflow uses a Transformer-based embedding model, e.g., sentence-transformers with normalized outputs for cosine:
import json
import arcadedb_embedded as arcadedb
from sentence_transformers import SentenceTransformer
from arcadedb_embedded import to_java_float_array
model = SentenceTransformer("all-MiniLM-L6-v2") # 384 dims
doc_text = "retrieval augmented generation"
vec = model.encode(doc_text, normalize_embeddings=True)
with arcadedb.create_database("./vector_demo") as db:
db.command("sql", "CREATE VERTEX TYPE Doc")
db.command("sql", "CREATE PROPERTY Doc.text STRING")
db.command("sql", "CREATE PROPERTY Doc.embedding ARRAY_OF_FLOATS")
metadata_json = json.dumps({"dimensions": len(vec), "similarity": "COSINE"})
db.command(
"sql",
"CREATE INDEX ON Doc (embedding) LSM_VECTOR METADATA " + metadata_json,
)
with db.transaction():
db.command(
"sql",
"INSERT INTO Doc SET text = ?, embedding = ?",
doc_text,
to_java_float_array(vec),
)
hits = db.query(
"sql",
"SELECT vectorNeighbors('Doc[embedding]', ?, 1) as res",
to_java_float_array(vec),
).to_list()
Notes:
to_java_float_arrayaccepts NumPy arrays directly.- For cosine, pass
normalize_embeddings=True(as above) to your model. - For euclidean, skip normalization if magnitude should matter. For dot_product, normalize: the engine expects unit-length vectors.
Preferred Search Surface: SQL / Cypher¶
For new code, prefer query APIs for search. Bind the query vector as a parameter
(to_java_float_array(query_vec)) rather than pasting it into the SQL text: a pasted
vector makes every query a new statement to parse (see
Parameters).
SQL filtered vector search with score shaping¶
from arcadedb_embedded import to_java_float_array
rows = db.query(
"sql",
(
"SELECT title, category, distance, (1 - distance) AS score "
"FROM (SELECT expand(vectorNeighbors('Article[embedding]', ?, 50))) "
"WHERE category = ? ORDER BY distance LIMIT 5"
),
to_java_float_array(query_vec),
"category_42",
).to_list()
SQL self-exclusion¶
rows = db.query(
"sql",
(
"SELECT title, distance, (1 - distance) AS score "
"FROM (SELECT expand(vectorNeighbors('Movie[embedding]', ?, 20))) "
"WHERE title <> ? ORDER BY distance LIMIT 10"
),
to_java_float_array(query_vec),
movie_title,
).to_list()
Cypher search with score shaping¶
rows = db.query(
"opencypher",
(
"CALL vector.neighbors('Doc[embedding]', $vec, $k) "
"YIELD name, distance RETURN name, (1 - distance) AS score ORDER BY score DESC"
),
{"vec": query_vec, "k": 5},
).to_list()
Quantization¶
quantizationaccepts"INT8","BINARY","PRODUCT"(PQ), orNone(full precision).- Default and recommended setting is
"INT8". - Use
quantization=Noneonly when you explicitly need full-precision vectors and can accept higher memory usage. - PQ tunables (require
quantization="PRODUCT"):pq_subspaces(M),pq_clusters(K),pq_center_globally,pq_training_limit. - In current ArcadeDB engine builds, PRODUCT/PQ also needs enough indexed vectors per
bucket for training. For very small corpora, set
pq_clustersto a value no larger than the number of indexed vectors in that bucket, or useINT8,BINARY, orNone. "PRODUCT"/PQ is currently not recommended for production workloads in these bindings.- SQL quantization helpers:
vector.quantizeInt8/vector.dequantizeInt8andvector.quantizeBinary/vector.dequantizeBinary(v[, low, high]). Binary quantization keeps only the sign of each element; dequantize reconstructs it aslow/high(default-1.0/1.0).
SQL Helpers¶
- Preferred path for embedded and server modes:
CREATE INDEX ON Doc (embedding) LSM_VECTOR METADATA {"dimensions": 128, "similarity": "COSINE"}(the key issimilarity;distanceFunctionis rejected by the index, though it remains the spellingIMPORT DATABASE ... WITH distanceFunction = cosineuses on that separate surface) - Search via SQL:
SELECT vectorNeighbors('Doc[embedding]', ?, 5) AS res, with the query vector bound (to_java_float_array(vec))
- Math/distance helpers:
vectorCosineSimilarity,vectorL2Distance,vectorDotProduct,vectorNormalize,vectorAdd,vectorSum, etc. - Every
vector.xxxfunction is also reachable as a camelCase alias (vector.l2Norm⇔vectorL2Norm), so both spellings resolve to the same function. - Distance:
vectorL2Distance(Euclidean),vectorManhattanDistance/vectorL1Distance(L1, sum of absolute differences). - Norms:
vectorL2Norm/vectorMagnitude(Euclidean length),vectorLInfNorm(max absolute element). - Element-wise math:
vectorAdd,vectorSubtract,vectorMultiply,vectorScale, andvectorClamp(v, min, max)to bound each element to a range.vectorAddandvectorSubtractalso broadcast a scalar operand (e.g.vectorAdd([1,2,3], 4)). - Quality checks:
vectorHasNull(v)returnstruewhen a vector contains a NaN/null element (useful for guarding embeddings before indexing). - Stats:
vectorStdDev,vectorVariance, andvector.sparsity(v[, threshold, mode])wheremodeis the default ratio,L0(count of significant elements), orGMEAN. - Score shaping:
vector.scoreTransform(score, mode)with modes such asLN/LOGandTANH, andvector.multiScore(scores, fusion)(e.g.MAX) to fuse score lists. - Conversions:
.asString(format)renders a vector as text. Formats areCOMPACT(default),PRETTY,PYTHON,JULIA,MATLAB,MATLAB_COLUMN, andNUMPY, whereNUMPYemits a bare comma-separated list ready fornp.array(s.split(','), dtype=np.float32)..asVector()is the inverse and parses any of those layouts (or a number/list) back into a vector, e.g.SELECT [1.0, 2.5].asString('NUMPY').asVector()..asSparse([threshold])converts a dense vector to a sparse one (see Sparse Vectors). - Quantization via SQL:
METADATA {"quantization": "INT8"}is the recommended path for embedded usage.
Native INT8 Storage¶
If your application already has INT8 vectors, store them in a BINARY property and
set encoding="INT8" on the vector index metadata.
import arcadedb_embedded as arcadedb
with arcadedb.create_database("./vector_demo_int8") as db:
db.command("sql", "CREATE VERTEX TYPE ByteDoc")
db.command("sql", "CREATE PROPERTY ByteDoc.id STRING")
db.command("sql", "CREATE PROPERTY ByteDoc.embedding BINARY")
db.command(
"sql",
"""
CREATE INDEX ON ByteDoc (embedding)
LSM_VECTOR
METADATA {
"dimensions": 4,
"similarity": "COSINE",
"quantization": "NONE",
"encoding": "INT8"
}
""",
)
with db.transaction():
db.command(
"sql",
"INSERT INTO ByteDoc SET id = ?, embedding = ?",
"doc_a",
arcadedb.to_java_byte_array([127, 0, 0, 0]),
)
Use encoding="INT8" only with quantization="NONE". Combining INT8 storage encoding
with INT8 quantization would quantize the same vector twice.
Sparse Vectors¶
ArcadeDB also supports sparse top-K retrieval through LSM_SPARSE_VECTOR and
vector.sparseNeighbors(...). To convert between representations in SQL, use
.asSparse([threshold]) (dense → sparse, keeping elements above the threshold)
and vector.sparseToDense(sv) (sparse → dense), e.g.
SELECT [0.0, 5.0, 0.0].asSparse().
import arcadedb_embedded as arcadedb
import jpype.types as jtypes
with arcadedb.create_database("./sparse_demo") as db:
db.command("sql", "CREATE DOCUMENT TYPE SparseDoc")
db.command("sql", "CREATE PROPERTY SparseDoc.tokens ARRAY_OF_INTEGERS")
db.command("sql", "CREATE PROPERTY SparseDoc.weights ARRAY_OF_FLOATS")
db.command(
"sql",
"""
CREATE INDEX ON SparseDoc (tokens, weights)
LSM_SPARSE_VECTOR
METADATA {"dimensions": 128}
""",
)
rows = db.query(
"sql",
"SELECT expand(`vector.sparseNeighbors`('SparseDoc[tokens,weights]', ?, ?, 5))",
jtypes.JArray(jtypes.JInt)([5]),
arcadedb.to_java_float_array([1.0]),
).to_list()
Weight precision: INT8 by default, FP32 on request¶
The sparse index stores posting weights quantized to 8 bits by default, which is what a search engine does and costs a small amount of recall. To keep exact 32-bit weights, say so in the index metadata:
CREATE INDEX ON SparseDoc (tokens, weights) LSM_SPARSE_VECTOR
METADATA {"dimensions": 30000, "weightQuantization": "FP32"}
Both forms answer vector.sparseNeighbors the same way. From 26.10.1 the INT8 index
only picks the candidates: it fetches k × rescoreOversample of them and ranks them by
the exact score computed from the records' own weights (ArcadeData/arcadedb#8576).
rescoreOversample is an index metadata key, 2 by default for INT8 and FP16 and off for
FP32; 0 turns it off. On 100,000 BigANN SPLADE documents that took INT8 recall@10 from
0.9948 to 1.0000, the same as FP32, and made the top-10 lists identical whatever the
commit size the index was loaded with, for about 16% more scorer time on Temurin 25
(43% on Temurin 21, laptop, 4 cores). So the choice is now on disk (FP32 about 20% larger
for a SPLADE-style corpus) and speed, not on the answers. Before 26.10.1, INT8 cost a few
tenths of a point of recall@10 and its answers depended on how the data was committed.
Leave the scorer's own settings at their defaults (posting block size 128,
arcadedb.sparseVectorScoringMaxPartitions=0); upstream has no other setting to
recommend (ArcadeData/arcadedb#8553).
The settle step: compact before you time queries¶
An LSM sparse index answers from every segment it has written until those segments are merged, and a freshly loaded index has many. Queries get faster once it is compacted, so compact after bulk loads and before benchmarks, the way you would force-merge a search engine:
The statement is synchronous and works embedded and over the server's HTTP
API alike. In-process there is also the Java handle,
db.get_java_database().getSchema().getIndexByName(...).compact(), which is
what the SQL form calls. At one million documents the compaction took about
two seconds and moved query p50 from 9.5 ms to 7.0 ms. From 26.10.1 it also flushes
the in-memory postings, so afterwards the whole index is one segment
(ArcadeData/arcadedb#8576); before, what was still in memory stayed there (422,190 of
12.7 million postings in a 100,000-document load committed 500 at a time).
Grouped Search¶
vector.neighbors accepts groupBy / groupSize options.
This is useful when you want diversity across a field such as source file, tenant, or
document family.
rows = db.query(
"sql",
(
"SELECT source_file, distance FROM "
"(SELECT expand(`vector.neighbors`(?, ?, ?, { groupBy: 'source_file', groupSize: 1 }))) "
"ORDER BY distance"
),
"GroupedDoc[embedding]",
arcadedb.to_java_float_array([1.0, 0.0, 0.0, 0.0]),
3,
).to_list()