Vector Search Tests (JVector / LSM)¶
The test suite exercises the current Java-native JVector + LSM vector index used by ArcadeDB (no Python hnswlib dependency). All tests run through the Python bindings.
The array-helper tests are module-level functions; the rest are in
TestLSMVectorIndex.
Overview¶
What the tests cover:
- ✅ Array helpers:
to_java_byte_array,to_java_float_array, andto_java_int_array - ✅ JVector
LSM_VECTORindex creation throughdb.create_vector_index(), including PQ, id-property, and graph-storage options - ✅ Nearest-neighbor search with
find_nearest,find_nearest_by_key, andfind_nearest_approximate - ✅ RID filtering using
allowed_rids - ✅ Search beam tuning (
ef_search) - ✅ Distance functions (cosine default, euclidean variants)
- ✅ Quantization (
INT8,BINARY,PRODUCT) - ✅ Persistence, size, and stats (the index survives reopen)
Test Cases¶
Array helpers (module level)¶
- test_to_java_byte_array_accepts_bytes_like_fast_paths:
bytes,bytearray, and NumPyint8/uint8arrays convert to Javabyte[]with signed values (255becomes-1). - test_to_java_float_array_accepts_numpy_directly: NumPy
float32andfloat64arrays convert to[1.0, 2.0, 3.0]. - test_to_java_int_array: lists, tuples,
range, an empty list, and values up to2147483647convert unchanged. - test_to_java_int_array_accepts_numpy_directly: NumPy arrays of each integer dtype convert to
[1, 2, 3]. - test_to_java_int_array_round_trips_through_a_sparse_vector: tokens stored from a NumPy array read back as
[7, 91, 4096], and the weights as[0.5, 0.25, 0.125].
Index creation¶
- test_create_vector_index_build_graph_now_default_true:
create_vector_index()callsVectorIndex.build_graph_now()once by default. - test_create_vector_index_build_graph_now_can_be_disabled: with
build_graph_now=Falseit is not called. - test_create_vector_index: returns a
VectorIndex, andschema:indexeslists an index onDocwhose name starts withDoc. - test_create_vector_index_with_pq_params: with
quantization="PRODUCT", the PQ arguments reach the metadata (pqSubspaces == 4,pqClusters == 128,pqCenterGloballyfalse,pqTrainingLimit == 5000). - test_create_vector_index_with_custom_id_property:
id_property="slug"reaches the metadata asidPropertyName. - test_create_vector_index_with_graph_storage:
store_vectors_in_graph=Truecreates an index. - test_graph_storage_with_quantization: graph storage plus
BINARYquantization creates an index whoseget_quantization()isBINARY. - test_get_vector_index_lsm:
SELECT name FROM schema:indexes WHERE typeName = 'Doc'lists an index onDoc. - test_lsm_vector_metadata:
get_metadata()returns the index and bucket names, type, property, dimensions, similarity, id property, quantization, cache settings, and flags as given, and has nolocation_cache_sizekey.
Search¶
- test_lsm_vector_search:
find_nearest(k=1)returns one result, the vector closest to the query. - test_lsm_vector_search_by_key:
find_nearest_by_key("doc-a", k=2)returnsdoc-afirst anddoc-bsecond. - test_lsm_vector_search_by_key_missing_record_raises: a key with no record raises
ArcadeDBError("No record found"). - test_lsm_vector_search_with_filter: with
allowed_rids, results come only from the allowed records, in distance order, for four different allow-lists. - test_lsm_vector_search_uses_database_lookup_by_rid: the two results are materialized through the
DatabaseRID lookup (called twice). - test_lsm_vector_search_uses_database_rid_conversion: the
allowed_ridsstrings are converted through theDatabasewrapper, in order. - test_lsm_vector_search_ef_search:
find_nearest(k=2, ef_search=32)returns 2 results with the closest vector first. - test_lsm_vector_search_rejects_invalid_ef_search:
ef_search=0raisesArcadeDBErrormentioningef_search. - test_lsm_vector_build_graph_now: after
build_graph_now(), a search returns the closest vector. - test_warm_up_loads_the_persisted_graph_before_the_first_search: 2,000 vectors, index built, database reopened:
get_stats()readsgraphState0 andgraphNodeCount0; afterwarm_up()the state is not 0 and all 2,000 nodes are resident; a secondwarm_up()changes nothing, and the search for a stored vector returns it at distance 0.0. Fails on 26.9.1, which has nowarm_up()(ArcadeData/arcadedb#8852). - test_lsm_vector_delete_and_search_others: of 100 random vectors, every tenth is deleted; each deleted vector is absent from its own search and each remaining one is found.
- test_lsm_vector_search_comprehensive: on small word embeddings, the neighbours of
kingincludequeenandmanorwomanbut notcatordog, and the neighbours ofcatincludedogbut notking. - test_document_vector_search: search on a document type returns
appleandbanana(notcar) for one query andcarandtruck(notapple) for another; results areMyDocrecords with 4-dimensional embeddings.
Approximate search (PRODUCT quantization)¶
- test_lsm_vector_search_approximate_product: on a
PRODUCTindex (aTypeIndexwrapper),find_nearest_approximate(k=1)returns the closest vector. - test_lsm_vector_search_approximate_typeindex: the same through the
TypeIndexwrapper path (asserts the build returns one). - test_lsm_vector_search_approximate_fallback: on an index without
PRODUCTquantization,find_nearest_approximate()raisesArcadeDBError(despite the test name, it asserts the error). - test_lsm_vector_search_approximate_product_requires_enough_vectors: a
PRODUCTindex on too few vectors raisesArcadeDBErrormentioningpq_clusters. - test_lsm_vector_search_approximate_returns_k: with 261 vectors (5 hand-picked plus 256 fillers),
find_nearest_approximate(k=2)returns exactly 2 results, each with a record and a distance. - test_lsm_vector_search_approximate_persistence: after closing and reopening,
vectorNeighbors(...)returns one neighbour with a record and a distance (it polls up to 50 times while the graph loads).
Size, stats, and persistence¶
- test_lsm_index_size:
get_size()is 0 before inserts and 2 after two inserts. - test_lsm_index_stats:
get_stats()is a non-empty dict of scalar values;totalVectorsis 0 before inserts and 2 after. - test_lsm_persistence: after a reopen, the type still has one record and
vectorNeighborsreturns one neighbour.
Distance functions¶
- test_lsm_cosine_distance_orthogonal_vectors: orthogonal vectors are at cosine distance 1.0.
- test_lsm_cosine_distance_parallel_vectors: parallel vectors are at distance below 0.01.
- test_lsm_cosine_distance_opposite_vectors: opposite vectors are at distance 2.0.
- test_lsm_cosine_distance_45_degree_vectors: vectors 45° apart are at the expected cosine distance.
- test_lsm_cosine_distance_3d_orthogonal_vectors: orthogonal 3-D vectors are at distance 1.0.
- test_lsm_cosine_distance_3d_parallel_and_opposite: parallel 3-D vectors are at distance below 0.01, opposite ones at 2.0.
- test_lsm_cosine_distance_high_dimensional: in 128 dimensions, parallel, opposite, and near-orthogonal vectors are at about 0, 2.0, and 1.0 (each check runs only for a vector that comes back; orthogonal tolerance 0.1; needs NumPy through
importorskip). - test_lsm_euclidean_distance: with
EUCLIDEAN, the reported distances are 0.0 for the origin and 25.0 for the point (3, 4) (the squared Euclidean distance from the origin), and the origin comes back first.
Quantization¶
- test_int8_quantization_boundary_condition: an
INT8index with 10 vectors of dimension 16 returns one result fork=1. - test_lsm_vector_quantization_int8_comprehensive: an
INT8index reportsget_quantization() == "INT8", and the top result's first component is above 0.9 at a distance below 0.2. - test_lsm_vector_quantization_binary_comprehensive: a
BINARYindex reportsget_quantization() == "BINARY"and returns one result with a full-length vector and a distance.
SQL Vector Functions Tests¶
SQL vector operations are tested separately in test_vector_sql.py, including vector math functions, distance calculations, aggregations, quantization (the INT8 SQL test does not check the returned vector's values), and SQL-based index creation and search.
Common Patterns¶
Create JVector (LSM-backed) index¶
with arcadedb.create_database("./test_db") as db:
db.command("sql", "CREATE VERTEX TYPE Doc")
db.command("sql", "CREATE PROPERTY Doc.embedding ARRAY_OF_FLOATS")
db.command(
"sql",
'''
CREATE INDEX ON Doc (embedding)
LSM_VECTOR
METADATA {
"dimensions": 384,
"similarity": "COSINE",
"maxConnections": 32,
"beamWidth": 100
}
''',
)
Search with filters and ef_search¶
with arcadedb.create_database("./test_db") as db:
db.command("sql", "CREATE VERTEX TYPE Doc")
db.command("sql", "CREATE PROPERTY Doc.docId INTEGER")
db.command("sql", "CREATE PROPERTY Doc.embedding ARRAY_OF_FLOATS")
db.command(
"sql",
'CREATE INDEX ON Doc (embedding) LSM_VECTOR METADATA {"dimensions": 3}',
)
# Insert test vertices with embeddings
with db.transaction():
doc1 = db.new_vertex("Doc")
doc1.set("docId", 1)
doc1.set("embedding", arcadedb.to_java_float_array([1.0, 0.0, 0.0]))
doc1.save()
doc2 = db.new_vertex("Doc")
doc2.set("docId", 2)
doc2.set("embedding", arcadedb.to_java_float_array([0.0, 1.0, 0.0]))
doc2.save()
# Search with filters (the WHERE goes on an outer query over the expanded
# neighbours; on the expand() query itself it matches nothing)
query = [1.0, 0.0, 0.0]
allowed_rids_sql = f"['{doc1.get_rid()}', '{doc2.get_rid()}']"
query_literal = "[" + ", ".join(str(float(v)) for v in query) + "]"
results = db.query(
"sql",
(
"SELECT FROM (SELECT expand(vectorNeighbors('Doc[embedding]', "
f"{query_literal}, 2, 100))) WHERE @rid IN {allowed_rids_sql}"
),
).to_list()
Chunked insert vectors (preferred)¶
import numpy as np
with arcadedb.create_database("./test_db") as db:
db.command("sql", "CREATE VERTEX TYPE Doc")
db.command("sql", "CREATE PROPERTY Doc.docId INTEGER")
db.command("sql", "CREATE PROPERTY Doc.embedding ARRAY_OF_FLOATS")
# Prefer chunked transactions for embedded (avoids batch_context overhead)
vectors = np.array(
[[1.0, 0.0, 0.0], [0.0, 1.0, 0.0], [0.0, 0.0, 1.0]], dtype=np.float32
)
chunk_size = 100
for start in range(0, len(vectors), chunk_size):
with db.transaction():
for idx, vec in enumerate(vectors[start : start + chunk_size]):
doc = db.new_vertex("Doc")
doc.set("docId", start + idx)
# One JVM crossing per vector; a Python list crosses element by element
doc.set("embedding", arcadedb.to_java_float_array(vec))
doc.save()
Key Takeaways¶
- JVector is fully Java-native and LSM-backed; no legacy hnswlib path remains.
max_connectionsandbeam_widthmap to JVector graph degree and search beam; tune per workload.- Prefer chunked
db.transaction()inserts for embedded workloads rather than a separate batching abstraction.
See Also¶
- Vector API – Full Python API reference
- NumPy Tests – NumPy integration
- Example 03: Vector Search – End-to-end usage
- Example 06: Movie Recommendations – Vector-powered recommender
- Vector Guide – Concepts and tuning