Skip to content

Vector API

Vector search capabilities in ArcadeDB use JVector (a graph-based index combining HNSW and DiskANN concepts) for fast approximate nearest neighbor search. Perfect for semantic search, recommendation systems, and similarity-based queries.

Overview

ArcadeDB's vector support enables:

  • Semantic Search: Find similar documents, images, or any embedded content
  • Recommendation Systems: Find similar items based on feature vectors
  • Clustering: Group similar vectors together
  • Anomaly Detection: Find outliers in vector space

Key Features:

  • Graph-based indexing for O(log N) search performance
  • Multiple distance metrics (cosine, euclidean, dot product)
  • Native NumPy integration (optional)
  • Configurable precision/performance trade-offs

What happens to a write after the index is built

The index is a hybrid, and knowing which half answers your query explains most of its timing behaviour.

A vector written after the graph has been built does not patch the graph. It is queued into an in-memory delta buffer, and the graph is left alone. The graph changes only when a rebuild runs, and a rebuild always re-indexes the whole graph from scratch: there is no incremental patch.

Your write is searchable immediately. Every query scores the graph's results and scans the delta buffer exhaustively, merges the two, drops duplicates, and filters anything deleted. Because the buffer is scanned by brute force rather than traversed approximately, a vector sitting in it is found exactly, if anything more reliably than one already in the graph.

What you pay for that is a linear per-query cost proportional to the buffer's size. On a 5,000-vector index at 16 dimensions, 900 buffered entries cost about 5% on median query latency; the cost grows with both the number of buffered vectors and the dimensionality.

When the rebuild actually fires

Not every mutationsBeforeRebuild writes. The effective threshold scales with the index:

threshold = max(mutationsBeforeRebuild,
                min(graphSize * rebuildGraphRatio, maxPendingMutations))

At the defaults (100, 0.2, 50_000):

index size rebuild after roughly
1,000 vectors 200 mutations
100,000 vectors 20,000 mutations
1,000,000 vectors and above 50,000 mutations (the ceiling binds)

A fixed threshold would make bulk loading quadratic (rebuilding a 200,000-node graph every 100 inserts), which is why it scales. The practical consequences:

  • Raising mutationsBeforeRebuild alone changes nothing on any index above roughly 500 vectors, because it is only the floor. Use rebuildGraphRatio to change how the threshold scales, and maxPendingMutations to move the ceiling.
  • A lower ratio means fresher graph structure and more rebuild CPU; a higher one means a longer delta scan on every query. Neither affects correctness.
  • maxPendingMutations caps the threshold, not the buffer. Writes keep appending between rebuilds, and rebuilds are asynchronous, so a sustained ingest can outrun them and grow the buffer past that number.

get_stats() reports deltaVectorsCount (how much is buffered now) and graphRebuildCount (how many rebuilds have run), which is the direct way to see which half of the index is answering your queries. VectorIndex.build_graph_now() drains the buffer on demand instead of waiting for the threshold.

Module Functions

Utility functions for converting between Python and Java vector representations:

to_java_float_array(vector)

Convert a Python array-like object to a Java float array compatible with ArcadeDB's vector indexing.

Parameters:

  • vector: Array-like object containing float values
    • Python list: [0.1, 0.2, 0.3]
    • NumPy array: np.array([0.1, 0.2, 0.3])
    • Any iterable: (0.1, 0.2, 0.3)

Returns:

  • Java float array (JArray<JFloat>)

Example:

import arcadedb_embedded as arcadedb
from arcadedb_embedded import to_java_float_array
import numpy as np

# From Python list
vec_list = [0.1, 0.2, 0.3, 0.4]
java_vec = to_java_float_array(vec_list)

# From NumPy array (if NumPy installed)
vec_np = np.array([0.5, 0.6, 0.7, 0.8], dtype=np.float32)
java_vec = to_java_float_array(vec_np)

# Use with SQL parameter binding
with db.transaction():
    db.command(
        "sql",
        "INSERT INTO Document SET embedding = ?",
        java_vec,
    )

to_python_array(java_vector, use_numpy=True)

Convert a Java array or ArrayList to a Python array.

Parameters:

  • java_vector: Java array or ArrayList of floats
  • use_numpy (bool): Return NumPy array if available (default: True)
    • If True and NumPy is installed: returns np.ndarray
    • If False or NumPy unavailable: returns Python list

Returns:

  • np.ndarray (if use_numpy=True and NumPy available)
  • list (otherwise)

Example:

from arcadedb_embedded import to_python_array

# Get vector from vertex
vertex = result_set.first()
java_vec = vertex.get("embedding")

# Convert to NumPy array
np_vec = to_python_array(java_vec, use_numpy=True)
print(type(np_vec))  # <class 'numpy.ndarray'>

# Convert to Python list
py_list = to_python_array(java_vec, use_numpy=False)
print(type(py_list))  # <class 'list'>

to_java_int_array(vector)

Convert a Python array-like object to a Java int[].

The natural use is the token-index side of a sparse vector, whose weights go through to_java_float_array.

import numpy as np
from arcadedb_embedded import to_java_int_array, to_java_float_array

tokens = np.array([7, 91, 4096], dtype=np.int32)
weights = np.array([0.5, 0.25, 0.125], dtype=np.float32)

db.command(
    "sql",
    "INSERT INTO SparseDoc SET tokens = ?, weights = ?",
    to_java_int_array(tokens),
    to_java_float_array(weights),
)

Prefer a NumPy array over a Python list. JPype copies an array across the JVM boundary in one crossing through the buffer protocol; a list is marshalled element by element. Measured at 150 non-zeros: 2.6 us from an array against 6.7 us from a list, and the gap widens with length. The dtype does not matter (int64, NumPy's default, converts as fast as int32), so there is no reason to cast before calling.

Parameters:

  • vector: Array-like object of integers. Accepts a Python list, a tuple or any iterable, and a NumPy array of any integer dtype.

Raises: OverflowError when an element does not fit in 32 bits, for a list and for an integer NumPy array alike (an int64 array used to wrap silently: 2**31 became -2**31).

Returns: a Java int[].


to_java_byte_array(vector)

Convert a Python byte-like or integer array-like object to a Java byte[].

Use this when inserting native INT8 vectors into a BINARY property for indexes created with encoding="INT8".

from arcadedb_embedded import to_java_byte_array

payload = to_java_byte_array([127, 0, -12, 5])

VectorIndex Class

Wrapper for ArcadeDB's vector index, providing similarity search capabilities.

For creation, prefer SQL CREATE INDEX ... LSM_VECTOR METADATA {...}. For search, prefer SQL or Cypher when you need filtering, projection, self-exclusion, or custom score shaping. The Python methods below are convenience helpers for simple embedded-mode workflows and for advanced/manual control. They are not the primary application-facing workflow this documentation recommends.

For SQL snippets in this documentation, use vectorNeighbors(...) by default. The engine also exposes an equivalent dotted canonical function name, but the alias is more ergonomic in SQL because it does not require backticks.

Preferred Creation: SQL

Use SQL as the default creation surface:

db.command(
    "sql",
    """
    CREATE INDEX ON Document (embedding)
    LSM_VECTOR
    METADATA {
        "dimensions": 384,
        "similarity": "COSINE"
    }
    """
)

SQL builds the vector graph immediately by default. Add "buildGraphNow": false only if you intentionally want lazy preparation.

Creation via Database

Database.create_vector_index() still exists, but it should be treated as a secondary, Python-driven helper for manual setup, tests, and API completeness rather than the primary documented workflow.

Signature:

db.create_vector_index(
    vertex_type: str,
    vector_property: str,
    dimensions: int,
    id_property: str | None = None,
    distance_function: str = "cosine",
    max_connections: int = 32,
    beam_width: int = 100,
    quantization: str = "INT8",
    encoding: str | None = None,
    location_cache_size: int | None = None,  # removed: any value raises ValueError
    graph_build_cache_size: int | None = None,
    mutations_before_rebuild: int | None = None,
    store_vectors_in_graph: bool = False,
    add_hierarchy: bool | None = True,
    pq_subspaces: int | None = None,
    pq_clusters: int | None = None,
    pq_center_globally: bool | None = None,
    pq_training_limit: int | None = None,
    build_graph_now: bool = True,
) -> VectorIndex

Parameters:

  • vertex_type (str): Vertex type containing vectors
  • vector_property (str): Property name storing vector arrays
  • dimensions (int): Vector dimensionality (must match your embeddings)
  • id_property (str | None): Optional property used for key-based lookup with find_nearest_by_key(). Defaults to the engine default ("id") when omitted.
  • distance_function (str): Distance metric (default: "cosine")
    • "cosine": Cosine distance (1 - cosine similarity)
    • "euclidean": Squared Euclidean distance
    • "dot_product": -(1 + A·B) / 2, for unit-length vectors
  • max_connections (int): Per-layer graph degree (default: 32; Vamana degree, NOT doubled at the base layer like hnswlib M, so use 2*M to match an hnswlib config)
    • Maps to maxConnections in JVector
    • Higher = better recall, more memory
      • Typical range: 8-64
  • beam_width (int): Beam width for search/construction (default: 100)
    • Maps to beamWidth in JVector
    • Higher = better recall, slower search
      • Typical range: 50-500
  • quantization (str | None): "INT8" (recommended), "BINARY", "PRODUCT", or None (default: "INT8")
    • In current ArcadeDB engine builds, "PRODUCT" also requires enough indexed vectors per bucket for PQ training. For tiny corpora, set pq_clusters explicitly to a small value or prefer "INT8", "BINARY", or None.
    • Prefer "INT8" for current production usage in these bindings.
    • "PRODUCT"/PQ is available but currently not recommended for production workloads.
  • encoding (str | None): Optional storage encoding for the vector property.
    • Use "INT8" with a BINARY property when your vectors are already stored as signed bytes.
    • Pair encoding="INT8" with quantization="NONE" to avoid double quantization.
  • build_graph_now (bool): If True (default), eagerly prepares the vector graph during index creation. Set to False to defer graph preparation until first query.

Returns:

  • VectorIndex: Index object for searching

Example:

import arcadedb_embedded as arcadedb

# Create database and schema
db = arcadedb.create_database("./vector_db")

db.command("sql", "CREATE VERTEX TYPE Document")
db.command("sql", "CREATE PROPERTY Document.id STRING")
db.command("sql", "CREATE PROPERTY Document.text STRING")
db.command("sql", "CREATE PROPERTY Document.embedding ARRAY_OF_FLOATS")
db.command("sql", "CREATE INDEX ON Document (id) UNIQUE_HASH")

# Secondary option: create vector index from Python
index = db.create_vector_index(
    vertex_type="Document",
    vector_property="embedding",
    dimensions=384,  # Match your embedding model
    id_property="id",
    distance_function="cosine",
    max_connections=16,
    beam_width=100
)

print(f"Created vector index: {index}")

VectorIndex.find_nearest(query_vector, k=10, ef_search=None, allowed_rids=None)

Find k-nearest neighbors to the query vector.

Treat this as a helper/manual API. For normal application queries, prefer SQL vectorNeighbors so search composes naturally with filtering, projection, and record exclusion.

Note: With default settings (build_graph_now=True in create_vector_index), graph preparation runs during index creation. In the preferred SQL path, this eager behavior is also the default. If you explicitly disable eager preparation, the first call to find_nearest may perform lazy graph preparation and therefore take longer.

Parameters:

  • query_vector: Query vector as:
    • Python list: [0.1, 0.2, ...]
    • NumPy array: np.array([0.1, 0.2, ...])
    • Any array-like iterable
  • k (int): Number of neighbors to return (default: 10)
  • ef_search (int | None): Optional exact-search beam width override. None uses ArcadeDB's default/adaptive search behavior.
  • allowed_rids (List[str]): Optional list of RID strings (e.g. ["#1:0", "#2:5"]) to restrict search (default: None)

Returns:

  • List[Tuple[record, float]]: List of (record, distance) tuples
    • record: Matched ArcadeDB record object (Vertex, Document, or Edge)
    • distance: Distance score (float), lower is better for all functions
      • Cosine: cosine distance, range [0, 2] (0 = identical)
      • Euclidean: squared Euclidean distance, range [0, ∞) (0 = identical)
      • Dot product: -(1 + A·B) / 2 (lower = more similar)
    • Range depends on distance_function

Example:

# Generate query vector (in practice, from your embedding model)
query_text = "machine learning tutorial"
query_vector = generate_embedding(query_text)  # Your embedding function

# Search for 5 most similar documents
neighbors = index.find_nearest(query_vector, k=5)

# Search with RID filtering
allowed_rids = ["#10:5", "#10:8", "#10:12"]
filtered_neighbors = index.find_nearest(query_vector, k=5, allowed_rids=allowed_rids)

for record, distance in neighbors:
    doc_id = record.get("id")
    text = record.get("text")
    print(f"Distance: {distance:.4f} | ID: {doc_id}")
    print(f"  Text: {text[:100]}...")

Preferred for richer query behavior:

from arcadedb_embedded import to_java_float_array

rows = db.query(
    "sql",
    (
        "SELECT id, distance, (1 - distance) AS score "
        "FROM (SELECT expand(vectorNeighbors('Document[embedding]', ?, 10))) "
        "WHERE id <> ? ORDER BY distance LIMIT 5"
    ),
    to_java_float_array(query_vector),
    "doc-42",
).to_list()

Distance Interpretation:

Function Range Similarity direction
cosine [0, 2] lower is better (0 = identical)
euclidean [0, ∞) lower is better (0 = identical, squared distance)
dot_product [-1, 0] for unit vectors lower is better (-(1 + A·B) / 2)
  • Vertex must have the vector property populated
  • Vector dimensionality must match index dimensions
  • A search is a read: it needs no transaction, and runs better outside one

VectorIndex.find_nearest_by_key(key, k=10, ef_search=None, allowed_rids=None)

Find nearest neighbors by reusing the vector stored on an existing record.

This is the Python wrapper for the common "search from an existing record" workflow, using the index's configured id_property to look up the source vector first.

Treat this as a convenience helper. If you need the recommended query surface, use SQL or Cypher instead.

Parameters:

  • key: Value of the configured ID property
  • k (int): Number of neighbors to return (default: 10)
  • ef_search (int | None): Optional exact-search beam width override
  • allowed_rids (List[str] | None): Optional RID whitelist to restrict search

Returns:

  • List[Tuple[record, float]]: Same shape as find_nearest()

Example:

neighbors = index.find_nearest_by_key("doc-42", k=5)

for record, distance in neighbors:
    print(record.get("id"), distance)

The helper keeps current nearest-neighbor semantics, so the source record may also be returned. If you want to exclude it, do that in SQL/Cypher with a WHERE clause.


VectorIndex.find_nearest_approximate(query_vector, k=10, allowed_rids=None)

Find k nearest neighbors using Product Quantization (PQ) approximate search.

Only available on indexes created with quantization="PRODUCT"; calling it on any other index raises ArcadeDBError.

Parameters:

  • query_vector: Query vector as Python list, NumPy array, or array-like
  • k (int): Number of nearest neighbors to return (default: 10)
  • allowed_rids (List[str] | None): Optional list of RID strings (e.g. ["#1:0", "#2:5"]) to restrict search (default: None)

Returns:

  • List[Tuple[record, float]]: Same (record, score) shape as find_nearest()

Raises:

  • ArcadeDBError: If the index quantization is not PRODUCT, or the search fails

Example:

index = db.create_vector_index(
    "Document", "embedding", dimensions=384, quantization="PRODUCT"
)

neighbors = index.find_nearest_approximate(query_vector, k=5)
for record, score in neighbors:
    print(record.get("id"), score)

VectorIndex.get_size()

Get the current number of items in the index.

Returns:

  • int: Number of items currently indexed

Raises:

  • ArcadeDBError: If the size cannot be read

Example:

print(f"Indexed vectors: {index.get_size()}")

VectorIndex.get_quantization()

Get the quantization type of the index.

Returns:

  • str: "NONE", "INT8", "BINARY", or "PRODUCT" (returns "NONE" if the quantization cannot be determined)

Example:

if index.get_quantization() == "PRODUCT":
    results = index.find_nearest_approximate(query_vector, k=10)
else:
    results = index.find_nearest(query_vector, k=10)

VectorIndex.get_metadata()

Return stable vector index metadata as a Python dictionary.

Returns:

  • dict with keys such as:
    • index_name
    • bucket_index_name
    • type_name
    • vector_property
    • dimensions
    • similarity_function
    • id_property
    • quantization
    • max_connections
    • beam_width
    • store_vectors_in_graph
    • build_state

Example:

meta = index.get_metadata()
print(meta["dimensions"], meta["similarity_function"], meta["id_property"])

VectorIndex.get_stats()

Return live runtime counters for the index as a Python dictionary.

Where get_metadata() returns the index's stable configuration, this returns counters that move as the index is built and queried. Keys follow the engine and may grow between releases, so read defensively.

Returns:

  • dict mapping counter name to a plain Python scalar, for the primary bucket index (see get_metadata()["bucket_index_name"]). Commonly useful keys:
    • totalVectors, activeVectors, deletedVectors, graphNodeCount
    • graphState, asyncRebuildInProgress, graphRebuildCount
    • searchVectorCacheCapacity, vectorCacheHits, vectorCacheMisses
    • vectorFetchFromDocuments, vectorFetchFromGraph, vectorFetchFromQuantized
    • searchOperations, insertOperations, avgSearchLatencyMs

Example:

stats = index.get_stats()

# Is the search cache holding the working set?
hits, misses = stats["vectorCacheHits"], stats["vectorCacheMisses"]
print(f"hit rate: {hits / max(hits + misses, 1):.1%}")

# Did the build have to re-read vectors from documents?
if stats["vectorFetchFromDocuments"] > 0:
    print("build cache did not hold the whole set; raise the heap or its share")

Where vectors are read from. With the default storeVectorsInGraph: false, the document is the only copy of an unquantized vector, so a cache miss costs a document read (vectorFetchFromDocuments). Quantized indexes keep codes in the index pages (vectorFetchFromQuantized), which is why INT8 tolerates a small cache far better than fp32 does.


VectorIndex.warm_up()

Load the index's graph now instead of on the first search. From 26.10.1.

After a database is opened, a vector index loads its persisted graph lazily, on the first search, and that search pays for it (ArcadeData/arcadedb#8852). After the 26.10.1 fix that is about 1 s at 1M 64-dimension vectors, growing with the index, and more on an index with deletions since its graph was saved, which still re-reads every vector's document. A service that restarts can call warm_up() right after opening the database, so the first user query does not pay it.

It loads the existing graph and does not rebuild it (that is build_graph_now()). It is a no-op once the graph is in memory, and safe to call while other threads search. On an engine before 26.10.1 it raises ArcadeDBError; there, one throwaway search does the same.

get_stats()["graphState"] reads 0 (loading) between the open and the first search or warm_up(), and graphNodeCount is 0 until the graph is resident.

Returns:

  • None

Example:

with arcadedb.open_database("./vectors") as db:
    index = db.schema.get_vector_index("Doc", "embedding")
    index.warm_up()  # pay the graph load now, not on the first query

# On engines before 26.10.1, the same with one throwaway search. The vector is
# wrapped in a list: a lone list argument is read as the parameter list itself.
#   db.query("sql", "SELECT vectorNeighbors('Doc[embedding]', ?, 1)", [probe_vector])

VectorIndex.build_graph_now()

Force an immediate rebuild/preparation of the vector graph.

This is a maintenance API, not part of the normal SQL-first creation workflow.

Use this when you want to control when rebuild cost is paid, for example:

  • after bulk inserts,
  • after bulk deletes/removals,
  • before opening traffic after large vector mutations.

This is especially useful if you created the index with build_graph_now=False or with SQL metadata "buildGraphNow": false and want to avoid rebuild work on the first query.

When you create an LSM_VECTOR index through SQL, ArcadeDB now builds the graph immediately by default. Use build_graph_now() only for explicit maintenance or when you intentionally deferred graph preparation.

Returns:

  • None

Example:

# After large vector changes, force rebuild now
index.build_graph_now()

Complete Examples

Semantic Search with Sentence Transformers

import arcadedb_embedded as arcadedb
from arcadedb_embedded import to_java_float_array
from sentence_transformers import SentenceTransformer
import numpy as np

# Load embedding model
model = SentenceTransformer('all-MiniLM-L6-v2')  # 384 dimensions

# Create database and schema
db = arcadedb.create_database("./semantic_search")

db.command("sql", "CREATE VERTEX TYPE Document")
db.command("sql", "CREATE PROPERTY Document.id STRING")
db.command("sql", "CREATE PROPERTY Document.title STRING")
db.command("sql", "CREATE PROPERTY Document.content STRING")
db.command("sql", "CREATE PROPERTY Document.embedding ARRAY_OF_FLOATS")
db.command("sql", "CREATE INDEX ON Document (id) UNIQUE_HASH")

# Preferred: create vector index in SQL
db.command(
    "sql",
    """
    CREATE INDEX ON Document (embedding)
    LSM_VECTOR
    METADATA {
        "dimensions": 384,
        "similarity": "COSINE",
        "maxConnections": 32,
        "beamWidth": 100
    }
    """
)

index = db.schema.get_vector_index("Document", "embedding")

# Sample documents
documents = [
    {"id": "doc1", "title": "Python Tutorial",
        "content": "Learn Python programming basics"},
    {"id": "doc2", "title": "Machine Learning Guide",
        "content": "Introduction to ML algorithms"},
    {"id": "doc3", "title": "Database Systems",
        "content": "Understanding relational databases"},
]

# Index documents
print("Indexing documents...")
with db.transaction():
    for doc in documents:
        # Generate embedding
        text = f"{doc['title']} {doc['content']}"
        embedding = model.encode(text)

        db.command(
            "sql",
            "INSERT INTO Document SET id = ?, title = ?, content = ?, embedding = ?",
            doc["id"],
            doc["title"],
            doc["content"],
            to_java_float_array(embedding),
        )

print(f"Indexed {len(documents)} documents")

# Search
query = "How to learn programming"
query_embedding = model.encode(query)

print(f"\nQuery: '{query}'")
results = index.find_nearest(query_embedding, k=3)

for vertex, distance in results:
    print(f"\nDistance: {distance:.4f}")
    print(f"Title: {vertex.get('title')}")
    print(f"Content: {vertex.get('content')}")

db.close()

Hybrid Search (Vector + Filters)

Combine vector similarity with property filters using SQL:

import arcadedb_embedded as arcadedb
from arcadedb_embedded import to_java_float_array
import numpy as np

db = arcadedb.open_database("./products_db")

# Create schema
db.command("sql", "CREATE VERTEX TYPE Product")
db.command("sql", "CREATE PROPERTY Product.id STRING")
db.command("sql", "CREATE PROPERTY Product.name STRING")
db.command("sql", "CREATE PROPERTY Product.category STRING")
db.command("sql", "CREATE PROPERTY Product.price DECIMAL")
db.command("sql", "CREATE PROPERTY Product.features ARRAY_OF_FLOATS")
db.command("sql", "CREATE INDEX ON Product (category) NOTUNIQUE")

# Create vector index in SQL
db.command(
    "sql",
    """
    CREATE INDEX ON Product (features)
    LSM_VECTOR
    METADATA {
        "dimensions": 128,
        "similarity": "COSINE"
    }
    """
)

index = db.schema.get_vector_index("Product", "features")

# Add products with feature vectors
products = [
    {"id": "p1", "name": "Laptop", "category": "Electronics",
        "price": 999.99, "features": np.random.rand(128)},
    {"id": "p2", "name": "Mouse", "category": "Electronics",
        "price": 29.99, "features": np.random.rand(128)},
    {"id": "p3", "name": "Desk", "category": "Furniture",
        "price": 299.99, "features": np.random.rand(128)},
]

with db.transaction():
    for prod in products:
        db.command(
            "sql",
            "INSERT INTO Product SET id = ?, name = ?, category = ?, price = ?, features = ?",
            prod["id"],
            prod["name"],
            prod["category"],
            prod["price"],
            to_java_float_array(prod["features"]),
        )
        # Note: LSM vector index automatically indexes new records

# Hybrid search: vector similarity + filters
query_features = np.random.rand(128)
candidates = index.find_nearest(query_features, k=100)  # Get many candidates

# Filter by category and price
filtered_results = []
for vertex, distance in candidates:
    category = vertex.get("category")
    price = float(vertex.get("price"))

    if category == "Electronics" and price < 500:
        filtered_results.append((vertex, distance))

    if len(filtered_results) >= 5:  # Want top 5 after filtering
        break

print("Filtered Results:")
for vertex, distance in filtered_results:
    print(f"{vertex.get('name')} - ${vertex.get('price')} - {distance:.4f}")

db.close()

import arcadedb_embedded as arcadedb
from arcadedb_embedded import to_java_float_array
from PIL import Image
import numpy as np

# Assuming you have a function to generate image embeddings
def get_image_embedding(image_path):
    """
    Generate embedding for image using your model
    (e.g., ResNet, CLIP, etc.)
    """
    # Placeholder - use your actual embedding model
    return np.random.rand(512)  # Example: 512-dim embedding

db = arcadedb.create_database("./image_search")

# Schema
db.command("sql", "CREATE VERTEX TYPE Image")
db.command("sql", "CREATE PROPERTY Image.id STRING")
db.command("sql", "CREATE PROPERTY Image.filename STRING")
db.command("sql", "CREATE PROPERTY Image.path STRING")
db.command("sql", "CREATE PROPERTY Image.embedding ARRAY_OF_FLOATS")

# Create index in SQL
db.command(
    "sql",
    """
    CREATE INDEX ON Image (embedding)
    LSM_VECTOR
    METADATA {
        "dimensions": 512,
        "similarity": "COSINE",
        "maxConnections": 24,
        "beamWidth": 200
    }
    """
)

index = db.schema.get_vector_index("Image", "embedding")

# Index images
image_files = ["img1.jpg", "img2.jpg", "img3.jpg"]

with db.transaction():
    for idx, img_file in enumerate(image_files):
        embedding = get_image_embedding(img_file)

        v = db.new_vertex("Image")
        v.set("id", f"img_{idx}")
        v.set("filename", img_file)
        v.set("path", f"/images/{img_file}")
        v.set("embedding", to_java_float_array(embedding))
        v.save()

        # Note: LSM vector index automatically indexes new records

# Search for similar images
query_image = "query.jpg"
query_embedding = get_image_embedding(query_image)

similar_images = index.find_nearest(query_embedding, k=5)

print(f"Similar images to {query_image}:")
for vertex, distance in similar_images:
    print(f"  {vertex.get('filename')} - similarity: {1 - distance:.4f}")

db.close()

Performance Tuning

Vector Index Parameters

max_connections (connections per node):

  • Lower (16): Faster build, less memory, lower recall
  • Medium (32): Balanced (default)
  • Higher (64): Better recall, more memory, slower build

ef_search (exact search beam width):

  • Unset (None): Use ArcadeDB's default/adaptive behavior
  • Lower (32): Faster search, lower recall
  • Medium (100): Balanced explicit override
  • Higher (200): Better recall, slower search

beam_width:

  • Lower (64): Faster build, lower quality
  • Medium (100): Balanced (default)
  • Higher (200): Better quality, slower build

Distance Functions

Cosine Distance:

  • Best for: Text embeddings, normalized vectors
  • Range: [0, 2], lower is better
  • Use when: Direction matters more than magnitude

Euclidean Distance:

  • Best for: Image embeddings, spatial data
  • Range: [0, ∞), lower is better
  • Use when: Absolute distance matters

Dot Product:

  • Best for: unit-length vectors, which it ranks the same way as cosine
  • Range: [-1, 0] for unit vectors, lower is better (the score is -(1 + A·B) / 2)
  • The engine expects unit-length vectors: when sampled vectors are not, it logs a warning that search quality is degraded. Normalize on ingest, or use cosine

Memory Considerations

Approximate memory per vertex:

memory_per_vertex = dimensions * 4 bytes + max_connections * 8 bytes + overhead

Example for 384-dim vectors with max_connections=16:

384 * 4 + 16 * 8 + ~100 bytes ≈ 1.8 KB per vertex

For 1 million vectors: ~1.8 GB RAM

Vector Caches

Two caches dominate large-index behavior. Both size themselves automatically from the index size and the heap, and both are settable as global configuration (for example through jvm_kwargs={"jvm_args": "-Darcadedb...=..."} before the first database is opened):

Setting Default Effect
arcadedb.vectorIndex.graphBuildCacheSize 0 (automatic) Vectors held in RAM while the graph is built. Too small and the build re-reads vectors from documents, which is visible as a non-zero vectorFetchFromDocuments in get_stats().
arcadedb.vectorIndex.graphBuildCacheMaxHeapPercent 25 Share of the heap the automatic build-cache sizing may use.
arcadedb.vectorIndex.searchCacheSize 0 (automatic) Vectors kept warm across queries. -1 disables it.
arcadedb.vectorIndex.searchCacheMaxHeapPercent 25 Share of the heap the automatic search-cache sizing may use.

Leave both at their defaults. Automatic build-cache sizing reads the heap the engine actually has free and caches the whole corpus when it fits, which is what a build wants; engines before 26.10 read a post-GC figure that included the evictable page cache and could settle on a fraction of a large corpus on a large heap (upstream #7146, fixed in #7147). Set an absolute size only to bound a build on a deliberately small heap.

The search cache is per index and stays warm between queries, so the first queries after a fresh build pay a cold-start cost while it fills; check vectorCacheHits against vectorCacheMisses to see when it has settled. Sizing headroom matters: a set of 10M 96-dimensional fp32 vectors is roughly 4.5 GB, so it only stays resident if 25% of the heap can hold it.


Error Handling

A vector whose length differs from the index's dimensions is refused by save(), which raises the Java java.lang.IllegalArgumentException ("Vector dimension does not match index dimension"), not ArcadeDBError. A record saved without the vector property raises nothing; it is not in the index.

from arcadedb_embedded import to_java_float_array
import numpy as np

db.command(
    "sql",
    'CREATE INDEX ON Doc (emb) LSM_VECTOR METADATA {"dimensions": 384}',
)

try:
    with db.transaction():
        v = db.new_vertex("Doc")
        v.set("emb", to_java_float_array(np.random.rand(512)))  # Wrong size!
        v.save()  # raises here; the transaction rolls back
except Exception as e:  # java.lang.IllegalArgumentException
    print(f"Error: {e}")

# Missing vector property: saved, not indexed, no error
with db.transaction():
    v = db.new_vertex("Doc")
    v.set("id", "doc1")
    v.save()

See Also