Skip to content

Example 06: Vector Search Movie Recommendations

View source code

Production-ready vector embeddings and HNSW (JVector) indexing for semantic movie search

Overview

This example demonstrates creating vector embeddings for movies and using HNSW (JVector) indexing for fast semantic similarity search. You'll learn:

  • Real embeddings - Using sentence-transformers for 384-dimensional vectors
  • HNSW (JVector) vector indexing - Fast approximate nearest neighbor search (cosine similarity)
  • Graph vs Vector - Compare collaborative filtering with semantic similarity
  • Multi-model comparison - Two embedding models with different semantic characteristics
  • Performance optimization - Graph query sampling, which makes collaborative filtering fast enough for real-time use on the large dataset

What You'll Learn

  • Generate embeddings using sentence-transformers (all-MiniLM-L6-v2, paraphrase-MiniLM-L6-v2)
  • Create and populate HNSW (JVector) vector indexes on Movie embedding properties
  • Graph-based collaborative filtering (full vs sampled modes)
  • Vector-based semantic similarity search
  • Performance comparison: 4 recommendation methods
  • Handling vector metadata persistence and property versioning

Prerequisites

1. Install dependencies:

pip install arcadedb-embedded sentence-transformers numpy

2. Source database:

This example requires a graph database from Example 05:

# Option A: Use existing database
python 05_csv_import_graph.py --dataset movielens-small --method java

# Option B: Import from JSONL export
python 05_csv_import_graph.py --dataset movielens-small --import-jsonl ./exports/movielens_small_db.jsonl.tgz

Two dataset sizes available:

  • movielens-small: 9,742 movies, ~100K ratings - Quick testing
  • movielens-large: 86,537 movies, ~33M ratings - Production testing

Usage

# Recommended: Import from JSONL export (fresh working database)
python 06_vector_search_recommendations.py \
    --import-jsonl ./exports/movielens_graph_small_db.jsonl.tgz \
    --db-path my_test_databases/movielens_vector_db

# Copy from an existing graph database (created by Example 05)
python 06_vector_search_recommendations.py \
    --source-db my_test_databases/movielens_graph_small_db \
    --db-path my_test_databases/movielens_vector_db

# Reuse an already-built working database at --db-path (no --import-jsonl/--source-db)
python 06_vector_search_recommendations.py \
    --db-path my_test_databases/movielens_vector_db

# Force re-generation of embeddings
python 06_vector_search_recommendations.py \
    --source-db my_test_databases/movielens_graph_small_db \
    --db-path my_test_databases/movielens_vector_db \
    --force-embed

# See all options
python 06_vector_search_recommendations.py --help

Key options:

  • --db-path DB_PATH - Working database path (default: ./my_test_databases/movielens_vector_db)
  • --import-jsonl IMPORT_JSONL - Create the working DB by importing this JSONL export
  • --source-db SOURCE_DB - Create the working DB by copying this existing graph database
  • --heap-size SIZE - JVM max heap size (e.g. 8g, 4096m)
  • --force-embed - Force re-generation of embeddings
  • --limit LIMIT - Limit number of movies to embed (for debugging)

Provide --import-jsonl or --source-db to (re)create the working database fresh; with neither, the script reuses an existing database at --db-path (and errors if none exists).

Recommendations:

  • Setup: Use fresh copy or import from JSONL to avoid conflicts
  • Memory: 8GB JVM heap for large dataset (--heap-size 8g)
  • Embeddings: Cached automatically, use --force-embed to regenerate
  • Models: Both models included for comparison

Vector Embedding Models

Model 1: all-MiniLM-L6-v2

  • Dimensions: 384
  • Best for: General-purpose semantic similarity

Model 2: paraphrase-MiniLM-L6-v2

  • Dimensions: 384
  • Best for: Paraphrase detection and semantic similarity

The script times encoding and index creation for each model and prints the results.

Recommendation Methods

1. Graph-Based Full (Collaborative Filtering - Comprehensive)

How it works:

  • Finds users who rated the query movie highly (≥4.0 stars)
  • Analyzes ALL movies those users also rated highly
  • Aggregates ratings and recommends top movies

Best for: Offline batch recommendations

Pros:

  • Most thorough analysis
  • High-quality recommendations based on user behavior

Cons:

  • Slow on large datasets (processes 100K+ intermediate results)
  • Cold start problem (new movies without ratings)

2. Graph-Based Fast (Collaborative Filtering - Sampled)

How it works:

  • Same as full mode, but limits intermediate results to 25K
  • Samples ~50 users' worth of ratings
  • Uses nested SELECT with LIMIT before GROUP BY aggregation

Best for: Real-time recommendations on large datasets, where full mode is slow

Pros:

  • Real-time recommendation speed
  • Still produces high-quality results
  • No cold start problem for existing movies

Cons:

  • Slightly less comprehensive than full mode
  • Still requires some ratings history

3. Vector (all-MiniLM-L6-v2)

How it works:

  • Encodes movie titles and genres into 384-dimensional vectors
  • Uses HNSW (JVector) index for fast approximate nearest neighbor search
  • Finds movies with similar semantic meaning

Pros:

  • Fast: one index lookup per query
  • No cold start problem (works for new movies)
  • Finds semantically similar content

Cons:

  • Different recommendations than collaborative filtering
  • Requires embedding generation and indexing

4. Vector (paraphrase-MiniLM-L6-v2)

How it works:

  • Same as Vector method 1, but with different embedding model
  • Optimized for paraphrase detection

Pros:

  • Fast: one index lookup per query
  • Different semantic characteristics than Model 1

Cons:

  • May find different movies than collaborative filtering
  • Embedding quality depends on model training

JVector Index Configuration

db.command(
    "sql",
    """
    CREATE INDEX ON Movie (embedding_v1)
    LSM_VECTOR
    METADATA {
        "dimensions": 384,
        "similarity": "COSINE"
    }
    """
)

Key parts of the statement:

  • Movie (embedding_v1): The vertex type and embedding property to index (the script also builds a second index on embedding_v2).
  • LSM_VECTOR: The vector index type (an HNSW-style JVector graph index).
  • dimensions: Vector dimensionality (384 for the sentence-transformers models used).
  • similarity: "COSINE" for cosine distance (lower distance is better).

create_sql_vector_index() also describes the underlying JVector graph build, printing something like metric=COSINE, max_connections=32, beam_width=100 (engine defaults). Those values are read back from the created index rather than hardcoded, because the METADATA above sets only dimensions and similarity and everything else comes from the engine's defaults. If a future release changes a default, the example reports the new one instead of quietly printing a stale number. The Python object API can still create indexes, but SQL is the cleaner default and is what this example uses.

Output

The script outputs:

  1. Dependency check - Verifies sentence-transformers is installed
  2. Database setup - Import from JSONL, copy from source, or reuse existing database
  3. Embedding generation - Progress bars for both models
  4. Index creation - JVector index building with timing
  5. Recommendation comparison - 4 methods × 5 query movies
  6. Summary - Method characteristics and use cases
  7. Overall timing - Total execution time

See Also