Skip to content

Examples Overview

Hands-on examples demonstrating ArcadeDB Python bindings in real-world scenarios. Each example is documented and ready to run; most run on their own, while 05 reads Example 04's database, 06 reads Example 05's, and 12 reuses Example 11's.

DSL-first examples

Current examples and docs use SQL/OpenCypher as the default approach for schema, CRUD, and graph operations.

Available Examples

🏁 Getting Started

Dataset Downloader Download and prepare datasets used by the examples (MovieLens, Stack Exchange, MSMARCO, TPC-H, and LDBC SNB).

01 - Simple Document Store Foundation example covering document types, CRUD operations, comprehensive data types (DATE, DATETIME, DECIMAL, FLOAT, INTEGER, STRING, BOOLEAN, and LIST), and NULL value handling (INSERT NULL, UPDATE to NULL, IS NULL queries).

02 - Social Network Graph Complete graph modeling with vertices, edges, NULL handling, and dual query languages (SQL MATCH vs Cypher). Demonstrates 8 people with optional fields, each friendship stored as two directed edges (one each way), graph traversal, and comprehensive queries.

03 - Vector Search Semantic similarity search with HNSW (JVector) indexing. Demonstrates vector storage, index creation, and nearest neighbor search.

04 - CSV Import - Documents Production CSV import with explicit schema mapping, batched parameterized INSERT, NULL handling, and index optimization. Imports MovieLens dataset (36M+ records) with comprehensive performance analysis and result validation with actual data samples.

05 - CSV Import - Graph Production graph creation from MovieLens dataset. Performance analysis of SQL pipelines, GraphBatch versus synchronous vertex transactions, and index effects. Includes benchmark configurations, validation queries, and export/import roundtrip testing.

06 - Vector Search - Movie Recommendations Production-ready vector embeddings and HNSW (JVector) indexing for semantic movie search.

07 - Stack Overflow Tables (OLTP) Table-oriented OLTP benchmark with mixed CRUD operations and deterministic single-thread verification.

08 - Stack Overflow Tables (OLAP) Table-oriented OLAP benchmark with fixed analytical queries, load/index timing, and repeated query runs.

09 - Stack Overflow Graph (OLTP) Graph OLTP benchmark with directed-edge semantics, result verification notes, and cross-backend workload comparison.

10 - Stack Overflow Graph (OLAP) Graph OLAP benchmark with a fixed query suite across multiple backends: OpenCypher where the backend runs it, an equivalent SQL or Python form elsewhere.

11 - Vector Index Build Build-only vector benchmark comparing ArcadeDB, pgvector, Qdrant, Milvus, FAISS, and LanceDB.

12 - Vector Search Search-only vector benchmark that reuses Example 11 output and sweeps backend-specific search parameters.

13 - Stack Overflow Hybrid Queries Standalone SQL + graph + vector workflow over Stack Overflow data.

14 - Lifecycle Timing Embedded lifecycle benchmark covering JVM startup, load, query, close, and reopen timing.

15 - Import Database vs Transactional Table Ingest Four-way table-ingest comparison. Repository guidance from these experiments is to prefer db.insert_many(...) for bulk table/document ingest; the async SQL arm is a comparison arm, not a recommendation.

16 - Import Database vs Transactional Graph Ingest Four-way graph-ingest comparison. Repository guidance from these experiments is to prefer GraphBatch for bulk graph ingest.

17 - Time Series End-to-End SQL-first time-series workflow covering type creation, tagged inserts, range queries, and hourly bucket aggregation.

18 - Geo Predicates With WKT Points And Polygons SQL-first geospatial workflow covering WKT storage, GEOSPATIAL indexes, indexed within / intersects, polygon overlap queries, and fallback after dropping the index.

19 - Hash Index Exact-Match Lookup Workflow SQL-first HASH index workflow covering unique and non-unique exact-match lookups, schema inspection, and duplicate-key rejection.

20 - Graph Algorithms Route Planning SQL-first graph algorithms workflow covering minimum-hop shortestPath, weighted dijkstra / astar, sqlscript variables, and route-cost comparison.

21 - Graph Analytical View SQL Workflow SQL-first Graph Analytical View workflow covering six-figure synthetic graph generation, CREATE / ALTER / REBUILD, schema metadata polling, stale-versus-ready lifecycle, and persisted GAV restoration.

22 - numpy Bulk I/O Batched Python/Java boundary crossings for bulk workloads: Database.insert_many (transactional and parallel), AsyncExecutor.append_samples from numpy arrays, time-bucketed aggregation, and columnar export with to_columns().

23 - Server Mode And HTTP Access Embedded-first server workflow covering create_server(...), HTTP auth (Basic and bearer token), server-managed database creation, and mixed embedded plus HTTP access to the same data.

24 - Transactions, Database Commands, and Time-Series Writes over HTTP The server HTTP features a second process needs next: one transaction across several requests through arcadedb-session-id, close database / open database server commands, and InfluxDB line-protocol writes to a TIMESERIES type read back with SQL.

25 - Sparse Vectors, Weight Precision, And Compaction Sparse retrieval on a synthetic SPLADE-style corpus built once per weight setting: INT8 rescored, INT8 without rescoring, and FP32 posting weights in LSM_SPARSE_VECTOR, and COMPACT INDEX as the settle step after a bulk load, with size, compaction time, query latency, and top-10 agreement.

26 - Cross-Model Transaction Atomicity Vector search, graph hop, and document update in one transaction, interrupted between the writes: rolled back cleanly inside a transaction, torn every time without one.

Quick Start

⚠️ Important: Always run examples from the examples/ directory.

cd bindings/python/examples/
python 01_simple_document_store.py

Learning Path

  1. Document Store (01) - Learn fundamentals
  2. Graph Operations (02) - Understand relationships
  3. Vector Search (03) - AI/ML integration
  4. CSV Import - Documents (04) - ETL to documents with MovieLens
  5. CSV Import - Graph (05) - Same data as graph with performance benchmarks
  6. Vector Search - Movies (06) - Semantic search and recommendations
  7. Stack Overflow Tables (OLTP/OLAP) (07/08) - Table benchmarks and fairness conventions
  8. Stack Overflow Graph (OLTP/OLAP) (09/10) - Directed graph benchmarks and query suites
  9. Vector Benchmarks (11/12) - Index build and search benchmarking across vector backends
  10. Hybrid Queries (13) - Combined SQL, graph, and vector workflow
  11. Lifecycle And Ingest Benchmarks (14/15/16) - Embedded lifecycle timing and ingest comparisons
  12. Time Series SQL Workflow (17) - Tagged samples, range queries, and bucket aggregation from Python
  13. Geo Predicate Workflow (18) - WKT points and polygons, indexed spatial filters, and fallback semantics
  14. Hash Index Workflow (19) - Exact-match HASH indexes, missing-key behavior, and duplicate protection
  15. Graph Algorithms Workflow (20) - Minimum-hop versus weighted routing with shortestPath, dijkstra, and astar
  16. Graph Analytical View Workflow (21) - Manage GAV lifecycle entirely through SQL and inspect schema:graphAnalyticalViews
  17. numpy Bulk I/O (22) - Batched bulk ingest with insert_many and append_samples, plus columnar numpy export with to_columns()
  18. Server Mode Workflow (23) - Start the in-process server, create schema over HTTP, and verify mixed embedded plus HTTP access
  19. Server HTTP Transactions And Time Series (24) - Multi-request transactions, server database commands, and line-protocol writes over HTTP
  20. Sparse Vectors And Compaction (25) - Sparse index weight precision and the COMPACT INDEX settle step
  21. Cross-Model Atomicity (26) - Search, hop, and update in one transaction that survives an interruption with nothing torn

Start with Simple Document Store!