Skip to content

Java API Coverage Analysis

This section provides a practical mapping between the ArcadeDB Java API and the Python bindings surface in this repository. It reflects the current code in arcadedb_embedded rather than a theoretical, full Java surface comparison.

Executive Summary

The Python bindings expose the core database, schema, graph, vector, async, import/export, and server workflows needed for typical application usage. Most omissions are low-level JVM internals (WAL details, bucket scanning, binary protocol, clustering) that are not typically used from Python. The package bundles the optional in-process server (HTTP API + Studio); see Server Mode. For a server whose lifetime is independent of your Python process, or for HA/TLS, use the official ArcadeDB distribution.

Coverage by Area (Qualitative)

Area Status Notes
Core Database ✅ Supported DatabaseFactory, Database, transactions, lookups, async helpers
Query Execution ✅ Supported SQL and OpenCypher
Schema & Indexes ✅ Supported Types, properties, LSM_TREE/HASH/FULL_TEXT/LSM_VECTOR/GEOSPATIAL indexes
Graph API ✅ Supported SQL/OpenCypher graph workflows plus Document/Vertex/Edge wrapper compatibility
Vector Search ✅ Supported JVector indexes + NumPy conversion helpers
Sparse Vectors ✅ Supported LSM_SPARSE_VECTOR indexes and vector.sparseNeighbors(...) through SQL
Time Series ✅ Supported TIMESERIES types through SQL, async_executor().append_samples(...) for bulk writes, and the server's /api/v1/ts/{db}/write endpoint
Async Execution ✅ Supported AsyncExecutor plus record-level and SQL/Cypher async flows
Data Import ✅ Supported SQL import workflows plus a narrow db.import_documents(...) wrapper for document files
Data Export ✅ Supported JSONL + CSV for query results
Server Mode ✅ Supported Embedded server lifecycle + Studio access
Advanced/Low-level ❌ Not exposed WAL internals, binary protocol, HA/replication

Detailed Coverage

1. Core Database Operations

DatabaseFactory:

  • ✅ create(), open(), exists()

Database:

  • ✅ query(language, query, *args) and command(language, command, *args)
  • ✅ Transactions: begin(), commit(), rollback(), transaction()
  • ✅ Transaction helpers: run_in_transaction() (retries on concurrent-modification conflicts), is_transaction_active()
  • ✅ Records: new_document(), new_vertex(), lookup_by_rid(), lookup_by_key()
  • ✅ Bulk ingest: insert_many() (documents), graph_batch() (vertices and edges)
  • ✅ Import: import_documents() for document-shaped files
  • ✅ Vector indexes: create_vector_index()
  • ✅ Utilities: count_type(), drop(), get_name(), get_database_path(), is_open(), close()
  • ✅ Configuration: set_auto_transaction(), set_read_your_writes(), is_read_your_writes(), set_wal_flush()
  • ✅ Async execution: async_executor()
  • ✅ Export helpers: export_database() and export_to_csv()

Not directly exposed: bucket scans, WAL internals, low-level binary protocol

2. Query Execution

SQL and OpenCypher run through db.query() and db.command():

  • ✅ SQL
  • ✅ OpenCypher

The GraphQL module ships in the wheel, but no test exercises it from Python. Gremlin and the MongoDB query language are not bundled (scripts/jar_exclusions.txt).

ResultSet & Results:

  • ✅ Pythonic iteration (ResultSet.__iter__, __next__)
  • ✅ ResultSet helpers: first(), one(), count(), close()
  • ✅ Row materialization: to_list() / iter_dicts() (full Python types), to_json_list() / iter_json_batches() (same list-of-dicts shape, ~5.5x faster than to_list() on a wide scan, JSON-native values), iter_chunks()
  • ✅ Columnar materialization: to_columns(), to_arrow(), to_dataframe()
  • ✅ Result.get(), has_property(), get_property_names()
  • ✅ Result.to_json(), to_dict() (Python enhancement)

3. Graph API

Recommended approach: SQL/OpenCypher for graph writes and traversals, with wrapper APIs available when you explicitly need record objects

Wrapper/record APIs available:

  • ✅ db.new_vertex(type) / db.new_document(type)
  • ✅ record.set(name, value) / record.save() / record.delete() / record.modify()
  • ✅ vertex.new_edge(label, target, **props) (bidirectionality controlled by EdgeType schema)
  • ✅ vertex.get_out_edges(), get_in_edges(), get_both_edges()
  • ✅ db.lookup_by_rid(rid) for direct record access

Graph Traversals & Queries:

  • ✅ SQL traversal: SELECT * FROM User WHERE out('Follows').name = 'Alice'
  • ✅ OpenCypher patterns: MATCH (a:User)-[:Follows]->(b) RETURN b
  • ✅ Path finding, shortest paths, pattern matching

Not exposed: event listeners/callback hooks, low-level graph internals

Recommended query-first approach:

# Create vertices via SQL
with db.transaction():
    db.command("sql", "INSERT INTO User SET id = 1, name = 'Alice'")
    db.command("sql", "INSERT INTO User SET id = 2, name = 'Bob'")
    db.command("sql", """
        CREATE EDGE Follows
        FROM (SELECT FROM User WHERE id = 1)
        TO (SELECT FROM User WHERE id = 2)
    """)

# Traverse via OpenCypher
result = db.query("opencypher", """
    MATCH (user:User {name: 'Alice'})-[:Follows]->(friend)
    RETURN friend.name
""")

Wrapper/object APIs still available:

with db.transaction():
    alice = db.new_vertex("Person").set("name", "Alice").save()
    bob = db.new_vertex("Person").set("name", "Bob").save()
    alice.new_edge("Follows", bob, since=date.today())  # saved by new_edge

4. Schema & Index API

Full Pythonic Schema API available via db.schema:

  • ✅ create_document_type(), create_vertex_type(), create_edge_type()
  • ✅ get_or_create_*() helpers
  • ✅ create_property(), drop_property()
  • ✅ drop_type(), exists_type(), get_type(), get_types()
  • ✅ Indexes: create_index(), drop_index(), get_indexes(), exists_index()
  • ✅ Vector indexes: SQL CREATE INDEX ... LSM_VECTOR, secondary/manual helper coverage, list_vector_indexes()

5. Server Mode

Supported:

  • ✅ ArcadeDBServer(root_path, root_password, config) / create_server(...) - Server initialization
  • ✅ start(), stop(), context manager support
  • ✅ get_database(), create_database() - Database management
  • ✅ get_studio_url(), get_http_port() - Python enhancements
  • ✅ Embedded and HTTP access to the same databases
  • ✅ The bundled Postgres, Redis, and Bolt protocol plugins, started through config={"server_plugins": ...} (see Wire Protocols)

Not exposed:

  • ❌ HA/replication, advanced user/security management: run the official ArcadeDB server for those

6. Data Import

Supported:

  • ✅ SQL IMPORT DATABASE for CSV/TSV documents
  • ✅ SQL IMPORT DATABASE for CSV graph vertices and edges with ID resolution
  • ✅ SQL IMPORT DATABASE for XML
  • ✅ SQL IMPORT DATABASE for ArcadeDB JSONL exports
  • ✅ SQL IMPORT DATABASE for RDF, Neo4j, and Word2Vec scenarios covered by tests
  • ❌ SQL IMPORT DATABASE into a TIMESERIES type (a TIMESERIES type owns no document buckets; use append_samples() or the /api/v1/ts/{db}/write endpoint)
  • ✅ db.import_documents(...) wrapper for document-shaped file imports via the Java importer
  • ✅ Batch processing and automatic type inference where supported by the Java importer

The importer surface is intentionally still described conservatively in this repository. Support exists, but the current repository guidance is:

  • bulk table/document ingest: db.insert_many(...)
  • bulk graph ingest: GraphBatch
  • importer-based paths: available, but not the recommended default for bulk loads

7. Data Export

  • ✅ JSONL export - Full database backup format
  • ✅ CSV export of query results via export_to_csv()
  • ✅ Type filtering via include_types / exclude_types
  • ✅ Compression when exporting JSONL (Java exporter)
  • ❌ GraphML and GraphSON export: they come from the arcadedb-gremlin module, which the wheel excludes, so they raise ArcadeDBError
  • ✅ Vector index creation - SQL CREATE INDEX ... LSM_VECTOR
  • ✅ NumPy array support - to_java_float_array(), to_java_int_array(), to_java_byte_array(), to_python_array()
  • ✅ Similarity search - SQL vectorNeighbors
  • ✅ Distance functions - cosine, euclidean, dot_product
  • ✅ Index tuning parameters (connections, beam width, quantization)
  • ✅ Automatic indexing of existing records
  • ✅ List vector indexes - schema.list_vector_indexes()

9. Advanced / Low-Level APIs Not Exposed

  • ❌ WAL and storage internals
  • ❌ Binary protocol and custom network stacks
  • ❌ HA/replication, distributed clustering
  • ❌ Plugins other than the bundled Postgres, Redis, and Bolt ones, and module management
  • ❌ Custom query engines and DSLs

Design Philosophy: Query-First Approach

The Python bindings follow a "query-first, API-second" philosophy, which is ideal for Python developers. Instead of exposing every Java object, operations are enabled through:

  • SQL DDL for schema management
  • SQL/OpenCypher for graph and document operations
  • Thin helper APIs for transactions, vector search, and targeted record access

This approach is actually cleaner and more maintainable than direct API exposure:

# Python way (clean):
db.command("sql", "CREATE INDEX ON User (email) UNIQUE_HASH")
db.query("opencypher", "MATCH (a)-[:Follows]->(b) RETURN b")

# vs. hypothetical direct API (complex):
schema = db.getSchema()
type = schema.getType("User")
index_builder = schema.buildTypeIndex("User", ["email"])
index = index_builder.withUnique(true).create()

Use Case Suitability

Use Case Suitable? Notes
Embedded database in Python app ✅ Excellent Core use case
Graph analytics with Cypher ✅ Excellent SQL and OpenCypher supported
Document store ✅ Excellent SQL and schema APIs
Vector similarity search ✅ Excellent JVector + NumPy integration
Development with Studio UI ✅ Excellent Server mode included
Data migration (CSV/XML/JSONL import) ✅ Good SQL import workflows exercised by tests
Async bulk ingestion ❌ Not recommended Before 26.10.1, async_executor().command(...) could silently drop records above parallel level 1 (ArcadeData/arcadedb#7615, fixed in #7625); see Bulk Ingest Recommendation. Use insert_many() or GraphBatch
Multi-master replication ❌ Not supported Java server only
Custom query language ❌ Not supported Use built-in languages

Conclusion

These bindings cover the primary workflows most Python developers need:

  • Embedded multi-model database
  • Graph, document, vector, and time-series data
  • SQL and OpenCypher queries

They intentionally do not expose low-level JVM internals and clustering. For those scenarios, use the Java APIs directly.

For development workflow and tests, see Contributing.