Architecture¶
Technical documentation for the ArcadeDB Python bindings architecture, JPype integration, and implementation details.
Overview¶
The ArcadeDB Python bindings are a thin wrapper around the ArcadeDB Java library using JPype for JVM integration. This design provides:
- Full API Coverage: Access to all ArcadeDB features
- Performance: Minimal Python overhead
- Maintenance: Automatic feature parity with Java releases
- Type Safety: Python type hints with Java type conversion
Module Structure¶
arcadedb_embedded/
├── __init__.py # Package exports and version
├── _logging.py # Internal logging helpers
├── async_executor.py # Async command/query + record wrapper
├── core.py # Database, DatabaseFactory, convenience helpers
├── exceptions.py # ArcadeDBError (unified exceptions)
├── exporter.py # Export (JSONL + CSV helper)
├── graph.py # Document, Vertex, Edge wrappers
├── graph_batch.py # High-throughput graph ingest wrapper
├── importer.py # Document import helpers and result payloads
├── jvm.py # JVM startup (bundled JRE, JAR discovery)
├── results.py # ResultSet, Result (query results)
├── schema.py # Schema/Index/Property helpers
├── server.py # ArcadeDBServer (HTTP/Studio)
├── transactions.py # TransactionContext (ACID guard)
├── type_conversion.py # Java ↔ Python value conversion
└── vector.py # VectorIndex + array helpers
Module Responsibilities¶
__init__.py
- Central export surface, defined by
__all__:Database,DatabaseFactory, and the module functionscreate_database,open_database, anddatabase_exists; the record wrappersDocument,Vertex, andEdge;ResultSetandResult;Schema,IndexType, andPropertyType;TransactionContext;AsyncExecutor,GraphBatch, andImportResult; the type converters;VectorIndexand the array helpers (to_java_float_array,to_java_int_array,to_java_byte_array, andto_python_array);ArcadeDBServerandcreate_server;export_databaseandexport_to_csv;jar_fingerprint; andArcadeDBError - Version metadata
_logging.py
- Internal logger access and swallowed-exception helpers for cleanup/finalizer paths
jvm.py
- Starts JVM using bundled JRE and packaged JARs
- Prefers programmatic configuration (
start_jvm(...),jvm_kwargs) - Supports explicit heap and common-pool thread limits via
heap_sizeandcommon_pool_parallelism - Always reads
ARCADEDB_JVM_ARGSand puts its flags beforejvm_args;ARCADEDB_JVM_ERROR_FILEsets the crash-log path - Adds default flags unless the merged arguments already set them:
--add-modules=jdk.incubator.vector,-Djava.awt.headless=true,--enable-native-access=ALL-UNNAMED,-Dfile.encoding=UTF8,--add-opensflags (forjava.util.concurrent.atomic,java.nio.channels.spi, andjava.lang),-Dpolyglot.engine.WarnInterpreterOnly=false,-XX:+UseCompactObjectHeaders, and-Xmx4gwhen no heap is given; thejdk.xmlentity limits are lifted whiledisable_xml_limitsis true, and-XX:ErrorFiledefaults to./log/hs_err_pid%p.log - Starts once per process: a later
start_jvm()with no settings, or the same ones, joins the running JVM; different settings raiseArcadeDBError - From a source checkout (no
jars/orjre/next to the package), extracts them from the newest wheel indist/intobindings/python/.runtime-cache/, stamped with that wheel and re-extracted when the wheel changes jar_fingerprint()hashes the JARs on disk (sha256over all of them,engine_sha256without the bridge JAR), so two installs can be compared by engine rather than by version stringshutdown_jvm()closes open databases and shuts the JVM down- Registers an
atexithook that closes any database still open when the interpreter exits - On Windows, disables Python's
faulthandlerright after the JVM starts, because the JVM's handled access violations would otherwise print as fatal exceptions
core.py
DatabaseFactory: create/open databasesDatabase: queries/commands, transactions, lookups, vector index builderinsert_many(): bulk document ingest via the bridge'sDocumentBatcher- Convenience:
async_executor(),schema, export helpers
graph.py
- Record wrappers:
Document,Vertex,Edge - Property helpers,
new_edge(), type-aware wrapping from Java records
graph_batch.py
GraphBatch: builder-backed high-throughput graph ingest API- Batch vertex/edge creation plus flush/close lifecycle helpers
importer.py
import_documents(): narrow Python wrapper around document import flowsImportResult: normalized import result payload and statistics accessor
schema.py
Schema: type/property/index managementIndexType,PropertyTypeenums
type_conversion.py
convert_java_to_python/convert_python_to_java- Datetime/Decimal/collection handling
async_executor.py
AsyncExecutor: async SQL/OpenCypher command/query flows plus parallel record helpers, commitEvery, WAL tuning
exporter.py
export_database: JSONL only; GraphML and GraphSON need the engine's arcadedb-gremlin module, which the wheel excludes, so they raiseArcadeDBErrorexport_to_csv: serialize ResultSet/list to CSV
vector.py
VectorIndex: JVector-based ANN searchto_java_float_array/to_java_int_array/to_java_byte_array/to_python_array
results.py
ResultSet: iterator, chunking, bulk materialization (to_json_list,iter_json_batches,to_columns), DataFrame and Arrow export (to_dataframe,to_arrow)ResultSet.close()and context-manager use: a set read to its end closes itself; one closed before its end (byfirst(),one(),close(), or leaving itswithblock) raisesArcadeDBErrorwhen read again- An unclosed result set can keep the engine's parallel-scan threads parked, and they stall later queries that need the pool (ArcadeData/arcadedb#8594), so close a set you stop reading early
Result: property access with conversion
transactions.py
TransactionContext: context-managed begin/commit/rollback
server.py
ArcadeDBServer: HTTP/Studio server lifecycle, db management
exceptions.py
ArcadeDBError: unified exception wrapper
Java Bridge Jar¶
Alongside the engine JARs, the wheel ships arcadedb-python-bridge.jar:
small Java helpers (RowBatcher, RowAccess, ColumnBatcher,
DocumentBatcher, EdgeBatcher, VertexBatcher, and TimeSeriesBatcher, sources in
bindings/python/src/java/com/arcadedb/python/)
that move per-row/per-record loops to the Java side so bulk operations cost
one JPype crossing per batch instead of several per row. See
Java Bridge for which Python APIs use it and which of them have no
pure-JPype fallback.
JPype Integration¶
JVM Lifecycle¶
def start_jvm(
heap_size="4g",
disable_xml_limits=True,
jvm_args=None,
common_pool_parallelism=None,
):
if jpype.isJVMStarted():
# No explicit settings: join the running JVM.
# Same settings as the first start: return.
# Different settings: raise ArcadeDBError (the JVM is configured once).
...
return
# Locate bundled JRE + packaged JARs
jvm_path = get_bundled_jre_lib_path()
jar_files = glob.glob(os.path.join(get_jar_path(), "*.jar"))
# ARCADEDB_JVM_ARGS first, then jvm_args, then any missing default flags
args = _build_jvm_args(
heap_size=heap_size,
disable_xml_limits=disable_xml_limits,
jvm_args=jvm_args,
common_pool_parallelism=common_pool_parallelism,
)
# Single-shot startup per process
jpype.startJVM(jvm_path, *args, classpath=os.pathsep.join(jar_files))
JVM Startup:
- Uses the bundled JRE inside the wheel (no system JVM required)
- Loads packaged ArcadeDB JARs from
arcadedb_embedded/jars - Configurable via Python API before first database or server creation (
start_jvm,jvm_kwargs) - JVM stays live for the process lifetime and cannot be restarted
Thread control example:
import arcadedb_embedded as arcadedb
db = arcadedb.create_database(
"./mydb",
jvm_kwargs={
"heap_size": "8g",
"common_pool_parallelism": 8,
},
)
Implications:
- Set JVM options before creating the first database or server in a process
- Tests that need different JVM args must run in separate processes
- Server and embedded modes share the same in-process JVM
Type Conversion¶
Python → Java:
# String
python_str = "hello"
java_str = jpype.JString(python_str)
# Array
python_array = [1.0, 2.0, 3.0]
java_array = jpype.JArray(jpype.JFloat)(python_array)
# NumPy → Java (vectors)
import numpy as np
from arcadedb_embedded import to_java_float_array
numpy_array = np.array([1.0, 2.0, 3.0], dtype=np.float32)
java_array = to_java_float_array(numpy_array)
Java → Python:
# Automatic for primitives
java_int = some_java_method() # Returns Java int
python_int = int(java_int) # Automatic conversion
# Manual for complex types
java_list = some_java_method()
python_list = [item for item in java_list]
# Java array → NumPy
from arcadedb_embedded import to_python_array
java_array = vertex.get("embedding")
numpy_array = to_python_array(java_array)
Type Mapping:
| Python Type | Java Type | Notes |
|---|---|---|
str |
String |
Automatic |
int |
Long |
Automatic |
float |
Double |
Automatic |
bool |
Boolean |
Automatic |
None |
null |
Automatic |
list |
ArrayList |
Converted by convert_python_to_java() (used by set(), and by a bound parameter that is one of several arguments) |
tuple |
ArrayList |
Converted by convert_python_to_java() |
set |
HashSet |
Converted by convert_python_to_java(); stored as a list, so it reads back as a list after the commit |
dict |
HashMap |
Converted by convert_python_to_java() |
Decimal |
BigDecimal |
Converted by convert_python_to_java() |
datetime |
java.util.Date |
Converted by convert_python_to_java() |
date |
LocalDate |
Converted by convert_python_to_java() |
bytes / bytearray |
byte[] |
Converted by convert_python_to_java() |
np.ndarray |
float[] |
via to_java_float_array(); a bound parameter is converted automatically |
np.ndarray (integer dtype) |
int[] |
via to_java_int_array() |
A single list or tuple passed as the only bound argument is not one parameter: it
expands into the positional parameters, one element per ?.
Memory Management¶
Garbage Collection:
- Python GC: Manages Python objects
- Java GC: Manages Java objects
- JPype: Bridges both, uses Java GC for wrapped objects
Best Practices:
# Good: Explicit cleanup
db = arcadedb.open_database("./mydb")
try:
# Use database
pass
finally:
db.close()
# Better: Context manager (closes the database on exit)
with arcadedb.open_database("./mydb") as db:
with db.transaction():
... # Work with database
# Long-running processes: Periodic GC
import gc
for batch in large_dataset:
process_batch(batch)
gc.collect() # Trigger Python GC
Memory Leaks:
- Holding references to Java objects prevents GC
- Large ResultSets should be consumed and released
- Server mode: Monitor JVM heap usage
Class Hierarchy¶
DatabaseFactory (core.py)
├─ create() / open() / exists()
└─ returns Database
create_database() / open_database() / database_exists() (core.py, module functions)
Database (core.py)
├─ query()/command() → ResultSet | None
├─ begin()/commit()/rollback()/transaction() → TransactionContext
├─ run_in_transaction(fn, retries=12) (retries on conflicts)
├─ is_transaction_active()
├─ new_vertex()/new_document() → Vertex | Document
├─ insert_many() → int
├─ import_documents() → ImportResult
├─ graph_batch() → GraphBatch (graph_batch.py)
├─ lookup_by_key()/lookup_by_rid()/count_type()
├─ create_vector_index() → VectorIndex
├─ async_executor() → AsyncExecutor (async_executor.py)
├─ schema → Schema (schema.py)
├─ set_wal_flush()/set_read_your_writes()/set_auto_transaction()
├─ export_database()/export_to_csv()
├─ get_name()
└─ close()/is_open()/drop()
Schema (schema.py)
├─ create_document_type()/create_vertex_type()/create_edge_type()
├─ get_or_create_* helpers
└─ create_property()/create_index()
AsyncExecutor (async_executor.py)
├─ set_parallel_level()/set_commit_every()/set_back_pressure()
├─ command()/query()/scan_type()/new_edge()/transaction()
└─ wait_completion()/close()
Record wrappers (graph.py)
├─ Vertex → new_edge(), modify(), get_out_edges()/get_in_edges()/get_both_edges(), property helpers
├─ Edge → get_out(), get_in(), modify()
└─ Document → get()/set()/save()/delete()/modify(), to_dict(), get_rid()
ResultSet (results.py)
├─ iterator protocol, context manager
├─ to_list()/to_dataframe()/iter_chunks()/count()/first()/one()
├─ to_json_list()/iter_json_batches()/to_columns()/to_arrow()
├─ close()
└─ wraps Result objects
Result (results.py)
├─ has_property()/get()
└─ to_dict()/to_json()
Threading Model¶
Thread Safety¶
Database:
- One
Databaseinstance can be shared by the threads of a process; open it once and pass it around rather than opening it again per thread - Transactions are per thread: each thread's
db.transaction()is its own - Two threads that update the same record can conflict: the losing commit raises
ArcadeDBErrorwithConcurrentModificationExceptionin the message, and the usual answer is to retry that transaction:db.run_in_transaction(fn, retries=12)rolls back and re-runsfnon a conflict
tests/test_concurrency.py covers this: test_thread_safety runs four threads against
one shared Database, and test_oltp_mixed_workload_threads mixes reads with retried
updates. See Concurrency Tests.
Example:
import threading
import arcadedb_embedded as arcadedb
db = arcadedb.open_database("./mydb") # one instance, shared (has a Worker vertex type)
def worker(worker_id):
"""Worker thread using the shared database."""
with db.transaction():
vertex = db.new_vertex("Worker")
vertex.set("id", worker_id)
vertex.save()
threads = [threading.Thread(target=worker, args=(i,)) for i in range(5)]
for t in threads:
t.start()
for t in threads:
t.join()
db.close()
Server Mode:
- The server shares the same in-process JVM and database instances
- HTTP requests are handled by the server's own thread pool
Multiprocessing¶
Only one process can open a database directory at a time: the engine holds an OS lock on
database.lck while the database is open, and a second process gets
ArcadeDBError: ... is locked by another process (asserted by
test_concurrency.py::test_concurrent_access_limitation). Separate processes can work
on separate databases. To share one database across processes, open it in one process
that runs a server, and have the others use its HTTP API (see
Server Patterns).
Performance Considerations¶
Bottlenecks¶
- JVM Boundary Crossing
- Cost: a fixed cost on every Java method call
- Impact: High-frequency calls (loops)
- Solution: Batch operations, use Java bulk APIs
- Type Conversion
- Cost: Varies by type (arrays expensive)
- Impact: Large data transfers
- Solution: Minimize conversions, use efficient formats
- Transaction Overhead
- Cost: a fixed cost on every commit
- Impact: Many small transactions
- Solution: Batch into larger transactions
Optimization Strategies¶
Batch Operations:
# Bad: Many small transactions
for record in records:
with db.transaction():
vertex = db.new_vertex("Data")
vertex.set("data", record)
vertex.save()
# 1000 records = 1000 transactions
# Good: One large transaction
with db.transaction():
for record in records:
vertex = db.new_vertex("Data")
vertex.set("data", record)
vertex.save()
# 1000 records = 1 transaction
Query Optimization:
# Bad: N+1 queries
users = db.query("sql", "SELECT FROM User")
for user in users:
# Separate query per user!
orders = db.query("sql", "SELECT FROM Order WHERE user_id = ?", user.get("id"))
# Good: Single query with traversal
result = db.query("sql", """
SELECT
name,
out('Placed').name as orders
FROM User
""")
ResultSet Streaming:
# Bad: Load all results
result = db.query("sql", "SELECT FROM LargeTable")
all_results = list(result) # Loads everything into memory
# Good: Stream results
result = db.query("sql", "SELECT FROM LargeTable")
for row in result:
process(row)
When you do need the whole result materialized, prefer the bulk APIs
(to_columns()/to_dataframe() or to_json_list()) over list(result) /
to_list(); see the Performance guide.
Profiling¶
Python Side:
import cProfile
import pstats
def benchmark():
db = arcadedb.create_database("./bench")
with db.transaction():
for i in range(10000):
vertex = db.new_vertex("Data")
vertex.set("id", i)
vertex.save()
db.close()
# Profile
cProfile.run('benchmark()', 'stats.prof')
stats = pstats.Stats('stats.prof')
stats.sort_stats('cumulative')
stats.print_stats(20)
Java Side:
The bindings start the JVM themselves (jvm.py), so pass JVM options through
start_jvm() (or jvm_kwargs, or the ARCADEDB_JVM_ARGS environment variable)
before the first database or server is created:
from arcadedb_embedded.jvm import start_jvm
# GC logging with JDK unified logging
start_jvm(jvm_args="-Xlog:gc*:file=gc.log")
Single Package Distribution¶
The Python binding is distributed as a single, self-contained package (arcadedb-embedded).
Features:
- Bundled JRE: Includes a minimal Java 25 Runtime Environment (JRE) bundled directly in the wheel.
- Query engines: SQL, OpenCypher, and GraphQL ship in the wheel.
scripts/jar_exclusions.txtdrops the Gremlin, MongoDB, and gRPC modules, Raft HA (so the server is single-node), metrics and tracing, the JavaScript stack (js, truffle, icu4j, and regex), commons-math3, Jackson, snappy-java, and jline. - Wire protocols: the Postgres, Redis, and Bolt plugins ship, and the server starts
each one only when its plugin is configured (see
create_server()). - Zero Configuration: No external Java installation required.
# pip install arcadedb-embedded
import arcadedb_embedded as arcadedb
db = arcadedb.create_database("./mydb")
db.query("sql", "SELECT FROM User")
db.query("opencypher", "MATCH (n) RETURN n")
Extension Points¶
Custom Vertex/Edge Classes¶
import jpype
class CustomVertex:
"""Custom vertex wrapper with helper methods."""
def __init__(self, java_vertex):
self._java_vertex = java_vertex
def get_friends(self):
"""Get the vertices reached by outgoing 'Knows' edges."""
direction = jpype.JClass("com.arcadedb.graph.Vertex$DIRECTION")
return list(self._java_vertex.getVertices(direction.OUT, "Knows"))
Custom Loaders¶
class CustomXmlLoader:
"""Custom XML loading helper."""
def __init__(self, db):
self.db = db
def load_xml(self, file_path, vertex_type):
"""Load XML records into a vertex type."""
import xml.etree.ElementTree as ET
tree = ET.parse(file_path)
root = tree.getroot()
with self.db.transaction():
for elem in root.findall('.//record'):
vertex = self.db.new_vertex(vertex_type)
for child in elem:
vertex.set(child.tag, child.text)
vertex.save()
# Usage
xml_loader = CustomXmlLoader(db)
xml_loader.load_xml("data.xml", "Data")
Testing¶
The test suite, its fixtures, and the patterns it uses are documented under Testing, with one page per test file and a Best Practices summary.
Build System¶
Package Build¶
# pyproject.toml configuration
[build-system]
requires = ["build>=0.7.0", "setuptools>=61.0", "wheel", "jpype1"]
build-backend = "setuptools.build_meta"
[project]
name = "arcadedb-embedded"
# Placeholder: the build overwrites this from the parent pom.xml
version = "0.0.0"
requires-python = ">=3.10"
dependencies = ["jpype1>=1.5.0"]
jpype1 is the only runtime dependency. numpy is an optional extra
(arcadedb-embedded[vector]) rather than a hard requirement, so a plain install
stays minimal. The extras' floors are audited in CI, so treat
bindings/python/pyproject.toml as authoritative rather than this excerpt.
JAR Management¶
scripts/setup_jars.py downloads nothing. It runs inside the Docker build
(scripts/Dockerfile.build, python-builder stage) and stages what earlier
stages produced into the package:
find_jar_files()looks for the already-filtered JARs in/build/jars(the Docker build location) or/home/arcadedb/libcopy_jars_to_package()clearssrc/arcadedb_embedded/jars/and copies those JARs into itcopy_jre()replacessrc/arcadedb_embedded/jre/with thejlinkJRE from/build/jremain()runs the two copies and exits non-zero if either fails
The JARs themselves come from the arcadedata/arcadedb image, unless
build.sh is given a local JAR directory (its third argument, JAR_LIB_DIR), and
jar_exclusions.txt is applied before this script runs. Native builds
(scripts/build-native.sh) do the same staging themselves and do not call it.
See Build Architecture.
See Also¶
- Database API Reference - Core database operations
- Troubleshooting - Common issues and solutions
- JPype Documentation - JPype library docs
- ArcadeDB Java API - Underlying Java API