Java Bridge (arcadedb-python-bridge.jar)¶
The bindings ship a small Java helper jar alongside the engine JARs. Its
sources live in bindings/python/src/java/com/arcadedb/python/:
| Class | Purpose |
|---|---|
RowBatcher |
Serializes up to N result rows into one JSON-array string per call (batched row transport) |
ColumnBatcher |
Encodes up to N rows into one byte[] of typed little-endian column buffers plus null bitmaps (binary columnar transport) |
DocumentBatcher |
Inserts a whole batch of documents from one JSON-rows string (transactional or async parallel writers); also boxes numpy numeric arrays for append_samples |
EdgeBatcher |
Buffers a whole batch of edges into GraphBatch from one call (RID strings, or JSON rows for edges with properties) |
VertexBatcher |
Creates a whole batch of vertices from one JSON-rows string, returning all RIDs as one joined string |
TimeSeriesBatcher |
Fills the engine's primitive TimeSeriesBatch one column per call, so numeric samples are never boxed |
RowAccess |
Hands a row's names and values to Python in one call (namesAndValues), and up to N such rows per call (nextRows); the values are the engine's own objects, so Python converts them with full type fidelity |
Why it exists¶
Every JPype call from Python into the JVM pays a fixed boundary-crossing tax
(on the order of a microsecond). That is invisible for engine-bound work, but
it dominates loops: materializing a wide row costs 2+C crossings (hasNext/next
plus one getProperty per column), and GraphBatch.newEdge() costs one
crossing per edge. Measured, that made 100k-row scans 15–21× slower than
Java-native iteration and bulk edge ingest 24× slower.
The bridge inverts the shape: the loop runs Java-side, and Python pays one
crossing per batch, receiving a bulk payload it can decode at C speed: the
json module for RowBatcher/EdgeBatcher/VertexBatcher, and
numpy.frombuffer for ColumnBatcher. See the
performance page for the resulting numbers.
Which Python APIs use it¶
| Python API | Bridge class |
|---|---|
ResultSet.to_json_list() / iter_json_batches() |
RowBatcher |
ResultSet.to_list() (rows in batches), Result.to_dict() (one row) |
RowAccess |
ResultSet.to_columns() / fast to_dataframe() / to_arrow() |
ColumnBatcher |
Database.insert_many() |
DocumentBatcher |
AsyncExecutor.append_samples() (numpy numeric-column boxing) |
DocumentBatcher |
AsyncExecutor.append_samples(..., primitive=True) |
TimeSeriesBatcher |
GraphBatch.new_edges() (with and without properties) |
EdgeBatcher |
GraphBatch.create_vertices() bulk path |
VertexBatcher |
Database.export_to_csv() (streams JSON batches) |
RowBatcher |
How it builds and ships¶
Both wheel builds compile the bridge with javac against the packaged engine
JARs and jar it up as arcadedb-python-bridge.jar:
scripts/Dockerfile.build(Linux wheels)scripts/build-native.sh(macOS/Windows wheels)
The jar lands in arcadedb_embedded/jars/ inside the wheel, next to the
engine JARs, so it is on the classpath automatically when jvm.py starts the
JVM. There is no separate release artifact or version: it is rebuilt from
source on every wheel build.
Design constraints¶
- No engine code is modified. The bridge is bindings-scoped glue over
Database,MutableDocument,ResultSet,Result,GraphBatch,RID, the asyncErrorCallback, and the engine's JSON serializer, so upstream syncs never conflict with it.TimeSeriesBatcheralso depends on internal time-series classes (TimeSeriesBatch,ColumnDefinition, andLocalTimeSeriesType), so an upstream refactor of those can break it. - Most callers have a fallback. Most Python APIs that ride the bridge
fall back to a pure-JPype implementation if the jar (or a required method)
is absent, so they still work, just slower.
to_columns()andto_arrow()returnNoneinstead, as they do without NumPy or pyarrow. Two need the jar:AsyncExecutor.append_samples(), whose numpy-column path andprimitive=Truepath load the bridge classes directly, andDatabase.insert_many(), which loadsDocumentBatcherfor any JSON-serializable rows and raisesArcadeDBErrorif it is missing (its per-row path runs only for rowsjson.dumpsrejects). - A missing jar is logged. When a bridge class cannot be loaded, the
arcadedb_embedded.resultslogger warnsbridge class <name> unavailable; falling back to the slow per-row path. To see whether the jar is in an install, look for an entry namedarcadedb-python-bridge.jarinjar_fingerprint(per_jar=True)["jars"]. RowBatcherserializes rows property-by-property rather than viaResult.toJSON()to work around upstream #4967 (primitive arrays rendered as"[F@...").