GraphBatch API¶
The GraphBatch helper exposes ArcadeDB's high-throughput graph-ingest path from Python.
Overview¶
Use GraphBatch when you need to load many vertices and edges efficiently.
This is the repository's current recommended bulk graph-ingest path from Python, and
the reason is not only throughput. The alternative of submitting per-record SQL through
db.async_executor().command(...) silently discarded records above parallel level 1
before 26.10.1 (ArcadeData/arcadedb#7615, fixed in #7625: a failed periodic commit is
now retried and otherwise reported through the error callback). GraphBatch dispatches
its edge flush through that same
executor and is measured exact: 20,000 vertices and 40,000 edges landed in full, with
and without parallel_flush.
You typically create it through db.graph_batch(...) rather than constructing the class directly.
A database runs one GraphBatch at a time: while one is open, db.graph_batch(...)
raises ArcadeDBError ("A GraphBatch is already in progress on this database"). Close
the open batch first, or use one batch with parallel_flush for parallel work.
Entry Point¶
db.graph_batch(...) -> GraphBatch¶
Create a configured batch helper tied to the current database.
Common options:
batch_size: buffered edge batch size before flush; usually leave it unset and giveexpected_edge_countinsteadexpected_edge_count: the number of edges you are about to load; with nobatch_size, the batch size is tuned to it (clamped to 100,000-5,000,000), and a single flush is the optimal shapeedge_list_initial_size: initial size in bytes of each vertex's edge segmentlight_edges: create property-less light edges when appropriatebidirectional: store each edge on both vertices (the default) or on its source only. PassFalseonly for an edge type declared one-way (CREATE EDGE TYPE ... UNIDIRECTIONAL). From 26.10.1 a one-way edge in a two-way type (the defaultCREATE EDGE TYPE) is refused:new_edgeraisesArcadeDBErrornaming the type and writes nothing. Before 26.10.1 it was accepted, and every query the planner walked from the target end returned 0 rows with no error (ArcadeData/arcadedb#8625). What a one-way edge is visible to is in Graphscommit_every: commit cadence during batch workuse_wal: write-ahead log during the import. Off by default: a crash in the middle of the import can lose its tail, with nothing to replay. Passuse_wal=Truefor an import that must survive a crashwal_flush: flush policy such asno,yes_nometadata,yes_fullpre_allocate_edge_chunks: allocate the edge chunks whencreate_vertex()creates the vertexparallel_flush: flush deferred work in parallelcommit_retries: retries for a vertex commit that hits a transientNeedRetryException(default 10,0fails fast)commit_retry_delay_ms: initial retry back-off, exponential thereafter and capped at 10000 ms (default 1000)chunk_cache_capacity: bound on each OUT/IN head-chunk RID cache, which keeps memory flat on a long-lived stream (default 1,000,000)max_deferred_incoming_edges: buffered deferred incoming edges before the connection pass runs early fromflush()instead of once atclose()(default 5,000,000,0defers everything to close)
Recommended settings for a crash-safe bulk load (ArcadeDB's maintainers,
ArcadeData/arcadedb#8287): use_wal=True and expected_edge_count, with
batch_size, commit_every, and parallel_flush left at their defaults
(commit_every is 50,000 with the WAL on and one commit per flush with it off;
parallel_flush is on). Measured on 26.10.1-dev with 50,000 vertices and
1,045,738 edges: about 7.5 s with the WAL on, against about 30 s for the same
graph through newVertex/newEdge one element at a time. The size hint made
no measurable difference at that size. The same options exist for a server as
query parameters of POST /api/v1/batch/{db} (wal=true,
expectedEdgeCount=...), the served bulk path.
Example:
with db.graph_batch(use_wal=True, expected_edge_count=50000) as batch:
alice = batch.create_vertex("Person", name="Alice")
bob = batch.create_vertex("Person", name="Bob")
batch.new_edge(alice, "Knows", bob, since=2024)
Transactions¶
Call the batch outside your own transactions. create_vertices(), flush(), and close()
manage their own transactions, and so do new_edge() and new_edges() when the buffer
reaches batch_size and flushes. Leaving a with db.graph_batch() block calls close().
From engine 26.10.1 (ArcadeDB
#9242, fixed) each of them raises
ArcadeDBError when you have a transaction open, and your transaction stays open with your
writes in it. A refused close() leaves the batch open with its edges pending: end your
transaction and call close() again. On 26.9.1 and earlier they commit the transaction open
on the thread, yours included. Commit your own writes before the batch's first call, or write them after it
closes, which is right on every version:
with db.transaction():
db.new_document("Note").set("text", "mine").save()
with db.graph_batch() as batch:
rids = batch.create_vertices("Person", [{"id": i} for i in range(1000)])
batch.new_edges(rids[:-1], "Knows", rids[1:])
create_vertex() and new_vertex() are the exceptions. Inside your transaction,
create_vertex() saves the vertex in it and does not commit, so your rollback() undoes
both; outside one, it commits its own. new_edge() and new_edges() with room left in the
buffer only buffer.
While a batch is open, the batch's WAL setting (use_wal=False by default) applies to the
batch's own calls only. On 26.9.1 and earlier it stayed on the thread from the batch's first
call until close(), and every commit on that thread used it, yours too: a transaction of
yours committed in that time wrote no WAL record.
Common Operations¶
create_vertex(type_name, **properties)¶
Create and persist a single vertex.
new_vertex(type_name)¶
Return an unsaved Vertex of that type from the batch; set its properties and call
save() inside a transaction, and end that transaction before the batch's next call that
commits (see Transactions).
create_vertices(type_name, count_or_properties)¶
Create many vertices efficiently and return their RIDs as strings. The call commits, in
the transaction open on the thread if there is one: call it outside your own transactions
(see Transactions).
count_or_properties is either an int, the number of vertices to create without
properties, or an iterable of property dicts (None or {} for a vertex without
properties). Only rows whose values are all scalars (str, int, float, bool, or
None) take the JSON bulk path, which sends them in chunks of 100,000 rows; a row with
any other value, a list or a Java array included, sends the whole call through the
per-value path, which converts each value on its own. So does a scalar the JSON text would
change on the way to the engine: an integer beyond 64 bits, NaN or Infinity, or a string with
a lone surrogate. The per-value path stores such a value exactly or raises.
Vector properties: pass to_java_float_array(vec), not a Python list. A plain
list is converted element by element (and, on a type with no declared vector property,
stored as a LIST rather than ARRAY_OF_FLOATS): 100,000 vectors of 128 floats loaded
at 3.6k vectors/s as lists against 31k/s as to_java_float_array values, 8.6x
(measured 2026-09-24 on 26.10.1-dev, same data, stored vectors identical). The list form
was also slower than inserting one vector per db.command(...).
new_edge(source, edge_type, destination, **properties)¶
Buffer an edge for creation during flush/close.
Declared edge properties before 26.10.1
Up to engine 26.9.1 an edge buffered with properties skipped the declared property's
conversion and the type's constraints: a None followed by another property was stored as
-1 in an INTEGER, and 40000 in a SHORT as -25536. From 26.10.1 a batched edge
stores a null as null and converts or refuses a declared value as Vertex.new_edge does;
on an older engine, write edges with declared properties through Vertex.new_edge.
new_edges(source_rids, edge_type, destination_rids, properties=None)¶
Buffer many edges with one JPype crossing per call: the bulk counterpart of
new_edge, which pays one boundary crossing per edge. RIDs may be strings
("#1:0") or objects with a string representation; properties is an optional
same-length sequence of per-edge property dicts. When every value is a scalar (str,
int, float, bool, or None) the call takes the bulk path; any other value,
a list included, sends the whole call through per-edge buffering, and so does a scalar the
JSON text would change (an integer beyond 64 bits, NaN or Infinity, a lone surrogate), which
per-edge buffering stores exactly or refuses. Returns the batch for chaining.
with db.graph_batch(use_wal=False) as batch:
rids = batch.create_vertices("Person", [{"id": i} for i in range(100)])
batch.new_edges(rids[:-1], "Knows", rids[1:])
flush()¶
Force buffered edge work to disk early. Commits the transaction open on the thread, yours included (see Transactions).
close()¶
Flush remaining work and finalize the batch. A second close() does nothing. Commits the
transaction open on the thread, yours included (see Transactions).
Counters¶
The helper also exposes counters such as:
get_total_edges_created()get_buffered_edge_count()get_deferred_incoming_edge_count()
After close(), the counters and every other method except close() raise
ArcadeDBError ("GraphBatch is closed"). Read the counters before the batch closes.
Notes¶
- Prefer
GraphBatchover importer-based graph loading for Python-managed bulk ingest. wal_flushvalidation is intentionally strict and raisesValueErrorfor invalid modes.- See the graph-ingest examples and tests for realistic usage patterns.