Example 16: Import Database vs Transactional Graph Ingest¶
This example compares four graph-ingest strategies against the same generated vertex and edge dataset shape.
Overview¶
Example 16 is the graph-ingest comparison harness for embedded Python.
- Generates synthetic graph CSV data for vertices and edges
- Runs four ingest modes:
- transactional SQL vertex and edge creation
- embedded
GraphBatch - async SQL vertex and edge creation, pinned to one worker
- SQL
IMPORT DATABASE
- Checks final vertex and edge count parity before trusting the timing result
Current Repository Guidance¶
- This example exists because ingest winners are workload-dependent
GraphBatchis the repository's recommended bulk graph ingest path from PythonIMPORT DATABASEis not the recommendation: its behavior varies by import path and data shape,commitEverytransaction splitting has not behaved as expected in some CSV-heavy runs, and very large imports can still hit transaction-buffer limits. Use it when its file-driven workflow is what you need, and when you have measured it on your own data- Async SQL is a comparison arm only, and it is not a bulk graph ingest path
Why the async arm accepts only --async-parallel 1
Before 26.10.1, async_executor().command(...) could silently drop records above
parallel level 1 (ArcadeData/arcadedb#7615, fixed in #7625); see
Bulk Ingest Recommendation.
run_async_sql_graph_load(...) therefore raises ValueError for any
--async-parallel other than 1, and counts stored vertices and edges against
submitted vertices and edges so a short load fails instead of being reported as a
fast one.
GraphBatch flushes its edges through that same executor and is measured exact,
with parallel_flush on or off.
Run¶
From bindings/python/examples:
python 16_import_database_vs_transactional_graph_ingest.py \
--vertices 100000 \
--edges 300000 \
--vertex-int-props 6 \
--vertex-str-props 4 \
--edge-int-props 2 \
--edge-str-props 1 \
--string-size 64 \
--batch-size 10000 \
--async-parallel 1 \
--parallel 1 \
--heap-size 4g
Key Options¶
--vertices: number of generated vertices--edges: number of generated edges--vertex-int-props/--vertex-str-props: vertex property counts--edge-int-props/--edge-str-props: edge property counts--string-size: generated string payload size--batch-size: ingest batch size--async-parallel: async SQL worker count; only1is accepted (#7615)--parallel: SQL import worker count and GraphBatch parallel-flush toggle--heap-size: JVM heap size
Parity Semantics¶
Timing comparisons only matter if all four modes produce matching final vertex and edge counts.