Data Import Guide¶
This guide covers the currently available import workflows in the Python bindings.
The current bindings are SQL-first. The import surface is SQL IMPORT DATABASE plus
the narrow db.import_documents(...) helper for document-file loads, but this
repository still does not encourage leaning on importer-based paths heavily from Python.
Use it when you need its supported file-import behavior or a full
EXPORT DATABASE + IMPORT DATABASE restore flow. For large Python-side table/document
ingest, prefer db.insert_many(...); for bulk graph ingest, prefer GraphBatch.
Overview¶
The currently exposed importer paths are ArcadeDB's native SQL importer plus a narrow Python wrapper for document imports.
- CSV document imports
- CSV graph imports for vertices and edges
- XML imports
- Neo4j imports
- Word2Vec imports for vector workflows
- RDF imports
- Full ArcadeDB JSONL restore
All of these are exercised through db.command("sql", "IMPORT DATABASE ...") in
the current test suite, but test coverage should not be read as a recommendation to make
this your default Python ingest path.
A TIMESERIES type cannot be an IMPORT DATABASE target: it owns no document buckets,
so the import fails (the test suite asserts this). Feed it with
db.async_executor().append_samples(...) or the server's /api/v1/ts/{db}/write
endpoint instead.
db.import_documents(...) is also covered by dedicated API tests, but it should be read
the same way: supported, not currently encouraged as the default Python ingest path.
Bulk Ingest Recommendation¶
IMPORT DATABASE is available, but it is not the recommended path for very large
Python-side bulk ingest workloads in this repository right now, and more broadly it is
not something we currently encourage as the default Python import story.
- Example 15 plus the larger table examples are the basis for the current repository guidance.
- For bulk document ingest from Python,
db.insert_many(...)is the recommended default: it batches rows across the FFI boundary; see Example 22.parallel=Truehands the rows to the async executor's writers; on a laptop (4 performance cores, parallel level 3, 1,000,000 rows, 6 runs per arm, engineb22b5e9954, 2026-10-04) it loaded 1.11x to 1.14x faster than the synchronous mode at 1, 3, 4, and 8 buckets alike (CREATE DOCUMENT TYPE T BUCKETS n). The maintainers' rule (ArcadeData/arcadedb#8478): as many buckets as the async executor has writers (async_executor().get_parallel_level(), default cores - 1), or a multiple of that, decided when the type is created; create indexes after the load where you can; a record the writers reject raisesArcadeDBErroronce the load completes. Each bucket has its own sub-index, so on a type with a key (a UNIQUE index) route records by it,ALTER TYPE T BucketSelectionStrategy `partitioned('id')`: then an insert's unique check and a keyed lookup touch one sub-index instead of every bucket's (the same laptop, 400,000 rows, 4 buckets, 6 runs per arm: the parallel load 3.15 s against 2.53 s partitioned, and a keyed lookup 16.2 against 14.6 us). - The async executor's SQL command path (
db.async_executor().command(...)) is not a bulk-write path at any parallel level. Above parallel level 1 it silently discarded records before 26.10.1 (ArcadeData/arcadedb#7615, fixed in #7625: a failed periodic commit is now retried and otherwise reported through the error callback). Observed on arcadedb-engine 26.9.1 and 26.6.1, measured 2026-09-15: how much was lost varied by run and by workload shape, and 9,742 single-recordINSERTcommands submitted at parallel level 4 stored 2,436, 5,742, and 7,742 rows across runs. No error reached the per-command callback, nothing was logged, andwait_completion()returned normally. Only the executor-wideon_errorhandler saw anything, oneConcurrentModificationExceptionper rolled-back batch.create_record,append_samples,db.insert_many(...), anddb.graph_batch(...)are unaffected. db.import_documents(...)exists for document-shaped file import convenience; likeIMPORT DATABASE, it is not the recommended default for bulk loads from Python.- Reserve
IMPORT DATABASEfor supported import formats, restore flows, and cases where you explicitly need that importer behavior from Python. - For bulk graph ingest, use
GraphBatchinstead: its bulkcreate_vertices()andnew_edges()methods run at or near Java speed; see Example 16 and the graph examples.
Example 15 and 16 Benchmark Structure¶
Example 15 and 16 are the repository's focused ingest comparison harnesses.
- Example 15 includes the
db.import_documents(...)wrapper in addition to the main SQL-based ingestion paths. - Example 15 targets multi-table/document ingest.
- Example 16 targets graph ingest with vertices plus edges.
- Both examples enforce parity checks on the final loaded data so timing comparisons are only accepted when the result counts match the expected shape.
- Example 15 is useful for comparison, but the repository recommendation for bulk
table/document ingest is
db.insert_many(...), and for bulk graph ingestGraphBatch.
Quick Start¶
Import CSV as Documents¶
from pathlib import Path
import arcadedb_embedded as arcadedb
def file_url(path: str) -> str:
return Path(path).resolve().as_uri()
with arcadedb.create_database("./mydb") as db:
db.command("sql", "CREATE DOCUMENT TYPE Movie")
db.command(
"sql",
f"IMPORT DATABASE {file_url('./movies.csv')} WITH documentType = 'Movie', commitEvery = 5000",
)
Import CSV as Vertices and Edges¶
from pathlib import Path
import arcadedb_embedded as arcadedb
def file_url(path: str) -> str:
return Path(path).resolve().as_uri()
with arcadedb.create_database("./graphdb") as db:
db.command("sql", "CREATE VERTEX TYPE Person")
db.command("sql", "CREATE EDGE TYPE Follows")
db.command(
"sql",
(
"IMPORT DATABASE WITH "
f"vertices = '{file_url('./people.csv')}', "
"vertexType = 'Person', "
"typeIdProperty = 'id', "
"typeIdType = 'Long', "
"typeIdUnique = true"
),
)
db.command(
"sql",
(
"IMPORT DATABASE WITH "
f"edges = '{file_url('./follows.csv')}', "
"edgeType = 'Follows', "
"typeIdProperty = 'id', "
"typeIdType = 'Long', "
"edgeFromField = 'from', "
"edgeToField = 'to'"
),
)
Restore an ArcadeDB Export¶
with arcadedb.create_database("./restored") as db:
db.command(
"sql",
"IMPORT DATABASE file:///exports/mydb.jsonl.tgz WITH commitEvery = 50000",
)
On 26.9.1 and earlier, ArcadeDB
#8871 makes this restore lossy: an
export that holds an infinite FLOAT or DOUBLE does not import at all (the import stops
with NumberFormatException: For input string: "NegInfinity"), a DECIMAL comes back
rounded to the 17 digits of a double, and NaN or an infinity inside a list is exported as
0. Fixed in 26.10.1 (PR #8878, verified on its merge) for declared FLOAT, DOUBLE, and
DECIMAL properties and for LIST OF FLOAT, LIST OF DOUBLE, and LIST OF DECIMAL. NaN
and the infinities still come back as the strings "NaN", "PosInfinity", and
"NegInfinity" from a property with no declared type, and are not restored inside MAP
properties, embedded documents, or nested lists.
When the copy has to be exact, take a backup instead: in the same test, BACKUP DATABASE and
its restore returned all 2,769 records of edge-case values unchanged on both versions, with
every index answer equal, and copying the closed database directory
(Database Backup Pattern) copies the files
themselves.
Choosing the Right Import Mode¶
CSV¶
Use CSV imports for flat tabular data and graph vertex/edge feeds.
Best fit:
- spreadsheet-style datasets
- relational exports
- graph nodes and edges split across files
- moderate-size file imports and format-driven imports
ArcadeDB JSONL Export¶
Use JSONL exports when you want a full database restore with schema and data intact.
Best fit:
- environment-to-environment migration
- backups and restore drills
- reproducible benchmark datasets
XML, RDF, Neo4j, Word2Vec¶
These formats are also driven through SQL IMPORT DATABASE. See the import tests for
working examples of the exact option sets used in the bindings.
Schema Strategy¶
Recommended: Pre-create the Schema¶
Create types, properties, and indexes before importing when you care about data types, constraints, and predictable query plans.
with arcadedb.create_database("./mydb") as db:
db.command("sql", "CREATE DOCUMENT TYPE User")
db.command("sql", "CREATE PROPERTY User.id LONG")
db.command("sql", "CREATE PROPERTY User.email STRING")
# id and email are read by equality only, so hash indexes (see "Index choice" in the queries guide)
db.command("sql", "CREATE INDEX ON User (id) UNIQUE_HASH")
db.command("sql", "CREATE INDEX ON User (email) UNIQUE_HASH")
db.command(
"sql",
"IMPORT DATABASE file:///data/users.csv WITH documentType = 'User', commitEvery = 10000",
)
Fast Start: Let the Import Create the Target Type¶
For exploratory work, you can import into a new type with minimal setup. This is quick, but you give up explicit control over schema shape and validation.
Performance Guidance¶
Tune commitEvery¶
Use larger commit batches for throughput and smaller ones for lower memory pressure when you do choose the SQL import path.
| Dataset Size | Recommended commitEvery |
|---|---|
| < 100K rows | 1,000 to 10,000 |
| 100K to 1M rows | 10,000 to 50,000 |
| > 1M rows | 50,000 to 100,000 |
Drop Heavy Indexes Before Bulk Loads¶
For large one-shot imports, remove expensive indexes first and recreate them afterward.
For the largest Python benchmark ingest paths in this repo, prefer db.insert_many(...)
for documents and GraphBatch for graphs instead of leaning on IMPORT DATABASE.
db.command("sql", "DROP INDEX `User[email]`")
db.command(
"sql",
"IMPORT DATABASE file:///data/users.csv WITH documentType = 'User', commitEvery = 50000",
)
db.command("sql", "CREATE INDEX ON User (email) UNIQUE_HASH")
Validate the Input Up Front¶
Check required columns and data cleanliness before starting a long import run. The SQL importer will fail fast on malformed configurations, but it is still cheaper to catch obvious CSV issues before JVM work starts.
Relationship Mapping¶
For graph imports, load vertices first, then edges, and use matching ID fields.
db.command(
"sql",
(
"IMPORT DATABASE WITH "
"vertices = 'file:///data/users.csv', "
"vertexType = 'User', "
"typeIdProperty = 'id', "
"typeIdType = 'Long', "
"typeIdUnique = true"
),
)
db.command(
"sql",
(
"IMPORT DATABASE WITH "
"edges = 'file:///data/follows.csv', "
"edgeType = 'Follows', "
"typeIdProperty = 'id', "
"typeIdType = 'Long', "
"edgeFromField = 'from', "
"edgeToField = 'to'"
),
)
See Also¶
- Import Workflow Reference - Supported SQL import surface
- Import Examples - Practical examples
- Database API - Database operations
- Transactions - Transaction management