Skip to content

Data Import Guide

This guide covers the currently available import workflows in the Python bindings.

The current bindings are SQL-first. The import surface is SQL IMPORT DATABASE plus the narrow db.import_documents(...) helper for document-file loads, but this repository still does not encourage leaning on importer-based paths heavily from Python.

Use it when you need its supported file-import behavior or a full EXPORT DATABASE + IMPORT DATABASE restore flow. For large Python-side table/document ingest, prefer db.insert_many(...); for bulk graph ingest, prefer GraphBatch.

Overview

The currently exposed importer paths are ArcadeDB's native SQL importer plus a narrow Python wrapper for document imports.

  • CSV document imports
  • CSV graph imports for vertices and edges
  • XML imports
  • Neo4j imports
  • Word2Vec imports for vector workflows
  • RDF imports
  • Full ArcadeDB JSONL restore

All of these are exercised through db.command("sql", "IMPORT DATABASE ...") in the current test suite, but test coverage should not be read as a recommendation to make this your default Python ingest path.

A TIMESERIES type cannot be an IMPORT DATABASE target: it owns no document buckets, so the import fails (the test suite asserts this). Feed it with db.async_executor().append_samples(...) or the server's /api/v1/ts/{db}/write endpoint instead.

db.import_documents(...) is also covered by dedicated API tests, but it should be read the same way: supported, not currently encouraged as the default Python ingest path.

Bulk Ingest Recommendation

IMPORT DATABASE is available, but it is not the recommended path for very large Python-side bulk ingest workloads in this repository right now, and more broadly it is not something we currently encourage as the default Python import story.

  • Example 15 plus the larger table examples are the basis for the current repository guidance.
  • For bulk document ingest from Python, db.insert_many(...) is the recommended default: it batches rows across the FFI boundary; see Example 22. parallel=True hands the rows to the async executor's writers; on a laptop (4 performance cores, parallel level 3, 1,000,000 rows, 6 runs per arm, engine b22b5e9954, 2026-10-04) it loaded 1.11x to 1.14x faster than the synchronous mode at 1, 3, 4, and 8 buckets alike (CREATE DOCUMENT TYPE T BUCKETS n). The maintainers' rule (ArcadeData/arcadedb#8478): as many buckets as the async executor has writers (async_executor().get_parallel_level(), default cores - 1), or a multiple of that, decided when the type is created; create indexes after the load where you can; a record the writers reject raises ArcadeDBError once the load completes. Each bucket has its own sub-index, so on a type with a key (a UNIQUE index) route records by it, ALTER TYPE T BucketSelectionStrategy `partitioned('id')`: then an insert's unique check and a keyed lookup touch one sub-index instead of every bucket's (the same laptop, 400,000 rows, 4 buckets, 6 runs per arm: the parallel load 3.15 s against 2.53 s partitioned, and a keyed lookup 16.2 against 14.6 us).
  • The async executor's SQL command path (db.async_executor().command(...)) is not a bulk-write path at any parallel level. Above parallel level 1 it silently discarded records before 26.10.1 (ArcadeData/arcadedb#7615, fixed in #7625: a failed periodic commit is now retried and otherwise reported through the error callback). Observed on arcadedb-engine 26.9.1 and 26.6.1, measured 2026-09-15: how much was lost varied by run and by workload shape, and 9,742 single-record INSERT commands submitted at parallel level 4 stored 2,436, 5,742, and 7,742 rows across runs. No error reached the per-command callback, nothing was logged, and wait_completion() returned normally. Only the executor-wide on_error handler saw anything, one ConcurrentModificationException per rolled-back batch. create_record, append_samples, db.insert_many(...), and db.graph_batch(...) are unaffected.
  • db.import_documents(...) exists for document-shaped file import convenience; like IMPORT DATABASE, it is not the recommended default for bulk loads from Python.
  • Reserve IMPORT DATABASE for supported import formats, restore flows, and cases where you explicitly need that importer behavior from Python.
  • For bulk graph ingest, use GraphBatch instead: its bulk create_vertices() and new_edges() methods run at or near Java speed; see Example 16 and the graph examples.

Example 15 and 16 Benchmark Structure

Example 15 and 16 are the repository's focused ingest comparison harnesses.

  • Example 15 includes the db.import_documents(...) wrapper in addition to the main SQL-based ingestion paths.
  • Example 15 targets multi-table/document ingest.
  • Example 16 targets graph ingest with vertices plus edges.
  • Both examples enforce parity checks on the final loaded data so timing comparisons are only accepted when the result counts match the expected shape.
  • Example 15 is useful for comparison, but the repository recommendation for bulk table/document ingest is db.insert_many(...), and for bulk graph ingest GraphBatch.

Quick Start

Import CSV as Documents

from pathlib import Path

import arcadedb_embedded as arcadedb


def file_url(path: str) -> str:
    return Path(path).resolve().as_uri()


with arcadedb.create_database("./mydb") as db:
    db.command("sql", "CREATE DOCUMENT TYPE Movie")
    db.command(
        "sql",
        f"IMPORT DATABASE {file_url('./movies.csv')} WITH documentType = 'Movie', commitEvery = 5000",
    )

Import CSV as Vertices and Edges

from pathlib import Path

import arcadedb_embedded as arcadedb


def file_url(path: str) -> str:
    return Path(path).resolve().as_uri()


with arcadedb.create_database("./graphdb") as db:
    db.command("sql", "CREATE VERTEX TYPE Person")
    db.command("sql", "CREATE EDGE TYPE Follows")

    db.command(
        "sql",
        (
            "IMPORT DATABASE WITH "
            f"vertices = '{file_url('./people.csv')}', "
            "vertexType = 'Person', "
            "typeIdProperty = 'id', "
            "typeIdType = 'Long', "
            "typeIdUnique = true"
        ),
    )

    db.command(
        "sql",
        (
            "IMPORT DATABASE WITH "
            f"edges = '{file_url('./follows.csv')}', "
            "edgeType = 'Follows', "
            "typeIdProperty = 'id', "
            "typeIdType = 'Long', "
            "edgeFromField = 'from', "
            "edgeToField = 'to'"
        ),
    )

Restore an ArcadeDB Export

with arcadedb.create_database("./restored") as db:
    db.command(
        "sql",
        "IMPORT DATABASE file:///exports/mydb.jsonl.tgz WITH commitEvery = 50000",
    )

On 26.9.1 and earlier, ArcadeDB #8871 makes this restore lossy: an export that holds an infinite FLOAT or DOUBLE does not import at all (the import stops with NumberFormatException: For input string: "NegInfinity"), a DECIMAL comes back rounded to the 17 digits of a double, and NaN or an infinity inside a list is exported as 0. Fixed in 26.10.1 (PR #8878, verified on its merge) for declared FLOAT, DOUBLE, and DECIMAL properties and for LIST OF FLOAT, LIST OF DOUBLE, and LIST OF DECIMAL. NaN and the infinities still come back as the strings "NaN", "PosInfinity", and "NegInfinity" from a property with no declared type, and are not restored inside MAP properties, embedded documents, or nested lists.

When the copy has to be exact, take a backup instead: in the same test, BACKUP DATABASE and its restore returned all 2,769 records of edge-case values unchanged on both versions, with every index answer equal, and copying the closed database directory (Database Backup Pattern) copies the files themselves.

Choosing the Right Import Mode

CSV

Use CSV imports for flat tabular data and graph vertex/edge feeds.

Best fit:

  • spreadsheet-style datasets
  • relational exports
  • graph nodes and edges split across files
  • moderate-size file imports and format-driven imports

ArcadeDB JSONL Export

Use JSONL exports when you want a full database restore with schema and data intact.

Best fit:

  • environment-to-environment migration
  • backups and restore drills
  • reproducible benchmark datasets

XML, RDF, Neo4j, Word2Vec

These formats are also driven through SQL IMPORT DATABASE. See the import tests for working examples of the exact option sets used in the bindings.

Schema Strategy

Create types, properties, and indexes before importing when you care about data types, constraints, and predictable query plans.

with arcadedb.create_database("./mydb") as db:
    db.command("sql", "CREATE DOCUMENT TYPE User")
    db.command("sql", "CREATE PROPERTY User.id LONG")
    db.command("sql", "CREATE PROPERTY User.email STRING")
    # id and email are read by equality only, so hash indexes (see "Index choice" in the queries guide)
    db.command("sql", "CREATE INDEX ON User (id) UNIQUE_HASH")
    db.command("sql", "CREATE INDEX ON User (email) UNIQUE_HASH")

    db.command(
        "sql",
        "IMPORT DATABASE file:///data/users.csv WITH documentType = 'User', commitEvery = 10000",
    )

Fast Start: Let the Import Create the Target Type

For exploratory work, you can import into a new type with minimal setup. This is quick, but you give up explicit control over schema shape and validation.

Performance Guidance

Tune commitEvery

Use larger commit batches for throughput and smaller ones for lower memory pressure when you do choose the SQL import path.

Dataset Size Recommended commitEvery
< 100K rows 1,000 to 10,000
100K to 1M rows 10,000 to 50,000
> 1M rows 50,000 to 100,000

Drop Heavy Indexes Before Bulk Loads

For large one-shot imports, remove expensive indexes first and recreate them afterward. For the largest Python benchmark ingest paths in this repo, prefer db.insert_many(...) for documents and GraphBatch for graphs instead of leaning on IMPORT DATABASE.

db.command("sql", "DROP INDEX `User[email]`")

db.command(
    "sql",
    "IMPORT DATABASE file:///data/users.csv WITH documentType = 'User', commitEvery = 50000",
)

db.command("sql", "CREATE INDEX ON User (email) UNIQUE_HASH")

Validate the Input Up Front

Check required columns and data cleanliness before starting a long import run. The SQL importer will fail fast on malformed configurations, but it is still cheaper to catch obvious CSV issues before JVM work starts.

Relationship Mapping

For graph imports, load vertices first, then edges, and use matching ID fields.

db.command(
    "sql",
    (
        "IMPORT DATABASE WITH "
        "vertices = 'file:///data/users.csv', "
        "vertexType = 'User', "
        "typeIdProperty = 'id', "
        "typeIdType = 'Long', "
        "typeIdUnique = true"
    ),
)

db.command(
    "sql",
    (
        "IMPORT DATABASE WITH "
        "edges = 'file:///data/follows.csv', "
        "edgeType = 'Follows', "
        "typeIdProperty = 'id', "
        "typeIdType = 'Long', "
        "edgeFromField = 'from', "
        "edgeToField = 'to'"
    ),
)

See Also