Data Import Examples¶
This page points to the examples that load data into ArcadeDB from Python. The Data Import Guide holds the recommendations, the import formats, and the code patterns; the examples below show them at scale.
Before running the MovieLens and Stack Overflow examples, download their datasets with the Dataset Downloader.
Which Example Covers What¶
Example 04 - CSV Import: Documents
- parses the MovieLens CSV files in Python and loads them into document types with
batched, parameterized
INSERTstatements - defines the schema explicitly (integer-like columns to LONG, decimals to DOUBLE, text to STRING) and imports empty cells as NULL
- benchmarks queries before and after indexes
- with
--export, exports the database to JSONL and re-imports it withIMPORT DATABASEas a round trip
Example 05 - CSV Import: Graph
- reads Example 04's document database (or an Example 04 JSONL export) and builds a graph from it: users and movies as vertices, ratings and tags as edges
- builds vertices with SQL,
db.graph_batch(...), or synchronous transactions, and edges with SQLCREATE EDGE - creates the indexes before the edges (unless
--no-index), and validates the graph with queries
Example 15 - Table Ingest Comparison and Example 16 - Graph Ingest Comparison
- compare transactional SQL, the async SQL path, SQL
IMPORT DATABASE, anddb.import_documents(...)(tables) orGraphBatch(graphs) on the same generated data, with count checks before any timing is trusted
- bulk document ingest with
db.insert_many(...), transactional and withparallel=Trueon a type created with one bucket per async writer - time-series ingest from numpy arrays with
AsyncExecutor.append_samples(...)
The Short Version¶
- For bulk document ingest, use
db.insert_many(...). Withparallel=Trueit raisesArcadeDBErrorif the async writers reject any record. - For bulk graph ingest, use
db.graph_batch(...). - Keep SQL
IMPORT DATABASEfor its supported file formats and for restoring exports. - Do not use
db.async_executor().command(...)for bulk writes. Before 26.10.1 it could silently drop records above parallel level 1 (ArcadeData/arcadedb#7615, fixed in #7625); see Bulk Ingest Recommendation.
Additional Resources¶
- Data Import Guide - Recommendations, formats, and code patterns
- Import Workflow Reference - Supported SQL import surface
- Database API: insert_many - Bulk document ingest
- GraphBatch API - Bulk graph ingest
Source Code¶
View the complete example source code: