Skip to content

Testing Overview

The ArcadeDB Python bindings have a comprehensive test suite covering all major functionality.

What's Tested

The test suite covers:

  • ✅ Core database operations - CRUD, transactions, queries
  • ✅ Server mode - HTTP API, multi-client access
  • ✅ Concurrency patterns - File locking, thread safety, multi-process
  • ✅ Graph operations - Vertices, edges, traversals
  • ✅ Query languages - SQL, OpenCypher
  • ✅ Vector search - JVector-based LSM_VECTOR indexes, similarity search
  • ✅ Data import - SQL IMPORT DATABASE across CSV, XML, Neo4j, Word2Vec, and RDF, plus the import_documents() wrapper
  • ✅ Graph ingest helper - GraphBatch buffering and flush behavior
  • ✅ Geospatial SQL - geo.within, geo.intersects, null input returns null; a boundary point returns a boolean (which one is not asserted)
  • ✅ Time series SQL - CREATE TIMESERIES TYPE, range queries, bucketing
  • ✅ Materialized views - create, refresh, alter, drop lifecycle
  • ✅ Graph algorithms - shortestPath, dijkstra, astar
  • ✅ HASH schema indexes - create, discover, idempotent get_or_create_index, force drop, and the UNIQUE_HASH id path
  • ✅ Unicode support - International characters, emoji
  • ✅ Schema introspection - Querying database metadata
  • ✅ Type conversions - Python/Java type mapping
  • ✅ Large result sets - Result sets of 1,000 records: ordered iteration, filtered and aggregate queries (no timing is asserted)
  • ✅ Async executor - async_executor() commands, queries, callbacks, and exact command-path counts at parallel levels 1 and 4
  • ✅ Bulk ingest - insert_many, AsyncExecutor.create_record, and numpy append_samples
  • ✅ Export - JSONL database export, CSV result export, and the error GraphML and GraphSON raise without arcadedb-gremlin
  • ✅ ResultSet API - to_list, to_json_list, to_dataframe, to_arrow, and release of the engine cursor
  • ✅ JVM - startup, arguments, and lifecycle (re-entry, reopen in one process, exit with an unclosed database)
  • ✅ Bundled wire protocols - PostgreSQL (including Arrow ADBC), Bolt, and the Redis port setting; a default server opens none of their ports
  • ✅ Packaging and provenance - server JARs in the wheel, jar_fingerprint(), the wheel platform tag, the dev-mode runtime cache, and __version__
  • ✅ Sparse vectors - LSM_SPARSE_VECTOR weight precision, the settle step (COMPACT INDEX), and INT8 rescoring
  • ✅ Cross-model atomicity - search, hop, and update in one transaction survive an interruption
  • ✅ RESTORE - RESTORE DOCUMENT and RESTORE VERTEX record counts and record integrity
  • ✅ Schema batching - schema statements apply at once; many batch in one transaction
  • ✅ Docs snippets - selected Python blocks from the documentation run as code
  • ✅ JVM payload check - a static check that no Python list crosses into the JVM

Quick Start

Install Test Dependencies

  1. Build the wheel: cd bindings/python && ./scripts/build.sh (Docker on Linux, Python 3.12 by default). It writes the wheel to bindings/python/dist/ and, outside CI, refreshes the repo-root uv environment.
  2. From the repository root (or bindings/python), run uv run pytest. The repo-root pyproject.toml is the test environment, pinned to Python 3.12, so it needs a cp312 wheel. To leave out the heavy example packages (torch, sentence-transformers), pass --no-group examples to uv sync and uv run.

Other Python versions and platforms are covered by CI (see CI/CD Setup). CI does not use the uv environment: it installs the built wheel and a hand-maintained list of test dependencies, then runs pytest tests/ from bindings/python (see CI Gates).

Run All Tests

# From the repository root or bindings/python (from any other directory a bare
# run collects only that directory)
uv run pytest

# With verbose output
uv run pytest -v

# With coverage report
uv run pytest --cov=arcadedb_embedded --cov-report=html

On Windows, run with --capture=sys: pytest's default fd capture can leave the JVM writing its log lines to a handle that no longer belongs to it, and the test then hangs in that write (issue #10).

Run Specific Tests

# Run a specific test file (paths are relative to the repository root)
uv run pytest bindings/python/tests/test_core.py

# Run a specific test function
uv run pytest bindings/python/tests/test_core.py::test_database_creation

# Run tests matching a keyword
uv run pytest -k "transaction"
uv run pytest -k "server"
uv run pytest -k "concurrency"

# Run with output (see print statements)
uv run pytest -v -s

Test Files Overview

Test counts evolve over time. For the latest per-file counts, run uv run pytest -v -rs.

Test File Description
test_async_executor.py Async command/query execution, callback behavior, and exact command-path counts at parallel levels 1 and 4
test_bulk_insert.py Recommended bulk paths land every row, plus Database.insert_many, AsyncExecutor.create_record, vector columns, and numpy append_samples bulk ingest
test_core.py Core database operations, CRUD, transactions, queries
test_database_utils.py count_type, is_transaction_active, and drop, plus error handling on a closed database
test_docs_examples.py Executes representative Python snippets from the documentation site
test_exporter.py Database export formats and CSV result export helpers
test_graph_api.py Graph wrapper behavior for vertices, edges, and traversal helpers
test_importer_api.py Narrow db.import_documents(...) wrapper coverage
test_logging_helper.py Internal _logging helper configuration behavior
test_numpy_support.py NumPy integration and array conversion behavior
test_resultset.py Result and ResultSet iteration, accessors, and export helpers
test_schema.py Schema, property, and index management behavior
test_schema_batching.py Schema statements apply immediately, a rollback does not undo them, and many batch in one transaction
test_server.py Server lifecycle, configuration, and databases through the Java API (no HTTP calls)
test_concurrency.py File locking, thread safety, multi-process behavior
test_server_patterns.py Best practices for embedded + server mode
test_import_database.py SQL IMPORT DATABASE scenarios and format coverage
test_cypher.py OpenCypher query language
test_graph_batch.py Bulk graph-ingest helper coverage
test_graph.py GraphBatch.new_edges and create_vertices bulk path coverage
test_geo_predicate_sql.py Geospatial SQL predicate semantics
test_timeseries_sql.py Time-series SQL type creation, range filters, and bucketing
test_materialized_view_sql.py Materialized view lifecycle and refresh behavior
test_restore_sql.py RESTORE DOCUMENT/VERTEX record-count and record integrity
test_graph_algorithms_sql.py SQL graph algorithm runtime coverage
test_hash_index_schema.py HASH index schema API behavior, the UNIQUE_HASH id path end to end, plus a named-list IN parameter on an LSM_TREE index
test_jvm_args.py JVM args handling
test_transaction_config.py WAL flush, read-your-writes, and auto-transaction settings
test_type_conversion.py Python/Java type conversion coverage
test_vector.py Vector API and nearest-neighbor search behavior
test_vector_params_verification.py Vector param validation
test_vector_sql.py SQL vector functions, index creation, and search flows
test_cross_model_atomicity.py Search, hop, and update in one transaction survive an interruption between the writes with nothing torn; with one transaction per write they are torn every time
test_example11_degree_matching.py hnsw_m_from_max_connections() halves maxConnections for hnswlib-derived backends, never below 1, and accepts a string; skips if examples/11 is absent
test_jar_provenance.py The wheel can say which engine it carries, not just which version it is.
test_java_package_shadowing.py A folder named java/ or com/ must not change what a query returns
test_jvm.py start_jvm() re-entry once the JVM is running, close and reopen in one process, and interpreter exit with an unclosed database
test_sigint.py Ctrl-C raises KeyboardInterrupt and runs cleanup once the JVM is started; interrupt=True keeps JPype's default
test_jvm_payload.py A Python list must never be what crosses into the JVM
test_resultset_arrow.py Tests for ResultSet.to_arrow().
test_columnar_readers.py to_columns, to_dataframe, and to_arrow on schemaless and DECIMAL data, at several batch sizes.
test_runtime_cache.py The dev-mode runtime cache must follow the wheel it was extracted from
test_server_http_endpoints.py The three server HTTP features the bindings document but do not wrap (multi-request transactions, server database commands, and line-protocol time-series writes), plus a projection read over HTTP
test_server_packaging.py The server stack is actually IN the wheel, and the API is reachable.
test_server_wire_protocols.py The wire protocols the wheel bundles are actually reachable.
test_sparse_quantization_compact.py Sparse index weight precision and the settle step, the dense search beam argument, and INT8 sparse rescoring
test_vector_delta_visibility.py Vectors written after an index build are searchable, exactly, before any rebuild
test_vector_second_pass.py A repeated query set returns the same neighbours as its first pass
test_wheel_platform_tag.py Built wheel manylinux platform tag verification (regression tests for ArcadeData/arcadedb#4037), and __version__ equals the installed distribution version

Common Testing Workflows

Development Workflow

# Run only failed tests from last run
uv run pytest --lf

Debugging Tests

# Stop on first failure
uv run pytest -x

# Drop into debugger on failure
uv run pytest --pdb

# Show local variables on failure
uv run pytest -l

# Verbose with full output
uv run pytest -vv -s

Test Markers

The markers are registered in bindings/python/pyproject.toml (server, server_wire, and integration); the repo-root pyproject.toml mirrors that block. These are the ones the suite uses:

Marker Tests
server test_server_creation, test_server_database_operations, test_server_custom_config, and test_server_context_manager in test_server.py, test_server_starts_and_serves_http in test_server_packaging.py, and test_docs_api_access_examples in test_docs_examples.py
server_wire Every test in test_server_wire_protocols.py (module-level pytestmark)

integration is registered, but no test uses it.

# Run only the tests marked server
uv run pytest -m server

# Run only OpenCypher tests (a keyword match, not a marker)
uv run pytest -k cypher

# Run all except the tests marked server
uv run pytest -m "not server"

-m "not server" skips only the tests marked server. Other tests that start a server still run: test_server_patterns.py, test_server_http_endpoints.py, and test_server_wire_protocols.py (marked server_wire, not server). To leave out every server-starting test:

uv run pytest -m "not server and not server_wire" \
  --ignore=bindings/python/tests/test_server_patterns.py \
  --ignore=bindings/python/tests/test_server_http_endpoints.py

Expected Output

A passing run ends with a summary of the form N passed, M skipped, K xfailed, with no failures or errors. The strict xfails are the engine bugs listed on the Known Engine Issues page, and nothing skips. Run with -rs to see why a test skipped locally.

Skips

CI runs the suite with no skips: scripts/check_test_skips.py fails the test job on any skip in the JUnit XML (tests/test_check_test_skips.py tests the gate). A skip is a test that should have run and did not, so the suite does not use one for anything it can state otherwise:

  • A test file that cannot run on a platform is left out of collection, not skipped: tests/conftest.py sets collect_ignore for test_sigint.py on Windows (it sends SIGINT to a child process) and for test_docs_examples.py on the upstream pull request branch, which has no docs/. A parametrized case Windows cannot create (a directory name with a question mark in test_importer_api.py) is not generated there.
  • An optional Python package (numpy, pandas, pyarrow, requests, psycopg, adbc-driver-postgresql, neo4j) uses pytest.importorskip without a custom reason. A run without the package skips locally; CI installs them all, so a missing one fails the job.
  • A bundled feature never skips: OpenCypher, the geo and graph-algorithm SQL functions, time series, the HASH index, vector encodings, the importers, and the server stack run and fail when they are missing. Earlier versions of these tests skipped on the engine's error or on an empty answer, which let a wrong empty answer pass as a skip.

A skip that is truly unavoidable (a Windows limitation, a case that needs an engine fix that is still upstream) is added to the list in scripts/check_test_skips.py with the platform, the reason, and the upstream issue; for an engine bug prefer a strict xfail, which fails the suite once the fix arrives. CI installs every optional package the tests import (numpy, pandas, pyarrow, requests, psycopg, neo4j, redis, adbc-driver-postgresql), so none of them skips there.

See CI Gates.

Next Steps