Skip to content

Schema API

The Schema provides type management, index creation, and property definitions for documents, vertices, and edges.

Overview

The Schema class enables:

  • Type Management: Create document, vertex, and edge types
  • Property Definitions: Define typed properties (constraints go through SQL DDL, see below)
  • Index Creation: Create indexes for query optimization
  • Schema Inspection: Query existing schema definitions

DSL-first recommendation

For application code, prefer ArcadeDB SQL DDL via db.command("sql", ...). This page documents the Schema wrapper API for reference and advanced typed schema workflows.

Getting Schema

import arcadedb_embedded as arcadedb

# Use context manager to ensure clean close
with arcadedb.create_database("./mydb") as db:
    # Create types (schema statements apply immediately)
    # Vertex type
    user_type = db.schema.create_vertex_type("User")

    # Edge type
    follows_type = db.schema.create_edge_type("Follows")

    # Document type
    log_type = db.schema.create_document_type("LogEntry")

Schema statements apply immediately

Creating or dropping a type, property, or index needs no transaction, and it is not transactional: it takes effect at once, and a rollback does not undo it. To create many types, run the statements inside one with db.transaction(): or send them as one sqlscript. The schema is then written to disk once, when the transaction ends, instead of once per statement (ArcadeData/arcadedb#8635). See Transactions.

Type Creation Methods

create_vertex_type

schema.create_vertex_type(
    name: str,
    buckets: Optional[int] = None
) -> Any

Create a new vertex type. Returns the underlying Java VertexType object.

Parameters:

  • name (str): Name of the vertex type
  • buckets (Optional[int]): Number of buckets (engine default 1 when omitted)

Returns:

  • VertexType: Created vertex type

Example:

# Schema statements apply immediately (no transaction needed)
# Basic vertex type
user_type = schema.create_vertex_type("User")

# With custom buckets
product_type = schema.create_vertex_type("Product", buckets=10)

create_edge_type

schema.create_edge_type(
    name: str,
    buckets: Optional[int] = None
) -> Any

Create a new edge type. Returns the underlying Java EdgeType object.

Parameters:

  • name (str): Name of the edge type
  • buckets (Optional[int]): Number of buckets (engine default 1 when omitted)

Returns:

  • EdgeType: Created edge type

Example:

# Schema statements apply immediately (no transaction needed)
# Basic edge type
follows_type = schema.create_edge_type("Follows")

# With custom buckets
purchased_type = schema.create_edge_type("Purchased", buckets=5)

create_document_type

schema.create_document_type(
    name: str,
    buckets: Optional[int] = None
) -> Any

Create a new document type. Returns the underlying Java DocumentType object.

Parameters:

  • name (str): Name of the document type
  • buckets (Optional[int]): Number of buckets (engine default 1 when omitted)

Returns:

  • DocumentType: Created document type

Example:

# Schema statements apply immediately (no transaction needed)
# Basic document type
log_type = schema.create_document_type("LogEntry")

# With custom buckets
event_type = schema.create_document_type("Event", buckets=8)

get_or_create_document_type

schema.get_or_create_document_type(
    name: str,
    buckets: Optional[int] = None
) -> Any

Get an existing document type or create it if it doesn't exist. Idempotent alternative to create_document_type. Returns the underlying Java DocumentType object.

Parameters:

  • name (str): Type name
  • buckets (Optional[int]): Number of buckets if creating a new type (engine default when omitted)

Returns:

  • DocumentType: Existing or newly created document type

Example:

doc_type = db.schema.get_or_create_document_type("Product")

get_or_create_vertex_type

schema.get_or_create_vertex_type(
    name: str,
    buckets: Optional[int] = None
) -> Any

Get an existing vertex type or create it if it doesn't exist. Idempotent alternative to create_vertex_type. Returns the underlying Java VertexType object.

Parameters:

  • name (str): Type name
  • buckets (Optional[int]): Number of buckets if creating a new type (engine default when omitted)

Returns:

  • VertexType: Existing or newly created vertex type

Example:

vertex_type = db.schema.get_or_create_vertex_type("User")

get_or_create_edge_type

schema.get_or_create_edge_type(
    name: str,
    buckets: Optional[int] = None
) -> Any

Get an existing edge type or create it if it doesn't exist. Idempotent alternative to create_edge_type. Returns the underlying Java EdgeType object.

Parameters:

  • name (str): Type name
  • buckets (Optional[int]): Number of buckets if creating a new type (engine default when omitted)

Returns:

  • EdgeType: Existing or newly created edge type

Example:

edge_type = db.schema.get_or_create_edge_type("Follows")

drop_type

schema.drop_type(name: str)

Drop a type and all its data.

Parameters:

  • name (str): Type name to drop

Raises:

  • ArcadeDBError: If the drop fails

Example:

db.schema.drop_type("OldType")

Property Definition

create_property

schema.create_property(
    type_name: str,
    property_name: str,
    property_type: Union[str, PropertyType],
    of_type: Optional[str] = None
) -> Any

Create a property on a type. Returns the underlying Java Property object.

Parameters:

  • type_name (str): Name of the type
  • property_name (str): Name of the property
  • property_type (str or PropertyType): ArcadeDB type (see types below)
  • of_type (Optional[str]): Element type for LIST/MAP collections

Returns:

  • Java Property object

Property Types:

The PropertyType enum (importable as from arcadedb_embedded import PropertyType) defines the supported values:

  • Primitives: STRING, INTEGER, LONG, SHORT, BYTE, BOOLEAN, FLOAT, DOUBLE, DECIMAL, DATE, DATETIME
  • Binary: BINARY
  • Collections: LIST, MAP, EMBEDDED
  • Links: LINK
  • Vectors: ARRAY_OF_FLOATS

Either the enum member or its string name may be passed. A string may also name an engine type the enum does not list: DATETIME_SECOND, DATETIME_MICROS, DATETIME_NANOS, ARRAY_OF_SHORTS, ARRAY_OF_INTEGERS, ARRAY_OF_LONGS, or ARRAY_OF_DOUBLES.

Example:

# Schema statements apply immediately (no transaction needed)
schema.create_vertex_type("User")

# String property
schema.create_property("User", "name", "STRING")

# Integer property (enum form)
from arcadedb_embedded import PropertyType
schema.create_property("User", "age", PropertyType.INTEGER)

# Date property
schema.create_property("User", "birthDate", "DATE")

# List property
schema.create_property("User", "tags", "LIST", of_type="STRING")

# Embedded property
schema.create_property("User", "profile", "EMBEDDED")

get_or_create_property

schema.get_or_create_property(
    type_name: str,
    property_name: str,
    property_type: Union[str, PropertyType],
    of_type: Optional[str] = None
) -> Any

Get an existing property or create it if it doesn't exist. Same parameters as create_property, except that of_type must be a string here: a PropertyType member raises ArcadeDBError. Returns the underlying Java Property object.

prop = schema.get_or_create_property("User", "email", "STRING")

drop_property

schema.drop_property(type_name: str, property_name: str)

Drop a property from a type.

schema.drop_property("User", "old_field")

Property constraints

Constraints such as mandatory, not-null, default, and min/max are configured on the returned Java Property object (for example prop.setMandatory(True)) or via SQL DDL. The Python Schema wrapper itself only exposes property creation/removal.

Index Creation

create_index

schema.create_index(
    type_name: str,
    property_names: List[str],
    unique: bool = False,
    index_type: Union[str, IndexType] = IndexType.LSM_TREE
) -> Index

Create an index on a type.

Parameters:

  • type_name (str): Name of the type
  • property_names (List[str]): List of property names to index
  • unique (bool): Whether the index should enforce uniqueness (default: False)
  • index_type (str or IndexType): Type of index ("LSM_TREE", "HASH", "FULL_TEXT", "LSM_VECTOR", "GEOSPATIAL"); default: IndexType.LSM_TREE

Returns:

  • Index: Created index object

Raises:

  • ArcadeDBError: If type doesn't exist or index creation fails

Example:

# Schema statements apply immediately (no transaction needed)
# Unique id read only by equality: a unique hash index (ArcadeData/arcadedb#9169)
schema.create_index("User", ["username"], unique=True, index_type="HASH")

# Unique key you also range over or sort by: the default LSM_TREE
schema.create_index("Ticket", ["number"], unique=True)

# Non-unique exact-match lookup index
schema.create_index("Order", ["customerId"], index_type="HASH")

# Composite index
schema.create_index("Event", ["userId", "timestamp"])

# Full-text index
schema.create_index("Article", ["content"], index_type="FULL_TEXT")

Index choice rules of thumb:

  • Use HASH for an id that is only read, updated and deleted by equality and is not bulk-loaded in key order. From 26.10.1 a unique hash index answers SQL id = ? 1.5 to 2.3 times faster than LSM_TREE and an Index.get() hit 1.9 to 3.1 times faster. Its insert is 1.14 to 1.23 times faster for shuffled ids but 9% to 16% slower for ids loaded in ascending order (ArcadeDB #9169, 200,000 and 2,000,000 entries). It cannot serve a range or an ORDER BY. On 26.9.1 its inserts are several times slower than LSM_TREE.
  • Use LSM_TREE when you need ranges, sorting, or a safe general-purpose default, and for a non-unique column with few distinct values.
  • Use FULL_TEXT, LSM_VECTOR, and GEOSPATIAL only for their specialized query types.
  • HASH can be unique or non-unique. The index structure and uniqueness constraint are separate choices.

SQL DSL equivalents:

  • CREATE INDEX ON User (email) UNIQUE -> unique LSM_TREE (ranges and ORDER BY)
  • CREATE INDEX ON User (email) NOTUNIQUE -> non-unique LSM_TREE
  • CREATE INDEX ON User (email) UNIQUE_HASH -> unique HASH (equality only, the choice for an id)
  • CREATE INDEX ON Order (customerId) NOTUNIQUE_HASH -> non-unique HASH
  • CREATE INDEX ON Article (content) FULL_TEXT -> FULL_TEXT
  • CREATE INDEX ON Doc (embedding) LSM_VECTOR ... -> LSM_VECTOR
  • CREATE INDEX ON Place (location) GEOSPATIAL -> GEOSPATIAL

Vector (JVector) Parameters:

Note

schema.create_index() only accepts type_name, property_names, unique, and index_type. It does not take the JVector tuning parameters below. To configure a vector index from Python, use db.create_vector_index() (which exposes max_connections, beam_width, dimensions, etc.), or SQL CREATE INDEX ... LSM_VECTOR METADATA {...}.

The parameters, their defaults, and tuning guidance are documented once, under create_vector_index and in the Vector API.


get_or_create_index

schema.get_or_create_index(
    type_name: str,
    property_names: List[str],
    unique: bool = False,
    index_type: Union[str, IndexType] = IndexType.LSM_TREE
) -> Any

Get an existing index or create it if it doesn't exist. Idempotent alternative to create_index with the same parameters. Returns the underlying Java Index object.

Parameters:

  • type_name (str): Name of the type
  • property_names (List[str]): List of property names to index
  • unique (bool): Whether the index should enforce uniqueness (default: False)
  • index_type (str or IndexType): Type of index (default: IndexType.LSM_TREE)

Returns:

  • Java Index object

Raises:

  • ArcadeDBError: If the index type is invalid or get/create fails

Example:

idx = db.schema.get_or_create_index("User", ["email"], unique=True)

drop_index

schema.drop_index(index_name: str, force: bool = False)

Drop an index by name.

Parameters:

  • index_name (str): Name of the index to drop (e.g. "User[email]")
  • force (bool): If True, skip the existence check (useful for corrupted/partial indexes; default: False)

Raises:

  • ArcadeDBError: If the index doesn't exist (when force=False) or the drop fails

Example:

db.schema.drop_index("User[email]")

# Force drop corrupted index
db.schema.drop_index("User[email]", force=True)

get_indexes

schema.get_indexes() -> List[Any]

Get all indexes in the schema.

Returns:

  • List[Any]: All indexes as Java Index objects (camelCase JPype methods, e.g. getName(), getType())

Example:

for idx in db.schema.get_indexes():
    print(idx.getName())

exists_index

schema.exists_index(index_name: str) -> bool

Check if an index exists.

Parameters:

  • index_name (str): Name of the index

Returns:

  • bool: True if the index exists

Example:

if db.schema.exists_index("User[email]"):
    print("Index exists")

get_index_by_name

schema.get_index_by_name(index_name: str) -> Optional[Any]

Get an index by name.

Parameters:

  • index_name (str): Name of the index

Returns:

  • Java Index object, or None if not found

Example:

idx = db.schema.get_index_by_name("User[email]")
if idx:
    print(f"Index type: {idx.getType()}")

get_vector_index

schema.get_vector_index(
    vertex_type: str,
    vector_property: str
) -> Optional[VectorIndex]

Get an existing vector (LSM_VECTOR/JVector) index for a type/property pair, wrapped as a Python VectorIndex. This is the standard way to obtain a search handle for an index created via SQL CREATE INDEX ... LSM_VECTOR.

Parameters:

  • vertex_type (str): Name of the vertex type
  • vector_property (str): Name of the vector property

Returns:

  • VectorIndex object, or None if no vector index covers that property

Example:

index = db.schema.get_vector_index("Document", "embedding")
if index:
    neighbors = index.find_nearest(query_vector, k=5)

list_vector_indexes

schema.list_vector_indexes() -> List[str]

List the names of all vector indexes in the database.

Returns:

  • List[str]: Vector index names (empty list if none)

Example:

for name in db.schema.list_vector_indexes():
    print(name)

Type Inspection

get_type

schema.get_type(name: str) -> Optional[Type]

Get a type by name.

Parameters:

  • name (str): Name of the type

Returns:

  • Optional[Type]: Java DocumentType/VertexType/EdgeType object, or None if not found. Methods on this object are the underlying Java methods exposed by JPype (camelCase, e.g. getName(), getProperties()). To count a type's records, use db.count_type("User").

Example:

# Check if type exists
user_type = schema.get_type("User")
if user_type:
    print(f"User type exists: {user_type.getName()}")
else:
    print("User type not found")

exists_type

schema.exists_type(name: str) -> bool

Check if type exists.

Parameters:

  • name (str): Name of the type

Returns:

  • bool: True if exists

Example:

if not schema.exists_type("User"):
    schema.create_vertex_type("User")  # schema statements apply immediately

get_types

schema.get_types() -> List[Type]

Get all types.

Returns:

  • List[Any]: All types in the schema as Java type objects

Example:

for type_obj in schema.get_types():
    print(f"Type: {type_obj.getName()}")

Type Methods

get_type() and get_types() return the raw Java type objects. Their methods are the underlying Java methods exposed by JPype using camelCase names (for example getName(), getProperties(), getIndexesByProperties(...)).

type_obj = schema.get_type("User")

# Get type info (Java methods via JPype)
name = type_obj.getName()

# Get properties (Java Property objects)
for prop in type_obj.getProperties():
    print(f"{prop.getName()}: {prop.getType()}")

Prefer SQL for inspection

For schema inspection in application code, prefer SQL such as db.query("sql", "SELECT FROM schema:types") / db.query("sql", "SELECT FROM schema:indexes"), which returns Python-friendly Result rows instead of raw Java objects.

Complete Example

import arcadedb_embedded as arcadedb

# Create database with context manager to ensure clean close
with arcadedb.create_database("./social_network") as db:
    # Create schema (applies immediately)
    # User vertex type
    db.schema.create_vertex_type("User")
    db.schema.create_property("User", "username", "STRING")
    db.schema.create_property("User", "email", "STRING")
    db.schema.create_property("User", "age", "INTEGER")
    db.schema.create_property("User", "tags", "LIST", of_type="STRING")
    db.schema.create_property("User", "createdAt", "DATETIME")

    # Post vertex type
    db.schema.create_vertex_type("Post")
    db.schema.create_property("Post", "title", "STRING")
    db.schema.create_property("Post", "content", "STRING")
    db.schema.create_property("Post", "timestamp", "DATETIME")

    # Follows edge type
    db.schema.create_edge_type("Follows")
    db.schema.create_property("Follows", "since", "DATETIME")

    # Likes edge type
    db.schema.create_edge_type("Likes")
    db.schema.create_property("Likes", "timestamp", "DATETIME")

    # Create indexes
    db.schema.create_index("User", ["username"], unique=True)
    db.schema.create_index("Post", ["timestamp"])

    # Verify schema (get_types() returns raw Java type objects)
    print("\n📋 Schema Summary:")
    for type_obj in db.schema.get_types():
        print(f"\nType: {type_obj.getName()}")
        for prop in type_obj.getProperties():
            print(f"    - {prop.getName()}: {prop.getType()}")

Schema Evolution

# Add property to an existing type (applies immediately)
if not db.schema.get_type("User").existsProperty("phoneNumber"):
    db.schema.create_property("User", "phoneNumber", "STRING")
    print("✅ Added phoneNumber property")

# get_or_create_* helpers make evolution idempotent
db.schema.get_or_create_property("User", "phoneNumber", "STRING")
db.schema.get_or_create_index("User", ["email"], unique=True)

Best Practices

1. Schema Statements Apply Immediately; Batch Many in One Transaction

# ✅ One or two statements: no transaction needed
schema.create_vertex_type("User")
schema.create_edge_type("Follows")

# ✅ Many statements: one transaction, so the schema is written to disk once, at the end
with db.transaction():
    for name in ("Author", "Book", "Review", "Shelf"):
        db.command("sql", f"CREATE DOCUMENT TYPE {name}")
        db.command("sql", f"CREATE PROPERTY {name}.id LONG")
        db.command("sql", f"CREATE INDEX ON {name} (id) UNIQUE_HASH")

A schema statement is not transactional: a rollback does not undo it, so a failed block can leave the types it already created behind.

2. Check Existence Before Creating

# ✅ Good: Check first
if not schema.exists_type("User"):
    schema.create_vertex_type("User")

# ❌ Bad: Don't check
schema.create_vertex_type("User")  # May error if exists

3. Create Indexes for Frequent Queries

# ✅ Good: Index frequently queried properties
schema.create_index("Event", ["timestamp"])
schema.create_index("Event", ["userId", "timestamp"])

4. Choose Bucket Counts for Parallel Loads

# ✅ Good: one bucket per async writer (or a multiple), decided when the type is created
buckets = db.async_executor().get_parallel_level()  # default: cores - 1
schema.create_document_type("Event", buckets=buckets)
  • One bucket (the default) is right unless the type is loaded with db.insert_many(..., parallel=True). A parallel load needs as many buckets as the async executor has writers, or a multiple of that (ArcadeData/arcadedb#8478).
  • Each bucket has its own sub-index, so on a multi-bucket type an index lookup touches every bucket's sub-index. On a type with a key, route records by it: ALTER TYPE Event BucketSelectionStrategy `partitioned('id')`. See the Bulk Ingest Recommendation.

5. Set Property Constraints

Property constraints are not exposed by the Python Schema wrapper. Apply them via SQL DDL, or on the Java Property object returned by create_property (camelCase methods):

# ✅ Recommended: SQL DDL for constraints (schema statements apply immediately)
db.schema.create_vertex_type("User")
db.command("sql", "CREATE PROPERTY User.username STRING (mandatory true, notnull true)")
db.command("sql", "CREATE PROPERTY User.age INTEGER (min 0, max 150)")

# Or configure the returned Java Property object directly
username = db.schema.create_property("User", "nickname", "STRING")
username.setMandatory(True)
username.setNotNull(True)

Common Patterns

1. Schema Initialization

def init_schema(db):
    """Initialize schema if not exists"""
    if db.schema.exists_type("User"):
        print("Schema already initialized")
        return

    # Schema statements apply immediately; several of them in one transaction write the schema once
    with db.transaction():
        db.schema.create_vertex_type("User")
        db.schema.create_property("User", "username", "STRING")
        db.schema.create_property("User", "email", "STRING")

        db.schema.create_index("User", ["username"], unique=True)

    print("✅ Schema initialized")

# Use it
init_schema(db)

2. Schema Export

def export_schema(db):
    """Export schema to dict (get_types() yields raw Java type objects)."""
    schema_dict = {}

    for type_obj in db.schema.get_types():
        type_name = str(type_obj.getName())
        schema_dict[type_name] = {"properties": {}, "indexes": []}

        # Export properties (Java Property objects via JPype)
        for prop in type_obj.getProperties():
            schema_dict[type_name]["properties"][str(prop.getName())] = str(
                prop.getType()
            )

    return schema_dict

# Use it
schema_export = export_schema(db)
print(json.dumps(schema_export, indent=2))

Tip

For schema export you can also query db.query("sql", "SELECT FROM schema:types") and db.query("sql", "SELECT FROM schema:indexes"), which return Python Result rows instead of raw Java objects.

3. Schema Migration

def migrate_schema_v1_to_v2(db):
    """Migrate schema from v1 to v2 (idempotent via get_or_create_*)."""
    # Add new property if missing
    db.schema.get_or_create_property("User", "status", "STRING")
    print("✅ Ensured status property")

    # Add new index if missing
    db.schema.get_or_create_index("User", ["status"])
    print("✅ Ensured status index")

# Use it
migrate_schema_v1_to_v2(db)

Troubleshooting

Type Already Exists Error

# ✅ Good: Check first
if not schema.exists_type("User"):
    schema.create_vertex_type("User")

Property Not Found

# ✅ Good: Check property exists on the Java type object (camelCase JPype methods)
user_type = schema.get_type("User")
if user_type is not None and user_type.existsProperty("email"):
    print(f"Email type: {user_type.getProperty('email').getType()}")
else:
    print("Email property not found")

Index Creation Fails

# ✅ Good: Create index (applies immediately)
schema.create_index("User", ["username"], unique=True)

See Also