Schema API¶
The Schema provides type management, index creation, and property definitions for documents, vertices, and edges.
Overview¶
The Schema class enables:
- Type Management: Create document, vertex, and edge types
- Property Definitions: Define typed properties (constraints go through SQL DDL, see below)
- Index Creation: Create indexes for query optimization
- Schema Inspection: Query existing schema definitions
DSL-first recommendation
For application code, prefer ArcadeDB SQL DDL via db.command("sql", ...).
This page documents the Schema wrapper API for reference and advanced typed schema workflows.
Getting Schema¶
import arcadedb_embedded as arcadedb
# Use context manager to ensure clean close
with arcadedb.create_database("./mydb") as db:
# Create types (schema statements apply immediately)
# Vertex type
user_type = db.schema.create_vertex_type("User")
# Edge type
follows_type = db.schema.create_edge_type("Follows")
# Document type
log_type = db.schema.create_document_type("LogEntry")
Schema statements apply immediately
Creating or dropping a type, property, or index needs no transaction, and it is not
transactional: it takes effect at once, and a rollback does not undo it. To create many
types, run the statements inside one with db.transaction(): or send them as one
sqlscript. The schema is then written to disk once, when the transaction ends, instead of
once per statement (ArcadeData/arcadedb#8635). See
Transactions.
Type Creation Methods¶
create_vertex_type¶
Create a new vertex type. Returns the underlying Java VertexType object.
Parameters:
name(str): Name of the vertex typebuckets(Optional[int]): Number of buckets (engine default 1 when omitted)
Returns:
VertexType: Created vertex type
Example:
# Schema statements apply immediately (no transaction needed)
# Basic vertex type
user_type = schema.create_vertex_type("User")
# With custom buckets
product_type = schema.create_vertex_type("Product", buckets=10)
create_edge_type¶
Create a new edge type. Returns the underlying Java EdgeType object.
Parameters:
name(str): Name of the edge typebuckets(Optional[int]): Number of buckets (engine default 1 when omitted)
Returns:
EdgeType: Created edge type
Example:
# Schema statements apply immediately (no transaction needed)
# Basic edge type
follows_type = schema.create_edge_type("Follows")
# With custom buckets
purchased_type = schema.create_edge_type("Purchased", buckets=5)
create_document_type¶
Create a new document type. Returns the underlying Java DocumentType object.
Parameters:
name(str): Name of the document typebuckets(Optional[int]): Number of buckets (engine default 1 when omitted)
Returns:
DocumentType: Created document type
Example:
# Schema statements apply immediately (no transaction needed)
# Basic document type
log_type = schema.create_document_type("LogEntry")
# With custom buckets
event_type = schema.create_document_type("Event", buckets=8)
get_or_create_document_type¶
Get an existing document type or create it if it doesn't exist. Idempotent alternative
to create_document_type. Returns the underlying Java DocumentType object.
Parameters:
name(str): Type namebuckets(Optional[int]): Number of buckets if creating a new type (engine default when omitted)
Returns:
DocumentType: Existing or newly created document type
Example:
get_or_create_vertex_type¶
Get an existing vertex type or create it if it doesn't exist. Idempotent alternative to
create_vertex_type. Returns the underlying Java VertexType object.
Parameters:
name(str): Type namebuckets(Optional[int]): Number of buckets if creating a new type (engine default when omitted)
Returns:
VertexType: Existing or newly created vertex type
Example:
get_or_create_edge_type¶
Get an existing edge type or create it if it doesn't exist. Idempotent alternative to
create_edge_type. Returns the underlying Java EdgeType object.
Parameters:
name(str): Type namebuckets(Optional[int]): Number of buckets if creating a new type (engine default when omitted)
Returns:
EdgeType: Existing or newly created edge type
Example:
drop_type¶
Drop a type and all its data.
Parameters:
name(str): Type name to drop
Raises:
ArcadeDBError: If the drop fails
Example:
Property Definition¶
create_property¶
schema.create_property(
type_name: str,
property_name: str,
property_type: Union[str, PropertyType],
of_type: Optional[str] = None
) -> Any
Create a property on a type. Returns the underlying Java Property object.
Parameters:
type_name(str): Name of the typeproperty_name(str): Name of the propertyproperty_type(str orPropertyType): ArcadeDB type (see types below)of_type(Optional[str]): Element type forLIST/MAPcollections
Returns:
- Java
Propertyobject
Property Types:
The PropertyType enum (importable as from arcadedb_embedded import PropertyType)
defines the supported values:
- Primitives:
STRING,INTEGER,LONG,SHORT,BYTE,BOOLEAN,FLOAT,DOUBLE,DECIMAL,DATE,DATETIME - Binary:
BINARY - Collections:
LIST,MAP,EMBEDDED - Links:
LINK - Vectors:
ARRAY_OF_FLOATS
Either the enum member or its string name may be passed. A string may also name an
engine type the enum does not list: DATETIME_SECOND, DATETIME_MICROS,
DATETIME_NANOS, ARRAY_OF_SHORTS, ARRAY_OF_INTEGERS, ARRAY_OF_LONGS, or
ARRAY_OF_DOUBLES.
Example:
# Schema statements apply immediately (no transaction needed)
schema.create_vertex_type("User")
# String property
schema.create_property("User", "name", "STRING")
# Integer property (enum form)
from arcadedb_embedded import PropertyType
schema.create_property("User", "age", PropertyType.INTEGER)
# Date property
schema.create_property("User", "birthDate", "DATE")
# List property
schema.create_property("User", "tags", "LIST", of_type="STRING")
# Embedded property
schema.create_property("User", "profile", "EMBEDDED")
get_or_create_property¶
schema.get_or_create_property(
type_name: str,
property_name: str,
property_type: Union[str, PropertyType],
of_type: Optional[str] = None
) -> Any
Get an existing property or create it if it doesn't exist. Same parameters as
create_property, except that of_type must be a string here: a PropertyType
member raises ArcadeDBError. Returns the underlying Java Property object.
drop_property¶
Drop a property from a type.
Property constraints
Constraints such as mandatory, not-null, default, and min/max are configured on the
returned Java Property object (for example prop.setMandatory(True)) or via SQL
DDL. The Python Schema wrapper itself only exposes property creation/removal.
Index Creation¶
create_index¶
schema.create_index(
type_name: str,
property_names: List[str],
unique: bool = False,
index_type: Union[str, IndexType] = IndexType.LSM_TREE
) -> Index
Create an index on a type.
Parameters:
type_name(str): Name of the typeproperty_names(List[str]): List of property names to indexunique(bool): Whether the index should enforce uniqueness (default:False)index_type(str or IndexType): Type of index ("LSM_TREE","HASH","FULL_TEXT","LSM_VECTOR","GEOSPATIAL"); default:IndexType.LSM_TREE
Returns:
Index: Created index object
Raises:
ArcadeDBError: If type doesn't exist or index creation fails
Example:
# Schema statements apply immediately (no transaction needed)
# Unique id read only by equality: a unique hash index (ArcadeData/arcadedb#9169)
schema.create_index("User", ["username"], unique=True, index_type="HASH")
# Unique key you also range over or sort by: the default LSM_TREE
schema.create_index("Ticket", ["number"], unique=True)
# Non-unique exact-match lookup index
schema.create_index("Order", ["customerId"], index_type="HASH")
# Composite index
schema.create_index("Event", ["userId", "timestamp"])
# Full-text index
schema.create_index("Article", ["content"], index_type="FULL_TEXT")
Index choice rules of thumb:
- Use
HASHfor an id that is only read, updated and deleted by equality and is not bulk-loaded in key order. From 26.10.1 a unique hash index answers SQLid = ?1.5 to 2.3 times faster thanLSM_TREEand anIndex.get()hit 1.9 to 3.1 times faster. Its insert is 1.14 to 1.23 times faster for shuffled ids but 9% to 16% slower for ids loaded in ascending order (ArcadeDB #9169, 200,000 and 2,000,000 entries). It cannot serve a range or anORDER BY. On 26.9.1 its inserts are several times slower thanLSM_TREE. - Use
LSM_TREEwhen you need ranges, sorting, or a safe general-purpose default, and for a non-unique column with few distinct values. - Use
FULL_TEXT,LSM_VECTOR, andGEOSPATIALonly for their specialized query types. HASHcan be unique or non-unique. The index structure and uniqueness constraint are separate choices.
SQL DSL equivalents:
CREATE INDEX ON User (email) UNIQUE-> uniqueLSM_TREE(ranges andORDER BY)CREATE INDEX ON User (email) NOTUNIQUE-> non-uniqueLSM_TREECREATE INDEX ON User (email) UNIQUE_HASH-> uniqueHASH(equality only, the choice for an id)CREATE INDEX ON Order (customerId) NOTUNIQUE_HASH-> non-uniqueHASHCREATE INDEX ON Article (content) FULL_TEXT->FULL_TEXTCREATE INDEX ON Doc (embedding) LSM_VECTOR ...->LSM_VECTORCREATE INDEX ON Place (location) GEOSPATIAL->GEOSPATIAL
Vector (JVector) Parameters:
Note
schema.create_index() only accepts type_name, property_names, unique, and
index_type. It does not take the JVector tuning parameters below. To configure a
vector index from Python, use db.create_vector_index()
(which exposes max_connections, beam_width, dimensions, etc.), or SQL
CREATE INDEX ... LSM_VECTOR METADATA {...}.
The parameters, their defaults, and tuning guidance are documented once, under
create_vector_index and in the
Vector API.
get_or_create_index¶
schema.get_or_create_index(
type_name: str,
property_names: List[str],
unique: bool = False,
index_type: Union[str, IndexType] = IndexType.LSM_TREE
) -> Any
Get an existing index or create it if it doesn't exist. Idempotent alternative to
create_index with the same parameters. Returns the underlying Java Index object.
Parameters:
type_name(str): Name of the typeproperty_names(List[str]): List of property names to indexunique(bool): Whether the index should enforce uniqueness (default:False)index_type(str or IndexType): Type of index (default:IndexType.LSM_TREE)
Returns:
- Java
Indexobject
Raises:
ArcadeDBError: If the index type is invalid or get/create fails
Example:
drop_index¶
Drop an index by name.
Parameters:
index_name(str): Name of the index to drop (e.g."User[email]")force(bool): IfTrue, skip the existence check (useful for corrupted/partial indexes; default:False)
Raises:
ArcadeDBError: If the index doesn't exist (whenforce=False) or the drop fails
Example:
db.schema.drop_index("User[email]")
# Force drop corrupted index
db.schema.drop_index("User[email]", force=True)
get_indexes¶
Get all indexes in the schema.
Returns:
List[Any]: All indexes as JavaIndexobjects (camelCase JPype methods, e.g.getName(),getType())
Example:
exists_index¶
Check if an index exists.
Parameters:
index_name(str): Name of the index
Returns:
bool: True if the index exists
Example:
get_index_by_name¶
Get an index by name.
Parameters:
index_name(str): Name of the index
Returns:
- Java
Indexobject, orNoneif not found
Example:
get_vector_index¶
Get an existing vector (LSM_VECTOR/JVector) index for a type/property pair, wrapped as
a Python VectorIndex. This is the standard way to obtain a search handle
for an index created via SQL CREATE INDEX ... LSM_VECTOR.
Parameters:
vertex_type(str): Name of the vertex typevector_property(str): Name of the vector property
Returns:
VectorIndexobject, orNoneif no vector index covers that property
Example:
index = db.schema.get_vector_index("Document", "embedding")
if index:
neighbors = index.find_nearest(query_vector, k=5)
list_vector_indexes¶
List the names of all vector indexes in the database.
Returns:
List[str]: Vector index names (empty list if none)
Example:
Type Inspection¶
get_type¶
Get a type by name.
Parameters:
name(str): Name of the type
Returns:
Optional[Type]: JavaDocumentType/VertexType/EdgeTypeobject, orNoneif not found. Methods on this object are the underlying Java methods exposed by JPype (camelCase, e.g.getName(),getProperties()). To count a type's records, usedb.count_type("User").
Example:
# Check if type exists
user_type = schema.get_type("User")
if user_type:
print(f"User type exists: {user_type.getName()}")
else:
print("User type not found")
exists_type¶
Check if type exists.
Parameters:
name(str): Name of the type
Returns:
bool: True if exists
Example:
if not schema.exists_type("User"):
schema.create_vertex_type("User") # schema statements apply immediately
get_types¶
Get all types.
Returns:
List[Any]: All types in the schema as Java type objects
Example:
Type Methods¶
get_type() and get_types() return the raw Java type objects. Their methods are the
underlying Java methods exposed by JPype using camelCase names (for example getName(),
getProperties(), getIndexesByProperties(...)).
type_obj = schema.get_type("User")
# Get type info (Java methods via JPype)
name = type_obj.getName()
# Get properties (Java Property objects)
for prop in type_obj.getProperties():
print(f"{prop.getName()}: {prop.getType()}")
Prefer SQL for inspection
For schema inspection in application code, prefer SQL such as
db.query("sql", "SELECT FROM schema:types") /
db.query("sql", "SELECT FROM schema:indexes"), which returns Python-friendly
Result rows instead of raw Java objects.
Complete Example¶
import arcadedb_embedded as arcadedb
# Create database with context manager to ensure clean close
with arcadedb.create_database("./social_network") as db:
# Create schema (applies immediately)
# User vertex type
db.schema.create_vertex_type("User")
db.schema.create_property("User", "username", "STRING")
db.schema.create_property("User", "email", "STRING")
db.schema.create_property("User", "age", "INTEGER")
db.schema.create_property("User", "tags", "LIST", of_type="STRING")
db.schema.create_property("User", "createdAt", "DATETIME")
# Post vertex type
db.schema.create_vertex_type("Post")
db.schema.create_property("Post", "title", "STRING")
db.schema.create_property("Post", "content", "STRING")
db.schema.create_property("Post", "timestamp", "DATETIME")
# Follows edge type
db.schema.create_edge_type("Follows")
db.schema.create_property("Follows", "since", "DATETIME")
# Likes edge type
db.schema.create_edge_type("Likes")
db.schema.create_property("Likes", "timestamp", "DATETIME")
# Create indexes
db.schema.create_index("User", ["username"], unique=True)
db.schema.create_index("Post", ["timestamp"])
# Verify schema (get_types() returns raw Java type objects)
print("\n📋 Schema Summary:")
for type_obj in db.schema.get_types():
print(f"\nType: {type_obj.getName()}")
for prop in type_obj.getProperties():
print(f" - {prop.getName()}: {prop.getType()}")
Schema Evolution¶
# Add property to an existing type (applies immediately)
if not db.schema.get_type("User").existsProperty("phoneNumber"):
db.schema.create_property("User", "phoneNumber", "STRING")
print("✅ Added phoneNumber property")
# get_or_create_* helpers make evolution idempotent
db.schema.get_or_create_property("User", "phoneNumber", "STRING")
db.schema.get_or_create_index("User", ["email"], unique=True)
Best Practices¶
1. Schema Statements Apply Immediately; Batch Many in One Transaction¶
# ✅ One or two statements: no transaction needed
schema.create_vertex_type("User")
schema.create_edge_type("Follows")
# ✅ Many statements: one transaction, so the schema is written to disk once, at the end
with db.transaction():
for name in ("Author", "Book", "Review", "Shelf"):
db.command("sql", f"CREATE DOCUMENT TYPE {name}")
db.command("sql", f"CREATE PROPERTY {name}.id LONG")
db.command("sql", f"CREATE INDEX ON {name} (id) UNIQUE_HASH")
A schema statement is not transactional: a rollback does not undo it, so a failed block can leave the types it already created behind.
2. Check Existence Before Creating¶
# ✅ Good: Check first
if not schema.exists_type("User"):
schema.create_vertex_type("User")
# ❌ Bad: Don't check
schema.create_vertex_type("User") # May error if exists
3. Create Indexes for Frequent Queries¶
# ✅ Good: Index frequently queried properties
schema.create_index("Event", ["timestamp"])
schema.create_index("Event", ["userId", "timestamp"])
4. Choose Bucket Counts for Parallel Loads¶
# ✅ Good: one bucket per async writer (or a multiple), decided when the type is created
buckets = db.async_executor().get_parallel_level() # default: cores - 1
schema.create_document_type("Event", buckets=buckets)
- One bucket (the default) is right unless the type is loaded with
db.insert_many(..., parallel=True). A parallel load needs as many buckets as the async executor has writers, or a multiple of that (ArcadeData/arcadedb#8478). - Each bucket has its own sub-index, so on a multi-bucket type an index lookup touches
every bucket's sub-index. On a type with a key, route records by it:
ALTER TYPE Event BucketSelectionStrategy `partitioned('id')`. See the Bulk Ingest Recommendation.
5. Set Property Constraints¶
Property constraints are not exposed by the Python Schema wrapper. Apply them via SQL
DDL, or on the Java Property object returned by create_property (camelCase methods):
# ✅ Recommended: SQL DDL for constraints (schema statements apply immediately)
db.schema.create_vertex_type("User")
db.command("sql", "CREATE PROPERTY User.username STRING (mandatory true, notnull true)")
db.command("sql", "CREATE PROPERTY User.age INTEGER (min 0, max 150)")
# Or configure the returned Java Property object directly
username = db.schema.create_property("User", "nickname", "STRING")
username.setMandatory(True)
username.setNotNull(True)
Common Patterns¶
1. Schema Initialization¶
def init_schema(db):
"""Initialize schema if not exists"""
if db.schema.exists_type("User"):
print("Schema already initialized")
return
# Schema statements apply immediately; several of them in one transaction write the schema once
with db.transaction():
db.schema.create_vertex_type("User")
db.schema.create_property("User", "username", "STRING")
db.schema.create_property("User", "email", "STRING")
db.schema.create_index("User", ["username"], unique=True)
print("✅ Schema initialized")
# Use it
init_schema(db)
2. Schema Export¶
def export_schema(db):
"""Export schema to dict (get_types() yields raw Java type objects)."""
schema_dict = {}
for type_obj in db.schema.get_types():
type_name = str(type_obj.getName())
schema_dict[type_name] = {"properties": {}, "indexes": []}
# Export properties (Java Property objects via JPype)
for prop in type_obj.getProperties():
schema_dict[type_name]["properties"][str(prop.getName())] = str(
prop.getType()
)
return schema_dict
# Use it
schema_export = export_schema(db)
print(json.dumps(schema_export, indent=2))
Tip
For schema export you can also query db.query("sql", "SELECT FROM schema:types")
and db.query("sql", "SELECT FROM schema:indexes"), which return Python Result
rows instead of raw Java objects.
3. Schema Migration¶
def migrate_schema_v1_to_v2(db):
"""Migrate schema from v1 to v2 (idempotent via get_or_create_*)."""
# Add new property if missing
db.schema.get_or_create_property("User", "status", "STRING")
print("✅ Ensured status property")
# Add new index if missing
db.schema.get_or_create_index("User", ["status"])
print("✅ Ensured status index")
# Use it
migrate_schema_v1_to_v2(db)
Troubleshooting¶
Type Already Exists Error¶
Property Not Found¶
# ✅ Good: Check property exists on the Java type object (camelCase JPype methods)
user_type = schema.get_type("User")
if user_type is not None and user_type.existsProperty("email"):
print(f"Email type: {user_type.getProperty('email').getType()}")
else:
print("Email property not found")
Index Creation Fails¶
See Also¶
- Database API - Database operations
- Transactions API - Transaction management
- Vector Search Guide - HNSW (JVector) indexes
- Example 03: Vector Search - A schema defined with SQL DDL, plus
get_vector_index - Example 02: Social Network - A graph schema defined with SQL DDL