Lifecycle after creation
Scale, backup, restore, upgrade and recovery are first-class actions.

DATABASE LIFECYCLE CONTROL PLANE
A lifecycle control plane for ten database engines with 1/2/3-node topologies, engine-native HA semantics, backup/restore, upgrade and runtime verification.
DBaaS Enterprise components: one control plane and ten database engines
Product problem
Replication, quorum, backup, restore, upgrade and failover differ by engine. DBaaS preserves those differences while presenting one durable operational surface.
Scale, backup, restore, upgrade and recovery are first-class actions.
InnoDB, Patroni, Sentinel and Galera keep their native quorum/election behavior.
Promotion, fencing and rejoin rules exist before an incident.
Membership, endpoint and client I/O must confirm the result.
Capabilities
The operator sees a coherent lifecycle while each engine keeps the coordination model it actually requires.
Engine, version, topology and nodes with preflight.
Add members without treating two nodes as automatic HA.
Engine-native backup, integrity evidence and post-restore verification.
Supported version edge, backup gate and runtime read-back.
Engine-aware promotion or Galera routing reconciliation.
Tenant FQDN remains stable across topology changes.
Host/database telemetry and service-level probes.
Access and every operation stay controlled and traceable in the control plane.
Deployment model
The relational and key-value profiles (MySQL, PostgreSQL, Redis, Valkey and MariaDB) start at one node, grow through two nodes without automatic failover, and reach automatic HA at three members using the engine's own coordination model.
| Profile | Engines | Failure behavior | Native mechanism |
|---|---|---|---|
| 1 node | MySQL · PostgreSQL · Redis · Valkey · MariaDB | Non-HA | Single member + stable endpoint |
| 2 nodes | MySQL · PostgreSQL · Redis · Valkey · MariaDB | Automatic failover disabled | Primary/Mirror or two-member Galera without majority |
| 3 nodes | MySQL · PostgreSQL · Redis · Valkey · MariaDB | Automatic HA | InnoDB Cluster · Patroni/etcd · Sentinel · Galera quorum |
Architecture
Intent and durable job state live in the manager; database runtime, membership and local routing stay on tenant database nodes.
| Boundary | Mechanism | Operational contract |
|---|---|---|
| Manager | Control plane only | Never a DB member, witness, quorum voter or failover target. |
| Stable endpoint | Tenant FQDN | Client identity survives topology changes. |
| Routing | Local Router/HAProxy on DB nodes | A client reaching any service node can still reach the correct role. |
| Coordination | Engine-native quorum/election | No external witness is added to invent HA. |
| Stage | Authority / Component | Technical contract |
|---|---|---|
| Admission | Panel / Product API | Engine, topology, permission and confirmation are checked; the request is only admitted. |
| Durable state | PostgreSQL | Operation and Job persist before mutation; only one active mutation is allowed per tenant; a second request is refused, not queued. |
| Dispatch | Transactional Outbox → RabbitMQ → Celery Worker | Work publication is coupled to transactional state and the worker receives serialized execution. |
| Execution fence | Execution lease + signed source manifest | The runtime tree must match expected release identity; drift blocks playbook execution. |
| Engine execution | Project-local Ansible | The native engine playbook runs only on tenant DB nodes; the Manager never enters the Data Plane. |
| Verification | Engine runtime + stable endpoint | Membership/quorum, replication, DNS/TLS and client read/write are read back from the real runtime. |
| Evidence | Job-scoped evidence | Result and correlation stay attached to the same Operation; HTTP acceptance never becomes completion by itself. |
Lifecycle
Every mutation is admitted, persisted, executed, read back and either verified or moved into an explicit recovery state.
Engine profiles
The control plane is shared, but replication, election/quorum, local routing, backup and failover/recovery follow each engine's native architecture.

| Engine | Native topology / routing | 1 node | 2 nodes | 3 nodes | Data protection |
|---|---|---|---|---|---|
| MySQL | InnoDB ReplicaSet (1–2) · InnoDB Cluster / Group Replication (3) · MySQL Router on DB nodes | Non-HA | Manual promotion | Automatic HA | XtraBackup |
| PostgreSQL | Patroni + Streaming Replication · etcd + PgBouncer + HAProxy on DB nodes | Non-HA | Manual promotion | Patroni automatic HA | pg_basebackup |
| Redis OSS | Replication + Sentinel + HAProxy; Sentinel and proxy stay on DB nodes | Non-HA | Manual promotion | Sentinel automatic HA | RDB / AOF |
| Valkey | Replication + Sentinel + HAProxy; a separate engine, not a Redis alias | Non-HA | Manual promotion | Sentinel automatic HA | RDB / AOF |
| MariaDB | Galera (wsrep) + HAProxy; multi-primary cluster with product-routed write target | Single-member Galera | No automatic failover | Galera quorum HA | MariaBackup / Logical Dump |
| MongoDB | Document | Topology, HA and data protection follow the engine's own native semantics, under the same control plane and lifecycle. | |||
| ClickHouse | Columnar analytics (OLAP) | Topology, HA and data protection follow the engine's own native semantics, under the same control plane and lifecycle. | |||
| OpenSearch | Search & log analytics | Topology, HA and data protection follow the engine's own native semantics, under the same control plane and lifecycle. | |||
| ScyllaDB | Wide-column | Topology, HA and data protection follow the engine's own native semantics, under the same control plane and lifecycle. | |||
| Qdrant | Vector | Topology, HA and data protection follow the engine's own native semantics, under the same control plane and lifecycle. | |||
State and flows
One active mutation per tenant prevents overlapping topology changes; runtime evidence closes the job.
The stable endpoint moves only after the new serving role is proven.
A recovery point must verify before it can enter the serving path.
Observed membership decides whether a node can rejoin or must be replaced.
From intent to service-level read-back.
Grow without inventing two-node HA.
Protection is only useful when recovery verifies.
Version change is staged and read back.
Fencing and quorum precede promotion.
Two Day-2 paths that matter as much as initial provisioning.
Return a recovered member without corrupting current membership.
Remove or replace hardware while protecting quorum and retained members.
| Action | Admission / Preconditions | Execution owner | Success criterion | Failure / recovery |
|---|---|---|---|---|
| Create Service | Valid tenant/project, engine/version, nodes | DBaaS worker + engine playbook | Membership + endpoint + client I/O | Failed job retains evidence; reconcile before retry |
| Scale | Healthy current cluster + exact new node | Serialized tenant mutation | Exact member count and replication/quorum | Partial join is reconciled |
| Backup | Compatible engine and destination | Engine-native backup | Artifact + integrity metadata | Incomplete backup is invalid |
| Restore | Verified backup + compatibility + maintenance gate | Cluster-locked restore | Service + replication + endpoint + I/O | Unknown outcome requires read-back |
| Upgrade | Supported edge + backup + health | Staged engine-aware upgrade | Exact version + health + client I/O | Unsupported/downgrade edge rejected; stops and recovers at the failed stage |
| Failover | Eligible topology + fencing/quorum | Promotion or Galera routing reconcile | New serving role + stable endpoint | Rejoin/reconcile remains separate |
Operator surface
Operators work with lifecycle intent rather than memorized shell sequences.
Engine-aware deployment with preflight and runtime verification.
Prepare, join, sync and verify topology.
Scheduled backup with retention and verify; native backup with integrity evidence.
Compatibility and service-level verification.
Supported edge and staged runtime read-back.
Planned switchover and failover with quorum/fencing-aware promotion or reconciliation.
Return a recovered node to valid membership.
Topology, replication, endpoint and telemetry state.
Create, grant, rotate passwords and drop through the same job path.
Certificate status, renewal, reissue and product CA rotation.
Resource-aware configuration profiles applied with ordered, safe restarts.
Remote MCP without a parallel authority. OAuth 2.1 Authorization Code + PKCE: the Panel is the Authorization Server and the MCP Addon the Resource Server; grants bind to the existing mcp_clients, RBAC, scope and rate limits instead of a second access model.
The control plane stays consistent while each engine retains its real topology and failure semantics.