Design quorum and failure domains
Cluster node count, quorum behaviour and physical placement should match expected failures. Understand what happens when a node, switch, rack or site is unavailable before relying on automatic recovery.
Match storage to workload behaviour
Local ZFS, shared storage and Ceph have very different performance, resilience and operational characteristics. Size IOPS, latency, capacity, replication overhead and recovery behaviour against the real VM estate rather than average utilisation alone.
Separate management and workload concerns
Plan management, cluster, migration, storage and guest traffic with appropriate redundancy and segmentation. Network design directly affects cluster stability and storage performance, especially during failure or rebalance activity.
Back up outside the failure domain
Use protected backup targets, retention and restore testing that survive the loss scenarios the cluster itself is designed to handle. High availability reduces some outages; it does not replace backup or disaster recovery.
AL Group can assess, design, implement and operate the underlying technology rather than stopping at advice.
Talk to an engineer