Description
Spot VMs in a Databricks Azure cluster can reduce cost but may be reclaimed. For important analysis workloads, consider on-demand nodes and fallback alongside availability requirements and the cost of rerunning work.
SPOT_AZURE applies to nodes after those specified by first_on_demand. The examples set first_on_demand = 1, so their driver is on demand and the cluster is not entirely spot-based.
Potential impact
- Spot reclamation can interrupt analysis jobs.
- Shared clusters and scheduled workloads may take longer to recover.
- Retries and on-demand fallback can increase costs.
Remediation
Choose SPOT_WITH_FALLBACK_AZURE or on-demand capacity according to workload importance. Review first_on_demand to keep the driver and necessary nodes on demand, and test retries and costs after reclamation. Automatic termination controls idle costs; it does not prevent spot reclamation.
Examples
These excerpts omit the data source definitions and use Azure attributes supported by Databricks provider v1.133.0. The before example requests SPOT_AZURE for nodes after the on-demand driver.
Before
resource "databricks_cluster" "example" {
cluster_name = "data"
spark_version = data.databricks_spark_version.latest.id
node_type_id = data.databricks_node_type.smallest.id
autotermination_minutes = 20
autoscale {
min_workers = 1
max_workers = 50
}
azure_attributes {
availability = "SPOT_AZURE"
first_on_demand = 1
}
}
After
resource "databricks_cluster" "example" {
cluster_name = "Shared Autoscaling"
spark_version = data.databricks_spark_version.latest.id
node_type_id = data.databricks_node_type.smallest.id
autotermination_minutes = 20
autoscale {
min_workers = 1
max_workers = 50
}
azure_attributes {
availability = "SPOT_WITH_FALLBACK_AZURE"
first_on_demand = 1
}
}
The after example permits on-demand fallback when spot capacity is unavailable. Fallback helps obtain capacity but does not guarantee that running work will avoid interruption.