Description
Spot nodes in a Databricks AWS cluster can be reclaimed, interrupting work or requiring it to run again. Choose the on-demand node count and fallback behavior according to the importance of the workload and its tolerance for interruption.
availability applies to nodes after those reserved by first_on_demand. A value of SPOT therefore does not by itself mean that every node is a spot instance.
Potential impact
- Spot reclamation can interrupt or restart work.
- Batch processing and shared analysis may take longer to recover.
- On-demand fallback and reprocessing can change the expected cost.
Remediation
Use on-demand nodes or consider SPOT_WITH_FALLBACK for workloads that cannot readily tolerate interruption. Set first_on_demand to at least 1 to keep the driver on demand, and check availability-zone selection and recovery procedures. Test costs and retry behavior because fallback does not guarantee uninterrupted execution.
Examples
These cluster excerpts omit the data source definitions. Both set first_on_demand = 1, placing the driver on demand; the before example uses SPOT for the remaining nodes.
Before
resource "databricks_cluster" "example" {
cluster_name = "data"
spark_version = data.databricks_spark_version.latest.id
node_type_id = data.databricks_node_type.smallest.id
autotermination_minutes = 20
autoscale {
min_workers = 1
max_workers = 50
}
aws_attributes {
availability = "SPOT"
zone_id = "auto"
first_on_demand = 1
spot_bid_price_percent = 100
}
}
After
resource "databricks_cluster" "example" {
cluster_name = "Shared Autoscaling"
spark_version = data.databricks_spark_version.latest.id
node_type_id = data.databricks_node_type.smallest.id
autotermination_minutes = 20
autoscale {
min_workers = 1
max_workers = 50
}
aws_attributes {
availability = "SPOT_WITH_FALLBACK"
zone_id = "auto"
first_on_demand = 1
spot_bid_price_percent = 100
}
}
The after example permits on-demand fallback when spot capacity cannot be obtained. zone_id = "auto" automatically selects a zone at launch; it does not distribute a running cluster across multiple availability zones.