Description
Preemptible instances in a Databricks GCP cluster can be reclaimed, interrupting work or requiring reprocessing. Balance cost savings with interruption tolerance, on-demand nodes and a fallback strategy.
Potential impact
- Preemption can interrupt running work.
- Cluster recovery and reprocessing can increase costs.
- Completion times for production analysis can become less predictable.
Remediation
Choose PREEMPTIBLE_WITH_FALLBACK_GCP or ON_DEMAND_GCP according to workload requirements. With a supported provider version, specify the required on-demand nodes through first_on_demand, and test recovery and costs after preemption. Fallback and automatic termination settings alone do not guarantee uninterrupted execution.
Examples
These Databricks provider v1.133.0 excerpts omit the data source definitions. The before example uses PREEMPTIBLE_GCP without explicitly setting an on-demand node count.
Before
resource "databricks_cluster" "example" {
cluster_name = "data"
spark_version = data.databricks_spark_version.latest.id
node_type_id = data.databricks_node_type.smallest.id
autotermination_minutes = 20
autoscale {
min_workers = 1
max_workers = 50
}
gcp_attributes {
availability = "PREEMPTIBLE_GCP"
zone_id = "AUTO"
}
}
After
resource "databricks_cluster" "example" {
cluster_name = "Shared Autoscaling"
spark_version = data.databricks_spark_version.latest.id
node_type_id = data.databricks_node_type.smallest.id
autotermination_minutes = 20
autoscale {
min_workers = 1
max_workers = 50
}
gcp_attributes {
availability = "PREEMPTIBLE_WITH_FALLBACK_GCP"
zone_id = "AUTO"
first_on_demand = 1
}
}
The after example places the driver on demand with first_on_demand = 1 and allows on-demand fallback for the remaining nodes if preemptible capacity is unavailable. zone_id = "AUTO" automatically selects one availability zone; it does not spread nodes across zones.