Review Databricks spark_submit_task usage

Use a task type suited to the Databricks workload where possible.

Description

Databricks spark_submit_task runs only on new clusters and has restrictions on how libraries and Spark settings are supplied. Where possible, use spark_jar_task, spark_python_task, or notebook_task to match the workload.

Potential impact

An incompatible cluster or argument configuration can cause the job to fail. Using spark_submit_task alone does not establish a security compromise.

Remediation

Migrate to a task type suited to the code while preserving its entry point, arguments, libraries, and execution requirements. Verify that the revised job produces the intended results.

Examples

Both excerpts run com.acme.data.Main from the same JAR. The first uses a new cluster; the second uses an existing shared cluster. The JAR-location input var.application_jar, data sources, and cluster definitions are omitted.

Before

hcl
resource "databricks_job" "example" {
  name = "Job with multiple tasks"

  task {
    task_key = "b"
    new_cluster {
      num_workers   = 1
      spark_version = data.databricks_spark_version.latest.id
      node_type_id  = data.databricks_node_type.smallest.id
    }

    spark_submit_task {
      parameters = ["--class", "com.acme.data.Main", var.application_jar]
    }
  }
}

After

hcl
resource "databricks_job" "example" {
  name = "Job with multiple tasks"

  task {
    task_key            = "b"
    existing_cluster_id = databricks_cluster.shared.id

    library {
      jar = var.application_jar
    }

    spark_jar_task {
      main_class_name = "com.acme.data.Main"
    }
  }
}

References