"Managed Service for Apache Spark" is the new name for the product formerly known as "Dataproc on Compute Engine" (cluster deployment) and "Google Cloud Serverless for Apache Spark" (serverless deployment).

Create a cluster

Create a cluster

Requirements:

  • Name: The cluster name must start with a lowercase letter followed by up to 51 lowercase letters, numbers, and hyphens, and cannot end with a hyphen.

  • Cluster region: You must specify a Compute Engine region for the cluster, such as us-east1 or europe-west1, to isolate cluster resources, such as VM instances and cluster metadata stored in Cloud Storage, within the region.

    • See Cluster region for more information on Compute Engine regions.
    • See Available regions & zones for information on selecting a region. You can also run the gcloud compute regions list command to display a listing of available regions.
  • Connectivity: Compute Engine Virtual Machine instances (VMs) in a Managed Service for Apache Spark cluster, consisting of master and worker VMs, require full internal IP networking cross connectivity. The default VPC network provides this connectivity (see Managed Service for Apache Spark Cluster Network Configuration).

  • Machine type (recommended): While specifying a machine type is optional, Google recommends that you explicitly select a machine type for the master and worker VMs in your cluster. If you don't specify a machine type, Managed Service for Apache Spark dynamically selects machine types based on resource availability. This dynamic selection can result in variations in both cost and performance.

    • For more information on choosing a machine type, see Supported machine types.
    • To mitigate potential resource unavailability issues, we recommend using Flexible VMs, which lets you specify a list of acceptable machine types.

Console

Open the Google Cloud console Create cluster page to display default cluster settings. You can confirm or change the displayed default settings, then click Additional configuration to further customize the cluster.

Click Create cluster to create the cluster. The cluster name appears in the Clusters page, and its status is updated to Running after the cluster is provisioned. Click the cluster name to open the cluster details page where you can examine jobs, instances, and configuration settings for your cluster and connect to web interfaces running on the cluster.

gcloud

To create a Managed Service for Apache Spark cluster on the command line, run the gcloud dataproc clusters create command locally in a terminal window or in Cloud Shell.

gcloud dataproc clusters create CLUSTER_NAME \
  --region=REGION \
  --master-machine-type=MASTER_MACHINE_TYPE \
  --worker-machine-type=WORKER_MACHINE_TYPE

The command creates a cluster. While master and worker machine types are optional, it is recommended to explicitly specify them using the --master-machine-type and --worker-machine-type flags (for example, n4-standard-4) to ensure consistent cost and performance. If you don't specify machine types, default machine types are selected dynamically based on resource availability. See the gcloud dataproc clusters create command for information on using command line flags to customize cluster settings.

Create a cluster with a YAML file

  1. Run the following gcloud command to export the configuration of an existing Managed Service for Apache Spark cluster into a cluster.yaml file.
    gcloud dataproc clusters export EXISTING_CLUSTER_NAME \
      --region=REGION \
      --destination=cluster.yaml
    
  2. Create a new cluster by importing the YAML file configuration.
    gcloud dataproc clusters import NEW_CLUSTER_NAME \
      --region=REGION \
      --source=cluster.yaml
    

**Note:** During the export operation, cluster-specific fields, such as cluster name, output-only fields, and automatically applied labels are filtered. These fields are disallowed in the imported YAML file used to create a cluster.

REST

This section shows how to create a cluster. While specifying machine types is optional, it is recommended to explicitly include machine_type_uri in your master_config and worker_config (for example, n4-standard-4) to ensure consistent cost and performance. If you don't specify machine types, default machine types are selected dynamically based on resource availability.

Before using any of the request data, make the following replacements:

  • CLUSTER_NAME: cluster name
  • PROJECT: Google Cloud project ID
  • REGION: An available Compute Engine region where the cluster will be created.
  • ZONE: An optional zone within the selected region where the cluster will be created.
  • MASTER_MACHINE_TYPE: (Recommended) The machine type for the master node (for example, n4-standard-4).
  • WORKER_MACHINE_TYPE: (Recommended) The machine type for worker nodes (for example, n4-standard-4).

HTTP method and URL:

POST https://dataproc.googleapis.com/v1/projects/PROJECT/regions/REGION/clusters

Request JSON body:

{
 "project_id":"PROJECT",
 "cluster_name":"CLUSTER_NAME",
 "config":{
 "master_config":{
 "num_instances":1,
 "machine_type_uri":"MASTER_MACHINE_TYPE",
 "image_uri":""
 },
 "softwareConfig": {
 "imageVersion": "",
 "properties": {},
 "optionalComponents": []
 },
 "worker_config":{
 "num_instances":2,
 "machine_type_uri":"WORKER_MACHINE_TYPE",
 "image_uri":""
 },
 "gce_cluster_config":{
 "zone_uri":"ZONE"
 }
 }
}

To send your request, expand one of these options:

curl (Linux, macOS, or Cloud Shell)

Save the request body in a file named request.json, and execute the following command:

curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json; charset=utf-8" \
-d @request.json \
"https://dataproc.googleapis.com/v1/projects/PROJECT/regions/REGION/clusters"

PowerShell (Windows)

Save the request body in a file named request.json, and execute the following command:

$cred = gcloud auth print-access-token
$headers = @{ "Authorization" = "Bearer $cred" }

Invoke-WebRequest `
-Method POST `
-Headers $headers `
-ContentType: "application/json; charset=utf-8" `
-InFile request.json `
-Uri "https://dataproc.googleapis.com/v1/projects/PROJECT/regions/REGION/clusters" | Select-Object -Expand Content

You should receive a JSON response similar to the following:

{
"name": "projects/PROJECT/regions/REGION/operations/b5706e31......",
 "metadata": {
 "@type": "type.googleapis.com/google.cloud.dataproc.v1.ClusterOperationMetadata",
 "clusterName": "CLUSTER_NAME",
 "clusterUuid": "5fe882b2-...",
 "status": {
 "state": "PENDING",
 "innerState": "PENDING",
 "stateStartTime": "2019-11-21T00:37:56.220Z"
 },
 "operationType": "CREATE",
 "description": "Create cluster with 2 workers",
 "warnings": [
 "For PD-Standard without local SSDs, we strongly recommend provisioning 1TB ...""
 ]
 }
}

Go

  1. Install the client library.
  2. Set up application default credentials.
  3. Run the code.

    Note: While specifying machine types is optional, it is recommended to explicitly set the master and worker machine types in your cluster configuration (for example, to n4-standard-4) to ensure consistent cost and performance. If omitted, default machine types are selected dynamically based on resource availability.

    import(
    "context"
    "fmt"
    "io"
    dataproc"cloud.google.com/go/dataproc/apiv1"
    "cloud.google.com/go/dataproc/apiv1/dataprocpb"
    "google.golang.org/api/option"
    )
    funccreateCluster(wio.Writer,projectID,region,clusterNamestring)error{
    // projectID := "your-project-id"
    // region := "us-central1"
    // clusterName := "your-cluster"
    ctx:=context.Background()
    // Create the cluster client.
    endpoint:=region+"-dataproc.googleapis.com:443"
    clusterClient,err:=dataproc.NewClusterControllerClient(ctx,option.WithEndpoint(endpoint))
    iferr!=nil{
    returnfmt.Errorf("dataproc.NewClusterControllerClient: %w",err)
    }
    deferclusterClient.Close()
    // Create the cluster config.
    req:=&dataprocpb.CreateClusterRequest{
    ProjectId:projectID,
    Region:region,
    Cluster:&dataprocpb.Cluster{
    ProjectId:projectID,
    ClusterName:clusterName,
    Config:&dataprocpb.ClusterConfig{
    MasterConfig:&dataprocpb.InstanceGroupConfig{
    NumInstances:1,
    MachineTypeUri:"n1-standard-2",
    },
    WorkerConfig:&dataprocpb.InstanceGroupConfig{
    NumInstances:2,
    MachineTypeUri:"n1-standard-2",
    },
    },
    },
    }
    // Create the cluster.
    op,err:=clusterClient.CreateCluster(ctx,req)
    iferr!=nil{
    returnfmt.Errorf("CreateCluster: %w",err)
    }
    resp,err:=op.Wait(ctx)
    iferr!=nil{
    returnfmt.Errorf("CreateCluster.Wait: %w",err)
    }
    // Output a success message.
    fmt.Fprintf(w,"Cluster created successfully: %s",resp.ClusterName)
    returnnil
    }
    

Java

  1. Install the client library.
  2. Set up application default credentials.
  3. Run the code.

    Note: While specifying machine types is optional, it is recommended to explicitly set the master and worker machine types in your cluster configuration (for example, to n4-standard-4) to ensure consistent cost and performance. If omitted, default machine types are selected dynamically based on resource availability.

    importcom.google.api.gax.longrunning.OperationFuture ;
    importcom.google.cloud.dataproc.v1.Cluster ;
    importcom.google.cloud.dataproc.v1.ClusterConfig ;
    importcom.google.cloud.dataproc.v1.ClusterControllerClient ;
    importcom.google.cloud.dataproc.v1.ClusterControllerSettings ;
    importcom.google.cloud.dataproc.v1.ClusterOperationMetadata ;
    importcom.google.cloud.dataproc.v1.InstanceGroupConfig ;
    importjava.io.IOException;
    importjava.util.concurrent.ExecutionException;
    publicclass CreateCluster{
    publicstaticvoidcreateCluster()throwsIOException,InterruptedException{
    // TODO(developer): Replace these variables before running the sample.
    StringprojectId="your-project-id";
    Stringregion="your-project-region";
    StringclusterName="your-cluster-name";
    createCluster(projectId,region,clusterName);
    }
    publicstaticvoidcreateCluster(StringprojectId,Stringregion,StringclusterName)
    throwsIOException,InterruptedException{
    StringmyEndpoint=String.format("%s-dataproc.googleapis.com:443",region);
    // Configure the settings for the cluster controller client.
    ClusterControllerSettings clusterControllerSettings=
    ClusterControllerSettings .newBuilder().setEndpoint(myEndpoint).build();
    // Create a cluster controller client with the configured settings. The client only needs to be
    // created once and can be reused for multiple requests. Using a try-with-resources
    // closes the client, but this can also be done manually with the .close() method.
    try(ClusterControllerClient clusterControllerClient=
    ClusterControllerClient .create(clusterControllerSettings)){
    // Configure the settings for our cluster.
    InstanceGroupConfig masterConfig=
    InstanceGroupConfig .newBuilder()
    .setMachineTypeUri ("n1-standard-2")
    .setNumInstances (1)
    .build();
    InstanceGroupConfig workerConfig=
    InstanceGroupConfig .newBuilder()
    .setMachineTypeUri ("n1-standard-2")
    .setNumInstances (2)
    .build();
    ClusterConfig clusterConfig=
    ClusterConfig .newBuilder()
    .setMasterConfig (masterConfig)
    .setWorkerConfig(workerConfig)
    .build();
    // Create the cluster object with the desired cluster config.
    Cluster cluster=
    Cluster .newBuilder().setClusterName(clusterName).setConfig(clusterConfig).build();
    // Create the Cloud Dataproc cluster.
    OperationFuture<Cluster,ClusterOperationMetadata>createClusterAsyncRequest=
    clusterControllerClient.createClusterAsync (projectId,region,cluster);
    Cluster response=createClusterAsyncRequest.get ();
    // Print out a success message.
    System.out.printf("Cluster created successfully: %s",response.getClusterName ());
    }catch(ExecutionExceptione){
    System.err.println(String.format("Error executing createCluster: %s ",e.getMessage()));
    }
    }
    }

Node.js

  1. Install the client library.
  2. Set up application default credentials.
  3. Run the code.

    Note: While specifying machine types is optional, it is recommended to explicitly set the master and worker machine types in your cluster configuration (for example, to n4-standard-4) to ensure consistent cost and performance. If omitted, default machine types are selected dynamically based on resource availability.

    constdataproc=require('@google-cloud/dataproc');
    // TODO(developer): Uncomment and set the following variables
    // projectId = 'YOUR_PROJECT_ID'
    // region = 'YOUR_CLUSTER_REGION'
    // clusterName = 'YOUR_CLUSTER_NAME'
    // Create a client with the endpoint set to the desired cluster region
    constclient=newdataproc.v1.ClusterControllerClient ({
    apiEndpoint:`${region}-dataproc.googleapis.com`,
    projectId:projectId,
    });
    asyncfunctioncreateCluster(){
    // Create the cluster config
    constrequest={
    projectId:projectId,
    region:region,
    cluster:{
    clusterName:clusterName,
    config:{
    masterConfig:{
    numInstances:1,
    machineTypeUri:'n1-standard-2',
    },
    workerConfig:{
    numInstances:2,
    machineTypeUri:'n1-standard-2',
    },
    },
    },
    };
    // Create the cluster
    const[operation]=awaitclient.createCluster(request);
    const[response]=awaitoperation.promise();
    // Output a success message
    console.log(`Cluster created successfully: ${response.clusterName}`);

Python

  1. Install the client library.
  2. Set up application default credentials.
  3. Run the code.

    Note: While specifying machine types is optional, it is recommended to explicitly set the master and worker machine types in your cluster configuration (for example, to n4-standard-4) to ensure consistent cost and performance. If omitted, default machine types are selected dynamically based on resource availability.

    fromgoogle.cloudimport dataproc_v1 as dataproc
    defcreate_cluster(project_id, region, cluster_name):
    """This sample walks a user through creating a Cloud Dataproc cluster
     using the Python client library.
     Args:
     project_id (string): Project to use for creating resources.
     region (string): Region where the resources should live.
     cluster_name (string): Name to use for creating a cluster.
     """
     # Create a client with the endpoint set to the desired cluster region.
     cluster_client = dataproc.ClusterControllerClient(
     client_options={"api_endpoint": f"{region}-dataproc.googleapis.com:443"}
     )
     # Create the cluster config.
     cluster = {
     "project_id": project_id,
     "cluster_name": cluster_name,
     "config": {
     "master_config": {"num_instances": 1, "machine_type_uri": "n1-standard-2"},
     "worker_config": {"num_instances": 2, "machine_type_uri": "n1-standard-2"},
     },
     }
     # Create the cluster.
     operation = cluster_client.create_cluster(
     request={"project_id": project_id, "region": region, "cluster": cluster}
     )
     result = operation.result()
     # Output a success message.
     print(f"Cluster created successfully: {result.cluster_name}")

Except as otherwise noted, the content of this page is licensed under the Creative Commons Attribution 4.0 License, and code samples are licensed under the Apache 2.0 License. For details, see the Google Developers Site Policies. Java is a registered trademark of Oracle and/or its affiliates.

Last updated 2026年08月26日 UTC.