# Upgrade ES version managed by ECK on terraform - all pods terminated

**URL:** <https://discuss.elastic.co/t/upgrade-es-version-managed-by-eck-on-terraform-all-pods-terminated/320897>\
**Category:** Elastic Cloud on Kubernetes (ECK)\
**Created:** [December 9, 2022, 4:06pm UTC](https://discuss.elastic.co/t/upgrade-es-version-managed-by-eck-on-terraform-all-pods-terminated/320897 "2022-12-09T16:06:16Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Roberto\_D\_Arco](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/roberto_d_arco/32/98729_2.png) [@Roberto\_D\_Arco](https://discuss.elastic.co/u/Roberto_D_Arco)\
**Post date:** [December 9, 2022, 4:06pm UTC](https://discuss.elastic.co/t/upgrade-es-version-managed-by-eck-on-terraform-all-pods-terminated/320897/1 "2022-12-09T16:06:16Z")

</div>

Hello everyone,

Our current Production Elasticsearch cluster for logs collection is manually managed and runs on AWS.  
I'm creating the same cluster using ECK deployed with Helm under Terraform.  
I was able to get all the features replicated (S3 repo for snapshots, ingest pipelines, index templates, etc) and deployed, but when I tried to update the cluster (changing the ES version from 8.3.2 to 8.5.2) I get a NEW elasticsearch cluster with version 8.5.2 in what doesn't appear as a rolling upgrade.

I can tell that it is a new cluster because the default 'elastic' superuser has a new password.

Also, when I check the kubernetes pods immediately after the terraform apply with updated ES version, the kibana pod doesn't even exists (probably normal) and all the ES nodes pods are simultaneously terminating.  
I'm not ingesting data on this new cluster at the moment, but I'm sure that if it was the case, I would get an ingest interruption, and red health status (or maybe not, since I have what it looks like a completely new cluster...).

Most probably the problem is in my elasticsearch manifest, But I couldn't pinpoint the problem.

Here my ES manifest:

```auto
apiVersion: elasticsearch.k8s.elastic.co/v1
kind: Elasticsearch
metadata:
  # copy the specified node labels as pod annotations and use it as an environment variable in the Pods; spreads a NodeSet across the availability zones of a Kubernetes cluster. Used for AZ awareness
  annotations:
    eck.k8s.elastic.co/downward-node-labels: "topology.kubernetes.io/zone"
  name: ${cluster_name}
  namespace: ${namespace}
spec:
  version: ${version}
  volumeClaimDeletePolicy: DeleteOnScaledown
  #updateStrategy:
  # changeBudget:
  # maxSurge: 1
  # maxUnavailable: 1
  # for monitoring see: https://www.elastic.co/guide/en/cloud-on-k8s/current/k8s-stack-monitoring.html
  monitoring:
    metrics:
      elasticsearchRefs:
        - name: ${cluster_name}
    logs:
      elasticsearchRefs:
        - name: ${cluster_name}
  nodeSets:
    - name: logging-nodes
      count: ${nodes}
      config:
        # logger.org.elasticsearch: DEBUG
        node.roles: ["master","data", "ingest", "ml", "transform", "remote_cluster_client"]
        # this allows ES to run on nodes even if their vm.max_map_count has not been increased, at a performance cost
        node.store.allow_mmap: false
        cluster:
        # name: "logging.elasticsearch" See: https://www.elastic.co/guide/en/cloud-on-k8s/current/k8s-reserved-settings.html
          routing:
            rebalance.enable: "all"
            allocation:
              enable: "all"
              allow_rebalance: "always"
              node_concurrent_recoveries: ${node_concurrent_recoveries}
        # use the zone attribute from the node labels. Used for AZ awareness; double $ is used to escape during templating
        node.attr.zone: $${ZONE}
        cluster.routing.allocation.awareness.attributes: k8s_node_name,zone
        gateway.expected_data_nodes: ${nodes}
        indices.recovery.max_bytes_per_sec: ${index_recovery_speed}
        # network.host: ["_ec2:publicDns_", "localhost"] See: https://www.elastic.co/guide/en/cloud-on-k8s/current/k8s-reserved-settings.html
        # xpack.security.enabled: true See: https://www.elastic.co/guide/en/cloud-on-k8s/current/k8s-reserved-settings.html
      podTemplate:
        metadata:
          namespace: ${namespace}
          labels:
            # additional labels for pods
            stack_name: ${stack_name}
            stack_repository: ${stack_repository}
        spec:
          volumes:
            - name: aws-iam-token-es
              projected:
                defaultMode: 420
                sources:
                - serviceAccountToken:
                    audience: sts.amazonaws.com
                    expirationSeconds: 86400
                    path: aws-web-identity-token-file
          serviceAccountName: ${service_account}
          containers:
            - name: elasticsearch
              # specify resource limits and requests
              resources:
                limits:
                  memory: 4Gi
                  cpu: "1"
              volumeMounts:
              - mountPath: /usr/share/elasticsearch/config/repository-s3
                name: aws-iam-token-es
                readOnly: true
              env:
                # Makes the topology.kubernetes.io/zone annotation available as an environment variable and
                # use it as a cluster routing allocation attribute.
                - name: AWS_ROLE_SESSION_NAME
                  value: elasticsearch-sts
                - name: ZONE
                  valueFrom:
                    fieldRef:
                      fieldPath: metadata.annotations['topology.kubernetes.io/zone']
          # used for availability zone awareness
          topologySpreadConstraints:
            - maxSkew: 1
              topologyKey: topology.kubernetes.io/zone
              whenUnsatisfiable: DoNotSchedule
              labelSelector:
                matchLabels:
                  elasticsearch.k8s.elastic.co/cluster-name: ${cluster_name}
                  elasticsearch.k8s.elastic.co/statefulset-name: ${cluster_name}-es-default
      # request 15Gi of persistent data storage for pods in this topology element
      volumeClaimTemplates:
        - metadata:
            name: elasticsearch-data # Do not change this name unless you set up a volume mount for the data path.
          spec:
            accessModes:
              - ReadWriteOnce
            resources:
              requests:
                storage: 15Gi
            storageClassName: gp2

```

I can also post the kibana manifest but I don't think it is relevant.  
To perform the upgrade, I just change the ${version} variable.

---

<div class="post-metadata">

**Author:** ![Roberto\_D\_Arco](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/roberto_d_arco/32/98729_2.png) [@Roberto\_D\_Arco](https://discuss.elastic.co/u/Roberto_D_Arco)\
**Post date:** [December 13, 2022, 10:07am UTC](https://discuss.elastic.co/t/upgrade-es-version-managed-by-eck-on-terraform-all-pods-terminated/320897/2 "2022-12-13T10:07:52Z")

</div>

I think I'm having the same problem as in [Deploy Elasticsearch Custom Resource with Terraform](https://discuss.elastic.co/t/deploy-elasticsearch-custom-resource-with-terraform/314699)

---

<div class="post-metadata">

**Author:** ![BenB196](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/benb196/32/83401_2.png) [@BenB196](https://discuss.elastic.co/u/BenB196)\
**Post date:** [December 13, 2022, 12:27pm UTC](https://discuss.elastic.co/t/upgrade-es-version-managed-by-eck-on-terraform-all-pods-terminated/320897/3 "2022-12-13T12:27:21Z")

</div>

I was never able to find a solution to this.

I did come across this post: [Problem with preventing deletion of elastic volumeClaimTemplate created by terraform - Kubernetes - HashiCorp Discuss](https://discuss.hashicorp.com/t/problem-with-preventing-deletion-of-elastic-volumeclaimtemplate-created-by-terraform/43983?u=benb196), which I believe is related. But I was never able to find a solution for this.

---

<div class="post-metadata">

**Author:** ![Roberto\_D\_Arco](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/roberto_d_arco/32/98729_2.png) [@Roberto\_D\_Arco](https://discuss.elastic.co/u/Roberto_D_Arco)\
**Post date:** [December 13, 2022, 1:54pm UTC](https://discuss.elastic.co/t/upgrade-es-version-managed-by-eck-on-terraform-all-pods-terminated/320897/4 "2022-12-13T13:54:31Z")

</div>

Thanks for jumping back here.  
Before seeing your post I was using gavinbunney/kubectl , but your post 'inspired me' to give another try to the 'official' kubernetes\_manifest.

Now, when I apply the elasticsearch version change I get the error:

```auto
│ The API returned the following conflict: "Apply failed with 1 conflict: conflict with \"elastic-operator\" using elasticsearch.k8s.elastic.co/v1: .spec.nodeSets"
│ 
│ You can override this conflict by setting "force_conflicts" to true in the "field_manager" block.

```

I tried to add the

```auto
  field_manager {
    force_conflicts = true
  }

```

but then I got:

```auto
│ Error: Provider produced inconsistent result after apply
│ 
│ When applying changes to kubernetes_manifest.kibana_deploy, provider "provider[\"registry.terraform.io/hashicorp/kubernetes\"]" produced an unexpected new value: .object: wrong final value type: incorrect object attributes.
│ 
│ This is a bug in the provider, which should be reported in the provider's own issue tracker.

```

so I think this is a no go.

But taking out the force\_conflict, even if I was getting an error, the 'plan part' of terraform was saying that it was going to **update** and not **replace** the resources, so it is a step closer.

In the kubernetes manifest I used:

```auto
computed_fields = ["metadata.labels", "metadata.annotations","spec.finalizers","status"]

```

that I found in some other post.  
Maybe if we found the complete list of required computed\_fields it could work.

Researching.....

---

<div class="post-metadata">

**Author:** ![Roberto\_D\_Arco](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/roberto_d_arco/32/98729_2.png) [@Roberto\_D\_Arco](https://discuss.elastic.co/u/Roberto_D_Arco)\
**Post date:** [December 13, 2022, 2:28pm UTC](https://discuss.elastic.co/t/upgrade-es-version-managed-by-eck-on-terraform-all-pods-terminated/320897/5 "2022-12-13T14:28:11Z")

</div>

> [@Roberto\_D\_Arco](#):
>
> `computed_fields = ["metadata.labels", "metadata.annotations","spec.finalizers","status"]`

this is where I found it:

> <https://github.com/hashicorp/terraform-provider-kubernetes/issues/1740>
>
> \## Terraform version, Kubernetes provider version and Kubernetes version
> \`\`\`
> T…erraform version: 1.2.1
> Go runtime version: go1.18.1
> hashicorp/kubernetes/2.11.0
> kubectl Version:"v1.24.1"
> \`\`\`
> \## Terraform configuration
> A lot is missing but you should get the idea.
> \`\`\`hcl
> \### GKE Module
> resource "google\_container\_cluster" "primary" {
> provider = google-beta
> 
> name = var.cluster\_name
> location = var.location
> project = var.project
> remove\_default\_node\_pool = true
> initial\_node\_count = 1
> ....
> }
> 
> output "cluster\_ca\_certificate" {
> value = google\_container\_cluster.primary.master\_auth.0.cluster\_ca\_certificate
> }
> 
> output "endpoint" {
> value = google\_container\_cluster.primary.endpoint
> }
> 
> \### Flux module
> 
> data "google\_client\_config" "default" {}
> 
> provider "kubernetes" {
> host = "https://${module.gke.endpoint}"
> token = data.google\_client\_config.default.access\_token
> cluster\_ca\_certificate = base64decode(
> module.gke.cluster\_ca\_certificate,
> )
> }
> 
> locals {
> raw\_emissary\_manifests = split("---", file("${path.root}/flux-config/conditionals/${var.env\_type}/ambassador.yaml"))
> hcl\_emissary\_manifests = \[for manifest in local.raw\_emissary\_manifests : yamldecode(manifest)\]
> emissary\_cfg = base64encode(\<\<YAML
> service:
> annotations:
> cloud.google.com/load-balancer-type: "${(var.external) ? "External" : "Internal"}"
> external-dns.alpha.kubernetes.io/hostname: ${(var.external) ? "amb.${var.cluster\_name}.bestsellerit.com" : "ambassador.${var.cluster\_name}.k8s.bestcorp.net"}
> tags.datadoghq.com/env: "${var.env\_type}"
> podLabels:
> tags.datadoghq.com/env: "${var.env\_type}"
> YAML
> )
> }
> resource "kubernetes\_cluster\_role\_binding" "admin" {
> metadata {
> name = "cluster-admin-binding"
> }
> role\_ref {
> api\_group = "rbac.authorization.k8s.io"
> kind = "ClusterRole"
> name = "cluster-admin"
> }
> subject {
> kind = "User"
> name = var.deployment\_account\_email
> api\_group = "rbac.authorization.k8s.io"
> }
> }
> \[...\]
> resource "kubernetes\_manifest" "emissary\_ns" {
> depends\_on = \[kubernetes\_cluster\_role\_binding.admin, helm\_release.gatekeeper\]
> computed\_fields = \["metadata.labels", "metadata.annotations","spec.finalizers","status"\]
> 
> manifest = yamldecode(file("${path.module}/conditionals/emissary/namespace.yaml"))
> }
> 
> resource "kubernetes\_manifest" "emissary\_cert" {
> count = var.ambassador ? 1 : 0
> depends\_on = \[kubectl\_manifest.sync\_flux, kubernetes\_manifest.emissary\_ns\]
> computed\_fields = \["metadata.labels", "metadata.annotations","spec.finalizers","status"\]
> manifest = yamldecode(templatefile("${path.module}/conditionals/emissary/certificate.yaml", {
> dns = (var.external) ? "amb.${var.cluster\_name}.bestsellerit.com" : "ambassador.${var.cluster\_name}.k8s.bestcorp.net"
> }))
> }
> 
> resource "kubernetes\_manifest" "emissary\_cfg" {
> count = var.ambassador ? 1 : 0
> depends\_on = \[kubectl\_manifest.sync\_flux, kubernetes\_manifest.emissary\_ns\]
> computed\_fields = \["metadata.labels", "metadata.annotations","spec.finalizers","status"\]
> manifest = yamldecode(templatefile("${path.module}/conditionals/emissary/helm-config.yaml", { dataBase64 : local.emissary\_cfg }))
> }
> \`\`\`
> 
> \## Question
> \`\`\`
> Hi, i have some kubernetes resources that i was managing using the old kubectl provider. I have removed them from the state of the old provider and imported them into the new one. 
> 
> I have two problems:
> 
> 1. Terraform wants to destroy my imported resources and there is no prompt stating what the reason is, this is not so important right now and maybe it's an improvement for a future version.
> 2. I am getting a cycle error on destroy. I am not sure why the gke part depends on the kubectl\_manifest
> 
> Error: Cycle: module.flux.kubernetes\_manifest.emissary\_docs\[0\] (destroy), module.flux.kubernetes\_manifest.emissary\_docs\[3\] (destroy), module.flux.kubernetes\_manifest.emissary\_docs\[1\] (destroy), module.gke.output.cluster\_ca\_certificate (expand), module.gke.output.endpoint (expand), provider\["registry.terraform.io/hashicorp/kubernetes"\], module.flux.kubernetes\_manifest.emissary\_docs\[2\] (destroy), module.gke.google\_container\_cluster.primary
> 
> \`\`\`
> Debug output:
> 
> \`\`\`
> 2022-06-10T06:13:31.649Z \[ERROR\] Graph validation failed. Graph:
> 
> \[...\]
> module.gke.output.cluster\_ca\_certificate (expand)
> module.gke (expand)
> module.gke.google\_container\_cluster.primary
> module.gke.google\_container\_cluster.primary (expand)
> \[...\]
> module.gke.output.endpoint (expand)
> module.gke (expand)
> module.gke.google\_container\_cluster.primary
> module.gke.google\_container\_cluster.primary (expand)
> \[...\]
> provider\["registry.terraform.io/hashicorp/kubernetes"\]
> module.gke.output.cluster\_ca\_certificate (expand)
> module.gke.output.endpoint (expand)
> 
> \[...\]
> module.flux.kubernetes\_manifest.emissary\_docs\[0\] (destroy)
> module.vault.null\_resource.create\_cluster\_policies (destroy)
> provider\["registry.terraform.io/hashicorp/kubernetes"\]
> 
> module.flux.kubernetes\_manifest.emissary\_docs\[1\] (destroy)
> module.vault.null\_resource.create\_cluster\_policies (destroy)
> provider\["registry.terraform.io/hashicorp/kubernetes"\]
> 
> \[...\]
> module.flux.kubernetes\_manifest.emissary\_docs\[2\] (destroy)
> module.vault.null\_resource.create\_cluster\_policies (destroy)
> provider\["registry.terraform.io/hashicorp/kubernetes"\]
> 
> \[...\]
> module.flux.kubernetes\_manifest.emissary\_docs\[3\] (destroy)
> module.vault.null\_resource.create\_cluster\_policies (destroy)
> provider\["registry.terraform.io/hashicorp/kubernetes"\]
> 
> \[...\]
> module.gke.google\_container\_cluster.primary
> module.flux.kubernetes\_manifest.emissary\_docs\[0\] (destroy)
> module.flux.kubernetes\_manifest.emissary\_docs\[1\] (destroy)
> module.flux.kubernetes\_manifest.emissary\_docs\[2\] (destroy)
> module.flux.kubernetes\_manifest.emissary\_docs\[3\] (destroy)
> module.flux.time\_sleep.wait\_60\_seconds (destroy)
> module.gke (expand)
> module.gke.data.google\_container\_engine\_versions.k8s\_version (expand)
> module.gke.data.google\_project.project (expand)
> module.gke.google\_bigquery\_dataset.dataset (expand)
> module.gke.google\_container\_cluster.primary (expand)
> module.gke.local.network (expand)
> module.gke.local.subnetwork (expand)
> module.gke.var.cluster\_name (expand)
> module.gke.var.cluster\_secondary\_range\_name (expand)
> module.gke.var.enable\_gke\_ingress (expand)
> module.gke.var.external (expand)
> module.gke.var.external\_auto\_subnet (expand)
> module.gke.var.location (expand)
> module.gke.var.master\_ipv4\_cidr\_block (expand)
> module.gke.var.namespaces (expand)
> module.gke.var.project (expand)
> module.gke.var.services\_secondary\_range\_name (expand)
> module.vault.null\_resource.create\_cluster\_policies (destroy)
> provider\["registry.terraform.io/hashicorp/google-beta"\]
> \`\`\`

---

<div class="post-metadata">

**Author:** ![Roberto\_D\_Arco](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/roberto_d_arco/32/98729_2.png) [@Roberto\_D\_Arco](https://discuss.elastic.co/u/Roberto_D_Arco)\
**Post date:** [December 13, 2022, 5:07pm UTC](https://discuss.elastic.co/t/upgrade-es-version-managed-by-eck-on-terraform-all-pods-terminated/320897/6 "2022-12-13T17:07:59Z")

</div>

New, related question asked:

> [@Bug in eck operator? Cluster upgrade fails (under terraform)](https://discuss.elastic.co/t/bug-in-eck-operator-cluster-upgrade-fails-under-terraform/321149):
>
> Hello everyone, Our current Production Elasticsearch cluster for logs collection is manually managed and runs on AWS. I'm creating the same cluster using ECK deployed with Helm under Terraform. I was able to get all the features replicated (S3 repo for snapshots, ingest pipelines, index templates, etc) and deployed, so, first deployment is perfectly working. But when I tried to update the cluster (changing the ES version from 8.3.2 to 8.5.2) I get this error: │ Error: Provider produced inco…

---

<div class="post-metadata">

**Author:** ![Roberto\_D\_Arco](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/roberto_d_arco/32/98729_2.png) [@Roberto\_D\_Arco](https://discuss.elastic.co/u/Roberto_D_Arco)\
**Post date:** [December 15, 2022, 11:19am UTC](https://discuss.elastic.co/t/upgrade-es-version-managed-by-eck-on-terraform-all-pods-terminated/320897/7 "2022-12-15T11:19:36Z")

</div>

In the end the problem was that the eck operator does a lot of changes in the spec session, so if you add the whole "spec" to the computed\_fields those changes are ignored and the upgrade proceed as intended:

```auto
resource "kubernetes_manifest" "elasticsearch_deploy" {
  field_manager {
    force_conflicts = true
  }
  computed_fields = ["metadata.labels", "metadata.annotations", "spec", "status"]
  manifest = yamldecode(templatefile("config/elasticsearch.yaml", {
    version = var.elastic_stack_version
    nodes = var.logging_elasticsearch_nodes_count
    cluster_name = local.cluster_name
  }))
}

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 12, 2023, 11:20am UTC](https://discuss.elastic.co/t/upgrade-es-version-managed-by-eck-on-terraform-all-pods-terminated/320897/8 "2023-01-12T11:20:36Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
