# Elasticsearch discovery issues kubernetes

**URL:** <https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465>\
**Category:** Elasticsearch\
**Created:** [August 5, 2022, 3:05am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465 "2022-08-05T03:05:28Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![willsc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/willsc/32/109430_2.png) [@willsc](https://discuss.elastic.co/u/willsc)\
**Post date:** [August 5, 2022, 3:05am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/1 "2022-08-05T03:05:28Z")

</div>

Deploying elasticsearch 7.17.4 into kubernetes {rancher}, and have noticed that elastic doesn't form a cluster, it partitions into 3 separate master nodes.

How do i increase the level of debugging on the masters, as all the docs seem to be wrong.

Secondly how can I verify that my baremetal cluster isn't the issue given, if deploy the same helm chart in eks, everything works like a charm.

Is there anything special I'd have to a baremeta cluster?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [August 5, 2022, 7:45am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/2 "2022-08-05T07:45:26Z")

</div>

By default Elasticsearch will emit enough logs to diagnose the problem, no need to look for debug logs.

If you need help understanding the logs, share them here. You'll need logs covering at least 5 mins from all nodes.

---

<div class="post-metadata">

**Author:** ![willsc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/willsc/32/109430_2.png) [@willsc](https://discuss.elastic.co/u/willsc)\
**Post date:** [August 8, 2022, 9:54pm UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/3 "2022-08-08T21:54:57Z")

</div>

So there's an issue, the logs only show one node joining the a cluster essentially itself. All 'nodes behave like this, it seems that they cannot resolve the respective node names and assume single node discovery.

Yet when I use getbyhostname to verify if resolution works, it appears to be fine. Explicitly setting the discovery mode makes no difference.

Is there a difference between the way the rpm and single non rpm binary works? My image uses the rpm instead of the compressed tar binary. I'm deploying 7.17.4.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 8, 2022, 9:58pm UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/4 "2022-08-08T21:58:29Z")

</div>

It'd be useful if you shared the logs and your config, otherwise we're really just making educated guesses.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [August 9, 2022, 6:41am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/5 "2022-08-09T06:41:25Z")

</div>

As Mark says we can only guess at the problem from a vague description of what you're seeing that you think to be relevant.

One possible guess is [this situation](https://www.elastic.co/guide/en/elasticsearch/reference/current/modules-discovery-bootstrap-cluster.html#modules-discovery-bootstrap-cluster-joining). If that describes what you're seeing then you can use the remedy in the docs:

> If you intended to form a new multi-node cluster but instead bootstrapped a collection of single-node clusters...

---

<div class="post-metadata">

**Author:** ![willsc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/willsc/32/109430_2.png) [@willsc](https://discuss.elastic.co/u/willsc)\
**Post date:** [August 9, 2022, 8:02am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/6 "2022-08-09T08:02:58Z")

</div>

Unfortunately due to company rules I can't send you the data you require, however what I can tell you is that the deployment is a statefulset in k8s and also the kubernetes deployment is bare metal running Rancher.

So my question is, given this isn't on a cloud provider how does the discovery seeding work, as there is no api as far as I'm aware of to leverage? I'm using the Elasticsearch helm chart for deployment for 7.17.x

Below is a section of the rendered chart related to discovery.

```auto
 env:
          - name: node.name
            valueFrom:
              fieldRef:
                fieldPath: metadata.name
          - name: cluster.initial_master_nodes
            value: "elasticsearch-master-0,elasticsearch-master-1,elasticsearch-master-2,"
          - name: discovery.seed_hosts
            value: "elasticsearch-master-headless"
          - name: cluster.name
            value: "elasticsearch"
          - name: network.host
            value: "0.0.0.0"
          - name: cluster.deprecation_indexing.enabled
            value: "false"

```

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [August 9, 2022, 8:16am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/7 "2022-08-09T08:16:19Z")

</div>

> [@willsc](#):
>
> how does the discovery seeding work

It looks like you set `discovery.seed_hosts: elasticsearch-master-headless` which means Elasticsearch will do a lookup for this name and use all the addresses in the response for discovery.

---

<div class="post-metadata">

**Author:** ![willsc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/willsc/32/109430_2.png) [@willsc](https://discuss.elastic.co/u/willsc)\
**Post date:** [August 9, 2022, 8:32am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/8 "2022-08-09T08:32:31Z")

</div>

The behaviour in a aws EKS deployment is somewhat different, in that the cluster is formed properly, but presumably in the case of AWS it takes advantage of api, in the case of a bare metal rancher cluster, I'm guessing this going to behave in a slightly different way?

Is the discovery config still valid ?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [August 9, 2022, 8:39am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/9 "2022-08-09T08:39:30Z")

</div>

> [@willsc](#):
>
> presumably in the case of AWS it takes advantage of api,

No, not unless you explicitly tell it to (e.g. install the `discovery-ec2` plugin and set `discovery.seed_providers: ec2`, see [these docs](https://www.elastic.co/guide/en/elasticsearch/plugins/current/discovery-ec2-usage.html)).

---

<div class="post-metadata">

**Author:** ![willsc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/willsc/32/109430_2.png) [@willsc](https://discuss.elastic.co/u/willsc)\
**Post date:** [August 9, 2022, 8:51am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/10 "2022-08-09T08:51:31Z")

</div>

Ah ok so in the case of bare metal clusters how does this work, I presume we rely on pod level resolution ?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [August 9, 2022, 9:05am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/11 "2022-08-09T09:05:13Z")

</div>

Elasticsearch doesn't do anything different, it does a lookup for the name(s) you configure and uses all the addresses in the response. The specific library function it calls for this lookup is [`getaddrinfo()`](https://man7.org/linux/man-pages/man3/getaddrinfo.3.html) which can be configured to behave differently in different environments (usually it uses `/etc/hosts` and DNS but many other options are available). You'll need to ask a local expert for the details of how name lookup works in your environment, sorry, that's not something I can help with here as it's not really anything to do with Elasticsearch.

---

<div class="post-metadata">

**Author:** ![willsc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/willsc/32/109430_2.png) [@willsc](https://discuss.elastic.co/u/willsc)\
**Post date:** [August 9, 2022, 9:07am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/12 "2022-08-09T09:07:32Z")

</div>

I guess the follow up question, given my cluster is baremetal should I configuring my chart differently in terms of discovery ?

Secondly, what else can I do from an elastic perspective to further debug this, the logs only show that a single node cluster has been formed it makes no mention of any other nodes. So I have to conclude that I need to treat a baremetal deployment in a different way ?

---

<div class="post-metadata">

**Author:** ![willsc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/willsc/32/109430_2.png) [@willsc](https://discuss.elastic.co/u/willsc)\
**Post date:** [August 9, 2022, 9:14am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/13 "2022-08-09T09:14:55Z")

</div>

Well given its k8s cluster that would be coreDNS . is there any other debug available to me which I can turn on to get a better idea about whats going on ? The logs are ok ish but are definitely light on the discovery process .

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [August 9, 2022, 9:25am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/14 "2022-08-09T09:25:58Z")

</div>

> [@willsc](#):
>
> the logs only show that a single node cluster has been formed it makes no mention of any other nodes

I think the docs I linked [earlier](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/5) say what to do here.

---

<div class="post-metadata">

**Author:** ![willsc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/willsc/32/109430_2.png) [@willsc](https://discuss.elastic.co/u/willsc)\
**Post date:** [August 9, 2022, 9:55am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/15 "2022-08-09T09:55:49Z")

</div>

Yes that clearly works in the context of a cluster built in a non kubernetes environment. I'm specifically talking about a kubernetes environment with ephemeral nodes ? I presume from the short replies, that elastic on kubernetes let alone bare metal k8s isn't something that a lot of people know a great deal about ?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [August 9, 2022, 10:18am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/16 "2022-08-09T10:18:51Z")

</div>

I don't really understand what you're asking here. The docs I linked apply to all operating environments. If you're having trouble getting Kubernetes and ES to play nicely by hand then maybe you'd be better off using something like [Elastic Cloud on Kubernetes | Deploy and Orchestrate Elasticsearch on Kubernetes](https://www.elastic.co/elastic-cloud-kubernetes)?

---

<div class="post-metadata">

**Author:** ![willsc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/willsc/32/109430_2.png) [@willsc](https://discuss.elastic.co/u/willsc)\
**Post date:** [August 9, 2022, 10:33am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/17 "2022-08-09T10:33:30Z")

</div>

and that's the problem, its a very simple question, given I have ephemeral containers in k8s how can I debug a discovery problem with the instructions from your docs which are clearly aimed at a non k8s deployment.

Second question how can I increase the debug levels in the logs, such that I can see what happens during discovery.

Unfortunately the cloud option is not a viable option for me.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [August 9, 2022, 10:43am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/18 "2022-08-09T10:43:58Z")

</div>

> [@willsc](#):
>
> how can I debug a discovery problem

I don't think you have a discovery problem, because ...

> [@willsc](#):
>
> how can I increase the debug levels in the logs, such that I can see what happens during discovery.

... there is no need to adjust any logging levels to diagnose discovery problems in 7.17. If you were having a discovery problem, the logs would already be full of debugging information about it.

Instead, I think you're having a cluster bootstrapping problem, and the docs I linked above tell you how to both diagnose and fix it. These docs apply to all environments, there's nothing about them which aims at any particular setup.

> [@willsc](#):
>
> Unfortunately the cloud option is not a viable option for me.

ECK is something you can run on your own local K8s environment, effectively a private cloud.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 6, 2022, 10:44am UTC](https://discuss.elastic.co/t/elasticsearch-discovery-issues-kubernetes/311465/19 "2022-09-06T10:44:17Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
