# Cluster setup questions and rollover logstash ingestion to elasticsearch

**URL:** <https://discuss.elastic.co/t/cluster-setup-questions-and-rollover-logstash-ingestion-to-elasticsearch/95796>\
**Category:** Elasticsearch\
**Created:** [August 4, 2017, 2:28am UTC](https://discuss.elastic.co/t/cluster-setup-questions-and-rollover-logstash-ingestion-to-elasticsearch/95796 "2017-08-04T02:28:37Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Drift](https://avatars.discourse-cdn.com/v4/letter/d/8797f3/32.png) [@Drift](https://discuss.elastic.co/u/Drift)\
**Post date:** [August 4, 2017, 2:28am UTC](https://discuss.elastic.co/t/cluster-setup-questions-and-rollover-logstash-ingestion-to-elasticsearch/95796/1 "2017-08-04T02:28:38Z")

</div>

Hi,

I am not sure if its possible to do this, but I am setting up a cluster setup and I wanted to make sure I do everything right. I have 3 nodes, all of them have same configuration (392gb ram, 32 core, 3TB SSD each).

I intend to configure it as follows:

Cluster A : Node A - node.master : True, node.data : False  
Cluster A : Node B - node.master : True, node.data : True  
Cluster A : Node C - node.master : false, node.data : True

My source of ingestion is logstash. Logstash will ingest the data into ClusterA with input as csv and output to elasticsearch.

I had a few questions for this setup:

a) Is the cluster config OK? i.e two master and two data nodes.  
b) What is the best way to run logstash?  
My options are:  
- Run Logstash on NodeA, and ingest data pointing output to the network IP of Node B?  
- Can I run Logstash on NodeA and ingest the csv to elasticsearch running on Node A although nodeA is not a data node? (i,e will the nodeA distribute data to other nodes?)  
- for efficiency, will it make a difference if I run logstash on an alltogether different host so that elastic indexing can do its best without hinderance from logstash? -- But in that case, logstash will point over network than localhost.

c) Which node is the best place to run kibana on?

d) Is it possible to have a rollover indexing in elastic? i.e I only want to ingest last 2 weeks of data. I want to remove the oldest two weeks and have a moving window of the data available. That way, I dont run up on disc space?

Is this something relevant ?  
[https://www.elastic.co/guide/en/elasticsearch/reference/master/indices-rollover-index.html](https://www.elastic.co/guide/en/elasticsearch/reference/master/indices-rollover-index.html)

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [August 4, 2017, 7:01am UTC](https://discuss.elastic.co/t/cluster-setup-questions-and-rollover-logstash-ingestion-to-elasticsearch/95796/2 "2017-08-04T07:01:18Z")

</div>

That's a ton of different questions.

First, you may want to have dedicated master nodes (own processes/instances), so that nodes doing the indexing work do not need to deal with master node tasks and vice versa. Also do not use dedicated master nodes to send indexing data two, always send to the master nodes (note: clients can handle this automatically if sniffing is enabled).

You can run logstash and elasticsearch on the same nodes, but this also implies they will potentially steal each other resources, meaning that performance issues will be hard to debug.

The elasticsearch rollover API is used to create a new index, but it does not delete old ones. You should use something like [curator](https://www.elastic.co/guide/en/elasticsearch/client/curator/5.1/index.html) to do those cleanup tasks.

--Alex

---

<div class="post-metadata">

**Author:** ![Drift](https://avatars.discourse-cdn.com/v4/letter/d/8797f3/32.png) [@Drift](https://discuss.elastic.co/u/Drift)\
**Post date:** [August 4, 2017, 8:53am UTC](https://discuss.elastic.co/t/cluster-setup-questions-and-rollover-logstash-ingestion-to-elasticsearch/95796/3 "2017-08-04T08:53:54Z")

</div>

Thankyou @spinscale.

> First, you may want to have dedicated master nodes (own processes/instances), so that nodes doing the indexing work do not need to deal with master node tasks and vice versa.

Yes, all three are running on independent hosts, with its own process/instance.

> Also do not use dedicated master nodes to send indexing data two, always send to the master nodes (note: clients can handle this automatically if sniffing is enabled).

Sorry I did not understand this (pardon me am a newbie to elastic). Does this mean that I should setup logstash to point to the master node only, and then master node will take care of the rest?

> You can run logstash and elasticsearch on the same nodes, but this also implies they will potentially steal each other resources, meaning that performance issues will be hard to debug.

Thanks for the suggestion, totally makes sense.

> The elasticsearch rollover API is used to create a new index, but it does not delete old ones. You should use something like curator1 to do those cleanup tasks.

How does kibana adapt to the changing indices. Say I use rollover API and curator to clean up old one, and all my dashboard/visualization are linked to indexA, is there a way to automatically link kibana dashboards to the rolled-over indexA\_1 ?

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [August 4, 2017, 8:58am UTC](https://discuss.elastic.co/t/cluster-setup-questions-and-rollover-logstash-ingestion-to-elasticsearch/95796/4 "2017-08-04T08:58:07Z")

</div>

you should configure your ingest nodes to **not** send data to the master nodes, but to data nodes only.

Kibana allows you to configure an index pattern, thats why you do not need to worry about that, the pattern should cover the indices created by the rollover API.

---

<div class="post-metadata">

**Author:** ![Drift](https://avatars.discourse-cdn.com/v4/letter/d/8797f3/32.png) [@Drift](https://discuss.elastic.co/u/Drift)\
**Post date:** [August 4, 2017, 9:00am UTC](https://discuss.elastic.co/t/cluster-setup-questions-and-rollover-logstash-ingestion-to-elasticsearch/95796/5 "2017-08-04T09:00:51Z")

</div>

> you should configure your ingest nodes to not send data to the master nodes, but to data nodes only.

Ok, so if I have two data nodes and one logstash ingest node, do I just need to point it to one of them?  
What is the role of the master here? Only for indexing?

Sorry for too many questions 🙂

> Kibana allows you to configure an index pattern, thats why you do not need to worry about that, the pattern should cover the indices created by the rollover API.

Ah, got it! Awesome.

---

<div class="post-metadata">

**Author:** ![Drift](https://avatars.discourse-cdn.com/v4/letter/d/8797f3/32.png) [@Drift](https://discuss.elastic.co/u/Drift)\
**Post date:** [August 5, 2017, 1:47am UTC](https://discuss.elastic.co/t/cluster-setup-questions-and-rollover-logstash-ingestion-to-elasticsearch/95796/6 "2017-08-05T01:47:08Z")

</div>

@spinscale, @warkolm so after a lot of reading and your suggestions, I have come up with this topology, and would like your opinion on this:

I have one logstash node, one dedicated master, one client node and three data nodes out of which two are also master-eligible.

When I configure logstash output, which data node should I point to? should it be the one data-only node?

---

<div class="post-metadata">

**Author:** ![Drift](https://avatars.discourse-cdn.com/v4/letter/d/8797f3/32.png) [@Drift](https://discuss.elastic.co/u/Drift)\
**Post date:** [August 5, 2017, 7:41am UTC](https://discuss.elastic.co/t/cluster-setup-questions-and-rollover-logstash-ingestion-to-elasticsearch/95796/7 "2017-08-05T07:41:00Z")

</div>

Will try answering my own question. would be great if anyone can confirm it.

I point the logstash to all the data nodes (output --\> hosts array)

Can curator/rollover be done dynamically after my data is ingested? Or are there any limitations there?

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [August 6, 2017, 8:58pm UTC](https://discuss.elastic.co/t/cluster-setup-questions-and-rollover-logstash-ingestion-to-elasticsearch/95796/8 "2017-08-06T20:58:21Z")

</div>

you cannot control which node is becoming master (the dedicated vs. the data nodes), so I dont think that this setup makes a lot of sense.

The logstash output should specify only data/client nodes, see [https://www.elastic.co/guide/en/logstash/5.5/plugins-outputs-elasticsearch.html#plugins-outputs-elasticsearch-hosts](https://www.elastic.co/guide/en/logstash/5.5/plugins-outputs-elasticsearch.html#plugins-outputs-elasticsearch-hosts)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 3, 2017, 9:11pm UTC](https://discuss.elastic.co/t/cluster-setup-questions-and-rollover-logstash-ingestion-to-elasticsearch/95796/9 "2017-09-03T21:11:46Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
