# Elasticsearch Architecture Problem

**URL:** <https://discuss.elastic.co/t/elasticsearch-architecture-problem/162915>\
**Category:** Elasticsearch\
**Created:** [January 4, 2019, 8:46am UTC](https://discuss.elastic.co/t/elasticsearch-architecture-problem/162915 "2019-01-04T08:46:45Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![josephmanalo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/josephmanalo/32/38163_2.png) [@josephmanalo](https://discuss.elastic.co/u/josephmanalo)\
**Post date:** [January 4, 2019, 8:46am UTC](https://discuss.elastic.co/t/elasticsearch-architecture-problem/162915/1 "2019-01-04T08:46:45Z")

</div>

Hi guys. The picture below is the existing data pipeline of a certain client.

Their MAIN PROBLEM are:

1. Not all the data are being ingested in Node 2 and Node 3
2. Slow retrieval of data

Do you have any idea what might be the cause of the problem?

1. Is it a hardware related? If they need to upgrade what is the ideal specs per Node?
2. Will adding another node will help in solving the problem?
3. Will it solve the problem by implementing 2nd picture or 3rd picture below?
4. Also, they want to implement Machine Learning in their data, what can you suggest for them?

Other Details:

- The machines and the nodes are just connected to a same network
- There are other machines which are connected from a different network but sends the data in logstash (Node 1) ![49206385_2237499219623044_8859078541510180864_o](https://us1.discourse-cdn.com/elastic/original/3X/e/3/e3ead12f5292158653eea2ae638dffea0bbb07c0.jpeg) ![49899279_2237578882948411_7035400178831982592_o](https://us1.discourse-cdn.com/elastic/original/3X/2/0/20b610f17786c50107f89f8eb390306c4e2e0f90.jpeg) ![49205917_2237625009610465_4643750758200639488_o](https://us1.discourse-cdn.com/elastic/original/3X/5/f/5f0f16eb6afa3e274a2934e03a55f57b14d69da6.jpeg)

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 4, 2019, 9:03am UTC](https://discuss.elastic.co/t/elasticsearch-architecture-problem/162915/2 "2019-01-04T09:03:00Z")

</div>

> [@josephmanalo](#):
>
> Hi guys.

Hey @josephmanalo

Please note that fortunately we are not all guys here 🙂

> [@josephmanalo](#):
>
> Do you have any idea what might be the cause of the problem?

Do you have any elasticsearch monitoring activated so you can understand may be better what's the cause of this?

It sounds like you are using HDD disks for some nodes and SSD on other nodes. That's an issue IMO.

I also wonder why are you using Logstash for?

May be you have also too many shards per node.  
What is the output of:

```
GET /_cat/health?v
GET /_cat/indices?v

```

May I suggest you look at the following resources about sizing:

> **[Quantitative Cluster Sizing](https://www.elastic.co/elasticon/conf/2016/sf/quantitative-cluster-sizing)**

> **[How many shards should I have in my Elasticsearch cluster?](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster)**

> **[NetSecureDay: Managing your Black Friday Logs](https://speakerdeck.com/elastic/netsecureday-managing-your-black-friday-logs)**
>
> Surveiller une application complexe n’est pas une tâche aisée, mais avec les bons outils, ce n’est pas si sorcier. Néanmoins, des périodes fortes telles que les opérations de type « Black Friday » (Vendredi noir) ou période de Noël peuvent pousser...

[![](https://us1.discourse-cdn.com/elastic/original/3X/7/c/7c2edebd5bac194c58a5df87ede6872cb41d61a8.jpeg "Managing your Black Friday Logs, Pablo Musa, Elastic, TechSummit Amsterdam") ](https://www.youtube.com/watch?v=ilP7tG6tabI)

And [https://www.elastic.co/webinars/using-rally-to-get-your-elasticsearch-cluster-size-right](https://www.elastic.co/webinars/using-rally-to-get-your-elasticsearch-cluster-size-right)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 4, 2019, 9:05am UTC](https://discuss.elastic.co/t/elasticsearch-architecture-problem/162915/3 "2019-01-04T09:05:56Z")

</div>

Elasticsearch assumes all nodes in the cluster are equal, so having different types of hardware can cause a problem. Are you sure your cluster has formed properly and that you have set `discovery.zen.minimum_master_nodes` correctly according to [these guidelines](https://www.elastic.co/guide/en/elasticsearch/reference/6.5/modules-node.html#split-brain)?

---

<div class="post-metadata">

**Author:** ![josephmanalo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/josephmanalo/32/38163_2.png) [@josephmanalo](https://discuss.elastic.co/u/josephmanalo)\
**Post date:** [January 4, 2019, 9:09am UTC](https://discuss.elastic.co/t/elasticsearch-architecture-problem/162915/4 "2019-01-04T09:09:41Z")

</div>

Hi! thanks for the reply.  
i'll check out what you've sent

Here's the result of the ff:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/5/b/5bb4bd423e9595c72d640c02b2f549fab4f260b9.png)

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/f/0/f02b963fc896b2783c5bf3c2e011ffe289ef1898.png)

---

<div class="post-metadata">

**Author:** ![josephmanalo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/josephmanalo/32/38163_2.png) [@josephmanalo](https://discuss.elastic.co/u/josephmanalo)\
**Post date:** [January 4, 2019, 9:23am UTC](https://discuss.elastic.co/t/elasticsearch-architecture-problem/162915/5 "2019-01-04T09:23:39Z")

</div>

> [@dadoonet](#):
>
> I also wonder why are you using Logstash for?

We are using logstash to append a new column to the data and push it to its corresponding node.  
We tried solving the problem by dividing the data some modules will be stored in Node 2 and other modules will be stored in Node 3.  
But still it doesn't solve the problem.

---

<div class="post-metadata">

**Author:** ![josephmanalo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/josephmanalo/32/38163_2.png) [@josephmanalo](https://discuss.elastic.co/u/josephmanalo)\
**Post date:** [January 4, 2019, 9:24am UTC](https://discuss.elastic.co/t/elasticsearch-architecture-problem/162915/6 "2019-01-04T09:24:09Z")

</div>

Hi. Thanks for the suggestion, I'll try get into this and I'll update you back. Thanks!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 4, 2019, 9:25am UTC](https://discuss.elastic.co/t/elasticsearch-architecture-problem/162915/7 "2019-01-04T09:25:53Z")

</div>

You seem to have a lot of shards being generated daily given the small data volumes. I would recommend reducing this significantly e.g. by changing to use only a single primary shard per index, use weekly or monthly indices or simply consolidate the data into fewer indices.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 4, 2019, 9:34am UTC](https://discuss.elastic.co/t/elasticsearch-architecture-problem/162915/8 "2019-01-04T09:34:27Z")

</div>

Please don't post images of text as they are hardly readable and not searchable.

Instead paste the text and format it with `</>` icon. Check the preview window.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 4, 2019, 9:35am UTC](https://discuss.elastic.co/t/elasticsearch-architecture-problem/162915/9 "2019-01-04T09:35:35Z")

</div>

I'm not sure I understood. Anyway, may be look at the node ingest feature which might be enough to replace your logstash pipeline.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 1, 2019, 9:35am UTC](https://discuss.elastic.co/t/elasticsearch-architecture-problem/162915/10 "2019-02-01T09:35:35Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
