# Elastic search Tuning Performance

**URL:** <https://discuss.elastic.co/t/elastic-search-tuning-performance/166528>\
**Category:** Elasticsearch\
**Created:** [January 31, 2019, 10:35am UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528 "2019-01-31T10:35:09Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![fadihaddad](https://avatars.discourse-cdn.com/v4/letter/f/4bbf92/32.png) [@fadihaddad](https://discuss.elastic.co/u/fadihaddad)\
**Post date:** [January 31, 2019, 10:35am UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/1 "2019-01-31T10:35:10Z")

</div>

I have 2 servers with same hardware on both:  
 ![cpu](https://us1.discourse-cdn.com/elastic/original/3X/5/d/5da27fe0138231cb3aae98817c0b1316d225eedc.png)  
16gb Ram  
Virtual Storage 500gb  
Unix System

I tried to index a 250mb file on one of them using logstash, it took 30 mins to index which is a really poor performance. Note I am using default settings in all configurations only changed JVM Heap to 8gb on each server. I installed metric beats to view the metrics on the server and they give

 ![02](https://us1.discourse-cdn.com/elastic/original/3X/2/4/243066d5b992995caa81e3cc7488553e7a9a84ba.png)  
and after like 15mins it became

 ![SnipImage](https://us1.discourse-cdn.com/elastic/original/3X/e/3/e3ae3795651c7be23683e6dfb36f6efb394ef6ba.jpeg)

Can you please tell me what to edit and follow and I already read the tuning speed guidelines but I got lost so I really need help in choosing the setting and which to change. Thank you in advanced

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 31, 2019, 10:39am UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/2 "2019-01-31T10:39:59Z")

</div>

It seems like you have swap enable and are using some of it. We do recommend disabling swap when running Elasticsearch.

How many indices and shards are you concurrently indexing into?

How many concurrent indexing threads/processes are you using?

---

<div class="post-metadata">

**Author:** ![fadihaddad](https://avatars.discourse-cdn.com/v4/letter/f/4bbf92/32.png) [@fadihaddad](https://discuss.elastic.co/u/fadihaddad)\
**Post date:** [January 31, 2019, 10:42am UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/3 "2019-01-31T10:42:29Z")

</div>

i am running now 3 pipelines on logstash and they are behaving with same speed as running one pipeline , and i am indexing them to the same index and 1 shard only, and i forgot to mention i made replicas to 0

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 31, 2019, 10:46am UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/4 "2019-01-31T10:46:41Z")

</div>

If you are indexing into a single index with only one primary shard and no replicas only one node will actually be indexing. If you change the number of primary shards to 2 so each node can get one I would expect you to see improved throughput.

If you are running Logstash on the same host as Elasticsearch, it would be interesting to see how CPU usage breaks down between these two processes.

---

<div class="post-metadata">

**Author:** ![fadihaddad](https://avatars.discourse-cdn.com/v4/letter/f/4bbf92/32.png) [@fadihaddad](https://discuss.elastic.co/u/fadihaddad)\
**Post date:** [January 31, 2019, 10:56am UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/5 "2019-01-31T10:56:23Z")

</div>

i made 1 node master and the other data so if I removed the master and made it 2 shards it will improve more and i am running logstash on the data node now

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 31, 2019, 11:01am UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/6 "2019-01-31T11:01:07Z")

</div>

If the nodes have the same specification, make them both data nodes. You can still keep just one node as master eligible. Performance issues are also often easier to troubleshoot of Logstash runs on a separate host.

---

<div class="post-metadata">

**Author:** ![fadihaddad](https://avatars.discourse-cdn.com/v4/letter/f/4bbf92/32.png) [@fadihaddad](https://discuss.elastic.co/u/fadihaddad)\
**Post date:** [January 31, 2019, 11:02am UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/7 "2019-01-31T11:02:37Z")

</div>

okay, and i read about resfresh intervals and node buffers any setting that can improve indexing speed

---

<div class="post-metadata">

**Author:** ![juka](https://avatars.discourse-cdn.com/v4/letter/j/8491ac/32.png) [@juka](https://discuss.elastic.co/u/juka)\
**Post date:** [January 31, 2019, 11:40am UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/8 "2019-01-31T11:40:35Z")

</div>

If indexing takes too long but Elasticsearch seems alright, maybe Logstash is on the slow side and you need to increase the amount of workers in your configfile.  
Check your config with  
`curl -XGET 'localhost:9600/_node/pipelines?pretty'`  
and look for workers.  
If workers is just one look here [https://www.elastic.co/guide/en/logstash/6.6/tuning-logstash.html](https://www.elastic.co/guide/en/logstash/6.6/tuning-logstash.html)  
and increase them.

Good Luck!

---

<div class="post-metadata">

**Author:** ![fadihaddad](https://avatars.discourse-cdn.com/v4/letter/f/4bbf92/32.png) [@fadihaddad](https://discuss.elastic.co/u/fadihaddad)\
**Post date:** [January 31, 2019, 11:43am UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/9 "2019-01-31T11:43:12Z")

</div>

i tried to increase the number of workers and it gave same results as before

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 31, 2019, 11:49am UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/10 "2019-01-31T11:49:14Z")

</div>

What is the full output of the [cluster stats API](https://www.elastic.co/guide/en/elasticsearch/reference/6.6/cluster-stats.html) and the [node stats API](https://www.elastic.co/guide/en/elasticsearch/reference/6.6/cluster-nodes-stats.html) For the data node?

---

<div class="post-metadata">

**Author:** ![juka](https://avatars.discourse-cdn.com/v4/letter/j/8491ac/32.png) [@juka](https://discuss.elastic.co/u/juka)\
**Post date:** [January 31, 2019, 11:50am UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/11 "2019-01-31T11:50:01Z")

</div>

Alright, what kind of data do you import? Does it contain a lot of rows. Is it a Json file? What does your logstash processing pipeline look like INPUT - FILTER - OUPUT?

---

<div class="post-metadata">

**Author:** ![fadihaddad](https://avatars.discourse-cdn.com/v4/letter/f/4bbf92/32.png) [@fadihaddad](https://discuss.elastic.co/u/fadihaddad)\
**Post date:** [January 31, 2019, 12:39pm UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/12 "2019-01-31T12:39:51Z")

</div>

its a csv file with a lot of columns and yeah it has a lot of rows  
here is the logstash config

input {  
file {  
path =\> "C:/Users/Fadi/Desktop/internship/procera\_Justice-\*.csv"  
start\_position =\> "beginning"

}  
}  
filter {  
csv {  
separator =\> "|"  
columns =\>  
["egressinterface", "ingressinterface", "applicationid", "destinationipaddress"  
, "destinationtransportport", "sourcetransportport","flowstartmilliseconds","flowendmilliseconds","flowstartseconds","flowendseconds","procerainternalrtt","proceraexternalrtt","proceraservice",  
"proceraserverhostname","proceraincomingoctets","proceraoutgoingoctets","octetdeltacount","proceraqoeincomingexternal"  
,"proceraqoeincominginternal","proceraqoeoutgoingexternal","proceraqoeoutgoinginternal","sourceipaddress","procerasubscriberidentifier","proceraggsn","procerasgsn",  
"procerarnc","proceraapn","procerauserlocationinformation"  
,"procerarat","observationpointid","proceradeviceid","proceraimsi","chargingid","ProceraServiceObject","protocol\_identifier","procerahttpresponsestatus","procerahttpuseragent","procerahttpurl"]  
convert =\> {

"egressinterface" =\> "integer"  
"ingressinterface" =\> "integer"  
"applicationid" =\> "integer"  
"destinationtransportport" =\> "integer"  
"sourcetransportport" =\> "integer"  
"procerainternalrtt" =\> "integer"  
"proceraexternalrtt" =\> "integer"  
"proceraincomingoctets" =\> "integer"  
"proceraoutgoingoctets" =\> "integer"  
"octetdeltacount" =\> "integer"  
"proceraqoeincomingexternal" =\> "integer"  
"proceraqoeincominginternal" =\> "integer"  
"proceraqoeoutgoingexternal" =\> "integer"  
"proceraqoeoutgoinginternal" =\> "integer"  
"observationpointid" =\> "integer"  
"chargingid" =\> "integer"  
"protocol\_identifier" =\> "integer"  
"procerahttpresponsestatus" =\> "integer"

```
}

```

}

date {  
match =\> ["flowstartseconds", "yyyy-M-d HH:mm:ss", "yyyy-M-dd HH:mm:ss", "yyyy-MM-d HH:m:ss", "yyyy-MM-dd HH:mm:ss"]  
target =\> "flowstartseconds"  
locale =\> "en"  
timezone =\> "UTC"

}  
date {  
match =\> ["flowendseconds", "yyyy-M-d HH:mm:ss", "yyyy-M-dd HH:mm:ss", "yyyy-MM-d HH:mm:ss", "yyyy-MM-dd HH:mm:ss"]  
target =\> "flowendseconds"  
locale =\> "en"  
timezone =\> "UTC"

}

}  
output {  
elasticsearch {  
hosts =\> "localhost"  
index =\> "procera"  
document\_type =\> "data"

}  
stdout {}  
}

---

<div class="post-metadata">

**Author:** ![fadihaddad](https://avatars.discourse-cdn.com/v4/letter/f/4bbf92/32.png) [@fadihaddad](https://discuss.elastic.co/u/fadihaddad)\
**Post date:** [January 31, 2019, 12:44pm UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/13 "2019-01-31T12:44:15Z")

</div>

this is the cluster api  
{  
"\_nodes" : {  
"total" : 2,  
"successful" : 2,  
"failed" : 0  
},  
"cluster\_name" : "elastic",  
"cluster\_uuid" : "d7Yg9IShRC6Q9-NbDAv1bw",  
"timestamp" : 1548938459707,  
"status" : "green",  
"indices" : {  
"count" : 27,  
"shards" : {  
"total" : 27,  
"primaries" : 27,  
"replication" : 0.0,  
"index" : {  
"shards" : {  
"min" : 1,  
"max" : 1,  
"avg" : 1.0  
},  
"primaries" : {  
"min" : 1,  
"max" : 1,  
"avg" : 1.0  
},  
"replication" : {  
"min" : 0.0,  
"max" : 0.0,  
"avg" : 0.0  
}  
}  
},  
"docs" : {  
"count" : 6816580,  
"deleted" : 1013  
},  
"store" : {  
"size" : "4.5gb",  
"size\_in\_bytes" : 4908194710  
},  
"fielddata" : {  
"memory\_size" : "32.4kb",  
"memory\_size\_in\_bytes" : 33256,  
"evictions" : 0  
},  
"query\_cache" : {  
"memory\_size" : "528kb",  
"memory\_size\_in\_bytes" : 540737,  
"total\_count" : 56974,  
"hit\_count" : 6179,  
"miss\_count" : 50795,  
"cache\_size" : 517,  
"cache\_count" : 722,  
"evictions" : 205  
},  
"completion" : {  
"size" : "0b",  
"size\_in\_bytes" : 0  
},  
"segments" : {  
"count" : 244,  
"memory" : "11.8mb",  
"memory\_in\_bytes" : 12373717,  
"terms\_memory" : "7.6mb",  
"terms\_memory\_in\_bytes" : 8025412,  
"stored\_fields\_memory" : "1.7mb",  
"stored\_fields\_memory\_in\_bytes" : 1848112,  
"term\_vectors\_memory" : "0b",  
"term\_vectors\_memory\_in\_bytes" : 0,  
"norms\_memory" : "199.3kb",  
"norms\_memory\_in\_bytes" : 204096,  
"points\_memory" : "784kb",  
"points\_memory\_in\_bytes" : 802865,  
"doc\_values\_memory" : "1.4mb",  
"doc\_values\_memory\_in\_bytes" : 1493232,  
"index\_writer\_memory" : "16.3mb",  
"index\_writer\_memory\_in\_bytes" : 17192532,  
"version\_map\_memory" : "520b",  
"version\_map\_memory\_in\_bytes" : 520,  
"fixed\_bit\_set" : "288.1kb",  
"fixed\_bit\_set\_memory\_in\_bytes" : 295024,  
"max\_unsafe\_auto\_id\_timestamp" : 1548927891218,  
"file\_sizes" : { }  
}  
},  
"nodes" : {  
"count" : {  
"total" : 2,  
"data" : 1,  
"coordinating\_only" : 0,  
"master" : 1,  
"ingest" : 2  
},  
"versions" : [  
"6.5.4"  
],  
"os" : {  
"available\_processors" : 8,  
"allocated\_processors" : 8,  
"names" : [  
{  
"name" : "Linux",  
"count" : 2  
}  
],  
"mem" : {  
"total" : "31gb",  
"total\_in\_bytes" : 33315569664,  
"free" : "1.4gb",  
"free\_in\_bytes" : 1608933376,  
"used" : "29.5gb",  
"used\_in\_bytes" : 31706636288,  
"free\_percent" : 5,  
"used\_percent" : 95  
}  
},  
"process" : {  
"cpu" : {  
"percent" : 9  
},  
"open\_file\_descriptors" : {  
"min" : 296,  
"max" : 414,  
"avg" : 355  
}  
},  
"jvm" : {  
"max\_uptime" : "3.8h",  
"max\_uptime\_in\_millis" : 14019209,  
"versions" : [  
{  
"version" : "1.8.0\_161",  
"vm\_name" : "OpenJDK 64-Bit Server VM",  
"vm\_version" : "25.161-b14",  
"vm\_vendor" : "Oracle Corporation",  
"count" : 2  
}  
],  
"mem" : {  
"heap\_used" : "2.6gb",  
"heap\_used\_in\_bytes" : 2883739064,  
"heap\_max" : "15.9gb",  
"heap\_max\_in\_bytes" : 17110138880  
},  
"threads" : 128  
},  
"fs" : {  
"total" : "982.3gb",  
"total\_in\_bytes" : 1054755577856,  
"free" : "907.5gb",  
"free\_in\_bytes" : 974524428288,  
"available" : "865.6gb",  
"available\_in\_bytes" : 929490112512  
},  
"plugins" : ,  
"network\_types" : {  
"transport\_types" : {  
"security4" : 2  
},  
"http\_types" : {  
"security4" : 2  
}  
}  
}

---

<div class="post-metadata">

**Author:** ![fadihaddad](https://avatars.discourse-cdn.com/v4/letter/f/4bbf92/32.png) [@fadihaddad](https://discuss.elastic.co/u/fadihaddad)\
**Post date:** [January 31, 2019, 12:46pm UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/14 "2019-01-31T12:46:34Z")

</div>

the nodes api is too big to post here

---

<div class="post-metadata">

**Author:** ![juka](https://avatars.discourse-cdn.com/v4/letter/j/8491ac/32.png) [@juka](https://discuss.elastic.co/u/juka)\
**Post date:** [January 31, 2019, 1:16pm UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/15 "2019-01-31T13:16:06Z")

</div>

when you tested more workers on logstash did you add the line  
workers =\> 20  
to your elasticsearch output section? I can't see it currently. I do not mean pipelines, just workers that would go under index =\> "procera" for example

Can you monitor your network throughput of logstash's port with tcpdump, dstat or something you are familiar with? Your Logstash surely has quite a bit to do. Maybe you can also check the processes with top, htop, etc. and see what's using how much on your system.

Your dataflow seems to me CSV File - \> Logastash -\> Elasticsearch  
One of these is the bottleneck, maybe you can split the file into smaller chunks of 50MB each?

---

<div class="post-metadata">

**Author:** ![fadihaddad](https://avatars.discourse-cdn.com/v4/letter/f/4bbf92/32.png) [@fadihaddad](https://discuss.elastic.co/u/fadihaddad)\
**Post date:** [January 31, 2019, 1:24pm UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/16 "2019-01-31T13:24:41Z")

</div>

i just added 10 workers to the pipeline of procera in pipelines.yml file the default was 4, and i will check this network throughput plus i cant split the file because my final result every 1 hour I receive 4 of this file and I want to index them in a really fast manner

---

<div class="post-metadata">

**Author:** ![fadihaddad](https://avatars.discourse-cdn.com/v4/letter/f/4bbf92/32.png) [@fadihaddad](https://discuss.elastic.co/u/fadihaddad)\
**Post date:** [January 31, 2019, 2:55pm UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/17 "2019-01-31T14:55:50Z")

</div>

![cpu](https://us1.discourse-cdn.com/elastic/original/3X/5/d/5da27fe0138231cb3aae98817c0b1316d225eedc.png)  
this is the top command

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 28, 2019, 3:10pm UTC](https://discuss.elastic.co/t/elastic-search-tuning-performance/166528/18 "2019-02-28T15:10:38Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
