# The problem about runing a big dataset

**URL:** <https://discuss.elastic.co/t/the-problem-about-runing-a-big-dataset/69862>\
**Category:** Elasticsearch\
**Tags:** rally\
**Created:** [December 23, 2016, 2:27am UTC](https://discuss.elastic.co/t/the-problem-about-runing-a-big-dataset/69862 "2016-12-23T02:27:06Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![a943013827](https://avatars.discourse-cdn.com/v4/letter/a/f475e1/32.png) [@a943013827](https://discuss.elastic.co/u/a943013827)\
**Post date:** [December 23, 2016, 2:27am UTC](https://discuss.elastic.co/t/the-problem-about-runing-a-big-dataset/69862/1 "2016-12-23T02:27:06Z")

</div>

hi:  
I am runing rally benchmark-only indexing a 365G(1 billions docs) dataset,  
Racing on track [geonames], challenge [append-only-no-conflicts-while-searching] and car [external]

I check the docs count every 30 seconds,when document count reach 286309645 , the indexing-speed  
reduce to 2000 docs/s from 9000 docs/s, and the cpu usage reduce to 300% from 1200% (12 core).

what can be cause? thank you.

---

<div class="post-metadata">

**Author:** ![a943013827](https://avatars.discourse-cdn.com/v4/letter/a/f475e1/32.png) [@a943013827](https://discuss.elastic.co/u/a943013827)\
**Post date:** [December 23, 2016, 2:59am UTC](https://discuss.elastic.co/t/the-problem-about-runing-a-big-dataset/69862/2 "2016-12-23T02:59:14Z")

</div>

![](https://us1.discourse-cdn.com/elastic/original/2X/d/d581a85bae888992216fc01ce1c0b8cd460ca66a.png)

the cpu usage of esrally

---

<div class="post-metadata">

**Author:** ![danielmitterdorfer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielmitterdorfer/32/110510_2.png) [@danielmitterdorfer](https://discuss.elastic.co/u/danielmitterdorfer)\
**Post date:** [December 23, 2016, 8:36am UTC](https://discuss.elastic.co/t/the-problem-about-runing-a-big-dataset/69862/3 "2016-12-23T08:36:17Z")

</div>

Hi @a943013827,

hard to tell, that could be anything from GC issues to the cluster simply merging segments. You should inspect the Elastcisearch logs for hints and also look at hot threads (no support for that in Rally).

If you use the defaults, then Rally starts Elasticsearch on port 39200, i.e. `curl http://localhost:39200/_nodes/hot_threads` should do the trick. If that does not reveal anything (but I'm pretty sure it reveals the cause), you can check the GC logs e.g. by starting Rally with the `--telemetry=gc` option.

Daniel

---

<div class="post-metadata">

**Author:** ![a943013827](https://avatars.discourse-cdn.com/v4/letter/a/f475e1/32.png) [@a943013827](https://discuss.elastic.co/u/a943013827)\
**Post date:** [December 24, 2016, 10:33am UTC](https://discuss.elastic.co/t/the-problem-about-runing-a-big-dataset/69862/4 "2016-12-24T10:33:15Z")

</div>

Hi Daniel,  
I check the hot\_threads,  
Normal time :

 ![](https://us1.discourse-cdn.com/elastic/original/2X/1/15e7136735eda6c6c34d8a6447761c7ea147f76e.png)  
 ![](https://us1.discourse-cdn.com/elastic/original/2X/4/452878c315adc5e5cd40913c01bccc333ffd6dc2.png)  
Abnormal time：  
 ![](https://us1.discourse-cdn.com/elastic/original/2X/3/3f2ecdc7b8513c1a6866f1aee10ac1479fb0d42c.png)  
 ![](https://us1.discourse-cdn.com/elastic/original/2X/2/23ff6a7b72329bba34ce93330d5a2acf180f2afc.png)  
 ![](https://us1.discourse-cdn.com/elastic/original/2X/5/51b9c2ce07985f4326657d18fd30d7673a014b16.png)  
The problem may be ...?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 24, 2016, 5:24pm UTC](https://discuss.elastic.co/t/the-problem-about-runing-a-big-dataset/69862/5 "2016-12-24T17:24:45Z")

</div>

What kind of hardware are you running on? Is there any change in disk IO patterns between the start of the run when speed is good and when it slows down? Is the slowdown gradual or sudden? What does the `_cat/indices` API report once it has slowed down?

---

<div class="post-metadata">

**Author:** ![a943013827](https://avatars.discourse-cdn.com/v4/letter/a/f475e1/32.png) [@a943013827](https://discuss.elastic.co/u/a943013827)\
**Post date:** [December 25, 2016, 1:49am UTC](https://discuss.elastic.co/t/the-problem-about-runing-a-big-dataset/69862/6 "2016-12-25T01:49:00Z")

</div>

Hi Christian\_Dahlqvist :

> [@Christian\_Dahlqvist](#):
>
> What kind of hardware are you running on?

hardware : cpu(12core) disk(2T remain).

> [@Christian\_Dahlqvist](#):
>
> Is there any change in disk IO patterns between the start of the run when speed is good and when it slows down?

I am sure that there no change in disk IO during running rally.

> [@Christian\_Dahlqvist](#):
>
> Is the slowdown gradual or sudden?

The indexing-speed reduce from 9000docs/s to 5000dos/s gradually and suddenly reduce to 2000docs/s at 280000000+ docs and I have try it twice and the problem reappeared.

> [@Christian\_Dahlqvist](#):
>
> What does the \_cat/indices API report once it has slowed down?

5 shards 0 replica, and the health shoud be green (I will check it again at Monday)

---

<div class="post-metadata">

**Author:** ![a943013827](https://avatars.discourse-cdn.com/v4/letter/a/f475e1/32.png) [@a943013827](https://discuss.elastic.co/u/a943013827)\
**Post date:** [December 26, 2016, 8:36am UTC](https://discuss.elastic.co/t/the-problem-about-runing-a-big-dataset/69862/7 "2016-12-26T08:36:10Z")

</div>

Hi Daniel:  
Now docs reach 0.45 billion , and the indexing-speed reduce to 400 docs/s .  
I list top 30 threads:  
`es.nodes.hot_threads(ignore_idle_threads=False,threads=40,timeout="30s")`

the hot\_threads info (now):

```
   11.0% (54.7ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][bulk][T#8]'
    9.5% (47.2ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][[geonames][0]: Lucene Merge Thread #41502]'
    8.4% (41.9ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][bulk][T#9]'
    7.9% (39.4ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][[geonames][2]: Lucene Merge Thread #41431]'
    7.2% (36.1ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][bulk][T#3]'
    7.1% (35.3ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][[geonames][0]: Lucene Merge Thread #41390]'
    6.6% (33ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][bulk][T#1]'
    6.6% (32.9ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][bulk][T#5]'
    6.6% (32.8ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][[geonames][4]: Lucene Merge Thread #41313]'
    6.6% (32.8ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][bulk][T#12]'
    6.2% (30.7ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][bulk][T#17]'
    5.9% (29.2ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][[geonames][3]: Lucene Merge Thread #41549]'
    5.8% (28.9ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][bulk][T#18]'
    5.7% (28.6ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][[geonames][1]: Lucene Merge Thread #41493]'
    5.7% (28.5ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][refresh][T#1]'
    5.6% (27.9ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][[geonames][2]: Lucene Merge Thread #41430]'
    5.5% (27.7ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][bulk][T#4]'
    5.5% (27.5ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][search][T#27]'
    5.5% (27.4ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][[geonames][0]: Lucene Merge Thread #41512]'
    5.3% (26.7ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][[geonames][3]: Lucene Merge Thread #41550]'
    5.3% (26.6ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][bulk][T#6]'
    5.3% (26.4ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][bulk][T#7]'
    5.0% (25.1ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][search][T#18]'
    5.0% (24.7ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][bulk][T#10]'
    4.9% (24.4ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][bulk][T#2]'
    4.8% (24.1ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][bulk][T#19]'
    4.8% (23.7ms out of 500ms) cpu usage by thread 'elasticsearch[M-Twins][[geonames][4]: Lucene Merge Thread #41327]'

```

the hot\_threads info used to be (just start):

```
   51.4% (257ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#5]'
   50.3% (251.6ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#3]'
   49.3% (246.4ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][[geonames][2]: Lucene Merge Thread #71]'
   47.4% (237.1ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#4]'
   46.4% (232ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#23]'
   45.2% (226ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#17]'
   44.7% (223.7ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#6]'
   44.4% (221.8ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#22]'
   44.2% (220.8ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#15]'
   44.1% (220.6ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][[geonames][1]: Lucene Merge Thread #62]'
   43.9% (219.7ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#14]'
   40.5% (202.5ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#1]'
   38.3% (191.3ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][[geonames][1]: Lucene Merge Thread #70]'
   34.5% (172.5ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#16]'
   34.1% (170.7ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#20]'
   33.3% (166.4ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#18]'
   33.0% (164.9ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#13]'
   32.4% (161.9ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][[geonames][3]: Lucene Merge Thread #70]'
   31.2% (156.2ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#9]'
   31.2% (155.7ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#10]'
   29.9% (149.4ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#8]'
   29.7% (148.6ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#2]'
   24.9% (124.2ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#7]'
   24.3% (121.5ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#19]'
   20.3% (101.5ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#24]'
   20.2% (100.8ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#12]'
   19.6% (97.9ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#11]'
   17.4% (86.8ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][bulk][T#21]'
   15.2% (75.9ms out of 500ms) cpu usage by thread 'elasticsearch[Toad][[geonames][0]: Lucene Merge Thread #70]'
```

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 26, 2016, 8:50am UTC](https://discuss.elastic.co/t/the-problem-about-runing-a-big-dataset/69862/8 "2016-12-26T08:50:26Z")

</div>

> [@a943013827](#):
>
> I am sure that there no change in disk IO during running rally.

Is this based on monitoring data? Based on your hot threads it looks like Elasticsearch is spending a fair amount of time merging Lucene segments. What type of storage do you have?

What is your index `refresh_interval` set to?

---

<div class="post-metadata">

**Author:** ![a943013827](https://avatars.discourse-cdn.com/v4/letter/a/f475e1/32.png) [@a943013827](https://discuss.elastic.co/u/a943013827)\
**Post date:** [December 26, 2016, 9:48am UTC](https://discuss.elastic.co/t/the-problem-about-runing-a-big-dataset/69862/9 "2016-12-26T09:48:42Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> Is this based on monitoring data?

sorry，I misunderstood what you mean  
the disk IO reduced ,but there is a large IO every 5-10 seconds.

> [@Christian\_Dahlqvist](#):
>
> What type of storage do you have? What is your index refresh\_interval set to?

All configurations are default. geonames's mapping，refresh\_interval=1。

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 26, 2016, 10:00am UTC](https://discuss.elastic.co/t/the-problem-about-runing-a-big-dataset/69862/10 "2016-12-26T10:00:54Z")

</div>

For maximum indexing performance you should set the refresh\_interval to a larger value in order to reduce the merging activity. Set it to 10 or 30 to see if it makes any difference. This will however increase the time it takes for records to become searchable.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 23, 2017, 10:01am UTC](https://discuss.elastic.co/t/the-problem-about-runing-a-big-dataset/69862/11 "2017-01-23T10:01:02Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
