# Need to test a index with terabytes of data, how can I do

**URL:** <https://discuss.elastic.co/t/need-to-test-a-index-with-terabytes-of-data-how-can-i-do/272787>\
**Category:** Elasticsearch\
**Tags:** rally\
**Created:** [May 12, 2021, 8:56am UTC](https://discuss.elastic.co/t/need-to-test-a-index-with-terabytes-of-data-how-can-i-do/272787 "2021-05-12T08:56:50Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![GITnewfish](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gitnewfish/32/87303_2.png) [@GITnewfish](https://discuss.elastic.co/u/GITnewfish)\
**Post date:** [May 12, 2021, 8:56am UTC](https://discuss.elastic.co/t/need-to-test-a-index-with-terabytes-of-data-how-can-i-do/272787/1 "2021-05-12T08:56:50Z")

</div>

I need to test performance of appand to a index with terabytes of data, and the index mapping is customed.  
Is there a solution to this scenario?

Refer to this topic: [Increase data size in Rally existing tracks](https://discuss.elastic.co/t/increase-data-size-in-rally-existing-tracks/116514)

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/f/3/f3252b147112974c08076ef3be7d5d70fb2a6ec4.png)

Whether to support looping to write a fixed index?

I tried, but got error.

---

<div class="post-metadata">

**Author:** ![GITnewfish](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gitnewfish/32/87303_2.png) [@GITnewfish](https://discuss.elastic.co/u/GITnewfish)\
**Post date:** [May 13, 2021, 6:35am UTC](https://discuss.elastic.co/t/need-to-test-a-index-with-terabytes-of-data-how-can-i-do/272787/2 "2021-05-13T06:35:17Z")

</div>

Eventually I did this by writing data to large files.  
However, it is too slow to compress large files.  
Is there a big difference between source-file performance measured using the original file documents.json and the compressed file documents.json.bz2?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 13, 2021, 7:02am UTC](https://discuss.elastic.co/t/need-to-test-a-index-with-terabytes-of-data-how-can-i-do/272787/3 "2021-05-13T07:02:33Z")

</div>

Have a look at the [rally-eventdata-track](https://github.com/elastic/rally-eventdata-track). Unlike other tracks this does not rely on data in files but does instead generate data at runtime based on a set of probability distributions. This makes it possible to generate very large amounts of data with just track configuration. You can use this as is and just create a modified config or use this as a base for generating your own track handling your particular type of data.

This [blog post](https://www.elastic.co/blog/querying-a-petabyte-of-cloud-storage-in-10-minutes) describes how it was used to generate 4 TB of indexed data for a set of storage benchmarks. [This video](https://www.elastic.co/webinars/using-rally-to-get-your-elasticsearch-cluster-size-right) also discusses this track and its use.

---

<div class="post-metadata">

**Author:** ![GITnewfish](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gitnewfish/32/87303_2.png) [@GITnewfish](https://discuss.elastic.co/u/GITnewfish)\
**Post date:** [May 13, 2021, 7:31am UTC](https://discuss.elastic.co/t/need-to-test-a-index-with-terabytes-of-data-how-can-i-do/272787/4 "2021-05-13T07:31:14Z")

</div>

This seems to be an extension based on the event-data log type, but my scenario requires testing with our actual business logs.Can rally-EventData-Track extend the data volume with custom index mapping?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 13, 2021, 8:23am UTC](https://discuss.elastic.co/t/need-to-test-a-index-with-terabytes-of-data-how-can-i-do/272787/5 "2021-05-13T08:23:40Z")

</div>

I suspect you may need to create a new custom track.

---

<div class="post-metadata">

**Author:** ![GITnewfish](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gitnewfish/32/87303_2.png) [@GITnewfish](https://discuss.elastic.co/u/GITnewfish)\
**Post date:** [May 14, 2021, 6:02am UTC](https://discuss.elastic.co/t/need-to-test-a-index-with-terabytes-of-data-how-can-i-do/272787/6 "2021-05-14T06:02:48Z")

</div>

Yes，does rally-EventData-Track support according custom track to generate terabytes of data?  
by the way, doce source-file support config more than one file? I think this will be convenient to Increase and decrease doc amount according to different requirements

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 14, 2021, 6:23am UTC](https://discuss.elastic.co/t/need-to-test-a-index-with-terabytes-of-data-how-can-i-do/272787/7 "2021-05-14T06:23:43Z")

</div>

If the event format created by the rally-eventdata-track can be used it is relatively easy to create a new challenge that can generate a very large amount of data. You may also be able to alter the mappings used if necessary. If you need a specific event format and mappings you probably customize the track or generate files.

I am still not sure I fully understand what you are looking to test. Could you please elaborate on what you are looking to test and achieve? Are you looking to index into a set of time-based indices and see hoe the cluster performs with large amounts of data or are you going to index into a single index that will grow very large? What is your use case?

---

<div class="post-metadata">

**Author:** ![GITnewfish](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gitnewfish/32/87303_2.png) [@GITnewfish](https://discuss.elastic.co/u/GITnewfish)\
**Post date:** [May 24, 2021, 8:30am UTC](https://discuss.elastic.co/t/need-to-test-a-index-with-terabytes-of-data-how-can-i-do/272787/8 "2021-05-24T08:30:53Z")

</div>

We want to test sending 2T of specific data to an index to see how large the cluster can reach.Now I have solved this problem, but I had a puzzle during the testing.

first question: Does the final result of throughput include replicas? I think the result was only primary shard and it seems according to samples to culculate: sum(bulk\_size)/time\_period. Does this right?  
I have read [How Write throughput is calculated in Rally - #2 by dliappis](https://discuss.elastic.co/t/how-write-throughput-is-calculated-in-rally/188997/2)

second question: When the index\_append start, i use iostat to monitor the io, sometimes the readbyte is zero, why? the bulk\_indexing\_client\_num is 64.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/f/a/fa91545b54c26cf600249db0f4facf9153a9dd47.png)

third question: How can I know rally has start the amount of clients that I set?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 21, 2021, 8:31am UTC](https://discuss.elastic.co/t/need-to-test-a-index-with-terabytes-of-data-how-can-i-do/272787/9 "2021-06-21T08:31:18Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
