# Increase data size in Rally with existing tracks

**URL:** <https://discuss.elastic.co/t/increase-data-size-in-rally-with-existing-tracks/206890>\
**Category:** Elasticsearch\
**Tags:** rally\
**Created:** [November 7, 2019, 4:58am UTC](https://discuss.elastic.co/t/increase-data-size-in-rally-with-existing-tracks/206890 "2019-11-07T04:58:11Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![suvarna](https://avatars.discourse-cdn.com/v4/letter/s/51bf81/32.png) [@suvarna](https://discuss.elastic.co/u/suvarna)\
**Post date:** [November 7, 2019, 4:58am UTC](https://discuss.elastic.co/t/increase-data-size-in-rally-with-existing-tracks/206890/1 "2019-11-07T04:58:11Z")

</div>

We are using Rally as benchmark tool for our experiments .

In Rally we have one larger track which is "nyc\_taxis", and it will give around 27GB indexing data. but I would require larger dataset so, I have followed the steps from below link.

> [@Increase data size in Rally existing tracks](https://discuss.elastic.co/t/increase-data-size-in-rally-existing-tracks/116514):
>
> Hi Daniel , I am using Rally's existing tracks to perform benchmarking. I noticed that nyc taxis is the largest track with 4.5GB compressed and 74.3 GB uncompressed docs. I want to test with larger data volume. Is there any option provided in rally to duplicate or triplicate the data in existing tracks ?

We are using ES7.3.0 version and if use the custom track from above link we are getting following error:

org.elasticsearch.index.mapper.MapperParsingException: failed to parse  
at org.elasticsearch.index.mapper.DocumentParser.wrapInMapperParsingException(DocumentParser.java:191) ~[elasticsearch-7.3.0.jar:7.3.0]  
at org.elasticsearch.index.mapper.DocumentParser.parseDocument(DocumentParser.java:74) ~[elasticsearch-7.3.0.jar:7.3.0]  
at org.elasticsearch.index.mapper.DocumentMapper.parse(DocumentMapper.java:267) ~[elasticsearch-7.3.0.jar:7.3.0]  
at org.elasticsearch.index.shard.IndexShard.prepareIndex(IndexShard.java:772) ~[elasticsearch-7.3.0.jar:7.3.0]  
at org.elasticsearch.index.shard.IndexShard.applyIndexOperation(IndexShard.java:749) ~[elasticsearch-7.3.0.jar:7.3.0]  
at org.elasticsearch.index.shard.IndexShard.applyIndexOperationOnPrimary(IndexShard.java:721) ~[elasticsearch-7.3.0.jar:7.3.0]  
at org.elasticsearch.action.bulk.TransportShardBulkAction.executeBulkItemRequest(TransportShardBulkAction.java:256) [elasticsearch-7.3.0.jar:7.3.0]  
at org.elasticsearch.action.bulk.TransportShardBulkAction$2.doRun(TransportShardBulkAction.java:159) [elasticsearch-7.3.0.jar:7.3.0]  
at org.elasticsearch.common.util.concurrent.AbstractRunnable.run(AbstractRunnable.java:37) [elasticsearch-7.3.0.jar:7.3.0]  
at org.elasticsearch.action.bulk.TransportShardBulkAction.performOnPrimary(TransportShardBulkAction.java:191) [elasticsearch-7.3.0.jar:7.3.0]  
at org.elasticsearch.action.bulk.TransportShardBulkAction.shardOperationOnPrimary(TransportShardBulkAction.java:116) [elasticsearch-7.3.0.jar:7.3.0]  
at org.elasticsearch.action.bulk.TransportShardBulkAction.shardOperationOnPrimary(TransportShardBulkAction.java:77) [elasticsearch-7.3.0.jar:7.3.0]

Can you please provide your inputs on this.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 7, 2019, 6:17am UTC](https://discuss.elastic.co/t/increase-data-size-in-rally-with-existing-tracks/206890/2 "2019-11-07T06:17:19Z")

</div>

You could look into using the [rally-eventdata-track](https://github.com/elastic/rally-eventdata-track) which generates data on the fly instead of working with a fized size corpora. I have used this to continously generate many terabytes of data over long periods.

---

<div class="post-metadata">

**Author:** ![suvarna](https://avatars.discourse-cdn.com/v4/letter/s/51bf81/32.png) [@suvarna](https://discuss.elastic.co/u/suvarna)\
**Post date:** [November 7, 2019, 9:40am UTC](https://discuss.elastic.co/t/increase-data-size-in-rally-with-existing-tracks/206890/3 "2019-11-07T09:40:02Z")

</div>

Thanks for your input.

We already using eventdata-track to index around 5Billion json documents..

Along with eventdata track, we need nyc\_taxis track where we can index at least 300GB indexing data from nyc\_taxis. (to have different dataset).

Can you please help me , how we can resolve the issue with nyc\_taxis custom track.

---

<div class="post-metadata">

**Author:** ![danielmitterdorfer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielmitterdorfer/32/110510_2.png) [@danielmitterdorfer](https://discuss.elastic.co/u/danielmitterdorfer)\
**Post date:** [November 11, 2019, 8:12am UTC](https://discuss.elastic.co/t/increase-data-size-in-rally-with-existing-tracks/206890/4 "2019-11-11T08:12:41Z")

</div>

Hi,

> [@suvarna](#):
>
> Can you please help me , how we can resolve the issue with nyc\_taxis custom track.

it appears as if there is a problem creating the index because the mapping is incorrect (also the error message shown here seems incomplete?). As a first step you can run Rally with `--on-error=abort` which should give you an idea what's wrong with your track. If that does not help you find the problem, please share the complete track including your changes. Thanks.

Daniel

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 9, 2019, 8:22am UTC](https://discuss.elastic.co/t/increase-data-size-in-rally-with-existing-tracks/206890/5 "2019-12-09T08:22:29Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
