# Rally throughput counter include data generation

**URL:** <https://discuss.elastic.co/t/rally-throughput-counter-include-data-generation/276366>\
**Category:** Elasticsearch\
**Tags:** rally\
**Created:** [June 18, 2021, 9:43am UTC](https://discuss.elastic.co/t/rally-throughput-counter-include-data-generation/276366 "2021-06-18T09:43:48Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Lasse\_Nedergaard](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lasse_nedergaard/32/81108_2.png) [@Lasse\_Nedergaard](https://discuss.elastic.co/u/Lasse_Nedergaard)\
**Post date:** [June 18, 2021, 9:43am UTC](https://discuss.elastic.co/t/rally-throughput-counter-include-data-generation/276366/1 "2021-06-18T09:43:48Z")

</div>

We have some large JSON doc’s we use for testing. The manipulation of the document before ingest takes some time. I can see Rally’s finally output score include throughputs but the time is Rally’s throughput and it’s including buffer array generation time.  
Anyone knows how to get the es ingest rate metric as it isn’t include in node-stats

Thanks in advance  
Lasse Nedergaard

---

<div class="post-metadata">

**Author:** ![RickBoyd](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rickboyd/32/80022_2.png) [@RickBoyd](https://discuss.elastic.co/u/RickBoyd)\
**Post date:** [June 21, 2021, 3:06pm UTC](https://discuss.elastic.co/t/rally-throughput-counter-include-data-generation/276366/2 "2021-06-21T15:06:56Z")

</div>

Hi @Lasse_Nedergaard ,

As Rally is keeping track of the number of documents it is indexing, the throughput from the client side and from the server side (the Elasticsearch ingest rate) will be the same.

The indication from your question seems to be that you believe you have a client-side (or network) bottleneck. We don't typically concern ourselves too much with this (as in real world scenarios, composing bulk requests also takes _some_ amount of time) unless:

- The data generation code is in rough shape and needs some optimization OR
- The client (Rally) machine is not powerful enough to generate load at the desired rate

If you are using a persistent data store (which is recommended) you can explore results in `rally-metrics-*` where the `name` field is "latency" and the `task` field is "bulk" (or whatever you have named your `bulk` task) and look at the `meta.took` field and compare to the `value` field, as both are expressed in milliseconds, to see what the latency overhead of your client and network roughly are, in order to assess if you need to optimize your track code, or upgrade your client machine.

Please let us know if this helps  
Rick B

---

<div class="post-metadata">

**Author:** ![Lasse\_Nedergaard](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lasse_nedergaard/32/81108_2.png) [@Lasse\_Nedergaard](https://discuss.elastic.co/u/Lasse_Nedergaard)\
**Post date:** [June 21, 2021, 3:32pm UTC](https://discuss.elastic.co/t/rally-throughput-counter-include-data-generation/276366/3 "2021-06-21T15:32:34Z")

</div>

Hi Rick

Thanks for cleaning this out it make sense. I will give it a try.  
And you are right my rally client do not perform 100% so my problem is likely there.

Thanks for helping out

Lasse Nedergaard

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 19, 2021, 3:32pm UTC](https://discuss.elastic.co/t/rally-throughput-counter-include-data-generation/276366/4 "2021-07-19T15:32:42Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
