# How Write throughput is calculated in Rally

**URL:** <https://discuss.elastic.co/t/how-write-throughput-is-calculated-in-rally/188997>\
**Category:** Elasticsearch\
**Tags:** rally\
**Created:** [July 5, 2019, 3:52am UTC](https://discuss.elastic.co/t/how-write-throughput-is-calculated-in-rally/188997 "2019-07-05T03:52:36Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![suvarna](https://avatars.discourse-cdn.com/v4/letter/s/51bf81/32.png) [@suvarna](https://discuss.elastic.co/u/suvarna)\
**Post date:** [July 5, 2019, 3:52am UTC](https://discuss.elastic.co/t/how-write-throughput-is-calculated-in-rally/188997/1 "2019-07-05T03:52:36Z")

</div>

Hi

I have sending 1 Billion event data with below parameters..

Clients :60  
Bulk indexing=10000  
iteration = 1666  
_index.translog.durability_= request

i can see the indexing throughput as 368997.56 doc/sec.

Can you please let me know how the throughput is calculated and what is formula for this.

Thanks and Regards,

---

<div class="post-metadata">

**Author:** ![dliappis](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dliappis/32/56174_2.png) [@dliappis](https://discuss.elastic.co/u/dliappis)\
**Post date:** [July 5, 2019, 3:43pm UTC](https://discuss.elastic.co/t/how-write-throughput-is-calculated-in-rally/188997/2 "2019-07-05T15:43:41Z")

</div>

This is a complicated topic because Rally can operated in [a distributed way](https://esrally.readthedocs.io/en/stable/recipes.html#distributing-the-load-test-driver) and thus needs to consider samples from all load drivers before it calculates global metrics, such as the total throughput.

The algorithm can be found in the [ThroughputCalculator class](https://github.com/elastic/rally/blob/master/esrally/driver/driver.py#L827-L940).  
Perhaps the best way to explain this is from Rally's own unit tests.

If you check [this unit test](https://github.com/elastic/rally/blob/8fda86b85703e74b97b16b0623680409fd629a15/tests/driver/driver_test.py#L404-L414) there is an example assuming two drivers.

```auto
samples = [
    driver.Sample(0, 1470838595, 21, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 1, 1 / 9),
    driver.Sample(0, 1470838596, 22, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 2, 2 / 9),
    driver.Sample(0, 1470838597, 23, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 3, 3 / 9),
    driver.Sample(0, 1470838598, 24, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 4, 4 / 9),
    driver.Sample(0, 1470838599, 25, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 5, 5 / 9),
    driver.Sample(0, 1470838600, 26, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 6, 6 / 9),
    driver.Sample(1, 1470838598.5, 24.5, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 4.5, 7 / 9),
    driver.Sample(1, 1470838599.5, 25.5, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 5.5, 8 / 9),
    driver.Sample(1, 1470838600.5, 26.5, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 6.5, 9 / 9)
]

```

The [signature for Sample](https://github.com/elastic/rally/blob/8fda86b85703e74b97b16b0623680409fd629a15/esrally/driver/driver.py#L791-L792) explains each of those parameters:

```auto
def __init__ (self, client_id, absolute_time, relative_time, task, sample_type, request_meta_data, latency_ms, service_time_ms,
                 total_ops, total_ops_unit, time_period, percent_completed):

```

Rally will first [sorting the samples by absolute time](https://github.com/elastic/rally/blob/8fda86b85703e74b97b16b0623680409fd629a15/esrally/driver/driver.py#L901):

```auto
    driver.Sample(0, 1470838595, 21, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 1, 1 / 9),
    driver.Sample(0, 1470838596, 22, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 2, 2 / 9),
    driver.Sample(0, 1470838597, 23, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 3, 3 / 9),
    driver.Sample(0, 1470838598, 24, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 4, 4 / 9),
    driver.Sample(1, 1470838598.5, 24.5, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 4.5, 7 / 9),
    driver.Sample(0, 1470838599, 25, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 5, 5 / 9),
    driver.Sample(1, 1470838599.5, 25.5, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 5.5, 8 / 9),
    driver.Sample(0, 1470838600, 26, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 6, 6 / 9),
    driver.Sample(1, 1470838600.5, 26.5, op, metrics.SampleType.Normal, None, -1, -1, 5000, "docs", 6.5, 9 / 9)

```

For each of the first four timestamps the calculated throughput is [5000 docs/s](https://github.com/elastic/rally/blob/8fda86b85703e74b97b16b0623680409fd629a15/tests/driver/driver_test.py#L423-L426) because each sample took 1s.

Starting with timestamp `1470838599` though, our calculation involves:

`4*5000 (total docs of first four samples) + 5000 (timestamp 1470838598.5) + 5000 (timestamp 1470838599) / 5 = 6000 doc/s`

Similarly for timestamp `1470838600` two more samples got collected each referencing 5000 docs/s so the throughput at that point is:

`(30000 (total docs at`1470838599`) + 5000 + 5000) / 6 = 6666.666666666667 docs/s`

This process keeps going for all samples and finally the summary output [calculates the min/median/max](https://github.com/elastic/rally/blob/8fda86b85703e74b97b16b0623680409fd629a15/esrally/reporter.py#L237-L255) throughput out of this list for **normal samples** only. Samples collected during the warm up period are not included in the summary report calculation.

---

<div class="post-metadata">

**Author:** ![suvarna](https://avatars.discourse-cdn.com/v4/letter/s/51bf81/32.png) [@suvarna](https://discuss.elastic.co/u/suvarna)\
**Post date:** [July 9, 2019, 4:30am UTC](https://discuss.elastic.co/t/how-write-throughput-is-calculated-in-rally/188997/3 "2019-07-09T04:30:52Z")

</div>

Thanks a lot for info.

We are using one cluster + single node +10 indices + each index having 10 shards.

Can you please consider this configuration as an example and please explain how we can calculate the Indexing throughput ..

Because its hard to understand above code.

---

<div class="post-metadata">

**Author:** ![dliappis](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dliappis/32/56174_2.png) [@dliappis](https://discuss.elastic.co/u/dliappis)\
**Post date:** [July 9, 2019, 6:26am UTC](https://discuss.elastic.co/t/how-write-throughput-is-calculated-in-rally/188997/4 "2019-07-09T06:26:39Z")

</div>

Hello,

Please take a look at the definition of the [throughput](https://esrally.readthedocs.io/en/stable/metrics.html?highlight=target-throughput%20metrics#metric-keys) metric key in the Rally docs, for details on what it corresponds to your track.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 6, 2019, 6:26am UTC](https://discuss.elastic.co/t/how-write-throughput-is-calculated-in-rally/188997/5 "2019-08-06T06:26:39Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
