# Running benchmarking on existing index

**URL:** <https://discuss.elastic.co/t/running-benchmarking-on-existing-index/114322>\
**Category:** Elasticsearch\
**Tags:** rally\
**Created:** [January 5, 2018, 6:01pm UTC](https://discuss.elastic.co/t/running-benchmarking-on-existing-index/114322 "2018-01-05T18:01:09Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Alp1](https://avatars.discourse-cdn.com/v4/letter/a/7bcc69/32.png) [@Alp1](https://discuss.elastic.co/u/Alp1)\
**Post date:** [January 5, 2018, 6:01pm UTC](https://discuss.elastic.co/t/running-benchmarking-on-existing-index/114322/1 "2018-01-05T18:01:09Z")

</div>

Hi

We have few already created indexes(with data) on elastic and i want to get search performance on existing indexes.  
As i understand from docs, we can use "auto-managed":false to stop rally creating new index.

I have few questions :

1. Can we run search operation on existing indexes ?
2. Do we need to provide mapping.json in 'indices' section always even for existing indexes?
3. What information should be included/excluded in track.json for such scenario?

Any help is appreciated !!

---

<div class="post-metadata">

**Author:** ![danielmitterdorfer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielmitterdorfer/32/110510_2.png) [@danielmitterdorfer](https://discuss.elastic.co/u/danielmitterdorfer)\
**Post date:** [January 8, 2018, 7:06am UTC](https://discuss.elastic.co/t/running-benchmarking-on-existing-index/114322/2 "2018-01-08T07:06:18Z")

</div>

Hi @Alp1,

if you set `auto-managed` to `false`, there is no need for you to provide a mapping file. In fact, you can just remove the entire `indices` section and just define an index and a type explicitly in the search operation. Here is a minimal example from the [docs](http://esrally.readthedocs.io/en/stable/track.html#examples):

```auto
{
  "challenge": {
    "name": "just-search",
    "schedule": [
      {
        "operation": {
          "operation-type": "search",
          "index": "_all",
          "body": {
            "query": {
              "match_all": {}
            }
          }
        },
        "warmup-iterations": 100,
        "iterations": 100,
        "target-throughput": 10
      }
    ]
  }
}

```

It will run a `match_all` query against all indices of the target cluster with one client.

To be clear: This is not just a snippet, it is the entire content of your track. You can save this e.g. as `search.json` and run it with

```auto
esrally --pipeline=benchmark-only --target-hosts=node_1_ip:9200,node_2_ip:9200 --track-path=search.json

```

(this requires that you are on the latest stable version of Rally though which is 0.8.1 at the moment, check with `esrally --version`).

---

<div class="post-metadata">

**Author:** ![Alp1](https://avatars.discourse-cdn.com/v4/letter/a/7bcc69/32.png) [@Alp1](https://discuss.elastic.co/u/Alp1)\
**Post date:** [January 8, 2018, 4:04pm UTC](https://discuss.elastic.co/t/running-benchmarking-on-existing-index/114322/3 "2018-01-08T16:04:07Z")

</div>

Thanks 🙂 It worked.  
Can throughput be more than the number specified in track file ? I can see min/max/med throughput as 11 if i specify it as 10 and 105 if i specify it as 100.

---

<div class="post-metadata">

**Author:** ![danielmitterdorfer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielmitterdorfer/32/110510_2.png) [@danielmitterdorfer](https://discuss.elastic.co/u/danielmitterdorfer)\
**Post date:** [January 8, 2018, 4:53pm UTC](https://discuss.elastic.co/t/running-benchmarking-on-existing-index/114322/4 "2018-01-08T16:53:22Z")

</div>

If you specify a target throughput Rally should very closely match it. While it is possible that it is slightly above the target throughput, it should not exceed more than 1 op/s in my experience. Here is an example lower / upper bound from our benchmarking environment where we have executed an operation with a target throughput of 200 operations / s:

- min: 200.055
- median: 200.092
- max: 200.164

But seeing 105 instead of 100 is surprising.

Can you tell me the output of the following?

- `uname -a`
- `python3 -c "import sys ; print(sys.implementation)"`

Also:

- Are you seeing this all the time?
- Do you have a lot of processes running on that machine so you have a lot of scheduling pressure? E.g. you simulate hundreds of clients with Rally?

---

<div class="post-metadata">

**Author:** ![Alp1](https://avatars.discourse-cdn.com/v4/letter/a/7bcc69/32.png) [@Alp1](https://discuss.elastic.co/u/Alp1)\
**Post date:** [January 8, 2018, 5:48pm UTC](https://discuss.elastic.co/t/running-benchmarking-on-existing-index/114322/5 "2018-01-08T17:48:08Z")

</div>

> [@danielmitterdorfer](#):
>
> python3 -c "import sys ; print(sys.implementation)"

Hi @danielmitterdorfer

It seems to be an intermittent issue. I tried just now with throughput as 10 and 100 and got below results. I don't think i had much processes running on machine at that time.

| All | Min Throughput | search | 10.04 | ops/s |  
| All | Median Throughput | search | 10.06 | ops/s |  
| All | Max Throughput | search | 10.09 | ops/s |

| All | Min Throughput | search | 100.26 | ops/s |  
| All | Median Throughput | search | 100.26 | ops/s |  
| All | Max Throughput | search | 100.26 | ops/s |

* * *

Earlier results -

| All | Min Throughput | search | 11.04 | ops/s |  
| All | Median Throughput | search | 11.04 | ops/s |  
| All | Max Throughput | search | 11.04 | ops/s |

| All | Min Throughput | search | 105.87 | ops/s |  
| All | Median Throughput | search | 105.87 | ops/s |  
| All | Max Throughput | search | 105.87 | ops/s |

Here is the output of commands :

1. uname -a

Linux IP 3.13.0-135-generic #184-Ubuntu SMP Wed Oct 18 11:55:51 UTC 2017 x86\_64 x86\_64 x86\_64 GNU/Linux

1. python3 -c "import sys ; print(sys.implementation)"

namespace(\_multiarch='x86\_64-linux-gnu', cache\_tag='cpython-34', hexversion=50594800, name='cpython', version=sys.version\_info(major=3, minor=4, micro=3, releaselevel='final', serial=0))

---

<div class="post-metadata">

**Author:** ![danielmitterdorfer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielmitterdorfer/32/110510_2.png) [@danielmitterdorfer](https://discuss.elastic.co/u/danielmitterdorfer)\
**Post date:** [January 9, 2018, 8:10am UTC](https://discuss.elastic.co/t/running-benchmarking-on-existing-index/114322/6 "2018-01-09T08:10:28Z")

</div>

Thanks for the feedback! This might have something to do with the system's timer accuracy. I need to see whether there is a chance I can reproduce this. One last question: Is this a bare-metal machine, running in VM or in the cloud?

---

<div class="post-metadata">

**Author:** ![Alp1](https://avatars.discourse-cdn.com/v4/letter/a/7bcc69/32.png) [@Alp1](https://discuss.elastic.co/u/Alp1)\
**Post date:** [January 11, 2018, 2:14am UTC](https://discuss.elastic.co/t/running-benchmarking-on-existing-index/114322/7 "2018-01-11T02:14:53Z")

</div>

No..It's not bare -metal instance.

---

<div class="post-metadata">

**Author:** ![danielmitterdorfer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielmitterdorfer/32/110510_2.png) [@danielmitterdorfer](https://discuss.elastic.co/u/danielmitterdorfer)\
**Post date:** [January 11, 2018, 7:07am UTC](https://discuss.elastic.co/t/running-benchmarking-on-existing-index/114322/8 "2018-01-11T07:07:58Z")

</div>

So, is it then running in a VM? Or is it running in a cloud environment (if yes: in which one and ideally also the instance type)? This may help me to reproduce this. Thank you! 🙂

---

<div class="post-metadata">

**Author:** ![Alp1](https://avatars.discourse-cdn.com/v4/letter/a/7bcc69/32.png) [@Alp1](https://discuss.elastic.co/u/Alp1)\
**Post date:** [January 12, 2018, 7:44pm UTC](https://discuss.elastic.co/t/running-benchmarking-on-existing-index/114322/9 "2018-01-12T19:44:14Z")

</div>

It is running in cloud environment and the EC2 instance type is t2.large on ubuntu 14.04. I am using this instance solely for Rally. Let me know if you need more information 🙂

---

<div class="post-metadata">

**Author:** ![danielmitterdorfer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielmitterdorfer/32/110510_2.png) [@danielmitterdorfer](https://discuss.elastic.co/u/danielmitterdorfer)\
**Post date:** [January 15, 2018, 1:59pm UTC](https://discuss.elastic.co/t/running-benchmarking-on-existing-index/114322/10 "2018-01-15T13:59:37Z")

</div>

Thanks, that helps! I have raised [https://github.com/elastic/rally/issues/393](https://github.com/elastic/rally/issues/393) to track the progress of the respective analysis.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 12, 2018, 2:12pm UTC](https://discuss.elastic.co/t/running-benchmarking-on-existing-index/114322/11 "2018-02-12T14:12:04Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
