# Split source-file into many indicies

**URL:** <https://discuss.elastic.co/t/split-source-file-into-many-indicies/176175>\
**Category:** Elasticsearch\
**Tags:** rally\
**Created:** [April 10, 2019, 8:46am UTC](https://discuss.elastic.co/t/split-source-file-into-many-indicies/176175 "2019-04-10T08:46:20Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![vmasarik](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vmasarik/32/44015_2.png) [@vmasarik](https://discuss.elastic.co/u/vmasarik)\
**Post date:** [April 10, 2019, 8:46am UTC](https://discuss.elastic.co/t/split-source-file-into-many-indicies/176175/1 "2019-04-10T08:46:21Z")

</div>

Hi,  
I have a case where I would like to split one huge source-file document into many (hundreds) indices. Could Rally help me? I might know why Rally would not support this, but I would be happy to be wrong.

My dream solution would be following: somehow tell Rally to `create 300 indices` (Rally will generate the names) then `split document.json into all created indices`.

Does Rally support this option or do I have to manually create 300 indices and split the document into 300 parts and then manually assign each document, in the corpora, its target index?

Thanks

---

<div class="post-metadata">

**Author:** ![dliappis](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dliappis/32/56174_2.png) [@dliappis](https://discuss.elastic.co/u/dliappis)\
**Post date:** [April 10, 2019, 6:29pm UTC](https://discuss.elastic.co/t/split-source-file-into-many-indicies/176175/2 "2019-04-10T18:29:31Z")

</div>

Hello @vmasarik,

Rally can't do this automatically for you, however, there are ways you can achieve it.

The easiest approach is to create a customized `document.json` file with a separate [action and metadata](https://www.elastic.co/guide/en/elasticsearch/reference/current/docs-bulk.html#docs-bulk) line. When you reference the json file in your corpora section you explicitly need to inform Rally that it includes [action-metadata](https://esrally.readthedocs.io/en/stable/track.html?highlight=action-metadata#corpora) using `"includes-action-and-meta-data": true`.

I used the [example track](https://esrally.readthedocs.io/en/stable/adding_tracks.html#example-track) from the official docs and modified the `toJSON.py` script as shown in [this gist](https://gist.github.com/dliappis/3a28eca9f2a280b517f6ae2105295d79#file-tojson-py); you can see in the script the variables `INDEX_FIRST=0` and `INDEX_LAST=299` that are used later to create the necessary action-and-metadata lines.

Running it creates a documents.json file like:

```auto
$ head -5 documents.json
{"index": {"_index": "geonames-054"}}
{"geonameid": 2986043, "name": "Pic de Font Blanca", "latitude": 42.64991, "longitude": 1.53335, "country_code": "AD", "population": 0}
{"index": {"_index": "geonames-068"}}
{"geonameid": 2994701, "name": "Roc Mélé", "latitude": 42.58765, "longitude": 1.74028, "country_code": "AD", "population": 0}
{"index": {"_index": "geonames-037"}}

```

Then I modified the [track.json](https://esrally.readthedocs.io/en/stable/adding_tracks.html) example in the docs to use a jinja2 loop to create 300 indices like `geonames-000` ... `geonames-299`; the modified code is [in this gist](https://gist.github.com/dliappis/3a28eca9f2a280b517f6ae2105295d79#file-track-json).

Then I simply ran the commands listed [in this gist](https://gist.github.com/dliappis/3a28eca9f2a280b517f6ae2105295d79#file-commands-sh) i.e. create the `documents.json` with action and metadata lines and then executed Rally using:

`esrally --distribution-version=7.0.0 --track-path=$PWD`

which ended up creating the specified 300 indices and split the documents.json across them.

Another approach would be using a [custom parameter source](https://esrally.readthedocs.io/en/stable/adding_tracks.html?highlight=custom%20parameter%20sources#custom-parameter-sources). This is more complicated and you can see an example in the bulk custom parameter source of the [eventdata-track](https://github.com/elastic/rally-eventdata-track/blob/master/eventdata/parameter_sources/elasticlogs_bulk_source.py).

Regards,  
Dimitris

---

<div class="post-metadata">

**Author:** ![vmasarik](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vmasarik/32/44015_2.png) [@vmasarik](https://discuss.elastic.co/u/vmasarik)\
**Post date:** [April 11, 2019, 3:16pm UTC](https://discuss.elastic.co/t/split-source-file-into-many-indicies/176175/3 "2019-04-11T15:16:22Z")

</div>

@dliappis

Well, I never expected a response in such a detail. Huge thanks Dimitris!

One more thing before I try this out. Is Rally able to evaluate the jinja2 on its own? Or do I have to preprocess it myself before using it?

Thank you!

---

<div class="post-metadata">

**Author:** ![dliappis](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dliappis/32/56174_2.png) [@dliappis](https://discuss.elastic.co/u/dliappis)\
**Post date:** [April 11, 2019, 3:56pm UTC](https://discuss.elastic.co/t/split-source-file-into-many-indicies/176175/4 "2019-04-11T15:56:28Z")

</div>

> [@vmasarik](#):
>
> One more thing before I try this out. Is Rally able to evaluate the jinja2 on its own? Or do I have to preprocess it myself before using it?

Yes Rally is capable of translating jinja2 by itself; in fact we are using this feature in our official tracks too (to organize things better), e.g. in [https://github.com/elastic/rally-tracks/blob/master/geonames/track.json](https://github.com/elastic/rally-tracks/blob/master/geonames/track.json).

---

<div class="post-metadata">

**Author:** ![vmasarik](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vmasarik/32/44015_2.png) [@vmasarik](https://discuss.elastic.co/u/vmasarik)\
**Post date:** [April 12, 2019, 4:37pm UTC](https://discuss.elastic.co/t/split-source-file-into-many-indicies/176175/5 "2019-04-12T16:37:42Z")

</div>

@dliappis

I tried to adapt your example to my situation. However, that did not work. As, after successful completion Rally announces:  
`error rate | bulk | 100 | % |`  
Which results into 300 empty indices.

After trying many different configurations and combinations of them I tried to copy pasta your example and that still resulted into 100% error rate. Which makes me think it is a versioning issue but I don't know how to deal with it. Mainly because I am not sure how to debug this. Any tips would be appreciated 🙂

I did not mention my environment as I never thought it would be that crucial. So, I am testing a remote cluster, which has version 5.6.13 ElasticSearch. I do not know what else might be important though.

Command:

```auto
esrally --track-path=$HOME --target-hosts=elasticsearch --pipeline=benchmark-only

```

Data that I use:

```auto
$ head docs.json
{"index": {"_index": "nasatre", "_type": "docs", "_id": "1"}}
{"ip": "199.72.81.55", "date": "[01/Jul/1995:00:00:01 -0400]", "request": "GET /history/apollo/ HTTP/1.0", "result": 200, "size": 6245}
{"index": {"_index": "nasaone", "_type": "docs", "_id": "2"}}
{"ip": "unicomp6.unicomp.net", "date": "[01/Jul/1995:00:00:06 -0400]", "request": "GET /shuttle/countdown/ HTTP/1.0", "result": 200, "size": 3985}
{"index": {"_index": "nasatre", "_type": "docs", "_id": "3"}}
{"ip": "199.120.110.21", "date": "[01/Jul/1995:00:00:09 -0400]", "request": "GET /shuttle/missions/sts-73/mission-sts-73.html HTTP/1.0", "result": 200, "size": 4085}

```

The tack.json:

```auto
{
  "version": 2,
  "description": "Desc of a track.",
  "indices": [
      {
        "name": "nasaone",
        "body": "index.json",
        "types": ["docs"]
      },
      {
        "name": "nasatwo",
        "body": "index.json",
        "types": ["docs"]
      },
      {
        "name": "nasatre",
        "body": "index.json",
        "types": ["docs"]
      }
  ],
  "corpora": [
    {
      "name": "nasa",
      "documents": [
        {
          "source-file": "docs.json",
          "includes-action-and-meta-data": true,
          "document-count": 1050
        }
      ]
    }
  ],
  "schedule": [
    {
      "operation": {
        "operation-type": "delete-index"
      }
    },
    {
      "operation": {
        "operation-type": "create-index"
      }
    },
    {
      "operation": {
        "operation-type": "cluster-health",
        "request-params": {
          "wait_for_status": "green"
        }
      }
    },
    {
      "operation": {
        "operation-type": "bulk",
        "bulk-size": 5000
      }
    }
  ]
}

```

---

<div class="post-metadata">

**Author:** ![danielmitterdorfer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielmitterdorfer/32/110510_2.png) [@danielmitterdorfer](https://discuss.elastic.co/u/danielmitterdorfer)\
**Post date:** [April 14, 2019, 7:03pm UTC](https://discuss.elastic.co/t/split-source-file-into-many-indicies/176175/6 "2019-04-14T19:03:14Z")

</div>

Hi,

you can add the command line parameter `--on-error=abort` when starting your benchmark. Then Rally will abort on the first erroneous request with an error message what went wrong. This should hopefully help you to diagnose the problem.

Daniel

---

<div class="post-metadata">

**Author:** ![vmasarik](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vmasarik/32/44015_2.png) [@vmasarik](https://discuss.elastic.co/u/vmasarik)\
**Post date:** [April 18, 2019, 12:34pm UTC](https://discuss.elastic.co/t/split-source-file-into-many-indicies/176175/7 "2019-04-18T12:34:10Z")

</div>

@danielmitterdorfer  
Thanks a lot, I was able to solve the problem using that parameter.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 16, 2019, 12:34pm UTC](https://discuss.elastic.co/t/split-source-file-into-many-indicies/176175/8 "2019-05-16T12:34:39Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
