# Parallel Iterations of Bulk

**URL:** <https://discuss.elastic.co/t/parallel-iterations-of-bulk/202271>\
**Category:** Elasticsearch\
**Tags:** rally\
**Created:** [October 4, 2019, 3:11am UTC](https://discuss.elastic.co/t/parallel-iterations-of-bulk/202271 "2019-10-04T03:11:59Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![David\_Samokovlisky](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/david_samokovlisky/32/44993_2.png) [@David\_Samokovlisky](https://discuss.elastic.co/u/David_Samokovlisky)\
**Post date:** [October 4, 2019, 3:11am UTC](https://discuss.elastic.co/t/parallel-iterations-of-bulk/202271/1 "2019-10-04T03:11:59Z")

</div>

Hi!

I'm trying to have iterations to a parallel bulk operations, basically each operation is ~100K docs, and it's finishing too soon, I want to load it for a longer period of time  
So I tried iterations and also target-throughput and also time-period but no success

It just finished 1 iteration and that's it, doesn't continue again, over and over.  
how can I achieve that ?

```
{
  "version": 2,
  "title": "a",
  "description": "a",
  "indices": [
{
  "name": "a1",
  "body": "mapping.json"
},
		{
  "name": "a2",
  "body": "mapping.json"
},
		{
  "name": "a3 ",
  "body": "mapping.json"
},
		{
  "name": "a4",
  "body": "mapping.json"
},
		{
  "name": "a5",
  "body": "mapping.json"
},
		{
  "name": "a6",
  "body": "mapping.json"
},
		{
  "name": "a7",
  "body": "mapping.json"
},
		{
  "name": "a8",
  "body": "mapping.json"
},
		{
  "name": "a9",
  "body": "mapping.json"
}
  ],
  "corpora": [
{
  "name": "data1",
  "documents": [
    {
      "source-file": "data.json",
					"target-index": "a1",
      "document-count": 100117
    }
  ]
},
		{
  "name": "data2",
  "documents": [
    {
      "source-file": "data.json",
					"target-index": "a2",
      "document-count": 100117
    }
  ]
},{
  "name": "data3",
  "documents": [
    {
      "source-file": "data.json",
					"target-index": "a3",
      "document-count": 100117
    }
  ]
},{
  "name": "data4",
  "documents": [
    {
      "source-file": "data.json",
					"target-index": "a4",
      "document-count": 100117
    }
  ]
},{
  "name": "data5",
  "documents": [
    {
      "source-file": "data.json",
					"target-index": "a5",
      "document-count": 100117
    }
  ]
},{
  "name": "data6",
  "documents": [
    {
      "source-file": "data.json",
					"target-index": "a6",
      "document-count": 100117
    }
  ]
},{
  "name": "data7",
  "documents": [
    {
      "source-file": "data.json",
					"target-index": "a7",
      "document-count": 100117
    }
  ]
},{
  "name": "data8",
  "documents": [
    {
      "source-file": "data.json",
					"target-index": "a8",
      "document-count": 100117
    }
  ]
},{
  "name": "data9",
  "documents": [
    {
      "source-file": "data.json",
					"target-index": "a9",
      "document-count": 100117
    }
  ]
}
  ],
  "schedule": [
{
  "parallel": {
				"iterations": 10000000,
    "tasks": [
      {
        "name": "bulk1",
        "clients": 3,
						"target-throughput": 50,
        "operation": {
          "operation-type": "bulk",
          "corpora": "data1",
          "bulk-size": 100
        }
      },
      {
        "name": "bulk",
        "clients": 3,
        "operation": {
          "operation-type": "bulk",
          "corpora": "data2",
          "bulk-size": 100
        }
      },
      {
        "name": "bulk3",
        "clients": 3,
        "operation": {
          "operation-type": "bulk",
          "corpora": "data3",
          "bulk-size": 100
        }
      },
      {
        "name": "bulk4",
        "clients": 3,
        "operation": {
          "operation-type": "bulk",
          "corpora": "data4",
          "bulk-size": 100
        }
      },
      {
        "name": "bulk5",
        "clients": 3,
        "operation": {
          "operation-type": "bulk",
          "corpora": "data5",
          "bulk-size": 100
        }
      },
      {
        "name": "bulk6",
        "clients": 3,
        "operation": {
          "operation-type": "bulk",
          "corpora": "data6",
          "bulk-size": 100
        }
      },
      {
        "name": "bulk7",
        "clients": 3,
        "operation": {
          "operation-type": "bulk",
          "corpora": "data7",
          "bulk-size": 100
        }
      },
      {
        "name": "bulk8",
        "clients": 3,
        "operation": {
          "operation-type": "bulk",
          "corpora": "data8",
          "bulk-size": 100
        }
      },
      {
        "name": "bulk9",
        "clients": 3,
        "operation": {
          "operation-type": "bulk",
          "corpora": "data9",
          "bulk-size": 100
        }
      }
    ]
  }
}
  ]
}

```

Tried upgrading to 1.3.0 same deal, it basically does only 1 iteration, even though the data file include only raw data, without metadata id's (from my understanding esrally parses the data and add metadata header to each request with id)

did some internet search, and found similar cases, and you told that esrally doesn't support more than 1 iteration on bulk operations, and suggested people duplicate the corpora or make bigger files, hope there is a new way to handle this more elegant.

would love to get some help !  
thank you all very much !

---

<div class="post-metadata">

**Author:** ![dliappis](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dliappis/32/56174_2.png) [@dliappis](https://discuss.elastic.co/u/dliappis)\
**Post date:** [October 4, 2019, 9:05am UTC](https://discuss.elastic.co/t/parallel-iterations-of-bulk/202271/2 "2019-10-04T09:05:03Z")

</div>

Hello,

As described in the [docs](https://esrally.readthedocs.io/en/stable/track.html#time-based-vs-iteration-based), `iterations` inside the parallel element:

> `iterations` (optional, defaults to 1): Allows to define a default value for all tasks of the `parallel` element.

So this just propagates the `iterations` property to operations that support it. See the [docs](https://esrally.readthedocs.io/en/stable/track.html#bulk) and [this reply](https://discuss.elastic.co/t/bulk-index-operation-with-iterations-failed/150355/2).

If you want to rerun bulk operations you can use a jinja2 [for](https://jinja.palletsprojects.com/en/2.10.x/templates/#loop-controls) loop together with a [comma joiner](https://jinja.palletsprojects.com/en/2.10.x/templates/#joiner); here is an [example](https://github.com/elastic/rally-eventdata-track/blob/3c60b78af4823679aa553e85f23b012e0d7b9184/eventdata/challenges/daily-log-volume-index-and-query.json#L18-L23) from another track.

Rgs,  
Dimitris

---

<div class="post-metadata">

**Author:** ![David\_Samokovlisky](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/david_samokovlisky/32/44993_2.png) [@David\_Samokovlisky](https://discuss.elastic.co/u/David_Samokovlisky)\
**Post date:** [October 11, 2019, 7:47pm UTC](https://discuss.elastic.co/t/parallel-iterations-of-bulk/202271/3 "2019-10-11T19:47:21Z")

</div>

@dliappis Thank you very much ! I ended up generating bigger and bigger files.

A question for the jinja solution, when I duplicate the same file, Is there any lag\delay between each time it reads the same file again? because continuity is very important for my testings, This way I can save a lot of disc space

Also a question about `target-throughput`, when I use it for `bulk` operation, and looking in my use case where I have around 10 concurrent bulk operations, if I set for each operation a `target-throughput`, will it slow down the whole thing? because it's a log of timers\sleeps concurrently across 10+ operations and each operation with multiple clients

---

<div class="post-metadata">

**Author:** ![David\_Samokovlisky](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/david_samokovlisky/32/44993_2.png) [@David\_Samokovlisky](https://discuss.elastic.co/u/David_Samokovlisky)\
**Post date:** [October 14, 2019, 7:27pm UTC](https://discuss.elastic.co/t/parallel-iterations-of-bulk/202271/4 "2019-10-14T19:27:50Z")

</div>

@dliappis  
I have tried playing with the `target-throughput`, and it seem to be connected\related to `bulk-size`, I've tried setting the `target-throughput`to either `1` `1000` and not use it at all, and the `Median Throughput` stays the same, for a parallel 1 bulk operation, with just 1 client.

---

<div class="post-metadata">

**Author:** ![dliappis](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dliappis/32/56174_2.png) [@dliappis](https://discuss.elastic.co/u/dliappis)\
**Post date:** [October 15, 2019, 7:14am UTC](https://discuss.elastic.co/t/parallel-iterations-of-bulk/202271/5 "2019-10-15T07:14:34Z")

</div>

> [@David\_Samokovlisky](#):
>
> A question for the jinja solution, when I duplicate the same file, Is there any lag\delay between each time it reads the same file again? because continuity is very important for my testings, This way I can save a lot of disc space

When you use a Jinja2 for loop there will be the same (small) delay between tasks to the delay you see between executing explicitly defined operations in a regular track.

> [@David\_Samokovlisky](#):
>
> if I set for each operation a `target-throughput` , will it slow down the whole thing? because it's a log of timers\sleeps concurrently across 10+ operations and each operation with multiple clients

Tasks inside the `parallel` element can have their own independent `clients` and `target-throughput` and are independent. There is a large number of examples about the parallel element in [this part of the documentation](https://esrally.readthedocs.io/en/stable/track.html#running-tasks-in-parallel) that I suggest you take a look at.

> [@David\_Samokovlisky](#):
>
> I have tried playing with the `target-throughput` , and it seem to be connected\related to `bulk-size` , I've tried setting the `target-throughput` to either `1` `1000` and not use it at all, and the `Median Throughput` stays the same, for a parallel 1 bulk operation, with just 1 client.

This sounds normal. `target-throughput` is not a property that will "accelerate" the execution of bulk by automatically increasing clients; if you've specified 1 client, it will stick to that to achieve the specified `target-throughput` with 1 client. If `target-throughput` is smaller that what 1 client can achieve, Rally will pause the schedule as required to honor this, but it won't automatically increase clients to achieve larger throughputs. You need to scale the number of clients/bulk size yourself (and read up on [sizing Elasticsearch](https://discuss.elastic.co/t/elasticsearch-best-practice-sizing/150782/2)) if the target-throughput can't be achieved.

---

<div class="post-metadata">

**Author:** ![David\_Samokovlisky](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/david_samokovlisky/32/44993_2.png) [@David\_Samokovlisky](https://discuss.elastic.co/u/David_Samokovlisky)\
**Post date:** [October 15, 2019, 6:31pm UTC](https://discuss.elastic.co/t/parallel-iterations-of-bulk/202271/6 "2019-10-15T18:31:06Z")

</div>

regarding the jinja question, I didn't mean how jinja work, I ment when I generate a track that uses the same file over and over again in the same corpus, will it affect the performance in terms of:

```
{
  "documents": [
    {
      "source-file": "data.json",
      "target-index": "a3",
      "document-count": 100117
    },
    {
      "source-file": "data.json",
      "target-index": "a3",
      "document-count": 100117
    },
    {
      "source-file": "data.json",
      "target-index": "a3",
      "document-count": 100117
    },
    {
      "source-file": "data.json",
      "target-index": "a3",
      "document-count": 100117
    },
    {
      "source-file": "data.json",
      "target-index": "a3",
      "document-count": 100117
    },
    {
      "source-file": "data.json",
      "target-index": "a3",
      "document-count": 100117
    }
  ]
}

```

Will there be any lag\performance impact between each document it goes through ?

Second, regarding the `target-throughput` I was only able to achieve it working with search operations, regarding `bulk` no-success, because in my opinion it seams that `bulk-size` over-affect the throughtput, in some way, but still I wasn't able to achieve any results with the `target-throughput`, as I wasn't trying to accelerate result, on the contrary, I was trying to limit them to a lower number that meets the requirements of my test.  
I was trying to meet an exact number for example 50 doc\s, but even by setting the `target-throughput` to `1`, still was getting huge number of throughput of 600-900 docs\s in the test results.

Another question I had in mind, in parallel operations, is there a way to know when a task completed ? because if I have multiple ones, I want to know who finishes earlier, to be able to balance them (in terms of data files sizes - each index have different doc-size that I'm using in the source-files)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 12, 2019, 6:31pm UTC](https://discuss.elastic.co/t/parallel-iterations-of-bulk/202271/7 "2019-11-12T18:31:39Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
