# Bulk indexing. Not all documents inserted, but no errors

**URL:** <https://discuss.elastic.co/t/bulk-indexing-not-all-documents-inserted-but-no-errors/24376>\
**Category:** Elasticsearch\
**Created:** [June 25, 2015, 8:13pm UTC](https://discuss.elastic.co/t/bulk-indexing-not-all-documents-inserted-but-no-errors/24376 "2015-06-25T20:13:48Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![javadevmtl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javadevmtl/32/45613_2.png) [@javadevmtl](https://discuss.elastic.co/u/javadevmtl)\
**Post date:** [June 25, 2015, 8:13pm UTC](https://discuss.elastic.co/t/bulk-indexing-not-all-documents-inserted-but-no-errors/24376/1 "2015-06-25T20:13:48Z")

</div>

Using E.S 1.6.0

Hi I'm running a bulk process with the following java/vertx code. I tried bulking 192,000 documents but only 20,000 got indexed. This used to be fine I have indexed over 1.3 billion documents.

I'm looking at Bulk Thread pool in Marvel...  
Bulk Thread Pool Count: 30 per node  
Bulk Thread Pool Reject: 0 per node  
Bulk Thread Pool Ops/sec: 3 per node  
Bulk Thread Pool Largest Count: 30 per node  
Bulk Thread Pool Queue Size: 0 per node

None of the logs report any throttling.

I have 20TB of disk storage of 13TB used (including replicas). So I'm way below the watermark for disk.

What else can I check?

Index settings:

```
{
"index" : {
"refresh_interval" : "30s",
"translog" : {
"flush_threshold_size" : "1000mb"
},
"number_of_shards" : "8",
"creation_date" : "1435262728426",
"analysis" : {
"analyzer" : {
"default" : {
"filter" : [
"icu_folding"
],
"type" : "custom",
"tokenizer" : "keyword"
}
}
},
"number_of_replicas" : "1",
"version" : {
"created" : "1060099"
},
"uuid" : "KtNOhb4qS6eFjM8BUxc9HA"
}
}

```

Java Bulk Code:

```
JsonObject body = message.body();
			JsonArray documents = body.getArray("documents");
	
			BulkRequestBuilder bulkRequest = client.prepareBulk();
			final Context ctx = getVertx().currentContext();
	
			for(int i = 0; i < documents.size(); i++)
			{	
				final JsonObject obj = documents.get(i);
				final JsonObject indexable = new JsonObject()
				
				.putString("action", "index")
				.putString("_index", obj.getString("index"))
				.putString("_type", obj.getString("type"))
				.putString("_id", obj.getString("id"))
				.putString("_route", obj.getString("routing"))
				.putObject("_source", obj);										
				
				final String index = getRequiredIndex(indexable, message);
				if (index == null) {
					return;
				}
		
				// type is optional
				String type = indexable.getString(CONST_TYPE);
				;
		
				JsonObject source = indexable.getObject(CONST_SOURCE);
				if (source == null) {
					sendError(message, CONST_SOURCE + " is required");
					return;
				}
		
				// id is optional
				String id = indexable.getString(CONST_ID);
				String route = indexable.getString(CONST_ROUTE);
	
				IndexRequestBuilder builder = client.prepareIndex(index, type, id).setSource(source.encode());
		        
				if(!route.isEmpty())
					builder.setRouting(route);				
				
				bulkRequest.add(builder);
			}
	
			bulkRequest.execute(new ActionListener<BulkResponse>(){
				@Override
				public void onResponse(BulkResponse resp) {
					
					
					message.reply(new JsonObject().putString("status", "Took: " + resp.getTookInMillis() + ", Indexed:" + documents.size() + "," + resp.getItems().length + ", Failed: " + resp.hasFailures()));
				}
	
				@Override
				public void onFailure(Throwable t) {
					ctx.runOnContext(new Handler<Void>() {
						@Override
						public void handle(Void event) {
							sendError(message,
									"Index error: " + t.getMessage(),
									new RuntimeException(t));
						}
					});
				}
			});

```

Basically I bulk a bunch of Vetx.io JsonObjects into an array and then finally bulk them to Elasticsearch.

bulkRequest.execute() does not return any error. resp.hasFailures is always false. And both my document.size matches resp.getItems.length.

---

<div class="post-metadata">

**Author:** ![javadevmtl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javadevmtl/32/45613_2.png) [@javadevmtl](https://discuss.elastic.co/u/javadevmtl)\
**Post date:** [June 26, 2015, 1:15pm UTC](https://discuss.elastic.co/t/bulk-indexing-not-all-documents-inserted-but-no-errors/24376/2 "2015-06-26T13:15:26Z")

</div>

I see the index disk size growing and shrinking while bulking as if there's merge activity, but the document count stays consistent. It indexes a few docs and then nada,

Again there's no error reported in logs or from the client side. Also I tried the same operation on existing index that has over 200 million documents. Only a couple of thousand got inserted...

Have I reached some kind of threshold and ES no longer accepting documents?

---

<div class="post-metadata">

**Author:** ![javadevmtl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javadevmtl/32/45613_2.png) [@javadevmtl](https://discuss.elastic.co/u/javadevmtl)\
**Post date:** [June 26, 2015, 3:24pm UTC](https://discuss.elastic.co/t/bulk-indexing-not-all-documents-inserted-but-no-errors/24376/3 "2015-06-26T15:24:37Z")

</div>

Some logs also

> **[Dropbox - Error](https://www.dropbox.com/sh/boauxml5j2zn7nm/AABavysHdnbyTQPHx6yBewOYa?dl=0)**
>
> Dropbox is a free service that lets you bring your photos, docs, and videos anywhere and share them easily. Never email yourself a file again!

I rebooted one node to mark a clean event. You will se there's no errors. Also added screen shot of marvel for last hour. You can see output of my job file I have done more then a few thousands records and it's based on the code posted above. hasFailures() returns false.

---

<div class="post-metadata">

**Author:** ![Harlin\_ES](https://avatars.discourse-cdn.com/v4/letter/h/977dab/32.png) [@Harlin\_ES](https://discuss.elastic.co/u/Harlin_ES)\
**Post date:** [June 26, 2015, 4:10pm UTC](https://discuss.elastic.co/t/bulk-indexing-not-all-documents-inserted-but-no-errors/24376/4 "2015-06-26T16:10:33Z")

</div>

What do you have your threadpool.bulk\_queue\_size set at? If its at the default still you might want to increase it.

---

<div class="post-metadata">

**Author:** ![javadevmtl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javadevmtl/32/45613_2.png) [@javadevmtl](https://discuss.elastic.co/u/javadevmtl)\
**Post date:** [June 26, 2015, 5:15pm UTC](https://discuss.elastic.co/t/bulk-indexing-not-all-documents-inserted-but-no-errors/24376/5 "2015-06-26T17:15:46Z")

</div>

Default.

See above. All the stats are there, pulled off Marvel 🙂  
I even posted a screen shot in dropbox.

30 per node and no rejections, no threads queued.

---

<div class="post-metadata">

**Author:** ![javadevmtl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javadevmtl/32/45613_2.png) [@javadevmtl](https://discuss.elastic.co/u/javadevmtl)\
**Post date:** [June 26, 2015, 5:48pm UTC](https://discuss.elastic.co/t/bulk-indexing-not-all-documents-inserted-but-no-errors/24376/6 "2015-06-26T17:48:02Z")

</div>

Even Marvel Index Request Rate on the main page is reporting 3000 inserts per second...

---

<div class="post-metadata">

**Author:** ![Harlin\_ES](https://avatars.discourse-cdn.com/v4/letter/h/977dab/32.png) [@Harlin\_ES](https://discuss.elastic.co/u/Harlin_ES)\
**Post date:** [June 26, 2015, 6:00pm UTC](https://discuss.elastic.co/t/bulk-indexing-not-all-documents-inserted-but-no-errors/24376/7 "2015-06-26T18:00:16Z")

</div>

Try increasing your threadpool.bulk\_queue\_size to 1000 (defaults to 50)

---

<div class="post-metadata">

**Author:** ![javadevmtl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javadevmtl/32/45613_2.png) [@javadevmtl](https://discuss.elastic.co/u/javadevmtl)\
**Post date:** [June 26, 2015, 6:04pm UTC](https://discuss.elastic.co/t/bulk-indexing-not-all-documents-inserted-but-no-errors/24376/8 "2015-06-26T18:04:51Z")

</div>

No difference.

Anyways nothing is being queued. At least Marvel is not reporting anything queued.

Would you be willing to get on a join.me session? I swear I'm going bonkers over this...

---

<div class="post-metadata">

**Author:** ![javadevmtl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javadevmtl/32/45613_2.png) [@javadevmtl](https://discuss.elastic.co/u/javadevmtl)\
**Post date:** [June 26, 2015, 6:29pm UTC](https://discuss.elastic.co/t/bulk-indexing-not-all-documents-inserted-but-no-errors/24376/9 "2015-06-26T18:29:47Z")

</div>

So I'm going through all the Marvel stats.

In INDICES STORE DELETED DOCUMENTS there's as many deletions as there is inserts from my bulk job...

Is ES evicting my documents? Have I reached some threshold?

---

<div class="post-metadata">

**Author:** ![Harlin\_ES](https://avatars.discourse-cdn.com/v4/letter/h/977dab/32.png) [@Harlin\_ES](https://discuss.elastic.co/u/Harlin_ES)\
**Post date:** [June 26, 2015, 9:19pm UTC](https://discuss.elastic.co/t/bulk-indexing-not-all-documents-inserted-but-no-errors/24376/10 "2015-06-26T21:19:31Z")

</div>

The only other thing I can think of is look closely at your IDs. If your document IDs are colliding in elasticsearch, it will count that as a deletion. Make sure all of your IDs are unique.

---

<div class="post-metadata">

**Author:** ![javadevmtl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javadevmtl/32/45613_2.png) [@javadevmtl](https://discuss.elastic.co/u/javadevmtl)\
**Post date:** [June 26, 2015, 9:53pm UTC](https://discuss.elastic.co/t/bulk-indexing-not-all-documents-inserted-but-no-errors/24376/11 "2015-06-26T21:53:57Z")

</div>

Yep. I swear I didn't change anything in my bulk logic, but who knows hehe! 😛

---

<div class="post-metadata">

**Author:** ![javadevmtl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javadevmtl/32/45613_2.png) [@javadevmtl](https://discuss.elastic.co/u/javadevmtl)\
**Post date:** [June 29, 2015, 1:47pm UTC](https://discuss.elastic.co/t/bulk-indexing-not-all-documents-inserted-but-no-errors/24376/12 "2015-06-29T13:47:07Z")

</div>

Yep! Nothing to see here. My stupidity. I had accidentally ticked something on my JMeter script that generates the data to cause it to recycle the ids per thread 😛

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:04am UTC](https://discuss.elastic.co/t/bulk-indexing-not-all-documents-inserted-but-no-errors/24376/13 "2017-07-06T00:04:46Z")

</div>


