# Missing documents after a bulk index

**URL:** <https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530>\
**Category:** Elasticsearch\
**Created:** [May 2, 2012, 4:21am UTC](https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530 "2012-05-02T04:21:27Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [May 2, 2012, 4:21am UTC](https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530/1 "2012-05-02T04:21:27Z")

</div>

Just finished finished bulk indexing 36 million documents to a single  
node with 5 shards. However, there are only 30 million products in the  
index. The node stats are:

docs": {  
"count": 30287500,  
"deleted": 0  
},  
"indexing": {  
"index\_total": 38177500,  
"index\_time": "1.6d",  
"index\_time\_in\_millis": 146190895,  
"index\_current": 0,  
"delete\_total": 0,  
"delete\_time": "0s",  
"delete\_time\_in\_millis": 0,  
"delete\_current": 0  
}

Why the large discrepancy between the expected count, the doc count,  
and the index\_total?

--  
Ivan

---

<div class="post-metadata">

**Author:** ![sujoysett](https://avatars.discourse-cdn.com/v4/letter/s/2acd7d/32.png) [@sujoysett](https://discuss.elastic.co/u/sujoysett)\
**Post date:** [May 2, 2012, 8:01am UTC](https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530/2 "2012-05-02T08:01:05Z")

</div>

Don't know your exact scenario or application, but I have faced such  
problem twice before.

Once, when gathering documents from multiple sources into a single index,  
there was overlap of document ids, which led to overwriting some documents.  
Another time, some UTF-8 characters in the documents was causing failure of  
some requests of every bulk. Removing UTF-8 chars using regex helped.

On Wednesday, May 2, 2012 9:51:27 AM UTC+5:30, Ivan Brusic wrote:

> Just finished finished bulk indexing 36 million documents to a single  
> node with 5 shards. However, there are only 30 million products in the  
> index. The node stats are:
> 
> docs": {  
> "count": 30287500,  
> "deleted": 0  
> },  
> "indexing": {  
> "index\_total": 38177500,  
> "index\_time": "1.6d",  
> "index\_time\_in\_millis": 146190895,  
> "index\_current": 0,  
> "delete\_total": 0,  
> "delete\_time": "0s",  
> "delete\_time\_in\_millis": 0,  
> "delete\_current": 0  
> }
> 
> Why the large discrepancy between the expected count, the doc count,  
> and the index\_total?
> 
> --  
> Ivan

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [May 2, 2012, 5:17pm UTC](https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530/3 "2012-05-02T17:17:32Z")

</div>

The document creation code has been running for Lucene for years, so I  
do not think it is re-using doc ids (although Lucene does not have doc  
ids).

My bulk indexer borrows heavily from existing Elasticsearch code:  
[https://gist.github.com/2577955](https://gist.github.com/2577955)

The log statements on lines 51 (throttling) and 78 (which keeps tracks  
of duplicate ids) are never called, so it appears that all the  
documents should have been indexed.

--  
Ivan

On Wed, May 2, 2012 at 1:01 AM, Sujoy Sett [sujoysett@gmail.com](mailto:sujoysett@gmail.com) wrote:

> Don't know your exact scenario or application, but I have faced such problem  
> twice before.
> 
> Once, when gathering documents from multiple sources into a single index,  
> there was overlap of document ids, which led to overwriting some documents.  
> Another time, some UTF-8 characters in the documents was causing failure of  
> some requests of every bulk. Removing UTF-8 chars using regex helped.
> 
> On Wednesday, May 2, 2012 9:51:27 AM UTC+5:30, Ivan Brusic wrote:
> 
> > Just finished finished bulk indexing 36 million documents to a single  
> > node with 5 shards. However, there are only 30 million products in the  
> > index. The node stats are:
> > 
> > docs": {  
> > "count": 30287500,  
> > "deleted": 0  
> > },  
> > "indexing": {  
> > "index\_total": 38177500,  
> > "index\_time": "1.6d",  
> > "index\_time\_in\_millis": 146190895,  
> > "index\_current": 0,  
> > "delete\_total": 0,  
> > "delete\_time": "0s",  
> > "delete\_time\_in\_millis": 0,  
> > "delete\_current": 0  
> > }
> > 
> > Why the large discrepancy between the expected count, the doc count,  
> > and the index\_total?
> > 
> > --  
> > Ivan

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [May 2, 2012, 5:20pm UTC](https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530/4 "2012-05-02T17:20:36Z")

</div>

Is your class being called from multiple threads? Maybe thats the problem?  
Also, you can mark the IndexRequest with a flag of create, in which case it  
will fail if it already exists in the index.

On Wed, May 2, 2012 at 8:17 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:

> The document creation code has been running for Lucene for years, so I  
> do not think it is re-using doc ids (although Lucene does not have doc  
> ids).
> 
> My bulk indexer borrows heavily from existing Elasticsearch code:  
> [https://gist.github.com/2577955](https://gist.github.com/2577955)
> 
> The log statements on lines 51 (throttling) and 78 (which keeps tracks  
> of duplicate ids) are never called, so it appears that all the  
> documents should have been indexed.
> 
> --  
> Ivan
> 
> On Wed, May 2, 2012 at 1:01 AM, Sujoy Sett [sujoysett@gmail.com](mailto:sujoysett@gmail.com) wrote:
> 
> > Don't know your exact scenario or application, but I have faced such  
> > problem  
> > twice before.
> > 
> > Once, when gathering documents from multiple sources into a single index,  
> > there was overlap of document ids, which led to overwriting some  
> > documents.  
> > Another time, some UTF-8 characters in the documents was causing failure  
> > of  
> > some requests of every bulk. Removing UTF-8 chars using regex helped.
> > 
> > On Wednesday, May 2, 2012 9:51:27 AM UTC+5:30, Ivan Brusic wrote:
> > 
> > > Just finished finished bulk indexing 36 million documents to a single  
> > > node with 5 shards. However, there are only 30 million products in the  
> > > index. The node stats are:
> > > 
> > > docs": {  
> > > "count": 30287500,  
> > > "deleted": 0  
> > > },  
> > > "indexing": {  
> > > "index\_total": 38177500,  
> > > "index\_time": "1.6d",  
> > > "index\_time\_in\_millis": 146190895,  
> > > "index\_current": 0,  
> > > "delete\_total": 0,  
> > > "delete\_time": "0s",  
> > > "delete\_time\_in\_millis": 0,  
> > > "delete\_current": 0  
> > > }
> > > 
> > > Why the large discrepancy between the expected count, the doc count,  
> > > and the index\_total?
> > > 
> > > --  
> > > Ivan

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [May 2, 2012, 6:38pm UTC](https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530/5 "2012-05-02T18:38:37Z")

</div>

I perform multithreaded bulk writes with ~2000-3000 docs per second over  
~90 minutes. Because such data load will cause high I/O load every 20-30  
minutes or so, there is a need to balance things out by setting a maximum  
limit of simultaneous multithreaded requests for that peak situations.  
Threads should wait until the ES indexer responds to outstanding bulk  
requests. Maybe you can get some inspiration from my code here - I also  
started from existing Elasticsearch code: [https://gist.github.com/2578923](https://gist.github.com/2578923)

Jörg

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [May 2, 2012, 6:41pm UTC](https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530/6 "2012-05-02T18:41:30Z")

</div>

Yes, the code is multi-threaded. It is the same code that has creates  
Lucene documents successfully.

Part of the problem might be our expected counts. We are in  
development phase and are pointing to a non-active table, so our  
expected count might be slightly off. In fact, it might be closer to  
30M and not 36M. However, I am still confused about the discrepancy  
between docs.count and indexing.index\_total. Now I don't remember if I  
indexed test documents prior to the bulk index!

--  
Ivan

On Wed, May 2, 2012 at 10:20 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> Is your class being called from multiple threads? Maybe thats the problem?  
> Also, you can mark the IndexRequest with a flag of create, in which case it  
> will fail if it already exists in the index.
> 
> On Wed, May 2, 2012 at 8:17 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:
> 
> > The document creation code has been running for Lucene for years, so I  
> > do not think it is re-using doc ids (although Lucene does not have doc  
> > ids).
> > 
> > My bulk indexer borrows heavily from existing Elasticsearch code:  
> > [https://gist.github.com/2577955](https://gist.github.com/2577955)
> > 
> > The log statements on lines 51 (throttling) and 78 (which keeps tracks  
> > of duplicate ids) are never called, so it appears that all the  
> > documents should have been indexed.
> > 
> > --  
> > Ivan
> > 
> > On Wed, May 2, 2012 at 1:01 AM, Sujoy Sett [sujoysett@gmail.com](mailto:sujoysett@gmail.com) wrote:
> > 
> > > Don't know your exact scenario or application, but I have faced such  
> > > problem  
> > > twice before.
> > > 
> > > Once, when gathering documents from multiple sources into a single  
> > > index,  
> > > there was overlap of document ids, which led to overwriting some  
> > > documents.  
> > > Another time, some UTF-8 characters in the documents was causing failure  
> > > of  
> > > some requests of every bulk. Removing UTF-8 chars using regex helped.
> > > 
> > > On Wednesday, May 2, 2012 9:51:27 AM UTC+5:30, Ivan Brusic wrote:
> > > 
> > > > Just finished finished bulk indexing 36 million documents to a single  
> > > > node with 5 shards. However, there are only 30 million products in the  
> > > > index. The node stats are:
> > > > 
> > > > docs": {  
> > > > "count": 30287500,  
> > > > "deleted": 0  
> > > > },  
> > > > "indexing": {  
> > > > "index\_total": 38177500,  
> > > > "index\_time": "1.6d",  
> > > > "index\_time\_in\_millis": 146190895,  
> > > > "index\_current": 0,  
> > > > "delete\_total": 0,  
> > > > "delete\_time": "0s",  
> > > > "delete\_time\_in\_millis": 0,  
> > > > "delete\_current": 0  
> > > > }
> > > > 
> > > > Why the large discrepancy between the expected count, the doc count,  
> > > > and the index\_total?
> > > > 
> > > > --  
> > > > Ivan

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [May 2, 2012, 6:45pm UTC](https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530/7 "2012-05-02T18:45:23Z")

</div>

Jörg,

Your code is not waiting for a index request response since you are  
using an ActionListener. I switched to using execute().actionGet()  
precisely because I was swamping ES with too much data.

--  
Ivan

On Wed, May 2, 2012 at 11:38 AM, Jörg Prante [joergprante@gmail.com](mailto:joergprante@gmail.com) wrote:

> I perform multithreaded bulk writes with ~2000-3000 docs per second over ~90  
> minutes. Because such data load will cause high I/O load every 20-30 minutes  
> or so, there is a need to balance things out by setting a maximum limit of  
> simultaneous multithreaded requests for that peak situations. Threads should  
> wait until the ES indexer responds to outstanding bulk requests. Maybe you  
> can get some inspiration from my code here - I also started from existing  
> Elasticsearch code: [Performing multithreaded asynchronous bulk writes to Elasticsearch · GitHub](https://gist.github.com/2578923)
> 
> Jörg

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [May 7, 2012, 6:25am UTC](https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530/8 "2012-05-07T06:25:00Z")

</div>

Ivan, any update, did you manage to solve this? One more thing, adding  
requests to the bulk request from multiple threads is problematic since the  
list there is not thread safe. I really need to write a BulkProcessor that  
allows for multiple threads to add requests and allows for simple  
throttling control (can use that in several places myself).

On Wed, May 2, 2012 at 9:45 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:

> Jörg,
> 
> Your code is not waiting for a index request response since you are  
> using an ActionListener. I switched to using execute().actionGet()  
> precisely because I was swamping ES with too much data.
> 
> --  
> Ivan
> 
> On Wed, May 2, 2012 at 11:38 AM, Jörg Prante [joergprante@gmail.com](mailto:joergprante@gmail.com)  
> wrote:
> 
> > I perform multithreaded bulk writes with ~2000-3000 docs per second over  
> > ~90  
> > minutes. Because such data load will cause high I/O load every 20-30  
> > minutes  
> > or so, there is a need to balance things out by setting a maximum limit  
> > of  
> > simultaneous multithreaded requests for that peak situations. Threads  
> > should  
> > wait until the ES indexer responds to outstanding bulk requests. Maybe  
> > you  
> > can get some inspiration from my code here - I also started from existing  
> > Elasticsearch code: [Performing multithreaded asynchronous bulk writes to Elasticsearch · GitHub](https://gist.github.com/2578923)
> > 
> > Jörg

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [May 7, 2012, 8:55am UTC](https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530/9 "2012-05-07T08:55:22Z")

</div>

This is by intention and ensures asynchronous multithreaded indexing. I  
submit a number of requests in parallel by a lot of threads, and I do not  
wait for ES bulk response. This number of asynchronous submits is bound  
(e.g. 30 open bulk requests), therefore I use a lot of resources, but I do  
not swamp ES with data. The ActionListener is invoked later by other  
response threads.

If you switch to execute.actionGet(), you just choose to perform bulk in  
synchronous mode, or 1 open bulk request at a time.

Jörg

On Wednesday, May 2, 2012 8:45:23 PM UTC+2, Ivan Brusic wrote:

> Jörg,
> 
> Your code is not waiting for a index request response since you are  
> using an ActionListener. I switched to using execute().actionGet()  
> precisely because I was swamping ES with too much data.
> 
> --  
> Ivan
> 
> On Wed, May 2, 2012 at 11:38 AM, Jörg Prante [joergprante@gmail.com](mailto:joergprante@gmail.com)  
> wrote:
> 
> > I perform multithreaded bulk writes with ~2000-3000 docs per second over  
> > ~90  
> > minutes. Because such data load will cause high I/O load every 20-30  
> > minutes  
> > or so, there is a need to balance things out by setting a maximum limit  
> > of  
> > simultaneous multithreaded requests for that peak situations. Threads  
> > should  
> > wait until the ES indexer responds to outstanding bulk requests. Maybe  
> > you  
> > can get some inspiration from my code here - I also started from  
> > existing  
> > Elasticsearch code: [Performing multithreaded asynchronous bulk writes to Elasticsearch · GitHub](https://gist.github.com/2578923)
> > 
> > Jörg

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [May 7, 2012, 5:16pm UTC](https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530/10 "2012-05-07T17:16:24Z")

</div>

As mentioned before, there might have been an issue with our expected  
count, so the results might be correct. Still confused about the  
difference between the doc count,and the index\_total. Will probably  
attempt another full index rebuild later today or tomorrow.

Each writer thread in the system has its own bulk indexer, which  
solves the thread safe issue. Synchronous calls work well of us since  
it allows use to collect detailed metrics in one place.

--  
Ivan

On Sun, May 6, 2012 at 11:25 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> Ivan, any update, did you manage to solve this? One more thing, adding  
> requests to the bulk request from multiple threads is problematic since the  
> list there is not thread safe. I really need to write a BulkProcessor that  
> allows for multiple threads to add requests and allows for simple throttling  
> control (can use that in several places myself).
> 
> On Wed, May 2, 2012 at 9:45 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:
> 
> > Jörg,
> > 
> > Your code is not waiting for a index request response since you are  
> > using an ActionListener. I switched to using execute().actionGet()  
> > precisely because I was swamping ES with too much data.
> > 
> > --  
> > Ivan
> > 
> > On Wed, May 2, 2012 at 11:38 AM, Jörg Prante [joergprante@gmail.com](mailto:joergprante@gmail.com)  
> > wrote:
> > 
> > > I perform multithreaded bulk writes with ~2000-3000 docs per second over  
> > > ~90  
> > > minutes. Because such data load will cause high I/O load every 20-30  
> > > minutes  
> > > or so, there is a need to balance things out by setting a maximum limit  
> > > of  
> > > simultaneous multithreaded requests for that peak situations. Threads  
> > > should  
> > > wait until the ES indexer responds to outstanding bulk requests. Maybe  
> > > you  
> > > can get some inspiration from my code here - I also started from  
> > > existing  
> > > Elasticsearch code: [Performing multithreaded asynchronous bulk writes to Elasticsearch · GitHub](https://gist.github.com/2578923)
> > > 
> > > Jörg

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [May 8, 2012, 1:49pm UTC](https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530/11 "2012-05-08T13:49:22Z")

</div>

I see, so just I am clear Ivan, each thread has its own bulk instance that  
you posted?

On Mon, May 7, 2012 at 8:16 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:

> As mentioned before, there might have been an issue with our expected  
> count, so the results might be correct. Still confused about the  
> difference between the doc count,and the index\_total. Will probably  
> attempt another full index rebuild later today or tomorrow.
> 
> Each writer thread in the system has its own bulk indexer, which  
> solves the thread safe issue. Synchronous calls work well of us since  
> it allows use to collect detailed metrics in one place.
> 
> --  
> Ivan
> 
> On Sun, May 6, 2012 at 11:25 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> > Ivan, any update, did you manage to solve this? One more thing, adding  
> > requests to the bulk request from multiple threads is problematic since  
> > the  
> > list there is not thread safe. I really need to write a BulkProcessor  
> > that  
> > allows for multiple threads to add requests and allows for simple  
> > throttling  
> > control (can use that in several places myself).
> > 
> > On Wed, May 2, 2012 at 9:45 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:
> > 
> > > Jörg,
> > > 
> > > Your code is not waiting for a index request response since you are  
> > > using an ActionListener. I switched to using execute().actionGet()  
> > > precisely because I was swamping ES with too much data.
> > > 
> > > --  
> > > Ivan
> > > 
> > > On Wed, May 2, 2012 at 11:38 AM, Jörg Prante [joergprante@gmail.com](mailto:joergprante@gmail.com)  
> > > wrote:
> > > 
> > > > I perform multithreaded bulk writes with ~2000-3000 docs per second  
> > > > over  
> > > > ~90  
> > > > minutes. Because such data load will cause high I/O load every 20-30  
> > > > minutes  
> > > > or so, there is a need to balance things out by setting a maximum  
> > > > limit  
> > > > of  
> > > > simultaneous multithreaded requests for that peak situations. Threads  
> > > > should  
> > > > wait until the ES indexer responds to outstanding bulk requests. Maybe  
> > > > you  
> > > > can get some inspiration from my code here - I also started from  
> > > > existing  
> > > > Elasticsearch code: [Performing multithreaded asynchronous bulk writes to Elasticsearch · GitHub](https://gist.github.com/2578923)
> > > > 
> > > > Jörg

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [May 8, 2012, 5:11pm UTC](https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530/12 "2012-05-08T17:11:12Z")

</div>

Correct. I originally wrote that bulk indexer code when I was running  
a single-threaded indexer (river), so I made minimal changes to  
support our existing multi-threaded Lucene indexing.

This Elasticsearch project has been put on hold for a while, but is  
finally starting up again. Most of the code is proof-of-concept at  
this stage and will get firmed up as time goes by. Still need to build  
another complete index.

--  
Ivan

On Tue, May 8, 2012 at 6:49 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> I see, so just I am clear Ivan, each thread has its own bulk instance that  
> you posted?

---

<div class="post-metadata">

**Author:** ![Santiago\_Trias](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/santiago_trias/32/1354_2.png) [@Santiago\_Trias](https://discuss.elastic.co/u/Santiago_Trias)\
**Post date:** [August 9, 2014, 7:02am UTC](https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530/13 "2014-08-09T07:02:47Z")

</div>

I got into lost documents when trying to do Bulk requests on my local  
server.  
I was doing 1000 per request and I was loosing around 80% of the documents.  
Changing to 10 solved it.  
Any other solution to this? I have to load 11 million documents and even  
multi threading is kind of slow doing it 10 at a time.

Thanks.

El martes, 1 de mayo de 2012 21:21:27 UTC-7, Ivan Brusic escribió:

> Just finished finished bulk indexing 36 million documents to a single  
> node with 5 shards. However, there are only 30 million products in the  
> index. The node stats are:
> 
> docs": {  
> "count": 30287500,  
> "deleted": 0  
> },  
> "indexing": {  
> "index\_total": 38177500,  
> "index\_time": "1.6d",  
> "index\_time\_in\_millis": 146190895,  
> "index\_current": 0,  
> "delete\_total": 0,  
> "delete\_time": "0s",  
> "delete\_time\_in\_millis": 0,  
> "delete\_current": 0  
> }
> 
> Why the large discrepancy between the expected count, the doc count,  
> and the index\_total?
> 
> --  
> Ivan

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/43b9865c-1a4f-47db-8410-6827a3af405e%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/43b9865c-1a4f-47db-8410-6827a3af405e%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:09am UTC](https://discuss.elastic.co/t/missing-documents-after-a-bulk-index/7530/14 "2017-07-06T01:09:51Z")

</div>


