# Indexing Performance vs Document Size

**URL:** <https://discuss.elastic.co/t/indexing-performance-vs-document-size/56805>\
**Category:** Elasticsearch\
**Created:** [July 30, 2016, 10:16pm UTC](https://discuss.elastic.co/t/indexing-performance-vs-document-size/56805 "2016-07-30T22:16:08Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![anjan](https://avatars.discourse-cdn.com/v4/letter/a/cc9497/32.png) [@anjan](https://discuss.elastic.co/u/anjan)\
**Post date:** [July 30, 2016, 10:16pm UTC](https://discuss.elastic.co/t/indexing-performance-vs-document-size/56805/1 "2016-07-30T22:16:08Z")

</div>

I see a big difference in indexing throughput between small documents and large documents. Is this expected and why under the below test conditions?

- Small documents are 1KB, large documents are 10KB to 30KB
- Observed throughput is 3MB/s vs 20MB/s (not able to beat 4000 documents/sec)
- Bulk size is 300 documents; There is no improvement in performance beyond this
- Refresh is disabled (-1)
- Not analyzing any field
- Index buffer and translog are sized appropriately
- Disk storage, no replication, 1 shard
- No big difference with auto-id

BTW, refresh still happens when index buffer is half full (index buffer must have a ping pong implementation).

I would like to understand what is the per document processing overhead (including per field) and where are the bottlenecks.

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [July 30, 2016, 10:37pm UTC](https://discuss.elastic.co/t/indexing-performance-vs-document-size/56805/2 "2016-07-30T22:37:05Z")

</div>

What ES version? What machine? What operating system? What network interface capacity? What Java VM version? How many clients are indexing?

Each field results in a Lucene index where you can search on. Also, enabled `_source` and `_all` contribute to the indexing. If you do not carefully specify your field mapping, you put more load on the document indexing than possibly required.

---

<div class="post-metadata">

**Author:** ![anjan](https://avatars.discourse-cdn.com/v4/letter/a/cc9497/32.png) [@anjan](https://discuss.elastic.co/u/anjan)\
**Post date:** [August 1, 2016, 5:49pm UTC](https://discuss.elastic.co/t/indexing-performance-vs-document-size/56805/3 "2016-08-01T17:49:02Z")

</div>

- ES: 2.3.2, Lucene 5.5.0
- OS: Ubuntu 14.04.1
- Java: 8u73  
java version "1.8.0\_73"  
Java(TM) SE Runtime Environment (build 1.8.0\_73-b02)  
Java HotSpot(TM) 64-Bit Server VM (build 25.73-b02, mixed mode)
- \_all is disabled
- \_source disabled
- Fields are explicitly mapped
- Same configs and hardware with only change being document size 1K vs 30K
- CPU is close to 100%
- The difference in throughput with size of documents is surprising (3MB/s vs 20MB/s)

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [August 1, 2016, 7:07pm UTC](https://discuss.elastic.co/t/indexing-performance-vs-document-size/56805/4 "2016-08-01T19:07:27Z")

</div>

Maybe the client is challenged? How do you generate the bulk input? What client language/tool? Is the client running on a separate machine?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:31pm UTC](https://discuss.elastic.co/t/indexing-performance-vs-document-size/56805/5 "2017-07-05T22:31:05Z")

</div>


