# Large zip content

**URL:** <https://discuss.elastic.co/t/large-zip-content/113143>\
**Category:** Elasticsearch\
**Created:** [December 25, 2017, 11:03am UTC](https://discuss.elastic.co/t/large-zip-content/113143 "2017-12-25T11:03:06Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Chelambarasan\_CR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chelambarasan_cr/32/62592_2.png) [@Chelambarasan\_CR](https://discuss.elastic.co/u/Chelambarasan_CR)\
**Post date:** [December 25, 2017, 11:03am UTC](https://discuss.elastic.co/t/large-zip-content/113143/1 "2017-12-25T11:03:06Z")

</div>

Large zip file 2gb content read by tika was not indexing. Indexing is hanging at bulk request indexing. Smaller zip has no issues.

Ingest attachment is used for indexing.

Is there any limit in ES per document?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 25, 2017, 1:28pm UTC](https://discuss.elastic.co/t/large-zip-content/113143/2 "2017-12-25T13:28:11Z")

</div>

There I memory limit of the JVM, http network limit (100mb IIRC), ...

Indexing too big binary documents in elasticsearch is not a good idea IMHO.

You should do the extraction of metadata and text outside elasticsearch. FSCrawler project could help.

---

<div class="post-metadata">

**Author:** ![Chelambarasan\_CR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chelambarasan_cr/32/62592_2.png) [@Chelambarasan\_CR](https://discuss.elastic.co/u/Chelambarasan_CR)\
**Post date:** [December 26, 2017, 3:03am UTC](https://discuss.elastic.co/t/large-zip-content/113143/3 "2017-12-26T03:03:33Z")

</div>

Hi @dadoonet ,

Can I use fscrawler in my indexing java application?

Fetching content is not a problem but committing to elastic via bulk request is hanging for large zip files.

I have some more data of same record coming from dB. Content field comes from extracting zip file and reading each file with tika.

Please share ur thoughts.

---

<div class="post-metadata">

**Author:** ![Chelambarasan\_CR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chelambarasan_cr/32/62592_2.png) [@Chelambarasan\_CR](https://discuss.elastic.co/u/Chelambarasan_CR)\
**Post date:** [December 26, 2017, 3:07am UTC](https://discuss.elastic.co/t/large-zip-content/113143/4 "2017-12-26T03:07:04Z")

</div>

A separate 64gb server is allocated. Is there a way of delta commit to same document ? Say 100 Mb at a time and doing it 20 times

---

<div class="post-metadata">

**Author:** ![Chelambarasan\_CR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chelambarasan_cr/32/62592_2.png) [@Chelambarasan\_CR](https://discuss.elastic.co/u/Chelambarasan_CR)\
**Post date:** [December 26, 2017, 1:17pm UTC](https://discuss.elastic.co/t/large-zip-content/113143/5 "2017-12-26T13:17:22Z")

</div>

The error message is

nested: EsRejectedExecutionException[rejected execution of org.elasticsearch.transport.TransportService$7@7b221dce on EsThreadPoolExecutor[bulk, queue capacity = 200, org.elasticsearch.common.util.concurrent.EsThreadPoolExecutor@49b77e41[Running, pool size = 12, active threads = 12, queued tasks = 5792, completed tasks = 455479]]];

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 23, 2018, 1:17pm UTC](https://discuss.elastic.co/t/large-zip-content/113143/6 "2018-01-23T13:17:34Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
