# Elasticsearch Logstash importing very large JSON files

**URL:** <https://discuss.elastic.co/t/elasticsearch-logstash-importing-very-large-json-files/37015>\
**Category:** Logstash\
**Created:** [December 11, 2015, 7:52pm UTC](https://discuss.elastic.co/t/elasticsearch-logstash-importing-very-large-json-files/37015 "2015-12-11T19:52:05Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![bgow](https://avatars.discourse-cdn.com/v4/letter/b/7993a0/32.png) [@bgow](https://discuss.elastic.co/u/bgow)\
**Post date:** [December 11, 2015, 7:52pm UTC](https://discuss.elastic.co/t/elasticsearch-logstash-importing-very-large-json-files/37015/1 "2015-12-11T19:52:05Z")

</div>

I'm trying to import some very large JSON files (up to 80GB / file) into Elasticsearch and have tried a couple different approaches but neither is giving me an efficient working solution:

1. Using the BULK API - I received heap memory errors (even with export ES\_HEAP\_SIZE=4g) so decided to write a bash script to break the JSON file up before sending the JSON information to Elasticsearch, see the marked solution [here](https://stackoverflow.com/questions/33535186/can-jq-perform-aggregation-across-files)for more info. This works but is extremely slow for files this large. To continue with this approach I'd need to improve the bash script or write a program to split the files more efficiently.

2. Using Logstash - I have also tried indexing/updating via Logstash. I set export LS\_HEAP\_SIZE=4g. It's working for my medium size files (~4GB) but not the larger (80GB) ones. When I try to send the largest files I get this error:

`HTTP content length exceeded 104857600 bytes`

My .conf is simply:

```
input {
stdin {
type => "stdin-type"
}
file {
	path => ["C:/path/file.json"]
	start_position => "beginning"
	}
}
filter {
	json{
		source => "message"
	}
}
output {
		elasticsearch {
			hosts => "localhost"
			index => "indexname"
			document_type=>"subject"
			document_id=>"%{id}"
			action=>"update"
	}
}

```

Prior to running this .conf for the large files I run a .conf which has every subjects id in it using the action=\>"index" command. For all files I've aggregated the documents under their appropriate id. For the largest files there can be thousands of documents under a given id.

From what I'm seeing online it seems unrealistic to try to import a single file around 80GB. Can I use the filters in Logstash to break the file up into smaller chunks prior to indexing/updating? If not, could you provide a suggestion for improving my bash script or a program to use for efficiently breaking up these files? One last note I'm using Windows.

---

<div class="post-metadata">

**Author:** ![CallumHolden](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/callumholden/32/16996_2.png) [@CallumHolden](https://discuss.elastic.co/u/CallumHolden)\
**Post date:** [April 3, 2017, 3:35pm UTC](https://discuss.elastic.co/t/elasticsearch-logstash-importing-very-large-json-files/37015/2 "2017-04-03T15:35:04Z")

</div>

Hi, just wondering if you found a solution to this problem? I am having the same issue so am hoping to get some light shed on it - cheers!

---

<div class="post-metadata">

**Author:** ![bgow](https://avatars.discourse-cdn.com/v4/letter/b/7993a0/32.png) [@bgow](https://discuss.elastic.co/u/bgow)\
**Post date:** [April 3, 2017, 5:05pm UTC](https://discuss.elastic.co/t/elasticsearch-logstash-importing-very-large-json-files/37015/3 "2017-04-03T17:05:15Z")

</div>

I ended up splitting my large JSON files into smaller files and using the BULK API. I haven't done much with Elastic recently so I'm not sure if there is a better solution now.

---

<div class="post-metadata">

**Author:** ![CallumHolden](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/callumholden/32/16996_2.png) [@CallumHolden](https://discuss.elastic.co/u/CallumHolden)\
**Post date:** [April 4, 2017, 8:29am UTC](https://discuss.elastic.co/t/elasticsearch-logstash-importing-very-large-json-files/37015/4 "2017-04-04T08:29:36Z")

</div>

Hi Brian, thanks for the quick response. Ah ok, yeah that's the solution we are currently working with - I am hoping to find a better solution however.

Cheers!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:27am UTC](https://discuss.elastic.co/t/elasticsearch-logstash-importing-very-large-json-files/37015/5 "2017-07-06T04:27:21Z")

</div>


