# Index size increase dramatically

**URL:** <https://discuss.elastic.co/t/index-size-increase-dramatically/62512>\
**Category:** Elasticsearch\
**Created:** [October 7, 2016, 5:15pm UTC](https://discuss.elastic.co/t/index-size-increase-dramatically/62512 "2016-10-07T17:15:43Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![AdaYang](https://avatars.discourse-cdn.com/v4/letter/a/97f17d/32.png) [@AdaYang](https://discuss.elastic.co/u/AdaYang)\
**Post date:** [October 7, 2016, 5:15pm UTC](https://discuss.elastic.co/t/index-size-increase-dramatically/62512/1 "2016-10-07T17:15:43Z")

</div>

I am running a 3-nodes v2.4.0 elasticsearch cluster mainly indexing pdf files. I was using Base64.getEncoder().encodeToString(bytes) for the pdf content. This method is using a deprecated String method which cause some some tika exception; So I change to Base64.getEncoder().encode(bytes) but the index size increase dramatically that run of disk space on my ec2 instance.

Anybody seen this before? OS is Ubuntu 14.04.3 LTS (GNU/Linux 3.13.0-74-generic x86\_64; java version "1.8.0\_91" Java(TM) SE Runtime Environment (build 1.8.0\_91-b14),Java HotSpot(TM) 64-Bit Server VM (build 25.91-b14, mixed mode)

---

<div class="post-metadata">

**Author:** ![abeyad](https://avatars.discourse-cdn.com/v4/letter/a/278dde/32.png) [@abeyad](https://discuss.elastic.co/u/abeyad)\
**Post date:** [October 7, 2016, 7:24pm UTC](https://discuss.elastic.co/t/index-size-increase-dramatically/62512/2 "2016-10-07T19:24:00Z")

</div>

You might want to try encoding your documents outside of ES to see what the differences are between the two method calls. I presume these are meant to be unanalyzed fields if they are base64 encoded?

---

<div class="post-metadata">

**Author:** ![AdaYang](https://avatars.discourse-cdn.com/v4/letter/a/97f17d/32.png) [@AdaYang](https://discuss.elastic.co/u/AdaYang)\
**Post date:** [October 7, 2016, 7:29pm UTC](https://discuss.elastic.co/t/index-size-increase-dramatically/62512/3 "2016-10-07T19:29:35Z")

</div>

Thanks for you reply.

I am using [https://www.elastic.co/guide/en/elasticsearch/plugins/current/mapper-attachments.html](https://www.elastic.co/guide/en/elasticsearch/plugins/current/mapper-attachments.html) with the type "attachment". Not sure if they are analyzed or unanalyzed. But I do want to search the content of the PDF.

---

<div class="post-metadata">

**Author:** ![abeyad](https://avatars.discourse-cdn.com/v4/letter/a/278dde/32.png) [@abeyad](https://discuss.elastic.co/u/abeyad)\
**Post date:** [October 7, 2016, 7:41pm UTC](https://discuss.elastic.co/t/index-size-increase-dramatically/62512/4 "2016-10-07T19:41:15Z")

</div>

I see, you are using the mapper attachment plugin to set attachment content, but that means you have to provide the base64 encoding of the attachment. When you run those two different methods to get the base64 encoding, are there any differences? It sounds like a tika issue, not an ES one?

---

<div class="post-metadata">

**Author:** ![AdaYang](https://avatars.discourse-cdn.com/v4/letter/a/97f17d/32.png) [@AdaYang](https://discuss.elastic.co/u/AdaYang)\
**Post date:** [October 7, 2016, 7:50pm UTC](https://discuss.elastic.co/t/index-size-increase-dramatically/62512/5 "2016-10-07T19:50:36Z")

</div>

in the PDF there are some diagrams. if I use the encodetoString(bytes[] ), I got error like this:

bulk has failure: failure in bulk execution:  
[0]: index [tech\_pdf0], type [pdf], id [ptv855.pdf], message [MapperParsingException[Failed to extract [-1] characters of text for [null] : Unexpected RuntimeException from org.apache.tika.parser.pdf.PDFParser@4f3926fb]; nested: NotSerializableExceptionWrapper[tika\_exception: Unexpected RuntimeException from org.apache.tika.parser.pdf.PDFParser@4f3926fb]; nested: NotSerializableExceptionWrapper[runtime\_exception: java.io.IOException: Value is not an integer: 38567420115579128561790115]; nested: NotSerializableExceptionWrapper[i\_o\_exception: Value is not an integer: 38567420115579128561790115];]

in Base64.java:  
@SuppressWarnings("deprecation")  
public String encodeToString(byte[] src) {  
byte[] encoded = encode(src);  
**return new String(encoded, 0, 0, encoded.length); -- this [new String] is deprecated method**  
}

So I changed to encode(bytes[]) it was able to index without failure. but the size blows up about 3 - 4 times for my documents.

You might be right, not an elasticsearch issue, but tika issue...

---

<div class="post-metadata">

**Author:** ![AdaYang](https://avatars.discourse-cdn.com/v4/letter/a/97f17d/32.png) [@AdaYang](https://discuss.elastic.co/u/AdaYang)\
**Post date:** [October 7, 2016, 7:57pm UTC](https://discuss.elastic.co/t/index-size-increase-dramatically/62512/6 "2016-10-07T19:57:21Z")

</div>

The encoded content sizes are same:

```
		byte[] bytes = IOUtils.toByteArray(content);			
		byte[] encoded = Base64.getEncoder().encode(bytes);
		logger.info("encoded: " + encoded.length);
		logger.info("string size " + Base64.getEncoder().encodeToString(bytes).length());

```

encoded: 5900292  
string size 5900292

encoded: 5903112  
string size 5903112

---

<div class="post-metadata">

**Author:** ![abeyad](https://avatars.discourse-cdn.com/v4/letter/a/278dde/32.png) [@abeyad](https://discuss.elastic.co/u/abeyad)\
**Post date:** [October 7, 2016, 8:15pm UTC](https://discuss.elastic.co/t/index-size-increase-dramatically/62512/7 "2016-10-07T20:15:35Z")

</div>

what happens if you call `Base64.getEncoder().encode(bytes)`, then wrap that in a string yourself: `new String(encoded, StandardCharsets.ISO_8859_1)`: see the Javadocs [here](https://docs.oracle.com/javase/8/docs/api/java/util/Base64.Encoder.html#encodeToString-byte:A-)

---

<div class="post-metadata">

**Author:** ![AdaYang](https://avatars.discourse-cdn.com/v4/letter/a/97f17d/32.png) [@AdaYang](https://discuss.elastic.co/u/AdaYang)\
**Post date:** [October 7, 2016, 8:43pm UTC](https://discuss.elastic.co/t/index-size-increase-dramatically/62512/8 "2016-10-07T20:43:36Z")

</div>

Thought about that. But the doc said it's the same. I will try and update  
here when I am back! Appreciate your help.

---

<div class="post-metadata">

**Author:** ![AdaYang](https://avatars.discourse-cdn.com/v4/letter/a/97f17d/32.png) [@AdaYang](https://discuss.elastic.co/u/AdaYang)\
**Post date:** [October 7, 2016, 10:43pm UTC](https://discuss.elastic.co/t/index-size-increase-dramatically/62512/9 "2016-10-07T22:43:34Z")

</div>

if I use String(encoded, StandardCharsets.ISO\_8859\_1) I will see errors like this:

[0]: index [tech\_pdf0], type [pdf], id [AVehT8YU7ThCmg0aRx0r], message [MapperParsingException[Failed to extract [-1] characters of text for [null] : Unable to extract PDF content]; nested: NotSerializableExceptionWrapper[tika\_exception: Unable to extract PDF content]; nested: NotSerializableExceptionWrapper[i\_o\_exception: null]; nested: NotSerializableExceptionWrapper[data\_format\_exception: invalid block type];]

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:13pm UTC](https://discuss.elastic.co/t/index-size-increase-dramatically/62512/10 "2017-07-05T22:13:58Z")

</div>


