# Advantages of base64 encoded content in ingest attachment plugin

**URL:** <https://discuss.elastic.co/t/advantages-of-base64-encoded-content-in-ingest-attachment-plugin/126452>\
**Category:** Elasticsearch\
**Created:** [April 2, 2018, 5:22pm UTC](https://discuss.elastic.co/t/advantages-of-base64-encoded-content-in-ingest-attachment-plugin/126452 "2018-04-02T17:22:17Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Hemanth\_Gowda](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hemanth_gowda/32/23886_2.png) [@Hemanth\_Gowda](https://discuss.elastic.co/u/Hemanth_Gowda)\
**Post date:** [April 2, 2018, 5:22pm UTC](https://discuss.elastic.co/t/advantages-of-base64-encoded-content-in-ingest-attachment-plugin/126452/1 "2018-04-02T17:22:17Z")

</div>

Hi,

I am using ingest attachment plugin to index my PDF files.In the query during ingestion the source field is expecting base64 content.Below are my questions.  
Why we need to pass base64 content as source to plugin.However plugin converts the base 64 to actual content. In that case we can directly index the actual content instead of converting to base 64 and pass it to ingest plugin.  
What are the advantages of encoding the source to base64?

---

<div class="post-metadata">

**Author:** ![shanec](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shanec/32/4004_2.png) [@shanec](https://discuss.elastic.co/u/shanec)\
**Post date:** [April 3, 2018, 2:10am UTC](https://discuss.elastic.co/t/advantages-of-base64-encoded-content-in-ingest-attachment-plugin/126452/2 "2018-04-03T02:10:36Z")

</div>

PDFs are basically just a big binary block of data, not text. Whenever you send arbitrary binary data to a server (e.g. Elasticsearch), you either have to encode the data in a way the server understands (e.g. Elasticsearch understands JSON, so getting the data into something that's JSON compatible) or teach the server how to accept arbitrary binary data (e.g. implement accepting multipart/form-data in the server). Here we've chosen base64 encoding. There are advantages like reducing complexity on the server by doing this vs accepting arbitrary binary uploads and the "cost" of doing this conversion isn't terribly high (4/3 the upload size vs the raw data and there are encoders/decoders for base64 in virtually every programming language). The plugin uses [Tika](https://tika.apache.org/) to actually extract the text out, and it's also perfectly fine to run Tika on the files in your own process and just send the text/metadata across to Elasticsearch instead of the binary data.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 3, 2018, 2:44am UTC](https://discuss.elastic.co/t/advantages-of-base64-encoded-content-in-ingest-attachment-plugin/126452/3 "2018-04-03T02:44:50Z")

</div>

> [@shanec](#):
>
> it's also perfectly fine to run Tika on the files in your own process and just send the text/metadata across to Elasticsearch instead of the binary data.

Have a look at FSCrawler project. It does that. Some code here: [https://github.com/dadoonet/fscrawler/tree/master/tika/src/main/java/fr/pilato/elasticsearch/crawler/fs/tika](https://github.com/dadoonet/fscrawler/tree/master/tika/src/main/java/fr/pilato/elasticsearch/crawler/fs/tika)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 1, 2018, 2:44am UTC](https://discuss.elastic.co/t/advantages-of-base64-encoded-content-in-ingest-attachment-plugin/126452/4 "2018-05-01T02:44:54Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
