# What is the curl command to convert pdf into base64 format?

**URL:** <https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751>\
**Category:** Elasticsearch\
**Created:** [March 11, 2019, 11:59am UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751 "2019-03-11T11:59:00Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![pyerunka](https://avatars.discourse-cdn.com/v4/letter/p/d6d6ee/32.png) [@pyerunka](https://discuss.elastic.co/u/pyerunka)\
**Post date:** [March 11, 2019, 11:59am UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/1 "2019-03-11T11:59:00Z")

</div>

I want to convert my pdf to base64.  
i am using below code but it is giving me an error:

curl -XPOST "[http://localhost:9200/test/xmlfile?pretty=1](http://localhost:9200/test/xmlfile?pretty=1)" -d '  
{  
"attachment" : "' `base64 /path/filename | perl -pe 's/\n/\\n/g'` '"  
}'

---

<div class="post-metadata">

**Author:** ![xavierfacq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xavierfacq/32/8744_2.png) [@xavierfacq](https://discuss.elastic.co/u/xavierfacq)\
**Post date:** [March 11, 2019, 3:45pm UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/2 "2019-03-11T15:45:21Z")

</div>

Hi,

NOTE: Please say "Hi / Hello", "Thank you" in your message to optimize your chance of response.

What do you want to do exactly ? The code you are given seems to be a Linux command. You  
can run it into a Shell, but it's not a Json command, isn't it ?

bye  
Xavier

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 12, 2019, 11:05am UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/4 "2019-03-12T11:05:13Z")

</div>

You need to transform your file binary to BASE64 before sending it to elasticsearch.  
This is to be made before calling elasticsearch.

You can do that with

> **[Base64 Encode and Decode - Online](https://www.base64encode.org/)**
>
> Decode from Base64 or Encode to Base64 - Here, with our simple online tool.

Or you can do that using some linux commands like `base64`.  
Or by writing some code like here:

> <https://github.com/dadoonet/fscrawler/blob/361495ddd03e3c40067157c1b92a454ca6247878/tika/src/main/java/fr/pilato/elasticsearch/crawler/fs/tika/TikaDocParser.java#L204>

You can also have a look at FSCrawler project. It has an upload endpoint where you can directly upload your binary document to elasticsearch. See [https://fscrawler.readthedocs.io/en/latest/admin/fs/rest.html#uploading-a-binary-document](https://fscrawler.readthedocs.io/en/latest/admin/fs/rest.html#uploading-a-binary-document)

---

<div class="post-metadata">

**Author:** ![pyerunka](https://avatars.discourse-cdn.com/v4/letter/p/d6d6ee/32.png) [@pyerunka](https://discuss.elastic.co/u/pyerunka)\
**Post date:** [March 12, 2019, 11:47am UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/5 "2019-03-12T11:47:44Z")

</div>

Hi,

Thanks for your reply!!!!🙂  
I have written JavaScript code to transform .pdf file to BASE64. I am getting value for DATA field that needs to be passed. but can i pass more than one data to ES? So currently i am indexing only one pdf document. I want to index more than one pdf document. so how can i pass it using below code?

PUT my\_index/\_doc/my\_id?pipeline=attachment  
{  
"data": "e1xydGYxXGFuc2kNCkxvcmVtIGlwc3VtIGRvbG9yIHNpdCBhbWV0DQpccGFyIH0="  
}

Thanks,  
Priyanka

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 12, 2019, 12:12pm UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/6 "2019-03-12T12:12:49Z")

</div>

Why do you want to index documents together and not individually? Are they related?

To answer your question, you can define multiple attachment processors within the same pipeline.

---

<div class="post-metadata">

**Author:** ![pyerunka](https://avatars.discourse-cdn.com/v4/letter/p/d6d6ee/32.png) [@pyerunka](https://discuss.elastic.co/u/pyerunka)\
**Post date:** [March 13, 2019, 4:48am UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/7 "2019-03-13T04:48:41Z")

</div>

Hi,

Thanks for your reply!!!  
Yes, I want to indexed documents together. Because it is our business requirement.  
after creating pipeline and passing data value to new index,when you create index pattern and discover it, you can see one pdf file that is indexed. I want more pdf indexed record under one index pattern. And want to search it.

Thanks,  
Priyanka

---

<div class="post-metadata">

**Author:** ![pyerunka](https://avatars.discourse-cdn.com/v4/letter/p/d6d6ee/32.png) [@pyerunka](https://discuss.elastic.co/u/pyerunka)\
**Post date:** [March 13, 2019, 6:56am UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/8 "2019-03-13T06:56:36Z")

</div>

Hello @dadoonet,

Thanks for your help!!!  
As per reply i tried for multiple attachment processors within the same pipeline. It is indexing documents together. But when i create index pattern and discover it, it is giving me one single record even if i have indexed 3 documents in one pipeline. If i have indexed 3 documents, i will be getting 3 different records. Correct me if i am wrong.

Thanks,  
Priyanka

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 13, 2019, 6:36pm UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/9 "2019-03-13T18:36:16Z")

</div>

So you won't get back one document when you search but an array of documents? Meaning that the user will have to guess in which document the text has been found.

Is that what you really want?

---

<div class="post-metadata">

**Author:** ![pyerunka](https://avatars.discourse-cdn.com/v4/letter/p/d6d6ee/32.png) [@pyerunka](https://discuss.elastic.co/u/pyerunka)\
**Post date:** [March 14, 2019, 3:30am UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/10 "2019-03-14T03:30:26Z")

</div>

Hello,

Yes, like google search. if user searches any word from attachment file, then it should give in which document the text has been found.

Thanks,  
Priyanka

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 14, 2019, 8:59am UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/11 "2019-03-14T08:59:32Z")

</div>

This won't be possible if you index an array of attachments. You need to index attachments individually.

---

<div class="post-metadata">

**Author:** ![pyerunka](https://avatars.discourse-cdn.com/v4/letter/p/d6d6ee/32.png) [@pyerunka](https://discuss.elastic.co/u/pyerunka)\
**Post date:** [March 14, 2019, 10:45am UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/12 "2019-03-14T10:45:37Z")

</div>

Hi @dadoonet ,

Thanks for your reply!!!!

If I indexed attachments individually every time, I have to create a new index. I want all the indexed attachments in one index only. So that I can see all the documents as a separate record and search through it.

Thanks,  
Priyanka

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 14, 2019, 11:59am UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/13 "2019-03-14T11:59:48Z")

</div>

No. All documents will go to the same index.

---

<div class="post-metadata">

**Author:** ![pyerunka](https://avatars.discourse-cdn.com/v4/letter/p/d6d6ee/32.png) [@pyerunka](https://discuss.elastic.co/u/pyerunka)\
**Post date:** [March 15, 2019, 5:05am UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/14 "2019-03-15T05:05:13Z")

</div>

Hi @dadoonet,

Thanks for reply!!!  
Could you please suggest me how I can indexed multiple documents with same index as I cannot use multiple attachment processors?

Thanks,  
Priyanka

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 15, 2019, 7:11am UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/15 "2019-03-15T07:11:30Z")

</div>

Like this:

```
PUT my_index/_doc/1?pipeline=attachment
{
  "data": "BASE64-doc1"
}
PUT my_index/_doc/2?pipeline=attachment
{
  "data": "BASE64-doc2"
}
```

---

<div class="post-metadata">

**Author:** ![pyerunka](https://avatars.discourse-cdn.com/v4/letter/p/d6d6ee/32.png) [@pyerunka](https://discuss.elastic.co/u/pyerunka)\
**Post date:** [March 15, 2019, 8:40am UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/16 "2019-03-15T08:40:35Z")

</div>

Hi @dadoonet,

Thanks for your quick help!!!!! 🙂

This solves my problem.

Thanks,  
Priyanka

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 12, 2019, 8:40am UTC](https://discuss.elastic.co/t/what-is-the-curl-command-to-convert-pdf-into-base64-format/171751/17 "2019-04-12T08:40:38Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
