# Is there a way to feed base64 encoded string to the ingest\_attachment plugin or fscrawler?

**URL:** <https://discuss.elastic.co/t/is-there-a-way-to-feed-base64-encoded-string-to-the-ingest-attachment-plugin-or-fscrawler/278016>\
**Category:** Elasticsearch\
**Created:** [July 7, 2021, 5:00am UTC](https://discuss.elastic.co/t/is-there-a-way-to-feed-base64-encoded-string-to-the-ingest-attachment-plugin-or-fscrawler/278016 "2021-07-07T05:00:55Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![dotel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dotel/32/77606_2.png) [@dotel](https://discuss.elastic.co/u/dotel)\
**Post date:** [July 7, 2021, 5:00am UTC](https://discuss.elastic.co/t/is-there-a-way-to-feed-base64-encoded-string-to-the-ingest-attachment-plugin-or-fscrawler/278016/1 "2021-07-07T05:00:55Z")

</div>

So I've been trying to come up with a poc for document content search.  
What I've found so far is, I can feed the documents to fscrawler and search for the content using something like

```auto
GET bookstore/_search
{
  "query": {
    "match_phrase": {
      "content": "coursera"
    }
  }
}

```

But my documents are actually base64 encoded strings and I didn't find a way to feed that into fscrawler. Is that possible? Seems like ingest attachment plugin is another thing I might look into(if I don't want image ocr and other cool fscrawler features later), but I couldn't find a way to use ingest attachment plugin. Of course, I looked the docs and tried doing something like this.

```auto
PUT _ingest/pipeline/attachment
{
  "description" : "Extract attachment information",
  "processors" : [
    {
      "attachment" : {
        "field" : "data"
      }
    }
  ]
}

PUT my-index-000001/_doc/bookstore?pipeline=attachment
{
  "data": "e1xydGYxXGFuc2kNCkxvcmVtIGlwc3VtIGRvbG9yIHNpdCBhbWV0DQpccGFyIH0="
}

PUT my-index-000001/_doc/bookstore?pipeline=attachment
{
  "data": "aGV5IHRoaXMgaXMgY29vbA=="
}

GET my-index-000001/_doc/bookstore

```

But this doesn't work. I don't much about elasticsearch but was just doing a poc to check if it can be done or not. What's happening in above request is the second put command is replacing the initial one. How can feed multiple documents that are base64 encoded to ingest attachment plugin and search the content?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 7, 2021, 5:08am UTC](https://discuss.elastic.co/t/is-there-a-way-to-feed-base64-encoded-string-to-the-ingest-attachment-plugin-or-fscrawler/278016/2 "2021-07-07T05:08:42Z")

</div>

Try removing bookstore from the URLs. I wonder if this with the removal of document types id interpreted as document id which causes an overwrite.

---

<div class="post-metadata">

**Author:** ![dotel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dotel/32/77606_2.png) [@dotel](https://discuss.elastic.co/u/dotel)\
**Post date:** [July 7, 2021, 5:28am UTC](https://discuss.elastic.co/t/is-there-a-way-to-feed-base64-encoded-string-to-the-ingest-attachment-plugin-or-fscrawler/278016/3 "2021-07-07T05:28:34Z")

</div>

That doesn't work either, unfortunately. Any other guesses? Why is it so hard to find an working example of searching through multiple documents using ingest attachment preprocessor plugin! I mean, isn't that one of the most common use case of this plugin.. ☹

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 7, 2021, 5:34am UTC](https://discuss.elastic.co/t/is-there-a-way-to-feed-base64-encoded-string-to-the-ingest-attachment-plugin-or-fscrawler/278016/4 "2021-07-07T05:34:25Z")

</div>

Try this:

```auto
PUT my-index-000001/_doc/1?pipeline=attachment
{
  "data": "e1xydGYxXGFuc2kNCkxvcmVtIGlwc3VtIGRvbG9yIHNpdCBhbWV0DQpccGFyIH0="
}

PUT my-index-000001/_doc/2?pipeline=attachment
{
  "data": "aGV5IHRoaXMgaXMgY29vbA=="
}

GET my-index-000001/_doc/1

GET my-index-000001/_doc/2

```

---

<div class="post-metadata">

**Author:** ![dotel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dotel/32/77606_2.png) [@dotel](https://discuss.elastic.co/u/dotel)\
**Post date:** [July 7, 2021, 5:39am UTC](https://discuss.elastic.co/t/is-there-a-way-to-feed-base64-encoded-string-to-the-ingest-attachment-plugin-or-fscrawler/278016/5 "2021-07-07T05:39:17Z")

</div>

I think you didn't get what I wanted to do. What I want to do is search amongst multiple documents. The queries you gave are just adding something to different indexes and retrieving them. Or if I misunderstood, can you please explain.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 7, 2021, 6:47am UTC](https://discuss.elastic.co/t/is-there-a-way-to-feed-base64-encoded-string-to-the-ingest-attachment-plugin-or-fscrawler/278016/6 "2021-07-07T06:47:17Z")

</div>

What you did was to index and then update a document with the document ID set to`bookstore`. My example showed how to index 2 different documents and then retrieve them separately so you can inspect them and the result of the ingest pipeline. Once you have ingested your documents and they do not overwrite eachother you can look at the indexed documents and start writing queries.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 7, 2021, 10:34am UTC](https://discuss.elastic.co/t/is-there-a-way-to-feed-base64-encoded-string-to-the-ingest-attachment-plugin-or-fscrawler/278016/7 "2021-07-07T10:34:27Z")

</div>

I don't understand everything. Could you clarify a bit?

> But my documents are actually base64 encoded strings

Do you mean that you have files on disk which contains a BASE64 text content? Or do you mean something else?

I don't understand what you are trying to do actually.

---

<div class="post-metadata">

**Author:** ![dotel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dotel/32/77606_2.png) [@dotel](https://discuss.elastic.co/u/dotel)\
**Post date:** [July 7, 2021, 12:43pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-feed-base64-encoded-string-to-the-ingest-attachment-plugin-or-fscrawler/278016/8 "2021-07-07T12:43:19Z")

</div>

The documents are obtained from an API as a base64 string and sent to amazon s3. I actually learned how to use fscrawler from your video on youtube. But on searching a lot, there's no functionality for indexing something that's in s3 as of now. So I thought I should index the documents in elastic search before sending it to s3(while it's in base64 format). Basically, I want lots of letters(obtained from an API as base64 string but are always pdf/docs) to be indexed by fscrawler or ingest attachment plugins and search keywords through them.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 7, 2021, 2:52pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-feed-base64-encoded-string-to-the-ingest-attachment-plugin-or-fscrawler/278016/9 "2021-07-07T14:52:51Z")

</div>

> [@dotel](#):
>
> But on searching a lot, there's no functionality for indexing something that's in s3 as of now.

True. That's something I'd like to have at some point after I do a big refactoring of the project for the version 3.x.

> <https://github.com/dadoonet/fscrawler/issues/263>
>
> I have documents stored in S3 - I need to use ElasticSearch for indexing and als…o for searching within the contents of the files stored in S3. I need a NodeJS example as well. Please help!!
> 
> Also if the s3 bucket contains files that cannot be content search it should skip it and search only the title - all the files need to be indexed in elasticsearch but the search on the content of the files need to return the results for content/filename

> [@dotel](#):
>
> Basically, I want lots of letters(obtained from an API as base64 string but are always pdf/docs) to be indexed by fscrawler or ingest attachment plugins and search keywords through them.

If I understand correctly the use case, you want to store files on S3 while being able to search for them.  
As you have very standard files, like pdfs and docs, I'd probably use the ingest attachment plugin.

In my code, I'd:

- Store the BASE64 content in S3
- Get the URL of the S3 bucket (URL)
- Send the BASE64 content to elasticsearch like this:

```auto
POST my-index-000001/_doc?pipeline=attachment
{
  "url": "URL",
  "data": "BASE64"
}

```

The pipeline I'd use for this is:

```auto
PUT _ingest/pipeline/attachment
{
  "description" : "Extract attachment information",
  "processors" : [
    {
      "attachment" : {
        "field" : "data"
      }
    }, 
    {
      "remove" : {
        "field" : "data"
      }
    }
  ]
}

```

HTH

---

<div class="post-metadata">

**Author:** ![dotel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dotel/32/77606_2.png) [@dotel](https://discuss.elastic.co/u/dotel)\
**Post date:** [July 8, 2021, 3:33am UTC](https://discuss.elastic.co/t/is-there-a-way-to-feed-base64-encoded-string-to-the-ingest-attachment-plugin-or-fscrawler/278016/10 "2021-07-08T03:33:50Z")

</div>

> [@dadoonet](#):
>
> - Send the BASE64 content to elasticsearch like this:
> 
> ```auto
> POST my-index-000001/_doc?pipeline=attachment
> {
> "url": "URL",
> "data": "BASE64"
> }
> 
> ```

Oh, I've actually tried doing this but didn't understand how to actually search after creating the pipeline. [https://www.elastic.co/guide/en/elasticsearch/plugins/current/ingest-attachment-with-arrays.html](https://www.elastic.co/guide/en/elasticsearch/plugins/current/ingest-attachment-with-arrays.html) Is this what I want to use? This basically creates everything inside of attachments field and when searching for something it either gives me both attachment or none at all. Thanks for the help. 🙂  
Also, my documents are available to the application before sending it to s3. Why get the URL and send to elasticsearch instead of doing the same things separately?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 8, 2021, 6:12am UTC](https://discuss.elastic.co/t/is-there-a-way-to-feed-base64-encoded-string-to-the-ingest-attachment-plugin-or-fscrawler/278016/11 "2021-07-08T06:12:34Z")

</div>

> [@dotel](#):
>
> This basically creates everything inside of attachments field and when searching for something it either gives me both attachment or none at all.

You need to provide an example of an indexed document and a query which does not match.

> [@dotel](#):
>
> Also, my documents are available to the application before sending it to s3. Why get the URL and send to elasticsearch instead of doing the same things separately?

I thought you would like to give a link to the user to the original document in the search response.  
If you don't need that, then remove that url I added.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 5, 2021, 6:13am UTC](https://discuss.elastic.co/t/is-there-a-way-to-feed-base64-encoded-string-to-the-ingest-attachment-plugin-or-fscrawler/278016/12 "2021-08-05T06:13:06Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
