# Is it possible to search in random files with random texts in it with Elastic Search?

**URL:** https://discuss.elastic.co/t/is-it-possible-to-search-in-random-files-with-random-texts-in-it-with-elastic-search/289863
**Category:** Elasticsearch
**Created:** [November 22, 2021, 6:41pm UTC](https://discuss.elastic.co/t/is-it-possible-to-search-in-random-files-with-random-texts-in-it-with-elastic-search/289863 "2021-11-22T18:41:18Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![Sahin\_Kasap](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sahin_kasap/32/97407_2.png) [@Sahin\_Kasap](https://discuss.elastic.co/u/Sahin_Kasap)
#### Post date: [November 22, 2021, 6:41pm UTC](https://discuss.elastic.co/t/is-it-possible-to-search-in-random-files-with-random-texts-in-it-with-elastic-search/289863/1 "2021-11-22T18:41:18Z")

</div>

Hello, I am developing a backend and I have big pdf files(~20 MB), which I want to search words inside.

I could upload the texts inside them to Sql or somewhere and search inside them, but this would be slow and pdf's have lots of images inside which makes me think of using OCR (and Elasticsearch gives the ability for this)

But these are not log files or something, there are only words in it. I tried to upload it via Kibana, tutoarial -upload file page but it gives me error "File structure cannot be determined".

1 - Is it possible to search inside files with random texts in it?  
2 - How to upload file and index if possible?

---

<div class="post-metadata">

### Author: ![ivar.ekman](https://avatars.discourse-cdn.com/v4/letter/i/b782af/32.png) [@ivar.ekman](https://discuss.elastic.co/u/ivar.ekman)
#### Post date: [November 22, 2021, 6:48pm UTC](https://discuss.elastic.co/t/is-it-possible-to-search-in-random-files-with-random-texts-in-it-with-elastic-search/289863/2 "2021-11-22T18:48:11Z")

</div>

You need to extract the text out from the PDFs somehow and push the text to Elasticsearch to be able to search it. There are multiple ways to get the text out from PDFs. However, often getting just the text is not enough but you want to more with it. Structure it according to the content and understand the meaning of the text. For example extract dates and document classifications and out from the text you often want to extract meaningful content, such as companies and persons as metadata also. It all goes down to what use case you really are solving.

---

<div class="post-metadata">

### Author: ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)
#### Post date: [November 22, 2021, 9:34pm UTC](https://discuss.elastic.co/t/is-it-possible-to-search-in-random-files-with-random-texts-in-it-with-elastic-search/289863/3 "2021-11-22T21:34:03Z")

</div>

You can use the [ingest attachment plugin](https://www.elastic.co/guide/en/elasticsearch/plugins/current/ingest-attachment.html).

There an example here: [Using the Attachment Processor in a Pipeline | Elasticsearch Plugins and Integrations [7.15] | Elastic](https://www.elastic.co/guide/en/elasticsearch/plugins/current/using-ingest-attachment.html)

```auto
PUT _ingest/pipeline/attachment
{
  "description" : "Extract attachment information",
  "processors" : [
    {
      "attachment" : {
        "field" : "data"
      }
    }
  ]
}
PUT my_index/_doc/my_id?pipeline=attachment
{
  "data": "e1xydGYxXGFuc2kNCkxvcmVtIGlwc3VtIGRvbG9yIHNpdCBhbWV0DQpccGFyIH0="
}
GET my_index/_doc/my_id

```

The `data` field is basically the BASE64 representation of your binary file.

You can use [FSCrawler](https://fscrawler.readthedocs.io). There's [a tutorial](https://fscrawler.readthedocs.io/en/latest/user/tutorial.html) to help you getting started.

---

<div class="post-metadata">

### Author: ![Sahin\_Kasap](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sahin_kasap/32/97407_2.png) [@Sahin\_Kasap](https://discuss.elastic.co/u/Sahin_Kasap)
#### Post date: [November 23, 2021, 6:57am UTC](https://discuss.elastic.co/t/is-it-possible-to-search-in-random-files-with-random-texts-in-it-with-elastic-search/289863/4 "2021-11-23T06:57:25Z")

</div>

Here is an example pdf I want to search in [https://www.resmigazete.gov.tr/eskiler/2021/11/20211123.pdf](https://www.resmigazete.gov.tr/eskiler/2021/11/20211123.pdf) .  
I want to

- Get if the word exists in the pdf
- Get the words (for example 100 chars before and after the word) around the word we search for.

---

<div class="post-metadata">

### Author: ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)
#### Post date: [November 23, 2021, 8:25am UTC](https://discuss.elastic.co/t/is-it-possible-to-search-in-random-files-with-random-texts-in-it-with-elastic-search/289863/5 "2021-11-23T08:25:48Z")

</div>

FSCrawler Will do that as it supports ocr.

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [December 21, 2021, 8:25am UTC](https://discuss.elastic.co/t/is-it-possible-to-search-in-random-files-with-random-texts-in-it-with-elastic-search/289863/6 "2021-12-21T08:25:48Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
