# How to index the PDF documents

**URL:** https://discuss.elastic.co/t/how-to-index-the-pdf-documents/327987
**Category:** Elasticsearch
**Created:** [March 18, 2023, 10:14am UTC](https://discuss.elastic.co/t/how-to-index-the-pdf-documents/327987 "2023-03-18T10:14:41Z")
**Posts on this page:** 10
**Page:** 1

<div class="post-metadata">

### Author: ![mruthyu](https://avatars.discourse-cdn.com/v4/letter/m/bb73d2/32.png) [@mruthyu](https://discuss.elastic.co/u/mruthyu)
#### Post date: [March 18, 2023, 10:14am UTC](https://discuss.elastic.co/t/how-to-index-the-pdf-documents/327987/1 "2023-03-18T10:14:41Z")

</div>

How to index the PDF and image documents into elasticsearch. Would like to extract the entities to enable the search on keywords. Whether the workplace search provide this functionality? Whether Apache Tika has been used within elasticsearch or the NLP modules to accomplish this functionality.?

Primarily would like to index few thousands of PDF/Image documents from

1. Local file system (Windows/Linux)
2. AWS S3 buscket.

---

<div class="post-metadata">

### Author: ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)
#### Post date: [March 18, 2023, 11:30am UTC](https://discuss.elastic.co/t/how-to-index-the-pdf-documents/327987/3 "2023-03-18T11:30:51Z")

</div>

You also asked in [IIndexing PDF and Image documents](https://discuss.elastic.co/t/iindexing-pdf-and-image-documents/327988). Let's keep the discussion in one single place.

---

<div class="post-metadata">

### Author: ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)
#### Post date: [March 18, 2023, 11:32am UTC](https://discuss.elastic.co/t/how-to-index-the-pdf-documents/327987/4 "2023-03-18T11:32:38Z")

</div>

You can use the [ingest attachment plugin](https://www.elastic.co/guide/en/elasticsearch/plugins/current/ingest-attachment.html).

There an example here: [https://www.elastic.co/guide/en/elasticsearch/plugins/current/using-ingest-attachment.html](https://www.elastic.co/guide/en/elasticsearch/plugins/current/using-ingest-attachment.html)

```auto
PUT _ingest/pipeline/attachment
{
  "description" : "Extract attachment information",
  "processors" : [
    {
      "attachment" : {
        "field" : "data"
      }
    }
  ]
}
PUT my_index/_doc/my_id?pipeline=attachment
{
  "data": "e1xydGYxXGFuc2kNCkxvcmVtIGlwc3VtIGRvbG9yIHNpdCBhbWV0DQpccGFyIH0="
}
GET my_index/_doc/my_id

```

The `data` field is basically the BASE64 representation of your binary file.

You can use [FSCrawler](https://fscrawler.readthedocs.io). There's [a tutorial](https://fscrawler.readthedocs.io/en/latest/user/tutorial.html) to help you getting started.

---

<div class="post-metadata">

### Author: ![mruthyu](https://avatars.discourse-cdn.com/v4/letter/m/bb73d2/32.png) [@mruthyu](https://discuss.elastic.co/u/mruthyu)
#### Post date: [March 18, 2023, 2:50pm UTC](https://discuss.elastic.co/t/how-to-index-the-pdf-documents/327987/5 "2023-03-18T14:50:48Z")

</div>

Thanks David for your quick response.

I have seen both these options. FSCrawler looks to be the best option. It can feed to Workplace search as well, which provides us with nice UI for search along with facets.  
If I want to use workplace search with the source as Onedrive or Sharepoint online, Whether the same functionality can be achieved?

If we want to use the ingest attachment plugin, how to feed the documents (PDF/IMAGE) in bulk?

Also I am looking into the NLP ML models which are being used in elasticsearch can help to tag these documents with relevant tags. That way the search can be done with the exact value of the identified tags.

---

<div class="post-metadata">

### Author: ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)
#### Post date: [March 19, 2023, 8:37pm UTC](https://discuss.elastic.co/t/how-to-index-the-pdf-documents/327987/6 "2023-03-19T20:37:39Z")

</div>

Have a look at [Connecting SharePoint Online | Workplace Search Guide [8.6] | Elastic](https://www.elastic.co/guide/en/workplace-search/current/workplace-search-sharepoint-online-connector.html) and [Connecting OneDrive | Workplace Search Guide [8.6] | Elastic](https://www.elastic.co/guide/en/workplace-search/current/workplace-search-onedrive-connector.html)

---

<div class="post-metadata">

### Author: ![mruthyu](https://avatars.discourse-cdn.com/v4/letter/m/bb73d2/32.png) [@mruthyu](https://discuss.elastic.co/u/mruthyu)
#### Post date: [March 20, 2023, 7:06am UTC](https://discuss.elastic.co/t/how-to-index-the-pdf-documents/327987/7 "2023-03-20T07:06:19Z")

</div>

Sure. Have checked those documents and not able to see the list of supported file formats from those data sources. So default PDF and image file formats are supported?

---

<div class="post-metadata">

### Author: ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)
#### Post date: [March 20, 2023, 7:14am UTC](https://discuss.elastic.co/t/how-to-index-the-pdf-documents/327987/8 "2023-03-20T07:14:38Z")

</div>

I believe so. 😊

---

<div class="post-metadata">

### Author: ![mruthyu](https://avatars.discourse-cdn.com/v4/letter/m/bb73d2/32.png) [@mruthyu](https://discuss.elastic.co/u/mruthyu)
#### Post date: [March 20, 2023, 9:28am UTC](https://discuss.elastic.co/t/how-to-index-the-pdf-documents/327987/9 "2023-03-20T09:28:54Z")

</div>

Sure. Thanks.

Can you please provide your inputs for the following query?

_Also I am looking into the NLP ML models which are being used in elasticsearch can help to tag these documents with relevant tags. That way the search can be done with the exact value of the identified tags._

---

<div class="post-metadata">

### Author: ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)
#### Post date: [March 20, 2023, 10:36am UTC](https://discuss.elastic.co/t/how-to-index-the-pdf-documents/327987/10 "2023-03-20T10:36:03Z")

</div>

I did not play with NLP yet. But I'd check: [Overview | Machine Learning in the Elastic Stack [8.6] | Elastic](https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-overview.html)

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [April 17, 2023, 10:37am UTC](https://discuss.elastic.co/t/how-to-index-the-pdf-documents/327987/11 "2023-04-17T10:37:03Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
