# IIndexing PDF and Image documents

**URL:** https://discuss.elastic.co/t/iindexing-pdf-and-image-documents/327988
**Category:** Elastic Search
**Tags:** elastic-workplace-search
**Created:** [March 18, 2023, 11:10am UTC](https://discuss.elastic.co/t/iindexing-pdf-and-image-documents/327988 "2023-03-18T11:10:25Z")
**Posts on this page:** 3
**Page:** 1

<div class="post-metadata">

### Author: ![mruthyu](https://avatars.discourse-cdn.com/v4/letter/m/bb73d2/32.png) [@mruthyu](https://discuss.elastic.co/u/mruthyu)
#### Post date: [March 18, 2023, 11:10am UTC](https://discuss.elastic.co/t/iindexing-pdf-and-image-documents/327988/1 "2023-03-18T11:10:25Z")

</div>

Would like to know whether PDF and Image documents (Stored in local file system/ AWS S3) can be indexed to enable the search on key entities depending on the nature of the documents (School Admission form, Examination Form, Purchase Orders, Telephone Bills, Electricity Bills..etc )

---

<div class="post-metadata">

### Author: ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)
#### Post date: [March 18, 2023, 11:31am UTC](https://discuss.elastic.co/t/iindexing-pdf-and-image-documents/327988/2 "2023-03-18T11:31:07Z")

</div>



---

<div class="post-metadata">

### Author: ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)
#### Post date: [March 19, 2023, 8:19pm UTC](https://discuss.elastic.co/t/iindexing-pdf-and-image-documents/327988/3 "2023-03-19T20:19:09Z")

</div>

Please don't post the same question more than once, it makes it harder to help you 🙂

> [@How to index the PDF documents](https://discuss.elastic.co/t/how-to-index-the-pdf-documents/327987):
>
> How to index the PDF and image documents into elasticsearch. Would like to extract the entities to enable the search on keywords. Whether the workplace search provide this functionality? Whether Apache Tika has been used within elasticsearch or the NLP modules to accomplish this functionality.? Primarily would like to index few thousands of PDF/Image documents from Local file system (Windows/Linux) AWS S3 buscket.
