# Word count from documents

**URL:** <https://discuss.elastic.co/t/word-count-from-documents/117379>\
**Category:** Elasticsearch\
**Created:** [January 28, 2018, 7:14pm UTC](https://discuss.elastic.co/t/word-count-from-documents/117379 "2018-01-28T19:14:36Z")\
**Posts on this page:** 1\
**Showing post:** 4

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 29, 2018, 3:28pm UTC](https://discuss.elastic.co/t/word-count-from-documents/117379/4 "2018-01-29T15:28:12Z")

</div>

I don't read Python code. So if you could you provide a full recreation script as described in [About the Elasticsearch category](https://discuss.elastic.co/t/about-the-elasticsearch-category/21). It will help to better understand what you are doing. Please, try to keep the example as simple as possible.

> The other solution you were saying ingest-attachment, am not familiar on how to do that !!

Not really another solution but part of it. If you want to extract text from a PDF document, you can use:

- ingest-attachment: [Ingest Attachment plugin | Elasticsearch Plugins and Integrations [8.11] | Elastic](https://www.elastic.co/guide/en/elasticsearch/plugins/current/ingest-attachment.html)
- FSCrawler: [GitHub - dadoonet/fscrawler: Elasticsearch File System Crawler (FS Crawler)](https://github.com/dadoonet/fscrawler)
- Apache Tika directly in Java: [https://tika.apache.org/](https://tika.apache.org/)

---

_[View the full topic](https://discuss.elastic.co/t/word-count-from-documents/117379)._
