# Visualizing the count of words in each document(pdf, word) in kibana using FSCRAWLER

**URL:** <https://discuss.elastic.co/t/visualizing-the-count-of-words-in-each-document-pdf-word-in-kibana-using-fscrawler/116897>\
**Category:** Kibana\
**Created:** [January 24, 2018, 4:25pm UTC](https://discuss.elastic.co/t/visualizing-the-count-of-words-in-each-document-pdf-word-in-kibana-using-fscrawler/116897 "2018-01-24T16:25:33Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![manasguduri](https://avatars.discourse-cdn.com/v4/letter/m/50afbb/32.png) [@manasguduri](https://discuss.elastic.co/u/manasguduri)\
**Post date:** [January 24, 2018, 4:25pm UTC](https://discuss.elastic.co/t/visualizing-the-count-of-words-in-each-document-pdf-word-in-kibana-using-fscrawler/116897/1 "2018-01-24T16:25:34Z")

</div>

![index_type](https://us1.discourse-cdn.com/elastic/original/3X/3/e/3ee798987e3f826b76a7a762476f1abfe85e3e24.png)

The screenshot I provided is the indexing done by fscrawler.

Suppose, from the screenshot if it says like content contains the words in document, when I try to visualize that in kibana, and I select aggregation as 'Terms' and the 'content' doesn't even show in the split series.

Am trying it for days, please help !

Thanks in advance.

Manas

---

<div class="post-metadata">

**Author:** ![Nathan\_Reese](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nathan_reese/32/84829_2.png) [@Nathan\_Reese](https://discuss.elastic.co/u/Nathan_Reese)\
**Post date:** [January 24, 2018, 5:25pm UTC](https://discuss.elastic.co/t/visualizing-the-count-of-words-in-each-document-pdf-word-in-kibana-using-fscrawler/116897/2 "2018-01-24T17:25:39Z")

</div>

The `content` field is getting stored in elastic search as [text](https://www.elastic.co/guide/en/elasticsearch/reference/current/text.html). Text fields can not be used for `terms` aggregation. The field must have the [keyword](https://www.elastic.co/guide/en/elasticsearch/reference/current/keyword.html) mapping type.

The [terms aggregation](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html) creates a bucket per unique value. A value is not a token from a larger string but rather the entire string.

---

<div class="post-metadata">

**Author:** ![manasguduri](https://avatars.discourse-cdn.com/v4/letter/m/50afbb/32.png) [@manasguduri](https://discuss.elastic.co/u/manasguduri)\
**Post date:** [January 24, 2018, 6:07pm UTC](https://discuss.elastic.co/t/visualizing-the-count-of-words-in-each-document-pdf-word-in-kibana-using-fscrawler/116897/3 "2018-01-24T18:07:45Z")

</div>

Hi Nathan, Thanks for responding, actually I have done this indexing using FSCRAWLER.

I am new to Elasticsearch and Kibana. I have then came to know about fscrawler and used it, as it supports pdf indexing. I am not sure how to change the keyword mapping, can u help me with that?

This is a screenshot of my fscrawler .json file. Should I give ""type" : keyword in this file.

 ![fscrawler](https://us1.discourse-cdn.com/elastic/original/3X/0/d/0db1d622f90562f61426d7de78987a63301c1ff1.png)

Thanks in advance !

Manas

---

<div class="post-metadata">

**Author:** ![manasguduri](https://avatars.discourse-cdn.com/v4/letter/m/50afbb/32.png) [@manasguduri](https://discuss.elastic.co/u/manasguduri)\
**Post date:** [January 24, 2018, 6:11pm UTC](https://discuss.elastic.co/t/visualizing-the-count-of-words-in-each-document-pdf-word-in-kibana-using-fscrawler/116897/4 "2018-01-24T18:11:37Z")

</div>

Previously, I used a python script to index the pdf file and it shows keyword mapping as you are saying but it still doesn't take the tokens from the 'file.keyword' but the entire string.

Thanks,  
Manas

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 21, 2018, 6:11pm UTC](https://discuss.elastic.co/t/visualizing-the-count-of-words-in-each-document-pdf-word-in-kibana-using-fscrawler/116897/5 "2018-02-21T18:11:47Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
