# Word count from documents

**URL:** <https://discuss.elastic.co/t/word-count-from-documents/117379>\
**Category:** Elasticsearch\
**Created:** [January 28, 2018, 7:14pm UTC](https://discuss.elastic.co/t/word-count-from-documents/117379 "2018-01-28T19:14:36Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![manasguduri](https://avatars.discourse-cdn.com/v4/letter/m/50afbb/32.png) [@manasguduri](https://discuss.elastic.co/u/manasguduri)\
**Post date:** [January 28, 2018, 7:14pm UTC](https://discuss.elastic.co/t/word-count-from-documents/117379/1 "2018-01-28T19:14:37Z")

</div>

Is it possible to index a pdf document to visualize the count of words or like top 10 words with their count ?

Thanks in advance.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 28, 2018, 7:45pm UTC](https://discuss.elastic.co/t/word-count-from-documents/117379/2 "2018-01-28T19:45:15Z")

</div>

You can do that by indexing the content (with `ingest-attachment`) in a `text` field with `fielddata: true`. Or may be add a `keyword` subfield but you might hit a limit.

My 2 cents.

---

<div class="post-metadata">

**Author:** ![manasguduri](https://avatars.discourse-cdn.com/v4/letter/m/50afbb/32.png) [@manasguduri](https://discuss.elastic.co/u/manasguduri)\
**Post date:** [January 28, 2018, 7:51pm UTC](https://discuss.elastic.co/t/word-count-from-documents/117379/3 "2018-01-28T19:51:50Z")

</div>

Hello dadoonet, thanks for your quick response.

I have tried using the keyword subfield but am unable to do that ! (I am using a python code to index my documents , link - [https://gist.github.com/stevehanson/7462063](https://gist.github.com/stevehanson/7462063)).

The other solution you were saying ingest-attachment, am not familiar on how to do that !!

Please help.

Thanks.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 29, 2018, 3:28pm UTC](https://discuss.elastic.co/t/word-count-from-documents/117379/4 "2018-01-29T15:28:12Z")

</div>

I don't read Python code. So if you could you provide a full recreation script as described in [About the Elasticsearch category](https://discuss.elastic.co/t/about-the-elasticsearch-category/21). It will help to better understand what you are doing. Please, try to keep the example as simple as possible.

> The other solution you were saying ingest-attachment, am not familiar on how to do that !!

Not really another solution but part of it. If you want to extract text from a PDF document, you can use:

- ingest-attachment: [Ingest Attachment plugin | Elasticsearch Plugins and Integrations [8.11] | Elastic](https://www.elastic.co/guide/en/elasticsearch/plugins/current/ingest-attachment.html)
- FSCrawler: [GitHub - dadoonet/fscrawler: Elasticsearch File System Crawler (FS Crawler)](https://github.com/dadoonet/fscrawler)
- Apache Tika directly in Java: [https://tika.apache.org/](https://tika.apache.org/)

---

<div class="post-metadata">

**Author:** ![manasguduri](https://avatars.discourse-cdn.com/v4/letter/m/50afbb/32.png) [@manasguduri](https://discuss.elastic.co/u/manasguduri)\
**Post date:** [January 30, 2018, 1:12am UTC](https://discuss.elastic.co/t/word-count-from-documents/117379/5 "2018-01-30T01:12:19Z")

</div>

I have tried this:

1. I have used tika with the python code i shared and it takes the data as '.keyword', but it doesn't show the count of individual words in a pdf file.

2. I have used fscrawler, it takes the data as content and not as '.keyword' format, so even the field doesn't show in visualization tab.

 ![index_type](https://us1.discourse-cdn.com/elastic/original/3X/3/e/3ee798987e3f826b76a7a762476f1abfe85e3e24.png)

1. Using ingest plugin, am still working on it, am not exactly finding a way to index a pdf file, am going through lot of issues. Will work on that.

You asked me to provide a script but from the types I went through doesn't require them. All I need to do is give the directory name in which files are stored, then it will do the work for me !

I have been working on this for days now, and I lost my belief that elasticsearch will be able to individually count the words in a pdf file.

Can you please give me some references where someone had did it really, becaause I don't want to waste anymore time on this !

You are my only hope. Please help !

Regards,  
Manas

---

<div class="post-metadata">

**Author:** ![manasguduri](https://avatars.discourse-cdn.com/v4/letter/m/50afbb/32.png) [@manasguduri](https://discuss.elastic.co/u/manasguduri)\
**Post date:** [January 31, 2018, 4:42pm UTC](https://discuss.elastic.co/t/word-count-from-documents/117379/6 "2018-01-31T16:42:42Z")

</div>

Here's a part of the script:

GET /test/\_search  
{  
"took": 0,  
"timed\_out": false,  
"\_shards": {  
"total": 5,  
"successful": 5,  
"skipped": 0,  
"failed": 0  
},  
"hits": {  
"total": 19,  
"max\_score": 1,  
"hits": [  
{  
"\_index": "test",  
"\_type": "attachment",  
"\_id": "gdYkmWABst9CE15hAvBX",  
"\_score": 1,  
"\_source": {  
"file": """  
Elastic Search Logstash Kibana

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 31, 2018, 5:21pm UTC](https://discuss.elastic.co/t/word-count-from-documents/117379/7 "2018-01-31T17:21:52Z")

</div>

My first answer was wrong. Sorry.

You can use term vectors I think. Like:

```auto
DELETE test
PUT test
{
  "mappings": {
    "doc": {
      "properties": {
        "foo": {
          "type": "text",
          "term_vector": "yes",
          "store": true
        }
      }
    }
  }
}
POST test/doc/1
{
  "foo": "a b c c"
}
GET test/doc/1/_termvectors

```

It gives:

```auto
{
  "_index": "test",
  "_type": "doc",
  "_id": "1",
  "_version": 1,
  "found": true,
  "took": 5,
  "term_vectors": {
    "foo": {
      "field_statistics": {
        "sum_doc_freq": 3,
        "doc_count": 1,
        "sum_ttf": 4
      },
      "terms": {
        "a": {
          "term_freq": 1
        },
        "b": {
          "term_freq": 1
        },
        "c": {
          "term_freq": 2
        }
      }
    }
  }
}

```

More on this at: [https://www.elastic.co/guide/en/elasticsearch/reference/current/docs-termvectors.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/docs-termvectors.html)

HTH

---

<div class="post-metadata">

**Author:** ![manasguduri](https://avatars.discourse-cdn.com/v4/letter/m/50afbb/32.png) [@manasguduri](https://discuss.elastic.co/u/manasguduri)\
**Post date:** [January 31, 2018, 5:34pm UTC](https://discuss.elastic.co/t/word-count-from-documents/117379/8 "2018-01-31T17:34:20Z")

</div>

Actually am able to get the term vectors even now, but am not able to visualize it on Kibana !!

But I'l create a new index using the mapping format you suggested and let u know.

Thanks.

---

<div class="post-metadata">

**Author:** ![manasguduri](https://avatars.discourse-cdn.com/v4/letter/m/50afbb/32.png) [@manasguduri](https://discuss.elastic.co/u/manasguduri)\
**Post date:** [January 31, 2018, 5:48pm UTC](https://discuss.elastic.co/t/word-count-from-documents/117379/9 "2018-01-31T17:48:38Z")

</div>

Hello David,

Even now it takes the foo field as whole string but not as an individual word !

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 31, 2018, 5:51pm UTC](https://discuss.elastic.co/t/word-count-from-documents/117379/10 "2018-01-31T17:51:04Z")

</div>

You need to index that information I guess within the the document.

So you could call elasticsearch like:

```auto
GET /test/doc/_termvectors
{
  "doc" : {
    "foo" : "a b c c"
  }
}

```

And get the result, enrich your document with that information somehow.

Not sure about what you would do then with that information though. May be sum the number of words and store it in your doc?

Or may be write an ingest script to do that computation at index time? [https://www.elastic.co/guide/en/elasticsearch/reference/master/script-processor.html](https://www.elastic.co/guide/en/elasticsearch/reference/master/script-processor.html)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 28, 2018, 5:51pm UTC](https://discuss.elastic.co/t/word-count-from-documents/117379/11 "2018-02-28T17:51:20Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
