# How can I ingest PDF and words files and extract keywords of these documents?

**URL:** <https://discuss.elastic.co/t/how-can-i-ingest-pdf-and-words-files-and-extract-keywords-of-these-documents/133194>\
**Category:** Elasticsearch\
**Created:** [May 24, 2018, 4:30pm UTC](https://discuss.elastic.co/t/how-can-i-ingest-pdf-and-words-files-and-extract-keywords-of-these-documents/133194 "2018-05-24T16:30:18Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Akimbo](https://avatars.discourse-cdn.com/v4/letter/a/b77776/32.png) [@Akimbo](https://discuss.elastic.co/u/Akimbo)\
**Post date:** [May 24, 2018, 4:30pm UTC](https://discuss.elastic.co/t/how-can-i-ingest-pdf-and-words-files-and-extract-keywords-of-these-documents/133194/1 "2018-05-24T16:30:19Z")

</div>

Hi all,  
Currently, I use FSCrawler to ingest files (PDF, words). All files are indexed in ES. Now, I want to do a search engine which uses Elasticsearch and looks like Google search engine. The first step of the search engine is to display keywords of the files in according to the text that i'm tapping.  
Is it possible, when I ingest files with FSCrawler, to index keywords of the files into the ES document ?

For example :  
"\_source": {  
"content": {...},  
"meta": {...} ,  
"file": { "keywords" : { "keyword": "buidling", "keyword": "floor" ... } , ... },  
"path": {...}

Thank you in advance

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 28, 2018, 10:35am UTC](https://discuss.elastic.co/t/how-can-i-ingest-pdf-and-words-files-and-extract-keywords-of-these-documents/133194/2 "2018-05-28T10:35:49Z")

</div>

You can probably run a terms aggregation on field `file.keywords.keyword` I guess.

---

<div class="post-metadata">

**Author:** ![Akimbo](https://avatars.discourse-cdn.com/v4/letter/a/b77776/32.png) [@Akimbo](https://discuss.elastic.co/u/Akimbo)\
**Post date:** [May 28, 2018, 12:28pm UTC](https://discuss.elastic.co/t/how-can-i-ingest-pdf-and-words-files-and-extract-keywords-of-these-documents/133194/3 "2018-05-28T12:28:09Z")

</div>

> [@dadoonet](#):
>
> You can probably run a terms aggregation on field `file.keywords.keyword` I guess.

The problem is that I have no keywords. "file.keywords" does not exist. I need index all terms of the document as keyword.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 28, 2018, 12:41pm UTC](https://discuss.elastic.co/t/how-can-i-ingest-pdf-and-words-files-and-extract-keywords-of-these-documents/133194/4 "2018-05-28T12:41:35Z")

</div>

Two things:

- If the document has "real" keywords, FSCrawler should be able to provide them.
- If you don't have any keyword, then you can only build a tag cloud like I guess which is going to be messy I believe. Anyway, in that case, the only way to build this from a raw text content is by enabling fielddata on field `content`. But this is going to put a lot of pressure I think on your JVM memory.

My 2 cents on this.

---

<div class="post-metadata">

**Author:** ![Akimbo](https://avatars.discourse-cdn.com/v4/letter/a/b77776/32.png) [@Akimbo](https://discuss.elastic.co/u/Akimbo)\
**Post date:** [May 28, 2018, 12:53pm UTC](https://discuss.elastic.co/t/how-can-i-ingest-pdf-and-words-files-and-extract-keywords-of-these-documents/133194/5 "2018-05-28T12:53:04Z")

</div>

The documents don't have "real" keywords... So what is the good way to do auto-completion with theses documents ?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 28, 2018, 1:13pm UTC](https://discuss.elastic.co/t/how-can-i-ingest-pdf-and-words-files-and-extract-keywords-of-these-documents/133194/6 "2018-05-28T13:13:47Z")

</div>

One way is the 2nd point I answered.  
The other way is may be by using suggesters: [https://www.elastic.co/guide/en/elasticsearch/reference/current/search-suggesters.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-suggesters.html)

---

<div class="post-metadata">

**Author:** ![Akimbo](https://avatars.discourse-cdn.com/v4/letter/a/b77776/32.png) [@Akimbo](https://discuss.elastic.co/u/Akimbo)\
**Post date:** [May 29, 2018, 9:48am UTC](https://discuss.elastic.co/t/how-can-i-ingest-pdf-and-words-files-and-extract-keywords-of-these-documents/133194/7 "2018-05-29T09:48:12Z")

</div>

I choose the term suggester to do the auto-completion. Currently, I can request only one index like that :

```auto
GET /_search
{
  "suggest" : {
    "my-suggestion" : {
      "text" : "build",
      "term" : {
        "field" : "content",
        "min_word_length": 2,
        "prefix_length": 5
      }
    }
  }
}

```

But is it possible to do a query on multi-fields of differents index ? I want to do the query on the "content" field of one index and on the "message" field of another index.

I have tried it withou success :

```auto
GET _search
{
  "suggest": {
    "text" : "build",
    "my-suggest-1" : {
      "term" : {
        "field" : "message"
      }
    },
    "my-suggest-2" : {
       "term" : {
        "field" : "content"
       }
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 29, 2018, 9:59am UTC](https://discuss.elastic.co/t/how-can-i-ingest-pdf-and-words-files-and-extract-keywords-of-these-documents/133194/8 "2018-05-29T09:59:58Z")

</div>

Please format your code, logs or configuration files using `</>` icon as explained in [this guide](https://discuss.elastic.co/t/about-the-elasticsearch-category/21) and not the citation button. It will make your post more readable.

Or use markdown style like:

````
```
CODE
```

````

There's a live preview panel for exactly this reasons.

Lots of people read these forums, and many of them will simply skip over a post that is difficult to read, because it's just too large an investment of their time to try and follow a wall of badly formatted text.  
If your goal is to get an answer to your questions, it's in your interest to make it as easy to read and understand as possible.  
Please update your post.

Also, could you create a new question for this as the title is not really related anymore?  
You can link to this question from the new one if you wish.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 26, 2018, 10:12am UTC](https://discuss.elastic.co/t/how-can-i-ingest-pdf-and-words-files-and-extract-keywords-of-these-documents/133194/9 "2018-06-26T10:12:49Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
