# How to index a certain number of words from each text?

**URL:** <https://discuss.elastic.co/t/how-to-index-a-certain-number-of-words-from-each-text/294506>\
**Category:** Elasticsearch\
**Created:** [January 16, 2022, 5:40pm UTC](https://discuss.elastic.co/t/how-to-index-a-certain-number-of-words-from-each-text/294506 "2022-01-16T17:40:41Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Amir\_Faramarzi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/amir_faramarzi/32/92919_2.png) [@Amir\_Faramarzi](https://discuss.elastic.co/u/Amir_Faramarzi)\
**Post date:** [January 16, 2022, 5:40pm UTC](https://discuss.elastic.co/t/how-to-index-a-certain-number-of-words-from-each-text/294506/1 "2022-01-16T17:40:41Z")

</div>

How can only the first two words of each text be indexed ??

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 16, 2022, 7:16pm UTC](https://discuss.elastic.co/t/how-to-index-a-certain-number-of-words-from-each-text/294506/2 "2022-01-16T19:16:17Z")

</div>

You can use a edge ngram token filter. See [Edge n-gram token filter | Elasticsearch Guide [7.16] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-edgengram-tokenfilter.html)

---

<div class="post-metadata">

**Author:** ![Amir\_Faramarzi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/amir_faramarzi/32/92919_2.png) [@Amir\_Faramarzi](https://discuss.elastic.co/u/Amir_Faramarzi)\
**Post date:** [January 16, 2022, 10:09pm UTC](https://discuss.elastic.co/t/how-to-index-a-certain-number-of-words-from-each-text/294506/3 "2022-01-16T22:09:38Z")

</div>

I want to index words and I do not want to index the first two words into small pieces. Although I still have access to the first two words, I do not need the rest of the indexes provided by ngram.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 16, 2022, 10:25pm UTC](https://discuss.elastic.co/t/how-to-index-a-certain-number-of-words-from-each-text/294506/4 "2022-01-16T22:25:48Z")

</div>

You probably need to create a custom analyser which discards everything after the two first works and then tokenizes e.g. based on whitespace.

---

<div class="post-metadata">

**Author:** ![Tomo\_M](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@Tomo\_M](https://discuss.elastic.co/u/Tomo_M)\
**Post date:** [January 17, 2022, 3:29am UTC](https://discuss.elastic.co/t/how-to-index-a-certain-number-of-words-from-each-text/294506/5 "2022-01-17T03:29:21Z")

</div>

[Limit token count token filter](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-limit-token-count-tokenfilter.html) will match the purpose.

There are many built-in tokenizers and filters, it's worth see everything for once.

```auto
POST _analyze
{
  "tokenizer": "standard", 
  "filter":[
    {
      "type": "limit",
      "max_token_count": 2
    }
  ],
  "text": "aaa bbb ccc ddd"
}

```

```auto
{
  "tokens" : [
    {
      "token" : "aaa",
      "start_offset" : 0,
      "end_offset" : 3,
      "type" : "<ALPHANUM>",
      "position" : 0
    },
    {
      "token" : "bbb",
      "start_offset" : 4,
      "end_offset" : 7,
      "type" : "<ALPHANUM>",
      "position" : 1
    }
  ]
}

```

---

<div class="post-metadata">

**Author:** ![Amir\_Faramarzi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/amir_faramarzi/32/92919_2.png) [@Amir\_Faramarzi](https://discuss.elastic.co/u/Amir_Faramarzi)\
**Post date:** [January 17, 2022, 5:31am UTC](https://discuss.elastic.co/t/how-to-index-a-certain-number-of-words-from-each-text/294506/7 "2022-01-17T05:31:55Z")

</div>

This filter indexes words from the beginning of the text  
I said for example the first two words  
But I may want to index only the third to fifth words

---

<div class="post-metadata">

**Author:** ![Tomo\_M](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@Tomo\_M](https://discuss.elastic.co/u/Tomo_M)\
**Post date:** [January 17, 2022, 9:54am UTC](https://discuss.elastic.co/t/how-to-index-a-certain-number-of-words-from-each-text/294506/8 "2022-01-17T09:54:13Z")

</div>

If it was an example, clear stating your requirement ("only the third to fifth words" is also an example?) would lead earlier solution without bothering those who answer here.

In case to filter tokens in any position, I suppose it is better to filter on the client side before indexing or create ingest pipeline with your own script.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 14, 2022, 9:54am UTC](https://discuss.elastic.co/t/how-to-index-a-certain-number-of-words-from-each-text/294506/9 "2022-02-14T09:54:33Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
