# Can we do : Analyser-\>Tokenizer-\>Token Filter-\>Re-tokenize and considers only these last tokens

**URL:** <https://discuss.elastic.co/t/can-we-do-analyser-tokenizer-token-filter-re-tokenize-and-considers-only-these-last-tokens/307298>\
**Category:** Elasticsearch\
**Created:** [June 15, 2022, 3:26pm UTC](https://discuss.elastic.co/t/can-we-do-analyser-tokenizer-token-filter-re-tokenize-and-considers-only-these-last-tokens/307298 "2022-06-15T15:26:01Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![fraf](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fraf/32/107072_2.png) [@fraf](https://discuss.elastic.co/u/fraf)\
**Post date:** [June 15, 2022, 3:26pm UTC](https://discuss.elastic.co/t/can-we-do-analyser-tokenizer-token-filter-re-tokenize-and-considers-only-these-last-tokens/307298/1 "2022-06-15T15:26:01Z")

</div>

Hello,

I want, given the input text :

```auto
{
  "analyzer": "parapheur_shingle",
  "text": "W 4.8.4.1 NI FNA NP 4.8.4 TEST 2"
}

```

having fllowing steps :

1. tokenize as standard (word break with spaces =\> 8 tokens)
2. filter each token ("lowercase", "pattern\_replace"). "pattern\_replace" replaces "." to " ".
3. so we obtain the token "4 8 4 1" after applying pattern\_replace to original token "4.8.4.1"
4. Re-do 1) =\> 8 + 5 more tokens
5. filter each token with shingle filter (this filter relies on generated token, so output will differ in comparison between after 3) or here in 5)).
6. End of token generation.

Thank you.

---

<div class="post-metadata">

**Author:** ![fraf](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fraf/32/107072_2.png) [@fraf](https://discuss.elastic.co/u/fraf)\
**Post date:** [June 15, 2022, 3:45pm UTC](https://discuss.elastic.co/t/can-we-do-analyser-tokenizer-token-filter-re-tokenize-and-considers-only-these-last-tokens/307298/2 "2022-06-15T15:45:24Z")

</div>

I managed to do what I want with "simple\_pattern\_split" tokenizer :

```auto
{
  "index": {
    "max_ngram_diff": 50,
    "analysis": {
      "analyzer": {
        "parapheur_shingle_new": {
          "tokenizer": "pattern_split_new",
          "filter": ["shingle"]
        }
      },
	  "tokenizer": {
	    "pattern_split_new": {
	      "type": "simple_pattern_split",
	      "pattern": "[\\s+ \\.]"
	    }
	  }
    }
  }
}

```

But I'd like the split pattern to perform the same split that "standard" tokenization, instead of my hard coded pattern.

---

<div class="post-metadata">

**Author:** ![RabBit\_BR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rabbit_br/32/82261_2.png) [@RabBit\_BR](https://discuss.elastic.co/u/RabBit_BR)\
**Post date:** [June 15, 2022, 6:23pm UTC](https://discuss.elastic.co/t/can-we-do-analyser-tokenizer-token-filter-re-tokenize-and-considers-only-these-last-tokens/307298/3 "2022-06-15T18:23:43Z")

</div>

Hi @fraf

Do you want the tokens that [way](https://gist.github.com/andreluiz1987/04d4d821494f3193e65e30f69cc903d3)?

---

<div class="post-metadata">

**Author:** ![fraf](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fraf/32/107072_2.png) [@fraf](https://discuss.elastic.co/u/fraf)\
**Post date:** [June 16, 2022, 7:35am UTC](https://discuss.elastic.co/t/can-we-do-analyser-tokenizer-token-filter-re-tokenize-and-considers-only-these-last-tokens/307298/4 "2022-06-16T07:35:42Z")

</div>

Yes. And even more token. I want max\_shingle\_size to be infinite ; that is to say ; the longest token is the input text (with dot replaced by space).  
But it should be possible given this error message from server :

`In Shingle TokenFilter the difference between max_shingle_size and min_shingle_size (and +1 if outputting unigrams) must be less than or equal to: [3] but was [19]. This limit can be set by changing the [index.max_shingle_diff] index level setting.`

Thx.

---

<div class="post-metadata">

**Author:** ![RabBit\_BR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rabbit_br/32/82261_2.png) [@RabBit\_BR](https://discuss.elastic.co/u/RabBit_BR)\
**Post date:** [June 16, 2022, 12:38pm UTC](https://discuss.elastic.co/t/can-we-do-analyser-tokenizer-token-filter-re-tokenize-and-considers-only-these-last-tokens/307298/5 "2022-06-16T12:38:02Z")

</div>

This is the analyzer that I used:

```auto
PUT teste
{
  "settings": {
    "analysis": {
      "analyzer": {
        "parapheur_shingle_new": {
          "tokenizer": "standard",
          "filter": [
            "lowercase",
            "my_filter",
            "shingle_filter"
          ]
        }
      },
      "filter": {
        "my_filter":{
          "type": "pattern_replace",
          "pattern": """\.""",
          "replacement": " "
        },
        "shingle_filter": {
          "type": "shingle",
          "min_shingle_size": 2,
          "max_shingle_size": 4
        }
      }
    }
  }
}

GET teste/_analyze
{
  "analyzer": "parapheur_shingle_new",
  "text": "W 4.8.4.1 NI FNA NP 4.8.4 TEST 2"
}

```

---

<div class="post-metadata">

**Author:** ![fraf](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fraf/32/107072_2.png) [@fraf](https://discuss.elastic.co/u/fraf)\
**Post date:** [June 16, 2022, 1:17pm UTC](https://discuss.elastic.co/t/can-we-do-analyser-tokenizer-token-filter-re-tokenize-and-considers-only-these-last-tokens/307298/6 "2022-06-16T13:17:18Z")

</div>

Thx, but be careful that's not exactly the same.

Given the input text "NP 4.8.1", you won't be able to match the input search "NP 4".  
With standard tokenizer, you obtain 2 tokens : "NP" and "4.8.1".  
On these two tokens, you apply your filters :  
"NP =\> unchanged =\> "NP"  
"4.8.1" =\> myfilter =\> "4 8 1" processed token  
shingle\_filter =\> it shingle two adjacent tokens ; so it generates the new token : "NP 4 8 1".

So you have three tokens in all : "NP", "4 8 1" and "NP 4 8 1"...... missing "NP 4" token ! So no matches.  
With shingle\_filter you have to operates on tokenizer, whatever the filter chain is. So I was using the "simple\_pattern\_split" tokenizer. But I need to enumerate word breaks characters.

---

<div class="post-metadata">

**Author:** ![RabBit\_BR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rabbit_br/32/82261_2.png) [@RabBit\_BR](https://discuss.elastic.co/u/RabBit_BR)\
**Post date:** [June 16, 2022, 1:53pm UTC](https://discuss.elastic.co/t/can-we-do-analyser-tokenizer-token-filter-re-tokenize-and-considers-only-these-last-tokens/307298/7 "2022-06-16T13:53:26Z")

</div>

Complicated, because even increasing the shingle limit you will get more tokens in addition to the 4 you want and I don't know if that's what you need.

---

<div class="post-metadata">

**Author:** ![fraf](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fraf/32/107072_2.png) [@fraf](https://discuss.elastic.co/u/fraf)\
**Post date:** [June 16, 2022, 2:36pm UTC](https://discuss.elastic.co/t/can-we-do-analyser-tokenizer-token-filter-re-tokenize-and-considers-only-these-last-tokens/307298/8 "2022-06-16T14:36:54Z")

</div>

I agree with you, but certainly much less than `min_gram` and `max_gram` parameters whose `NGram tokenizer` relies on. Cutting a phrase into words, generates much less token than cutting into characters for sure.  
The dilemma is : how much word the user can enter in its input ? It follows then directly the param value `max_shingle_size`

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 14, 2022, 2:37pm UTC](https://discuss.elastic.co/t/can-we-do-analyser-tokenizer-token-filter-re-tokenize-and-considers-only-these-last-tokens/307298/9 "2022-07-14T14:37:49Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
