# Alternative option of shingle token filter

**URL:** <https://discuss.elastic.co/t/alternative-option-of-shingle-token-filter/285787>\
**Category:** Elasticsearch\
**Created:** [October 4, 2021, 7:18am UTC](https://discuss.elastic.co/t/alternative-option-of-shingle-token-filter/285787 "2021-10-04T07:18:59Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![ankurpatel](https://avatars.discourse-cdn.com/v4/letter/a/278dde/32.png) [@ankurpatel](https://discuss.elastic.co/u/ankurpatel)\
**Post date:** [October 4, 2021, 7:18am UTC](https://discuss.elastic.co/t/alternative-option-of-shingle-token-filter/285787/1 "2021-10-04T07:18:59Z")

</div>

I would like to implement auto suggest functionality just like google search. Shingle Token Filter is best suitable option for my requirement. Let me explain my index details.

1. Index size is 16gb - 20 million documents. Each document average size is 0.53 Kb.
2. Shingle filter is applied on WorkDescripton field. Average size of WorkDescripton field is 0.45kb.
3. Below Analyzer, filter and mapping used at Index creation time  
a. Filter

```auto
 "shingle-filter" : {
                       "max_shingle_size" : "3",
                       "min_shingle_size" : "2",
                       "output_unigrams" : "false",
                       "type" : "shingle"
                  }  

```

```
      b. Analyzer-

```

```auto
                    "ana_autocomplete" : {
                       "filter" : [
                                "lowercase",
                               "shingle-filter"
                        ],
                       "tokenizer" : "standard"
                  }

```

```
 c. Mapping with work description field is

```

```auto
 "workDesc" : 
                 {
                     "type" : "text",
                     "fields" : 
                     {
                          "suggestions" : 
                          {
                               "type" : "text",
                               "analyzer" : "ana_autocomplete",
                               "fielddata" : true,
                               "fielddata_frequency_filter" : 
                                {
                                       "min" : 0.001,
                                       "max" : 0.1,
                                      "min_segment_size" : 500
                                }
                           },
                           "workDesc" : 
                          {
                                 "type" : "text",
                                 "analyzer" : "ana_tenderinfo"
                          }
                    }
             }

```

1. I used aggregate query to get the result

```auto
GET tenderinfo_version_9/_search
      {
            "aggs":
           {
                "workDesc_111":
                {
                     "terms":
                     {
                          "field":"workDesc.suggestions",
                           "include":"civil.*",
                           "order":[
                           {
                                "_count":"desc"
                           }],
                           "size":10
                     }
                }
           }
           ,"size":0,
          "_source":false
    }

```

1. Auto complete functionality work like charm (result coming in millisecond) when documents are 0.1 million, But I got time out error when documents are 20 million.

2. Elasticsearch server configuration is

Is there any other options to get the same result? Please advise

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 1, 2021, 7:19am UTC](https://discuss.elastic.co/t/alternative-option-of-shingle-token-filter/285787/2 "2021-11-01T07:19:57Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
