# What is default index analyzer?

**URL:** <https://discuss.elastic.co/t/what-is-default-index-analyzer/69402>\
**Category:** Elasticsearch\
**Created:** [December 19, 2016, 7:44am UTC](https://discuss.elastic.co/t/what-is-default-index-analyzer/69402 "2016-12-19T07:44:04Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Youxu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/youxu/32/117375_2.png) [@Youxu](https://discuss.elastic.co/u/Youxu)\
**Post date:** [December 19, 2016, 7:44am UTC](https://discuss.elastic.co/t/what-is-default-index-analyzer/69402/1 "2016-12-19T07:44:04Z")

</div>

I though default analyzer is "standard" analyzer, but per my following experimentation, seems not.

1. Create index with customized standard analyzer which included a pattern\_capture filter to split words by "." or "\_"

```auto
POST / myindex
{
      "settings" : {
        "analysis" : {
          "filter" : {
            "customsplit" : {
              "type" : "pattern_capture",
              "preserve_original" : 1,
              "patterns" : [
                "([^_.]+)"
              ]
            }
          },
          "analyzer" : {
            "standard" : {
              "tokenizer" : "standard",
              "filter" : [
                "lowercase",
                "customsplit"
              ]
            }
          }
        }
      },
      "mappings" : {
        "docs" : {
          "properties" : {
            "Url" : {
              "type" : "string"
            }
          }
        }
      }
    }

```

1. Insert one doc to myindex

```auto
POST /myindex/docs/1
{
	"Url": "www.xyz.com"
}

```

Per \_analyze API, the standard analyzer used by myindex DOES split the word by "."

```auto
GET /myindex/_analyze?analyzer=standard&text=www.xyz.com

```

output:

```auto
{
      "tokens": [
        {
          "token": "www.xyz.com",
          "start_offset": 0,
          "end_offset": 11,
          "type": "<ALPHANUM>",
          "position": 1
        },
        {
          "token": "www",
          "start_offset": 0,
          "end_offset": 11,
          "type": "<ALPHANUM>",
          "position": 1
        },
        {
          "token": "xyz",
          "start_offset": 0,
          "end_offset": 11,
          "type": "<ALPHANUM>",
          "position": 1
        },
        {
          "token": "com",
          "start_offset": 0,
          "end_offset": 11,
          "type": "<ALPHANUM>",
          "position": 1
        }
      ]
    }

```

BUT, the problem is, if I search "xyz" from myindex, nothing returned:

```auto
POST /myindex/_search
{
 "query": {
  "match": {
    "Url": "xyz"
  }
 }
}

```

BUT, if I explicitly set the analyzer to "standard" in index mapping:

```auto
mappings": {
  "docs": {
    "properties": {
      "Url": {
        "type": "string",
        "analyzer": "standard"
      }
    }
  }

```

Then searching "xyz" can return the documents.

SO my question is: Is "standard" really default analyzer of ES index? if NOT, how to set default analyzer?

Or anything wrong in my above testing steps, if standard is indeed the default analyzer?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 19, 2016, 8:04am UTC](https://discuss.elastic.co/t/what-is-default-index-analyzer/69402/2 "2016-12-19T08:04:38Z")

</div>

Please format your code using `</>` icon as explained in [this guide](https://discuss.elastic.co/t/about-the-elasticsearch-category/21). It will make your post more readable.

The standard analyzer is explained here: [https://www.elastic.co/guide/en/elasticsearch/reference/5.1/analysis-standard-analyzer.html](https://www.elastic.co/guide/en/elasticsearch/reference/5.1/analysis-standard-analyzer.html)

As per [https://www.elastic.co/guide/en/elasticsearch/reference/5.1/analysis.html#\_specifying\_an\_index\_time\_analyzer](https://www.elastic.co/guide/en/elasticsearch/reference/5.1/analysis.html#_specifying_an_index_time_analyzer), the default analyzer is the standard one.

---

<div class="post-metadata">

**Author:** ![Youxu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/youxu/32/117375_2.png) [@Youxu](https://discuss.elastic.co/u/Youxu)\
**Post date:** [December 19, 2016, 8:47am UTC](https://discuss.elastic.co/t/what-is-default-index-analyzer/69402/3 "2016-12-19T08:47:20Z")

</div>

Thanks David.  
Yes, per ES reference, "standard" analyzer should be default analyzer.

But then, anything wrong with my testing? Does the default "standard" used just mean the built-in standard analyzer, but not the customized "standard" as defined in my setting?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 19, 2016, 9:55am UTC](https://discuss.elastic.co/t/what-is-default-index-analyzer/69402/4 "2016-12-19T09:55:13Z")

</div>

That's interesting.

Indeed, you can't here "overwrite" the `standard` analyzer which is built in elasticsearch.

The proper way to solve your issue for now is to do something like:

```auto
DELETE myindex
PUT myindex
{
  "settings": {
    "analysis": {
      "filter": {
        "customsplit": {
          "type": "pattern_capture",
          "preserve_original": 1,
          "patterns": [
            "([^_.]+)"
          ]
        }
      },
      "analyzer": {
        "my_analyzer": {
          "tokenizer": "standard",
          "filter": [
            "lowercase",
            "customsplit"
          ]
        }
      }
    }
  },
  "mappings": {
    "docs": {
      "properties": {
        "Url": {
          "type": "string",
          "analyzer": "my_analyzer"
        }
      }
    }
  }
}
PUT /myindex/docs/1
{
	"Url": "www.xyz.com"
}
GET /myindex/_analyze?analyzer=my_analyzer&text=www.xyz.com
POST /myindex/_search
{
 "query": {
  "match": {
    "Url": "xyz"
  }
 }
}

```

May be open an issue on github and refer to this ticket? I think that we should either reject that you are using `standard` as an analyzer name or pick the right one when running `_search`. Here I think we are using the built-in one at search time instead of the one which is defined within your index.

Thanks for reporting!

---

<div class="post-metadata">

**Author:** ![Youxu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/youxu/32/117375_2.png) [@Youxu](https://discuss.elastic.co/u/Youxu)\
**Post date:** [December 19, 2016, 1:04pm UTC](https://discuss.elastic.co/t/what-is-default-index-analyzer/69402/5 "2016-12-19T13:04:32Z")

</div>

Thanks David.  
Yes, I can resolve this issue by explicitly setting the "standard" analyzer to "Url" field.

This issue seems to me a bug of Elasticsearch. As you said, ES should either reject "standard" as an analyzer name in customized analyzer setting or pick the right one when running \_search.

I will open this issue on GitHub, if not opened yet.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 16, 2017, 1:04pm UTC](https://discuss.elastic.co/t/what-is-default-index-analyzer/69402/6 "2017-01-16T13:04:50Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
