# POST analyzed, Index

**URL:** <https://discuss.elastic.co/t/post-analyzed-index/125718>\
**Category:** Elasticsearch\
**Created:** [March 27, 2018, 9:15am UTC](https://discuss.elastic.co/t/post-analyzed-index/125718 "2018-03-27T09:15:12Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Varan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/varan/32/32120_2.png) [@Varan](https://discuss.elastic.co/u/Varan)\
**Post date:** [March 27, 2018, 9:15am UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/1 "2018-03-27T09:15:12Z")

</div>

Hi,  
Sorry for this basic question, but I can't find any sense to it.

I have an analyzer and wish it to analyze the data I index.  
But when I do this:

```
POST my_index/_analyze
{
  "analyzer": "my_analyzer",
  "text": "This is a test sentence"
}
GET my_index/_search?q=*

```

I got this:  
{  
"took": 1,  
"timed\_out": false,  
"\_shards": {  
"total": 5,  
"successful": 5,  
"skipped": 0,  
"failed": 0  
},  
"hits": {  
"total": 0,  
"max\_score": null,  
"hits": []  
}  
Why does my query not return anything?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 27, 2018, 11:40am UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/2 "2018-03-27T11:40:18Z")

</div>

`_analyze` just analyze the text and does not index anything.

---

<div class="post-metadata">

**Author:** ![Varan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/varan/32/32120_2.png) [@Varan](https://discuss.elastic.co/u/Varan)\
**Post date:** [March 27, 2018, 12:10pm UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/3 "2018-03-27T12:10:19Z")

</div>

Merci, @dadoonet  
Mais comment puis-je indexer quelque chose en l'analysant au préalable?  
Je ne vois pas l'intérêt de la fonction analyze si le résultat n'est pas stocker.

Thanks, but how could I index something while analyze it beforehand?  
I don't see the point of the analyze function if the result goes to waste.

---

<div class="post-metadata">

**Author:** ![Varan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/varan/32/32120_2.png) [@Varan](https://discuss.elastic.co/u/Varan)\
**Post date:** [March 27, 2018, 12:29pm UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/4 "2018-03-27T12:29:29Z")

</div>

I mean when I do this I can index but the analyzer doesn't work:

```
POST my_index/_analyze
{
  "analyzer": "my_analyzer",
  "field" : "this_field",
  "text": "This is a test sentence"
}
```

---

<div class="post-metadata">

**Author:** ![val](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/val/32/138203_2.png) [@val](https://discuss.elastic.co/u/val)\
**Post date:** [March 27, 2018, 12:39pm UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/5 "2018-03-27T12:39:40Z")

</div>

> [@Varan](#):
>
> I don't see the point of the analyze function if the result goes to waste.

To test how your analyzer works before actually indexing stuff and potentially messing up your index?

---

<div class="post-metadata">

**Author:** ![Varan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/varan/32/32120_2.png) [@Varan](https://discuss.elastic.co/u/Varan)\
**Post date:** [March 27, 2018, 12:50pm UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/6 "2018-03-27T12:50:12Z")

</div>

@val Thanks I didn't think of that since I just test whatever I want on an index I invent beforehand.

But do you have an example of a POST query that analyze and index something?  
It would be very helpfull.

---

<div class="post-metadata">

**Author:** ![val](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/val/32/138203_2.png) [@val](https://discuss.elastic.co/u/val)\
**Post date:** [March 27, 2018, 1:25pm UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/7 "2018-03-27T13:25:31Z")

</div>

If you want to index your data you can simply POST/PUT it into your index, provided that `this_field` is properly mapped with your analyzer

So first create your index with the mapping:

```
PUT my-index
{
    "mappings": {
        "doc": {
            "properties": {
               "this_field": {
                   "type": "text",
                   "analyzer": "my_analyzer"
               }
            }
        }
    }
}

```

And then index your data

```
PUT my-index/doc/1
{
    "this_field": "This is a test sentence"
}
```

---

<div class="post-metadata">

**Author:** ![Varan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/varan/32/32120_2.png) [@Varan](https://discuss.elastic.co/u/Varan)\
**Post date:** [March 27, 2018, 3:03pm UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/9 "2018-03-27T15:03:05Z")

</div>

Thanks again @val , but I don't seem to make it work.  
Could I show it to you?  
I define a token filter to clean a phone number, and put a keyword tokenizer to let the whole entry in a single term.

```
    PUT phone
    {
      "settings": {
        "analysis": {
           "char_filter": {
            "my_char_filter": {
              "type": "mapping",
              "mappings": [
                    "( => ",
                    ") => ",
                    ", => ",
                    ". => ",
                    "; => ",
                    "\\u0020 => ",
                    "+ => 00"
              ]
            }
          },
          "analyzer": {
            "one": {
              "tokenizer": "keyword",
              "char_filter": [
                "my_char_filter"]
            }
          }
        }
      },
      "mapping": {
        "_doc": {
          "properties": {
            "phone_number": {
              "type": "text",
              "analyzer": "one"
            }
          }
        }
      }
    }

PUT /phone/doc/1
{         
  "phone_number": "077 , 1.436;25 "
}

```

With this, the number is indexed but no filter have been made on it.

The goal would be to index "077143625" instead of "077 , 1.436;25 ".

PS : \u0020 is for space

---

<div class="post-metadata">

**Author:** ![val](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/val/32/138203_2.png) [@val](https://discuss.elastic.co/u/val)\
**Post date:** [March 27, 2018, 3:10pm UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/10 "2018-03-27T15:10:32Z")

</div>

There's a much simpler way to do it using a `pattern_replace` character filter and removing all non-digit characters. Try it out

```
            "digit_only": {
                "type": "pattern_replace",
                "pattern": "\\D+",
                "replacement": ""
            },

```

Also note that the value you have in your source will never be changed, i.e. the source will still contain `077 , 1.436;25` even though `077143625` is indexed.

---

<div class="post-metadata">

**Author:** ![Varan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/varan/32/32120_2.png) [@Varan](https://discuss.elastic.co/u/Varan)\
**Post date:** [March 28, 2018, 8:48am UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/11 "2018-03-28T08:48:51Z")

</div>

Thanks @val, but I prefer my filter since it allow to the exception of this rule: "+ =\> 00".  
But still, when I run this:

```
PUT /phone/doc/1
{         
  "phone_number": "0777 , 1.436;25 "
}

```

I manage to index something unanalyzed even thought the field is mapped with the correct analyzer.  
And when I run this:

```
POST /phone/_analyze
{         
  "analyzer":"one",
  "text": "0777 , 1.436;25 "
}

```

It return an analyzed answer but it's not indexed.

I'm sorry to continue to bother you.

---

<div class="post-metadata">

**Author:** ![val](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/val/32/138203_2.png) [@val](https://discuss.elastic.co/u/val)\
**Post date:** [March 28, 2018, 9:30am UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/12 "2018-03-28T09:30:36Z")

</div>

> [@Varan](#):
>
> but I prefer my filter since it allow to the exception of this rule: "+ =\> 00".

Fair enough

At this point, please show your index settings and mappings, i.e. what you get when running:

```
GET phone

```

---

<div class="post-metadata">

**Author:** ![Varan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/varan/32/32120_2.png) [@Varan](https://discuss.elastic.co/u/Varan)\
**Post date:** [March 28, 2018, 9:37am UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/13 "2018-03-28T09:37:20Z")

</div>

```
{
  "phone": {
    "aliases": {},
    "mappings": {
      "doc": {
        "properties": {
          "phone_number": {
            "type": "text",
            "fields": {
              "keyword": {
                "type": "keyword",
                "ignore_above": 256
              }
            }
          }
        }
      }
    },
    "settings": {
      "index": {
        "number_of_shards": "5",
        "provided_name": "phone",
        "creation_date": "***",
        "analysis": {
          "analyzer": {
            "one": {
              "char_filter": [
                "my_char_filter"
              ],
              "tokenizer": "two"
            }
          },
          "char_filter": {
            "my_char_filter": {
              "type": "mapping",
              "mappings": [
                "( => ",
                ") => ",
                ", => ",
                ". => ",
                "; => ",
                "\\u0020 => ",
                "+ => 00"
              ]
            }
          },
          "tokenizer": {
            "two": {
              "type": "keyword",
              "min_gram": "3",
              "max_gram": "4"
            }
          }
        },
        "number_of_replicas": "1",
        "uuid": "c-XyKl__Q9GTF2bkNqdgbQ",
        "version": {
          "created": "6020299"
        }
      }
    }
  }
}

```

Here the whole thing, I just added the ngram tokenizer, but that's out of the subject.

---

<div class="post-metadata">

**Author:** ![val](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/val/32/138203_2.png) [@val](https://discuss.elastic.co/u/val)\
**Post date:** [March 28, 2018, 9:52am UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/14 "2018-03-28T09:52:36Z")

</div>

The `phone_number` field doesn't have any analyzer set, which could explain the problem.

Change your mapping to this instead:

```
      "phone_number": {
        "type": "text",
        "analyzer": "one", <--- add this line
        "fields": {
          "keyword": {
            "type": "keyword",
            "ignore_above": 256
          }
        }
      }
```

---

<div class="post-metadata">

**Author:** ![Varan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/varan/32/32120_2.png) [@Varan](https://discuss.elastic.co/u/Varan)\
**Post date:** [March 28, 2018, 10:07am UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/15 "2018-03-28T10:07:57Z")

</div>

Please ignore the precedent message, I messed up because I was putting in the ngram tokenizer and then going back to keyword to keep it simple for you to read it. But I just put a mix of both in the end, which doesn't make sense.  
Here the real thing:

```
{
  "phone": {
    "aliases": {},
    "mappings": {},
    "settings": {
      "index": {
        "number_of_shards": "5",
        "provided_name": "phone",
        "creation_date": "***",
        "analysis": {
          "analyzer": {
            "one": {
              "char_filter": [
                "my_char_filter"
              ],
              "tokenizer": "two"
            }
          },
          "char_filter": {
            "my_char_filter": {
              "type": "mapping",
              "mappings": [
                "( => ",
                ") => ",
                ", => ",
                ". => ",
                "; => ",
                "\\u0020 => ",
                "+ => 00"
              ]
            }
          },
          "tokenizer": {
            "two": {
              "type": "ngram",
              "min_gram": "3",
              "max_gram": "4"
            }
          }
        },
        "number_of_replicas": "1",
        "uuid": "a6xJMLUOQAeBbI7wFKQYcQ",
        "version": {
          "created": "6020299"
        }
      }
    }
  }
}

```

I notice that the mapping is empty, and if i PUT this query:

```
PUT /phone/doc/1
{         
  "phone_number": "0777 , 1.436;25 "
}

```

then the `GET phone` returned a mapping that was added dynamically.

```
{
  "phone": {
    "aliases": {},
    "mappings": {
      "doc": {
        "properties": {
          "phone_number": {
            "type": "text",
            "fields": {
              "keyword": {
                "type": "keyword",
                "ignore_above": 256
              }
            }
          }
        }
      }
    },
    "settings": {
      "index": {
        "number_of_shards": "5",
        "provided_name": "phone",
        "creation_date": "***",
        "analysis": {
          "analyzer": {
            "one": {
              "char_filter": [
                "my_char_filter"
              ],
              "tokenizer": "two"
            }
          },
          "char_filter": {
            "my_char_filter": {
              "type": "mapping",
              "mappings": [
                "( => ",
                ") => ",
                ", => ",
                ". => ",
                "; => ",
                "\\u0020 => ",
                "+ => 00"
              ]
            }
          },
          "tokenizer": {
            "two": {
              "type": "ngram",
              "min_gram": "3",
              "max_gram": "4"
            }
          }
        },
        "number_of_replicas": "1",
        "uuid": "a6xJMLUOQAeBbI7wFKQYcQ",
        "version": {
          "created": "6020299"
        }
      }
    }
  }
}

```

**Thanks again for all of your answers.**

---

<div class="post-metadata">

**Author:** ![val](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/val/32/138203_2.png) [@val](https://discuss.elastic.co/u/val)\
**Post date:** [March 28, 2018, 10:59am UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/16 "2018-03-28T10:59:48Z")

</div>

There's no mappings in the first code snippet and in the third one there's still no analyzer set on the `phone_number` field. You need to set the mapping yourself, if you let ES create it for you, there's no way it would know to apply the analyzer to your field.

---

<div class="post-metadata">

**Author:** ![Varan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/varan/32/32120_2.png) [@Varan](https://discuss.elastic.co/u/Varan)\
**Post date:** [March 28, 2018, 12:50pm UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/17 "2018-03-28T12:50:28Z")

</div>

Yeah, I saw that there was no mapping recognized when i `GET phone` but one was hardcoded nervertheless. I didn't let ES do the mapping for me, I just saw that it ignore my mapping and put one by itself.

Anyway, I broke my index in two part, one where I put the settings filled with the char\_filter and the tokenizer composing the analyzer, and one with the mapping (in this order, it doesn't work the other way since I call an analyzer that's not declared yet, in the mapping).

And the `GET phone` is good this time. It show my mapping. I don't know why it didn't appear earlier.  
But still; my data won't get analyzed.  
Here the split index :

```
PUT phone
{ "settings": {
    "analysis": {
       "char_filter": {
        "my_char_filter": {
          "type": "mapping",
          "mappings": [
                "( => ",
                ") => ",
                ", => ",
                ". => ",
                "; => ",
                "\\u0020 => ",
                "+ => 00"
          ]
        }
      },
       "tokenizer": {
        "two": {
          "type": "ngram",
          "min_gram": 3,
          "max_gram": 4
        }
      },
      "analyzer": {
        "one": {
          "tokenizer": "two",
          "char_filter": [
            "my_char_filter"]
        }
      }
    }
  }}
PUT phone/_mapping/_doc
{
      "properties": {
        "phone_number": {
          "type": "text",
          "analyzer": "one"},
        "favoris": {
          "type":"boolean"}}
}

```

and the answer for the `GET phone` command:

```
{
  "phone": {
    "aliases": {},
    "mappings": {
      "_doc": {
        "properties": {
          "favoris": {
            "type": "boolean"
          },
          "phone_number": {
            "type": "text",
            "analyzer": "one"
          }
        }
      }
    },
    "settings": {
      "index": {
        "number_of_shards": "5",
        "provided_name": "phone",
        "creation_date": "1522240891958",
        "analysis": {
          "analyzer": {
            "one": {
              "char_filter": [
                "my_char_filter"
              ],
              "tokenizer": "two"
            }
          },
          "char_filter": {
            "my_char_filter": {
              "type": "mapping",
              "mappings": [
                "( => ",
                ") => ",
                ", => ",
                ". => ",
                "; => ",
                "\\u0020 => ",
                "+ => 00"
              ]
            }
          },
          "tokenizer": {
            "two": {
              "type": "ngram",
              "min_gram": "3",
              "max_gram": "4"
            }
          }
        },
        "number_of_replicas": "1",
        "uuid": "tEKKrMe7Qu2TKXlPzzCm-Q",
        "version": {
          "created": "6020299"
        }
      }
    }
  }
}

```

I don't think I'm far from it, I just probably missed a part where I should better declare things.

---

<div class="post-metadata">

**Author:** ![val](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/val/32/138203_2.png) [@val](https://discuss.elastic.co/u/val)\
**Post date:** [March 28, 2018, 12:54pm UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/18 "2018-03-28T12:54:02Z")

</div>

After modifying the mapping to include the analyzer, you need to reindex your document in order to analyze it again using the new analyzer.

---

<div class="post-metadata">

**Author:** ![Varan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/varan/32/32120_2.png) [@Varan](https://discuss.elastic.co/u/Varan)\
**Post date:** [March 28, 2018, 1:08pm UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/19 "2018-03-28T13:08:03Z")

</div>

I think it work! 😀

So for the posterity, splitting the mapping and the settings in two differents index made it work.  
Wrong:

```
PUT phone
{
  "mapping": {
    "_doc": {
      "properties": {
        "phone_number": {
          "type": "text",
          "analyzer": "one"},
        "favoris": {
          "type":"boolean"}
        }
      }
    },
  "settings": {
    "analysis": {
       "char_filter": {
        "my_char_filter": {
          "type": "mapping",
          "mappings": [
                "( => ",
                ") => ",
                ", => ",
                ". => ",
                "; => ",
                "\\u0020 => ",
                "+ => 00"
          ]
        }
      },
       "tokenizer": {
        "two": {
          "type": "ngram",
          "min_gram": 3,
          "max_gram": 4
        }
      },
      "analyzer": {
        "one": {
          "tokenizer": "two",
          "char_filter": [
            "my_char_filter"]
        }
      }
    }
  }
}

```

Good:

```
PUT phone
{
    "settings": {
    "analysis": {
       "char_filter": {
        "my_char_filter": {
          "type": "mapping",
          "mappings": [
                "( => ",
                ") => ",
                ", => ",
                ". => ",
                "; => ",
                "\\u0020 => ",
                "+ => 00"
          ]
        }
      },
       "tokenizer": {
        "two": {
          "type": "ngram",
          "min_gram": 3,
          "max_gram": 4
        }
      },
      "analyzer": {
        "one": {
          "tokenizer": "two",
          "char_filter": [
            "my_char_filter"]
        }
      }
    }
  }
}
PUT phone/_mapping/_doc
{
      "properties": {
        "phone_number": {
          "type": "text",
          "analyzer": "one"},
        "favoris": {
          "type":"boolean"}
        }
}

```

So when i

```
PUT /phone/_doc/1
{         
  "phone_number": "0777 , 1.436;25 "
}

```

And then :

`GET phone/_search?q=07771`

It returned the original number "0777 , 1.436;25 ".

Thanks @val for the help and help me to better declare the mapping!

---

<div class="post-metadata">

**Author:** ![val](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/val/32/138203_2.png) [@val](https://discuss.elastic.co/u/val)\
**Post date:** [March 28, 2018, 1:18pm UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/21 "2018-03-28T13:18:58Z")

</div>

Awesome, glad it worked. You can make it work in one single call, though, you just must not index a document before installing the mapping. You had a typo `mapping` should read `mappings`:

This would work:

```
PUT phone
{
  "mappings": { <--- there was a typo here in your last post
    "_doc": {
      "properties": {
        "phone_number": {
          "type": "text",
          "analyzer": "one"
        },
        "favoris": {
          "type": "boolean"
        }
      }
    }
  },
  "settings": {
    "analysis": {
      "char_filter": {
        "my_char_filter": {
          "type": "mapping",
          "mappings": [
            "( => ",
            ") => ",
            ", => ",
            ". => ",
            "; => ",
            "\\u0020 => ",
            "+ => 00"
          ]
        }
      },
      "tokenizer": {
        "two": {
          "type": "ngram",
          "min_gram": 3,
          "max_gram": 4
        }
      },
      "analyzer": {
        "one": {
          "tokenizer": "two",
          "char_filter": [
            "my_char_filter"
          ]
        }
      }
    }
  }
}

```

Then only index your document

```
PUT /phone/_doc/1
{         
  "phone_number": "0777 , 1.436;25 "
}
```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 25, 2018, 1:19pm UTC](https://discuss.elastic.co/t/post-analyzed-index/125718/22 "2018-04-25T13:19:14Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
