# How to replace multiple spaces for the field value before creating the document using pattern analyzers

**URL:** <https://discuss.elastic.co/t/how-to-replace-multiple-spaces-for-the-field-value-before-creating-the-document-using-pattern-analyzers/360658>\
**Category:** Elastic Search\
**Created:** [June 1, 2024, 11:43am UTC](https://discuss.elastic.co/t/how-to-replace-multiple-spaces-for-the-field-value-before-creating-the-document-using-pattern-analyzers/360658 "2024-06-01T11:43:33Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Neeraj10](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/neeraj10/32/125404_2.png) [@Neeraj10](https://discuss.elastic.co/u/Neeraj10)\
**Post date:** [June 1, 2024, 11:43am UTC](https://discuss.elastic.co/t/how-to-replace-multiple-spaces-for-the-field-value-before-creating-the-document-using-pattern-analyzers/360658/1 "2024-06-01T11:43:33Z")

</div>

```auto
PUT candidate-index
{
  "settings": {
  "index": {
    "max_ngram_diff": 20,
    "number_of_replicas": 0,
    "number_of_shards": 1
  },
  "analysis": {
    "normalizer": {
      "lowercase_normalize": {
        "type": "custom",
        "filter": ["lowercase"]
      }
    },
    "tokenizer": {
      "grams_tokenizer": {
        "type": "ngram",
        "min_gram": 2,
        "max_gram": 20,
        "token_chars": ["letter", "digit", "punctuation", "symbol"]
      },
      "phone_grams_tokenizer": {
        "type": "ngram",
        "min_gram": 2,
        "max_gram": 3,
        "token_chars": ["digit"]
      },
      "search_keywords_tokenizer": {
        "type": "char_group",
        "tokenize_on_chars": ["whitespace", ","]
      }
    },
    "filter": {
      "non_space_tokens_filter": {
        "type": "pattern_capture",
        "preserve_original": true,
        "patterns": ["([\\d\\w]+)", "([^\\w]+)", "([^\\W_]+)"]
      },
      "email_filter": {
        "type": "pattern_capture",
        "preserve_original": true,
        "patterns": ["([^@]+)", "(\\p{L}+)", "(\\d+)", "@(.+)", "([^-@]+)"]
      },
      "rm_empty_string_tokens": {
        "type": "length",
        "min": 3,
        "max": 30
      }
    },
    "char_filter": {
      "rm_non_digits_char_filter": {
        "type": "pattern_replace",
        "pattern": "([^0-9]+)",
        "replacement": ""
      }
    },
    "analyzer": {
      "default": {
        "tokenizer": "whitespace",
        "type": "custom",
        "filter": ["non_space_tokens_filter", "lowercase", "unique"]
      },
      "default_search": {
        "filter": ["lowercase"],
        "type": "custom",
        "tokenizer": "search_keywords_tokenizer"
      },
      "email_analyzer": {
        "type": "custom",
        "tokenizer": "uax_url_email",
        "filter": ["email_filter", "lowercase", "unique"]
      },
      "search_analyzer": {
        "type": "custom",
        "filter": ["lowercase"],
        "tokenizer": "search_keywords_tokenizer"
      },
      "grams_analyzer": {
        "type": "custom",
        "filter": ["lowercase"],
        "tokenizer": "grams_tokenizer"
      },
      "phone_grams_analyzer": {
        "type": "custom",
        "char_filter": ["rm_non_digits_char_filter"],
        "tokenizer": "phone_grams_tokenizer"
      },
      "phone_keyword_analyzer": {
        "type": "custom",
        "char_filter": ["rm_non_digits_char_filter"],
        "filter": ["rm_empty_string_tokens"],
        "tokenizer": "keyword"
      }
    }
  }
},
  "mappings": {
    "dynamic": "false",
    "properties": {
      "currentlocation": {
        "type": "text",
        "fields": {
          "keyword": {
            "type": "keyword",
            "ignore_above": 256
          },
          "lowercase_keyword": {
            "type": "keyword",
            "normalizer": "lowercase_normalize",
            "ignore_above": 256
          }
        }
      }
    }
  }
}

```

There is an existing index with the mentioned settings and there is a field named "currentlocation" with the mentioned mapping. Now I want to use pattern analyzer and trim the currentlocation field value before storing a document. Need a help with usage of pattern analyzer and not to affect any existing mapping of the field

Example:

If I'm creating a document as mentioned below

```auto
POST candidate-index/_doc/1
{
  "currentlocation": " Hosur, Bangalore, India "
}

```

It should be stored in the index as mentioned below

```auto
{
  "took" : 0,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 1,
      "relation" : "eq"
    },
    "max_score" : 1.0,
    "hits" : [
      {
        "_index" : "candidate-index",
        "_type" : "_doc",
        "_id" : "1",
        "_score" : 1.0,
        "_source" : {
          "currentlocation" : "Hosur, Bangalore, India"
        }
      }
    ]
  }
}

```

I want to replace the multiple spaces in the string with single space.

Case 2:  
Say document stored as mentioned below:

```auto
POST candidate-index/_doc/2
{
  "currentlocation": "Qatar , As ia "
}

```

It should be stored as mentioned below

```auto
{
  "took" : 0,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 1,
      "relation" : "eq"
    },
    "max_score" : 1.0,
    "hits" : [
      {
        "_index" : "candidate-index",
        "_type" : "_doc",
        "_id" : "1",
        "_score" : 1.0,
        "_source" : {
          "currentlocation" : "Qatar , Asia"
        }
      }
    ]
  }
}

```

---

<div class="post-metadata">

**Author:** ![Artem\_Shelkovnikov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/artem_shelkovnikov/32/134322_2.png) [@Artem\_Shelkovnikov](https://discuss.elastic.co/u/Artem_Shelkovnikov)\
**Post date:** [June 3, 2024, 9:22am UTC](https://discuss.elastic.co/t/how-to-replace-multiple-spaces-for-the-field-value-before-creating-the-document-using-pattern-analyzers/360658/2 "2024-06-03T09:22:58Z")

</div>

Do you have control over the code that you use to ingest the data into Elasticsearch? If so, it's probably better to do normalization/trimming before ingesting data.

Additionally - have you seen this post: [Keyword analyzer but allow redundant white spaces](https://discuss.elastic.co/t/keyword-analyzer-but-allow-redundant-white-spaces/110441)?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 1, 2024, 9:23am UTC](https://discuss.elastic.co/t/how-to-replace-multiple-spaces-for-the-field-value-before-creating-the-document-using-pattern-analyzers/360658/3 "2024-07-01T09:23:21Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
