# How to handle special characters in span\_near query?

**URL:** <https://discuss.elastic.co/t/how-to-handle-special-characters-in-span-near-query/363119>\
**Category:** Elastic Search\
**Created:** [July 15, 2024, 6:21am UTC](https://discuss.elastic.co/t/how-to-handle-special-characters-in-span-near-query/363119 "2024-07-15T06:21:29Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![shubham\_gupta4](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shubham_gupta4/32/126062_2.png) [@shubham\_gupta4](https://discuss.elastic.co/u/shubham_gupta4)\
**Post date:** [July 15, 2024, 6:21am UTC](https://discuss.elastic.co/t/how-to-handle-special-characters-in-span-near-query/363119/1 "2024-07-15T06:21:29Z")

</div>

Hi Team,  
I'm having very huge boolean query which contains NEAR operator with more than 80k phrase limit, so I'm using span\_near to handle NEAR operator in my boolean query, but I can see that if I pass keyword with special characters, it simply getting ignored by elasticsearch, please can anyone help me out on how to handle special characters in span\_near query?  
Example:

```auto

{
 "span_near": {
              "clauses": [
                  {
                      "span_or": {
                          "clauses": [
                              {
                                  "span_term": {
                                      "body": "covid*"
                                  }
                              }
                          ]
                      }
                  },
                  {
                      "span_or": {
                          "clauses": [
                              {
                                  "span_term": {
                                      "body": "9-8-8"
                                  }
                              }
                          ]
                      }
                  }
              ],
              "slop": "10",
              "in_order": false
          }
      }

```

---

<div class="post-metadata">

**Author:** ![Kathleen\_DeRusso](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kathleen_derusso/32/132039_2.png) [@Kathleen\_DeRusso](https://discuss.elastic.co/u/Kathleen_DeRusso)\
**Post date:** [July 15, 2024, 12:10pm UTC](https://discuss.elastic.co/t/how-to-handle-special-characters-in-span-near-query/363119/2 "2024-07-15T12:10:41Z")

</div>

Hey there, special characters are getting stripped during tokenization. You can customize this by specifying the [tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-tokenizers.html) you want to use, creating a [custom analyzer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-custom-analyzer.html) and reindexing your data. Hope that helps!

---

<div class="post-metadata">

**Author:** ![shubham\_gupta4](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shubham_gupta4/32/126062_2.png) [@shubham\_gupta4](https://discuss.elastic.co/u/shubham_gupta4)\
**Post date:** [July 15, 2024, 12:29pm UTC](https://discuss.elastic.co/t/how-to-handle-special-characters-in-span-near-query/363119/3 "2024-07-15T12:29:13Z")

</div>

Hey Kathleen,  
Thanks for the quick reply, really appreciated!!  
I already have some custom analyzers indexed on my data, they are working fine with normal query i.e. if I specify them in default fields inside query string then its working and returning matching results.  
Please can you share syntax or any other way to implement it on span\_near query?  
Thanks in Advance!!

---

<div class="post-metadata">

**Author:** ![Kathleen\_DeRusso](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kathleen_derusso/32/132039_2.png) [@Kathleen\_DeRusso](https://discuss.elastic.co/u/Kathleen_DeRusso)\
**Post date:** [July 15, 2024, 2:53pm UTC](https://discuss.elastic.co/t/how-to-handle-special-characters-in-span-near-query/363119/4 "2024-07-15T14:53:45Z")

</div>

Sure, here's a quick demo script using the `whitespace` analyzer - this will not do anything but tokenize on whitespace, but will return a matching document you want using the query you provided. You can play with this script using additional data and different analyzers to see what works for your data and use case.

```auto
PUT my-span-test
{
  "mappings": {
    "properties": {
      "body": {
        "type": "text",
        "analyzer": "whitespace"
      }
    }
  }
}

PUT my-span-test/_doc/1
{
  "body": "covid* 9-8-8"
}

POST my-span-test/_search
{
  "query": {
    "span_near": {
      "clauses": [
        {
          "span_or": {
            "clauses": [
              {
                "span_term": {
                  "body": "covid*"
                }
              }
            ]
          }
        },
        {
          "span_or": {
            "clauses": [
              {
                "span_term": {
                  "body": "9-8-8"
                }
              }
            ]
          }
        }
      ],
      "slop": "10",
      "in_order": false
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![shubham\_gupta4](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shubham_gupta4/32/126062_2.png) [@shubham\_gupta4](https://discuss.elastic.co/u/shubham_gupta4)\
**Post date:** [July 16, 2024, 7:46am UTC](https://discuss.elastic.co/t/how-to-handle-special-characters-in-span-near-query/363119/5 "2024-07-16T07:46:26Z")

</div>

Thanks for the demo!!

Custom analyzer is only working if I'm specifying it  
For example: I've created a custom analyzer for special characters i.e. cs\_special\_characters

It'll only work if I specify it:

```auto
"query": {
    "query_string": {
        "default_field": "body.cs_special_characters",
        "query": "(covid* AND 9-8-8)",
        "analyzer":"whitespace"
    }
}

```

But here in span\_near we are not specifying it anywhere and if I give body.cs\_special\_characters inside span\_term instead of body then it is throwing error

```auto
{
  "query": {
    "span_near": {
      "clauses": [
        {
          "span_or": {
            "clauses": [
              {
                "span_term": {
                  "body": "covid*"
                }
              }
            ]
          }
        },
        {
          "span_or": {
            "clauses": [
              {
                "span_term": {
                  "body": "9-8-8"
                }
              }
            ]
          }
        }
      ],
      "slop": "10",
      "in_order": false
    }
  }
}

```

Is there any syntax or way by which we can specify to use particular analyzer inside span\_near query?  
Or if you have any other alternative to handle NEAR other than span\_near?

---

<div class="post-metadata">

**Author:** ![Kathleen\_DeRusso](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kathleen_derusso/32/132039_2.png) [@Kathleen\_DeRusso](https://discuss.elastic.co/u/Kathleen_DeRusso)\
**Post date:** [July 16, 2024, 12:16pm UTC](https://discuss.elastic.co/t/how-to-handle-special-characters-in-span-near-query/363119/6 "2024-07-16T12:16:58Z")

</div>

You can try an index analyzer and reindexing.

Alternately you could see if a [match phrase](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-match-query-phrase.html) query with a very high slop would work for your needs.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 13, 2024, 12:17pm UTC](https://discuss.elastic.co/t/how-to-handle-special-characters-in-span-near-query/363119/7 "2024-08-13T12:17:35Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
