# Can't get stop words working

**URL:** <https://discuss.elastic.co/t/cant-get-stop-words-working/9828>\
**Category:** Elasticsearch\
**Created:** [November 26, 2012, 7:50pm UTC](https://discuss.elastic.co/t/cant-get-stop-words-working/9828 "2012-11-26T19:50:36Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![racedo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/racedo/32/2499_2.png) [@racedo](https://discuss.elastic.co/u/racedo)\
**Post date:** [November 26, 2012, 7:50pm UTC](https://discuss.elastic.co/t/cant-get-stop-words-working/9828/1 "2012-11-26T19:50:36Z")

</div>

I'm trying to add some stopwords to the default settings that haystack is  
using and the settings look like this (added "esto", "que" and "de" just  
for testing purposes:

```
    'settings': {
        "analysis": {
            "analyzer": {
                "ngram_analyzer": {
                    "type": "custom",
                    "tokenizer": "lowercase",
                    "filter": ["ramon_stopwords", "haystack_ngram"]
                },
                "edgengram_analyzer": {
                    "type": "custom",
                    "tokenizer": "lowercase",
                    "filter": ["ramon_stopwords", "haystack_edgengram"]
                }
            },
            "tokenizer": {
                "haystack_ngram_tokenizer": {
                    "type": "nGram",
                    "min_gram": 3,
                    "max_gram": 15,
                },
                "haystack_edgengram_tokenizer": {
                    "type": "edgeNGram",
                    "min_gram": 2,
                    "max_gram": 15,
                    "side": "front"
                }
            },
            "filter": {
                "haystack_ngram": {
                    "type": "nGram",
                    "min_gram": 3,
                    "max_gram": 15
                },
                "haystack_edgengram": {
                    "type": "edgeNGram",
                    "min_gram": 2,
                    "max_gram": 15
                },
                "ramon_stopwords": {
                    "type": "stop",
                    "stopwords": ["esto","de","que"]
                }
            }
        }
    }
}

```

The settings look like this for the haystack index:

$ curl -XGET '[http://localhost:9200/haystack/\_settings?pretty=true](http://localhost:9200/haystack/_settings?pretty=true)'  
{  
"haystack" : {  
"settings" : {  
"index.analysis.filter.haystack\_edgengram.min\_gram" : "2",  
"index.analysis.filter.haystack\_ngram.max\_gram" : "15",  
"index.analysis.tokenizer.haystack\_ngram\_tokenizer.max\_gram" : "15",  
"index.analysis.analyzer.edgengram\_analyzer.type" : "custom",  
"index.analysis.tokenizer.haystack\_edgengram\_tokenizer.min\_gram" :  
"2",  
"index.analysis.filter.ramon\_stopwords.stopwords.2" : "que",  
"index.analysis.filter.ramon\_stopwords.stopwords.1" : "de",  
"index.analysis.filter.ramon\_stopwords.stopwords.0" : "esto",  
"index.analysis.tokenizer.haystack\_ngram\_tokenizer.min\_gram" : "3",  
"index.analysis.analyzer.ngram\_analyzer.tokenizer" : "lowercase",  
"index.analysis.filter.haystack\_ngram.min\_gram" : "3",  
"index.analysis.analyzer.edgengram\_analyzer.tokenizer" : "lowercase",  
"index.analysis.filter.haystack\_edgengram.max\_gram" : "15",  
"index.analysis.filter.haystack\_ngram.type" : "nGram",  
"index.analysis.analyzer.edgengram\_analyzer.filter.1" :  
"haystack\_edgengram",  
"index.analysis.tokenizer.haystack\_edgengram\_tokenizer.type" :  
"edgeNGram",  
"index.analysis.analyzer.edgengram\_analyzer.filter.0" :  
"ramon\_stopwords",  
"index.analysis.tokenizer.haystack\_edgengram\_tokenizer.side" :  
"front",  
"index.analysis.filter.ramon\_stopwords.type" : "stop",  
"index.analysis.filter.haystack\_edgengram.type" : "edgeNGram",  
"index.analysis.tokenizer.haystack\_ngram\_tokenizer.type" : "nGram",  
"index.analysis.tokenizer.haystack\_edgengram\_tokenizer.max\_gram" :  
"15",  
"index.analysis.analyzer.ngram\_analyzer.filter.1" : "haystack\_ngram",  
"index.analysis.analyzer.ngram\_analyzer.filter.0" : "ramon\_stopwords",  
"index.analysis.analyzer.ngram\_analyzer.type" : "custom",  
"index.number\_of\_shards" : "5",  
"index.number\_of\_replicas" : "1",  
"index.version.created" : "191199"  
}  
}

Which looks right to me. But when testing it the stopwords that are applied  
are only the ones for English and the ones I add remain ignored. See how  
"is" is filtered here:

$ curl -XGET 'localhost:9200/haystack/\_analyze?text=esto+is+a+test+que  
&pretty=true'  
{  
"tokens" : [ {  
"token" : "esto",  
"start\_offset" : 0,  
"end\_offset" : 4,  
"type" : "",  
"position" : 1  
}, {  
"token" : "test",  
"start\_offset" : 10,  
"end\_offset" : 14,  
"type" : "",  
"position" : 4  
}, {  
"token" : "que",  
"start\_offset" : 15,  
"end\_offset" : 18,  
"type" : "",  
"position" : 5  
} ]

The only way I manage to change the stopwords is changing the analyzer in  
the query, but I have tried in the settings too and it doesn't work either.  
This example with the Spanish analyzer works:

$ curl -XGET  
'localhost:9200/haystack/\_analyze?text=esto+is+a+test+que&analyzer=spanish&pr  
etty=true'  
{  
"tokens" : [ {  
"token" : "is",  
"start\_offset" : 5,  
"end\_offset" : 7,  
"type" : "",  
"position" : 2  
}, {  
"token" : "test",  
"start\_offset" : 10,  
"end\_offset" : 14,  
"type" : "",  
"position" : 4  
} ]

Any hint to where this might be failing?

Many thanks.

--

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 27, 2012, 10:06am UTC](https://discuss.elastic.co/t/cant-get-stop-words-working/9828/2 "2012-11-27T10:06:07Z")

</div>

Hi,

You have just defined an analyzer. Fine.  
Now you have to apply it on your mapping [1].  
By default, ES use the standard analyzer. You can change the default analyzer :  
[2]

[1]

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

[2]

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

HTH  
David.

Le 26 novembre 2012 à 20:50, racedo [ramon@linux-labs.net](mailto:ramon@linux-labs.net) a écrit :

> I'm trying to add some stopwords to the default settings that haystack is  
> using and the settings look like this (added "esto", "que" and "de" just for  
> testing purposes:
> 
> ```
> 'settings': {
> "analysis": {
> "analyzer": {
> "ngram_analyzer": {
> "type": "custom",
> "tokenizer": "lowercase",
> "filter": ["ramon_stopwords", "haystack_ngram"]
> },
> "edgengram_analyzer": {
> "type": "custom",
> "tokenizer": "lowercase",
> "filter": ["ramon_stopwords", "haystack_edgengram"]
> }
> },
> "tokenizer": {
> "haystack_ngram_tokenizer": {
> "type": "nGram",
> "min_gram": 3,
> "max_gram": 15,
> },
> "haystack_edgengram_tokenizer": {
> "type": "edgeNGram",
> "min_gram": 2,
> "max_gram": 15,
> "side": "front"
> }
> },
> "filter": {
> "haystack_ngram": {
> "type": "nGram",
> "min_gram": 3,
> "max_gram": 15
> },
> "haystack_edgengram": {
> "type": "edgeNGram",
> "min_gram": 2,
> "max_gram": 15
> },
> "ramon_stopwords": {
> "type": "stop",
> "stopwords": ["esto","de","que"]
> }
> }
> }
> }
> }
> 
> ```
> 
> The settings look like this for the haystack index:
> 
> $ curl -XGET '[http://localhost:9200/haystack/\_settings?pretty=true](http://localhost:9200/haystack/_settings?pretty=true)'  
> {  
> "haystack" : {  
> "settings" : {  
> "index.analysis.filter.haystack\_edgengram.min\_gram" : "2",  
> "index.analysis.filter.haystack\_ngram.max\_gram" : "15",  
> "index.analysis.tokenizer.haystack\_ngram\_tokenizer.max\_gram" : "15",  
> "index.analysis.analyzer.edgengram\_analyzer.type" : "custom",  
> "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.min\_gram" : "2",  
> "index.analysis.filter.ramon\_stopwords.stopwords.2" : "que",  
> "index.analysis.filter.ramon\_stopwords.stopwords.1" : "de",  
> "index.analysis.filter.ramon\_stopwords.stopwords.0" : "esto",  
> "index.analysis.tokenizer.haystack\_ngram\_tokenizer.min\_gram" : "3",  
> "index.analysis.analyzer.ngram\_analyzer.tokenizer" : "lowercase",  
> "index.analysis.filter.haystack\_ngram.min\_gram" : "3",  
> "index.analysis.analyzer.edgengram\_analyzer.tokenizer" : "lowercase",  
> "index.analysis.filter.haystack\_edgengram.max\_gram" : "15",  
> "index.analysis.filter.haystack\_ngram.type" : "nGram",  
> "index.analysis.analyzer.edgengram\_analyzer.filter.1" :  
> "haystack\_edgengram",  
> "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.type" :  
> "edgeNGram",  
> "index.analysis.analyzer.edgengram\_analyzer.filter.0" :  
> "ramon\_stopwords",  
> "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.side" : "front",  
> "index.analysis.filter.ramon\_stopwords.type" : "stop",  
> "index.analysis.filter.haystack\_edgengram.type" : "edgeNGram",  
> "index.analysis.tokenizer.haystack\_ngram\_tokenizer.type" : "nGram",  
> "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.max\_gram" :  
> "15",  
> "index.analysis.analyzer.ngram\_analyzer.filter.1" : "haystack\_ngram",  
> "index.analysis.analyzer.ngram\_analyzer.filter.0" : "ramon\_stopwords",  
> "index.analysis.analyzer.ngram\_analyzer.type" : "custom",  
> "index.number\_of\_shards" : "5",  
> "index.number\_of\_replicas" : "1",  
> "index.version.created" : "191199"  
> }  
> }
> 
> Which looks right to me. But when testing it the stopwords that are applied  
> are only the ones for English and the ones I add remain ignored. See how "is"  
> is filtered here:
> 
> $ curl -XGET  
> 'localhost:9200/haystack/\_analyze?text=esto+is+a+test+que&pretty=true'  
> {  
> "tokens" : [ {  
> "token" : "esto",  
> "start\_offset" : 0,  
> "end\_offset" : 4,  
> "type" : "",  
> "position" : 1  
> }, {  
> "token" : "test",  
> "start\_offset" : 10,  
> "end\_offset" : 14,  
> "type" : "",  
> "position" : 4  
> }, {  
> "token" : "que",  
> "start\_offset" : 15,  
> "end\_offset" : 18,  
> "type" : "",  
> "position" : 5  
> } ]
> 
> The only way I manage to change the stopwords is changing the analyzer in the  
> query, but I have tried in the settings too and it doesn't work either. This  
> example with the Spanish analyzer works:
> 
> $ curl -XGET  
> 'localhost:9200/haystack/\_analyze?text=esto+is+a+test+que&analyzer=spanish&pr  
> etty=true'  
> {  
> "tokens" : [ {  
> "token" : "is",  
> "start\_offset" : 5,  
> "end\_offset" : 7,  
> "type" : "",  
> "position" : 2  
> }, {  
> "token" : "test",  
> "start\_offset" : 10,  
> "end\_offset" : 14,  
> "type" : "",  
> "position" : 4  
> } ]
> 
> Any hint to where this might be failing?
> 
> Many thanks.
> 
> --

--  
David Pilato  
[http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
[http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

**Author:** ![racedo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/racedo/32/2499_2.png) [@racedo](https://discuss.elastic.co/u/racedo)\
**Post date:** [November 27, 2012, 6:13pm UTC](https://discuss.elastic.co/t/cant-get-stop-words-working/9828/3 "2012-11-27T18:13:57Z")

</div>

Hi David,

Your feedback really helps me to understand it better (I only started with  
ES last weekend!). So, after defining an analyzer, is a mapping mandatory?  
As per the ES help page [1] I understood that it was needed when a  
different analyzer would be applied to different document fields.

So far I've managed to get it working by using "default" on my index and  
without a mapping:

```
    'settings': {
       "analysis": {
          "analyzer": {
             "default": {
                 "type": "spanish",
                 "stopwords": ["_spanish_","quot"]
             }
          },
       }
    }

```

Should I rather use a mapping like the one in [1] then? Apologies if this  
is a basic question, I must be misinterpreting the documentation.

Thanks again.

[1] [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/help/)

On Tuesday, November 27, 2012 10:06:18 AM UTC, David Pilato wrote:

> Hi,
> 
> You have just defined an analyzer. Fine.  
> Now you have to apply it on your mapping [1].  
> By default, ES use the standard analyzer. You can change the default  
> analyzer : [2]
> 
> [1]  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/api/admin-indices-put-mapping.html)  
> [2]  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/index.html)
> 
> HTH  
> David.
> 
> Le 26 novembre 2012 à 20:50, racedo \<[ra...@linux-labs.net](mailto:ra...@linux-labs.net) \<javascript:\>\>  
> a écrit :
> 
> I'm trying to add some stopwords to the default settings that haystack is  
> using and the settings look like this (added "esto", "que" and "de" just  
> for testing purposes:
> 
> ```
> 'settings': { 
> "analysis": { 
> "analyzer": { 
> "ngram_analyzer": { 
> "type": "custom", 
> "tokenizer": "lowercase", 
> "filter": ["ramon_stopwords", "haystack_ngram"] 
> }, 
> "edgengram_analyzer": { 
> "type": "custom", 
> "tokenizer": "lowercase", 
> "filter": ["ramon_stopwords", 
> 
> ```
> 
> "haystack\_edgengram"]  
> }  
> },  
> "tokenizer": {  
> "haystack\_ngram\_tokenizer": {  
> "type": "nGram",  
> "min\_gram": 3,  
> "max\_gram": 15,  
> },  
> "haystack\_edgengram\_tokenizer": {  
> "type": "edgeNGram",  
> "min\_gram": 2,  
> "max\_gram": 15,  
> "side": "front"  
> }  
> },  
> "filter": {  
> "haystack\_ngram": {  
> "type": "nGram",  
> "min\_gram": 3,  
> "max\_gram": 15  
> },  
> "haystack\_edgengram": {  
> "type": "edgeNGram",  
> "min\_gram": 2,  
> "max\_gram": 15  
> },  
> "ramon\_stopwords": {  
> "type": "stop",  
> "stopwords": ["esto","de","que"]  
> }  
> }  
> }  
> }  
> }
> 
> The settings look like this for the haystack index:
> 
> $ curl -XGET '[http://localhost:9200/haystack/\_settings?pretty=true](http://localhost:9200/haystack/_settings?pretty=true)'  
> {  
> "haystack" : {  
> "settings" : {  
> "index.analysis.filter.haystack\_edgengram.min\_gram" : "2",  
> "index.analysis.filter.haystack\_ngram.max\_gram" : "15",  
> "index.analysis.tokenizer.haystack\_ngram\_tokenizer.max\_gram" :  
> "15",  
> "index.analysis.analyzer.edgengram\_analyzer.type" : "custom",  
> "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.min\_gram" :  
> "2",  
> "index.analysis.filter.ramon\_stopwords.stopwords.2" : "que",  
> "index.analysis.filter.ramon\_stopwords.stopwords.1" : "de",  
> "index.analysis.filter.ramon\_stopwords.stopwords.0" : "esto",  
> "index.analysis.tokenizer.haystack\_ngram\_tokenizer.min\_gram" : "3",  
> "index.analysis.analyzer.ngram\_analyzer.tokenizer" : "lowercase",  
> "index.analysis.filter.haystack\_ngram.min\_gram" : "3",  
> "index.analysis.analyzer.edgengram\_analyzer.tokenizer" :  
> "lowercase",  
> "index.analysis.filter.haystack\_edgengram.max\_gram" : "15",  
> "index.analysis.filter.haystack\_ngram.type" : "nGram",  
> "index.analysis.analyzer.edgengram\_analyzer.filter.1" :  
> "haystack\_edgengram",  
> "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.type" :  
> "edgeNGram",  
> "index.analysis.analyzer.edgengram\_analyzer.filter.0" :  
> "ramon\_stopwords",  
> "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.side" :  
> "front",  
> "index.analysis.filter.ramon\_stopwords.type" : "stop",  
> "index.analysis.filter.haystack\_edgengram.type" : "edgeNGram",  
> "index.analysis.tokenizer.haystack\_ngram\_tokenizer.type" : "nGram",  
> "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.max\_gram" :  
> "15",  
> "index.analysis.analyzer.ngram\_analyzer.filter.1" :  
> "haystack\_ngram",  
> "index.analysis.analyzer.ngram\_analyzer.filter.0" :  
> "ramon\_stopwords",  
> "index.analysis.analyzer.ngram\_analyzer.type" : "custom",  
> "index.number\_of\_shards" : "5",  
> "index.number\_of\_replicas" : "1",  
> "index.version.created" : "191199"  
> }  
> }
> 
> Which looks right to me. But when testing it the stopwords that are  
> applied are only the ones for English and the ones I add remain ignored.  
> See how "is" is filtered here:
> 
> $ curl -XGET 'localhost:9200/haystack/\_analyze?text=esto+is+a+test+que  
> &pretty=true'  
> {  
> "tokens" : [ {  
> "token" : "esto",  
> "start\_offset" : 0,  
> "end\_offset" : 4,  
> "type" : "",  
> "position" : 1  
> }, {  
> "token" : "test",  
> "start\_offset" : 10,  
> "end\_offset" : 14,  
> "type" : "",  
> "position" : 4  
> }, {  
> "token" : "que",  
> "start\_offset" : 15,  
> "end\_offset" : 18,  
> "type" : "",  
> "position" : 5  
> } ]
> 
> The only way I manage to change the stopwords is changing the analyzer  
> in the query, but I have tried in the settings too and it doesn't work  
> either. This example with the Spanish analyzer works:
> 
> $ curl -XGET  
> 'localhost:9200/haystack/\_analyze?text=esto+is+a+test+que&analyzer=spanish&pr  
> etty=true'  
> {  
> "tokens" : [ {  
> "token" : "is",  
> "start\_offset" : 5,  
> "end\_offset" : 7,  
> "type" : "",  
> "position" : 2  
> }, {  
> "token" : "test",  
> "start\_offset" : 10,  
> "end\_offset" : 14,  
> "type" : "",  
> "position" : 4  
> } ]
> 
> Any hint to where this might be failing?
> 
> Many thanks.
> 
> --
> 
> --  
> David Pilato  
> [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 27, 2012, 8:44pm UTC](https://discuss.elastic.co/t/cant-get-stop-words-working/9828/4 "2012-11-27T20:44:33Z")

</div>

When you have defined a custom analyzer, you have to set where you want to apply  
it.  
When you send your first document, ES compute a mapping automagicaly.

You can get back this mapping with a curl localhost:9200/index/type/\_mapping

Then you can adapt it (add your analyzer to one or many field, as you need) and  
send it back to ES:

First delete all documents:  
curl -XDELETE localhost:9200/index/type

Then, send it again:

curl -XPUT '[http://localhost:9200/twitter/tweet/\_mapping](http://localhost:9200/twitter/tweet/_mapping)' -d '  
{  
"tweet" : {  
"properties" : {  
"message" : {"type" : "string", "analyzer" : "youranalyzername"}  
}  
}  
}  
'

Then send your first document.

Does it help?

David

Le 27 novembre 2012 à 19:13, racedo [ramon@linux-labs.net](mailto:ramon@linux-labs.net) a écrit :

> Hi David,
> 
> Your feedback really helps me to understand it better (I only started with ES  
> last weekend!). So, after defining an analyzer, is a mapping mandatory? As per  
> the ES help page [1] I understood that it was needed when a different analyzer  
> would be applied to different document fields.
> 
> So far I've managed to get it working by using "default" on my index and  
> without a mapping:
> 
> ```
> 'settings': {
> "analysis": {
> "analyzer": {
> "default": {
> "type": "spanish",
> "stopwords": ["_spanish_","quot"]
> }
> },
> }
> }
> 
> ```
> 
> Should I rather use a mapping like the one in [1] then? Apologies if this is  
> a basic question, I must be misinterpreting the documentation.
> 
> Thanks again.
> 
> [1] [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/help/)
> 
> On Tuesday, November 27, 2012 10:06:18 AM UTC, David Pilato wrote:
> 
> > > Hi,
> > 
> > You have just defined an analyzer. Fine.  
> > Now you have to apply it on your mapping [1].  
> > By default, ES use the standard analyzer. You can change the default  
> > analyzer : [2]
> > 
> > [1]  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/api/admin-indices-put-mapping.html)  
> > [http://www.elasticsearch.org/guide/reference/api/admin-indices-put-mapping.html](http://www.elasticsearch.org/guide/reference/api/admin-indices-put-mapping.html)  
> > [2]  
> > [http://www.elasticsearch.org/guide/reference/api/admin-indices-put-mapping.html](http://www.elasticsearch.org/guide/reference/api/admin-indices-put-mapping.html)  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/index.html)  
> > [http://www.elasticsearch.org/guide/reference/index-modules/analysis/index.html](http://www.elasticsearch.org/guide/reference/index-modules/analysis/index.html)
> > 
> > HTH  
> > David.
> > 
> > Le 26 novembre 2012 à 20:50, racedo \<  
> > [http://www.elasticsearch.org/guide/reference/index-modules/analysis/index.html](http://www.elasticsearch.org/guide/reference/index-modules/analysis/index.html)  
> > [ra...@linux-labs.net](mailto:ra...@linux-labs.net)\> a écrit :
> > 
> > ```
> > > > > I'm trying to add some stopwords to the default settings that
> > > > > haystack is using and the settings look like this (added "esto",
> > > > > "que" and "de" just for testing purposes:
> > 
> > ```
> > 
> > > ```
> > > 'settings': {
> > > "analysis": {
> > > "analyzer": {
> > > "ngram_analyzer": {
> > > "type": "custom",
> > > "tokenizer": "lowercase",
> > > "filter": ["ramon_stopwords",
> > > 
> > > ```
> > > 
> > > "haystack\_ngram"]  
> > > },  
> > > "edgengram\_analyzer": {  
> > > "type": "custom",  
> > > "tokenizer": "lowercase",  
> > > "filter": ["ramon\_stopwords",  
> > > "haystack\_edgengram"]  
> > > }  
> > > },  
> > > "tokenizer": {  
> > > "haystack\_ngram\_tokenizer": {  
> > > "type": "nGram",  
> > > "min\_gram": 3,  
> > > "max\_gram": 15,  
> > > },  
> > > "haystack\_edgengram\_tokenizer": {  
> > > "type": "edgeNGram",  
> > > "min\_gram": 2,  
> > > "max\_gram": 15,  
> > > "side": "front"  
> > > }  
> > > },  
> > > "filter": {  
> > > "haystack\_ngram": {  
> > > "type": "nGram",  
> > > "min\_gram": 3,  
> > > "max\_gram": 15  
> > > },  
> > > "haystack\_edgengram": {  
> > > "type": "edgeNGram",  
> > > "min\_gram": 2,  
> > > "max\_gram": 15  
> > > },  
> > > "ramon\_stopwords": {  
> > > "type": "stop",  
> > > "stopwords": ["esto","de","que"]  
> > > }  
> > > }  
> > > }  
> > > }  
> > > }
> > > 
> > > ```
> > > The settings look like this for the haystack index:
> > > 
> > > $ curl -XGET 'http://localhost:9200/haystack/_settings?pretty=true'
> > > 
> > > ```
> > > 
> > > [http://localhost:9200/haystack/\_settings?pretty=true](http://localhost:9200/haystack/_settings?pretty=true)  
> > > {  
> > > "haystack" : {  
> > > "settings" : {  
> > > "index.analysis.filter.haystack\_edgengram.min\_gram" : "2",  
> > > "index.analysis.filter.haystack\_ngram.max\_gram" : "15",  
> > > "index.analysis.tokenizer.haystack\_ngram\_tokenizer.max\_gram" :  
> > > "15",  
> > > "index.analysis.analyzer.edgengram\_analyzer.type" : "custom",  
> > > "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.min\_gram"  
> > > : "2",  
> > > "index.analysis.filter.ramon\_stopwords.stopwords.2" : "que",  
> > > "index.analysis.filter.ramon\_stopwords.stopwords.1" : "de",  
> > > "index.analysis.filter.ramon\_stopwords.stopwords.0" : "esto",  
> > > "index.analysis.tokenizer.haystack\_ngram\_tokenizer.min\_gram" :  
> > > "3",  
> > > "index.analysis.analyzer.ngram\_analyzer.tokenizer" :  
> > > "lowercase",  
> > > "index.analysis.filter.haystack\_ngram.min\_gram" : "3",  
> > > "index.analysis.analyzer.edgengram\_analyzer.tokenizer" :  
> > > "lowercase",  
> > > "index.analysis.filter.haystack\_edgengram.max\_gram" : "15",  
> > > "index.analysis.filter.haystack\_ngram.type" : "nGram",  
> > > "index.analysis.analyzer.edgengram\_analyzer.filter.1" :  
> > > "haystack\_edgengram",  
> > > "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.type" :  
> > > "edgeNGram",  
> > > "index.analysis.analyzer.edgengram\_analyzer.filter.0" :  
> > > "ramon\_stopwords",  
> > > "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.side" :  
> > > "front",  
> > > "index.analysis.filter.ramon\_stopwords.type" : "stop",  
> > > "index.analysis.filter.haystack\_edgengram.type" : "edgeNGram",  
> > > "index.analysis.tokenizer.haystack\_ngram\_tokenizer.type" :  
> > > "nGram",  
> > > "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.max\_gram"  
> > > : "15",  
> > > "index.analysis.analyzer.ngram\_analyzer.filter.1" :  
> > > "haystack\_ngram",  
> > > "index.analysis.analyzer.ngram\_analyzer.filter.0" :  
> > > "ramon\_stopwords",  
> > > "index.analysis.analyzer.ngram\_analyzer.type" : "custom",  
> > > "index.number\_of\_shards" : "5",  
> > > "index.number\_of\_replicas" : "1",  
> > > "index.version.created" : "191199"  
> > > }  
> > > }
> > > 
> > > ```
> > > Which looks right to me. But when testing it the stopwords that are
> > > 
> > > ```
> > > 
> > > applied are only the ones for English and the ones I add remain ignored.  
> > > See how "is" is filtered here:
> > > 
> > > ```
> > > $ curl -XGET
> > > 
> > > ```
> > > 
> > > 'localhost:9200/haystack/\_analyze?text=esto+is+a+test+que&pretty=true'  
> > > {  
> > > "tokens" : [ {  
> > > "token" : "esto",  
> > > "start\_offset" : 0,  
> > > "end\_offset" : 4,  
> > > "type" : "",  
> > > "position" : 1  
> > > }, {  
> > > "token" : "test",  
> > > "start\_offset" : 10,  
> > > "end\_offset" : 14,  
> > > "type" : "",  
> > > "position" : 4  
> > > }, {  
> > > "token" : "que",  
> > > "start\_offset" : 15,  
> > > "end\_offset" : 18,  
> > > "type" : "",  
> > > "position" : 5  
> > > } ]
> > > 
> > > ```
> > > The only way I manage to change the stopwords is changing the analyzer
> > > 
> > > ```
> > > 
> > > in the query, but I have tried in the settings too and it doesn't work  
> > > either. This example with the Spanish analyzer works:
> > > 
> > > ```
> > > $ curl -XGET
> > > 
> > > ```
> > > 
> > > 'localhost:9200/haystack/\_analyze?text=esto+is+a+test+que&analyzer=spanish&pr  
> > > etty=true'  
> > > {  
> > > "tokens" : [ {  
> > > "token" : "is",  
> > > "start\_offset" : 5,  
> > > "end\_offset" : 7,  
> > > "type" : "",  
> > > "position" : 2  
> > > }, {  
> > > "token" : "test",  
> > > "start\_offset" : 10,  
> > > "end\_offset" : 14,  
> > > "type" : "",  
> > > "position" : 4  
> > > } ]
> > > 
> > > ```
> > > Any hint to where this might be failing?
> > > 
> > > Many thanks.
> > > 
> > > --
> > > 
> > > ```
> > > 
> > > > > ```
> > > > > <http://localhost:9200/haystack/_settings?pretty=true>
> > > > > 
> > > > > ```
> > > > > 
> > > > > [http://localhost:9200/haystack/\_settings?pretty=true](http://localhost:9200/haystack/_settings?pretty=true)
> > 
> > --  
> > David Pilato  
> > [http://localhost:9200/haystack/\_settings?pretty=true](http://localhost:9200/haystack/_settings?pretty=true)  
> > [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/) [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> > [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/) [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> > Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> > 
> > >
> 
> --

--  
David Pilato  
[http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
[http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

**Author:** ![racedo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/racedo/32/2499_2.png) [@racedo](https://discuss.elastic.co/u/racedo)\
**Post date:** [November 28, 2012, 2:08am UTC](https://discuss.elastic.co/t/cant-get-stop-words-working/9828/5 "2012-11-28T02:08:38Z")

</div>

It does indeed help. Actually my problem was a bug in django haystack  
([https://github.com/toastdriven/django-haystack/issues/686](https://github.com/toastdriven/django-haystack/issues/686)) which manages  
the mappings through pyelasticsearch.

Your suggestions are great, thanks to the above bug (which had me  
researching on mappings non-stop) and your feedback I got up to speed with  
elasticsearch in three intensive days. Really appreciated.

Ramon

On Tuesday, 27 November 2012 20:44:33 UTC, David Pilato wrote:

> When you have defined a custom analyzer, you have to set where you want  
> to apply it.  
> When you send your first document, ES compute a mapping automagicaly.
> 
> You can get back this mapping with a curl  
> localhost:9200/index/type/\_mapping
> 
> Then you can adapt it (add your analyzer to one or many field, as you  
> need) and send it back to ES:
> 
> First delete all documents:  
> curl -XDELETE localhost:9200/index/type
> 
> Then, send it again:
> 
> curl -XPUT '[http://localhost:9200/twitter/tweet/\_mapping](http://localhost:9200/twitter/tweet/_mapping)' -d '  
> {  
> "tweet" : {  
> "properties" : {  
> "message" : {"type" : "string", "analyzer" : "youranalyzername"}  
> }  
> }  
> }  
> '
> 
> Then send your first document.
> 
> Does it help?
> 
> David
> 
> Le 27 novembre 2012 à 19:13, racedo \<[ra...@linux-labs.net](mailto:ra...@linux-labs.net) \<javascript:\>\>  
> a écrit :
> 
> Hi David,
> 
> Your feedback really helps me to understand it better (I only started  
> with ES last weekend!). So, after defining an analyzer, is a mapping  
> mandatory? As per the ES help page [1] I understood that it was needed when  
> a different analyzer would be applied to different document fields.
> 
> So far I've managed to get it working by using "default" on my index and  
> without a mapping:
> 
> ```
> 'settings': { 
> "analysis": { 
> "analyzer": { 
> "default": { 
> "type": "spanish", 
> "stopwords": ["_spanish_","quot"] 
> } 
> }, 
> } 
> } 
> 
> ```
> 
> Should I rather use a mapping like the one in [1] then? Apologies if  
> this is a basic question, I must be misinterpreting the documentation.
> 
> Thanks again.
> 
> [1] [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/help/)
> 
> On Tuesday, November 27, 2012 10:06:18 AM UTC, David Pilato wrote:
> 
> Hi,
> 
> You have just defined an analyzer. Fine.  
> Now you have to apply it on your mapping [1].  
> By default, ES use the standard analyzer. You can change the default  
> analyzer : [2]
> 
> [1] [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/api/admin-indices-put-mapping.html)
> 
> [2]  
> [http://www.elasticsearch.org/guide/reference/api/admin-indices-put-mapping.html](http://www.elasticsearch.org/guide/reference/api/admin-indices-put-mapping.html)  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/index.html)
> 
> HTH  
> David.
> 
> Le 26 novembre 2012 à 20:50, racedo \<[http://www.elasticsearch.org/guide/reference/index-modules/analysis/index.html](http://www.elasticsearch.org/guide/reference/index-modules/analysis/index.html)  
> [ra...@linux-labs.net](mailto:ra...@linux-labs.net)\> a écrit :
> 
> I'm trying to add some stopwords to the default settings that haystack is  
> using and the settings look like this (added "esto", "que" and "de" just  
> for testing purposes:
> 
> ```
> 'settings': { 
> "analysis": { 
> "analyzer": { 
> "ngram_analyzer": { 
> "type": "custom", 
> "tokenizer": "lowercase", 
> "filter": ["ramon_stopwords", "haystack_ngram"] 
> }, 
> "edgengram_analyzer": { 
> "type": "custom", 
> "tokenizer": "lowercase", 
> "filter": ["ramon_stopwords", 
> 
> ```
> 
> "haystack\_edgengram"]  
> }  
> },  
> "tokenizer": {  
> "haystack\_ngram\_tokenizer": {  
> "type": "nGram",  
> "min\_gram": 3,  
> "max\_gram": 15,  
> },  
> "haystack\_edgengram\_tokenizer": {  
> "type": "edgeNGram",  
> "min\_gram": 2,  
> "max\_gram": 15,  
> "side": "front"  
> }  
> },  
> "filter": {  
> "haystack\_ngram": {  
> "type": "nGram",  
> "min\_gram": 3,  
> "max\_gram": 15  
> },  
> "haystack\_edgengram": {  
> "type": "edgeNGram",  
> "min\_gram": 2,  
> "max\_gram": 15  
> },  
> "ramon\_stopwords": {  
> "type": "stop",  
> "stopwords": ["esto","de","que"]  
> }  
> }  
> }  
> }  
> }
> 
> The settings look like this for the haystack index:
> 
> $ curl -XGET '[http://localhost:9200/haystack/\_settings?pretty=true](http://localhost:9200/haystack/_settings?pretty=true)'  
> [http://localhost:9200/haystack/\_settings?pretty=true](http://localhost:9200/haystack/_settings?pretty=true)  
> {  
> "haystack" : {  
> "settings" : {  
> "index.analysis.filter.haystack\_edgengram.min\_gram" : "2",  
> "index.analysis.filter.haystack\_ngram.max\_gram" : "15",  
> "index.analysis.tokenizer.haystack\_ngram\_tokenizer.max\_gram" :  
> "15",  
> "index.analysis.analyzer.edgengram\_analyzer.type" : "custom",  
> "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.min\_gram" :  
> "2",  
> "index.analysis.filter.ramon\_stopwords.stopwords.2" : "que",  
> "index.analysis.filter.ramon\_stopwords.stopwords.1" : "de",  
> "index.analysis.filter.ramon\_stopwords.stopwords.0" : "esto",  
> "index.analysis.tokenizer.haystack\_ngram\_tokenizer.min\_gram" : "3",  
> "index.analysis.analyzer.ngram\_analyzer.tokenizer" : "lowercase",  
> "index.analysis.filter.haystack\_ngram.min\_gram" : "3",  
> "index.analysis.analyzer.edgengram\_analyzer.tokenizer" :  
> "lowercase",  
> "index.analysis.filter.haystack\_edgengram.max\_gram" : "15",  
> "index.analysis.filter.haystack\_ngram.type" : "nGram",  
> "index.analysis.analyzer.edgengram\_analyzer.filter.1" :  
> "haystack\_edgengram",  
> "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.type" :  
> "edgeNGram",  
> "index.analysis.analyzer.edgengram\_analyzer.filter.0" :  
> "ramon\_stopwords",  
> "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.side" :  
> "front",  
> "index.analysis.filter.ramon\_stopwords.type" : "stop",  
> "index.analysis.filter.haystack\_edgengram.type" : "edgeNGram",  
> "index.analysis.tokenizer.haystack\_ngram\_tokenizer.type" :  
> "nGram",  
> "index.analysis.tokenizer.haystack\_edgengram\_tokenizer.max\_gram" :  
> "15",  
> "index.analysis.analyzer.ngram\_analyzer.filter.1" :  
> "haystack\_ngram",  
> "index.analysis.analyzer.ngram\_analyzer.filter.0" :  
> "ramon\_stopwords",  
> "index.analysis.analyzer.ngram\_analyzer.type" : "custom",  
> "index.number\_of\_shards" : "5",  
> "index.number\_of\_replicas" : "1",  
> "index.version.created" : "191199"  
> }  
> }
> 
> Which looks right to me. But when testing it the stopwords that are  
> applied are only the ones for English and the ones I add remain ignored.  
> See how "is" is filtered here:
> 
> $ curl -XGET 'localhost:9200/haystack/\_analyze?text=esto+is+a+test+  
> que&pretty=true'  
> {  
> "tokens" : [ {  
> "token" : "esto",  
> "start\_offset" : 0,  
> "end\_offset" : 4,  
> "type" : "",  
> "position" : 1  
> }, {  
> "token" : "test",  
> "start\_offset" : 10,  
> "end\_offset" : 14,  
> "type" : "",  
> "position" : 4  
> }, {  
> "token" : "que",  
> "start\_offset" : 15,  
> "end\_offset" : 18,  
> "type" : "",  
> "position" : 5  
> } ]
> 
> The only way I manage to change the stopwords is changing the analyzer  
> in the query, but I have tried in the settings too and it doesn't work  
> either. This example with the Spanish analyzer works:
> 
> $ curl -XGET  
> 'localhost:9200/haystack/\_analyze?text=esto+is+a+test+que&analyzer=spanish&pr  
> etty=true'  
> {  
> "tokens" : [ {  
> "token" : "is",  
> "start\_offset" : 5,  
> "end\_offset" : 7,  
> "type" : "",  
> "position" : 2  
> }, {  
> "token" : "test",  
> "start\_offset" : 10,  
> "end\_offset" : 14,  
> "type" : "",  
> "position" : 4  
> } ]
> 
> Any hint to where this might be failing?
> 
> Many thanks.
> 
> --
> 
> [http://localhost:9200/haystack/\_settings?pretty=true](http://localhost:9200/haystack/_settings?pretty=true)  
> [http://localhost:9200/haystack/\_settings?pretty=true](http://localhost:9200/haystack/_settings?pretty=true)
> 
> --  
> David Pilato  
> [http://localhost:9200/haystack/\_settings?pretty=true](http://localhost:9200/haystack/_settings?pretty=true)  
> [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> 
> --
> 
> --  
> David Pilato  
> [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:02am UTC](https://discuss.elastic.co/t/cant-get-stop-words-working/9828/6 "2017-07-06T03:02:32Z")

</div>


