# Stopwords file format

**URL:** <https://discuss.elastic.co/t/stopwords-file-format/6228>\
**Category:** Elasticsearch\
**Created:** [December 23, 2011, 2:42am UTC](https://discuss.elastic.co/t/stopwords-file-format/6228 "2011-12-23T02:42:59Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Eugene\_Strokin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eugene_strokin/32/1238_2.png) [@Eugene\_Strokin](https://discuss.elastic.co/u/Eugene_Strokin)\
**Post date:** [December 23, 2011, 2:42am UTC](https://discuss.elastic.co/t/stopwords-file-format/6228/1 "2011-12-23T02:42:59Z")

</div>

I want to specify my own stop-words. This is what I found so far:  
[http://www.elasticsearch.org/guide/reference/index-modules/analysis/stop-tokenfilter.html](http://www.elasticsearch.org/guide/reference/index-modules/analysis/stop-tokenfilter.html)  
In elasticsearch.yml I'd have such analyzer specified:

index :  
analysis:  
analyzer:  
string\_lowercase:  
tokenizer : keyword  
filter : lowercase  
stopwords\_path : stopwords.txt  
ignore\_case : true

How should I specify the stop-words in the stopwords.txt file? Just a  
word in a line, or somehow else?

Also, I don't care which language users will use to index data, so if  
I'd put stopwords from different languages into the same file, it  
should be no problem, but should I use just UTF-8 encoding, or should  
I use encoding like we use in .properties files, e.q. "de art  
\u00edculos"?

Thank you,  
Eugene S.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [December 25, 2011, 4:35pm UTC](https://discuss.elastic.co/t/stopwords-file-format/6228/2 "2011-12-25T16:35:29Z")

</div>

Each stop word should be in its own "line" (separated by \n). The file is  
read in UTF8 format.

On Fri, Dec 23, 2011 at 4:42 AM, Eugene Strokin [eugene@strokin.info](mailto:eugene@strokin.info) wrote:

> I want to specify my own stop-words. This is what I found so far:
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/stop-tokenfilter.html)  
> In elasticsearch.yml I'd have such analyzer specified:
> 
> index :  
> analysis:  
> analyzer:  
> string\_lowercase:  
> tokenizer : keyword  
> filter : lowercase  
> stopwords\_path : stopwords.txt  
> ignore\_case : true
> 
> How should I specify the stop-words in the stopwords.txt file? Just a  
> word in a line, or somehow else?
> 
> Also, I don't care which language users will use to index data, so if  
> I'd put stopwords from different languages into the same file, it  
> should be no problem, but should I use just UTF-8 encoding, or should  
> I use encoding like we use in .properties files, e.q. "de art  
> \u00edculos"?
> 
> Thank you,  
> Eugene S.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:44am UTC](https://discuss.elastic.co/t/stopwords-file-format/6228/3 "2017-07-06T03:44:23Z")

</div>


