# Re-Index Strategies

**URL:** <https://discuss.elastic.co/t/re-index-strategies/3071>\
**Category:** Elasticsearch\
**Created:** [July 7, 2010, 2:16am UTC](https://discuss.elastic.co/t/re-index-strategies/3071 "2010-07-07T02:16:09Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Andrew\_Harvey\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrew_harvey_2/32/3328_2.png) [@Andrew\_Harvey\_2](https://discuss.elastic.co/u/Andrew_Harvey_2)\
**Post date:** [July 7, 2010, 2:16am UTC](https://discuss.elastic.co/t/re-index-strategies/3071/1 "2010-07-07T02:16:09Z")

</div>

Hi All,

I'm wondering what re-index strategies you guys are using with elasticsearch.

I primarily interact with ES through the REST api, and my java-fu is not exactly strong, but I'm not exactly keen on the overhead required by pulling all my docs out via HTTP and putting them back in via HTTP. Is there a better way, or should I just bite the bullet and do a scrolled search to re-index all my documents?

Why am I re-indexing you might ask? I need to add stopwords to my analyser, and I'm fairly sure that's going to require a re-index, but I'm willing to be proven wrong.

Any help would be most appreciated.

## Andrew Andrew Harvey / Developer lexer m/ t/ +61 2 9019 6379 w/ [http://lexer.com.au](http://lexer.com.au) Help put an end to whaling. Visit [http://www.givewhalesavoice.com.au/](http://www.givewhalesavoice.com.au/)

Please consider the environment before printing this email  
This email transmission is confidential and intended solely for the person or organisation to whom it is addressed. If you are not the intended recipient, you must not copy, distribute or disseminate the information, or take any action in relation to it and please delete this e-mail. Any views expressed in this message are those of the individual sender, except where the send specifically states them to be the views of any organisation or employer. If you have received this message in error, do not open any attachment but please notify the sender (above). This message has been checked for all known viruses powered by McAfee.

For further information visit [http://www.mcafee.com/us/threat\_center/default.asp](http://www.mcafee.com/us/threat_center/default.asp)  
Please rely on your own virus check as no responsibility is taken by the sender for any damage rising out of any virus infection this communication may contain.

This message has been scanned for malware by Websense. [www.websense.com](http://www.websense.com)

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 7, 2010, 7:12am UTC](https://discuss.elastic.co/t/re-index-strategies/3071/2 "2010-07-07T07:12:36Z")

</div>

A simple thing that can improve the indexing performance without doing java  
would be to use the memcached protocol, which have shown to be faster. But I  
would also say measure first before you say that the HTTP is too slow. It  
might be good enough, especially if you fork the indexing into several  
processes / threads (for example, by executing a search which segments the  
data using a filter for each filter / thread).

Regarding adding stopwords without reindexing, if you change them, they will  
only be applied to feature indexing and search requests (which might be  
enough for you, especially on the search side).

-shay.banon

On Wed, Jul 7, 2010 at 5:16 AM, Andrew Harvey [Andrew.Harvey@lexer.com.au](mailto:Andrew.Harvey@lexer.com.au)wrote:

> Hi All,
> 
> I'm wondering what re-index strategies you guys are using with  
> elasticsearch.
> 
> I primarily interact with ES through the REST api, and my java-fu is not  
> exactly strong, but I'm not exactly keen on the overhead required by pulling  
> all my docs out via HTTP and putting them back in via HTTP. Is there a  
> better way, or should I just bite the bullet and do a scrolled search to  
> re-index all my documents?
> 
> Why am I re-indexing you might ask? I need to add stopwords to my analyser,  
> and I'm fairly sure that's going to require a re-index, but I'm willing to  
> be proven wrong.
> 
> Any help would be most appreciated.
> 
> ## Andrew Andrew Harvey / Developer lexer m/ t/ +61 2 9019 6379 w/ [http://lexer.com.au](http://lexer.com.au) Help put an end to whaling. Visit [http://www.givewhalesavoice.com.au/](http://www.givewhalesavoice.com.au/)
> 
> Please consider the environment before printing this email  
> This email transmission is confidential and intended solely for the person  
> or organisation to whom it is addressed. If you are not the intended  
> recipient, you must not copy, distribute or disseminate the information, or  
> take any action in relation to it and please delete this e-mail. Any views  
> expressed in this message are those of the individual sender, except where  
> the send specifically states them to be the views of any organisation or  
> employer. If you have received this message in error, do not open any  
> attachment but please notify the sender (above). This message has been  
> checked for all known viruses powered by McAfee.
> 
> For further information visit  
> [Advanced Research Center | Trellix](http://www.mcafee.com/us/threat_center/default.asp)  
> Please rely on your own virus check as no responsibility is taken by the  
> sender for any damage rising out of any virus infection this communication  
> may contain.
> 
> This message has been scanned for malware by Websense. [www.websense.com](http://www.websense.com)

---

<div class="post-metadata">

**Author:** ![Andrew\_Harvey\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrew_harvey_2/32/3328_2.png) [@Andrew\_Harvey\_2](https://discuss.elastic.co/u/Andrew_Harvey_2)\
**Post date:** [July 7, 2010, 7:15am UTC](https://discuss.elastic.co/t/re-index-strategies/3071/3 "2010-07-07T07:15:09Z")

</div>

Right. I'll just go down the HTTP route. It shouldn't be too slow.

The reason I need to change the stop words is for term facets. I'd wager they _will_ require a re-index.

Andrew

On 07/07/2010, at 5:12 PM, Shay Banon wrote:

A simple thing that can improve the indexing performance without doing java would be to use the memcached protocol, which have shown to be faster. But I would also say measure first before you say that the HTTP is too slow. It might be good enough, especially if you fork the indexing into several processes / threads (for example, by executing a search which segments the data using a filter for each filter / thread).

Regarding adding stopwords without reindexing, if you change them, they will only be applied to feature indexing and search requests (which might be enough for you, especially on the search side).

-shay.banon

On Wed, Jul 7, 2010 at 5:16 AM, Andrew Harvey \<Andrew.Harvey@lexer.com.au[mailto:Andrew.Harvey@lexer.com.au](mailto:Andrew.Harvey@lexer.com.au)\> wrote:  
Hi All,

I'm wondering what re-index strategies you guys are using with elasticsearch.

I primarily interact with ES through the REST api, and my java-fu is not exactly strong, but I'm not exactly keen on the overhead required by pulling all my docs out via HTTP and putting them back in via HTTP. Is there a better way, or should I just bite the bullet and do a scrolled search to re-index all my documents?

Why am I re-indexing you might ask? I need to add stopwords to my analyser, and I'm fairly sure that's going to require a re-index, but I'm willing to be proven wrong.

Any help would be most appreciated.

## Andrew Andrew Harvey / Developer lexer m/ t/ +61 2 9019 6379 w/ [http://lexer.com.au](http://lexer.com.au)[http://lexer.com.au/](http://lexer.com.au/) Help put an end to whaling. Visit [http://www.givewhalesavoice.com.au/](http://www.givewhalesavoice.com.au/)

Please consider the environment before printing this email  
This email transmission is confidential and intended solely for the person or organisation to whom it is addressed. If you are not the intended recipient, you must not copy, distribute or disseminate the information, or take any action in relation to it and please delete this e-mail. Any views expressed in this message are those of the individual sender, except where the send specifically states them to be the views of any organisation or employer. If you have received this message in error, do not open any attachment but please notify the sender (above). This message has been checked for all known viruses powered by McAfee.

For further information visit [http://www.mcafee.com/us/threat\_center/default.asp](http://www.mcafee.com/us/threat_center/default.asp)  
Please rely on your own virus check as no responsibility is taken by the sender for any damage rising out of any virus infection this communication may contain.

This message has been scanned for malware by Websense. [www.websense.com](http://www.websense.com)[http://www.websense.com/](http://www.websense.com/)

Click here[https://www.mailcontrol.com/sr/1X9s1s6k5KzTndxI!oX7UrpakrQMGuaSaLC5GvEXgaH3tvJPQukjtl1!sQzokgGawjM0KpQ9XUMsWPB8dn7asg==](https://www.mailcontrol.com/sr/1X9s1s6k5KzTndxI!oX7UrpakrQMGuaSaLC5GvEXgaH3tvJPQukjtl1!sQzokgGawjM0KpQ9XUMsWPB8dn7asg==) to report this email as spam.

Andrew Harvey / Developer  
lexer

m/  
t/ +61 2 9019 6379  
w/ [http://lexer.com.au](http://lexer.com.au)

Help put an end to whaling. Visit www.givewhalesavoice.com.au[http://www.givewhalesavoice.com.au/](http://www.givewhalesavoice.com.au/)

* * *

Please consider the environment before printing this email  
This email transmission is confidential and intended solely for the person or organisation to whom it is addressed. If you are not the intended recipient, you must not copy, distribute or disseminate the information, or take any action in relation to it and please delete this e-mail. Any views expressed in this message are those of the individual sender, except where the send specifically states them to be the views of any organisation or employer. If you have received this message in error, do not open any attachment but please notify the sender (above). This message has been checked for all known viruses powered by McAfee.

For further information visit [http://www.mcafee.com/us/threat\_center/default.asp](http://www.mcafee.com/us/threat_center/default.asp)  
Please rely on your own virus check as no responsibility is taken by the sender for any damage rising out of any virus infection this communication may contain.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 7, 2010, 11:24am UTC](https://discuss.elastic.co/t/re-index-strategies/3071/4 "2010-07-07T11:24:49Z")

</div>

Yes, 0.9 will require reindex, but you raise a very good point, I think that  
a nice feature would be to define an exclude term set as an optional setting  
for terms facets (can be set on the request side, or on per request, or  
both). This means that the stop words can remain independent of the terms  
facets behavior. What do you think? If it make sense, can you open an issue  
for this?

-shay.banon

On Wed, Jul 7, 2010 at 10:15 AM, Andrew Harvey  
[Andrew.Harvey@lexer.com.au](mailto:Andrew.Harvey@lexer.com.au)wrote:

> Right. I'll just go down the HTTP route. It shouldn't be too slow.
> 
> The reason I need to change the stop words is for term facets. I'd wager  
> they _will_ require a re-index.
> 
> Andrew
> 
> On 07/07/2010, at 5:12 PM, Shay Banon wrote:
> 
> A simple thing that can improve the indexing performance without doing java  
> would be to use the memcached protocol, which have shown to be faster. But I  
> would also say measure first before you say that the HTTP is too slow. It  
> might be good enough, especially if you fork the indexing into several  
> processes / threads (for example, by executing a search which segments the  
> data using a filter for each filter / thread).
> 
> Regarding adding stopwords without reindexing, if you change them, they  
> will only be applied to feature indexing and search requests (which might be  
> enough for you, especially on the search side).
> 
> -shay.banon
> 
> On Wed, Jul 7, 2010 at 5:16 AM, Andrew Harvey [Andrew.Harvey@lexer.com.au](mailto:Andrew.Harvey@lexer.com.au)wrote:
> 
> > Hi All,
> > 
> > I'm wondering what re-index strategies you guys are using with  
> > elasticsearch.
> > 
> > I primarily interact with ES through the REST api, and my java-fu is not  
> > exactly strong, but I'm not exactly keen on the overhead required by pulling  
> > all my docs out via HTTP and putting them back in via HTTP. Is there a  
> > better way, or should I just bite the bullet and do a scrolled search to  
> > re-index all my documents?
> > 
> > Why am I re-indexing you might ask? I need to add stopwords to my  
> > analyser, and I'm fairly sure that's going to require a re-index, but I'm  
> > willing to be proven wrong.
> > 
> > Any help would be most appreciated.
> > 
> > ## Andrew Andrew Harvey / Developer lexer m/ t/ +61 2 9019 6379 w/ [http://lexer.com.au](http://lexer.com.au) Help put an end to whaling. Visit [http://www.givewhalesavoice.com.au/](http://www.givewhalesavoice.com.au/)
> > 
> > Please consider the environment before printing this email  
> > This email transmission is confidential and intended solely for the person  
> > or organisation to whom it is addressed. If you are not the intended  
> > recipient, you must not copy, distribute or disseminate the information, or  
> > take any action in relation to it and please delete this e-mail. Any views  
> > expressed in this message are those of the individual sender, except where  
> > the send specifically states them to be the views of any organisation or  
> > employer. If you have received this message in error, do not open any  
> > attachment but please notify the sender (above). This message has been  
> > checked for all known viruses powered by McAfee.
> > 
> > For further information visit  
> > [Advanced Research Center | Trellix](http://www.mcafee.com/us/threat_center/default.asp)  
> > Please rely on your own virus check as no responsibility is taken by the  
> > sender for any damage rising out of any virus infection this communication  
> > may contain.
> > 
> > This message has been scanned for malware by Websense. [www.websense.com](http://www.websense.com)
> 
> Click here[https://www.mailcontrol.com/sr/1X9s1s6k5KzTndxI!oX7UrpakrQMGuaSaLC5GvEXgaH3tvJPQukjtl1!sQzokgGawjM0KpQ9XUMsWPB8dn7asg==](https://www.mailcontrol.com/sr/1X9s1s6k5KzTndxI!oX7UrpakrQMGuaSaLC5GvEXgaH3tvJPQukjtl1!sQzokgGawjM0KpQ9XUMsWPB8dn7asg==)to report this email as spam.
> 
> Andrew Harvey / Developer _lexer_
> 
> _m/__t/_ +61 2 9019 6379 _w/_[http://lexer.com.au](http://lexer.com.au)
> 
> ## Help put an end to whaling. Visit www.givewhalesavoice.com.au
> 
> Please consider the environment before printing this email  
> This email transmission is confidential and intended solely for the person  
> or organisation to whom it is addressed. If you are not the intended  
> recipient, you must not copy, distribute or disseminate the information, or  
> take any action in relation to it and please delete this e-mail. Any views  
> expressed in this message are those of the individual sender, except where  
> the send specifically states them to be the views of any organisation or  
> employer. If you have received this message in error, do not open any  
> attachment but please notify the sender (above). This message has been  
> checked for all known viruses powered by McAfee.
> 
> For further information visit  
> [Advanced Research Center | Trellix](http://www.mcafee.com/us/threat_center/default.asp)  
> Please rely on your own virus check as no responsibility is taken by the  
> sender for any damage rising out of any virus infection this communication  
> may contain.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 7, 2010, 11:42am UTC](https://discuss.elastic.co/t/re-index-strategies/3071/5 "2010-07-07T11:42:09Z")

</div>

ok, opened [Terms Facets: Allow to specify a set of terms to exclude in the request · Issue #246 · elastic/elasticsearch · GitHub](http://github.com/elasticsearch/elasticsearch/issues/issue/246),  
the request part terms to exclude is simple to implement, should be in  
master laster today.

-shay.banon

On Wed, Jul 7, 2010 at 2:24 PM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> Yes, 0.9 will require reindex, but you raise a very good point, I think  
> that a nice feature would be to define an exclude term set as an optional  
> setting for terms facets (can be set on the request side, or on per request,  
> or both). This means that the stop words can remain independent of the terms  
> facets behavior. What do you think? If it make sense, can you open an issue  
> for this?
> 
> -shay.banon
> 
> On Wed, Jul 7, 2010 at 10:15 AM, Andrew Harvey \<Andrew.Harvey@lexer.com.au
> 
> > wrote:
> 
> > Right. I'll just go down the HTTP route. It shouldn't be too slow.
> > 
> > The reason I need to change the stop words is for term facets. I'd wager  
> > they _will_ require a re-index.
> > 
> > Andrew
> > 
> > On 07/07/2010, at 5:12 PM, Shay Banon wrote:
> > 
> > A simple thing that can improve the indexing performance without doing  
> > java would be to use the memcached protocol, which have shown to be faster.  
> > But I would also say measure first before you say that the HTTP is too slow.  
> > It might be good enough, especially if you fork the indexing into several  
> > processes / threads (for example, by executing a search which segments the  
> > data using a filter for each filter / thread).
> > 
> > Regarding adding stopwords without reindexing, if you change them, they  
> > will only be applied to feature indexing and search requests (which might be  
> > enough for you, especially on the search side).
> > 
> > -shay.banon
> > 
> > On Wed, Jul 7, 2010 at 5:16 AM, Andrew Harvey \<Andrew.Harvey@lexer.com.au
> > 
> > > wrote:
> > 
> > > Hi All,
> > > 
> > > I'm wondering what re-index strategies you guys are using with  
> > > elasticsearch.
> > > 
> > > I primarily interact with ES through the REST api, and my java-fu is not  
> > > exactly strong, but I'm not exactly keen on the overhead required by pulling  
> > > all my docs out via HTTP and putting them back in via HTTP. Is there a  
> > > better way, or should I just bite the bullet and do a scrolled search to  
> > > re-index all my documents?
> > > 
> > > Why am I re-indexing you might ask? I need to add stopwords to my  
> > > analyser, and I'm fairly sure that's going to require a re-index, but I'm  
> > > willing to be proven wrong.
> > > 
> > > Any help would be most appreciated.
> > > 
> > > Andrew  
> > > Andrew Harvey / Developer  
> > > lexer  
> > > m/  
> > > t/ +61 2 9019 6379  
> > > w/ [http://lexer.com.au](http://lexer.com.au)  
> > > Help put an end to whaling. Visit [http://www.givewhalesavoice.com.au/](http://www.givewhalesavoice.com.au/)
> > > 
> > > * * *
> > > 
> > > Please consider the environment before printing this email  
> > > This email transmission is confidential and intended solely for the  
> > > person or organisation to whom it is addressed. If you are not the intended  
> > > recipient, you must not copy, distribute or disseminate the information, or  
> > > take any action in relation to it and please delete this e-mail. Any views  
> > > expressed in this message are those of the individual sender, except where  
> > > the send specifically states them to be the views of any organisation or  
> > > employer. If you have received this message in error, do not open any  
> > > attachment but please notify the sender (above). This message has been  
> > > checked for all known viruses powered by McAfee.
> > > 
> > > For further information visit  
> > > [Advanced Research Center | Trellix](http://www.mcafee.com/us/threat_center/default.asp)  
> > > Please rely on your own virus check as no responsibility is taken by the  
> > > sender for any damage rising out of any virus infection this communication  
> > > may contain.
> > > 
> > > This message has been scanned for malware by Websense. [www.websense.com](http://www.websense.com)
> > 
> > Click here[https://www.mailcontrol.com/sr/1X9s1s6k5KzTndxI!oX7UrpakrQMGuaSaLC5GvEXgaH3tvJPQukjtl1!sQzokgGawjM0KpQ9XUMsWPB8dn7asg==](https://www.mailcontrol.com/sr/1X9s1s6k5KzTndxI!oX7UrpakrQMGuaSaLC5GvEXgaH3tvJPQukjtl1!sQzokgGawjM0KpQ9XUMsWPB8dn7asg==)to report this email as spam.
> > 
> > Andrew Harvey / Developer _lexer_
> > 
> > _m/__t/_ +61 2 9019 6379 _w/_[http://lexer.com.au](http://lexer.com.au)
> > 
> > ## Help put an end to whaling. Visit www.givewhalesavoice.com.au
> > 
> > Please consider the environment before printing this email  
> > This email transmission is confidential and intended solely for the person  
> > or organisation to whom it is addressed. If you are not the intended  
> > recipient, you must not copy, distribute or disseminate the information, or  
> > take any action in relation to it and please delete this e-mail. Any views  
> > expressed in this message are those of the individual sender, except where  
> > the send specifically states them to be the views of any organisation or  
> > employer. If you have received this message in error, do not open any  
> > attachment but please notify the sender (above). This message has been  
> > checked for all known viruses powered by McAfee.
> > 
> > For further information visit  
> > [Advanced Research Center | Trellix](http://www.mcafee.com/us/threat_center/default.asp)  
> > Please rely on your own virus check as no responsibility is taken by the  
> > sender for any damage rising out of any virus infection this communication  
> > may contain.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:22am UTC](https://discuss.elastic.co/t/re-index-strategies/3071/6 "2017-07-06T04:22:36Z")

</div>


