# Best way to offload indexing from reading node

**URL:** <https://discuss.elastic.co/t/best-way-to-offload-indexing-from-reading-node/7238>\
**Category:** Elasticsearch\
**Created:** [April 5, 2012, 7:09am UTC](https://discuss.elastic.co/t/best-way-to-offload-indexing-from-reading-node/7238 "2012-04-05T07:09:02Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mikhail\_Sayapin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mikhail_sayapin/32/2758_2.png) [@Mikhail\_Sayapin](https://discuss.elastic.co/u/Mikhail_Sayapin)\
**Post date:** [April 5, 2012, 7:09am UTC](https://discuss.elastic.co/t/best-way-to-offload-indexing-from-reading-node/7238/1 "2012-04-05T07:09:02Z")

</div>

Hello,

I have an ecommerce store with a relatively large index of products. About  
10 000 000 products, 15 GB index size. This index gets updated very often,  
maybe 100-1000 updates per second.

What I'm trying to do is to setup two servers. One for rapidly updating the  
index, and another one just for querying. There is no problem if the data  
on the "querying" server is relatively stale (up to 4 hours is ok).

However I can't figure out a way to implement this. When I do replication,  
querying server is also suffering from intense writing i/o (even on XFS  
with a large buffer, or ext4 with commit=60,data=writeback), and sometimes  
the querying gets redirected to the "indexing" server.

So far my best guesses are:

1. Setup both nodes over the same Shared FS or Hadoop Gateway, write to  
"indexing" node, read from "querying" node (maybe even disable data  
alteration with 0.19.2 new APIs), don't let them join into cluster (?).  
Sometimes reboot "querying" node so it will recover from the gateway (maybe  
there is a better way to propagate changes?). Use "memory" index on  
"querying" node to prevent confusion over changed disk data (?).

2. Setup a simple cluster replication with 1 replica, and just use  
"preference=\_local" Search parameter to query local replica.

Would gladly accept any advice on how to achieve this. Thanks!

---

<div class="post-metadata">

**Author:** ![Berkay\_Mollamustafao](https://avatars.discourse-cdn.com/v4/letter/b/22d042/32.png) [@Berkay\_Mollamustafao](https://discuss.elastic.co/u/Berkay_Mollamustafao)\
**Post date:** [April 5, 2012, 8:32pm UTC](https://discuss.elastic.co/t/best-way-to-offload-indexing-from-reading-node/7238/2 "2012-04-05T20:32:12Z")

</div>

With 2 servers and 1 replica, both servers will have the same write load,  
regardless of which you use for queries. In your use case, you can disable  
auto refresh, refresh manually and use bulk API for writes.ES does not have  
write vs read server concept, so by far your best option is to improve  
performance is adding a 3rd server or more disks/cpu/memory, etc. I'd make  
sure you have a performance problem with queries before diving into  
attempting to force ES to work they way you're describing.

If you do have to go that route, you may try handling writes from your  
client code rather than relying on replication. Write to 1st server while  
querying the other for x minutes and then start reading from first server  
while bringing the second up to date, etc.

Regards,  
Berkay Mollamustafaoglu  
mberkay on yahoo, google and skype

On Thu, Apr 5, 2012 at 3:09 AM, Mikhail Sayapin  
[mikhail.sayapin@gmail.com](mailto:mikhail.sayapin@gmail.com)wrote:

> Hello,
> 
> I have an ecommerce store with a relatively large index of products. About  
> 10 000 000 products, 15 GB index size. This index gets updated very often,  
> maybe 100-1000 updates per second.
> 
> What I'm trying to do is to setup two servers. One for rapidly updating  
> the index, and another one just for querying. There is no problem if the  
> data on the "querying" server is relatively stale (up to 4 hours is ok).
> 
> However I can't figure out a way to implement this. When I do replication,  
> querying server is also suffering from intense writing i/o (even on XFS  
> with a large buffer, or ext4 with commit=60,data=writeback), and sometimes  
> the querying gets redirected to the "indexing" server.
> 
> So far my best guesses are:
> 
> 1. Setup both nodes over the same Shared FS or Hadoop Gateway, write to  
> "indexing" node, read from "querying" node (maybe even disable data  
> alteration with 0.19.2 new APIs), don't let them join into cluster (?).  
> Sometimes reboot "querying" node so it will recover from the gateway (maybe  
> there is a better way to propagate changes?). Use "memory" index on  
> "querying" node to prevent confusion over changed disk data (?).
> 
> 2. Setup a simple cluster replication with 1 replica, and just use  
> "preference=\_local" Search parameter to query local replica.
> 
> Would gladly accept any advice on how to achieve this. Thanks!

---

<div class="post-metadata">

**Author:** ![Mikhail\_Sayapin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mikhail_sayapin/32/2758_2.png) [@Mikhail\_Sayapin](https://discuss.elastic.co/u/Mikhail_Sayapin)\
**Post date:** [April 6, 2012, 3:53am UTC](https://discuss.elastic.co/t/best-way-to-offload-indexing-from-reading-node/7238/3 "2012-04-06T03:53:26Z")

</div>

Hello Berkay,

Thanks for the advice. I use bulk indexing with auto refresh disabled,  
and refresh from time to time based on 2nd server la. I mostly use  
constant\_score queries with terms filter, simple two-field numeric  
ordering. From time to time I build relatively large facets  
(3000-4000) with regex filters. Xms = Xmx = 16 GB. The disk is two  
software raid-1 SSDs, XFS mounted with noatime. Tried other gc options  
as well (settled on "-XX:+UseConcMarkSweepGC -XX:ParallelGCThreads=2  
-XX:+AggressiveOpts"). Queries take about 1.5-2 seconds, after that  
become cached, but 2 seconds is a bit too much anyway, I was able to  
reproduce dog piling effect.

What do you think about 1st option - using 2 disconnected (in terms of  
replication) servers over the same shared gateway location, 2nd server  
disabled for writing?

On Fri, Apr 6, 2012 at 4:32 AM, Berkay Mollamustafaoglu  
[mberkay@gmail.com](mailto:mberkay@gmail.com) wrote:

> With 2 servers and 1 replica, both servers will have the same write load,  
> regardless of which you use for queries. In your use case, you can disable  
> auto refresh, refresh manually and use bulk API for writes.ES does not have  
> write vs read server concept, so by far your best option is to improve  
> performance is adding a 3rd server or more disks/cpu/memory, etc. I'd make  
> sure you have a performance problem with queries before diving into  
> attempting to force ES to work they way you're describing.
> 
> If you do have to go that route, you may try handling writes from your  
> client code rather than relying on replication. Write to 1st server while  
> querying the other for x minutes and then start reading from first server  
> while bringing the second up to date, etc.
> 
> Regards,  
> Berkay Mollamustafaoglu  
> mberkay on yahoo, google and skype
> 
> On Thu, Apr 5, 2012 at 3:09 AM, Mikhail Sayapin [mikhail.sayapin@gmail.com](mailto:mikhail.sayapin@gmail.com)  
> wrote:
> 
> > Hello,
> > 
> > I have an ecommerce store with a relatively large index of products. About  
> > 10 000 000 products, 15 GB index size. This index gets updated very often,  
> > maybe 100-1000 updates per second.
> > 
> > What I'm trying to do is to setup two servers. One for rapidly updating  
> > the index, and another one just for querying. There is no problem if the  
> > data on the "querying" server is relatively stale (up to 4 hours is ok).
> > 
> > However I can't figure out a way to implement this. When I do replication,  
> > querying server is also suffering from intense writing i/o (even on XFS with  
> > a large buffer, or ext4 with commit=60,data=writeback), and sometimes the  
> > querying gets redirected to the "indexing" server.
> > 
> > So far my best guesses are:
> > 
> > 1. Setup both nodes over the same Shared FS or Hadoop Gateway, write to  
> > "indexing" node, read from "querying" node (maybe even disable data  
> > alteration with 0.19.2 new APIs), don't let them join into cluster (?).  
> > Sometimes reboot "querying" node so it will recover from the gateway (maybe  
> > there is a better way to propagate changes?). Use "memory" index on  
> > "querying" node to prevent confusion over changed disk data (?).
> > 
> > 2. Setup a simple cluster replication with 1 replica, and just use  
> > "preference=\_local" Search parameter to query local replica.
> > 
> > Would gladly accept any advice on how to achieve this. Thanks!

--  
Regards,  
Mikhail Sayapin  
["I recommend the art of slow reading."]

---

<div class="post-metadata">

**Author:** ![Berkay\_Mollamustafao](https://avatars.discourse-cdn.com/v4/letter/b/22d042/32.png) [@Berkay\_Mollamustafao](https://discuss.elastic.co/u/Berkay_Mollamustafao)\
**Post date:** [April 7, 2012, 1:51pm UTC](https://discuss.elastic.co/t/best-way-to-offload-indexing-from-reading-node/7238/4 "2012-04-07T13:51:51Z")

</div>

Hi Michail,

Don't think that would help. With shared gateway, ES still writes to local  
file system so both servers would have the same load plus the additional  
overhead of writing to the gateway.  
You may want to try the following:

- have a separate ES cluster at each server.
- writes go to 1 server and timestamp all writes
- have code on the 2nd server that runs periodically retrieves only changes  
and writes to the second cluster. you can use scan search type and bulk API  
to make this process efficient.
- use the 2nd server for queries

This way you can isolate most write traffic to the first server and update  
second server when you want to. Hope this helps ..

Regards,  
Berkay Mollamustafaoglu  
mberkay on yahoo, google and skype

On Thu, Apr 5, 2012 at 11:53 PM, Mikhail Sayapin  
[mikhail.sayapin@gmail.com](mailto:mikhail.sayapin@gmail.com)wrote:

> Hello Berkay,
> 
> Thanks for the advice. I use bulk indexing with auto refresh disabled,  
> and refresh from time to time based on 2nd server la. I mostly use  
> constant\_score queries with terms filter, simple two-field numeric  
> ordering. From time to time I build relatively large facets  
> (3000-4000) with regex filters. Xms = Xmx = 16 GB. The disk is two  
> software raid-1 SSDs, XFS mounted with noatime. Tried other gc options  
> as well (settled on "-XX:+UseConcMarkSweepGC -XX:ParallelGCThreads=2  
> -XX:+AggressiveOpts"). Queries take about 1.5-2 seconds, after that  
> become cached, but 2 seconds is a bit too much anyway, I was able to  
> reproduce dog piling effect.
> 
> What do you think about 1st option - using 2 disconnected (in terms of  
> replication) servers over the same shared gateway location, 2nd server  
> disabled for writing?
> 
> On Fri, Apr 6, 2012 at 4:32 AM, Berkay Mollamustafaoglu  
> [mberkay@gmail.com](mailto:mberkay@gmail.com) wrote:
> 
> > With 2 servers and 1 replica, both servers will have the same write load,  
> > regardless of which you use for queries. In your use case, you can  
> > disable  
> > auto refresh, refresh manually and use bulk API for writes.ES does not  
> > have  
> > write vs read server concept, so by far your best option is to improve  
> > performance is adding a 3rd server or more disks/cpu/memory, etc. I'd  
> > make  
> > sure you have a performance problem with queries before diving into  
> > attempting to force ES to work they way you're describing.
> > 
> > If you do have to go that route, you may try handling writes from your  
> > client code rather than relying on replication. Write to 1st server while  
> > querying the other for x minutes and then start reading from first server  
> > while bringing the second up to date, etc.
> > 
> > Regards,  
> > Berkay Mollamustafaoglu  
> > mberkay on yahoo, google and skype
> > 
> > On Thu, Apr 5, 2012 at 3:09 AM, Mikhail Sayapin \<  
> > [mikhail.sayapin@gmail.com](mailto:mikhail.sayapin@gmail.com)\>  
> > wrote:
> > 
> > > Hello,
> > > 
> > > I have an ecommerce store with a relatively large index of products.  
> > > About  
> > > 10 000 000 products, 15 GB index size. This index gets updated very  
> > > often,  
> > > maybe 100-1000 updates per second.
> > > 
> > > What I'm trying to do is to setup two servers. One for rapidly updating  
> > > the index, and another one just for querying. There is no problem if the  
> > > data on the "querying" server is relatively stale (up to 4 hours is ok).
> > > 
> > > However I can't figure out a way to implement this. When I do  
> > > replication,  
> > > querying server is also suffering from intense writing i/o (even on XFS  
> > > with  
> > > a large buffer, or ext4 with commit=60,data=writeback), and sometimes  
> > > the  
> > > querying gets redirected to the "indexing" server.
> > > 
> > > So far my best guesses are:
> > > 
> > > 1. Setup both nodes over the same Shared FS or Hadoop Gateway, write to  
> > > "indexing" node, read from "querying" node (maybe even disable data  
> > > alteration with 0.19.2 new APIs), don't let them join into cluster (?).  
> > > Sometimes reboot "querying" node so it will recover from the gateway  
> > > (maybe  
> > > there is a better way to propagate changes?). Use "memory" index on  
> > > "querying" node to prevent confusion over changed disk data (?).
> > > 
> > > 2. Setup a simple cluster replication with 1 replica, and just use  
> > > "preference=\_local" Search parameter to query local replica.
> > > 
> > > Would gladly accept any advice on how to achieve this. Thanks!
> 
> --  
> Regards,  
> Mikhail Sayapin  
> ["I recommend the art of slow reading."]

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [April 7, 2012, 3:33pm UTC](https://discuss.elastic.co/t/best-way-to-offload-indexing-from-reading-node/7238/5 "2012-04-07T15:33:14Z")

</div>

There isn't an option to have a slave that does not "index" documents in  
elasticsearch. There is a concept of replicating index segment files to  
slaves, but effectively, this is usually more expensive compared to  
indexing (which you can solve by adding more boxes). I talk about it a bit  
here:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

.

On Thu, Apr 5, 2012 at 10:09 AM, Mikhail Sayapin  
[mikhail.sayapin@gmail.com](mailto:mikhail.sayapin@gmail.com)wrote:

> Hello,
> 
> I have an ecommerce store with a relatively large index of products. About  
> 10 000 000 products, 15 GB index size. This index gets updated very often,  
> maybe 100-1000 updates per second.
> 
> What I'm trying to do is to setup two servers. One for rapidly updating  
> the index, and another one just for querying. There is no problem if the  
> data on the "querying" server is relatively stale (up to 4 hours is ok).
> 
> However I can't figure out a way to implement this. When I do replication,  
> querying server is also suffering from intense writing i/o (even on XFS  
> with a large buffer, or ext4 with commit=60,data=writeback), and sometimes  
> the querying gets redirected to the "indexing" server.
> 
> So far my best guesses are:
> 
> 1. Setup both nodes over the same Shared FS or Hadoop Gateway, write to  
> "indexing" node, read from "querying" node (maybe even disable data  
> alteration with 0.19.2 new APIs), don't let them join into cluster (?).  
> Sometimes reboot "querying" node so it will recover from the gateway (maybe  
> there is a better way to propagate changes?). Use "memory" index on  
> "querying" node to prevent confusion over changed disk data (?).
> 
> 2. Setup a simple cluster replication with 1 replica, and just use  
> "preference=\_local" Search parameter to query local replica.
> 
> Would gladly accept any advice on how to achieve this. Thanks!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:33am UTC](https://discuss.elastic.co/t/best-way-to-offload-indexing-from-reading-node/7238/6 "2017-07-06T03:33:22Z")

</div>


