# Warmers and IO

**URL:** <https://discuss.elastic.co/t/warmers-and-io/10716>\
**Category:** Elasticsearch\
**Created:** [February 12, 2013, 8:11pm UTC](https://discuss.elastic.co/t/warmers-and-io/10716 "2013-02-12T20:11:45Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Justin\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/justin_2/32/1977_2.png) [@Justin\_2](https://discuss.elastic.co/u/Justin_2)\
**Post date:** [February 12, 2013, 8:11pm UTC](https://discuss.elastic.co/t/warmers-and-io/10716/1 "2013-02-12T20:11:45Z")

</div>

Hello all,

I'm wondering if there are any pointers you could give me for search  
queries that return large result sets. The number of hits could be anywhere  
from 10,000 - 2,000,000. As indicated by the slowlog, the consuming part is  
fetching the data. I presume this fetching phase also includes sorting the  
data?

The initial query invocation may take upwards of 5 minutes. If I initiate  
the same subsequent query, it returns \< 500ms.

Would warmers be an appropriate solution here?

Thanks

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [February 13, 2013, 4:49am UTC](https://discuss.elastic.co/t/warmers-and-io/10716/2 "2013-02-13T04:49:49Z")

</div>

Hi,

Yes, if you know you will be running a query X and it is expensive or loads  
some data that will be reused by other queries, or initializes some data  
structures, and so on, then yes, using such a query for warming up is a  
good idea.

## Otis

ELASTICSEARCH Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)

On Tuesday, February 12, 2013 3:11:45 PM UTC-5, Justin wrote:

> Hello all,
> 
> I'm wondering if there are any pointers you could give me for search  
> queries that return large result sets. The number of hits could be anywhere  
> from 10,000 - 2,000,000. As indicated by the slowlog, the consuming part is  
> fetching the data. I presume this fetching phase also includes sorting the  
> data?
> 
> The initial query invocation may take upwards of 5 minutes. If I initiate  
> the same subsequent query, it returns \< 500ms.
> 
> Would warmers be an appropriate solution here?
> 
> Thanks

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Justin\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/justin_2/32/1977_2.png) [@Justin\_2](https://discuss.elastic.co/u/Justin_2)\
**Post date:** [February 13, 2013, 6:18am UTC](https://discuss.elastic.co/t/warmers-and-io/10716/3 "2013-02-13T06:18:48Z")

</div>

Thanks Otis

Unfortunately, we do not know the queries beforehand. BTW, they are text  
queries.

We had to act fast. It may not be pretty but this is the current solution  
we're using to "force" the data into the file system cache.

find /var/lib/elasticsearch/data -type f -exec cat {} ; \> /dev/null

We've yet to see any queries taking \> 500 ms - it seems to be working well.

Any thoughts on this approach?

On Tuesday, February 12, 2013 11:49:49 PM UTC-5, Otis Gospodnetic wrote:

> Hi,
> 
> Yes, if you know you will be running a query X and it is expensive or  
> loads some data that will be reused by other queries, or initializes some  
> data structures, and so on, then yes, using such a query for warming up is  
> a good idea.
> 
> ## Otis
> 
> ELASTICSEARCH Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> 
> On Tuesday, February 12, 2013 3:11:45 PM UTC-5, Justin wrote:
> 
> > Hello all,
> > 
> > I'm wondering if there are any pointers you could give me for search  
> > queries that return large result sets. The number of hits could be anywhere  
> > from 10,000 - 2,000,000. As indicated by the slowlog, the consuming part is  
> > fetching the data. I presume this fetching phase also includes sorting the  
> > data?
> > 
> > The initial query invocation may take upwards of 5 minutes. If I initiate  
> > the same subsequent query, it returns \< 500ms.
> > 
> > Would warmers be an appropriate solution here?
> > 
> > Thanks

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [February 13, 2013, 10:57am UTC](https://discuss.elastic.co/t/warmers-and-io/10716/4 "2013-02-13T10:57:17Z")

</div>

> We had to act fast. It may not be pretty but this is the current  
> solution we're using to "force" the data into the file system cache.
> 
> find /var/lib/elasticsearch/data -type f -exec cat {} ; \> /dev/null

That's just part of it. You also need to build field and filter caches  
inside ES.

The queries themselves are not cached, but there are bound to be filters  
that you use regularly which you can know in advance, eg

{ "term": { "status": "active"}}  
{ "date": { "range": { "from": "2013-01-0"}}}

Also any fields that you use in:

- sorting
- scripts
- facets

need to be loaded into the field cache, so include those in your warmers  
as well

clint

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [February 13, 2013, 11:33am UTC](https://discuss.elastic.co/t/warmers-and-io/10716/5 "2013-02-13T11:33:48Z")

</div>

Hi Justin,

What Clinton said. Plus, this is a good old trick, but if your index is  
larger than your RAM then it's only partially effective.

Otis  
Solr & Elasticsearch Support

> **[Sematext | IT System Monitoring Tools for DevOps](https://sematext.com/)**
>
> IT system monitoring and management tools for DevOps who need 24x7 live visibility into their infrastructure. Get started now with a 14-day free trial!

On Feb 13, 2013 1:18 AM, "Justin" [tcpandip@gmail.com](mailto:tcpandip@gmail.com) wrote:

> Thanks Otis
> 
> Unfortunately, we do not know the queries beforehand. BTW, they are text  
> queries.
> 
> We had to act fast. It may not be pretty but this is the current solution  
> we're using to "force" the data into the file system cache.
> 
> find /var/lib/elasticsearch/data -type f -exec cat {} ; \> /dev/null
> 
> We've yet to see any queries taking \> 500 ms - it seems to be working  
> well.
> 
> Any thoughts on this approach?
> 
> On Tuesday, February 12, 2013 11:49:49 PM UTC-5, Otis Gospodnetic wrote:
> 
> > Hi,
> > 
> > Yes, if you know you will be running a query X and it is expensive or  
> > loads some data that will be reused by other queries, or initializes some  
> > data structures, and so on, then yes, using such a query for warming up is  
> > a good idea.
> > 
> > ## Otis
> > 
> > ELASTICSEARCH Performance Monitoring - [http://sematext.com/spm/index.](http://sematext.com/spm/index.)\*\*  
> > html [http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > 
> > On Tuesday, February 12, 2013 3:11:45 PM UTC-5, Justin wrote:
> > 
> > > Hello all,
> > > 
> > > I'm wondering if there are any pointers you could give me for search  
> > > queries that return large result sets. The number of hits could be anywhere  
> > > from 10,000 - 2,000,000. As indicated by the slowlog, the consuming part is  
> > > fetching the data. I presume this fetching phase also includes sorting the  
> > > data?
> > > 
> > > The initial query invocation may take upwards of 5 minutes. If I  
> > > initiate the same subsequent query, it returns \< 500ms.
> > > 
> > > Would warmers be an appropriate solution here?
> > > 
> > > Thanks
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [February 13, 2013, 1:01pm UTC](https://discuss.elastic.co/t/warmers-and-io/10716/6 "2013-02-13T13:01:56Z")

</div>

A recommended method is

index.store.type: mmapfs

Your "find loop" loads everything from the data folder, but only once  
for the "cat" process, and even the files that may be not used by your  
elasticsearch workload. In contrast, mmapfs reads and keeps just the  
relevant files in memory and continue to let the OS VM manage the cache.  
Together with bootstrap.mlockall: true, such a page cache will stay  
perfectly in RAM until eviction (mostly the exit of elasticsearch  
process, assuming enough RAM), while your "find loop" loaded files will  
only be loaded in RAM once before they are evicted, so the find command  
would have to be repeated over and over again during the lifetime of the  
ES process. And this would boggle down your overall system performance.

If you want another "no need to think" solution you could also use ZFS  
with L2ARC (adaptive replacement cache, intro  
[http://dtrace.org/blogs/brendan/2008/07/22/zfs-l2arc/](http://dtrace.org/blogs/brendan/2008/07/22/zfs-l2arc/) ) where you don't  
have to tinker with your resources like RAM, fs cache, disks, SSD and so  
on - ZFS manages it for you.

Best regards,

Jörg

Am 13.02.13 07:18, schrieb Justin:

> We had to act fast. It may not be pretty but this is the current  
> solution we're using to "force" the data into the file system cache.
> 
> find /var/lib/elasticsearch/data -type f -exec cat {} ; \> /dev/null
> 
> We've yet to see any queries taking \> 500 ms - it seems to be working  
> well.
> 
> Any thoughts on this approach?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:51am UTC](https://discuss.elastic.co/t/warmers-and-io/10716/7 "2017-07-06T02:51:41Z")

</div>


