# Elasticsearch query performance

**URL:** <https://discuss.elastic.co/t/elasticsearch-query-performance/13235>\
**Category:** Elasticsearch\
**Created:** [August 16, 2013, 2:04am UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235 "2013-08-16T02:04:35Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![VB1](https://avatars.discourse-cdn.com/v4/letter/v/2acd7d/32.png) [@VB1](https://discuss.elastic.co/u/VB1)\
**Post date:** [August 16, 2013, 2:04am UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/1 "2013-08-16T02:04:35Z")

</div>

I'm using elasticsearch to index two types of objects -

Data details -

Contract object ~ 60 properties (Object size - 120 bytes)  
Risk Item Object ~ 125 properties (Object size - 250 bytes)

Contract is parent of risk item (\_parent)

I'm storing 240 million such objects in single index (210 million risk  
items, 30 million contracts)  
Index size is - 322 gb

Cluster details -

11 m2.4x.large EC2 boxes [68 gb memory, 1.6 TB storage, 8 cores](1 box is a  
load balancer node with node.data = false)  
50 shards  
1 replica

===  
elasticsearch.yml -

node.data: true

http.enabled: false

index.number\_of\_shards: 50

index.number\_of\_replicas: 1

index.translog.flush\_threshold\_ops: 10000

index.merge.policy.use\_compound\_files: false

indices.memory.index\_buffer\_size: 30%

index.refresh\_interval: 30s

index.store.type: mmapfs

path.data: /data-xvdf,/data-xvdg

===

I'm starting the elasticsearch nodes with following command -  
/home/ec2-user/elasticsearch-0.90.2/bin/elasticsearch -f -Xms30g -Xmx30g

My problem is that I'm running following query on risk item type and it is  
taking about 10-15 seconds to return data.

I'm running this with a load of 50 concurrent users and a bulk index load  
of about 5000 risk items happening in parallel.

Query -

http://:9200/contractindex/riskitem/\_search

{  
"query": {  
"has\_parent": {  
"parent\_type": "contract",  
"query": {  
"range": {  
"ContractDate": {  
"gte": "2010-01-01"  
}  
}  
}  
}  
},  
"filter": {  
"and": [{  
"query": {  
"bool": {  
"must": [{  
"query\_string": {  
"fields": ["RiskItemProperty1"],  
"query": "abc"  
}  
},  
{  
"query\_string": {  
"fields": ["RiskItemProperty2"],  
"query": "xyz"  
}  
}]  
}  
}  
}]  
}  
}

Can somebody please help me with how I can improve this query performance ?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [August 16, 2013, 5:28pm UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/2 "2013-08-16T17:28:09Z")

</div>

Can you profile the query without the indexing process happening in  
parallel? The index\_buffer\_size setting seems high compared to the default  
and your bulk load should only be just over a MB.

The has\_parent query could easily be turned into a filter so that you can  
take advantage of filtering caching. Is scoring important for that query? I  
am assuming it is not since it is a range query.

Cheers,

Ivan

On Thu, Aug 15, 2013 at 7:04 PM, VB [vishal.batghare@gmail.com](mailto:vishal.batghare@gmail.com) wrote:

> I'm using elasticsearch to index two types of objects -
> 
> Data details -
> 
> Contract object ~ 60 properties (Object size - 120 bytes)  
> Risk Item Object ~ 125 properties (Object size - 250 bytes)
> 
> Contract is parent of risk item (\_parent)
> 
> I'm storing 240 million such objects in single index (210 million risk  
> items, 30 million contracts)  
> Index size is - 322 gb
> 
> Cluster details -
> 
> 11 m2.4x.large EC2 boxes [68 gb memory, 1.6 TB storage, 8 cores](1 box is  
> a load balancer node with node.data = false)  
> 50 shards  
> 1 replica
> 
> ===  
> elasticsearch.yml -
> 
> node.data: true
> 
> http.enabled: false
> 
> index.number\_of\_shards: 50
> 
> index.number\_of\_replicas: 1
> 
> index.translog.flush\_threshold\_ops: 10000
> 
> index.merge.policy.use\_compound\_files: false
> 
> indices.memory.index\_buffer\_size: 30%
> 
> index.refresh\_interval: 30s
> 
> index.store.type: mmapfs
> 
> path.data: /data-xvdf,/data-xvdg
> 
> ===
> 
> I'm starting the elasticsearch nodes with following command -  
> /home/ec2-user/elasticsearch-0.90.2/bin/elasticsearch -f -Xms30g -Xmx30g
> 
> My problem is that I'm running following query on risk item type and it is  
> taking about 10-15 seconds to return data.
> 
> I'm running this with a load of 50 concurrent users and a bulk index load  
> of about 5000 risk items happening in parallel.
> 
> Query -
> 
> http://:9200/contractindex/riskitem/\_search
> 
> {  
> "query": {  
> "has\_parent": {  
> "parent\_type": "contract",  
> "query": {  
> "range": {  
> "ContractDate": {  
> "gte": "2010-01-01"  
> }  
> }  
> }  
> }  
> },  
> "filter": {  
> "and": [{  
> "query": {  
> "bool": {  
> "must": [{  
> "query\_string": {  
> "fields": ["RiskItemProperty1"],  
> "query": "abc"  
> }  
> },  
> {  
> "query\_string": {  
> "fields": ["RiskItemProperty2"],  
> "query": "xyz"  
> }  
> }]  
> }  
> }  
> }]  
> }  
> }
> 
> Can somebody please help me with how I can improve this query performance ?
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![VB1](https://avatars.discourse-cdn.com/v4/letter/v/2acd7d/32.png) [@VB1](https://discuss.elastic.co/u/VB1)\
**Post date:** [August 16, 2013, 8:50pm UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/3 "2013-08-16T20:50:51Z")

</div>

Ivan,

Thanks for the reply.

We are new to elasticsearch, and yes we did run search queries without  
indexing and and it still takes around 10 secs.

We can reduce buffer size or remove that setting from yml. Can we  
remove/change it after indexes are created or we need to create indexes  
again. Does it need server restart or we can call update setting API?

It would be highly appreciated if you can provide a filter version of the  
query and scoring is also not important.

Regards,  
VB

On Friday, 16 August 2013 10:28:09 UTC-7, Ivan Brusic wrote:

> Can you profile the query without the indexing process happening in  
> parallel? The index\_buffer\_size setting seems high compared to the default  
> and your bulk load should only be just over a MB.
> 
> The has\_parent query could easily be turned into a filter so that you can  
> take advantage of filtering caching. Is scoring important for that query? I  
> am assuming it is not since it is a range query.
> 
> Cheers,
> 
> Ivan
> 
> On Thu, Aug 15, 2013 at 7:04 PM, VB \<[vishal....@gmail.com](mailto:vishal....@gmail.com) \<javascript:\>\>wrote:
> 
> > I'm using elasticsearch to index two types of objects -
> > 
> > Data details -
> > 
> > Contract object ~ 60 properties (Object size - 120 bytes)  
> > Risk Item Object ~ 125 properties (Object size - 250 bytes)
> > 
> > Contract is parent of risk item (\_parent)
> > 
> > I'm storing 240 million such objects in single index (210 million risk  
> > items, 30 million contracts)  
> > Index size is - 322 gb
> > 
> > Cluster details -
> > 
> > 11 m2.4x.large EC2 boxes [68 gb memory, 1.6 TB storage, 8 cores](1 box is  
> > a load balancer node with node.data = false)  
> > 50 shards  
> > 1 replica
> > 
> > ===  
> > elasticsearch.yml -
> > 
> > node.data: true
> > 
> > http.enabled: false
> > 
> > index.number\_of\_shards: 50
> > 
> > index.number\_of\_replicas: 1
> > 
> > index.translog.flush\_threshold\_ops: 10000
> > 
> > index.merge.policy.use\_compound\_files: false
> > 
> > indices.memory.index\_buffer\_size: 30%
> > 
> > index.refresh\_interval: 30s
> > 
> > index.store.type: mmapfs
> > 
> > path.data: /data-xvdf,/data-xvdg
> > 
> > ===
> > 
> > I'm starting the elasticsearch nodes with following command -  
> > /home/ec2-user/elasticsearch-0.90.2/bin/elasticsearch -f -Xms30g -Xmx30g
> > 
> > My problem is that I'm running following query on risk item type and it  
> > is taking about 10-15 seconds to return data.
> > 
> > I'm running this with a load of 50 concurrent users and a bulk index load  
> > of about 5000 risk items happening in parallel.
> > 
> > Query -
> > 
> > http://:9200/contractindex/riskitem/\_search
> > 
> > {  
> > "query": {  
> > "has\_parent": {  
> > "parent\_type": "contract",  
> > "query": {  
> > "range": {  
> > "ContractDate": {  
> > "gte": "2010-01-01"  
> > }  
> > }  
> > }  
> > }  
> > },  
> > "filter": {  
> > "and": [{  
> > "query": {  
> > "bool": {  
> > "must": [{  
> > "query\_string": {  
> > "fields": ["RiskItemProperty1"],  
> > "query": "abc"  
> > }  
> > },  
> > {  
> > "query\_string": {  
> > "fields": ["RiskItemProperty2"],  
> > "query": "xyz"  
> > }  
> > }]  
> > }  
> > }  
> > }]  
> > }  
> > }
> > 
> > Can somebody please help me with how I can improve this query performance  
> > ?
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![VB1](https://avatars.discourse-cdn.com/v4/letter/v/2acd7d/32.png) [@VB1](https://discuss.elastic.co/u/VB1)\
**Post date:** [August 16, 2013, 10:20pm UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/4 "2013-08-16T22:20:26Z")

</div>

We have removed buffer size and restarted cluster/nodes.

Query is still taking around 10 seconds, CPU on all server is maxing out.

And tried looking for documentation to change has\_parent/has\_child queries  
to normal filter queries. Could not find anything, any inputs will be  
useful.

Regards,  
VB

On Friday, 16 August 2013 13:50:51 UTC-7, VB wrote:

> Ivan,
> 
> Thanks for the reply.
> 
> We are new to elasticsearch, and yes we did run search queries without  
> indexing and and it still takes around 10 secs.
> 
> We can reduce buffer size or remove that setting from yml. Can we  
> remove/change it after indexes are created or we need to create indexes  
> again. Does it need server restart or we can call update setting API?
> 
> It would be highly appreciated if you can provide a filter version of the  
> query and scoring is also not important.
> 
> Regards,  
> VB
> 
> On Friday, 16 August 2013 10:28:09 UTC-7, Ivan Brusic wrote:
> 
> > Can you profile the query without the indexing process happening in  
> > parallel? The index\_buffer\_size setting seems high compared to the default  
> > and your bulk load should only be just over a MB.
> > 
> > The has\_parent query could easily be turned into a filter so that you can  
> > take advantage of filtering caching. Is scoring important for that query? I  
> > am assuming it is not since it is a range query.
> > 
> > Cheers,
> > 
> > Ivan
> > 
> > On Thu, Aug 15, 2013 at 7:04 PM, VB [vishal....@gmail.com](mailto:vishal....@gmail.com) wrote:
> > 
> > > I'm using elasticsearch to index two types of objects -
> > > 
> > > Data details -
> > > 
> > > Contract object ~ 60 properties (Object size - 120 bytes)  
> > > Risk Item Object ~ 125 properties (Object size - 250 bytes)
> > > 
> > > Contract is parent of risk item (\_parent)
> > > 
> > > I'm storing 240 million such objects in single index (210 million risk  
> > > items, 30 million contracts)  
> > > Index size is - 322 gb
> > > 
> > > Cluster details -
> > > 
> > > 11 m2.4x.large EC2 boxes [68 gb memory, 1.6 TB storage, 8 cores](1 box  
> > > is a load balancer node with node.data = false)  
> > > 50 shards  
> > > 1 replica
> > > 
> > > ===  
> > > elasticsearch.yml -
> > > 
> > > node.data: true
> > > 
> > > http.enabled: false
> > > 
> > > index.number\_of\_shards: 50
> > > 
> > > index.number\_of\_replicas: 1
> > > 
> > > index.translog.flush\_threshold\_ops: 10000
> > > 
> > > index.merge.policy.use\_compound\_files: false
> > > 
> > > indices.memory.index\_buffer\_size: 30%
> > > 
> > > index.refresh\_interval: 30s
> > > 
> > > index.store.type: mmapfs
> > > 
> > > path.data: /data-xvdf,/data-xvdg
> > > 
> > > ===
> > > 
> > > I'm starting the elasticsearch nodes with following command -  
> > > /home/ec2-user/elasticsearch-0.90.2/bin/elasticsearch -f -Xms30g -Xmx30g
> > > 
> > > My problem is that I'm running following query on risk item type and it  
> > > is taking about 10-15 seconds to return data.
> > > 
> > > I'm running this with a load of 50 concurrent users and a bulk index  
> > > load of about 5000 risk items happening in parallel.
> > > 
> > > Query -
> > > 
> > > http://:9200/contractindex/riskitem/\_search
> > > 
> > > {  
> > > "query": {  
> > > "has\_parent": {  
> > > "parent\_type": "contract",  
> > > "query": {  
> > > "range": {  
> > > "ContractDate": {  
> > > "gte": "2010-01-01"  
> > > }  
> > > }  
> > > }  
> > > }  
> > > },  
> > > "filter": {  
> > > "and": [{  
> > > "query": {  
> > > "bool": {  
> > > "must": [{  
> > > "query\_string": {  
> > > "fields": ["RiskItemProperty1"],  
> > > "query": "abc"  
> > > }  
> > > },  
> > > {  
> > > "query\_string": {  
> > > "fields": ["RiskItemProperty2"],  
> > > "query": "xyz"  
> > > }  
> > > }]  
> > > }  
> > > }  
> > > }]  
> > > }  
> > > }
> > > 
> > > Can somebody please help me with how I can improve this query  
> > > performance ?
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google  
> > > Groups "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send  
> > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).  
> > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![roytmana](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/roytmana/32/44855_2.png) [@roytmana](https://discuss.elastic.co/u/roytmana)\
**Post date:** [August 17, 2013, 8:59pm UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/5 "2013-08-17T20:59:56Z")

</div>

Hi VB,

I do not know your use case but have you considered denormalizing your data? I other words storing parent object as part of its children json. Your query and facet performance wull be much better but the price to pay is having to update every child record if parent changes. Plus if majority of your searches need to return parent you would need to have some way of distincting single parent record out of potentially multiple hits on this parent/child. Still it may worth it. Your parent record is pretty small so the key is just how often it changes. Maybe bring only some of the parent fields you really need for serching into the child records ?  
Alex

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![VB1](https://avatars.discourse-cdn.com/v4/letter/v/2acd7d/32.png) [@VB1](https://discuss.elastic.co/u/VB1)\
**Post date:** [August 19, 2013, 9:23pm UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/6 "2013-08-19T21:23:35Z")

</div>

Alex we cannot go with denormalizing data, as you mentioned it would need  
to update each parent document for any change any attribute of the child  
document. Is there anything else you can propose.

In general also our queries from one table are also slower

This query takes around 8 seconds.

{  
"query": {  
"constant\_score": {  
"filter": {  
"and": [{  
"term": {  
"CommonCharacteristic\_BuildingScheme": "BuildingScheme1"  
}  
},  
{  
"term": {  
"Address\_Admin2Name": "Admin2Name1"  
}  
}]  
}  
}  
}  
}

This query takes around 6.5 seconds for Top 10 records ( but has sort on  
top of it)

{  
"query": {  
"constant\_score": {  
"filter": {  
"and": [{  
"term": {  
"Insurer": "Insurer1"  
}  
},  
{  
"term": {  
"Status": "Status1"  
}  
}]  
}  
}  
}  
}

But all our queries are with random values with few random set of data.

On Saturday, 17 August 2013 13:59:56 UTC-7, AlexR wrote:

> Hi VB,
> 
> I do not know your use case but have you considered denormalizing your  
> data? I other words storing parent object as part of its children json.  
> Your query and facet performance wull be much better but the price to pay  
> is having to update every child record if parent changes. Plus if majority  
> of your searches need to return parent you would need to have some way of  
> distincting single parent record out of potentially multiple hits on this  
> parent/child. Still it may worth it. Your parent record is pretty small so  
> the key is just how often it changes. Maybe bring only some of the parent  
> fields you really need for serching into the child records ?  
> Alex

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![mattweber](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mattweber/32/44940_2.png) [@mattweber](https://discuss.elastic.co/u/mattweber)\
**Post date:** [August 19, 2013, 9:37pm UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/7 "2013-08-19T21:37:38Z")

</div>

You should use a Bool Filter with must clauses, read this:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

{  
"query": {  
"constant\_score": {  
"filter": {  
"bool": {  
"must": [  
{"term": {"CommonCharacteristic\_BuildingScheme":  
"BuildingScheme1"}},  
{"term": {"Address\_Admin2Name": "Admin2Name1"}}  
]  
}  
}  
}  
}  
}

Thanks,  
Matt Weber

On Mon, Aug 19, 2013 at 2:23 PM, VB [vishal.batghare@gmail.com](mailto:vishal.batghare@gmail.com) wrote:

> Alex we cannot go with denormalizing data, as you mentioned it would need  
> to update each parent document for any change any attribute of the child  
> document. Is there anything else you can propose.
> 
> In general also our queries from one table are also slower
> 
> This query takes around 8 seconds.
> 
> {  
> "query": {  
> "constant\_score": {  
> "filter": {  
> "and": [{  
> "term": {  
> "CommonCharacteristic\_BuildingScheme": "BuildingScheme1"  
> }  
> },  
> {  
> "term": {  
> "Address\_Admin2Name": "Admin2Name1"  
> }  
> }]  
> }  
> }  
> }  
> }
> 
> This query takes around 6.5 seconds for Top 10 records ( but has sort on  
> top of it)
> 
> {  
> "query": {  
> "constant\_score": {  
> "filter": {  
> "and": [{  
> "term": {  
> "Insurer": "Insurer1"  
> }  
> },  
> {  
> "term": {  
> "Status": "Status1"  
> }  
> }]  
> }  
> }  
> }  
> }
> 
> But all our queries are with random values with few random set of data.
> 
> On Saturday, 17 August 2013 13:59:56 UTC-7, AlexR wrote:
> 
> > Hi VB,
> > 
> > I do not know your use case but have you considered denormalizing your  
> > data? I other words storing parent object as part of its children json.  
> > Your query and facet performance wull be much better but the price to pay  
> > is having to update every child record if parent changes. Plus if majority  
> > of your searches need to return parent you would need to have some way of  
> > distincting single parent record out of potentially multiple hits on this  
> > parent/child. Still it may worth it. Your parent record is pretty small so  
> > the key is just how often it changes. Maybe bring only some of the parent  
> > fields you really need for serching into the child records ?  
> > Alex
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![VB1](https://avatars.discourse-cdn.com/v4/letter/v/2acd7d/32.png) [@VB1](https://discuss.elastic.co/u/VB1)\
**Post date:** [August 19, 2013, 9:46pm UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/8 "2013-08-19T21:46:59Z")

</div>

Thanks Matt.

Can we use this on our parent child queries? and how to write parent child  
queries without using has\_parent/has\_child?

And is there a thumb rule about about when to use bool and when not use it?

Regards,  
VB

On Monday, 19 August 2013 14:37:38 UTC-7, Matt Weber wrote:

> You should use a Bool Filter with must clauses, read this:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/blog/all-about-elasticsearch-filter-bitsets/)
> 
> {  
> "query": {  
> "constant\_score": {  
> "filter": {  
> "bool": {  
> "must": [  
> {"term": {"CommonCharacteristic\_BuildingScheme":  
> "BuildingScheme1"}},  
> {"term": {"Address\_Admin2Name": "Admin2Name1"}}  
> ]  
> }  
> }  
> }  
> }  
> }
> 
> Thanks,  
> Matt Weber
> 
> On Mon, Aug 19, 2013 at 2:23 PM, VB \<[vishal....@gmail.com](mailto:vishal....@gmail.com) \<javascript:\>\>wrote:
> 
> > Alex we cannot go with denormalizing data, as you mentioned it would need  
> > to update each parent document for any change any attribute of the child  
> > document. Is there anything else you can propose.
> > 
> > In general also our queries from one table are also slower
> > 
> > This query takes around 8 seconds.
> > 
> > {  
> > "query": {  
> > "constant\_score": {  
> > "filter": {  
> > "and": [{  
> > "term": {  
> > "CommonCharacteristic\_BuildingScheme": "BuildingScheme1"  
> > }  
> > },  
> > {  
> > "term": {  
> > "Address\_Admin2Name": "Admin2Name1"  
> > }  
> > }]  
> > }  
> > }  
> > }  
> > }
> > 
> > This query takes around 6.5 seconds for Top 10 records ( but has sort on  
> > top of it)
> > 
> > {  
> > "query": {  
> > "constant\_score": {  
> > "filter": {  
> > "and": [{  
> > "term": {  
> > "Insurer": "Insurer1"  
> > }  
> > },  
> > {  
> > "term": {  
> > "Status": "Status1"  
> > }  
> > }]  
> > }  
> > }  
> > }  
> > }
> > 
> > But all our queries are with random values with few random set of data.
> > 
> > On Saturday, 17 August 2013 13:59:56 UTC-7, AlexR wrote:
> > 
> > > Hi VB,
> > > 
> > > I do not know your use case but have you considered denormalizing your  
> > > data? I other words storing parent object as part of its children json.  
> > > Your query and facet performance wull be much better but the price to pay  
> > > is having to update every child record if parent changes. Plus if majority  
> > > of your searches need to return parent you would need to have some way of  
> > > distincting single parent record out of potentially multiple hits on this  
> > > parent/child. Still it may worth it. Your parent record is pretty small so  
> > > the key is just how often it changes. Maybe bring only some of the parent  
> > > fields you really need for serching into the child records ?  
> > > Alex
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![roytmana](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/roytmana/32/44855_2.png) [@roytmana](https://discuss.elastic.co/u/roytmana)\
**Post date:** [August 19, 2013, 9:50pm UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/9 "2013-08-19T21:50:03Z")

</div>

I meant the opposite denormalize parent into child. You will not need to update parent on child change but all child on parent change which hopefully will be less frequent. And perhaps you only need some parent fields in the child which would make relevant pare nt changes less frequent.

I wonder if bool filter would make dramatic diff please let us know

Alex

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [August 19, 2013, 9:54pm UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/10 "2013-08-19T21:54:49Z")

</div>

With your ES cluster node config, you tell ES that it should fill 30g of  
heap for filter/cache. Do you use warming?

Another observation is that your index is 322g across 11 nodes, which makes  
~30g per node and you have assigned 64g - 30g = 34g to file system and  
other so your whole 322g files will fit into the file system cache.

My opinion is that 10s is blazingly fast to fill ~30g from the file system,  
prepare your filter query in the heap which may use up to another 30g, and  
execute the query plus delivering results.

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![mattweber](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mattweber/32/44940_2.png) [@mattweber](https://discuss.elastic.co/u/mattweber)\
**Post date:** [August 19, 2013, 10:00pm UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/11 "2013-08-19T22:00:42Z")

</div>

Yes, you can and should use bool query/filter with your parent/child  
queries. Read the article that I linked to know when and when not to use  
them. Looking at your original query, I would probably go with something  
like this:

{  
"query": {  
"filtered" : {  
"query": {  
"bool": {  
"must": [  
{"match": {"RiskItemProperty1": "abc"}},  
{"match": {"RiskItemProperty2": "xyz"}}  
]  
}  
},  
"filter": {  
"has\_parent": {  
"parent\_type": "contract",  
"filter": {  
"range": {  
"ContractDate": {  
"gte": "2010-01-01"  
}  
}  
}  
}  
}  
}  
}  
}

Remember that your first couple has\_parent or has\_child filters and queries  
are going to be slower due to id cache being loaded into memory.

Thanks,  
Matt Weber

On Mon, Aug 19, 2013 at 2:46 PM, VB [vishal.batghare@gmail.com](mailto:vishal.batghare@gmail.com) wrote:

> Thanks Matt.
> 
> Can we use this on our parent child queries? and how to write parent child  
> queries without using has\_parent/has\_child?
> 
> And is there a thumb rule about about when to use bool and when not use it?
> 
> Regards,  
> VB
> 
> On Monday, 19 August 2013 14:37:38 UTC-7, Matt Weber wrote:
> 
> > You should use a Bool Filter with must clauses, read this:  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/**blog/all-about-elasticsearch-)\*\*  
> > filter-bitsets/[http://www.elasticsearch.org/blog/all-about-elasticsearch-filter-bitsets/](http://www.elasticsearch.org/blog/all-about-elasticsearch-filter-bitsets/)
> > 
> > {  
> > "query": {  
> > "constant\_score": {  
> > "filter": {  
> > "bool": {  
> > "must": [  
> > {"term": {"CommonCharacteristic\_\*\*BuildingScheme":  
> > "BuildingScheme1"}},  
> > {"term": {"Address\_Admin2Name": "Admin2Name1"}}  
> > ]  
> > }  
> > }  
> > }  
> > }  
> > }
> > 
> > Thanks,  
> > Matt Weber
> > 
> > On Mon, Aug 19, 2013 at 2:23 PM, VB [vishal....@gmail.com](mailto:vishal....@gmail.com) wrote:
> > 
> > > Alex we cannot go with denormalizing data, as you mentioned it would  
> > > need to update each parent document for any change any attribute of the  
> > > child document. Is there anything else you can propose.
> > > 
> > > In general also our queries from one table are also slower
> > > 
> > > This query takes around 8 seconds.
> > > 
> > > {  
> > > "query": {  
> > > "constant\_score": {  
> > > "filter": {  
> > > "and": [{  
> > > "term": {  
> > > "CommonCharacteristic\_\*\*BuildingScheme": "BuildingScheme1"  
> > > }  
> > > },  
> > > {  
> > > "term": {  
> > > "Address\_Admin2Name": "Admin2Name1"  
> > > }  
> > > }]  
> > > }  
> > > }  
> > > }  
> > > }
> > > 
> > > This query takes around 6.5 seconds for Top 10 records ( but has sort on  
> > > top of it)
> > > 
> > > {  
> > > "query": {  
> > > "constant\_score": {  
> > > "filter": {  
> > > "and": [{  
> > > "term": {  
> > > "Insurer": "Insurer1"  
> > > }  
> > > },  
> > > {  
> > > "term": {  
> > > "Status": "Status1"  
> > > }  
> > > }]  
> > > }  
> > > }  
> > > }  
> > > }
> > > 
> > > But all our queries are with random values with few random set of data.
> > > 
> > > On Saturday, 17 August 2013 13:59:56 UTC-7, AlexR wrote:
> > > 
> > > > Hi VB,
> > > > 
> > > > I do not know your use case but have you considered denormalizing your  
> > > > data? I other words storing parent object as part of its children json.  
> > > > Your query and facet performance wull be much better but the price to pay  
> > > > is having to update every child record if parent changes. Plus if majority  
> > > > of your searches need to return parent you would need to have some way of  
> > > > distincting single parent record out of potentially multiple hits on this  
> > > > parent/child. Still it may worth it. Your parent record is pretty small so  
> > > > the key is just how often it changes. Maybe bring only some of the parent  
> > > > fields you really need for serching into the child records ?  
> > > > Alex
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google  
> > > Groups "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send  
> > > an email to elasticsearc...@\*\*[googlegroups.com](http://googlegroups.com).
> > > 
> > > For more options, visit [https://groups.google.com/\*\*groups/opt\_out](https://groups.google.com/**groups/opt_out)[https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> > > .
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![VB1](https://avatars.discourse-cdn.com/v4/letter/v/2acd7d/32.png) [@VB1](https://discuss.elastic.co/u/VB1)\
**Post date:** [August 19, 2013, 10:04pm UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/12 "2013-08-19T22:04:58Z")

</div>

Thanks Matt. I will run thorough this and post my observation.

On Monday, 19 August 2013 15:00:42 UTC-7, Matt Weber wrote:

> Yes, you can and should use bool query/filter with your parent/child  
> queries. Read the article that I linked to know when and when not to use  
> them. Looking at your original query, I would probably go with something  
> like this:
> 
> {  
> "query": {  
> "filtered" : {  
> "query": {  
> "bool": {  
> "must": [  
> {"match": {"RiskItemProperty1": "abc"}},  
> {"match": {"RiskItemProperty2": "xyz"}}  
> ]  
> }  
> },  
> "filter": {  
> "has\_parent": {  
> "parent\_type": "contract",  
> "filter": {  
> "range": {  
> "ContractDate": {  
> "gte": "2010-01-01"  
> }  
> }  
> }  
> }  
> }  
> }  
> }  
> }
> 
> Remember that your first couple has\_parent or has\_child filters and  
> queries are going to be slower due to id cache being loaded into memory.
> 
> Thanks,  
> Matt Weber
> 
> On Mon, Aug 19, 2013 at 2:46 PM, VB \<[vishal....@gmail.com](mailto:vishal....@gmail.com) \<javascript:\>\>wrote:
> 
> > Thanks Matt.
> > 
> > Can we use this on our parent child queries? and how to write parent  
> > child queries without using has\_parent/has\_child?
> > 
> > And is there a thumb rule about about when to use bool and when not use  
> > it?
> > 
> > Regards,  
> > VB
> > 
> > On Monday, 19 August 2013 14:37:38 UTC-7, Matt Weber wrote:
> > 
> > > You should use a Bool Filter with must clauses, read this:  
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/**blog/all-about-elasticsearch-)\*\*  
> > > filter-bitsets/[http://www.elasticsearch.org/blog/all-about-elasticsearch-filter-bitsets/](http://www.elasticsearch.org/blog/all-about-elasticsearch-filter-bitsets/)
> > > 
> > > {  
> > > "query": {  
> > > "constant\_score": {  
> > > "filter": {  
> > > "bool": {  
> > > "must": [  
> > > {"term": {"CommonCharacteristic\_\*\*BuildingScheme":  
> > > "BuildingScheme1"}},  
> > > {"term": {"Address\_Admin2Name": "Admin2Name1"}}  
> > > ]  
> > > }  
> > > }  
> > > }  
> > > }  
> > > }
> > > 
> > > Thanks,  
> > > Matt Weber
> > > 
> > > On Mon, Aug 19, 2013 at 2:23 PM, VB [vishal....@gmail.com](mailto:vishal....@gmail.com) wrote:
> > > 
> > > > Alex we cannot go with denormalizing data, as you mentioned it would  
> > > > need to update each parent document for any change any attribute of the  
> > > > child document. Is there anything else you can propose.
> > > > 
> > > > In general also our queries from one table are also slower
> > > > 
> > > > This query takes around 8 seconds.
> > > > 
> > > > {  
> > > > "query": {  
> > > > "constant\_score": {  
> > > > "filter": {  
> > > > "and": [{  
> > > > "term": {  
> > > > "CommonCharacteristic\_\*\*BuildingScheme": "BuildingScheme1"  
> > > > }  
> > > > },  
> > > > {  
> > > > "term": {  
> > > > "Address\_Admin2Name": "Admin2Name1"  
> > > > }  
> > > > }]  
> > > > }  
> > > > }  
> > > > }  
> > > > }
> > > > 
> > > > This query takes around 6.5 seconds for Top 10 records ( but has sort  
> > > > on top of it)
> > > > 
> > > > {  
> > > > "query": {  
> > > > "constant\_score": {  
> > > > "filter": {  
> > > > "and": [{  
> > > > "term": {  
> > > > "Insurer": "Insurer1"  
> > > > }  
> > > > },  
> > > > {  
> > > > "term": {  
> > > > "Status": "Status1"  
> > > > }  
> > > > }]  
> > > > }  
> > > > }  
> > > > }  
> > > > }
> > > > 
> > > > But all our queries are with random values with few random set of data.
> > > > 
> > > > On Saturday, 17 August 2013 13:59:56 UTC-7, AlexR wrote:
> > > > 
> > > > > Hi VB,
> > > > > 
> > > > > I do not know your use case but have you considered denormalizing your  
> > > > > data? I other words storing parent object as part of its children json.  
> > > > > Your query and facet performance wull be much better but the price to pay  
> > > > > is having to update every child record if parent changes. Plus if majority  
> > > > > of your searches need to return parent you would need to have some way of  
> > > > > distincting single parent record out of potentially multiple hits on this  
> > > > > parent/child. Still it may worth it. Your parent record is pretty small so  
> > > > > the key is just how often it changes. Maybe bring only some of the parent  
> > > > > fields you really need for serching into the child records ?  
> > > > > Alex
> > > > 
> > > > --  
> > > > You received this message because you are subscribed to the Google  
> > > > Groups "elasticsearch" group.  
> > > > To unsubscribe from this group and stop receiving emails from it, send  
> > > > an email to elasticsearc...@\*\*[googlegroups.com](http://googlegroups.com).
> > > > 
> > > > For more options, visit [https://groups.google.com/\*\*groups/opt\_out](https://groups.google.com/**groups/opt_out)[https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> > > > .
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google Groups  
> > > "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send an  
> > > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![VB1](https://avatars.discourse-cdn.com/v4/letter/v/2acd7d/32.png) [@VB1](https://discuss.elastic.co/u/VB1)\
**Post date:** [August 20, 2013, 5:24pm UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/13 "2013-08-20T17:24:48Z")

</div>

Hi all,

We made changes as suggested by Matt to use bitsets.

We ran 50 concurrent users (Read Only) for an hour. All our queries are  
performing 4 to 5 times faster, except parent child query (query in  
question) it has gone down from 7 seconds to 3 seconds.

Matt, thank you so much fort helping us. Is there anything else we can do  
in parent child one or in general.

I have one more query with has\_child in it. Do you think we can further  
improve this one?

{  
"query": {  
"filtered": {  
"query": {  
"bool": {  
"must": [{  
"match": {  
"LineOfBusiness": "LOBValue1"  
}  
}]  
}  
},  
"filter": {  
"has\_child": {  
"type": "riskitem",  
"filter": {  
"bool": {  
"must": [{  
"term": {  
"Address\_Admin1Name": "Admin1Name1"  
}  
}]  
}  
}  
}  
}  
}  
}  
}

Regards,  
VB.

On Monday, 19 August 2013 15:04:58 UTC-7, VB wrote:

> Thanks Matt. I will run thorough this and post my observation.
> 
> On Monday, 19 August 2013 15:00:42 UTC-7, Matt Weber wrote:
> 
> > Yes, you can and should use bool query/filter with your parent/child  
> > queries. Read the article that I linked to know when and when not to use  
> > them. Looking at your original query, I would probably go with something  
> > like this:
> > 
> > {  
> > "query": {  
> > "filtered" : {  
> > "query": {  
> > "bool": {  
> > "must": [  
> > {"match": {"RiskItemProperty1": "abc"}},  
> > {"match": {"RiskItemProperty2": "xyz"}}  
> > ]  
> > }  
> > },  
> > "filter": {  
> > "has\_parent": {  
> > "parent\_type": "contract",  
> > "filter": {  
> > "range": {  
> > "ContractDate": {  
> > "gte": "2010-01-01"  
> > }  
> > }  
> > }  
> > }  
> > }  
> > }  
> > }  
> > }
> > 
> > Remember that your first couple has\_parent or has\_child filters and  
> > queries are going to be slower due to id cache being loaded into memory.
> > 
> > Thanks,  
> > Matt Weber
> > 
> > On Mon, Aug 19, 2013 at 2:46 PM, VB [vishal....@gmail.com](mailto:vishal....@gmail.com) wrote:
> > 
> > > Thanks Matt.
> > > 
> > > Can we use this on our parent child queries? and how to write parent  
> > > child queries without using has\_parent/has\_child?
> > > 
> > > And is there a thumb rule about about when to use bool and when not use  
> > > it?
> > > 
> > > Regards,  
> > > VB
> > > 
> > > On Monday, 19 August 2013 14:37:38 UTC-7, Matt Weber wrote:
> > > 
> > > > You should use a Bool Filter with must clauses, read this:  
> > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/**blog/all-about-elasticsearch-)\*\*  
> > > > filter-bitsets/[http://www.elasticsearch.org/blog/all-about-elasticsearch-filter-bitsets/](http://www.elasticsearch.org/blog/all-about-elasticsearch-filter-bitsets/)
> > > > 
> > > > {  
> > > > "query": {  
> > > > "constant\_score": {  
> > > > "filter": {  
> > > > "bool": {  
> > > > "must": [  
> > > > {"term": {"CommonCharacteristic\_\*\*BuildingScheme":  
> > > > "BuildingScheme1"}},  
> > > > {"term": {"Address\_Admin2Name": "Admin2Name1"}}  
> > > > ]  
> > > > }  
> > > > }  
> > > > }  
> > > > }  
> > > > }
> > > > 
> > > > Thanks,  
> > > > Matt Weber
> > > > 
> > > > On Mon, Aug 19, 2013 at 2:23 PM, VB [vishal....@gmail.com](mailto:vishal....@gmail.com) wrote:
> > > > 
> > > > > Alex we cannot go with denormalizing data, as you mentioned it would  
> > > > > need to update each parent document for any change any attribute of the  
> > > > > child document. Is there anything else you can propose.
> > > > > 
> > > > > In general also our queries from one table are also slower
> > > > > 
> > > > > This query takes around 8 seconds.
> > > > > 
> > > > > {  
> > > > > "query": {  
> > > > > "constant\_score": {  
> > > > > "filter": {  
> > > > > "and": [{  
> > > > > "term": {  
> > > > > "CommonCharacteristic\_\*\*BuildingScheme": "BuildingScheme1"  
> > > > > }  
> > > > > },  
> > > > > {  
> > > > > "term": {  
> > > > > "Address\_Admin2Name": "Admin2Name1"  
> > > > > }  
> > > > > }]  
> > > > > }  
> > > > > }  
> > > > > }  
> > > > > }
> > > > > 
> > > > > This query takes around 6.5 seconds for Top 10 records ( but has sort  
> > > > > on top of it)
> > > > > 
> > > > > {  
> > > > > "query": {  
> > > > > "constant\_score": {  
> > > > > "filter": {  
> > > > > "and": [{  
> > > > > "term": {  
> > > > > "Insurer": "Insurer1"  
> > > > > }  
> > > > > },  
> > > > > {  
> > > > > "term": {  
> > > > > "Status": "Status1"  
> > > > > }  
> > > > > }]  
> > > > > }  
> > > > > }  
> > > > > }  
> > > > > }
> > > > > 
> > > > > But all our queries are with random values with few random set of data.
> > > > > 
> > > > > On Saturday, 17 August 2013 13:59:56 UTC-7, AlexR wrote:
> > > > > 
> > > > > > Hi VB,
> > > > > > 
> > > > > > I do not know your use case but have you considered denormalizing  
> > > > > > your data? I other words storing parent object as part of its children  
> > > > > > json. Your query and facet performance wull be much better but the price to  
> > > > > > pay is having to update every child record if parent changes. Plus if  
> > > > > > majority of your searches need to return parent you would need to have some  
> > > > > > way of distincting single parent record out of potentially multiple hits on  
> > > > > > this parent/child. Still it may worth it. Your parent record is pretty  
> > > > > > small so the key is just how often it changes. Maybe bring only some of the  
> > > > > > parent fields you really need for serching into the child records ?  
> > > > > > Alex
> > > > > 
> > > > > --  
> > > > > You received this message because you are subscribed to the Google  
> > > > > Groups "elasticsearch" group.  
> > > > > To unsubscribe from this group and stop receiving emails from it, send  
> > > > > an email to elasticsearc...@\*\*[googlegroups.com](http://googlegroups.com).
> > > > > 
> > > > > For more options, visit [https://groups.google.com/\*\*groups/opt\_out](https://groups.google.com/**groups/opt_out)[https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> > > > > .
> > > > 
> > > > --  
> > > > You received this message because you are subscribed to the Google  
> > > > Groups "elasticsearch" group.  
> > > > To unsubscribe from this group and stop receiving emails from it, send  
> > > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).  
> > > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![VB1](https://avatars.discourse-cdn.com/v4/letter/v/2acd7d/32.png) [@VB1](https://discuss.elastic.co/u/VB1)\
**Post date:** [August 20, 2013, 6:33pm UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/14 "2013-08-20T18:33:39Z")

</div>

And one more of this type which needs improvement.

{  
"query": {  
"bool": {  
"must": [{  
"range": {  
"InceptionDate": {  
"gt": "2009-01-01"  
}  
}  
},  
{  
"range": {  
"ExpirationDate": {  
"lt": "2013-01-01"  
}  
}  
},  
{  
"has\_child": {  
"type": "riskitem",  
"query": {  
"filtered": {  
"filter": {  
"or": [{  
"bool": {  
"must": [{  
"term": {  
"Address\_Admin2Name": "tureni"  
}  
},  
{  
"term": {  
"Address\_Admin2Name\_US": "burlington"  
}  
},  
{  
"term": {  
"CommonCharacteristic\_BuildingClass": "62"  
}  
}]  
}  
},  
{  
"bool": {  
"must": [{  
"term": {  
"CommonCharacteristic\_ConstructionName": "heavy"  
}  
},  
{  
"term": {  
"CommonCharacteristic\_BuildingScheme": "rms"  
}  
},  
{  
"terms": {  
"CommonCharacteristic\_ValuationType": ["reported",  
"reported"]  
}  
}]  
}  
}]  
}  
}  
}  
}  
}]  
}  
}  
}

On Tuesday, 20 August 2013 10:24:48 UTC-7, VB wrote:

> Hi all,
> 
> We made changes as suggested by Matt to use bitsets.
> 
> We ran 50 concurrent users (Read Only) for an hour. All our queries are  
> performing 4 to 5 times faster, except parent child query (query in  
> question) it has gone down from 7 seconds to 3 seconds.
> 
> Matt, thank you so much fort helping us. Is there anything else we can do  
> in parent child one or in general.
> 
> I have one more query with has\_child in it. Do you think we can further  
> improve this one?
> 
> {  
> "query": {  
> "filtered": {  
> "query": {  
> "bool": {  
> "must": [{  
> "match": {  
> "LineOfBusiness": "LOBValue1"  
> }  
> }]  
> }  
> },  
> "filter": {  
> "has\_child": {  
> "type": "riskitem",  
> "filter": {  
> "bool": {  
> "must": [{  
> "term": {  
> "Address\_Admin1Name": "Admin1Name1"  
> }  
> }]  
> }  
> }  
> }  
> }  
> }  
> }  
> }
> 
> Regards,  
> VB.
> 
> On Monday, 19 August 2013 15:04:58 UTC-7, VB wrote:
> 
> > Thanks Matt. I will run thorough this and post my observation.
> > 
> > On Monday, 19 August 2013 15:00:42 UTC-7, Matt Weber wrote:
> > 
> > > Yes, you can and should use bool query/filter with your parent/child  
> > > queries. Read the article that I linked to know when and when not to use  
> > > them. Looking at your original query, I would probably go with something  
> > > like this:
> > > 
> > > {  
> > > "query": {  
> > > "filtered" : {  
> > > "query": {  
> > > "bool": {  
> > > "must": [  
> > > {"match": {"RiskItemProperty1": "abc"}},  
> > > {"match": {"RiskItemProperty2": "xyz"}}  
> > > ]  
> > > }  
> > > },  
> > > "filter": {  
> > > "has\_parent": {  
> > > "parent\_type": "contract",  
> > > "filter": {  
> > > "range": {  
> > > "ContractDate": {  
> > > "gte": "2010-01-01"  
> > > }  
> > > }  
> > > }  
> > > }  
> > > }  
> > > }  
> > > }  
> > > }
> > > 
> > > Remember that your first couple has\_parent or has\_child filters and  
> > > queries are going to be slower due to id cache being loaded into memory.
> > > 
> > > Thanks,  
> > > Matt Weber
> > > 
> > > On Mon, Aug 19, 2013 at 2:46 PM, VB [vishal....@gmail.com](mailto:vishal....@gmail.com) wrote:
> > > 
> > > > Thanks Matt.
> > > > 
> > > > Can we use this on our parent child queries? and how to write parent  
> > > > child queries without using has\_parent/has\_child?
> > > > 
> > > > And is there a thumb rule about about when to use bool and when not use  
> > > > it?
> > > > 
> > > > Regards,  
> > > > VB
> > > > 
> > > > On Monday, 19 August 2013 14:37:38 UTC-7, Matt Weber wrote:
> > > > 
> > > > > You should use a Bool Filter with must clauses, read this:  
> > > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/**blog/all-about-elasticsearch-)\*\*  
> > > > > filter-bitsets/[http://www.elasticsearch.org/blog/all-about-elasticsearch-filter-bitsets/](http://www.elasticsearch.org/blog/all-about-elasticsearch-filter-bitsets/)
> > > > > 
> > > > > {  
> > > > > "query": {  
> > > > > "constant\_score": {  
> > > > > "filter": {  
> > > > > "bool": {  
> > > > > "must": [  
> > > > > {"term": {"CommonCharacteristic\_\*\*BuildingScheme":  
> > > > > "BuildingScheme1"}},  
> > > > > {"term": {"Address\_Admin2Name": "Admin2Name1"}}  
> > > > > ]  
> > > > > }  
> > > > > }  
> > > > > }  
> > > > > }  
> > > > > }
> > > > > 
> > > > > Thanks,  
> > > > > Matt Weber
> > > > > 
> > > > > On Mon, Aug 19, 2013 at 2:23 PM, VB [vishal....@gmail.com](mailto:vishal....@gmail.com) wrote:
> > > > > 
> > > > > > Alex we cannot go with denormalizing data, as you mentioned it  
> > > > > > would need to update each parent document for any change any attribute of  
> > > > > > the child document. Is there anything else you can propose.
> > > > > > 
> > > > > > In general also our queries from one table are also slower
> > > > > > 
> > > > > > This query takes around 8 seconds.
> > > > > > 
> > > > > > {  
> > > > > > "query": {  
> > > > > > "constant\_score": {  
> > > > > > "filter": {  
> > > > > > "and": [{  
> > > > > > "term": {  
> > > > > > "CommonCharacteristic\_\*\*BuildingScheme": "BuildingScheme1"  
> > > > > > }  
> > > > > > },  
> > > > > > {  
> > > > > > "term": {  
> > > > > > "Address\_Admin2Name": "Admin2Name1"  
> > > > > > }  
> > > > > > }]  
> > > > > > }  
> > > > > > }  
> > > > > > }  
> > > > > > }
> > > > > > 
> > > > > > This query takes around 6.5 seconds for Top 10 records ( but has sort  
> > > > > > on top of it)
> > > > > > 
> > > > > > {  
> > > > > > "query": {  
> > > > > > "constant\_score": {  
> > > > > > "filter": {  
> > > > > > "and": [{  
> > > > > > "term": {  
> > > > > > "Insurer": "Insurer1"  
> > > > > > }  
> > > > > > },  
> > > > > > {  
> > > > > > "term": {  
> > > > > > "Status": "Status1"  
> > > > > > }  
> > > > > > }]  
> > > > > > }  
> > > > > > }  
> > > > > > }  
> > > > > > }
> > > > > > 
> > > > > > But all our queries are with random values with few random set of  
> > > > > > data.
> > > > > > 
> > > > > > On Saturday, 17 August 2013 13:59:56 UTC-7, AlexR wrote:
> > > > > > 
> > > > > > > Hi VB,
> > > > > > > 
> > > > > > > I do not know your use case but have you considered denormalizing  
> > > > > > > your data? I other words storing parent object as part of its children  
> > > > > > > json. Your query and facet performance wull be much better but the price to  
> > > > > > > pay is having to update every child record if parent changes. Plus if  
> > > > > > > majority of your searches need to return parent you would need to have some  
> > > > > > > way of distincting single parent record out of potentially multiple hits on  
> > > > > > > this parent/child. Still it may worth it. Your parent record is pretty  
> > > > > > > small so the key is just how often it changes. Maybe bring only some of the  
> > > > > > > parent fields you really need for serching into the child records ?  
> > > > > > > Alex
> > > > > > 
> > > > > > --  
> > > > > > You received this message because you are subscribed to the Google  
> > > > > > Groups "elasticsearch" group.  
> > > > > > To unsubscribe from this group and stop receiving emails from it,  
> > > > > > send an email to elasticsearc...@\*\*[googlegroups.com](http://googlegroups.com).
> > > > > > 
> > > > > > For more options, visit [https://groups.google.com/\*\*groups/opt\_out](https://groups.google.com/**groups/opt_out)[https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> > > > > > .
> > > > > 
> > > > > --  
> > > > > You received this message because you are subscribed to the Google  
> > > > > Groups "elasticsearch" group.  
> > > > > To unsubscribe from this group and stop receiving emails from it, send  
> > > > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).  
> > > > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![VB1](https://avatars.discourse-cdn.com/v4/letter/v/2acd7d/32.png) [@VB1](https://discuss.elastic.co/u/VB1)\
**Post date:** [August 21, 2013, 5:13pm UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/15 "2013-08-21T17:13:14Z")

</div>

Can anyone please comment/help?

On Tuesday, 20 August 2013 11:33:39 UTC-7, VB wrote:

> And one more of this type which needs improvement.
> 
> {  
> "query": {  
> "bool": {  
> "must": [{  
> "range": {  
> "InceptionDate": {  
> "gt": "2009-01-01"  
> }  
> }  
> },  
> {  
> "range": {  
> "ExpirationDate": {  
> "lt": "2013-01-01"  
> }  
> }  
> },  
> {  
> "has\_child": {  
> "type": "riskitem",  
> "query": {  
> "filtered": {  
> "filter": {  
> "or": [{  
> "bool": {  
> "must": [{  
> "term": {  
> "Address\_Admin2Name": "tureni"  
> }  
> },  
> {  
> "term": {  
> "Address\_Admin2Name\_US": "burlington"  
> }  
> },  
> {  
> "term": {  
> "CommonCharacteristic\_BuildingClass": "62"  
> }  
> }]  
> }  
> },  
> {  
> "bool": {  
> "must": [{  
> "term": {  
> "CommonCharacteristic\_ConstructionName": "heavy"  
> }  
> },  
> {  
> "term": {  
> "CommonCharacteristic\_BuildingScheme": "rms"  
> }  
> },  
> {  
> "terms": {  
> "CommonCharacteristic\_ValuationType": ["reported",  
> "reported"]  
> }  
> }]  
> }  
> }]  
> }  
> }  
> }  
> }  
> }]  
> }  
> }  
> }
> 
> On Tuesday, 20 August 2013 10:24:48 UTC-7, VB wrote:
> 
> > Hi all,
> > 
> > We made changes as suggested by Matt to use bitsets.
> > 
> > We ran 50 concurrent users (Read Only) for an hour. All our queries are  
> > performing 4 to 5 times faster, except parent child query (query in  
> > question) it has gone down from 7 seconds to 3 seconds.
> > 
> > Matt, thank you so much fort helping us. Is there anything else we can do  
> > in parent child one or in general.
> > 
> > I have one more query with has\_child in it. Do you think we can further  
> > improve this one?
> > 
> > {  
> > "query": {  
> > "filtered": {  
> > "query": {  
> > "bool": {  
> > "must": [{  
> > "match": {  
> > "LineOfBusiness": "LOBValue1"  
> > }  
> > }]  
> > }  
> > },  
> > "filter": {  
> > "has\_child": {  
> > "type": "riskitem",  
> > "filter": {  
> > "bool": {  
> > "must": [{  
> > "term": {  
> > "Address\_Admin1Name": "Admin1Name1"  
> > }  
> > }]  
> > }  
> > }  
> > }  
> > }  
> > }  
> > }  
> > }
> > 
> > Regards,  
> > VB.
> > 
> > On Monday, 19 August 2013 15:04:58 UTC-7, VB wrote:
> > 
> > > Thanks Matt. I will run thorough this and post my observation.
> > > 
> > > On Monday, 19 August 2013 15:00:42 UTC-7, Matt Weber wrote:
> > > 
> > > > Yes, you can and should use bool query/filter with your parent/child  
> > > > queries. Read the article that I linked to know when and when not to use  
> > > > them. Looking at your original query, I would probably go with something  
> > > > like this:
> > > > 
> > > > {  
> > > > "query": {  
> > > > "filtered" : {  
> > > > "query": {  
> > > > "bool": {  
> > > > "must": [  
> > > > {"match": {"RiskItemProperty1": "abc"}},  
> > > > {"match": {"RiskItemProperty2": "xyz"}}  
> > > > ]  
> > > > }  
> > > > },  
> > > > "filter": {  
> > > > "has\_parent": {  
> > > > "parent\_type": "contract",  
> > > > "filter": {  
> > > > "range": {  
> > > > "ContractDate": {  
> > > > "gte": "2010-01-01"  
> > > > }  
> > > > }  
> > > > }  
> > > > }  
> > > > }  
> > > > }  
> > > > }  
> > > > }
> > > > 
> > > > Remember that your first couple has\_parent or has\_child filters and  
> > > > queries are going to be slower due to id cache being loaded into memory.
> > > > 
> > > > Thanks,  
> > > > Matt Weber
> > > > 
> > > > On Mon, Aug 19, 2013 at 2:46 PM, VB [vishal....@gmail.com](mailto:vishal....@gmail.com) wrote:
> > > > 
> > > > > Thanks Matt.
> > > > > 
> > > > > Can we use this on our parent child queries? and how to write parent  
> > > > > child queries without using has\_parent/has\_child?
> > > > > 
> > > > > And is there a thumb rule about about when to use bool and when not  
> > > > > use it?
> > > > > 
> > > > > Regards,  
> > > > > VB
> > > > > 
> > > > > On Monday, 19 August 2013 14:37:38 UTC-7, Matt Weber wrote:
> > > > > 
> > > > > > You should use a Bool Filter with must clauses, read this:  
> > > > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/**blog/all-about-elasticsearch-)\*\*  
> > > > > > filter-bitsets/[http://www.elasticsearch.org/blog/all-about-elasticsearch-filter-bitsets/](http://www.elasticsearch.org/blog/all-about-elasticsearch-filter-bitsets/)
> > > > > > 
> > > > > > {  
> > > > > > "query": {  
> > > > > > "constant\_score": {  
> > > > > > "filter": {  
> > > > > > "bool": {  
> > > > > > "must": [  
> > > > > > {"term": {"CommonCharacteristic\_\*\*BuildingScheme":  
> > > > > > "BuildingScheme1"}},  
> > > > > > {"term": {"Address\_Admin2Name":  
> > > > > > "Admin2Name1"}}  
> > > > > > ]  
> > > > > > }  
> > > > > > }  
> > > > > > }  
> > > > > > }  
> > > > > > }
> > > > > > 
> > > > > > Thanks,  
> > > > > > Matt Weber
> > > > > > 
> > > > > > On Mon, Aug 19, 2013 at 2:23 PM, VB [vishal....@gmail.com](mailto:vishal....@gmail.com) wrote:
> > > > > > 
> > > > > > > Alex we cannot go with denormalizing data, as you mentioned it  
> > > > > > > would need to update each parent document for any change any attribute of  
> > > > > > > the child document. Is there anything else you can propose.
> > > > > > > 
> > > > > > > In general also our queries from one table are also slower
> > > > > > > 
> > > > > > > This query takes around 8 seconds.
> > > > > > > 
> > > > > > > {  
> > > > > > > "query": {  
> > > > > > > "constant\_score": {  
> > > > > > > "filter": {  
> > > > > > > "and": [{  
> > > > > > > "term": {  
> > > > > > > "CommonCharacteristic\_\*\*BuildingScheme": "BuildingScheme1"  
> > > > > > > }  
> > > > > > > },  
> > > > > > > {  
> > > > > > > "term": {  
> > > > > > > "Address\_Admin2Name": "Admin2Name1"  
> > > > > > > }  
> > > > > > > }]  
> > > > > > > }  
> > > > > > > }  
> > > > > > > }  
> > > > > > > }
> > > > > > > 
> > > > > > > This query takes around 6.5 seconds for Top 10 records ( but has  
> > > > > > > sort on top of it)
> > > > > > > 
> > > > > > > {  
> > > > > > > "query": {  
> > > > > > > "constant\_score": {  
> > > > > > > "filter": {  
> > > > > > > "and": [{  
> > > > > > > "term": {  
> > > > > > > "Insurer": "Insurer1"  
> > > > > > > }  
> > > > > > > },  
> > > > > > > {  
> > > > > > > "term": {  
> > > > > > > "Status": "Status1"  
> > > > > > > }  
> > > > > > > }]  
> > > > > > > }  
> > > > > > > }  
> > > > > > > }  
> > > > > > > }
> > > > > > > 
> > > > > > > But all our queries are with random values with few random set of  
> > > > > > > data.
> > > > > > > 
> > > > > > > On Saturday, 17 August 2013 13:59:56 UTC-7, AlexR wrote:
> > > > > > > 
> > > > > > > > Hi VB,
> > > > > > > > 
> > > > > > > > I do not know your use case but have you considered denormalizing  
> > > > > > > > your data? I other words storing parent object as part of its children  
> > > > > > > > json. Your query and facet performance wull be much better but the price to  
> > > > > > > > pay is having to update every child record if parent changes. Plus if  
> > > > > > > > majority of your searches need to return parent you would need to have some  
> > > > > > > > way of distincting single parent record out of potentially multiple hits on  
> > > > > > > > this parent/child. Still it may worth it. Your parent record is pretty  
> > > > > > > > small so the key is just how often it changes. Maybe bring only some of the  
> > > > > > > > parent fields you really need for serching into the child records ?  
> > > > > > > > Alex
> > > > > > > 
> > > > > > > --  
> > > > > > > You received this message because you are subscribed to the Google  
> > > > > > > Groups "elasticsearch" group.  
> > > > > > > To unsubscribe from this group and stop receiving emails from it,  
> > > > > > > send an email to elasticsearc...@\*\*[googlegroups.com](http://googlegroups.com).
> > > > > > > 
> > > > > > > For more options, visit [https://groups.google.com/\*\*groups/opt\_out](https://groups.google.com/**groups/opt_out)[https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> > > > > > > .
> > > > > > 
> > > > > > --  
> > > > > > You received this message because you are subscribed to the Google  
> > > > > > Groups "elasticsearch" group.  
> > > > > > To unsubscribe from this group and stop receiving emails from it, send  
> > > > > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).  
> > > > > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:20am UTC](https://discuss.elastic.co/t/elasticsearch-query-performance/13235/16 "2017-07-06T02:20:18Z")

</div>


