# Geopolygon testing

**URL:** <https://discuss.elastic.co/t/geopolygon-testing/4740>\
**Category:** Elasticsearch\
**Created:** [June 30, 2011, 8:30am UTC](https://discuss.elastic.co/t/geopolygon-testing/4740 "2011-06-30T08:30:56Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![ian\_clark](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ian_clark/32/3129_2.png) [@ian\_clark](https://discuss.elastic.co/u/ian_clark)\
**Post date:** [June 30, 2011, 8:30am UTC](https://discuss.elastic.co/t/geopolygon-testing/4740/1 "2011-06-30T08:30:56Z")

</div>

Hi,

I've been looking into writing a search application powered by  
elasticsearch using user-defined polygons. I can do simple cases and  
get fairly good performance, but in the cases of complex polygons(256  
points for example) the search times get slower. No surprise it's  
having to do more calculations. Here are my results:-

[https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-VaryingNodes.png](https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-VaryingNodes.png)

I repeated the test adding more nodes, and because it's splitting the  
work out the results are coming down. The nodes are 4 core windows  
machines, running a 12 shard index of about 736 documents. Is there  
anyway I can optimise the search besides throwing more nodes at it?  
This is an example my query:-

> <https://gist.github.com/ianAndrewClark/1055826>

I've been thinking of wrapping the polygon filter in an "and" filter  
with the first clause being a bounded box filter(http://  
[www.elasticsearch.org/guide/reference/query-dsl/geo-bounding-box-filter.html](http://www.elasticsearch.org/guide/reference/query-dsl/geo-bounding-box-filter.html)),  
as this should be a faster calculation to discount hits, but I'm not  
sure if the order of filters means anything. Would this help? Is there  
anything else I can do?

cheers

Ian

---

<div class="post-metadata">

**Author:** ![ian\_clark](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ian_clark/32/3129_2.png) [@ian\_clark](https://discuss.elastic.co/u/ian_clark)\
**Post date:** [June 30, 2011, 8:32am UTC](https://discuss.elastic.co/t/geopolygon-testing/4740/2 "2011-06-30T08:32:35Z")

</div>

sorry that is 736k documents... not 736.

On Jun 30, 9:30 am, Ian [ian.andrew.cl...@gmail.com](mailto:ian.andrew.cl...@gmail.com) wrote:

> Hi,
> 
> I've been looking into writing a search application powered by  
> elasticsearch using user-defined polygons. I can do simple cases and  
> get fairly good performance, but in the cases of complex polygons(256  
> points for example) the search times get slower. No surprise it's  
> having to do more calculations. Here are my results:-
> 
> [https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-V](https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-V)...
> 
> I repeated the test adding more nodes, and because it's splitting the  
> work out the results are coming down. The nodes are 4 core windows  
> machines, running a 12 shard index of about 736 documents. Is there  
> anyway I can optimise the search besides throwing more nodes at it?  
> This is an example my query:-
> 
> [3 times 10 point polygon query · GitHub](https://gist.github.com/1055826)
> 
> I've been thinking of wrapping the polygon filter in an "and" filter  
> with the first clause being a bounded box filter([Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/geo-bounding-box-filt)...),  
> as this should be a faster calculation to discount hits, but I'm not  
> sure if the order of filters means anything. Would this help? Is there  
> anything else I can do?
> 
> cheers
> 
> Ian

---

<div class="post-metadata">

**Author:** ![ian\_clark](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ian_clark/32/3129_2.png) [@ian\_clark](https://discuss.elastic.co/u/ian_clark)\
**Post date:** [July 1, 2011, 11:44am UTC](https://discuss.elastic.co/t/geopolygon-testing/4740/3 "2011-07-01T11:44:41Z")

</div>

I got my test results to be even better, er by moving the test runner  
to be closer to the cluster, but also adding a bounded box improved  
the results further

[https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-CloserWithBounded.png](https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-CloserWithBounded.png)

With the exception of multiple polygon filters where the bounded box  
actually made the query time longer. That makes the structure of the  
filters:- and(bbox, or(poly1, poly2, poly3)), as opposed to single  
polygons which are and(bbox, poly1) and I guess that extra level takes  
it toll in the calculation.

My queries now look like this:-

> <https://gist.github.com/ianAndrewClark/1058361>

IC

On Jun 30, 9:32 am, Ian [ian.andrew.cl...@gmail.com](mailto:ian.andrew.cl...@gmail.com) wrote:

> sorry that is 736k documents... not 736.
> 
> On Jun 30, 9:30 am, Ian [ian.andrew.cl...@gmail.com](mailto:ian.andrew.cl...@gmail.com) wrote:
> 
> > Hi,
> 
> > I've been looking into writing a search application powered by  
> > elasticsearch using user-defined polygons. I can do simple cases and  
> > get fairly good performance, but in the cases of complex polygons(256  
> > points for example) the search times get slower. No surprise it's  
> > having to do more calculations. Here are my results:-
> 
> > [https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-V](https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-V)...
> 
> > I repeated the test adding more nodes, and because it's splitting the  
> > work out the results are coming down. The nodes are 4 core windows  
> > machines, running a 12 shard index of about 736 documents. Is there  
> > anyway I can optimise the search besides throwing more nodes at it?  
> > This is an example my query:-
> 
> > [3 times 10 point polygon query · GitHub](https://gist.github.com/1055826)
> 
> > I've been thinking of wrapping the polygon filter in an "and" filter  
> > with the first clause being a bounded box filter([Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/geo-bounding-b)......),  
> > as this should be a faster calculation to discount hits, but I'm not  
> > sure if the order of filters means anything. Would this help? Is there  
> > anything else I can do?
> 
> > cheers
> 
> > Ian

---

<div class="post-metadata">

**Author:** ![Stephane\_Bastian](https://avatars.discourse-cdn.com/v4/letter/s/35a633/32.png) [@Stephane\_Bastian](https://discuss.elastic.co/u/Stephane_Bastian)\
**Post date:** [July 1, 2011, 1:07pm UTC](https://discuss.elastic.co/t/geopolygon-testing/4740/4 "2011-07-01T13:07:08Z")

</div>

Hi Ian,

Thanks a lot for sharing. Very interesting and useful to ES users

Stephane Bastian

On Fri, 2011-07-01 at 04:44 -0700, Ian wrote:

> I got my test results to be even better, er by moving the test runner  
> to be closer to the cluster, but also adding a bounded box improved  
> the results further
> 
> [https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-CloserWithBounded.png](https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-CloserWithBounded.png)
> 
> With the exception of multiple polygon filters where the bounded box  
> actually made the query time longer. That makes the structure of the  
> filters:- and(bbox, or(poly1, poly2, poly3)), as opposed to single  
> polygons which are and(bbox, poly1) and I guess that extra level takes  
> it toll in the calculation.
> 
> My queries now look like this:-
> 
> [3 point polygon with bounded box query · GitHub](https://gist.github.com/1058361)
> 
> IC
> 
> On Jun 30, 9:32 am, Ian [ian.andrew.cl...@gmail.com](mailto:ian.andrew.cl...@gmail.com) wrote:
> 
> > sorry that is 736k documents... not 736.
> > 
> > On Jun 30, 9:30 am, Ian [ian.andrew.cl...@gmail.com](mailto:ian.andrew.cl...@gmail.com) wrote:
> > 
> > > Hi,
> > 
> > > I've been looking into writing a search application powered by  
> > > elasticsearch using user-defined polygons. I can do simple cases and  
> > > get fairly good performance, but in the cases of complex polygons(256  
> > > points for example) the search times get slower. No surprise it's  
> > > having to do more calculations. Here are my results:-
> > 
> > > [https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-V](https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-V)...
> > 
> > > I repeated the test adding more nodes, and because it's splitting the  
> > > work out the results are coming down. The nodes are 4 core windows  
> > > machines, running a 12 shard index of about 736 documents. Is there  
> > > anyway I can optimise the search besides throwing more nodes at it?  
> > > This is an example my query:-
> > 
> > > [3 times 10 point polygon query · GitHub](https://gist.github.com/1055826)
> > 
> > > I've been thinking of wrapping the polygon filter in an "and" filter  
> > > with the first clause being a bounded box filter([Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/geo-bounding-b)......),  
> > > as this should be a faster calculation to discount hits, but I'm not  
> > > sure if the order of filters means anything. Would this help? Is there  
> > > anything else I can do?
> > 
> > > cheers
> > 
> > > Ian

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 2, 2011, 8:46pm UTC](https://discuss.elastic.co/t/geopolygon-testing/4740/5 "2011-07-02T20:46:24Z")

</div>

Yea, thanks a lot for the effort!. Yea, the computation is heavy as the number of points grow, might give it another round of check to see if maybe it can be further optimized...

On Friday, July 1, 2011 at 4:07 PM, stephane wrote:

> Hi Ian,
> 
> Thanks a lot for sharing. Very interesting and useful to ES users
> 
> Stephane Bastian
> 
> On Fri, 2011-07-01 at 04:44 -0700, Ian wrote:
> 
> > I got my test results to be even better, er by moving the test runner  
> > to be closer to the cluster, but also adding a bounded box improved  
> > the results further
> > 
> > [https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-CloserWithBounded.png](https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-CloserWithBounded.png)
> > 
> > With the exception of multiple polygon filters where the bounded box  
> > actually made the query time longer. That makes the structure of the  
> > filters:- and(bbox, or(poly1, poly2, poly3)), as opposed to single  
> > polygons which are and(bbox, poly1) and I guess that extra level takes  
> > it toll in the calculation.
> > 
> > My queries now look like this:-
> > 
> > [3 point polygon with bounded box query · GitHub](https://gist.github.com/1058361)
> > 
> > IC
> > 
> > On Jun 30, 9:32 am, Ian \<[ian.andrew.cl...@gmail.com](mailto:ian.andrew.cl...@gmail.com) ([http://gmail.com](http://gmail.com))\> wrote:
> > 
> > > sorry that is 736k documents... not 736.
> > > 
> > > On Jun 30, 9:30 am, Ian \<[ian.andrew.cl...@gmail.com](mailto:ian.andrew.cl...@gmail.com) ([http://gmail.com](http://gmail.com))\> wrote:
> > > 
> > > > Hi,
> > > 
> > > > I've been looking into writing a search application powered by  
> > > > elasticsearch using user-defined polygons. I can do simple cases and  
> > > > get fairly good performance, but in the cases of complex polygons(256  
> > > > points for example) the search times get slower. No surprise it's  
> > > > having to do more calculations. Here are my results:-
> > > 
> > > > [https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-V](https://s3-eu-west-1.amazonaws.com/es-results/ComplexPolygonResults-V)...
> > > 
> > > > I repeated the test adding more nodes, and because it's splitting the  
> > > > work out the results are coming down. The nodes are 4 core windows  
> > > > machines, running a 12 shard index of about 736 documents. Is there  
> > > > anyway I can optimise the search besides throwing more nodes at it?  
> > > > This is an example my query:-
> > > 
> > > > [3 times 10 point polygon query · GitHub](https://gist.github.com/1055826)
> > > 
> > > > I've been thinking of wrapping the polygon filter in an "and" filter  
> > > > with the first clause being a bounded box filter([Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/geo-bounding-b)......),  
> > > > as this should be a faster calculation to discount hits, but I'm not  
> > > > sure if the order of filters means anything. Would this help? Is there  
> > > > anything else I can do?
> > > 
> > > > cheers
> > > 
> > > > Ian

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:02am UTC](https://discuss.elastic.co/t/geopolygon-testing/4740/6 "2017-07-06T04:02:00Z")

</div>


