# Alternative approaches to a query

**URL:** <https://discuss.elastic.co/t/alternative-approaches-to-a-query/3323>\
**Category:** Elasticsearch\
**Created:** [September 12, 2010, 5:53pm UTC](https://discuss.elastic.co/t/alternative-approaches-to-a-query/3323 "2010-09-12T17:53:51Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [September 12, 2010, 5:53pm UTC](https://discuss.elastic.co/t/alternative-approaches-to-a-query/3323/1 "2010-09-12T17:53:51Z")

</div>

I have a multilingual datastore of documents. After filtering by other  
criteria, I need to find the document than best matches a particular locale  
specified by the user. As an example, assume I have narrowed by choices down  
to four documents, each with one of the following locale properties:

- doc1.locale = 'en\_US';
- doc2.locale = 'ar';
- doc3.locale = 'en\_GB';
- doc4.locale = 'ar\_SA';

I'd like to devise a query which will return the results ranked where a  
specific language/country beats out just a match on language.

1. A query for 'en' will score 'en\_US' and 'en\_GB' the same and higher  
than the rest.
2. A query for 'en\_GB' will score 'en\_GB' the highest, 'en\_US' next.
3. A query for 'en\_GU' will score 'en\_US' and 'en\_GB' the same and higher  
than the rest.
4. A query for 'ar' will score 'ar' the highest, followed by 'ar\_SA'.
5. A query for 'ar\_SA' will score 'ar\_SA' the highest, followed by 'ar'.
6. A query for 'ar\_LB' will score 'ar' the highest, followed by 'ar\_SA'.

I suppose I can achieve this sort order by breaking the locale stored on the  
document into two terms when the language and country are included. For  
example:

- doc1.locale = 'en\_US en';
- doc2.locale = 'ar';
- doc3.locale = 'en\_GB en';
- doc4.locale = 'ar\_SA ar';

Then perhaps I can do something similar with the search term, so the term  
'en\_GB' becomes 'en en\_GB' with a boost on the more precise match.

But I was wondering if there was a better way to perform this query than me  
manually mangling the search term and the locales stored in each of my  
objects.

Thanks for any help.

---

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [September 12, 2010, 7:28pm UTC](https://discuss.elastic.co/t/alternative-approaches-to-a-query/3323/2 "2010-09-12T19:28:38Z")

</div>

My cursory first attempt to investigate this type of matching seems to  
indicate that the out-of-the-box behavior might be close to what I need  
after all. Assuming I leave the search term and the local properties alone,  
I am seeing the desired matches and scoring I need. For example, if the  
locales are ['en', 'en\_US', 'en\_GB', 'ar', 'ar\_SA', and 'en' (again), the  
following searches return these results:

term results in order

* * *

en en, en, en\_US, en\_GB  
en\_GB en\_GB  
ar ar, ar\_SA  
ar\_SA ar\_SA

So, I guess I got lucky and the out-of-the-box experience is nearly what I  
need. I might be getting fooled by my simple test however. I thought that  
when I ran an early test for 'en', I wasn't getting the exact match ('en')  
to be scored higher than 'en\_US', but I can't seem to duplicate that issue.

I would probably want en\_GB to match 'en' in the absence of an 'en\_GB' exact  
match. So, adding the language to the search terms when the user specifies a  
lang\_country term would alleviate the problem. For example, a search for '  
en\_GB' would become a search for 'en en\_GB'.

The only problem at the moment is the query can't differentiate between  
matching on the language or the country. For example, if the language I am  
searching for is Sanskrit (sa), it will match on my country name for Saudi  
Arabia (also SA but capitalized). If the searches were case sensitive, I  
suppose this could be managed.

My gut feel is that doing this right might require a custom analyzer, but I  
am willing to explore other approaches if that is unnecessarily involved.

On Sun, Sep 12, 2010 at 1:53 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com) wrote:

> I have a multilingual datastore of documents. After filtering by other  
> criteria, I need to find the document than best matches a particular locale  
> specified by the user. As an example, assume I have narrowed by choices down  
> to four documents, each with one of the following locale properties:
> 
> - doc1.locale = 'en\_US';
> - doc2.locale = 'ar';
> - doc3.locale = 'en\_GB';
> - doc4.locale = 'ar\_SA';
> 
> I'd like to devise a query which will return the results ranked where a  
> specific language/country beats out just a match on language.
> 
> 1. A query for 'en' will score 'en\_US' and 'en\_GB' the same and higher  
> than the rest.
> 2. A query for 'en\_GB' will score 'en\_GB' the highest, 'en\_US' next.
> 3. A query for 'en\_GU' will score 'en\_US' and 'en\_GB' the same and  
> higher than the rest.
> 4. A query for 'ar' will score 'ar' the highest, followed by 'ar\_SA'.
> 5. A query for 'ar\_SA' will score 'ar\_SA' the highest, followed by  
> 'ar'.
> 6. A query for 'ar\_LB' will score 'ar' the highest, followed by  
> 'ar\_SA'.
> 
> I suppose I can achieve this sort order by breaking the locale stored on  
> the document into two terms when the language and country are included. For  
> example:
> 
> - doc1.locale = 'en\_US en';
> - doc2.locale = 'ar';
> - doc3.locale = 'en\_GB en';
> - doc4.locale = 'ar\_SA ar';
> 
> Then perhaps I can do something similar with the search term, so the term  
> 'en\_GB' becomes 'en en\_GB' with a boost on the more precise match.
> 
> But I was wondering if there was a better way to perform this query than me  
> manually mangling the search term and the locales stored in each of my  
> objects.
> 
> Thanks for any help.

---

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [September 12, 2010, 7:31pm UTC](https://discuss.elastic.co/t/alternative-approaches-to-a-query/3323/3 "2010-09-12T19:31:59Z")

</div>

In case someone is trying to help, here is a set of curls to set up and test  
a couple queries:

curl -XDELETE '[http://localhost:9200/twitter/tweet/\_query?q=user:](http://localhost:9200/twitter/tweet/_query?q=user:)\*'

curl -XPOST '[http://localhost:9200/twitter/tweet/](http://localhost:9200/twitter/tweet/)' -d '  
{  
"user": "user0",  
"locale": "en"  
}'

curl -XPOST '[http://localhost:9200/twitter/tweet/](http://localhost:9200/twitter/tweet/)' -d '  
{  
"user": "user1",  
"locale": "en\_US"  
}'

curl -XPOST '[http://localhost:9200/twitter/tweet/](http://localhost:9200/twitter/tweet/)' -d '  
{  
"user": "user2",  
"locale": "en\_GB"  
}'

curl -XPOST '[http://localhost:9200/twitter/tweet/](http://localhost:9200/twitter/tweet/)' -d '  
{  
"user": "user3",  
"locale": "ar"  
}'

curl -XPOST '[http://localhost:9200/twitter/tweet/](http://localhost:9200/twitter/tweet/)' -d '  
{  
"user": "user4",  
"locale": "ar\_SA"  
}'

curl -XPOST '[http://localhost:9200/twitter/tweet/](http://localhost:9200/twitter/tweet/)' -d '  
{  
"user": "user5",  
"locale": "en"  
}'

curl -XGET '[http://localhost:9200/twitter/tweet/\_count?q=user:](http://localhost:9200/twitter/tweet/_count?q=user:)\*'

curl -XGET '  
[http://localhost:9200/twitter/tweet/\_search?q=locale:ar&pretty=true](http://localhost:9200/twitter/tweet/_search?q=locale:ar&pretty=true)'

curl -XGET '  
[http://localhost:9200/twitter/tweet/\_search?q=locale:en&pretty=true](http://localhost:9200/twitter/tweet/_search?q=locale:en&pretty=true)'

Thanks

On Sun, Sep 12, 2010 at 3:28 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com) wrote:

> My cursory first attempt to investigate this type of matching seems to  
> indicate that the out-of-the-box behavior might be close to what I need  
> after all. Assuming I leave the search term and the local properties alone,  
> I am seeing the desired matches and scoring I need. For example, if the  
> locales are ['en', 'en\_US', 'en\_GB', 'ar', 'ar\_SA', and 'en' (again), the  
> following searches return these results:
> 
> term results in order
> 
> * * *
> 
> en en, en, en\_US, en\_GB  
> en\_GB en\_GB  
> ar ar, ar\_SA  
> ar\_SA ar\_SA
> 
> So, I guess I got lucky and the out-of-the-box experience is nearly what I  
> need. I might be getting fooled by my simple test however. I thought that  
> when I ran an early test for 'en', I wasn't getting the exact match ('en')  
> to be scored higher than 'en\_US', but I can't seem to duplicate that issue.
> 
> I would probably want en\_GB to match 'en' in the absence of an 'en\_GB'  
> exact match. So, adding the language to the search terms when the user  
> specifies a lang\_country term would alleviate the problem. For example, a  
> search for 'en\_GB' would become a search for 'en en\_GB'.
> 
> The only problem at the moment is the query can't differentiate between  
> matching on the language or the country. For example, if the language I am  
> searching for is Sanskrit (sa), it will match on my country name for Saudi  
> Arabia (also SA but capitalized). If the searches were case sensitive, I  
> suppose this could be managed.
> 
> My gut feel is that doing this right might require a custom analyzer, but I  
> am willing to explore other approaches if that is unnecessarily involved.
> 
> On Sun, Sep 12, 2010 at 1:53 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com) wrote:
> 
> > I have a multilingual datastore of documents. After filtering by other  
> > criteria, I need to find the document than best matches a particular locale  
> > specified by the user. As an example, assume I have narrowed by choices down  
> > to four documents, each with one of the following locale properties:
> > 
> > - doc1.locale = 'en\_US';
> > - doc2.locale = 'ar';
> > - doc3.locale = 'en\_GB';
> > - doc4.locale = 'ar\_SA';
> > 
> > I'd like to devise a query which will return the results ranked where a  
> > specific language/country beats out just a match on language.
> > 
> > 1. A query for 'en' will score 'en\_US' and 'en\_GB' the same and higher  
> > than the rest.
> > 2. A query for 'en\_GB' will score 'en\_GB' the highest, 'en\_US' next.
> > 3. A query for 'en\_GU' will score 'en\_US' and 'en\_GB' the same and  
> > higher than the rest.
> > 4. A query for 'ar' will score 'ar' the highest, followed by 'ar\_SA'.
> > 5. A query for 'ar\_SA' will score 'ar\_SA' the highest, followed by  
> > 'ar'.
> > 6. A query for 'ar\_LB' will score 'ar' the highest, followed by  
> > 'ar\_SA'.
> > 
> > I suppose I can achieve this sort order by breaking the locale stored on  
> > the document into two terms when the language and country are included. For  
> > example:
> > 
> > - doc1.locale = 'en\_US en';
> > - doc2.locale = 'ar';
> > - doc3.locale = 'en\_GB en';
> > - doc4.locale = 'ar\_SA ar';
> > 
> > Then perhaps I can do something similar with the search term, so the term  
> > 'en\_GB' becomes 'en en\_GB' with a boost on the more precise match.
> > 
> > But I was wondering if there was a better way to perform this query than  
> > me manually mangling the search term and the locales stored in each of my  
> > objects.
> > 
> > Thanks for any help.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:19am UTC](https://discuss.elastic.co/t/alternative-approaches-to-a-query/3323/4 "2017-07-06T04:19:27Z")

</div>


