# In case of Massive number of Filter items, which is faster? Filter by \_id or Filter by a long field..?

**URL:** <https://discuss.elastic.co/t/in-case-of-massive-number-of-filter-items-which-is-faster-filter-by--id-or-filter-by-a-long-field/8977>\
**Category:** Elasticsearch\
**Created:** [September 10, 2012, 10:06pm UTC](https://discuss.elastic.co/t/in-case-of-massive-number-of-filter-items-which-is-faster-filter-by--id-or-filter-by-a-long-field/8977 "2012-09-10T22:06:06Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Pradeep\_2](https://avatars.discourse-cdn.com/v4/letter/p/9de053/32.png) [@Pradeep\_2](https://discuss.elastic.co/u/Pradeep_2)\
**Post date:** [September 10, 2012, 10:06pm UTC](https://discuss.elastic.co/t/in-case-of-massive-number-of-filter-items-which-is-faster-filter-by--id-or-filter-by-a-long-field/8977/1 "2012-09-10T22:06:06Z")

</div>

In my query I have to pass many values to filter on.. like

"ids" : {  
"values" : [1, 4, 100 ....... (upto 300k such values)]  
}

as \_id is of type String.. I am wondering if ,I store the value of \_id as a  
long filed and filter on it, will it be faster?

please help...

Thanks a lot..

an elaborate explanation is here  
[https://groups.google.com/forum/?fromgroups=#!searchin/elasticsearch/fastest$20way$20to/elasticsearch/zgO\_qK6kbxE/C26ltlWt7WsJ](https://groups.google.com/forum/?fromgroups=#!searchin/elasticsearch/fastest%2420way%2420to/elasticsearch/zgO_qK6kbxE/C26ltlWt7WsJ)

--

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [September 11, 2012, 12:08pm UTC](https://discuss.elastic.co/t/in-case-of-massive-number-of-filter-items-which-is-faster-filter-by--id-or-filter-by-a-long-field/8977/2 "2012-09-11T12:08:21Z")

</div>

On Mon, 2012-09-10 at 15:06 -0700, Pradeep wrote:

> In my query I have to pass many values to filter on.. like
> 
> "ids" : {  
> "values" : [1, 4, 100 ....... (upto 300k such values)]  
> }
> 
> as \_id is of type String.. I am wondering if ,I store the value of \_id  
> as a long filed and filter on it, will it be faster?

Very difficult to answer - your use case is not clearly described I'm  
afraid.

clint

> please help...
> 
> Thanks a lot..
> 
> an elaborate explanation is here  
> [Redirecting to Google Groups](https://groups.google.com/forum/?fromgroups=#)!  
> searchin/elasticsearch/fastest$20way  
> $20to/elasticsearch/zgO\_qK6kbxE/C26ltlWt7WsJ
> 
> --

--

---

<div class="post-metadata">

**Author:** ![Pradeep\_2](https://avatars.discourse-cdn.com/v4/letter/p/9de053/32.png) [@Pradeep\_2](https://discuss.elastic.co/u/Pradeep_2)\
**Post date:** [September 11, 2012, 1:54pm UTC](https://discuss.elastic.co/t/in-case-of-massive-number-of-filter-items-which-is-faster-filter-by--id-or-filter-by-a-long-field/8977/3 "2012-09-11T13:54:20Z")

</div>

Hi Clinton,

I have a use-case where I have to query elastic with multiple (huge number  
of) values of a field. ( I get these values from another system, so no  
other way but query with all these values)

This field also happens to be the \_id of the documents I have indexed.

So in my query, I have a filter like below...

"filter": { "ids" :{  
"values" : [1, 3, 4, 6, 7,... (upto 300,000 such  
values) ]  
}  
}

( ...eventually I do a facet with term\_stats ... on the docs matching this  
filter)

I noticed for 300,000 such values, the query was taking ( on my single  
node) .. 3-4 secs to return..

So I am looking for ways to reduce this time.

One thought is, as \_id is 'string' type it is probably slower than,  
filtering on a long type ?.. am not sure. Actually that's my question.. is  
it faster ?

If it is faster, I can duplicate the value of \_id as a filed of 'long' type  
in the document and have a "terms" filter

Hope it's clear now... Thanks a lot.  
....

( Btw.. I have tested solr with this use case, and it was many times slower  
than Elasticsearch for such massive list of query terms !! )

On Monday, September 10, 2012 6:06:06 PM UTC-4, Pradeep wrote:

> In my query I have to pass many values to filter on.. like
> 
> "ids" : {  
> "values" : [1, 4, 100 ....... (upto 300k such values)]  
> }
> 
> as \_id is of type String.. I am wondering if ,I store the value of \_id as  
> a long filed and filter on it, will it be faster?
> 
> please help...
> 
> Thanks a lot..
> 
> an elaborate explanation is here
> 
> [Redirecting to Google Groups](https://groups.google.com/forum/?fromgroups=#!searchin/elasticsearch/fastest$20way$20to/elasticsearch/zgO_qK6kbxE/C26ltlWt7WsJ)

--

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [September 11, 2012, 2:05pm UTC](https://discuss.elastic.co/t/in-case-of-massive-number-of-filter-items-which-is-faster-filter-by--id-or-filter-by-a-long-field/8977/4 "2012-09-11T14:05:55Z")

</div>

Hi Pradeep

> I have a use-case where I have to query elastic with multiple (huge  
> number of) values of a field. ( I get these values from another  
> system, so no other way but query with all these values)
> 
> This field also happens to be the \_id of the documents I have indexed.

OK - I understand now. Unfortunately, I don't know what the answer  
is 🙂 I'd suggest just trying it out and seeing what happens.

I think that the ids filter is using an internal UIDs cache, and is  
probably going to be faster than reindexing the values. But that's just  
a guess.

try it and see. would be interested in hearing the results.

clint

> So in my query, I have a filter like below...
> 
> "filter": { "ids" :{  
> "values" : [1, 3, 4, 6, 7,... (upto 300,000  
> such values) ]  
> }  
> }
> 
> ( ...eventually I do a facet with term\_stats ... on the docs matching  
> this filter)
> 
> I noticed for 300,000 such values, the query was taking ( on my single  
> node) .. 3-4 secs to return..
> 
> So I am looking for ways to reduce this time.
> 
> One thought is, as \_id is 'string' type it is probably slower than,  
> filtering on a long type ?.. am not sure. Actually that's my  
> question.. is it faster ?
> 
> If it is faster, I can duplicate the value of \_id as a filed of  
> 'long' type in the document and have a "terms" filter
> 
> Hope it's clear now... Thanks a lot.  
> ....
> 
> ( Btw.. I have tested solr with this use case, and it was many times  
> slower than Elasticsearch for such massive list of query terms !! )
> 
> On Monday, September 10, 2012 6:06:06 PM UTC-4, Pradeep wrote:  
> In my query I have to pass many values to filter on.. like
> 
> ```
> "ids" : {
> "values" : [1, 4, 100 ....... (upto 300k such
> values) ]
> } 
>     
>     
> as _id is of type String.. I am wondering if ,I store the
> value of _id as a long filed and filter on it, will it be
> faster?
>     
>     
>     
>     
> please help...
>     
>     
> Thanks a lot..
>     
>     
>     
>     
> an elaborate explanation is here 
> https://groups.google.com/forum/?fromgroups=#!
> searchin/elasticsearch/fastest$20way
> $20to/elasticsearch/zgO_qK6kbxE/C26ltlWt7WsJ
> 
> ```
> 
> --

--

---

<div class="post-metadata">

**Author:** ![phill](https://avatars.discourse-cdn.com/v4/letter/p/779978/32.png) [@phill](https://discuss.elastic.co/u/phill)\
**Post date:** [September 11, 2012, 4:05pm UTC](https://discuss.elastic.co/t/in-case-of-massive-number-of-filter-items-which-is-faster-filter-by--id-or-filter-by-a-long-field/8977/5 "2012-09-11T16:05:40Z")

</div>

> On Monday, September 10, 2012 6:06:06 PM UTC-4, Pradeep wrote:
> 
> ```
> In my query I have to pass many values to filter on.. like
> 
> "ids" : {
> "values" : [1, 4, 100 ....... (upto 300k such values)]
> }
> 
> ```

Is it possible you could do some pre-processing to work out your large  
sets and already have that information in the index?  
If you can pre-process and generate data before the time of the users  
query do it! I'd bet that finding fewer terms in an (untokenized)  
additional field would be even faster than sending Ks of IDs.

-Paul

--

---

<div class="post-metadata">

**Author:** ![Pradeep\_2](https://avatars.discourse-cdn.com/v4/letter/p/9de053/32.png) [@Pradeep\_2](https://discuss.elastic.co/u/Pradeep_2)\
**Post date:** [September 11, 2012, 4:48pm UTC](https://discuss.elastic.co/t/in-case-of-massive-number-of-filter-items-which-is-faster-filter-by--id-or-filter-by-a-long-field/8977/6 "2012-09-11T16:48:31Z")

</div>

Thanks Clinton. I'll try that out and post.

Hi Paul, I am afraid, that's not possible, this list of values is a  
result of another search, which is served by a different search engine..  
(this data cant be put into that engine. nor the other way around) .. so  
only way is to pass so many values...

On Tuesday, September 11, 2012 12:03:43 PM UTC-4, P Hill wrote:

> > On Monday, September 10, 2012 6:06:06 PM UTC-4, Pradeep wrote:
> > 
> > ```
> > In my query I have to pass many values to filter on.. like 
> > 
> > "ids" : { 
> > "values" : [1, 4, 100 ....... (upto 300k such values)] 
> > } 
> > 
> > ```
> 
> Is it possible you could do some pre-processing to work out your large  
> sets and already have that information in the index?  
> If you can pre-process and generate data before the time of the users  
> query do it! I'd bet that finding fewer terms in an (untokenized)  
> additional field would be even faster than sending Ks of IDs.
> 
> -Paul

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:13am UTC](https://discuss.elastic.co/t/in-case-of-massive-number-of-filter-items-which-is-faster-filter-by--id-or-filter-by-a-long-field/8977/7 "2017-07-06T03:13:09Z")

</div>


