# Few Querys related to ElasticSearch

**URL:** <https://discuss.elastic.co/t/few-querys-related-to-elasticsearch/3606>\
**Category:** Elasticsearch\
**Created:** [November 29, 2010, 11:08am UTC](https://discuss.elastic.co/t/few-querys-related-to-elasticsearch/3606 "2010-11-29T11:08:19Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Gautam](https://avatars.discourse-cdn.com/v4/letter/g/7ab992/32.png) [@Gautam](https://discuss.elastic.co/u/Gautam)\
**Post date:** [November 29, 2010, 11:08am UTC](https://discuss.elastic.co/t/few-querys-related-to-elasticsearch/3606/1 "2010-11-29T11:08:19Z")

</div>

Hi  
Would appreciate if any of you can share your experience / thoughts on below  
questions:

1. REST Api vs Java api - Have read that Java api is much faster as it works  
at a lower level protocol. Do you guys have any comparison?
2. What approach do you suggest for the below mentioned use case:  
I plan to index a stream of short messages (like Tweets) into ES.  
Now I don't want to keep say more than a month old data. How do I flush it?
3. If I create 3 index files say a, b, c. How do I tell ES to search on all  
these indexes?
4. ES seems to have good shard support. Is there a way to control these  
shards on capacity?

Thanks in advance for your help.

Regards  
Gautam

---

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [November 29, 2010, 12:27pm UTC](https://discuss.elastic.co/t/few-querys-related-to-elasticsearch/3606/2 "2010-11-29T12:27:02Z")

</div>

Hi,

let me try to answer (inlining)"

On Mon, Nov 29, 2010 at 12:08 PM, Gautam Mr [mrgautamsam@gmail.com](mailto:mrgautamsam@gmail.com) wrote:

> Hi  
> Would appreciate if any of you can share your experience / thoughts on  
> below questions:
> 
> 1. REST Api vs Java api - Have read that Java api is much faster as it  
> works at a lower level protocol. Do you guys have any comparison?

It depends on what exactly you measure and also on your use case. First, as  
of writing the REST API can be used via  
HTTP[http://www.elasticsearch.com/docs/elasticsearch/modules/http/](http://www.elasticsearch.com/docs/elasticsearch/modules/http/)or  
Memcached[http://www.elasticsearch.com/docs/elasticsearch/modules/memcached/](http://www.elasticsearch.com/docs/elasticsearch/modules/memcached/)protocols.  
Memcached protocol should be faster then HTTP (and has also some  
minor downsides) but it depends on your client implementation (e.g. client  
can be using slow implementation of HTTP client module under the hood). When  
using Java API there different  
options[http://www.elasticsearch.com/docs/elasticsearch/java\_api/client/](http://www.elasticsearch.com/docs/elasticsearch/java_api/client/):  
TransportClient or NodeClient. TransportClient is slower then NodeClient;  
however, NodeClient joins directly the cluster while TransportCient does  
not. Both (Java) clients use optimized binary protocol so they are faster  
then HTTP and Memcached protocols.

> 1. What approach do you suggest for the below mentioned use case:  
> I plan to index a stream of short messages (like Tweets) into ES.  
> Now I don't want to keep say more than a month old data. How do I flush it?

You can index your data by weeks (days, hours, ... etc, you name it) and  
have each data bucket indexed into a specific index. You will end up with  
more indices like: twitter-ww31, twitter-ww32, twitter-ww33 (...). Then you  
can search across more  
indices[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/search/indices\_types/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/search/indices_types/)and  
drop old indices (see index  
delete[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/admin/indices/delete\_index/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/admin/indices/delete_index/)).  
Also note that each index can have  
aliases[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/admin/indices/aliases/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/admin/indices/aliases/)so  
this could help you to just search in one "index alias" while this  
will  
span to multiple indices automatically.

> 1. If I create 3 index files say a, b, c. How do I tell ES to search on all  
> these indexes?

See index aliases[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/admin/indices/aliases/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/admin/indices/aliases/)  
.

> 1. ES seems to have good shard support. Is there a way to control these  
> shards on capacity?

You mean if shards are of the same size? As far as I understand the data is  
split among shards evenly by edfault. So if you have 10MB of data and you  
have 5 shards, then each shard would have around 2MB. However, there has  
been implemented a new  
routing[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/index/#Routing](http://www.elasticsearch.com/docs/elasticsearch/rest_api/index/#Routing)API  
in 0.13.0 which gives you a chance to control shard routing (see this  
ticket for details:  
[API: Allow to control document shard routing, and search shard routing · Issue #470 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/issue/470)).

> Thanks in advance for your help.
> 
> Regards  
> Gautam

Regards,  
Lukas

---

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [November 29, 2010, 12:30pm UTC](https://discuss.elastic.co/t/few-querys-related-to-elasticsearch/3606/3 "2010-11-29T12:30:54Z")

</div>

On Mon, Nov 29, 2010 at 1:27 PM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com) wrote:

> Hi,
> 
> let me try to answer (inlining)"
> 
> On Mon, Nov 29, 2010 at 12:08 PM, Gautam Mr [mrgautamsam@gmail.com](mailto:mrgautamsam@gmail.com) wrote:
> 
> > Hi  
> > Would appreciate if any of you can share your experience / thoughts on  
> > below questions:
> > 
> > 1. REST Api vs Java api - Have read that Java api is much faster as it  
> > works at a lower level protocol. Do you guys have any comparison?
> 
> It depends on what exactly you measure and also on your use case. First, as  
> of writing the REST API can be used via HTTP[http://www.elasticsearch.com/docs/elasticsearch/modules/http/](http://www.elasticsearch.com/docs/elasticsearch/modules/http/)or  
> Memcached[http://www.elasticsearch.com/docs/elasticsearch/modules/memcached/](http://www.elasticsearch.com/docs/elasticsearch/modules/memcached/)protocols. Memcached protocol should be faster then HTTP (and has also some  
> minor downsides) but it depends on your client implementation (e.g. client  
> can be using slow implementation of HTTP client module under the hood). When  
> using Java API there different options[http://www.elasticsearch.com/docs/elasticsearch/java\_api/client/](http://www.elasticsearch.com/docs/elasticsearch/java_api/client/):  
> TransportClient or NodeClient. TransportClient is slower then NodeClient;  
> however, NodeClient joins directly the cluster while TransportCient does  
> not. Both (Java) clients use optimized binary protocol so they are faster  
> then HTTP and Memcached protocols.
> 
> > 1. What approach do you suggest for the below mentioned use case:  
> > I plan to index a stream of short messages (like Tweets) into ES.  
> > Now I don't want to keep say more than a month old data. How do I flush it?
> 
> You can index your data by weeks (days, hours, ... etc, you name it) and  
> have each data bucket indexed into a specific index. You will end up with  
> more indices like: twitter-ww31, twitter-ww32, twitter-ww33 (...). Then you  
> can search across more indices[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/search/indices\_types/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/search/indices_types/)and drop old indices (see index  
> delete[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/admin/indices/delete\_index/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/admin/indices/delete_index/)).  
> Also note that each index can have aliases[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/admin/indices/aliases/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/admin/indices/aliases/)so this could help you to just search in one "index alias" while this will  
> span to multiple indices automatically.
> 
> > 1. If I create 3 index files say a, b, c. How do I tell ES to search on  
> > all these indexes?
> 
> See index aliases[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/admin/indices/aliases/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/admin/indices/aliases/)  
> .

Oops, I meant, see searching multiple  
indices[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/search/indices\_types/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/search/indices_types/),  
but as I said above you can consider index aliases in this case.

> > 1. ES seems to have good shard support. Is there a way to control these  
> > shards on capacity?
> 
> You mean if shards are of the same size? As far as I understand the data is  
> split among shards evenly by edfault. So if you have 10MB of data and you  
> have 5 shards, then each shard would have around 2MB. However, there has  
> been implemented a new routing[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/index/#Routing](http://www.elasticsearch.com/docs/elasticsearch/rest_api/index/#Routing)API in 0.13.0 which gives you a chance to control shard routing (see this  
> ticket for details:  
> [API: Allow to control document shard routing, and search shard routing · Issue #470 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/issue/470)).
> 
> > Thanks in advance for your help.
> > 
> > Regards  
> > Gautam
> 
> Regards,  
> Lukas

---

<div class="post-metadata">

**Author:** ![Gautam](https://avatars.discourse-cdn.com/v4/letter/g/7ab992/32.png) [@Gautam](https://discuss.elastic.co/u/Gautam)\
**Post date:** [November 29, 2010, 6:42pm UTC](https://discuss.elastic.co/t/few-querys-related-to-elasticsearch/3606/4 "2010-11-29T18:42:24Z")

</div>

Thanks a lot Lucas for the detailed reply. This helps a lot!!

Best Regards  
Gautam

On Mon, Nov 29, 2010 at 6:00 PM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com) wrote:

> On Mon, Nov 29, 2010 at 1:27 PM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com)wrote:
> 
> > Hi,
> > 
> > let me try to answer (inlining)"
> > 
> > On Mon, Nov 29, 2010 at 12:08 PM, Gautam Mr [mrgautamsam@gmail.com](mailto:mrgautamsam@gmail.com)wrote:
> > 
> > > Hi  
> > > Would appreciate if any of you can share your experience / thoughts on  
> > > below questions:
> > > 
> > > 1. REST Api vs Java api - Have read that Java api is much faster as it  
> > > works at a lower level protocol. Do you guys have any comparison?
> > 
> > It depends on what exactly you measure and also on your use case. First,  
> > as of writing the REST API can be used via HTTP[http://www.elasticsearch.com/docs/elasticsearch/modules/http/](http://www.elasticsearch.com/docs/elasticsearch/modules/http/)or  
> > Memcached[http://www.elasticsearch.com/docs/elasticsearch/modules/memcached/](http://www.elasticsearch.com/docs/elasticsearch/modules/memcached/)protocols. Memcached protocol should be faster then HTTP (and has also some  
> > minor downsides) but it depends on your client implementation (e.g. client  
> > can be using slow implementation of HTTP client module under the hood). When  
> > using Java API there different options[http://www.elasticsearch.com/docs/elasticsearch/java\_api/client/](http://www.elasticsearch.com/docs/elasticsearch/java_api/client/):  
> > TransportClient or NodeClient. TransportClient is slower then NodeClient;  
> > however, NodeClient joins directly the cluster while TransportCient does  
> > not. Both (Java) clients use optimized binary protocol so they are faster  
> > then HTTP and Memcached protocols.
> > 
> > > 1. What approach do you suggest for the below mentioned use case:  
> > > I plan to index a stream of short messages (like Tweets) into  
> > > ES. Now I don't want to keep say more than a month old data. How do I flush  
> > > it?
> > 
> > You can index your data by weeks (days, hours, ... etc, you name it) and  
> > have each data bucket indexed into a specific index. You will end up with  
> > more indices like: twitter-ww31, twitter-ww32, twitter-ww33 (...). Then you  
> > can search across more indices[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/search/indices\_types/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/search/indices_types/)and drop old indices (see index  
> > delete[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/admin/indices/delete\_index/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/admin/indices/delete_index/)).  
> > Also note that each index can have aliases[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/admin/indices/aliases/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/admin/indices/aliases/)so this could help you to just search in one "index alias" while this will  
> > span to multiple indices automatically.
> > 
> > > 1. If I create 3 index files say a, b, c. How do I tell ES to search on  
> > > all these indexes?
> > 
> > See index aliases[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/admin/indices/aliases/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/admin/indices/aliases/)  
> > .
> 
> Oops, I meant, see searching multiple indices[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/search/indices\_types/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/search/indices_types/),  
> but as I said above you can consider index aliases in this case.
> 
> > > 1. ES seems to have good shard support. Is there a way to control these  
> > > shards on capacity?
> > 
> > You mean if shards are of the same size? As far as I understand the data  
> > is split among shards evenly by edfault. So if you have 10MB of data and you  
> > have 5 shards, then each shard would have around 2MB. However, there has  
> > been implemented a new routing[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/index/#Routing](http://www.elasticsearch.com/docs/elasticsearch/rest_api/index/#Routing)API in 0.13.0 which gives you a chance to control shard routing (see this  
> > ticket for details:  
> > [API: Allow to control document shard routing, and search shard routing · Issue #470 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/issue/470)).
> > 
> > > Thanks in advance for your help.
> > > 
> > > Regards  
> > > Gautam
> > 
> > Regards,  
> > Lukas

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 30, 2010, 12:00pm UTC](https://discuss.elastic.co/t/few-querys-related-to-elasticsearch/3606/5 "2010-11-30T12:00:31Z")

</div>

just a note regarding the HTTP vs. the native transport performance. I can  
get really good performance with HTTP as well compared to the native one  
using Java, it really depends on the http lib used by your programming  
language of choice.

On Mon, Nov 29, 2010 at 8:42 PM, Gautam Mr [mrgautamsam@gmail.com](mailto:mrgautamsam@gmail.com) wrote:

> Thanks a lot Lucas for the detailed reply. This helps a lot!!
> 
> Best Regards  
> Gautam
> 
> On Mon, Nov 29, 2010 at 6:00 PM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com)wrote:
> 
> > On Mon, Nov 29, 2010 at 1:27 PM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com)wrote:
> > 
> > > Hi,
> > > 
> > > let me try to answer (inlining)"
> > > 
> > > On Mon, Nov 29, 2010 at 12:08 PM, Gautam Mr [mrgautamsam@gmail.com](mailto:mrgautamsam@gmail.com)wrote:
> > > 
> > > > Hi  
> > > > Would appreciate if any of you can share your experience / thoughts on  
> > > > below questions:
> > > > 
> > > > 1. REST Api vs Java api - Have read that Java api is much faster as it  
> > > > works at a lower level protocol. Do you guys have any comparison?
> > > 
> > > It depends on what exactly you measure and also on your use case. First,  
> > > as of writing the REST API can be used via HTTP[http://www.elasticsearch.com/docs/elasticsearch/modules/http/](http://www.elasticsearch.com/docs/elasticsearch/modules/http/)or  
> > > Memcached[http://www.elasticsearch.com/docs/elasticsearch/modules/memcached/](http://www.elasticsearch.com/docs/elasticsearch/modules/memcached/)protocols. Memcached protocol should be faster then HTTP (and has also some  
> > > minor downsides) but it depends on your client implementation (e.g. client  
> > > can be using slow implementation of HTTP client module under the hood). When  
> > > using Java API there different options[http://www.elasticsearch.com/docs/elasticsearch/java\_api/client/](http://www.elasticsearch.com/docs/elasticsearch/java_api/client/):  
> > > TransportClient or NodeClient. TransportClient is slower then NodeClient;  
> > > however, NodeClient joins directly the cluster while TransportCient does  
> > > not. Both (Java) clients use optimized binary protocol so they are faster  
> > > then HTTP and Memcached protocols.
> > > 
> > > > 1. What approach do you suggest for the below mentioned use case:  
> > > > I plan to index a stream of short messages (like Tweets) into  
> > > > ES. Now I don't want to keep say more than a month old data. How do I flush  
> > > > it?
> > > 
> > > You can index your data by weeks (days, hours, ... etc, you name it) and  
> > > have each data bucket indexed into a specific index. You will end up with  
> > > more indices like: twitter-ww31, twitter-ww32, twitter-ww33 (...). Then you  
> > > can search across more indices[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/search/indices\_types/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/search/indices_types/)and drop old indices (see index  
> > > delete[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/admin/indices/delete\_index/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/admin/indices/delete_index/)).  
> > > Also note that each index can have aliases[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/admin/indices/aliases/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/admin/indices/aliases/)so this could help you to just search in one "index alias" while this will  
> > > span to multiple indices automatically.
> > > 
> > > > 1. If I create 3 index files say a, b, c. How do I tell ES to search on  
> > > > all these indexes?
> > > 
> > > See index aliases[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/admin/indices/aliases/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/admin/indices/aliases/)  
> > > .
> > 
> > Oops, I meant, see searching multiple indices[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/search/indices\_types/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/search/indices_types/),  
> > but as I said above you can consider index aliases in this case.
> > 
> > > > 1. ES seems to have good shard support. Is there a way to control these  
> > > > shards on capacity?
> > > 
> > > You mean if shards are of the same size? As far as I understand the data  
> > > is split among shards evenly by edfault. So if you have 10MB of data and you  
> > > have 5 shards, then each shard would have around 2MB. However, there has  
> > > been implemented a new routing[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/index/#Routing](http://www.elasticsearch.com/docs/elasticsearch/rest_api/index/#Routing)API in 0.13.0 which gives you a chance to control shard routing (see this  
> > > ticket for details:  
> > > [API: Allow to control document shard routing, and search shard routing · Issue #470 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/issue/470)).
> > > 
> > > > Thanks in advance for your help.
> > > > 
> > > > Regards  
> > > > Gautam
> > > 
> > > Regards,  
> > > Lukas

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:15am UTC](https://discuss.elastic.co/t/few-querys-related-to-elasticsearch/3606/6 "2017-07-06T04:15:59Z")

</div>


