# Exporting big number of entries from elasticsearch

**URL:** <https://discuss.elastic.co/t/exporting-big-number-of-entries-from-elasticsearch/6412>\
**Category:** Elasticsearch\
**Created:** [January 17, 2012, 4:11pm UTC](https://discuss.elastic.co/t/exporting-big-number-of-entries-from-elasticsearch/6412 "2012-01-17T16:11:47Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Simeon\_Zaharici](https://avatars.discourse-cdn.com/v4/letter/s/fbc32d/32.png) [@Simeon\_Zaharici](https://discuss.elastic.co/u/Simeon_Zaharici)\
**Post date:** [January 17, 2012, 4:11pm UTC](https://discuss.elastic.co/t/exporting-big-number-of-entries-from-elasticsearch/6412/1 "2012-01-17T16:11:47Z")

</div>

Hello

I am a new user of elasticsearch. We are using graylog2, a log  
management solution that stores the log messages in elasticsearch.

I would like to be able to extract all log messages that are created  
in one day for archiving purposes. Our servers and applications  
generate around 5 million entries a day. What would be the best way to  
extract these entries from elasticsearch ?

When doing a search that would match all these messages the query  
hangs forever and the load on the elasticserch node against which the  
query is executed goes up.

Here is the query I am executing

curl -XGET '[http://server:9200/graylog2/message/\_search](http://server:9200/graylog2/message/_search)' -d "{  
"size" : "$MESSAGES",  
"query" : {  
"range" : {  
"created\_at" : {  
"from" : "$DATE\_YESTERDAY",  
"to" : "$DATE\_NOW"  
}  
}  
}  
}"|/usr/bin/bzip2 - \> /var/lib/lc-archive/$date\_today.json.bz2

Thanks

---

<div class="post-metadata">

**Author:** ![Berkay\_Mollamustafao](https://avatars.discourse-cdn.com/v4/letter/b/22d042/32.png) [@Berkay\_Mollamustafao](https://discuss.elastic.co/u/Berkay_Mollamustafao)\
**Post date:** [January 17, 2012, 4:17pm UTC](https://discuss.elastic.co/t/exporting-big-number-of-entries-from-elasticsearch/6412/2 "2012-01-17T16:17:25Z")

</div>

Hi,  
You may want to take a look at the scan search type

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

Regards,  
Berkay Mollamustafaoglu  
mberkay on yahoo, google and skype

On Tue, Jan 17, 2012 at 11:11 AM, Simeon Zaharici \<[simeon.zaharici@gmail.com](mailto:simeon.zaharici@gmail.com)

> wrote:

> Hello
> 
> I am a new user of elasticsearch. We are using graylog2, a log  
> management solution that stores the log messages in elasticsearch.
> 
> I would like to be able to extract all log messages that are created  
> in one day for archiving purposes. Our servers and applications  
> generate around 5 million entries a day. What would be the best way to  
> extract these entries from elasticsearch ?
> 
> When doing a search that would match all these messages the query  
> hangs forever and the load on the elasticserch node against which the  
> query is executed goes up.
> 
> Here is the query I am executing
> 
> curl -XGET '[http://server:9200/graylog2/message/\_search](http://server:9200/graylog2/message/_search)' -d "{  
> "size" : "$MESSAGES",  
> "query" : {  
> "range" : {  
> "created\_at" : {  
> "from" : "$DATE\_YESTERDAY",  
> "to" : "$DATE\_NOW"  
> }  
> }  
> }  
> }"|/usr/bin/bzip2 - \> /var/lib/lc-archive/$date\_today.json.bz2
> 
> Thanks

---

<div class="post-metadata">

**Author:** ![Simeon\_Zaharici](https://avatars.discourse-cdn.com/v4/letter/s/fbc32d/32.png) [@Simeon\_Zaharici](https://discuss.elastic.co/u/Simeon_Zaharici)\
**Post date:** [January 17, 2012, 9:43pm UTC](https://discuss.elastic.co/t/exporting-big-number-of-entries-from-elasticsearch/6412/3 "2012-01-17T21:43:43Z")

</div>

Hello

Thanks for your answer, I tried using the scan search type but the  
behavior is the same, the curl request against the search\_id hangs  
forever and the elasticsearch node against which the query was run  
becomes non-responsive...

Thanks

On Jan 17, 11:17 am, Berkay Mollamustafaoglu [mber...@gmail.com](mailto:mber...@gmail.com)  
wrote:

> Hi,  
> You may want to take a look at the scan search typehttp://www.elasticsearch.org/guide/reference/api/search/search-type.html
> 
> Regards,  
> Berkay Mollamustafaoglu  
> mberkay on yahoo, google and skype
> 
> On Tue, Jan 17, 2012 at 11:11 AM, Simeon Zaharici \<[simeon.zahar...@gmail.com](mailto:simeon.zahar...@gmail.com)
> 
> > wrote:  
> > Hello
> 
> > I am a new user of elasticsearch. We are using graylog2, a log  
> > management solution that stores the log messages in elasticsearch.
> 
> > I would like to be able to extract all log messages that are created  
> > in one day for archiving purposes. Our servers and applications  
> > generate around 5 million entries a day. What would be the best way to  
> > extract these entries from elasticsearch ?
> 
> > When doing a search that would match all these messages the query  
> > hangs forever and the load on the elasticserch node against which the  
> > query is executed goes up.
> 
> > Here is the query I am executing
> 
> > curl -XGET '[http://server:9200/graylog2/message/\_search'-d](http://server:9200/graylog2/message/_search'-d) "{  
> > "size" : "$MESSAGES",  
> > "query" : {  
> > "range" : {  
> > "created\_at" : {  
> > "from" : "$DATE\_YESTERDAY",  
> > "to" : "$DATE\_NOW"  
> > }  
> > }  
> > }  
> > }"|/usr/bin/bzip2 - \> /var/lib/lc-archive/$date\_today.json.bz2
> 
> > Thanks

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [January 17, 2012, 10:39pm UTC](https://discuss.elastic.co/t/exporting-big-number-of-entries-from-elasticsearch/6412/4 "2012-01-17T22:39:31Z")

</div>

On Tue, 2012-01-17 at 13:43 -0800, Simeon Zaharici wrote:

> Hello
> 
> Thanks for your answer, I tried using the scan search type but the  
> behavior is the same, the curl request against the search\_id hangs  
> forever and the elasticsearch node against which the query was run  
> becomes non-responsive...

You didn't specify what value $MESSAGES has.

The idea is to use search\_type=scan and to use a scrolled search, with a  
reasonable size (eg 1000).

So you pull the first 1000 (x no of primary shards), then keep pulling  
until there are no more records left to pull

clint

> Thanks
> 
> On Jan 17, 11:17 am, Berkay Mollamustafaoglu [mber...@gmail.com](mailto:mber...@gmail.com)  
> wrote:
> 
> > Hi,  
> > You may want to take a look at the scan search typehttp://www.elasticsearch.org/guide/reference/api/search/search-type.html
> > 
> > Regards,  
> > Berkay Mollamustafaoglu  
> > mberkay on yahoo, google and skype
> > 
> > On Tue, Jan 17, 2012 at 11:11 AM, Simeon Zaharici \<[simeon.zahar...@gmail.com](mailto:simeon.zahar...@gmail.com)
> > 
> > > wrote:  
> > > Hello
> > 
> > > I am a new user of elasticsearch. We are using graylog2, a log  
> > > management solution that stores the log messages in elasticsearch.
> > 
> > > I would like to be able to extract all log messages that are created  
> > > in one day for archiving purposes. Our servers and applications  
> > > generate around 5 million entries a day. What would be the best way to  
> > > extract these entries from elasticsearch ?
> > 
> > > When doing a search that would match all these messages the query  
> > > hangs forever and the load on the elasticserch node against which the  
> > > query is executed goes up.
> > 
> > > Here is the query I am executing
> > 
> > > curl -XGET '[http://server:9200/graylog2/message/\_search'-d](http://server:9200/graylog2/message/_search'-d) "{  
> > > "size" : "$MESSAGES",  
> > > "query" : {  
> > > "range" : {  
> > > "created\_at" : {  
> > > "from" : "$DATE\_YESTERDAY",  
> > > "to" : "$DATE\_NOW"  
> > > }  
> > > }  
> > > }  
> > > }"|/usr/bin/bzip2 - \> /var/lib/lc-archive/$date\_today.json.bz2
> > 
> > > Thanks

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [January 18, 2012, 9:28am UTC](https://discuss.elastic.co/t/exporting-big-number-of-entries-from-elasticsearch/6412/5 "2012-01-18T09:28:30Z")

</div>

> I would like to be able to extract all log messages that are created in one day for archiving purposes.

Why not backup one full index? It would be much faster and cheaper in  
terms of CPU? (not sure if graylog has an option to move to a new  
index per day or sth.)

Peter.

On 17 Jan., 17:11, Simeon Zaharici [simeon.zahar...@gmail.com](mailto:simeon.zahar...@gmail.com) wrote:

> Hello
> 
> I am a new user of elasticsearch. We are using graylog2, a log  
> management solution that stores the log messages in elasticsearch.
> 
> I would like to be able to extract all log messages that are created  
> in one day for archiving purposes. Our servers and applications  
> generate around 5 million entries a day. What would be the best way to  
> extract these entries from elasticsearch ?
> 
> When doing a search that would match all these messages the query  
> hangs forever and the load on the elasticserch node against which the  
> query is executed goes up.
> 
> Here is the query I am executing
> 
> curl -XGET '[http://server:9200/graylog2/message/\_search'-d](http://server:9200/graylog2/message/_search'-d) "{  
> "size" : "$MESSAGES",  
> "query" : {  
> "range" : {  
> "created\_at" : {  
> "from" : "$DATE\_YESTERDAY",  
> "to" : "$DATE\_NOW"  
> }  
> }  
> }
> 
> }"|/usr/bin/bzip2 - \> /var/lib/lc-archive/$date\_today.json.bz2
> 
> Thanks

---

<div class="post-metadata">

**Author:** ![Simeon\_Zaharici](https://avatars.discourse-cdn.com/v4/letter/s/fbc32d/32.png) [@Simeon\_Zaharici](https://discuss.elastic.co/u/Simeon_Zaharici)\
**Post date:** [January 18, 2012, 11:50pm UTC](https://discuss.elastic.co/t/exporting-big-number-of-entries-from-elasticsearch/6412/6 "2012-01-18T23:50:48Z")

</div>

Ah, thanks a lot I was using the scroll all wrong, I had 3million  
entries as size. It works very well by scrolling a reasonable amount  
of entries

On Jan 17, 5:39 pm, Clinton Gormley [cl...@traveljury.com](mailto:cl...@traveljury.com) wrote:

> On Tue, 2012-01-17 at 13:43 -0800, Simeon Zaharici wrote:
> 
> > Hello
> 
> > Thanks for your answer, I tried using the scan search type but the  
> > behavior is the same, the curl request against the search\_id hangs  
> > forever and the elasticsearch node against which the query was run  
> > becomes non-responsive...
> 
> You didn't specify what value $MESSAGES has.
> 
> The idea is to use search\_type=scan and to use a scrolled search, with a  
> reasonable size (eg 1000).
> 
> So you pull the first 1000 (x no of primary shards), then keep pulling  
> until there are no more records left to pull
> 
> clint
> 
> > Thanks
> 
> > On Jan 17, 11:17 am, Berkay Mollamustafaoglu [mber...@gmail.com](mailto:mber...@gmail.com)  
> > wrote:
> > 
> > > Hi,  
> > > You may want to take a look at the scan search typehttp://www.elasticsearch.org/guide/reference/api/search/search-type.html
> 
> > > Regards,  
> > > Berkay Mollamustafaoglu  
> > > mberkay on yahoo, google and skype
> 
> > > On Tue, Jan 17, 2012 at 11:11 AM, Simeon Zaharici \<[simeon.zahar...@gmail.com](mailto:simeon.zahar...@gmail.com)
> 
> > > > wrote:  
> > > > Hello
> 
> > > > I am a new user of elasticsearch. We are using graylog2, a log  
> > > > management solution that stores the log messages in elasticsearch.
> 
> > > > I would like to be able to extract all log messages that are created  
> > > > in one day for archiving purposes. Our servers and applications  
> > > > generate around 5 million entries a day. What would be the best way to  
> > > > extract these entries from elasticsearch ?
> 
> > > > When doing a search that would match all these messages the query  
> > > > hangs forever and the load on the elasticserch node against which the  
> > > > query is executed goes up.
> 
> > > > Here is the query I am executing
> 
> > > > curl -XGET '[http://server:9200/graylog2/message/\_search'-d](http://server:9200/graylog2/message/_search'-d)"{  
> > > > "size" : "$MESSAGES",  
> > > > "query" : {  
> > > > "range" : {  
> > > > "created\_at" : {  
> > > > "from" : "$DATE\_YESTERDAY",  
> > > > "to" : "$DATE\_NOW"  
> > > > }  
> > > > }  
> > > > }  
> > > > }"|/usr/bin/bzip2 - \> /var/lib/lc-archive/$date\_today.json.bz2
> 
> > > > Thanks

---

<div class="post-metadata">

**Author:** ![Simeon\_Zaharici](https://avatars.discourse-cdn.com/v4/letter/s/fbc32d/32.png) [@Simeon\_Zaharici](https://discuss.elastic.co/u/Simeon_Zaharici)\
**Post date:** [January 18, 2012, 11:52pm UTC](https://discuss.elastic.co/t/exporting-big-number-of-entries-from-elasticsearch/6412/7 "2012-01-18T23:52:56Z")

</div>

Hi,

No graylog does not have that option ( yet ? ) so that's why I am  
implementing the daily archive

On Jan 18, 4:28 am, Karussell [tableyourt...@googlemail.com](mailto:tableyourt...@googlemail.com) wrote:

> > I would like to be able to extract all log messages that are created in one day for archiving purposes.
> 
> Why not backup one full index? It would be much faster and cheaper in  
> terms of CPU? (not sure if graylog has an option to move to a new  
> index per day or sth.)
> 
> Peter.
> 
> On 17 Jan., 17:11, Simeon Zaharici [simeon.zahar...@gmail.com](mailto:simeon.zahar...@gmail.com) wrote:
> 
> > Hello
> 
> > I am a new user of elasticsearch. We are using graylog2, a log  
> > management solution that stores the log messages in elasticsearch.
> 
> > I would like to be able to extract all log messages that are created  
> > in one day for archiving purposes. Our servers and applications  
> > generate around 5 million entries a day. What would be the best way to  
> > extract these entries from elasticsearch ?
> 
> > When doing a search that would match all these messages the query  
> > hangs forever and the load on the elasticserch node against which the  
> > query is executed goes up.
> 
> > Here is the query I am executing
> 
> > curl -XGET '[http://server:9200/graylog2/message/\_search'-d](http://server:9200/graylog2/message/_search'-d)"{  
> > "size" : "$MESSAGES",  
> > "query" : {  
> > "range" : {  
> > "created\_at" : {  
> > "from" : "$DATE\_YESTERDAY",  
> > "to" : "$DATE\_NOW"  
> > }  
> > }  
> > }
> 
> > }"|/usr/bin/bzip2 - \> /var/lib/lc-archive/$date\_today.json.bz2
> 
> > Thanks

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:42am UTC](https://discuss.elastic.co/t/exporting-big-number-of-entries-from-elasticsearch/6412/8 "2017-07-06T03:42:11Z")

</div>


