# How to get data more than 10000 in elasticsearch

**URL:** <https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869>\
**Category:** Elasticsearch\
**Created:** [November 16, 2017, 6:36am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869 "2017-11-16T06:36:28Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![kanchan](https://avatars.discourse-cdn.com/v4/letter/k/8797f3/32.png) [@kanchan](https://discuss.elastic.co/u/kanchan)\
**Post date:** [November 16, 2017, 6:36am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/1 "2017-11-16T06:36:28Z")

</div>

Hi Team,  
I am trying to fetch data using rest client in java side, but not able to fetch more than 10000 and even if i am trying to fetch data less then 10000 like 5000 or 7000 it is taking too much time.  
please let me know hoe we can achieve it.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 16, 2017, 7:43am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/2 "2017-11-16T07:43:36Z")

</div>

You need to use the scroll API to extract data.

---

<div class="post-metadata">

**Author:** ![kanchan](https://avatars.discourse-cdn.com/v4/letter/k/8797f3/32.png) [@kanchan](https://discuss.elastic.co/u/kanchan)\
**Post date:** [November 21, 2017, 7:29am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/3 "2017-11-21T07:29:26Z")

</div>

I used scroll api that worked f9. Thanks!!!!!!!!!!!  
I checked one example and follow that but now I want to use raw query.Like:

`final Scroll scroll = new Scroll(TimeValue.timeValueMinutes(1L));`  
SearchRequest searchRequest = new SearchRequest("entity\_fact\_test4");  
searchRequest.scroll(scroll);  
SearchSourceBuilder searchSourceBuilder = new SearchSourceBuilder();  
QueryBuilder qb = QueryBuilders.termQuery("client\_id", "262");  
searchSourceBuilder.size(100000);  
searchSourceBuilder.query(qb);  
searchRequest.source(searchSourceBuilder);  
SearchResponse searchResponse = restClient.search(searchRequest);

here currently i am using :  
QueryBuilder qb = QueryBuilders.termQuery("client\_id", "262");

but I want to replace this using raw or customize query like:

String queryProductElk ="{\r\n" +  
" "query": { \r\n" +  
" "bool": {\r\n" +  
" "must": [\r\n" +  
" {\r\n" +  
" "term": {\r\n" +  
" "client\_id": {\r\n" +  
" "value": "262"\r\n" +  
" }\r\n" +  
" }\r\n" +  
" }\r\n" +  
" ]\r\n" +  
" }\r\n" +  
" }\r\n" +  
"}";  
How could I achieve this.Please help me out.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 21, 2017, 7:49am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/4 "2017-11-21T07:49:26Z")

</div>

That’s another question. Could you open a new discussion?

Please format your code using `</>` icon as explained in [this guide](https://discuss.elastic.co/t/about-the-elasticsearch-category/21). It will make your post more readable.

Or use markdown style like:

````
```
CODE
```
````

---

<div class="post-metadata">

**Author:** ![kanchan](https://avatars.discourse-cdn.com/v4/letter/k/8797f3/32.png) [@kanchan](https://discuss.elastic.co/u/kanchan)\
**Post date:** [December 7, 2017, 7:30am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/5 "2017-12-07T07:30:05Z")

</div>

Thanks dadoonet, It works fine.But performance is not good.For 4 lack records it is taking approx 2 min which is too much.please let me know if i need to do any configuration changes or required to use any api like bulk or any setting.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 7, 2017, 8:41am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/6 "2017-12-07T08:41:02Z")

</div>

You meant for 4 documents?

---

<div class="post-metadata">

**Author:** ![kanchan](https://avatars.discourse-cdn.com/v4/letter/k/8797f3/32.png) [@kanchan](https://discuss.elastic.co/u/kanchan)\
**Post date:** [December 7, 2017, 9:48am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/7 "2017-12-07T09:48:24Z")

</div>

400000 records

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 7, 2017, 10:07am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/8 "2017-12-07T10:07:56Z")

</div>

What does a typical document looks like? What is its size?

---

<div class="post-metadata">

**Author:** ![kanchan](https://avatars.discourse-cdn.com/v4/letter/k/8797f3/32.png) [@kanchan](https://discuss.elastic.co/u/kanchan)\
**Post date:** [December 7, 2017, 10:21am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/9 "2017-12-07T10:21:20Z")

</div>

size parameter I set as 50000. and each document looks like this:

"\_source": {  
"product": "G10 Rates",  
"time\_id": 20121,  
"wallet": 0.000057,  
"entity\_name": "Brigade Capital Management",  
"country\_hq": "USA",  
"gap\_tox": "1-3",  
"entity\_id": 208625,  
"revenue": 0,  
"gap": 0.00001,  
"rank": "8+",  
"region\_hq": "Americas",  
"sow": 0,  
"region": "APAC",  
"sector": "Hedge Fund Managers"  
}

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 7, 2017, 10:39am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/10 "2017-12-07T10:39:53Z")

</div>

Can you try with less documents like 1000?  
Also what is the exact scroll query you are running?

Have a look at [https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-scroll.html#sliced-scroll](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-scroll.html#sliced-scroll)

Might help doing things in parallel.

Some hardware questions:

Do you have ssd drives?  
Is there anything in logs like gc information?

---

<div class="post-metadata">

**Author:** ![kanchan](https://avatars.discourse-cdn.com/v4/letter/k/8797f3/32.png) [@kanchan](https://discuss.elastic.co/u/kanchan)\
**Post date:** [December 8, 2017, 7:38am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/11 "2017-12-08T07:38:10Z")

</div>

I used sliced scrolling but unable to connect transport [client.my](http://client.my) code is:

```
TransportClient client = new PreBuiltTransportClient(settings)
			.addTransportAddress(new InetSocketTransportAddress(InetAddress.getByName("172.21.153.176"),9300));

```

**Error:**  
Exception in thread "main" java.lang.AbstractMethodError: org.elasticsearch.transport.TcpTransport.connectToChannels(Lorg/elasticsearch/cluster/node/DiscoveryNode;Lorg/elasticsearch/transport/ConnectionProfile;Ljava/util/function/Consumer;)Lorg/elasticsearch/transport/TcpTransport$NodeChannels;

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 8, 2017, 8:28am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/12 "2017-12-08T08:28:26Z")

</div>

This is not related.

How did you connect previously?

The fact that will use slice scroll in the future is totally unrelated IMO.  
Or I'm missing something in which case it could help if you share the full code and the full logs (stack trace).

---

<div class="post-metadata">

**Author:** ![kanchan](https://avatars.discourse-cdn.com/v4/letter/k/8797f3/32.png) [@kanchan](https://discuss.elastic.co/u/kanchan)\
**Post date:** [December 11, 2017, 11:35am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/13 "2017-12-11T11:35:29Z")

</div>

I used slice scroll api which works t9.Thanks!!! but performance is still slow.  
Any suggestion to improve performance of scroll api.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 11, 2017, 2:25pm UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/14 "2017-12-11T14:25:51Z")

</div>

What did you do at the end?

How many parallel scrolls did you run? And how?  
What does "slow" mean?

Do you monitor elasticsearch and your application to understand where is the bottleneck?

---

<div class="post-metadata">

**Author:** ![kanchan](https://avatars.discourse-cdn.com/v4/letter/k/8797f3/32.png) [@kanchan](https://discuss.elastic.co/u/kanchan)\
**Post date:** [December 15, 2017, 6:04am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/15 "2017-12-15T06:04:27Z")

</div>

Hi, We are fetching 40000 documents from elasticsearch using scroll API. We have given 10 slices and scroll size is 4000. But slice API is taking more time to fetch all data from elasticsearch. After fetching all data from ES, we have java code to iterate data but that code is very faster. We need to fetch all data very fast including connection to elasticsearch. We have used Transport Client to connect to ES.

IntStream.range(0, slices).parallel().forEach(i -\> {  
SearchSourceBuilder searchSourceBuilder = SearchSourceBuilder.searchSource();  
WrapperQueryBuilder qb = QueryBuilders.wrapperQuery(queryELKPart1);  
searchSourceBuilder.query(qb);

```
			SliceBuilder sliceBuilder = new SliceBuilder(i, slices);
			SearchResponse response = transclient.prepareSearch("entity_fact").setTypes("logs").
					setSource(searchSourceBuilder).
					setScroll(scrollTimeout).
					slice(sliceBuilder).
					setSize(scrollSize).
					setFetchSource(reqFields, null).
					setExplain(false).
					get();
			List<String> r = Arrays.stream(response.getHits().getHits())
					.map(SearchHit::getSourceAsString).collect(Collectors.toList());
			dataCollectionList.add(r);
		} );

```

Above code we have implemented for slice API to fetch all data from ES. How can we increase the performance of fetching data. Please provide solution. We need to fetch bulk data in few milliseconds. How can we achieve this?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 15, 2017, 8:48am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/16 "2017-12-15T08:48:42Z")

</div>

How long does it take?

---

<div class="post-metadata">

**Author:** ![kanchan](https://avatars.discourse-cdn.com/v4/letter/k/8797f3/32.png) [@kanchan](https://discuss.elastic.co/u/kanchan)\
**Post date:** [December 15, 2017, 8:53am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/17 "2017-12-15T08:53:31Z")

</div>

2 seconds we want in millisecond

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 15, 2017, 8:57am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/18 "2017-12-15T08:57:12Z")

</div>

How much data do you have on the each node? How much RAM and heap do you have per node? What type of storage do you have?

If you want to reduce the response time as much as possible, you probably want to make sure that the full data set can fit in the operating system file cache. If this is not feasible, using SSDs if you are not already will probably help too.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 15, 2017, 9:06am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/19 "2017-12-15T09:06:29Z")

</div>

2 seconds looks very good to me.  
What kind of use case are you trying to solve here?

---

<div class="post-metadata">

**Author:** ![kanchan](https://avatars.discourse-cdn.com/v4/letter/k/8797f3/32.png) [@kanchan](https://discuss.elastic.co/u/kanchan)\
**Post date:** [December 15, 2017, 10:29am UTC](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869/20 "2017-12-15T10:29:43Z")

</div>

In ignite it is taking 0.3 [millisecond.so](http://millisecond.so) could I improve perfomance if I create replicas of my index in multiple nodes.

[Next page](https://discuss.elastic.co/t/how-to-get-data-more-than-10000-in-elasticsearch/107869.md?page=2)
