# Scroll API(C# NEST) fails

**URL:** <https://discuss.elastic.co/t/scroll-api-c-nest-fails/129320>\
**Category:** Elasticsearch\
**Created:** [April 24, 2018, 2:11pm UTC](https://discuss.elastic.co/t/scroll-api-c-nest-fails/129320 "2018-04-24T14:11:49Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![mww1](https://avatars.discourse-cdn.com/v4/letter/m/ccd318/32.png) [@mww1](https://discuss.elastic.co/u/mww1)\
**Post date:** [April 24, 2018, 2:11pm UTC](https://discuss.elastic.co/t/scroll-api-c-nest-fails/129320/1 "2018-04-24T14:11:49Z")

</div>

Hello guys, I am using scan and scroll api for ES nest 5.5 to retrieve documents over 1.3 million in size(6 fields/document)  
my scroll api works fine if I retrieve less number of documents(around 800000) but when I want to retrieve large documents(1.3 million) it just fails. I dont know what is going wrong. I am trying to do this on C# visual studio and also make a web api and host as a microservice on service fabric app. both places it is failing. on visual studio, I get error: unexpected ES exception, sometimes I get web request aborted and sometimes " search context lost" even though I am giving new scrollid in each loop. below is my code. please help me. its urgent.

```auto
ISearchResponse<MyDocument> initialResponse = this.ElasticClient.Search<MyDocument>(scr => scr
    .Index(indexName)
    .From(0)
    .Take(scrollSize)
    .MatchAll()
    .Scroll(scrollTimeout)
    .SearchType(dfs));
    
List<MyDocument> results = new List<MyDocument>();

if (!initialResponse.IsValid || string.IsNullOrEmpty(initialResponse.ScrollId))
    throw new Exception(initialResponse.ServerError.Error.Reason);
if (initialResponse.Documents.Any())
    results.AddRange(initialResponse.Documents);
    
string scrollid = initialResponse.ScrollId;
bool isScrollSetHasData = true;

while (isScrollSetHasData)
{
    ISearchResponse<MyDocument> loopingResponse = this.ElasticClient.Scroll<MyDocument>(scrollTimeout, scrollid);
    if (loopingResponse.IsValid)
    {
        results.AddRange(loopingResponse.Documents);
        scrollid = loopingResponse.ScrollId;
    }
    isScrollSetHasData = loopingResponse.Documents.Any();
}

this.ElasticClient.ClearScroll(new ClearScrollRequest(scrollid));

```

---

<div class="post-metadata">

**Author:** ![forloop](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/forloop/32/9021_2.png) [@forloop](https://discuss.elastic.co/u/forloop)\
**Post date:** [May 1, 2018, 6:47am UTC](https://discuss.elastic.co/t/scroll-api-c-nest-fails/129320/2 "2018-05-01T06:47:25Z")

</div>

I've formatted your code this time; please use the `</>` button to format your code in future, or surround in ``` (triple backticks) with optional language after the first block. Trying to read unformatted code makes it much harder to help 🙂

Have you seen the `ScrollAll()` helper method? I think this would fit well with your use case:

```auto
var index = "my_index_name";
var numberOfShards = 3;
var seenDocuments = 0;
var documents = new ConcurrentBag<IReadOnlyCollection<MyDocument>>();

client.ScrollAll<MyDocument>("1m", numberOfShards, s => s
    .MaxDegreeOfParallelism(numberOfShards / 2)
    .Search(search => search
        .Index(index)
        .AllTypes()
        .MatchAll()
    )
).Wait(TimeSpan.FromMinutes(5), r =>
{
    documents.Add(r.SearchResponse.Documents);       
    Interlocked.Add(ref seenDocuments, r.SearchResponse.Hits.Count);
});

```

`ScrollAll()` can send concurrent scroll requests using slicing, and you can adjust the overall time to wait to fetch all documents with the `TimeSpan` passed to `.Wait()`. Using `ScrollAll()`, you can process documents as they are scrolled.

You can have more control with

```auto
var index = "my_index_name";
var numberOfShards = 3;
var seenDocuments = 0;
var documents = new ConcurrentBag<IReadOnlyCollection<MyDocument>>();

var observable = client.ScrollAll<MyDocument>("1m", numberOfShards, s => s
    .MaxDegreeOfParallelism(numberOfShards / 2)
    .Search(search => search
        .Index(index)
        .AllTypes()
        .MatchAll()
    )
);

var waitHandle = new ManualResetEvent(false);
Exception exception = null;
observable.Subscribe(new ScrollAllObserver<MyDocument>(
    onNext: r =>
    {
        documents.Add(r.SearchResponse.Documents);
        Interlocked.Add(ref seenDocuments, r.SearchResponse.Hits.Count);
    },
    onError: e => 
    { 
        exception = e;
        waitHandle.Set();
    },
    onCompleted: () => waitHandle.Set()
));

waitHandle.WaitOne();

if (exception != null)
    throw exception;

```

---

<div class="post-metadata">

**Author:** ![mww1](https://avatars.discourse-cdn.com/v4/letter/m/ccd318/32.png) [@mww1](https://discuss.elastic.co/u/mww1)\
**Post date:** [May 1, 2018, 4:33pm UTC](https://discuss.elastic.co/t/scroll-api-c-nest-fails/129320/3 "2018-05-01T16:33:39Z")

</div>

thanks a lot Russ. this is great. but I dont understand how is it scrolling n batches? is it trying to pull all the data at once? if so, will it not put extra load on the ES cluster?  
can you please see the below code which I am trying to execute.? do you think this should work fine for extracting documents over a million in size?

```auto
var index = "index name";
				var numberOfShards = 5;
				var seenDocuments = 0;
				var documents = new ConcurrentBag<IReadOnlyCollection<object>>();

				 client.ScrollAll<object>("5m", numberOfShards, g =>g.MaxDegreeOfParallelism(numberOfShards / 2).Search(	
					scr => scr.Index(DefaultIndex)
				   .From(0)
				   .Take(30000)
				   .Source(r => r.Includes(i => i.Fields(Fields)))
				   .Query(q => q
				   	.Nested(n => n
				   	.Path("A_field_name")
				   	.Query(q1 => q1.MatchAll()
				   )
				   )
				   )
				   .Type("punchouts")
				  // .SearchType(Elasticsearch.Net.SearchType.DfsQueryThenFetch)
				   .Scroll("15m"))).Wait(TimeSpan.FromMinutes(5), z =>
				   {
					   documents.Add(z.SearchResponse.Documents);
					   Interlocked.Add(ref seenDocuments, z.SearchResponse.Hits.Count);
				   }); ;

```

do you think this will work fine? if not can you please help me figure out an optimum soltion for the same.

thanks a ton!

---

<div class="post-metadata">

**Author:** ![mww1](https://avatars.discourse-cdn.com/v4/letter/m/ccd318/32.png) [@mww1](https://discuss.elastic.co/u/mww1)\
**Post date:** [May 2, 2018, 3:52am UTC](https://discuss.elastic.co/t/scroll-api-c-nest-fails/129320/4 "2018-05-02T03:52:36Z")

</div>

Russ I did try both the method and they seem to work on dataset of hundred thousand. one question why does the document store it in array of documents? what is the relevance of number of shard?  
Finally, will this work for fetching data over a million documents? (1 doc = 6 fields)

Again,  
Thanks a ton!

---

<div class="post-metadata">

**Author:** ![forloop](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/forloop/32/9021_2.png) [@forloop](https://discuss.elastic.co/u/forloop)\
**Post date:** [May 2, 2018, 12:07pm UTC](https://discuss.elastic.co/t/scroll-api-c-nest-fails/129320/5 "2018-05-02T12:07:24Z")

</div>

> [@mww1](#):
>
> I dont understand how is it scrolling n batches

The `ScrollAllObservable` that `ScrollAll()` uses takes advantage of [Sliced Scrolling within Elasticsearch's Scroll API](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-scroll.html#sliced-scroll) to slice each scroll request into a number of slices that can be processed concurrently.

> [@mww1](#):
>
> is it trying to pull all the data at once?

No, it is still issuing scroll requests to scroll through all of the documents that satisfy the query. Within the `onNext` delegate of the `ScrollAllObserver<T>`, you can determine what to do with the documents returned from each scroll response. To reduce memory footprint, you may be able to process documents as they are returned, in this delegate.

> [@mww1](#):
>
> if so, will it not put extra load on the ES cluster?

No more load than issuing scroll requests using the Scroll API.

> [@mww1](#):
>
> do you think this should work fine for extracting documents over a million in size?

"should work fine" is very subjective 😄 This will scroll through a million documents, and if you need to retrieve that many documents, the scroll API is a good candidate for doing this.

---

<div class="post-metadata">

**Author:** ![mww1](https://avatars.discourse-cdn.com/v4/letter/m/ccd318/32.png) [@mww1](https://discuss.elastic.co/u/mww1)\
**Post date:** [May 4, 2018, 6:00pm UTC](https://discuss.elastic.co/t/scroll-api-c-nest-fails/129320/6 "2018-05-04T18:00:14Z")

</div>

oops you tricked me. 😃 which one is better? ScrollAll or Scroll?

```auto
**Scroll => (which I originally used)**
ISearchResponse<MyDocument> initialResponse = this.ElasticClient.Search<MyDocument>(scr => scr
    .Index(indexName)
    .From(0)
    .Take(scrollSize)
    .MatchAll()
    .Scroll(scrollTimeout)
    .SearchType(dfs));
    
List<MyDocument> results = new List<MyDocument>();

if (!initialResponse.IsValid || string.IsNullOrEmpty(initialResponse.ScrollId))
    throw new Exception(initialResponse.ServerError.Error.Reason);
if (initialResponse.Documents.Any())
    results.AddRange(initialResponse.Documents);
    
string scrollid = initialResponse.ScrollId;
bool isScrollSetHasData = true;

while (isScrollSetHasData)
{
    ISearchResponse<MyDocument> loopingResponse = this.ElasticClient.Scroll<MyDocument>(scrollTimeout, scrollid);
    if (loopingResponse.IsValid)
    {
        results.AddRange(loopingResponse.Documents);
        scrollid = loopingResponse.ScrollId;
    }
    isScrollSetHasData = loopingResponse.Documents.Any();
}

this.ElasticClient.ClearScroll(new ClearScrollRequest(scrollid));

```

**or this one ?**  
**ScrollAll =\> which you suggested**

```auto
var index = "my_index_name";
var numberOfShards = 3;
var seenDocuments = 0;
var documents = new ConcurrentBag<IReadOnlyCollection<MyDocument>>();

client.ScrollAll<MyDocument>("1m", numberOfShards, s => s
    .MaxDegreeOfParallelism(numberOfShards / 2)
    .Search(search => search
        .Index(index)
        .AllTypes()
        .MatchAll()
    )
).Wait(TimeSpan.FromMinutes(5), r =>
{
    documents.Add(r.SearchResponse.Documents);       
    Interlocked.Add(ref seenDocuments, r.SearchResponse.Hits.Count);
});

```

**My target is to retrieve over a million documents**

---

<div class="post-metadata">

**Author:** ![forloop](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/forloop/32/9021_2.png) [@forloop](https://discuss.elastic.co/u/forloop)\
**Post date:** [May 5, 2018, 4:13am UTC](https://discuss.elastic.co/t/scroll-api-c-nest-fails/129320/7 "2018-05-05T04:13:36Z")

</div>

> [@mww1](#):
>
> which one is better?

I'd suggest trying both approaches and seeing which one works better for you. I would expect `ScrollAll` to perform better since it performs concurrent scrolling. `ScrollAll` is provided as a helper with NEST so that you don't need to write a concurrent scroll implementation yourself 🙂

---

<div class="post-metadata">

**Author:** ![mww1](https://avatars.discourse-cdn.com/v4/letter/m/ccd318/32.png) [@mww1](https://discuss.elastic.co/u/mww1)\
**Post date:** [May 5, 2018, 5:17pm UTC](https://discuss.elastic.co/t/scroll-api-c-nest-fails/129320/8 "2018-05-05T17:17:51Z")

</div>

perfect. thanks a lot.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 2, 2018, 5:17pm UTC](https://discuss.elastic.co/t/scroll-api-c-nest-fails/129320/9 "2018-06-02T17:17:51Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
