# Error while indexing documents into ES using Fscrawler

**URL:** <https://discuss.elastic.co/t/error-while-indexing-documents-into-es-using-fscrawler/155999>\
**Category:** Elasticsearch\
**Created:** [November 9, 2018, 8:15am UTC](https://discuss.elastic.co/t/error-while-indexing-documents-into-es-using-fscrawler/155999 "2018-11-09T08:15:34Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jasmeet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jasmeet/32/69393_2.png) [@Jasmeet](https://discuss.elastic.co/u/Jasmeet)\
**Post date:** [November 9, 2018, 8:15am UTC](https://discuss.elastic.co/t/error-while-indexing-documents-into-es-using-fscrawler/155999/1 "2018-11-09T08:15:35Z")

</div>

Hi, I am using Fscrawler to index a large set of documents kept in varous folders. I have created separate jobs for all the major folders and i run each job in Fscrawler. Some of the folders are quite large (\>180 Gb) and contain some sub folders also for which creating individual jobs is very cumbersome process. In one such folder, I ran FScrawaler and after running for the entire day it gave an error which I am reproducing below. Can someone pl guide as to why the error is coming and how to resolve it.  
2. I had earlier also run the crawler on the same folder and got an error, so Fscrawler tries to reindex every document/folder from the beginning every time it is started as there is no status.json file created if the crawler exits with an error.  
Thanks  
JS  
// ERROR

16:37:08,809 DEBUG [f.p.e.c.f.c.FileAbstractor] Listing local files from E:\MISC FILES\IT Attendance System\Record of Attendance\IT attendance\2018\December 2018  
16:37:08,809 DEBUG [f.p.e.c.f.c.FileAbstractor] 0 local files found  
16:37:08,809 DEBUG [f.p.e.c.f.FsParser] Looking for removed files in [E:\MISC FILES\IT Attendance System\Record of Attendance\IT attendance\2018\December 2018]...  
16:37:08,809 TRACE [f.p.e.c.f.FsParser] Querying elasticsearch for files in dir [path.root:4fcf95a09d4a7831e5910733dbaa]  
16:37:12,982 TRACE [f.p.e.c.f.c.ElasticsearchClientManager] Sending a bulk request of [1] requests  
16:37:13,581 TRACE [f.p.e.c.f.c.ElasticsearchClientManager] Sending a bulk request of [1] requests  
16:37:34,027 WARN [f.p.e.c.f.c.ElasticsearchClientManager] **Got a hard failure when executing the bulk request**  
java.net.SocketException: null  
at org.apache.http.nio.protocol.HttpAsyncRequestExecutor.timeout(HttpAsyncRequestExecutor.java:375) [httpcore-nio-4.4.5.jar:4.4.5]  
at org.apache.http.impl.nio.client.InternalIODispatch.onTimeout(InternalIODispatch.java:92) [httpasyncclient-4.1.2.jar:4.1.2]  
at org.apache.http.impl.nio.client.InternalIODispatch.onTimeout(InternalIODispatch.java:39) [httpasyncclient-4.1.2.jar:4.1.2]  
at org.apache.http.impl.nio.reactor.AbstractIODispatch.timeout(AbstractIODispatch.java:175) [httpcore-nio-4.4.5.jar:4.4.5]  
at org.apache.http.impl.nio.reactor.BaseIOReactor.sessionTimedOut(BaseIOReactor.java:263) [httpcore-nio-4.4.5.jar:4.4.5]  
at org.apache.http.impl.nio.reactor.AbstractIOReactor.timeoutCheck(AbstractIOReactor.java:492) [httpcore-nio-4.4.5.jar:4.4.5]  
at org.apache.http.impl.nio.reactor.BaseIOReactor.validate(BaseIOReactor.java:213) [httpcore-nio-4.4.5.jar:4.4.5]  
at org.apache.http.impl.nio.reactor.AbstractIOReactor.execute(AbstractIOReactor.java:280) [httpcore-nio-4.4.5.jar:4.4.5]  
at org.apache.http.impl.nio.reactor.BaseIOReactor.execute(BaseIOReactor.java:104) [httpcore-nio-4.4.5.jar:4.4.5]  
at org.apache.http.impl.nio.reactor.AbstractMultiworkerIOReactor$Worker.run(AbstractMultiworkerIOReactor.java:588) [httpcore-nio-4.4.5.jar:4.4.5]  
at java.lang.Thread.run(Unknown Source) [?:1.8.0\_171]  
16:37:38,927 WARN [f.p.e.c.f.FsParser] **Error while crawling E:\MISC FILES: listener timeout after waiting for [30000] ms**  
16:37:38,999 WARN [f.p.e.c.f.FsParser] Full stacktrace  
java.io.IOException: listener timeout after waiting for [30000] ms  
at org.elasticsearch.client.RestClient$SyncResponseListener.get(RestClient.java:684) ~[elasticsearch-rest-client-6.3.2.jar:6.3.2]  
at org.elasticsearch.client.RestClient.performRequest(RestClient.java:235) ~[elasticsearch-rest-client-6.3.2.jar:6.3.2]  
at org.elasticsearch.client.RestClient.performRequest(RestClient.java:198) ~[elasticsearch-rest-client-6.3.2.jar:6.3.2]  
at org.elasticsearch.client.RestHighLevelClient.performRequest(RestHighLevelClient.java:522) ~[elasticsearch-rest-high-level-client-6.3.2.jar:6.3.2]  
at org.elasticsearch.client.RestHighLevelClient.performRequestAndParseEntity(RestHighLevelClient.java:508) ~[elasticsearch-rest-high-level-client-6.3.2.jar:6.3.2]  
at org.elasticsearch.client.RestHighLevelClient.search(RestHighLevelClient.java:404) ~[elasticsearch-rest-high-level-client-6.3.2.jar:6.3.2]  
at fr.pilato.elasticsearch.crawler.fs.FsParser.getFileDirectory(FsParser.java:356) ~[fscrawler-core-2.5.jar:?]  
at fr.pilato.elasticsearch.crawler.fs.FsParser.addFilesRecursively(FsParser.java:307) ~[fscrawler-core-2.5.jar:?]  
at fr.pilato.elasticsearch.crawler.fs.FsParser.addFilesRecursively(FsParser.java:290) ~[fscrawler-core-2.5.jar:?]  
at fr.pilato.elasticsearch.crawler.fs.FsParser.addFilesRecursively(FsParser.java:290) ~[fscrawler-core-2.5.jar:?]  
at fr.pilato.elasticsearch.crawler.fs.FsParser.addFilesRecursively(FsParser.java:290) ~[fscrawler-core-2.5.jar:?]  
at fr.pilato.elasticsearch.crawler.fs.FsParser.addFilesRecursively(FsParser.java:290) ~[fscrawler-core-2.5.jar:?]  
at fr.pilato.elasticsearch.crawler.fs.FsParser.addFilesRecursively(FsParser.java:290) ~[fscrawler-core-2.5.jar:?]  
at fr.pilato.elasticsearch.crawler.fs.FsParser.run(FsParser.java:167) [fscrawler-core-2.5.jar:?]  
at java.lang.Thread.run(Unknown Source) [?:1.8.0\_171]  
16:37:39,139 INFO [f.p.e.c.f.FsParser] FS crawler is stopping after 1 run  
16:37:39,428 DEBUG [f.p.e.c.f.FsCrawlerImpl] Closing FS crawler [e\_drive\_misc]  
16:37:39,505 DEBUG [f.p.e.c.f.FsCrawlerImpl] FS crawler thread is now stopped  
16:37:39,505 DEBUG [f.p.e.c.f.c.ElasticsearchClientManager] Closing Elasticsearch client manager  
16:37:44,226 WARN [f.p.e.c.f.c.ElasticsearchClientManager] Got a hard failure when executing the bulk request  
java.net.SocketTimeoutException: null  
at org.apache.http.nio.protocol.HttpAsyncRequestExecutor.timeout(HttpAsyncRequestExecutor.java:375) [httpcore-nio-4.4.5.jar:4.4.5]  
at org.apache.http.impl.nio.client.InternalIODispatch.onTimeout(InternalIODispatch.java:92) [httpasyncclient-4.1.2.jar:4.1.2]  
at org.apache.http.impl.nio.client.InternalIODispatch.onTimeout(InternalIODispatch.java:39) [httpasyncclient-4.1.2.jar:4.1.2]

16:38:02,138 TRACE [f.p.e.c.f.c.ElasticsearchClientManager] Executed bulk request with [1] requests  
16:38:02,328 DEBUG [f.p.e.c.f.c.ElasticsearchClient] Closing REST client  
16:38:02,629 DEBUG [f.p.e.c.f.FsCrawlerImpl] ES Client Manager stopped  
16:38:02,629 INFO [f.p.e.c.f.FsCrawlerImpl] FS crawler [e\_drive\_misc] stopped  
16:38:03,321 DEBUG [f.p.e.c.f.FsCrawlerImpl] Closing FS crawler [e\_drive\_misc]  
16:38:03,422 DEBUG [f.p.e.c.f.FsCrawlerImpl] FS crawler thread is now stopped  
16:38:03,422 DEBUG [f.p.e.c.f.c.ElasticsearchClientManager] Closing Elasticsearch client manager  
16:38:03,422 DEBUG [f.p.e.c.f.c.ElasticsearchClient] Closing REST client  
16:38:03,422 DEBUG [f.p.e.c.f.FsCrawlerImpl] ES Client Manager stopped  
16:38:03,422 INFO [f.p.e.c.f.FsCrawlerImpl] FS crawler [e\_drive\_misc] stopped  
//

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 9, 2018, 8:09pm UTC](https://discuss.elastic.co/t/error-while-indexing-documents-into-es-using-fscrawler/155999/2 "2018-11-09T20:09:28Z")

</div>

What are your FSCrawler settings for this job?

Please format your code, logs or configuration files using `</>` icon as explained in [this guide](https://discuss.elastic.co/t/about-the-elasticsearch-category/21) and not the citation button. It will make your post more readable.

Or use markdown style like:

````
```
CODE
```

````

This is the icon to use if you are not using markdown format:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/7/e/7e6e239431ec2d71cbf1beef741f2e93e7cc762c.jpg)

There's a live preview panel for exactly this reasons.

---

<div class="post-metadata">

**Author:** ![Jasmeet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jasmeet/32/69393_2.png) [@Jasmeet](https://discuss.elastic.co/u/Jasmeet)\
**Post date:** [November 11, 2018, 5:13pm UTC](https://discuss.elastic.co/t/error-while-indexing-documents-into-es-using-fscrawler/155999/3 "2018-11-11T17:13:57Z")

</div>

Hi, sorry for the delay in responding. My settings are given below:

```auto

  "name" : "test", "fs" : { "url" :"E:\\MISC FILES"

```

```auto
"update_rate" : "5m",

```

```auto
    "excludes" : ["*/~*"], "json_support" : false, "filename_as_id" : false, "add_filesize" : true, "remove_deleted" : true, "add_as_inner_object" : false, "store_source" : false, "index_content" : true, "attributes_support" : false, "raw_metadata" : true, "xml_support" : false, "index_folders" : true, "lang_detect" : false, "continue_on_error" : false, "pdf_ocr" : true, "ocr" : { "language" : "eng" } }, "elasticsearch" : { "nodes" : [{ "host" : "127.0.0.1", "port" : 9200, "scheme" : "HTTP" }], "bulk_size" : 100, "flush_interval" : "5s", "byte_size" : "10mb" }, "rest" : { "scheme" : "HTTP", "host" : "127.0.0.1", "port" : 8080, "endpoint" : "fscrawler" } }

```

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 11, 2018, 6:30pm UTC](https://discuss.elastic.co/t/error-while-indexing-documents-into-es-using-fscrawler/155999/4 "2018-11-11T18:30:38Z")

</div>

Could you please format correctly your json so it will be readable?

---

<div class="post-metadata">

**Author:** ![Jasmeet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jasmeet/32/69393_2.png) [@Jasmeet](https://discuss.elastic.co/u/Jasmeet)\
**Post date:** [November 11, 2018, 7:33pm UTC](https://discuss.elastic.co/t/error-while-indexing-documents-into-es-using-fscrawler/155999/5 "2018-11-11T19:33:05Z")

</div>

I am actually doing it from a mobile phone as I am traveling and don't have access to a computer. Tried to format using \</\>

`Preformatted text`{  
"name" : "test",  
"fs" : {  
"url" : "C:\Users\Sun\Documents\test",  
"update\_rate" : "5m",  
"excludes" : ["_/~_"],  
"json\_support" : false,  
"filename\_as\_id" : false,  
"add\_filesize" : true,  
"remove\_deleted" : true,  
"add\_as\_inner\_object" : false,  
"store\_source" : false,  
"index\_content" : true,  
"attributes\_support" : false,  
"raw\_metadata" : true,  
"xml\_support" : false,  
"index\_folders" : true,  
"lang\_detect" : false,  
"continue\_on\_error" : false,  
"pdf\_ocr" : true,  
"ocr" : {  
"language" : "eng"  
}  
},  
"elasticsearch" : {  
"nodes" : [ {  
"host" : "127.0.0.1",  
"port" : 9200,  
"scheme" : "HTTP"  
} ],  
"bulk\_size" : 100,  
"flush\_interval" : "5s",  
"byte\_size" : "10mb"  
},  
"rest" : {  
"scheme" : "HTTP",  
"host" : "127.0.0.1",  
"port" : 8080,  
"endpoint" : "fscrawler"  
}  
}  
indent preformatted text by 4 spaces

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 11, 2018, 8:20pm UTC](https://discuss.elastic.co/t/error-while-indexing-documents-into-es-using-fscrawler/155999/6 "2018-11-11T20:20:03Z")

</div>

Please format your code, logs or configuration files using `</>` icon as explained in [this guide](https://discuss.elastic.co/t/about-the-elasticsearch-category/21) and not the citation button. It will make your post more readable.

Or use markdown style like:

````
```
CODE
```

````

This is the icon to use if you are not using markdown format:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/7/e/7e6e239431ec2d71cbf1beef741f2e93e7cc762c.jpg)

There's a live preview panel for exactly this reasons.

Lots of people read these forums, and many of them will simply skip over a post that is difficult to read, because it's just too large an investment of their time to try and follow a wall of badly formatted text.  
If your goal is to get an answer to your questions, it's in your interest to make it as easy to read and understand as possible.  
Please update your post.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 9, 2018, 8:20pm UTC](https://discuss.elastic.co/t/error-while-indexing-documents-into-es-using-fscrawler/155999/7 "2018-12-09T20:20:06Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
