# FSCrawler errors with large PDF files - Hard failure and timeout

**URL:** <https://discuss.elastic.co/t/fscrawler-errors-with-large-pdf-files-hard-failure-and-timeout/192082>\
**Category:** Elasticsearch\
**Created:** [July 24, 2019, 3:55pm UTC](https://discuss.elastic.co/t/fscrawler-errors-with-large-pdf-files-hard-failure-and-timeout/192082 "2019-07-24T15:55:15Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![erick\_andrade](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/erick_andrade/32/50829_2.png) [@erick\_andrade](https://discuss.elastic.co/u/erick_andrade)\
**Post date:** [July 24, 2019, 3:55pm UTC](https://discuss.elastic.co/t/fscrawler-errors-with-large-pdf-files-hard-failure-and-timeout/192082/1 "2019-07-24T15:55:15Z")

</div>

Hi, folks!

I'm trying to index some large PDF files (~70MB) in Elasticsearch using FSCrawler, but not all of them are being indexed, and some exceptions are thrown, like following:

```
DEBUG [f.p.e.c.f.FsParserAbstract] [/04/dejt-jud_26-04-2019.pdf] can be indexed: [true]
DEBUG [f.p.e.c.f.FsParserAbstract] - file: /04/dejt-jud_26-04-2019.pdf
DEBUG [f.p.e.c.f.FsParserAbstract] fetching content from [/home/erick/Elastic/DEJT/04],[dejt-jud_26-04-2019.pdf]
DEBUG [f.p.e.c.f.f.FsCrawlerUtil] computeVirtualPathName(/home/erick/Elastic/DEJT, /home/erick/Elastic/DEJT/04/dejt-jud_26-04-2019.pdf) = /04/dejt-jud_26-04-2019.pdf
WARN [f.p.e.c.f.c.v.ElasticsearchClientV7] Got a hard failure when executing the bulk request
java.net.SocketTimeoutException: 30.000 milliseconds timeout on connection http-outgoing-6 [ACTIVE]

```

My \_settings.yaml file:

```
---
name: "desenv_dejt_jud"
fs:
  url: "/home/erick/Elastic/DEJT"
  update_rate: "30m"
  excludes:
  - ".*dejt-adm.*"
  json_support: false
  filename_as_id: false
  add_filesize: true
  remove_deleted: true
  add_as_inner_object: false
  store_source: false
  index_content: true
  attributes_support: false
  raw_metadata: false
  xml_support: false
  index_folders: true
  lang_detect: false
  continue_on_error: false
  indexed_chars: "100%"
  ocr:
    language: "eng"
    enabled: false
    pdf_strategy: "ocr_and_text"
elasticsearch:
  nodes:
  - url: "http://127.0.0.1:9200"
  bulk_size: 1
  flush_interval: "60s"
  byte_size: "10m"
  pipeline: "dejt_jud_pipeline"

```

Without "indexed\_chars: 100%" it works ok.

I'm running a single Elasticsearch node with Xms10g/Xmx10g, and starting fscrawler like this:

```
FS_JAVA_OPTS="-Xmx4g -Xms4g" ./fscrawler-es7-2.7-SNAPSHOT/bin/fscrawler desenv_dejt_jud --debug

```

If i try to index file by file, sometimes the exception are thrown, but document gets indexed.  
With all documents in folder, many exceptions "hard failure/timeout 30.000" are thrown, and few documents gets indexed.

I tried some different values for bulk\_size, byte\_size e flush\_interval, but result is the same.

I also looked for a way to increase the Elasticsearch connection timeout value, but found nothing about.

Have any of you seen anything like this before?

Thank you!!

--  
Erick

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 21, 2019, 3:55pm UTC](https://discuss.elastic.co/t/fscrawler-errors-with-large-pdf-files-hard-failure-and-timeout/192082/2 "2019-08-21T15:55:16Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
