# With FSCrawler 2.7 I am not able to index pdf and other types of documents which worked fine with 2.6

**URL:** <https://discuss.elastic.co/t/with-fscrawler-2-7-i-am-not-able-to-index-pdf-and-other-types-of-documents-which-worked-fine-with-2-6/206246>\
**Category:** Elasticsearch\
**Created:** [November 2, 2019, 4:31pm UTC](https://discuss.elastic.co/t/with-fscrawler-2-7-i-am-not-able-to-index-pdf-and-other-types-of-documents-which-worked-fine-with-2-6/206246 "2019-11-02T16:31:09Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![rkmohapatra](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rkmohapatra/32/54250_2.png) [@rkmohapatra](https://discuss.elastic.co/u/rkmohapatra)\
**Post date:** [November 2, 2019, 4:31pm UTC](https://discuss.elastic.co/t/with-fscrawler-2-7-i-am-not-able-to-index-pdf-and-other-types-of-documents-which-worked-fine-with-2-6/206246/1 "2019-11-02T16:31:09Z")

</div>

FSCrawler log says this.

```
[Elasticsearch exception [type=mapper_parsing_exception, reason=failed to parse]]; nested: ElasticsearchException[Elasticsearch exception [type=illegal_argument_exception, reason=Malformed content, found extra data after parsing: FIELD_NAME]];
16:10:16,190 ^[[36mDEBUG^[[m [f.p.e.c.f.c.v.ElasticsearchClientV7] Error caught for [ips-internal-doc-index]/[_doc]/[8f531bfbb22847e4c87c31a17a6284]: ElasticsearchException[Elasticsearch exception [type=mapper_parsing_exception, reason=failed to parse]]; nested: ElasticsearchException[Elasticsearch exception [type=illegal_argument_exception, reason=Malformed content, found extra data after parsing: FIELD_NAME]];
16:10:16,191 ^[[33mWARN ^[[m [f.p.e.c.f.c.v.ElasticsearchClientV7] Got [3] failures of [4] requests

```

Elasticsearch log says this.

```
"Caused by: java.lang.IllegalArgumentException: Malformed content, found extra data after parsing: FIELD_NAME",
"at org.elasticsearch.index.mapper.DocumentParser.validateEnd(DocumentParser.java:146) ~[elasticsearch-7.3.0.jar:7.3.0]",
"at org.elasticsearch.index.mapper.DocumentParser.parseDocument(DocumentParser.java:72) ~[elasticsearch-7.3.0.jar:7.3.0]",
"... 34 more"] }

```

The same configuration worked fine for me with the same documents in FSCrawler 2.6. I am using all default configuration of Elasticsearch.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 2, 2019, 4:48pm UTC](https://discuss.elastic.co/t/with-fscrawler-2-7-i-am-not-able-to-index-pdf-and-other-types-of-documents-which-worked-fine-with-2-6/206246/3 "2019-11-02T16:48:28Z")

</div>

Could you share a way to reproduce it (FSCrawler settings and a pdf file causing this?)

---

<div class="post-metadata">

**Author:** ![rkmohapatra](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rkmohapatra/32/54250_2.png) [@rkmohapatra](https://discuss.elastic.co/u/rkmohapatra)\
**Post date:** [November 3, 2019, 1:22pm UTC](https://discuss.elastic.co/t/with-fscrawler-2-7-i-am-not-able-to-index-pdf-and-other-types-of-documents-which-worked-fine-with-2-6/206246/4 "2019-11-03T13:22:34Z")

</div>

Here is the job configuration file.  
{  
"name" : "job-internal-doc-index",  
"fs" : {  
"url" : "/path/to/docs",  
"update\_rate" : "30m",  
"includes" : ["_.pdf", "_.xls", "_.xlsx","_.ppt", "_.doc","_.docx"],  
"json\_support" : false,  
"filename\_as\_id" : false,  
"add\_filesize" : true,  
"remove\_deleted" : false,  
"add\_as\_inner\_object" : false,  
"store\_source" : true,  
"index\_content" : true,  
"indexed\_chars" : "-1",  
"attributes\_support" : false,  
"raw\_metadata" : true,  
"xml\_support" : false,  
"index\_folders" : true,  
"lang\_detect" : true,  
"continue\_on\_error" : false  
},  
"elasticsearch" : {  
"nodes" : [{  
"url" : "ES\_URL"  
}],  
"index" : "index1",  
"index\_folder": "index\_folder1",  
"pipeline": "ips",  
"bulk\_size" : 1000,  
"flush\_interval" : "5s",  
"byte\_size" : "10mb"  
}  
}

---

<div class="post-metadata">

**Author:** ![rkmohapatra](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rkmohapatra/32/54250_2.png) [@rkmohapatra](https://discuss.elastic.co/u/rkmohapatra)\
**Post date:** [November 3, 2019, 1:24pm UTC](https://discuss.elastic.co/t/with-fscrawler-2-7-i-am-not-able-to-index-pdf-and-other-types-of-documents-which-worked-fine-with-2-6/206246/5 "2019-11-03T13:24:10Z")

</div>

How do I attach the sample pdf file here? It's not allowing to upload pdf here.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 3, 2019, 2:14pm UTC](https://discuss.elastic.co/t/with-fscrawler-2-7-i-am-not-able-to-index-pdf-and-other-types-of-documents-which-worked-fine-with-2-6/206246/6 "2019-11-03T14:14:47Z")

</div>

Could you share it somewhere and paste the link here?

---

<div class="post-metadata">

**Author:** ![rkmohapatra](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rkmohapatra/32/54250_2.png) [@rkmohapatra](https://discuss.elastic.co/u/rkmohapatra)\
**Post date:** [November 3, 2019, 2:32pm UTC](https://discuss.elastic.co/t/with-fscrawler-2-7-i-am-not-able-to-index-pdf-and-other-types-of-documents-which-worked-fine-with-2-6/206246/8 "2019-11-03T14:32:11Z")

</div>

See if you can access the sample document. In fact, almost all documents are failing with the same error. Though FSCrawler can parse the document, extract metadata and create the JSON.

---

<div class="post-metadata">

**Author:** ![rkmohapatra](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rkmohapatra/32/54250_2.png) [@rkmohapatra](https://discuss.elastic.co/u/rkmohapatra)\
**Post date:** [November 4, 2019, 2:40pm UTC](https://discuss.elastic.co/t/with-fscrawler-2-7-i-am-not-able-to-index-pdf-and-other-types-of-documents-which-worked-fine-with-2-6/206246/9 "2019-11-04T14:40:06Z")

</div>

Any clue to debug it further? Appreciate your help as we need to resolve it asap.

---

<div class="post-metadata">

**Author:** ![rkmohapatra](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rkmohapatra/32/54250_2.png) [@rkmohapatra](https://discuss.elastic.co/u/rkmohapatra)\
**Post date:** [November 5, 2019, 9:57am UTC](https://discuss.elastic.co/t/with-fscrawler-2-7-i-am-not-able-to-index-pdf-and-other-types-of-documents-which-worked-fine-with-2-6/206246/10 "2019-11-05T09:57:08Z")

</div>

@dadoonet It worked for me after I deleted the existing index and recreated it as per the discussion in [https://github.com/dadoonet/fscrawler/issues/755](https://github.com/dadoonet/fscrawler/issues/755). Will come back if I get any more issues. For now, the issue seems resolved.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 5, 2019, 12:35pm UTC](https://discuss.elastic.co/t/with-fscrawler-2-7-i-am-not-able-to-index-pdf-and-other-types-of-documents-which-worked-fine-with-2-6/206246/11 "2019-11-05T12:35:44Z")

</div>

Great. If you have any idea of what happened and how to reproduce the problem please open an issue in Fscrawler project. Thanks

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 3, 2019, 12:35pm UTC](https://discuss.elastic.co/t/with-fscrawler-2-7-i-am-not-able-to-index-pdf-and-other-types-of-documents-which-worked-fine-with-2-6/206246/12 "2019-12-03T12:35:45Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
