# Pdf Parsing Issues when using FS crawler

**URL:** <https://discuss.elastic.co/t/pdf-parsing-issues-when-using-fs-crawler/364488>\
**Category:** Elasticsearch\
**Created:** [August 6, 2024, 3:32pm UTC](https://discuss.elastic.co/t/pdf-parsing-issues-when-using-fs-crawler/364488 "2024-08-06T15:32:52Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![saifstech26](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/saifstech26/32/126560_2.png) [@saifstech26](https://discuss.elastic.co/u/saifstech26)\
**Post date:** [August 6, 2024, 3:32pm UTC](https://discuss.elastic.co/t/pdf-parsing-issues-when-using-fs-crawler/364488/1 "2024-08-06T15:32:52Z")

</div>

Hi @dadoonet,

I have started using FS crawler recently. We have a requirement to crawl all files on the local file system, most of which are PDFs along with other doc formats. The thing with these PDFs is that all PDFs are not generated from same tool. In particular we have around 15K+ files generated using the **Acrobat Capture 3.0** tool. And all the content which we have crawled and indexed into ES are malformed. The FS crawler works like a charm with other PDFs and doc formats when generated using - **Acrobat PDFMaker 9.1 for Word** or even MS Word.

We tried to troubleshoot on this behavior and could not find any work around for those PDFs.

The example output looks like this:

`S a m p l e t e x t is not a r e a l l y w o r k i n g .`

some words are correctly formed and some are not.

We have also copied the entire text of PDF and pasted on text editors like notepad++ and found the same output as that of FS crawler. We have also tried with other PDF readers like pdfreader from pypi, though the words are formed well, but bullets are breaking out.

The sample output looks like this:

```auto
1.
2.
3.
Dummy text
Lorem Ipsum
Hello World

```

Please let me know how can we resolve these parsing issues.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 6, 2024, 4:20pm UTC](https://discuss.elastic.co/t/pdf-parsing-issues-when-using-fs-crawler/364488/2 "2024-08-06T16:20:07Z")

</div>

Welcome!

It would help if you could share a sample file so I can test it and be some options in Tika if any.

---

<div class="post-metadata">

**Author:** ![saifstech26](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/saifstech26/32/126560_2.png) [@saifstech26](https://discuss.elastic.co/u/saifstech26)\
**Post date:** [August 7, 2024, 8:30am UTC](https://discuss.elastic.co/t/pdf-parsing-issues-when-using-fs-crawler/364488/3 "2024-08-07T08:30:26Z")

</div>

Thanks for the swift reply. I am waiting for a sample PDF that can be shared here. I really appreciate your patience.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 7, 2024, 8:46am UTC](https://discuss.elastic.co/t/pdf-parsing-issues-when-using-fs-crawler/364488/4 "2024-08-07T08:46:30Z")

</div>

Ideally share it within a new issue in [GitHub - dadoonet/fscrawler: Elasticsearch File System Crawler (FS Crawler)](https://github.com/dadoonet/fscrawler). So we can track this 😊

---

<div class="post-metadata">

**Author:** ![saifstech26](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/saifstech26/32/126560_2.png) [@saifstech26](https://discuss.elastic.co/u/saifstech26)\
**Post date:** [August 7, 2024, 9:07am UTC](https://discuss.elastic.co/t/pdf-parsing-issues-when-using-fs-crawler/364488/5 "2024-08-07T09:07:28Z")

</div>

Surely will follow the suggestion. 😀
