# FScrawler: perform OCR selectively only on PDF files that do not have text

**URL:** <https://discuss.elastic.co/t/fscrawler-perform-ocr-selectively-only-on-pdf-files-that-do-not-have-text/236016>\
**Category:** Elasticsearch\
**Created:** [June 6, 2020, 1:19am UTC](https://discuss.elastic.co/t/fscrawler-perform-ocr-selectively-only-on-pdf-files-that-do-not-have-text/236016 "2020-06-06T01:19:16Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![equj](https://avatars.discourse-cdn.com/v4/letter/e/65b543/32.png) [@equj](https://discuss.elastic.co/u/equj)\
**Post date:** [June 6, 2020, 1:19am UTC](https://discuss.elastic.co/t/fscrawler-perform-ocr-selectively-only-on-pdf-files-that-do-not-have-text/236016/1 "2020-06-06T01:19:17Z")

</div>

Hello,  
I'm using FScrawler (2.7) to load text from PDFs into Elasticsearch (7.6.x). Most of PDF files have text, but some of PDF files contain images of scanned text and need to be OCRed. Is there a way to configure FScrawler such as that it performs OCR only on PDF files that contain images of scanned text, but not on files that already have text?

So far I can configure it to either not to do OCR on any files (case 1) or to do it on all files (case 2). In the first case, FScrawler skips all files with images of scanned text, but loads all files with text very quickly. In the second case, it takes really long time because it OCRs all the files, including those that already have text.

Here is OCR options setting for FScrawler: [https://fscrawler.readthedocs.io/en/latest/user/ocr.html](https://fscrawler.readthedocs.io/en/latest/user/ocr.html)

Config for case 1:

```auto
name: "test"
fs:
  url: "/path/to/data/dir"
  ocr:
    enabled: false
    pdf_strategy: 'no_ocr'

```

Config of case 2:

```auto
name: "test"
fs:
  url: "/path/to/data/dir"
  ocr:
    enabled: true
    pdf_strategy: 'ocr_and_text'`

```

P.S. I can sort them OCRed and non-OCRed files using other means and have two separate FScrawler jobs for each pile of PDF files, but before I do this, I want to check if there is an easier way to use FScrawler native features.

Thank you!

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [June 6, 2020, 8:32am UTC](https://discuss.elastic.co/t/fscrawler-perform-ocr-selectively-only-on-pdf-files-that-do-not-have-text/236016/2 "2020-06-06T08:32:13Z")

</div>

Welcome!

I don't think there's a way to do that in Tika (so in FSCrawler).  
How do you know which files should be OCRed and the others? Is there a technical way to implement that?

---

<div class="post-metadata">

**Author:** ![equj](https://avatars.discourse-cdn.com/v4/letter/e/65b543/32.png) [@equj](https://discuss.elastic.co/u/equj)\
**Post date:** [June 6, 2020, 2:50pm UTC](https://discuss.elastic.co/t/fscrawler-perform-ocr-selectively-only-on-pdf-files-that-do-not-have-text/236016/3 "2020-06-06T14:50:26Z")

</div>

Thank you for your response, David!  
This clarifies things perfectly!

My thinking was to do the following operation for every file:

1. Extract text from a file and check how many characters in the text and how many pages the file has
2. If there are fewer than X characters per page, perform OCR on that file.

I was uploading PDFs to Elasticsearch using Python before and this is how I did it. I'd like to use FSCrawler from now one - let me explore if I can solve it by having two FSCrawler jobs and sorting files on whether they need to be OCRed.

P.S. Thank you so much for your work on FSCrawler!

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [June 6, 2020, 3:39pm UTC](https://discuss.elastic.co/t/fscrawler-perform-ocr-selectively-only-on-pdf-files-that-do-not-have-text/236016/4 "2020-06-06T15:39:20Z")

</div>

Smart idea. Could you open an issue so I could try to come with a solution?

Adding an option in ocr like `"run_above": 500`.

---

<div class="post-metadata">

**Author:** ![tallison](https://avatars.discourse-cdn.com/v4/letter/t/258eb7/32.png) [@tallison](https://discuss.elastic.co/u/tallison)\
**Post date:** [June 8, 2020, 1:23pm UTC](https://discuss.elastic.co/t/fscrawler-perform-ocr-selectively-only-on-pdf-files-that-do-not-have-text/236016/5 "2020-06-08T13:23:10Z")

</div>

We have a rudimentary "auto" mode for OCR'ing of PDFs. I just updated our wiki to include this -- [https://cwiki.apache.org/confluence/pages/viewpage.action?pageId=109454066](https://cwiki.apache.org/confluence/pages/viewpage.action?pageId=109454066). See "Option 2: Configuring OCR on Rendered Pages".

This will trigger OCR on pages if \< 10 characters were extracted or more than 10 characters lack unicode mappings.

If there are better heuristics we should add, let us know!

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [June 18, 2020, 1:14pm UTC](https://discuss.elastic.co/t/fscrawler-perform-ocr-selectively-only-on-pdf-files-that-do-not-have-text/236016/6 "2020-06-18T13:14:36Z")

</div>

This is amazing! Thanks a ton @tallison.

PR is on its way here: [https://github.com/dadoonet/fscrawler/pull/965](https://github.com/dadoonet/fscrawler/pull/965)

@equj you can already use the `auto` option by setting:

```auto
name: "test"
fs:
  url: "/path/to/data/dir"
  ocr:
    enabled: true
    pdf_strategy: 'auto'

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 16, 2020, 1:14pm UTC](https://discuss.elastic.co/t/fscrawler-perform-ocr-selectively-only-on-pdf-files-that-do-not-have-text/236016/7 "2020-07-16T13:14:38Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
