# Unable to extract PDF content

**URL:** <https://discuss.elastic.co/t/unable-to-extract-pdf-content/354481>\
**Category:** Elasticsearch\
**Created:** [February 29, 2024, 9:15pm UTC](https://discuss.elastic.co/t/unable-to-extract-pdf-content/354481 "2024-02-29T21:15:04Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![NitzaAg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nitzaag/32/132245_2.png) [@NitzaAg](https://discuss.elastic.co/u/NitzaAg)\
**Post date:** [February 29, 2024, 9:15pm UTC](https://discuss.elastic.co/t/unable-to-extract-pdf-content/354481/1 "2024-02-29T21:15:04Z")

</div>

I'm trying to extract text from a pdf, but I get the following:

```auto
Unable to extract PDF content -> Unable to end a page -> I regret that I couldn't find an OCR parser to handle image/ocr-png.Please set the OCR_STRATEGY to NO_OCR or configure yourOCR parser correctly

```

My configuration looks like this:

```auto
---
name: "ocr_docs"
fs:
  url: "C:\\Users\\Documents\\docs2ocr"
  update_rate: "15m"
  excludes:
  - "*\\~*"
  json_support: false
  filename_as_id: false
  add_filesize: true
  remove_deleted: true
  add_as_inner_object: false
  store_source: false
  index_content: true
  attributes_support: true
  raw_metadata: true
  xml_support: false
  index_folders: true
  lang_detect: false
  continue_on_error: false
  ocr:
    language: "eng"
    enabled: true
    pdf_strategy: "ocr_and_text"
    path: "C:\\Program Files\\https%3a%2f%2fmirrors.163.com%2fcygwin%2f\\x86_64\\release\\tesseract-ocr\\tesseract-ocr-5.3.3-1\\usr\\bin"
    data_path: "C:\\Program Files\\https%3a%2f%2fmirrors.163.com%2fcygwin%2f\\x86_64\\release\\tesseract-ocr\\tesseract-ocr-5.3.3-1\\usr\\share\\tessdata"
  follow_symlinks: false
elasticsearch:
  nodes:
  - cloud_id: "cloud_id"
  username: "username"
  password: "password"
  bulk_size: 100
  flush_interval: "5s"
  byte_size: "10mb"

```

I'm using fscrawler-distribution-2.10-20231023.160816-291

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 29, 2024, 11:10pm UTC](https://discuss.elastic.co/t/unable-to-extract-pdf-content/354481/2 "2024-02-29T23:10:06Z")

</div>

Welcome!

I think you should find a more recent build than "20231023".  
The latest I can see is [fscrawler-distribution-2.10-20240213.145447-315.zip](https://s01.oss.sonatype.org/content/repositories/snapshots/fr/pilato/elasticsearch/crawler/fscrawler-distribution/2.10-SNAPSHOT/fscrawler-distribution-2.10-20240213.145447-315.zip).

But I don't think that will fix your problem. Could you share your PDF document so I can try it out? If you can't share it publicly, you can send me a private message.

---

<div class="post-metadata">

**Author:** ![NitzaAg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nitzaag/32/132245_2.png) [@NitzaAg](https://discuss.elastic.co/u/NitzaAg)\
**Post date:** [March 1, 2024, 4:43pm UTC](https://discuss.elastic.co/t/unable-to-extract-pdf-content/354481/3 "2024-03-01T16:43:02Z")

</div>

Thank you! I sent you a dm with the pdf doc. Please let me know if more information is needed

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 3, 2024, 12:33pm UTC](https://discuss.elastic.co/t/unable-to-extract-pdf-content/354481/4 "2024-03-03T12:33:49Z")

</div>

So I tried your document and was able to get its content. Something like ` ***18 de septiembre de 2013*** ` (skipping the rest)...

May be this does not work on windows or that was caused by an old build of FSCrawler?

Could you try with the most recent build?  
If this does not work, could you try with the FSCrawler Docker images?

---

<div class="post-metadata">

**Author:** ![NitzaAg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nitzaag/32/132245_2.png) [@NitzaAg](https://discuss.elastic.co/u/NitzaAg)\
**Post date:** [March 12, 2024, 10:41pm UTC](https://discuss.elastic.co/t/unable-to-extract-pdf-content/354481/5 "2024-03-12T22:41:40Z")

</div>

> [@dadoonet](#):
>
> cument and was able to get its content. Something like ` ***18 de septiembre de 2013*** ` (skipping the rest)...
> 
> May be this does not work on windows or that was caused by an old build of FSCrawler?
> 
> Could you try with the most recent build?  
> If this does not work, could you try with the FSCrawler Docker images?

I tried that build and didn't work. However, it worked after using the FSCrawler Docker images. Thanks for the help 😀

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 14, 2024, 9:17pm UTC](https://discuss.elastic.co/t/unable-to-extract-pdf-content/354481/8 "2024-04-14T21:17:52Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
