# How to use OCR in Elasticsearch ingest attachment plugin?

**URL:** <https://discuss.elastic.co/t/how-to-use-ocr-in-elasticsearch-ingest-attachment-plugin/263007>\
**Category:** Elasticsearch\
**Tags:** ingest-pipeline\
**Created:** [February 2, 2021, 3:36pm UTC](https://discuss.elastic.co/t/how-to-use-ocr-in-elasticsearch-ingest-attachment-plugin/263007 "2021-02-02T15:36:54Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![OB1290](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ob1290/32/83316_2.png) [@OB1290](https://discuss.elastic.co/u/OB1290)\
**Post date:** [February 2, 2021, 3:36pm UTC](https://discuss.elastic.co/t/how-to-use-ocr-in-elasticsearch-ingest-attachment-plugin/263007/1 "2021-02-02T15:36:54Z")

</div>

I was able to make the plugin work with PDFs that contain searchable text. When I give it a PDF or PNG with non searchable text, it fails to extract the text from the binary data.  
Unfortunately, I couldn't find anything useful in the documentation and forums.

Here are the steps I followed to get the plugin up and running ( PHP ):

Executed this command in the bin directory of Elasticsearch

```auto
    elasticsearch-plugin install ingest-attachment

```

Next I created the pipeline

```auto
            $client = ClientBuilder::create()->build();
            $params = [
                'id' => 'attachment',
                'body' => [
                    'description' => 'Extract attachment information',
                    'processors' => [
                        [
                            'attachment' => [
                                'field' => 'data'
                            ]
                        ]
                    ]
                ],
            ];
            return $client->ingest()->putPipeline($params);

```

Then I got the file, encoded it in base64 and attached it to an ES document

```auto
            $client = ClientBuilder::create()->build();
            $myfiles = array_diff(scandir('pdf_files'), array('.', '..'));
            $params = [
                'index' => 'candidates',
                'type' => '_doc',
                'id' => 'e9AuBXcBC0zZvKKfMaH9',
                'pipeline' => 'attachment',
                'body' => [
                    'data' => base64_encode(file_get_contents('./pdf_files/'.$myfiles[2]))
                ]                
            ];
            $response = $client->index($params);

```

When I fetch the document through the kibana console, I get this response

```auto
    "attachment" : {
          "content_type" : "application/pdf",
          "language" : "lt",
          "content" : "",
          "content_length" : 2
        }

```

As you can see the content property is empty. Any ideas on how to make it work ?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 2, 2021, 4:05pm UTC](https://discuss.elastic.co/t/how-to-use-ocr-in-elasticsearch-ingest-attachment-plugin/263007/3 "2021-02-02T16:05:56Z")

</div>

Oh sorry. OCR is not supported by ingest attachment plugin.

---

<div class="post-metadata">

**Author:** ![OB1290](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ob1290/32/83316_2.png) [@OB1290](https://discuss.elastic.co/u/OB1290)\
**Post date:** [February 2, 2021, 4:22pm UTC](https://discuss.elastic.co/t/how-to-use-ocr-in-elasticsearch-ingest-attachment-plugin/263007/4 "2021-02-02T16:22:46Z")

</div>

Oh, ok thanks. Do you have a suggestion on the best way to implement OCR with Elasticsearch ?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 2, 2021, 7:41pm UTC](https://discuss.elastic.co/t/how-to-use-ocr-in-elasticsearch-ingest-attachment-plugin/263007/5 "2021-02-02T19:41:00Z")

</div>

You can use [FSCrawler](https://fscrawler.readthedocs.io). There's [a tutorial](https://fscrawler.readthedocs.io/en/latest/user/tutorial.html) to help you getting started.

---

<div class="post-metadata">

**Author:** ![OB1290](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ob1290/32/83316_2.png) [@OB1290](https://discuss.elastic.co/u/OB1290)\
**Post date:** [February 3, 2021, 3:49pm UTC](https://discuss.elastic.co/t/how-to-use-ocr-in-elasticsearch-ingest-attachment-plugin/263007/6 "2021-02-03T15:49:24Z")

</div>

FSCrawler worked like a charm, it pushed every file in the specified index.  
The problem is that I wanted to push only relevant files in a specific document within the index. I couldn't find anything in the documentation that would help me do that. Any suggestion ?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 3, 2021, 9:01pm UTC](https://discuss.elastic.co/t/how-to-use-ocr-in-elasticsearch-ingest-attachment-plugin/263007/7 "2021-02-03T21:01:07Z")

</div>

Not sure I understand the question.

But did you look at the REST service of FSCrawler? Might be what you want.

---

<div class="post-metadata">

**Author:** ![russellmenezes](https://avatars.discourse-cdn.com/v4/letter/r/f4b2a3/32.png) [@russellmenezes](https://discuss.elastic.co/u/russellmenezes)\
**Post date:** [February 4, 2021, 2:24am UTC](https://discuss.elastic.co/t/how-to-use-ocr-in-elasticsearch-ingest-attachment-plugin/263007/8 "2021-02-04T02:24:40Z")

</div>

It's easier to setup a custom pipeline in Python. I have used OCRmyPDF to process the pdfs. Then use Tika to extract the text. This can then be bulk indexed into ES

---

<div class="post-metadata">

**Author:** ![OB1290](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ob1290/32/83316_2.png) [@OB1290](https://discuss.elastic.co/u/OB1290)\
**Post date:** [February 4, 2021, 8:35am UTC](https://discuss.elastic.co/t/how-to-use-ocr-in-elasticsearch-ingest-attachment-plugin/263007/9 "2021-02-04T08:35:48Z")

</div>

> [@dadoonet](#):
>
> Not sure I understand the question.

I have a list of candidates stored in an ES index and would like to index each PDF file to a specific candidate. How do I go about doing this ? Can it be done with FSCrawler at all ?

> [@dadoonet](#):
>
> But did you look at the REST service of FSCrawler? Might be what you want.

That might be it. But it's unclear to me how I would direct FSCrawler to update specific ES documents using the REST service.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 4, 2021, 1:24pm UTC](https://discuss.elastic.co/t/how-to-use-ocr-in-elasticsearch-ingest-attachment-plugin/263007/10 "2021-02-04T13:24:40Z")

</div>

Not really. Actually not directly.

What you could do is the suggestion @russellmenezes made.  
Another option is to run FSCrawler as a REST Service and use its simulate API: [REST service — FSCrawler 2.7-SNAPSHOT documentation](https://fscrawler.readthedocs.io/en/latest/admin/fs/rest.html#simulate-upload)

That way you can get the JSON back and update the candidate document with that content.

Or you can also think of it in another way:

- Create one index for the candidates
- Create one index for the resumes

Update the candidate document with just a link to the resume.

Highly depends on the use case: how you want to search? What do you want to search for? (I assume candidates)...

---

<div class="post-metadata">

**Author:** ![OB1290](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ob1290/32/83316_2.png) [@OB1290](https://discuss.elastic.co/u/OB1290)\
**Post date:** [February 4, 2021, 1:59pm UTC](https://discuss.elastic.co/t/how-to-use-ocr-in-elasticsearch-ingest-attachment-plugin/263007/11 "2021-02-04T13:59:57Z")

</div>

> [@dadoonet](#):
>
> Another option is to run FSCrawler as a REST Service and use its simulate API: [REST service — FSCrawler 2.7-SNAPSHOT documentation](https://fscrawler.readthedocs.io/en/latest/admin/fs/rest.html#simulate-upload)

I liked the simulate API approach, I got it to work via curl :

```auto
curl -F "file=@pdf_files/non-searchable-text.pdf" "http://127.0.0.1:8080/fscrawler/_upload?debug=true&simulate=true"

```

Which gave me this response :

```auto
{
  "ok": true,
  "filename": "non-searchable-text.pdf",
  "url": "http://127.0.0.1:9200/resumes/_doc/f39614d4716aed76167498ac4945ed7",
  "doc": {
    "content": "\n \n\nTABLE OF CONTENTS\n\n \n\nIntroduction 1\nChapter 1: The ABC of Programming 11\nChapter 2: Basic JavaScript Instructions 53\nChapter 3: Functions, Methods & Objects 85\n(i aY-] 0) (=) ae Sam DY -101 |} (0) aioe\" Mole) 0-) 145\nChapter 5: Document Object Model 183\nChapter 6: Events 243\nChapter 7: jQuery 293\nChapter 8: Ajax & JSON 367\nChapter 9: APIs 409\nChapter 10: Error Handling & Debugging 449\nChapter 11: Content Panels 487\nChapter 12: Filtering, Searching & Sorting 527\nChapter 13: Form Enhancement & Validation 567\nTare toy 623\n\n \n\nTry out & download the code in this book\nwww. javascriptbook.com\n\n \n\n \n\n \n\n\n",
    "meta": {
      "date": "2021-02-02T11:44:56.000+00:00",
      "format": "application/pdf; version=1.6"
    },
    "file": {
      "extension": "pdf",
      "content_type": "application/pdf",
      "indexing_date": "2021-02-04T13:49:29.353+00:00",
      "filename": "non-searchable-text.pdf"
    },
    "path": {
      "virtual": "non-searchable-text.pdf",
      "real": "non-searchable-text.pdf"
    }
  }
}

```

I tried making it work through PHP like this :

```auto
$curl = curl_init();

curl_setopt($curl, CURLOPT_URL, 'http://127.0.0.1:8080/fscrawler/_upload?debug=true&simulate=true');
curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1);
curl_setopt($curl, CURLOPT_POST, 1);
$args['file'] = '@/pdf_files/non-searchable-text.pdf';
curl_setopt($curl, CURLOPT_POSTFIELDS, $args);

$result = curl_exec($curl);
if (curl_errno($curl)) {
    echo 'Error:' . curl_error($curl);
} else {
    echo '<pre> response : ', print_r($result, true) ,'</pre>';
}
curl_close($curl);

```

But this was the response :

![curl response](https://us1.discourse-cdn.com/elastic/original/3X/a/d/adde5eca4df426194f28ee555bfb8185ef2fd916.png)

Any idea what I'm doing wrong ?

---

<div class="post-metadata">

**Author:** ![OB1290](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ob1290/32/83316_2.png) [@OB1290](https://discuss.elastic.co/u/OB1290)\
**Post date:** [February 4, 2021, 4:08pm UTC](https://discuss.elastic.co/t/how-to-use-ocr-in-elasticsearch-ingest-attachment-plugin/263007/12 "2021-02-04T16:08:00Z")

</div>

The problem is just in the request, it's working now. Thank you for your help @dadoonet , FSCrawler is an amazing solution to extracting and indexing content from files in an optimal way to Elasticsearch. I don't know why there is no official ES support for it yet but it certainly deserves it.

Here is the PHP curl request for the curious :

```auto
$curl = curl_init();
curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1);

$mimetype = mime_content_type('FULLPATH GOES HERE');
$curlFile = new CURLFile('FULLPATH GOES HERE', $mimetype); 

$postFields = array('id' => 'PDFTESTESEARCH', 'file' => $curlFile);

curl_setopt($curl, CURLOPT_URL, 'http://127.0.0.1:8080/fscrawler/_upload?debug=true&simulate=true');
curl_setopt($curl, CURLOPT_POST, 1);
curl_setopt($curl, CURLOPT_POSTFIELDS, $postFields);

$result = curl_exec($curl);

if (curl_errno($curl)) {
    echo 'Error:' . curl_error($curl);
} else {
    echo '<pre> response : ', print_r( json_decode($result), true) ,'</pre>';
    var_dump(json_decode($result));
}
curl_close($curl);

```

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 4, 2021, 4:27pm UTC](https://discuss.elastic.co/t/how-to-use-ocr-in-elasticsearch-ingest-attachment-plugin/263007/13 "2021-02-04T16:27:31Z")

</div>

Thanks for giving a closure, sharing your solution and the kind words. 😉🤗

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 4, 2021, 4:28pm UTC](https://discuss.elastic.co/t/how-to-use-ocr-in-elasticsearch-ingest-attachment-plugin/263007/14 "2021-03-04T16:28:07Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
