# Problem when using Elasticsearch and Tesseract-OCR

**URL:** <https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770>\
**Category:** Elasticsearch\
**Created:** [June 26, 2020, 3:10am UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770 "2020-06-26T03:10:30Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![hanguyen.uet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hanguyen.uet/32/71151_2.png) [@hanguyen.uet](https://discuss.elastic.co/u/hanguyen.uet)\
**Post date:** [June 26, 2020, 3:10am UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/1 "2020-06-26T03:10:30Z")

</div>

Hi everybody!  
I have a directory containing text data at /opt/data/. When a new file is uploaded to that directory. Is there a way to automatically use tesseract-ocr to convert to another language, then use elasticsearch to automatically index to search for content in it?  
I'm using Ubuntu Server 16.04, Elasticsearch version 7.7.1and Tesseract-OCR 3.04.01.  
Look forward to your help.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [June 26, 2020, 5:57am UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/2 "2020-06-26T05:57:44Z")

</div>

You can use [FSCrawler](https://fscrawler.readthedocs.io). There's [a tutorial](https://fscrawler.readthedocs.io/en/latest/user/tutorial.html) to help you getting started.

---

<div class="post-metadata">

**Author:** ![hanguyen.uet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hanguyen.uet/32/71151_2.png) [@hanguyen.uet](https://discuss.elastic.co/u/hanguyen.uet)\
**Post date:** [July 14, 2020, 8:09am UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/3 "2020-07-14T08:09:45Z")

</div>

Thanks for the suggestion.  
When I use fscrawler, I index on elasticsearch. The files are all stored in tmp/es directory. But every time there is a new file in the tmp/es directory, I see it is not automatically updated, but I have to run the "bin/fscrawler job-name --restart" command again. I find this really inconvenient. Is there any way to run fscrawler forever?  
Best regards!

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 14, 2020, 5:59pm UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/4 "2020-07-14T17:59:11Z")

</div>

It's probably because you're moving a file to dir instead of copying it. So it has an old creation date and is not picked up by FSCrawler.

You can activate the debug mode to see why it's ignored.

Otherwise on Linux, you can also `touch` the file.

---

<div class="post-metadata">

**Author:** ![hanguyen.uet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hanguyen.uet/32/71151_2.png) [@hanguyen.uet](https://discuss.elastic.co/u/hanguyen.uet)\
**Post date:** [July 15, 2020, 3:42am UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/5 "2020-07-15T03:42:34Z")

</div>

I have applied it to my project. I have the directory structure as shown and install the \_setting.yaml file as shown below. And when I ran the command "bin / fscrawler job-name --restart", I saw that fscrawler started indexing but it happened very slowly and there were some files missing. I have tried to run the above command again and wait for a long time but still do not see the change in the number of files in elasticsearch.  
Look forward to the help. Thank you!

 ![form](https://us1.discourse-cdn.com/elastic/original/3X/2/c/2cfdf85750318bbd035b6350ef57f32a6f2a524d.png)

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 15, 2020, 6:45am UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/6 "2020-07-15T06:45:03Z")

</div>

Could you share the output of the following command:

```
GET /indexname/_search

```

How many files are you expecting?  
Run FSCrawler with `--debug` option and share the full logs.

Please format your code, logs or configuration files using `</>` icon as explained in [this guide](https://discuss.elastic.co/t/about-the-elasticsearch-category/21) and not the citation button. It will make your post more readable.

Or use markdown style like:

````
```
CODE
```

````

This is the icon to use if you are not using markdown format:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/7/e/7e6e239431ec2d71cbf1beef741f2e93e7cc762c.jpg)

If some outputs are too big, please share them on [gist.github.com](http://gist.github.com) and link them here.

---

<div class="post-metadata">

**Author:** ![hanguyen.uet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hanguyen.uet/32/71151_2.png) [@hanguyen.uet](https://discuss.elastic.co/u/hanguyen.uet)\
**Post date:** [July 15, 2020, 8:31am UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/7 "2020-07-15T08:31:41Z")

</div>

I realize that some xlsx files are not indexed. I see only one of all indexed xlsx files.  
GET /indexname/\_search

 ![index](https://us1.discourse-cdn.com/elastic/original/3X/e/6/e61e641ff8d78a9f1c060fd7f8dc11ac5094cd2f.png)

And file debug fscrawler: [https://gist.github.com/hanguyenuet96/b7413659993d434ca6e869a4ebbcaa17](https://gist.github.com/hanguyenuet96/b7413659993d434ca6e869a4ebbcaa17)

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 15, 2020, 8:47am UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/8 "2020-07-15T08:47:14Z")

</div>

Please don't post images of text as they are hard to read, may not display correctly for everyone, and are not searchable.

Instead, paste the text and format it with `</>` icon or pairs of triple backticks (```), and check the preview window to make sure it's properly formatted before posting it. This makes it more likely that your question will receive a useful answer.

There are 31 documents indexed here. How many did you add to the folder? What are the file which are not indexed? What's their names?

---

<div class="post-metadata">

**Author:** ![hanguyen.uet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hanguyen.uet/32/71151_2.png) [@hanguyen.uet](https://discuss.elastic.co/u/hanguyen.uet)\
**Post date:** [July 15, 2020, 9:34am UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/9 "2020-07-15T09:34:27Z")

</div>

There are a total of 36 documents, and has added 31 files to elasticsearch. The names of the files not indexed are:

1. 20\_8\_73734\_TestCase.xls

2. \_380\_filename (2).xls

3. 1\_44548\_73744\_lichcoquan.xls

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 15, 2020, 2:18pm UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/10 "2020-07-15T14:18:31Z")

</div>

I did not see `20_8_73734_TestCase.xls` in the logs. What is its full path?

---

<div class="post-metadata">

**Author:** ![hanguyen.uet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hanguyen.uet/32/71151_2.png) [@hanguyen.uet](https://discuss.elastic.co/u/hanguyen.uet)\
**Post date:** [July 16, 2020, 4:20am UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/11 "2020-07-16T04:20:15Z")

</div>

Thank you for helping me.  
The full path is: /opt/lampp/htdocs/selab/Contents/OfficialDispatch/2020/07/15/20\_8\_73734\_TestCase.xls

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 16, 2020, 7:53am UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/12 "2020-07-16T07:53:55Z")

</div>

In the logs, it sounds like the `07` dir is not available when the crawler ran.

I searched for `OfficialDispatch/2020/07` in the logs and it is not there. But `OfficialDispatch/2020/05` is in logs.

Could you `ls -l /opt/lampp/htdocs/selab/Contents/OfficialDispatch/2020/`?

Also share the fscrawler job settings please.

---

<div class="post-metadata">

**Author:** ![hanguyen.uet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hanguyen.uet/32/71151_2.png) [@hanguyen.uet](https://discuss.elastic.co/u/hanguyen.uet)\
**Post date:** [July 21, 2020, 10:56am UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/13 "2020-07-21T10:56:40Z")

</div>

I have checked and missing files, this is one of my shortcomings. Thank you for your help.  
By the way, may I ask, is there a way to run FScrawler forever?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 21, 2020, 1:48pm UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/14 "2020-07-21T13:48:47Z")

</div>

On windows, you can follow this maybe ? [https://fscrawler.readthedocs.io/en/latest/installation.html#running-as-a-service-on-windows](https://fscrawler.readthedocs.io/en/latest/installation.html#running-as-a-service-on-windows)

On Linux, i guess you need to create a service.

---

<div class="post-metadata">

**Author:** ![hanguyen.uet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hanguyen.uet/32/71151_2.png) [@hanguyen.uet](https://discuss.elastic.co/u/hanguyen.uet)\
**Post date:** [July 22, 2020, 3:48am UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/15 "2020-07-22T03:48:12Z")

</div>

Yeahh. Thank you so much.  
Best regards!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 19, 2020, 3:52am UTC](https://discuss.elastic.co/t/problem-when-using-elasticsearch-and-tesseract-ocr/238770/16 "2020-08-19T03:52:25Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
