# Achieving a books dot google dot com workflow with Elasticsearch as search engine, any indexer and tesseract OCR tool?

**URL:** <https://discuss.elastic.co/t/achieving-a-books-dot-google-dot-com-workflow-with-elasticsearch-as-search-engine-any-indexer-and-tesseract-ocr-tool/375380>\
**Category:** Elasticsearch\
**Created:** [March 4, 2025, 2:55pm UTC](https://discuss.elastic.co/t/achieving-a-books-dot-google-dot-com-workflow-with-elasticsearch-as-search-engine-any-indexer-and-tesseract-ocr-tool/375380 "2025-03-04T14:55:03Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![davecomputertips](https://avatars.discourse-cdn.com/v4/letter/d/e495f1/32.png) [@davecomputertips](https://discuss.elastic.co/u/davecomputertips)\
**Post date:** [March 4, 2025, 2:55pm UTC](https://discuss.elastic.co/t/achieving-a-books-dot-google-dot-com-workflow-with-elasticsearch-as-search-engine-any-indexer-and-tesseract-ocr-tool/375380/1 "2025-03-04T14:55:03Z")

</div>

Chatgpt and deepseek hallucinated on this so asking here:

- 10s of pdfs is downloaded.
- I want to search for a particular topic inside the content of pdf.
- I search it.
- Then I want to read that pdf going to that exact page where the "content" that I wanted was found.

Something similar to the search engine of [books.google.com](http://books.google.com). I researched a bit and found google uses tesseract to convert pdfs to text. Now, my concern is how to integerate this workflow with an indexer like Lucene and searching engine like Elasticsearch. I am absolutely new to all these three things. I'd want a starting point. And how complicated will this gonna be An general idea. I am not looking for step by step guide.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 4, 2025, 3:26pm UTC](https://discuss.elastic.co/t/achieving-a-books-dot-google-dot-com-workflow-with-elasticsearch-as-search-engine-any-indexer-and-tesseract-ocr-tool/375380/2 "2025-03-04T15:26:37Z")

</div>

You can use the [ingest attachment plugin](https://www.elastic.co/guide/en/elasticsearch/plugins/current/ingest-attachment.html).

There an example here: [https://www.elastic.co/guide/en/elasticsearch/plugins/current/using-ingest-attachment.html](https://www.elastic.co/guide/en/elasticsearch/plugins/current/using-ingest-attachment.html)

```auto
PUT _ingest/pipeline/attachment
{
  "description" : "Extract attachment information",
  "processors" : [
    {
      "attachment" : {
        "field" : "data"
      }
    }
  ]
}
PUT my_index/_doc/my_id?pipeline=attachment
{
  "data": "e1xydGYxXGFuc2kNCkxvcmVtIGlwc3VtIGRvbG9yIHNpdCBhbWV0DQpccGFyIH0="
}
GET my_index/_doc/my_id

```

The `data` field is basically the BASE64 representation of your binary file.

You can use [FSCrawler](https://fscrawler.readthedocs.io). There's [a tutorial](https://fscrawler.readthedocs.io/en/latest/user/tutorial.html) to help you getting started.

But none of those features will tell you exactly the page number where the text was found. At least, not yet with FSCrawler. This would require this to be implemented:

> <https://github.com/dadoonet/fscrawler/issues/767>
>
> When indexing large documents you may hit limits not only on the indexing part, …but also when doing searches. 
> 
> Splitting documents into one entry per page helps slice up large documents into bite-size chunks and help performance of indexing and searching in the documents.

And also Kibana now supports directly uploading PDF files. See this very nice blog post:

> **[Chat with your PDFs using Elastic Playground - Elasticsearch Labs](https://www.elastic.co/search-labs/blog/chat-with-pdf-elastic-playground)**
>
> Learn how to upload PDF files into Kibana and interact with them using Elastic Playground. This blog showcases a practical example of chatting with PDFs in Playground.
