# Parse PDF catalogs to extract products informations

**URL:** https://discuss.elastic.co/t/parse-pdf-catalogs-to-extract-products-informations/322145
**Category:** Elasticsearch
**Created:** [December 29, 2022, 10:33am UTC](https://discuss.elastic.co/t/parse-pdf-catalogs-to-extract-products-informations/322145 "2022-12-29T10:33:53Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![Valentin\_Hirson](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/valentin_hirson/32/60872_2.png) [@Valentin\_Hirson](https://discuss.elastic.co/u/Valentin_Hirson)
#### Post date: [December 29, 2022, 10:33am UTC](https://discuss.elastic.co/t/parse-pdf-catalogs-to-extract-products-informations/322145/1 "2022-12-29T10:33:53Z")

</div>

Hello,

I need to parse some PDF files to find the most relevant words and trying to find what PDF are talking about.

Those PDF are vendors catalog, I need for example to extract products to propose them in a web search engine (actually powered by Elasticsearch with entities parsed from database).

Some catalogs could also be EXCEL files and those files are not standardized.

I used Elasticsearch a long time ago and I am not a data scientist 🙃

My project uses Elasticsearch as standalone, with a PHP client :

curl -XGET '[http://localhost:9200](http://localhost:9200)'  
{  
"name" : "n1es6kx",  
"cluster\_name" : "elasticsearch",  
"cluster\_uuid" : "5ya3CLzuQYKWMtqLk4Libw",  
"version" : {  
"number" : "6.8.10",  
"build\_flavor" : "default",  
"build\_type" : "deb",  
"build\_hash" : "537cb22",  
"build\_date" : "2020-05-28T14:47:19.882936Z",  
"build\_snapshot" : false,  
"lucene\_version" : "7.7.3",  
"minimum\_wire\_compatibility\_version" : "5.6.0",  
"minimum\_index\_compatibility\_version" : "5.0.0"  
},  
"tagline" : "You Know, for Search"  
}

"friendsofsymfony/elastica-bundle": "~5.2",  
"ruflin/elastica": "~6.1"

I am currently reading documentation but any advices/recommandations would be useful.

Thank you.

---

<div class="post-metadata">

### Author: ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)
#### Post date: [December 29, 2022, 11:21am UTC](https://discuss.elastic.co/t/parse-pdf-catalogs-to-extract-products-informations/322145/2 "2022-12-29T11:21:45Z")

</div>

First, upgrade! At least to 7.17.8 but better to 8.5.3.

You can use the [ingest attachment plugin](https://www.elastic.co/guide/en/elasticsearch/plugins/current/ingest-attachment.html).

There an example here: [https://www.elastic.co/guide/en/elasticsearch/plugins/current/using-ingest-attachment.html](https://www.elastic.co/guide/en/elasticsearch/plugins/current/using-ingest-attachment.html)

```auto
PUT _ingest/pipeline/attachment
{
  "description" : "Extract attachment information",
  "processors" : [
    {
      "attachment" : {
        "field" : "data"
      }
    }
  ]
}
PUT my_index/_doc/my_id?pipeline=attachment
{
  "data": "e1xydGYxXGFuc2kNCkxvcmVtIGlwc3VtIGRvbG9yIHNpdCBhbWV0DQpccGFyIH0="
}
GET my_index/_doc/my_id

```

The `data` field is basically the BASE64 representation of your binary file.

You can use [FSCrawler](https://fscrawler.readthedocs.io). There's [a tutorial](https://fscrawler.readthedocs.io/en/latest/user/tutorial.html) to help you getting started.

---

<div class="post-metadata">

### Author: ![Valentin\_Hirson](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/valentin_hirson/32/60872_2.png) [@Valentin\_Hirson](https://discuss.elastic.co/u/Valentin_Hirson)
#### Post date: [December 29, 2022, 1:08pm UTC](https://discuss.elastic.co/t/parse-pdf-catalogs-to-extract-products-informations/322145/3 "2022-12-29T13:08:15Z")

</div>

Hello,

Thank you I will check ingest attachment plugin documentation.

Any advice how my search into PDF could be more relevant ? I thought maybe I could count words to find out which words are the most relevant but I don't think it'will be the best solution.

Maybe there is a kind of IA or something ? Or should I parse my PDF first with an other application ?

---

<div class="post-metadata">

### Author: ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)
#### Post date: [January 11, 2023, 11:22am UTC](https://discuss.elastic.co/t/parse-pdf-catalogs-to-extract-products-informations/322145/4 "2023-01-11T11:22:52Z")

</div>

> [@Valentin\_Hirson](#):
>
> Any advice how my search into PDF could be more relevant ? I thought maybe I could count words to find out which words are the most relevant but I don't think it'will be the best solution.

What do you mean? Do you have an example for this question?

Ideally open a new discussion about it because this one is marked as solved and your original question is not directly related to the new question IMO. 😉

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [February 8, 2023, 11:23am UTC](https://discuss.elastic.co/t/parse-pdf-catalogs-to-extract-products-informations/322145/5 "2023-02-08T11:23:47Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
