# Elastic Search injest attachment cannot index multiple pdf

**URL:** <https://discuss.elastic.co/t/elastic-search-injest-attachment-cannot-index-multiple-pdf/299234>\
**Category:** Elasticsearch\
**Created:** [March 9, 2022, 3:19pm UTC](https://discuss.elastic.co/t/elastic-search-injest-attachment-cannot-index-multiple-pdf/299234 "2022-03-09T15:19:08Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Leo\_Baby\_Jacob](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leo_baby_jacob/32/87425_2.png) [@Leo\_Baby\_Jacob](https://discuss.elastic.co/u/Leo_Baby_Jacob)\
**Post date:** [March 9, 2022, 3:19pm UTC](https://discuss.elastic.co/t/elastic-search-injest-attachment-cannot-index-multiple-pdf/299234/1 "2022-03-09T15:19:08Z")

</div>

I have an mysql table with multiple pdfs and an associated item\_name for

```auto
title pdf_s3_link
    -------- --------------
    Harry Potter	linktopdfons3
    Batman s3_link_to_pdfons3

```

I am trying to injest these data into my Elasticsearch index so that if there's a match I need to display the title. I am trying to use injest api but I dont know how to run this automatically. (I am doing a POC so this is not on escloud yet but in the future the whole system, injestion should be serverless + es cloud).

I came across fscrawler but even with that I cannot download all those pdf to a local directory. Whats the best way out of this ?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [March 9, 2022, 9:48pm UTC](https://discuss.elastic.co/t/elastic-search-injest-attachment-cannot-index-multiple-pdf/299234/2 "2022-03-09T21:48:55Z")

</div>

Are the PDFs on S3 or in the database as blobs?

---

<div class="post-metadata">

**Author:** ![Leo\_Baby\_Jacob](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leo_baby_jacob/32/87425_2.png) [@Leo\_Baby\_Jacob](https://discuss.elastic.co/u/Leo_Baby_Jacob)\
**Post date:** [March 9, 2022, 9:59pm UTC](https://discuss.elastic.co/t/elastic-search-injest-attachment-cannot-index-multiple-pdf/299234/3 "2022-03-09T21:59:46Z")

</div>

free to access pdf links

---

<div class="post-metadata">

**Author:** ![Leo\_Baby\_Jacob](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leo_baby_jacob/32/87425_2.png) [@Leo\_Baby\_Jacob](https://discuss.elastic.co/u/Leo_Baby_Jacob)\
**Post date:** [March 9, 2022, 10:01pm UTC](https://discuss.elastic.co/t/elastic-search-injest-attachment-cannot-index-multiple-pdf/299234/4 "2022-03-09T22:01:03Z")

</div>

I wrote a python script using pypdf2 which I could run as a cron job, but its slow and I am not sure if there is a better way around it. fscrawler does not allow links

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [March 9, 2022, 10:01pm UTC](https://discuss.elastic.co/t/elastic-search-injest-attachment-cannot-index-multiple-pdf/299234/5 "2022-03-09T22:01:57Z")

</div>

There's nothing in the Elastic Stack that would do this for you unfortunately.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 6, 2022, 10:02pm UTC](https://discuss.elastic.co/t/elastic-search-injest-attachment-cannot-index-multiple-pdf/299234/6 "2022-04-06T22:02:34Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
