# Indexing large pdf document

**URL:** <https://discuss.elastic.co/t/indexing-large-pdf-document/22917>\
**Category:** Elasticsearch\
**Created:** [March 26, 2015, 9:51am UTC](https://discuss.elastic.co/t/indexing-large-pdf-document/22917 "2015-03-26T09:51:33Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jakko\_Sikkar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jakko_sikkar/32/738_2.png) [@Jakko\_Sikkar](https://discuss.elastic.co/u/Jakko_Sikkar)\
**Post date:** [March 26, 2015, 9:51am UTC](https://discuss.elastic.co/t/indexing-large-pdf-document/22917/1 "2015-03-26T09:51:33Z")

</div>

Hi,

I'm trying to index big document with ES and Mapper Attachment plugin  
([https://github.com/elastic/elasticsearch-mapper-attachments](https://github.com/elastic/elasticsearch-mapper-attachments)). Document has  
719 pages, but after indexing I can search phrases only up to page 33. When  
I index a document I'm base64 encoding the file contents and file get  
successfully added to the index. Is there some limits of the size of the  
file?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/04dd35e4-1caf-4a30-8f24-13cf47907067%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/04dd35e4-1caf-4a30-8f24-13cf47907067%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 26, 2015, 10:51am UTC](https://discuss.elastic.co/t/indexing-large-pdf-document/22917/2 "2015-03-26T10:51:08Z")

</div>

There is a limit of the number of extracted characters.

See [GitHub - elastic/elasticsearch-mapper-attachments: Mapper Attachments Type plugin for Elasticsearch](https://github.com/elastic/elasticsearch-mapper-attachments#indexed-characters) [https://github.com/elastic/elasticsearch-mapper-attachments#indexed-characters](https://github.com/elastic/elasticsearch-mapper-attachments#indexed-characters)

--  
David Pilato - Developer | Evangelist

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co/)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

@dadoonet [https://twitter.com/dadoonet](https://twitter.com/dadoonet) | @elasticsearchfr [https://twitter.com/elasticsearchfr](https://twitter.com/elasticsearchfr) | @scrutmydocs [https://twitter.com/scrutmydocs](https://twitter.com/scrutmydocs)

> Le 26 mars 2015 à 10:51, Jakko Sikkar [jakko.sikkar@gmail.com](mailto:jakko.sikkar@gmail.com) a écrit :
> 
> Hi,
> 
> I'm trying to index big document with ES and Mapper Attachment plugin ([GitHub - elastic/elasticsearch-mapper-attachments: Mapper Attachments Type plugin for Elasticsearch](https://github.com/elastic/elasticsearch-mapper-attachments)). Document has 719 pages, but after indexing I can search phrases only up to page 33. When I index a document I'm base64 encoding the file contents and file get successfully added to the index. Is there some limits of the size of the file?
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com) [mailto:elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/04dd35e4-1caf-4a30-8f24-13cf47907067%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/04dd35e4-1caf-4a30-8f24-13cf47907067%40googlegroups.com) [https://groups.google.com/d/msgid/elasticsearch/04dd35e4-1caf-4a30-8f24-13cf47907067%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/04dd35e4-1caf-4a30-8f24-13cf47907067%40googlegroups.com?utm_medium=email&utm_source=footer).  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout) [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/68560613-23C5-4398-A7F0-FEFBACF83DEA%40pilato.fr](https://groups.google.com/d/msgid/elasticsearch/68560613-23C5-4398-A7F0-FEFBACF83DEA%40pilato.fr).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Jakko\_Sikkar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jakko_sikkar/32/738_2.png) [@Jakko\_Sikkar](https://discuss.elastic.co/u/Jakko_Sikkar)\
**Post date:** [March 26, 2015, 11:40am UTC](https://discuss.elastic.co/t/indexing-large-pdf-document/22917/3 "2015-03-26T11:40:37Z")

</div>

Thank you very much for pointing that out, I read documentation but skipped  
that part somehow 🙂

neljapäev, 26. märts 2015 12:51.50 UTC+2 kirjutas David Pilato:

> There is a limit of the number of extracted characters.
> 
> See  
> [GitHub - elastic/elasticsearch-mapper-attachments: Mapper Attachments Type plugin for Elasticsearch](https://github.com/elastic/elasticsearch-mapper-attachments#indexed-characters)
> 
> --  
> _David Pilato_ - Developer | Evangelist  
> _[elastic.co](http://elastic.co) [http://elastic.co](http://elastic.co)_  
> @dadoonet [https://twitter.com/dadoonet](https://twitter.com/dadoonet) | @elasticsearchfr  
> [https://twitter.com/elasticsearchfr](https://twitter.com/elasticsearchfr) | @scrutmydocs  
> [https://twitter.com/scrutmydocs](https://twitter.com/scrutmydocs)
> 
> Le 26 mars 2015 à 10:51, Jakko Sikkar \<[jakko....@gmail.com](mailto:jakko....@gmail.com) \<javascript:\>\>  
> a écrit :
> 
> Hi,
> 
> I'm trying to index big document with ES and Mapper Attachment plugin (  
> [GitHub - elastic/elasticsearch-mapper-attachments: Mapper Attachments Type plugin for Elasticsearch](https://github.com/elastic/elasticsearch-mapper-attachments)). Document  
> has 719 pages, but after indexing I can search phrases only up to page 33.  
> When I index a document I'm base64 encoding the file contents and file get  
> successfully added to the index. Is there some limits of the size of the  
> file?
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/04dd35e4-1caf-4a30-8f24-13cf47907067%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/04dd35e4-1caf-4a30-8f24-13cf47907067%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/04dd35e4-1caf-4a30-8f24-13cf47907067%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/04dd35e4-1caf-4a30-8f24-13cf47907067%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/ff655c88-1e8a-4703-935a-f0136deee442%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ff655c88-1e8a-4703-935a-f0136deee442%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Meenal.Luktuke](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/meenal.luktuke/32/7795_2.png) [@Meenal.Luktuke](https://discuss.elastic.co/u/Meenal.Luktuke)\
**Post date:** [February 10, 2016, 9:26am UTC](https://discuss.elastic.co/t/indexing-large-pdf-document/22917/4 "2016-02-10T09:26:25Z")

</div>

Hi,

Can you please share your ES configuration? I want to index many large PDFs, but not sure how to include them together. Do I have to write curl for each separately? Also, need help with this syntax --

PUT /test-mapping/person/1  
{  
"my\_attachment" : {  
"\_name" : "/home/ubuntu/test.pdf",  
"\_language" : "en",  
**"\_content" : "... base64 encoded attachment ..."** ---\> Do I have to write content even if I know I want to index the complete file? How to specify the location to read from?  
}  
}

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 10, 2016, 9:33am UTC](https://discuss.elastic.co/t/indexing-large-pdf-document/22917/5 "2016-02-10T09:33:17Z")

</div>

This thread is 10 months old. I think it would be better if you started a new thread for this.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 10, 2016, 10:04am UTC](https://discuss.elastic.co/t/indexing-large-pdf-document/22917/6 "2016-02-10T10:04:37Z")

</div>

> Do I have to write content even if I know I want to index the complete file? How to specify the location to read from?

You can't specify the location to read from with the mapper attachments plugin.

---

<div class="post-metadata">

**Author:** ![hubert](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hubert/32/9848_2.png) [@hubert](https://discuss.elastic.co/u/hubert)\
**Post date:** [June 10, 2016, 8:53am UTC](https://discuss.elastic.co/t/indexing-large-pdf-document/22917/7 "2016-06-10T08:53:33Z")

</div>

Hello Sir?  
I am new in elasticsearch but I like the way it is a power ful tool  
Can you help me please see the documentation we have been using ti index even one pdf file?  
Best regards!

---

<div class="post-metadata">

**Author:** ![Meenal.Luktuke](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/meenal.luktuke/32/7795_2.png) [@Meenal.Luktuke](https://discuss.elastic.co/u/Meenal.Luktuke)\
**Post date:** [June 23, 2016, 12:37pm UTC](https://discuss.elastic.co/t/indexing-large-pdf-document/22917/8 "2016-06-23T12:37:52Z")

</div>

Hi,

You can definitely index PDF document provided they are encoded in base-64

[https://discuss.elastic.co/t/logstash-parsing-for-rich-text-documents/40640/3](https://discuss.elastic.co/t/logstash-parsing-for-rich-text-documents/40640/3)

---

<div class="post-metadata">

**Author:** ![hubert](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hubert/32/9848_2.png) [@hubert](https://discuss.elastic.co/u/hubert)\
**Post date:** [June 24, 2016, 12:13pm UTC](https://discuss.elastic.co/t/indexing-large-pdf-document/22917/9 "2016-06-24T12:13:36Z")

</div>

Hello,  
Thank you so much I have tried it works!  
Cheers!!!!!!!!!!!!!!!!!!!!!!!!!!

---

<div class="post-metadata">

**Author:** ![hubert](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hubert/32/9848_2.png) [@hubert](https://discuss.elastic.co/u/hubert)\
**Post date:** [August 18, 2016, 10:21am UTC](https://discuss.elastic.co/t/indexing-large-pdf-document/22917/10 "2016-08-18T10:21:22Z")

</div>

Hello again?

I met a problem with my elasticsearch in production it is stopping after two days and I am not able to figure out what I am missing.

 ![](https://us1.discourse-cdn.com/elastic/original/2X/c/c823e66d371f23ebe5017b8956c97ab00e63f250.png)  
Any idea would be welcomed!

Best regards

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:26pm UTC](https://discuss.elastic.co/t/indexing-large-pdf-document/22917/11 "2017-07-05T22:26:57Z")

</div>


