# Possible to Index PDFs by page?

**URL:** <https://discuss.elastic.co/t/possible-to-index-pdfs-by-page/8883>\
**Category:** Elasticsearch\
**Created:** [August 29, 2012, 5:13pm UTC](https://discuss.elastic.co/t/possible-to-index-pdfs-by-page/8883 "2012-08-29T17:13:40Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Meltemi](https://avatars.discourse-cdn.com/v4/letter/m/a87d85/32.png) [@Meltemi](https://discuss.elastic.co/u/Meltemi)\
**Post date:** [August 29, 2012, 5:13pm UTC](https://discuss.elastic.co/t/possible-to-index-pdfs-by-page/8883/1 "2012-08-29T17:13:40Z")

</div>

Can elasticsearch index an attachment (PDF specifically) so the  
parent/child relationship between the _document_ (PDF) and the _page_ are  
preserved?

Our requirement dictates that matches should initially return the title of  
the PDF where the match occurred. Then if user wants to drill down further  
that _only_ the actual page where the hit occurred (with highlighting)  
should be presented. From there user should be able to page forward (or  
back) to continue reading. We should _not_ return the entire 100+ page  
documents but _only_ individual pages from within each document. Anyone  
know how to do this with elasticsearch?

--

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [August 29, 2012, 5:41pm UTC](https://discuss.elastic.co/t/possible-to-index-pdfs-by-page/8883/2 "2012-08-29T17:41:51Z")

</div>

Have a look at this:

> <https://stackoverflow.com/questions/10854858/best-practices-for-searchable-archive-of-thousands-of-documents-pdf-and-or-xml>

clint

On Wed, Aug 29, 2012 at 7:13 PM, Meltemi [mdemetrios@gmail.com](mailto:mdemetrios@gmail.com) wrote:

> Can elasticsearch index an attachment (PDF specifically) so the  
> parent/child relationship between the _document_ (PDF) and the _page_ are  
> preserved?
> 
> Our requirement dictates that matches should initially return the title of  
> the PDF where the match occurred. Then if user wants to drill down further  
> that _only_ the actual page where the hit occurred (with highlighting)  
> should be presented. From there user should be able to page forward (or  
> back) to continue reading. We should _not_ return the entire 100+ page  
> documents but _only_ individual pages from within each document. Anyone  
> know how to do this with elasticsearch?
> 
> --

--

---

<div class="post-metadata">

**Author:** ![Meltemi](https://avatars.discourse-cdn.com/v4/letter/m/a87d85/32.png) [@Meltemi](https://discuss.elastic.co/u/Meltemi)\
**Post date:** [August 29, 2012, 6:28pm UTC](https://discuss.elastic.co/t/possible-to-index-pdfs-by-page/8883/3 "2012-08-29T18:28:42Z")

</div>

Yeah, that's my post from a few months ago (lingering project, don't  
ask)...and I got a _very_ helpful answer on it _but_ it doesn't answer \*this  
_question: How to get elasticsearch to index the PDFs and include the \*  
page_ information so we can then use the advice in that post to serve the  
individual _pages_?!?

Do we need to break the PDFs up into individual pages and _then_ feed them  
into ES and somehow associate those individual pages back to a parent? Or  
is there a way to have ES, when it indexes a whole PDF(parent), add some  
kind of page meta-data to the text as it indexes each page(child)? Or is  
there a better way to do this?

Thanks for any & all advice!

On Wednesday, August 29, 2012 10:41:56 AM UTC-7, Clinton Gormley wrote:

> Have a look at this:
> 
> [Best practices for searchable archive of thousands of documents (pdf and/or xml) - Stack Overflow](http://stackoverflow.com/questions/10854858/best-practices-for-searchable-archive-of-thousands-of-documents-pdf-and-or-xml)
> 
> clint
> 
> On Wed, Aug 29, 2012 at 7:13 PM, Meltemi \<[mdeme...@gmail.com](mailto:mdeme...@gmail.com) \<javascript:\>
> 
> > wrote:
> 
> > Can elasticsearch index an attachment (PDF specifically) so the  
> > parent/child relationship between the _document_ (PDF) and the _page_ are  
> > preserved?
> > 
> > Our requirement dictates that matches should initially return the title  
> > of the PDF where the match occurred. Then if user wants to drill down  
> > further that _only_ the actual page where the hit occurred (with  
> > highlighting) should be presented. From there user should be able to page  
> > forward (or back) to continue reading. We should _not_ return the entire  
> > 100+ page documents but _only_ individual pages from within each  
> > document. Anyone know how to do this with elasticsearch?
> > 
> > --

--

---

<div class="post-metadata">

**Author:** ![phill](https://avatars.discourse-cdn.com/v4/letter/p/779978/32.png) [@phill](https://discuss.elastic.co/u/phill)\
**Post date:** [August 30, 2012, 12:27am UTC](https://discuss.elastic.co/t/possible-to-index-pdfs-by-page/8883/4 "2012-08-30T00:27:47Z")

</div>

I would like the same information and was wondering if Lucene payloads  
could somehow be leveraged (but those are a long way away when using ES).  
Here are a few problems with one page in each document. If there is  
sentence that continues on the next page, a phrase won't be matched.  
Another question: is a combined score of all pages for all terms  
equivalent to the whole document?

recall that  
idf = inverse document frequency, a formula based on the number of  
documents (not pages), but it is trying to give scores to rare vs common  
words, so maybe it all works out.  
and  
tf = term frequency in a document (not in a page)

I don't know the answer to these questions.

-Paul

On 8/29/2012 11:28 AM, Meltemi wrote:

> Yeah, that's my post from a few months ago (lingering project, don't  
> ask)...and I got a /very/ helpful answer on it /but/ it doesn't answer  
> /this/question: How to get elasticsearch to index the PDFs and  
> /include/ the _page_ information so we can then use the advice in that  
> post to serve the individual _pages_?!?
> 
> Do we need to break the PDFs up into individual pages and /then/ feed  
> them into ES and somehow associate those individual pages back to a  
> parent? Or is there a way to have ES, when it indexes a whole  
> PDF(parent), add some kind of page meta-data to the text as it indexes  
> each page(child)? Or is there a better way to do this?
> 
> Thanks for any & all advice!

--

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [August 31, 2012, 2:09pm UTC](https://discuss.elastic.co/t/possible-to-index-pdfs-by-page/8883/5 "2012-08-31T14:09:41Z")

</div>

Hi Meltemi

On Wed, 2012-08-29 at 11:28 -0700, Meltemi wrote:

> Yeah, that's my post from a few months ago (lingering project, don't  
> ask)...and I got a very helpful answer on it but it doesn't answer  
> thisquestion: How to get elasticsearch to index the PDFs and include  
> the page information so we can then use the advice in that post to  
> serve the individual pages?!?
> 
> Do we need to break the PDFs up into individual pages and then feed  
> them into ES and somehow associate those individual pages back to a  
> parent?

Yes, you need to do what you describe above.

Reread the answer I gave on

> <https://stackoverflow.com/questions/10854858/best-practices-for-searchable-archive-of-thousands-of-documents-pdf-and-or-xml/10861308#10861308>

starting from "First the indexing part: storing your docs in  
Elasticsearch:"

I give a step-by-step guid explaining how to do it.

If this doesn't answer your question, them I'm missing the bit you don't  
understand.

clint

> 

--

---

<div class="post-metadata">

**Author:** ![Santosh\_B](https://avatars.discourse-cdn.com/v4/letter/s/ea666f/32.png) [@Santosh\_B](https://discuss.elastic.co/u/Santosh_B)\
**Post date:** [August 27, 2014, 6:31am UTC](https://discuss.elastic.co/t/possible-to-index-pdfs-by-page/8883/6 "2014-08-27T06:31:30Z")

</div>

Hi,  
So what design approach did you follow ?  
Am thinking of storing the contents of pdf and indexing it in  
Elasticsearch and storing the link in filesystem/s3 or some NOSQL.  
When querying Elasticsearch use term vector to extract position offset and  
then extract the contents from the file system(may be some extra bytes  
before and after offset.)

On Wednesday, 29 August 2012 23:58:42 UTC+5:30, Meltemi wrote:

> Yeah, that's my post from a few months ago (lingering project, don't  
> ask)...and I got a _very_ helpful answer on it _but_ it doesn't answer  
> _this_question: How to get elasticsearch to index the PDFs and _include_  
> the _page_ information so we can then use the advice in that post to  
> serve the individual _pages_?!?
> 
> Do we need to break the PDFs up into individual pages and _then_ feed  
> them into ES and somehow associate those individual pages back to a parent?  
> Or is there a way to have ES, when it indexes a whole PDF(parent), add some  
> kind of page meta-data to the text as it indexes each page(child)? Or is  
> there a better way to do this?
> 
> Thanks for any & all advice!
> 
> On Wednesday, August 29, 2012 10:41:56 AM UTC-7, Clinton Gormley wrote:
> 
> > Have a look at this:
> > 
> > [Best practices for searchable archive of thousands of documents (pdf and/or xml) - Stack Overflow](http://stackoverflow.com/questions/10854858/best-practices-for-searchable-archive-of-thousands-of-documents-pdf-and-or-xml)
> > 
> > clint
> > 
> > On Wed, Aug 29, 2012 at 7:13 PM, Meltemi [mdeme...@gmail.com](mailto:mdeme...@gmail.com) wrote:
> > 
> > > Can elasticsearch index an attachment (PDF specifically) so the  
> > > parent/child relationship between the _document_ (PDF) and the _page_ are  
> > > preserved?
> > > 
> > > Our requirement dictates that matches should initially return the title  
> > > of the PDF where the match occurred. Then if user wants to drill down  
> > > further that _only_ the actual page where the hit occurred (with  
> > > highlighting) should be presented. From there user should be able to page  
> > > forward (or back) to continue reading. We should _not_ return the  
> > > entire 100+ page documents but _only_ individual pages from within each  
> > > document. Anyone know how to do this with elasticsearch?
> > > 
> > > --

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/f8fab5a5-d7ac-4aaf-bdff-a0b12035a516%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f8fab5a5-d7ac-4aaf-bdff-a0b12035a516%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:06am UTC](https://discuss.elastic.co/t/possible-to-index-pdfs-by-page/8883/7 "2017-07-06T01:06:06Z")

</div>


