# Nested objects and fragment highlighting

**URL:** <https://discuss.elastic.co/t/nested-objects-and-fragment-highlighting/4715>\
**Category:** Elasticsearch\
**Created:** [June 27, 2011, 11:39am UTC](https://discuss.elastic.co/t/nested-objects-and-fragment-highlighting/4715 "2011-06-27T11:39:38Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![johno\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/johno_2/32/3191_2.png) [@johno\_2](https://discuss.elastic.co/u/johno_2)\
**Post date:** [June 27, 2011, 11:39am UTC](https://discuss.elastic.co/t/nested-objects-and-fragment-highlighting/4715/1 "2011-06-27T11:39:38Z")

</div>

Hi there,

I am indexing documents created from pdf files. These files are  
processed and broken down to pages with text contained on them. Sample  
document looks like this:

{:name =\> "A name",  
:attachments =\> [  
{:name =\> "first attachment", :pages =\> [  
{:number =\> 1, :text =\> "page 1 contents"},  
{:number =\> 2, :text =\> "page 2 contents"},  
]  
]}

When I search for text with highlights i get response with  
field :highlight =\> {"attachments.pages.text" =\> array\_of\_fragments}.

Now the problem is, that I lose information about on which page/  
attachment the highlight is in. (Or any other possible page fields.)

I've also tried creating/indexing pages as separate documents, where i  
get all the fields back. However by doing so I cannot find a way to  
group them by attachment\_id/document\_id and thus I lose to ability to  
score attachments/documents with multiple matching pages better.

Any ideas how to solve this?

johno

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [June 27, 2011, 11:53am UTC](https://discuss.elastic.co/t/nested-objects-and-fragment-highlighting/4715/2 "2011-06-27T11:53:39Z")

</div>

Perhaps, you could index pages instead of a full doc containing many pages as you search for pages.

Using \_parent / \_childs, you could perhaps make the link between the document and its pages.

I didn't do it yet by myself but thinking of it. So I don't know if it will work or not.

Hope this could help  
David 😉

Le 27 juin 2011 à 13:39, johno [johno@jsmf.net](mailto:johno@jsmf.net) a écrit :

> Hi there,
> 
> I am indexing documents created from pdf files. These files are  
> processed and broken down to pages with text contained on them. Sample  
> document looks like this:
> 
> {:name =\> "A name",  
> :attachments =\> [  
> {:name =\> "first attachment", :pages =\> [  
> {:number =\> 1, :text =\> "page 1 contents"},  
> {:number =\> 2, :text =\> "page 2 contents"},  
> ]  
> ]}
> 
> When I search for text with highlights i get response with  
> field :highlight =\> {"attachments.pages.text" =\> array\_of\_fragments}.
> 
> Now the problem is, that I lose information about on which page/  
> attachment the highlight is in. (Or any other possible page fields.)
> 
> I've also tried creating/indexing pages as separate documents, where i  
> get all the fields back. However by doing so I cannot find a way to  
> group them by attachment\_id/document\_id and thus I lose to ability to  
> score attachments/documents with multiple matching pages better.
> 
> Any ideas how to solve this?
> 
> johno

---

<div class="post-metadata">

**Author:** ![johno\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/johno_2/32/3191_2.png) [@johno\_2](https://discuss.elastic.co/u/johno_2)\
**Post date:** [June 27, 2011, 12:15pm UTC](https://discuss.elastic.co/t/nested-objects-and-fragment-highlighting/4715/3 "2011-06-27T12:15:11Z")

</div>

Hi, David. Thanks for the response, but as I written in my first post.  
I've already tried that, but there is another problem with scoring.

On Jun 27, 1:53 pm, David Pilato [da...@pilato.fr](mailto:da...@pilato.fr) wrote:

> Perhaps, you could index pages instead of a full doc containing many pages as you search for pages.
> 
> Using \_parent / \_childs, you could perhaps make the link between the document and its pages.
> 
> I didn't do it yet by myself but thinking of it. So I don't know if it will work or not.
> 
> Hope this could help  
> David 😉
> 
> Le 27 juin 2011 à 13:39, johno [jo...@jsmf.net](mailto:jo...@jsmf.net) a écrit :
> 
> > Hi there,
> 
> > I am indexing documents created from pdf files. These files are  
> > processed and broken down to pages with text contained on them. Sample  
> > document looks like this:
> 
> > {:name =\> "A name",  
> > :attachments =\> [  
> > {:name =\> "first attachment", :pages =\> [  
> > {:number =\> 1, :text =\> "page 1 contents"},  
> > {:number =\> 2, :text =\> "page 2 contents"},  
> > ]  
> > ]}
> 
> > When I search for text with highlights i get response with  
> > field :highlight =\> {"attachments.pages.text" =\> array\_of\_fragments}.
> 
> > Now the problem is, that I lose information about on which page/  
> > attachment the highlight is in. (Or any other possible page fields.)
> 
> > I've also tried creating/indexing pages as separate documents, where i  
> > get all the fields back. However by doing so I cannot find a way to  
> > group them by attachment\_id/document\_id and thus I lose to ability to  
> > score attachments/documents with multiple matching pages better.
> 
> > Any ideas how to solve this?
> 
> > johno

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [June 27, 2011, 12:26pm UTC](https://discuss.elastic.co/t/nested-objects-and-fragment-highlighting/4715/4 "2011-06-27T12:26:00Z")

</div>

Sorry. Didn't read well the end of your post ☹ shame on me !

So, i have no other idea but I would love to see a solution for that.

On my project, we have made a "group" like solution. We have two field :  
Field1  
Field2

We create a EsField which contains the content of field1 + separator + content of field2.  
Such as content1KKKKcontent2

So if i want documents containing a line with content1 AND content2, we can search for content1KKKKcontent2.

I'm not convinced that it can be useful for your use case.

Cheers  
David 😉

Le 27 juin 2011 à 14:15, johno [johno@jsmf.net](mailto:johno@jsmf.net) a écrit :

> Hi, David. Thanks for the response, but as I written in my first post.  
> I've already tried that, but there is another problem with scoring.
> 
> On Jun 27, 1:53 pm, David Pilato [da...@pilato.fr](mailto:da...@pilato.fr) wrote:
> 
> > Perhaps, you could index pages instead of a full doc containing many pages as you search for pages.
> > 
> > Using \_parent / \_childs, you could perhaps make the link between the document and its pages.
> > 
> > I didn't do it yet by myself but thinking of it. So I don't know if it will work or not.
> > 
> > Hope this could help  
> > David 😉
> > 
> > Le 27 juin 2011 à 13:39, johno [jo...@jsmf.net](mailto:jo...@jsmf.net) a écrit :
> > 
> > > Hi there,
> > 
> > > I am indexing documents created from pdf files. These files are  
> > > processed and broken down to pages with text contained on them. Sample  
> > > document looks like this:
> > 
> > > {:name =\> "A name",  
> > > :attachments =\> [  
> > > {:name =\> "first attachment", :pages =\> [  
> > > {:number =\> 1, :text =\> "page 1 contents"},  
> > > {:number =\> 2, :text =\> "page 2 contents"},  
> > > ]  
> > > ]}
> > 
> > > When I search for text with highlights i get response with  
> > > field :highlight =\> {"attachments.pages.text" =\> array\_of\_fragments}.
> > 
> > > Now the problem is, that I lose information about on which page/  
> > > attachment the highlight is in. (Or any other possible page fields.)
> > 
> > > I've also tried creating/indexing pages as separate documents, where i  
> > > get all the fields back. However by doing so I cannot find a way to  
> > > group them by attachment\_id/document\_id and thus I lose to ability to  
> > > score attachments/documents with multiple matching pages better.
> > 
> > > Any ideas how to solve this?
> > 
> > > johno

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:02am UTC](https://discuss.elastic.co/t/nested-objects-and-fragment-highlighting/4715/5 "2017-07-06T04:02:32Z")

</div>


