# Any suggestion for duplicate data?

**URL:** <https://discuss.elastic.co/t/any-suggestion-for-duplicate-data/10188>\
**Category:** Elasticsearch\
**Created:** [December 28, 2012, 2:56am UTC](https://discuss.elastic.co/t/any-suggestion-for-duplicate-data/10188 "2012-12-28T02:56:56Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Burak\_Emre\_Kabakci](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/burak_emre_kabakci/32/2212_2.png) [@Burak\_Emre\_Kabakci](https://discuss.elastic.co/u/Burak_Emre_Kabakci)\
**Post date:** [December 28, 2012, 2:56am UTC](https://discuss.elastic.co/t/any-suggestion-for-duplicate-data/10188/1 "2012-12-28T02:56:56Z")

</div>

Here is the mapping of my index: [https://gist.github.com/4394060](https://gist.github.com/4394060)

I have used parent & child mapping to normalize data but as far as I  
understand there is no way to get any fields from \_parent document. Now,  
I'm trying to find the best way of storing artist\_name and release\_name in  
song type. I won't query these fields but I should be able to get them when  
I query other fields like name. There will be millions of song and I don't  
have much memory so I suspect that these duplicate values may cause out of  
memory. For now, artist\_name and release\_name fields are has "index": "no"  
property and I turned on compression for \_source field. Do you have any  
efficient suggestion for avoiding duplicate values like querying multiple  
queries or hacky way to get fields from \_parent document or denormalized  
data is the only way to handle this kindle of problem?

--

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [December 31, 2012, 1:20pm UTC](https://discuss.elastic.co/t/any-suggestion-for-duplicate-data/10188/2 "2012-12-31T13:20:02Z")

</div>

Are you talking about memory or disk space? Storing additional fields  
without indexing shouldn't affect memory in a any way, unless you  
are retrieving all these songs in a single request. If you are really  
concerned about denormalization you can always use mget or search to  
retrieve fields from parents.

On Thursday, December 27, 2012 9:56:56 PM UTC-5, Burak Emre Kabakcı wrote:

> Here is the mapping of my index: [elasticsearch mapping · GitHub](https://gist.github.com/4394060)
> 
> I have used parent & child mapping to normalize data but as far as I  
> understand there is no way to get any fields from \_parent document. Now,  
> I'm trying to find the best way of storing artist\_name and release\_name in  
> song type. I won't query these fields but I should be able to get them when  
> I query other fields like name. There will be millions of song and I  
> don't have much memory so I suspect that these duplicate values may cause  
> out of memory. For now, artist\_name and release\_name fields are has  
> "index": "no" property and I turned on compression for \_source field. Do  
> you have any efficient suggestion for avoiding duplicate values like  
> querying multiple queries or hacky way to get fields from \_parent document  
> or denormalized data is the only way to handle this kindle of problem?

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:58am UTC](https://discuss.elastic.co/t/any-suggestion-for-duplicate-data/10188/3 "2017-07-06T02:58:08Z")

</div>


