# ES data redundancy VS Kibana visualizations

**URL:** <https://discuss.elastic.co/t/es-data-redundancy-vs-kibana-visualizations/270022>\
**Category:** Kibana\
**Created:** [April 13, 2021, 3:00pm UTC](https://discuss.elastic.co/t/es-data-redundancy-vs-kibana-visualizations/270022 "2021-04-13T15:00:58Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![MartinKolar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/martinkolar/32/78690_2.png) [@MartinKolar](https://discuss.elastic.co/u/MartinKolar)\
**Post date:** [April 13, 2021, 3:00pm UTC](https://discuss.elastic.co/t/es-data-redundancy-vs-kibana-visualizations/270022/1 "2021-04-13T15:00:58Z")

</div>

Hi everyone, I would love to hear some tips on how to avoid data redundancy in ES and Kibana.

We are using ES and Kibana to collect and visualize videogame analytics. The game client indexes an "event" - a document with lots of fields like "EventTimestamp", "EventName", "UserID", "SessionID", "ScreenResolution" etc. - when the player performs some significant actions as launching the game, loading the gameplay, dying, returning to the menu etc. All events until quitting the game are considered as one "session", marked by a unique "SessionID".

The problem is that there is dozens of fields (like "GPU\_Name") whose values are constant throughout the whole session, but we still send them with every event in the session. That significantly raises the storage size of all individual event documents and makes the index size grow quickly. It feels wasteful to use the storage and memory for that much redundant information.

But we are sending the information with every individual event to be able to leverage it in Kibana. If I want to visualize for example the number of crashes by "GPU\_Name", it's very straightforward if I have the gpu field on the "GameCrash" event document", but as far as I know it is very difficult or impossible to do if the gpu field is only on the "SessionStart" event document sent half an hour prior.

So, my question is: Should we just settle with redundant fields taking our storage space? Or is there some trick to structure our data differently to avoid duplicate information across the events in a session, but still be able to use it in Kibana's visualizations?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [April 13, 2021, 10:44pm UTC](https://discuss.elastic.co/t/es-data-redundancy-vs-kibana-visualizations/270022/2 "2021-04-13T22:44:17Z")

</div>

Welcome to our community! 😃

There's no tricks available, no.  
Are you using `best_compression` on your indices?

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [April 14, 2021, 12:41am UTC](https://discuss.elastic.co/t/es-data-redundancy-vs-kibana-visualizations/270022/3 "2021-04-14T00:41:05Z")

</div>

Hi @MartinKolar welcome to the community.

Also remember a proper mapping will help reduce storage for all those `keywords` GPU\_Name if you set the datatype as a `keyword` then the storage will be optimized if you leave it as default it will be saved as both `text` and `keyword` using more storage.

AND this feels weird but that remember that for `keywords` that Elasticsearch using an inverted index so technically the GPU\_Name is stored once then the inverted document tells which document it belongs too...

Like you said There is also other benefits of indexing that field as it will let you do important aggregations across that fields with other dimensions like How many Crashes for GPU\_Name = \< gpu\_name\>

You can also chose not to store i.e. [disable `_source`](https://www.elastic.co/guide/en/elasticsearch/reference/7.12/mapping-source-field.html) which will reduce storage (but I would be careful with that there are some serious downsides to that).

There is a setting to prune some of source but as the docs say that is a "Expert Setting" and it still has ramifications.

All that may save very little, I would get started, elasticsearch is pretty efficient... and see where you get whether it is really an issue or not.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 12, 2021, 12:41am UTC](https://discuss.elastic.co/t/es-data-redundancy-vs-kibana-visualizations/270022/4 "2021-05-12T00:41:59Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
