# Join multiple Independent Indices

**URL:** <https://discuss.elastic.co/t/join-multiple-independent-indices/35796>\
**Category:** Elasticsearch\
**Created:** [November 28, 2015, 2:12am UTC](https://discuss.elastic.co/t/join-multiple-independent-indices/35796 "2015-11-28T02:12:22Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![mvenkat\_in](https://avatars.discourse-cdn.com/v4/letter/m/58f4c7/32.png) [@mvenkat\_in](https://discuss.elastic.co/u/mvenkat_in)\
**Post date:** [November 28, 2015, 2:12am UTC](https://discuss.elastic.co/t/join-multiple-independent-indices/35796/1 "2015-11-28T02:12:22Z")

</div>

Environment: File Beat (1.0.0-rc2) --\> Log Stash (2.1.0)--\> Elastic Search (2.1.0)--\> Kibana (4.3)  
Use Case: Real time application transactions logs metrics analysis & monitoring, Search  
Domain: Telecom IT  
Description:  
The Logstash collects logs from multiple applications and indexes to ES in a different Index for each application. For example  
Logs from App 1 would be indexed to Index\_1  
Logs from App 2 would be indexed to Index\_2  
Logs from App 3 would be indexed to Index\_3

These logs can/may not be inserted to the same Index, as these are with different formats.  
Each Index has a common filed say subscriber ID.

Requirement:  
I want to search for the information of a subscriber joining the data for multiple Indexes. How can I achieve this !! I have gone through few options like parent-child relationship, de normalize at index time.. However those may not be applicable in my use case.

Kindly suggest any approach.

Thanks  
Venkatesh

---

<div class="post-metadata">

**Author:** ![vtst2412](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vtst2412/32/6228_2.png) [@vtst2412](https://discuss.elastic.co/u/vtst2412)\
**Post date:** [November 28, 2015, 2:37am UTC](https://discuss.elastic.co/t/join-multiple-independent-indices/35796/2 "2015-11-28T02:37:14Z")

</div>

You can absolutely have data in different format / with different fields in the same index. You just need to define them as different types. And since they share common subscriber ID you can query on those.

---

<div class="post-metadata">

**Author:** ![mvenkat\_in](https://avatars.discourse-cdn.com/v4/letter/m/58f4c7/32.png) [@mvenkat\_in](https://discuss.elastic.co/u/mvenkat_in)\
**Post date:** [November 30, 2015, 3:39am UTC](https://discuss.elastic.co/t/join-multiple-independent-indices/35796/3 "2015-11-30T03:39:35Z")

</div>

Thanks Vicent for the Idea. I have modeled the data to be in different types for each different source.  
However I am not able to breakthrough the following.  
Say in the normal RDBMS world, I have 2 tables for 2 sources of data.  
Table A - Common ID, Col 1A, Col 2A, Col 3A  
Table B - Common ID, Col 1B, Col 2B  
To Join them - Select Col 1A, Col 2A, Col 1B, Col 2B From Table A, Table B where A.CommonID =B.Common ID;

How Can I achieve this in Kibana Discover. I have seen numerous discussions on this idea in the internet but I couldn't get a straight forward/simplified solution.

It would great if you can throw some pointers/tips. 🙂

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 30, 2015, 6:18am UTC](https://discuss.elastic.co/t/join-multiple-independent-indices/35796/4 "2015-11-30T06:18:29Z")

</div>

You can't really join in elasticsearch unless you use parent child but Kibana does not support it.

It's definitely better to model your documents in a different way.  
Just index everything in a single doc and forget relations.

---

<div class="post-metadata">

**Author:** ![mvenkat\_in](https://avatars.discourse-cdn.com/v4/letter/m/58f4c7/32.png) [@mvenkat\_in](https://discuss.elastic.co/u/mvenkat_in)\
**Post date:** [November 30, 2015, 7:13am UTC](https://discuss.elastic.co/t/join-multiple-independent-indices/35796/5 "2015-11-30T07:13:29Z")

</div>

Thanks David. OK In such case can you give me a high level idea how to do it.

I have multiple source applications from which the data is collected by Logstash using various plugin such as File beat, File, JDBC and currently indexing to ES in different mapping types.

In order to Index all the info to single doc, where and how can Join the data from various sources before indexing to ES. Or is there any way to generate a "global" index with all the data from different indices !!

I am really finding it very interesting to think beyond the typical RDBMS ideas.  
Thanks

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 30, 2015, 7:37am UTC](https://discuss.elastic.co/t/join-multiple-independent-indices/35796/6 "2015-11-30T07:37:55Z")

</div>

Typically index something like:

```auto
{
 "Col1A": "",
 "Col2A": "",
 "Col3A": "",
 "Col1B": "",
 "Col2B": ""
}

```

But you can't really do join in LS IMO or it could be hard to do so. May be you could fetch an existing data from elasticsearch using [https://www.elastic.co/guide/en/logstash/2.1/plugins-inputs-http.html](https://www.elastic.co/guide/en/logstash/2.1/plugins-inputs-http.html) but unsure.

If you want to display multiple time series data on Kibana, you should give a look at [TimeLion](https://www.elastic.co/blog/timelion-timeline).

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 30, 2015, 8:34am UTC](https://discuss.elastic.co/t/join-multiple-independent-indices/35796/7 "2015-11-30T08:34:25Z")

</div>

When using time-based indices, it is often difficult to use parent-child relationships, as a limitation is that it requires all documents in a hierarchy to be present in the same shard. For time-based data it is therefore generally best to denormalise your data at index time. This typically works well when you have reasonably static data, e.g. customer or subscriber information, that can be added onto the time-based events.

As long as you control the names and mappings for fields that are common across different types of data, e.g. customer id, you can create visualisations in Kibana based on different underlying indices and place these in a single dashboard. You can then filter on these common fields in the dashboard and the filter will be applied to all visualisations, irrespective of underlying index.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [November 30, 2015, 2:52pm UTC](https://discuss.elastic.co/t/join-multiple-independent-indices/35796/8 "2015-11-30T14:52:02Z")

</div>

See [How can I use aggregations to query distinct values across all time grouped by first seen](https://discuss.elastic.co/t/how-can-i-use-aggregations-to-query-distinct-values-across-all-time-grouped-by-first-seen/25482/16) and the links to using an entity-centric indexing approach.

In your case the central "entity" would be a subscriber and you would be fusing data from multiple "event" indices.

Cheers  
Mark

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:34pm UTC](https://discuss.elastic.co/t/join-multiple-independent-indices/35796/9 "2017-07-05T23:34:50Z")

</div>


