# Recommendations on "schema" for an ElasticSearch project

**URL:** <https://discuss.elastic.co/t/recommendations-on-schema-for-an-elasticsearch-project/33847>\
**Category:** Elasticsearch\
**Created:** [November 5, 2015, 9:53am UTC](https://discuss.elastic.co/t/recommendations-on-schema-for-an-elasticsearch-project/33847 "2015-11-05T09:53:00Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Cylindric](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cylindric/32/5660_2.png) [@Cylindric](https://discuss.elastic.co/u/Cylindric)\
**Post date:** [November 5, 2015, 9:53am UTC](https://discuss.elastic.co/t/recommendations-on-schema-for-an-elasticsearch-project/33847/1 "2015-11-05T09:53:00Z")

</div>

I've build an ELK cluster for an internal proof-of-concept project, consisting of log-analysis (which is straight forward enough in an ELK context) but also for a data-mining exercise, for which I'm just using ElasticSearch (and possibly Kibana as the front-end).

I'm hoping that if I describe the scenario here, I'll get some tips'n'tricks and guidance to save me from some nightmares later when I have ALL the data in there and can't easily re-index it all...

I'm storing a few million contact and company records into ElasticSearch, with the requirement for later on extracting a subset of them based on various criteria. The query will be of the form

Show me all `contacts` that have a `job-level` of "_CEO_" at a `company` with a `countryCode` of "UK", `revenue` \> £1m and `industryCodes` of "_healthcare_" or "_finance_" but not "_marketing_".

We currently hold the data in various SQL Server databases, and I'm in the process of pumping them all into ES using a C# project with NEST. The aim is to combine various disparate data-sources into one definitive source.

The questions:

1. I'm putting the `company` and `contact` objects into a single index. Because Kibana doesn't appear to support the parent/child relationships of ES, I'm duplicating a lot of the data by including the company data on the contact. For example, the field-list in Kibana shows both "companyname" (from all the company records) and also "company.companyname" (from all the contact records). Should I split these into two indexes? Can I even have a parent-child relationship across indexes? I could also ditch the parent-child concept and just flatten the "contact" object to contain all the fields on "company", and have a separate (much smaller!) index for just searching company info.

2. I'm using a 5\*2 shard setup. Is there any advantage to splitting the index like the logstash defaults do? For example, corpdata-a, corpdata-b, etc. Or would increasing the number of shards be worthwhile?

Currently running on two Dell PowerEdge R710 servers with 12 cores and 98Gb of memory. The hosts are an ESX environment running a total of 4 ES nodes with 20G ES\_HEAP\_SIZE.

Apologies for the ramble, I just thought it might help to provide some context 😄

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:40pm UTC](https://discuss.elastic.co/t/recommendations-on-schema-for-an-elasticsearch-project/33847/2 "2017-07-05T23:40:23Z")

</div>


