# Need some help / idea about architecture

**URL:** https://discuss.elastic.co/t/need-some-help-idea-about-architecture/6144
**Category:** Elasticsearch
**Created:** [December 13, 2011, 5:05pm UTC](https://discuss.elastic.co/t/need-some-help-idea-about-architecture/6144 "2011-12-13T17:05:19Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![TheKnto](https://avatars.discourse-cdn.com/v4/letter/t/b9bd4f/32.png) [@TheKnto](https://discuss.elastic.co/u/TheKnto)
#### Post date: [December 13, 2011, 5:05pm UTC](https://discuss.elastic.co/t/need-some-help-idea-about-architecture/6144/1 "2011-12-13T17:05:19Z")

</div>

Hi everyone,  
I'am working on a new webapp which is dedicated to search documents on  
several criterias.

We intend to use elasticsearch as a search engine and are very keen on it (  
@Kimchy : fantastic job ! ; ) ) .

- the documents we want to index are quite complex : several data levels ,  
so we use nested fields in our mapping;

- the space used in the source (3 databases couchDB) is 900 Go by year,

- these 3 differents databases in couchDB are indexed in ES on 1 cluster:  
the space used by all the indexes in ES is about 2.1 To a year,

- we have an index per month per database source (12 indexes by year per  
database source)

- each index has 2 types and 1 replica, 5 shards

- the total number of indexed documents is 15 Millions

- the data indexed are stored on a disk bay (RAID-5)

- several fields (about 150) are opened for search

- we intend to use facets in queries (this will be a new functionnality and  
could increase the numbers of queries done and users logged),

- each query is limited to a search period of only 1 year

- 1000 differents users log every day to execute a mean of 2 queries

* * *

We are facing the problem of the architecture to start (number of servers,  
number of CPUs, power, RAM, etc ..)  
We need to have a scalable solution because we will have to index 4 years  
of datas without decreasing perfs.

Has anyone an idea of the best approach ?  
Help will be very usefull and appreciated.  
Many thanks

---

<div class="post-metadata">

### Author: ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)
#### Post date: [December 13, 2011, 5:11pm UTC](https://discuss.elastic.co/t/need-some-help-idea-about-architecture/6144/2 "2011-12-13T17:11:38Z")

</div>

Hi,

You can sample portion of the data and do estimate based on this, no?

--  
Regards,  
Lukas

On Tuesday, December 13, 2011 at 6:05 PM, TheKnto wrote:

> Hi everyone,  
> I'am working on a new webapp which is dedicated to search documents on several criterias.
> 
> We intend to use elasticsearch as a search engine and are very keen on it ( @Kimchy : fantastic job ! ; ) ) .
> 
> - the documents we want to index are quite complex : several data levels , so we use nested fields in our mapping;
> 
> - the space used in the source (3 databases couchDB) is 900 Go by year,
> 
> - these 3 differents databases in couchDB are indexed in ES on 1 cluster: the space used by all the indexes in ES is about 2.1 To a year,
> 
> - we have an index per month per database source (12 indexes by year per database source)
> 
> - each index has 2 types and 1 replica, 5 shards
> 
> - the total number of indexed documents is 15 Millions
> 
> - the data indexed are stored on a disk bay (RAID-5)
> 
> - several fields (about 150) are opened for search
> 
> - we intend to use facets in queries (this will be a new functionnality and could increase the numbers of queries done and users logged),
> 
> - each query is limited to a search period of only 1 year
> 
> - 1000 differents users log every day to execute a mean of 2 queries
> 
> * * *
> 
> We are facing the problem of the architecture to start (number of servers, number of CPUs, power, RAM, etc ..)  
> We need to have a scalable solution because we will have to index 4 years of datas without decreasing perfs.
> 
> Has anyone an idea of the best approach ?  
> Help will be very usefull and appreciated.  
> Many thanks

---

<div class="post-metadata">

### Author: ![TheKnto](https://avatars.discourse-cdn.com/v4/letter/t/b9bd4f/32.png) [@TheKnto](https://discuss.elastic.co/u/TheKnto)
#### Post date: [December 13, 2011, 5:35pm UTC](https://discuss.elastic.co/t/need-some-help-idea-about-architecture/6144/3 "2011-12-13T17:35:02Z")

</div>

Hello Lukáš,  
at this time we have only succeeded on indexing 1 month of 1 database  
source on a cluster with 2 nodes (5Go RAM). Each nodes was running on a  
separate server but these servers were also running the couchdb job. So  
it's not the good config and the results may be wrong.

As for the production architecture, we have to define it in order to buy  
the machines, so it's quite risky to evaluate the results based on one  
month. I would hope that someone has experiment to share ...

best regards,  
Cyril

2011/12/13 Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com)

> Hi,
> 
> You can sample portion of the data and do estimate based on this, no?
> 
> --  
> Regards,  
> Lukas
> 
> On Tuesday, December 13, 2011 at 6:05 PM, TheKnto wrote:
> 
> Hi everyone,  
> I'am working on a new webapp which is dedicated to search documents on  
> several criterias.
> 
> We intend to use elasticsearch as a search engine and are very keen on it  
> ( @Kimchy : fantastic job ! ; ) ) .
> 
> - the documents we want to index are quite complex : several data levels ,  
> so we use nested fields in our mapping;
> 
> - the space used in the source (3 databases couchDB) is 900 Go by year,
> 
> - these 3 differents databases in couchDB are indexed in ES on 1 cluster:  
> the space used by all the indexes in ES is about 2.1 To a year,
> 
> - we have an index per month per database source (12 indexes by year per  
> database source)
> 
> - each index has 2 types and 1 replica, 5 shards
> 
> - the total number of indexed documents is 15 Millions
> 
> - the data indexed are stored on a disk bay (RAID-5)
> 
> - several fields (about 150) are opened for search
> 
> - we intend to use facets in queries (this will be a new functionnality  
> and could increase the numbers of queries done and users logged),
> 
> - each query is limited to a search period of only 1 year
> 
> - 1000 differents users log every day to execute a mean of 2 queries
> 
> * * *
> 
> We are facing the problem of the architecture to start (number of servers,  
> number of CPUs, power, RAM, etc ..)  
> We need to have a scalable solution because we will have to index 4 years  
> of datas without decreasing perfs.
> 
> Has anyone an idea of the best approach ?  
> Help will be very usefull and appreciated.  
> Many thanks

---

<div class="post-metadata">

### Author: ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)
#### Post date: [December 13, 2011, 8:30pm UTC](https://discuss.elastic.co/t/need-some-help-idea-about-architecture/6144/4 "2011-12-13T20:30:01Z")

</div>

Sadly, I do not think anybody can give you guaranteed answer because it depends on many factors, it is not only about number of documents and number of queries. It is also important how exactly you structure the documents before indexing and which analysis you execute on it, and on the type of queries (simple, facets, prefix, … etc). Also I bet you if you just started with the ES stuff then I am sure sooner or later you will find that you want change your documents mapping and analysis or rewrite some queries… this will have impact as well. Also you should consider how you want to go about upgrades if complete reindexing will be necessary.  
Can't you for example rent AWS for some time and do capacity planning there? May be that would help.

But may be someone can share some experience, I personally do not have experience with that big data in ES.

--  
Regards,  
Lukas

On Tuesday, December 13, 2011 at 6:35 PM, TheKnto wrote:

> Hello Lukáš,  
> at this time we have only succeeded on indexing 1 month of 1 database source on a cluster with 2 nodes (5Go RAM). Each nodes was running on a separate server but these servers were also running the couchdb job. So it's not the good config and the results may be wrong.
> 
> As for the production architecture, we have to define it in order to buy the machines, so it's quite risky to evaluate the results based on one month. I would hope that someone has experiment to share ...
> 
> best regards,  
> Cyril
> 
> 2011/12/13 Lukáš Vlček \<[lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com) ([mailto:lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com))\>
> 
> > Hi,
> > 
> > You can sample portion of the data and do estimate based on this, no?
> > 
> > --  
> > Regards,  
> > Lukas
> > 
> > On Tuesday, December 13, 2011 at 6:05 PM, TheKnto wrote:
> > 
> > > Hi everyone,  
> > > I'am working on a new webapp which is dedicated to search documents on several criterias.
> > > 
> > > We intend to use elasticsearch as a search engine and are very keen on it ( @Kimchy : fantastic job ! ; ) ) .
> > > 
> > > - the documents we want to index are quite complex : several data levels , so we use nested fields in our mapping;
> > > 
> > > - the space used in the source (3 databases couchDB) is 900 Go by year,
> > > 
> > > - these 3 differents databases in couchDB are indexed in ES on 1 cluster: the space used by all the indexes in ES is about 2.1 To a year,
> > > 
> > > - we have an index per month per database source (12 indexes by year per database source)
> > > 
> > > - each index has 2 types and 1 replica, 5 shards
> > > 
> > > - the total number of indexed documents is 15 Millions
> > > 
> > > - the data indexed are stored on a disk bay (RAID-5)
> > > 
> > > - several fields (about 150) are opened for search
> > > 
> > > - we intend to use facets in queries (this will be a new functionnality and could increase the numbers of queries done and users logged),
> > > 
> > > - each query is limited to a search period of only 1 year
> > > 
> > > - 1000 differents users log every day to execute a mean of 2 queries
> > > 
> > > * * *
> > > 
> > > We are facing the problem of the architecture to start (number of servers, number of CPUs, power, RAM, etc ..)  
> > > We need to have a scalable solution because we will have to index 4 years of datas without decreasing perfs.
> > > 
> > > Has anyone an idea of the best approach ?  
> > > Help will be very usefull and appreciated.  
> > > Many thanks

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 3:45am UTC](https://discuss.elastic.co/t/need-some-help-idea-about-architecture/6144/5 "2017-07-06T03:45:32Z")

</div>


