# Evaluating ES and need some input

**URL:** <https://discuss.elastic.co/t/evaluating-es-and-need-some-input/7285>\
**Category:** Elasticsearch\
**Created:** [April 10, 2012, 8:37pm UTC](https://discuss.elastic.co/t/evaluating-es-and-need-some-input/7285 "2012-04-10T20:37:09Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Beau\_Keogh](https://avatars.discourse-cdn.com/v4/letter/b/848f3c/32.png) [@Beau\_Keogh](https://discuss.elastic.co/u/Beau_Keogh)\
**Post date:** [April 10, 2012, 8:37pm UTC](https://discuss.elastic.co/t/evaluating-es-and-need-some-input/7285/1 "2012-04-10T20:37:09Z")

</div>

Hi all, I'm evaluating ES for our project and am wondering if it's a good  
idea to use it for storage and indexing, or should we store the documents  
elsewhere and use it just for indexing. We're going to be storing and  
indexing data that clients commit so it's absolutely imperative that we  
don't lose anything if the ES index goes south.

Should we store the documents in say S3 and index in ES, or is there a 100%  
safe way to store documents in ES? Or should I say, a 100% reliable way to  
recover if/when something goes wrong...

If an index does get corrupted which scenario will offer the best recovery  
options?

Also, thinking about having 1 index for each client. Any notable pros/cons  
to setting things up that way? Or should we do one index and reference  
documents by clientid?

Thanks!

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [April 10, 2012, 9:56pm UTC](https://discuss.elastic.co/t/evaluating-es-and-need-some-input/7285/2 "2012-04-10T21:56:19Z")

</div>

You might find this helpful:  
[http://stackoverflow.com/questions/6636508/elasticsearch-as-a-database](http://stackoverflow.com/questions/6636508/elasticsearch-as-a-database)

On Tuesday, April 10, 2012 4:37:09 PM UTC-4, Beau Keogh wrote:

> Hi all, I'm evaluating ES for our project and am wondering if it's a good  
> idea to use it for storage and indexing, or should we store the documents  
> elsewhere and use it just for indexing. We're going to be storing and  
> indexing data that clients commit so it's absolutely imperative that we  
> don't lose anything if the ES index goes south.
> 
> Should we store the documents in say S3 and index in ES, or is there a  
> 100% safe way to store documents in ES? Or should I say, a 100% reliable  
> way to recover if/when something goes wrong...
> 
> If an index does get corrupted which scenario will offer the best recovery  
> options?
> 
> Also, thinking about having 1 index for each client. Any notable pros/cons  
> to setting things up that way? Or should we do one index and reference  
> documents by clientid?
> 
> Thanks!

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [April 11, 2012, 11:39am UTC](https://discuss.elastic.co/t/evaluating-es-and-need-some-input/7285/3 "2012-04-11T11:39:38Z")

</div>

My recommendation is to also have an option to reindex the data when using  
elasticsearch, at least for the time being. The aim is definitely to  
eventually be a stable storage layer, but I recommend either doing backups  
or storing the data elsewhere for now.

Regarding using an index per user, here is a good thread:  
[https://groups.google.com/forum/?fromgroups#!searchin/elasticsearch/data$20flow/elasticsearch/49q-\_AgQCp8/MRol0t9asEcJ](https://groups.google.com/forum/?fromgroups#!searchin/elasticsearch/data$20flow/elasticsearch/49q-_AgQCp8/MRol0t9asEcJ)  
.

On Tue, Apr 10, 2012 at 11:37 PM, Beau Keogh [beaukeogh@gmail.com](mailto:beaukeogh@gmail.com) wrote:

> Hi all, I'm evaluating ES for our project and am wondering if it's a good  
> idea to use it for storage and indexing, or should we store the documents  
> elsewhere and use it just for indexing. We're going to be storing and  
> indexing data that clients commit so it's absolutely imperative that we  
> don't lose anything if the ES index goes south.
> 
> Should we store the documents in say S3 and index in ES, or is there a  
> 100% safe way to store documents in ES? Or should I say, a 100% reliable  
> way to recover if/when something goes wrong...
> 
> If an index does get corrupted which scenario will offer the best recovery  
> options?
> 
> Also, thinking about having 1 index for each client. Any notable pros/cons  
> to setting things up that way? Or should we do one index and reference  
> documents by clientid?
> 
> Thanks!

---

<div class="post-metadata">

**Author:** ![Beau\_Keogh](https://avatars.discourse-cdn.com/v4/letter/b/848f3c/32.png) [@Beau\_Keogh](https://discuss.elastic.co/u/Beau_Keogh)\
**Post date:** [April 11, 2012, 9:38pm UTC](https://discuss.elastic.co/t/evaluating-es-and-need-some-input/7285/4 "2012-04-11T21:38:55Z")

</div>

Thanks, that's very helpful.

So the post says: "A "user" based data flow, in theory, is perfect for an  
index per user case. If you have enough nodes (each shard is a Lucene  
index, which has a cost) in the cluster to do that, thats great, and  
several very large scale ES users actually do that."

What is very large scale? How many nodes would you recommend in order to  
support 100, 500, 1000, or 5000 clients using an "index per client" setup  
where index size could range from 0 to 15,000,000 documents which are  
typically 10-100 kb in size.

>

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [April 13, 2012, 12:11pm UTC](https://discuss.elastic.co/t/evaluating-es-and-need-some-input/7285/5 "2012-04-13T12:11:34Z")

</div>

Its really hard to tell based on the numbers you gave, since it also  
relates to what type of queries you execute, faceting / sorting, and what  
type of machines you have the nodes are running on.

On Thu, Apr 12, 2012 at 12:38 AM, Beau Keogh [beaukeogh@gmail.com](mailto:beaukeogh@gmail.com) wrote:

> Thanks, that's very helpful.
> 
> So the post says: "A "user" based data flow, in theory, is perfect for an  
> index per user case. If you have enough nodes (each shard is a Lucene  
> index, which has a cost) in the cluster to do that, thats great, and  
> several very large scale ES users actually do that."
> 
> What is very large scale? How many nodes would you recommend in order to  
> support 100, 500, 1000, or 5000 clients using an "index per client" setup  
> where index size could range from 0 to 15,000,000 documents which are  
> typically 10-100 kb in size.
> 
> >

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:32am UTC](https://discuss.elastic.co/t/evaluating-es-and-need-some-input/7285/6 "2017-07-06T03:32:42Z")

</div>


