# Usage of Elastic Search in Production Environment

**URL:** https://discuss.elastic.co/t/usage-of-elastic-search-in-production-environment/5049
**Category:** Elasticsearch
**Created:** [August 3, 2011, 6:33pm UTC](https://discuss.elastic.co/t/usage-of-elastic-search-in-production-environment/5049 "2011-08-03T18:33:52Z")
**Posts on this page:** 12
**Page:** 1

<div class="post-metadata">

### Author: ![Jason\_N](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jason_n/32/3149_2.png) [@Jason\_N](https://discuss.elastic.co/u/Jason_N)
#### Post date: [August 3, 2011, 6:33pm UTC](https://discuss.elastic.co/t/usage-of-elastic-search-in-production-environment/5049/1 "2011-08-03T18:33:52Z")

</div>

Does anyone know if Elastic Search is/has been used in a production  
environment? Specifically looking at a case where one index would be  
receiving ~3m new records indexed into the cluster per week with a starting  
set of ~600m. While we wouldn't have a huge user base, downtime and data  
loss would be a huge issue.

I guess I'm slightly discouraged by the amount of issues (likely operator  
error) I've been having with both ingest and general cluster health w/ ES  
(lost an unreplicated shard today), and was hoping for a "yeah we've been  
there, and we are processing 10x your data needs no issues..."

Thanks,

J

---

<div class="post-metadata">

### Author: ![felipera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/felipera/32/3038_2.png) [@felipera](https://discuss.elastic.co/u/felipera)
#### Post date: [August 3, 2011, 8:55pm UTC](https://discuss.elastic.co/t/usage-of-elastic-search-in-production-environment/5049/2 "2011-08-03T20:55:13Z")

</div>

"yeah we've been there, and we are processing 10x your data needs no  
issues..." on a few cases.

---

<div class="post-metadata">

### Author: ![Jay\_Taylor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jay_taylor/32/3151_2.png) [@Jay\_Taylor](https://discuss.elastic.co/u/Jay_Taylor)
#### Post date: [August 3, 2011, 8:58pm UTC](https://discuss.elastic.co/t/usage-of-elastic-search-in-production-environment/5049/3 "2011-08-03T20:58:07Z")

</div>

I've been ingesting a few hundred GB per day now with ES without any major  
issues.

---

<div class="post-metadata">

### Author: ![dottom](https://avatars.discourse-cdn.com/v4/letter/d/c68b51/32.png) [@dottom](https://discuss.elastic.co/u/dottom)
#### Post date: [August 3, 2011, 11:21pm UTC](https://discuss.elastic.co/t/usage-of-elastic-search-in-production-environment/5049/4 "2011-08-03T23:21:50Z")

</div>

I am up to 200M documents inserted per day using HTTP REST API bulk  
insertion with refresh settings from 5s to 20s.

The only problems I've had so far:

- Transaction logs take up a lot of space, which I believe 0.17.3 will  
address that.
- Some operational issues had catastrophic impact, such as not having enough  
open files and the index getting corrupted.
- Smaller java heap size that ran out of memory, when crashing, sometimes  
corrupted the index. I tried to reproduce but corruption on crash doesn't  
occur in all cases..
- Index and search response can be slow for a few minutes when starting up  
ES and you have really large indexes.

I would like a way to recover an index short of re-indexing, as 6 months  
down the road, I do not want to re-index everything should something  
unexpected happen.

On Wed, Aug 3, 2011 at 1:58 PM, Jay Taylor [outtatime@gmail.com](mailto:outtatime@gmail.com) wrote:

> I've been ingesting a few hundred GB per day now with ES without any major  
> issues.

---

<div class="post-metadata">

### Author: ![Eran\_Kutner\_2](https://avatars.discourse-cdn.com/v4/letter/e/dc4da7/32.png) [@Eran\_Kutner\_2](https://discuss.elastic.co/u/Eran_Kutner_2)
#### Post date: [August 4, 2011, 6:13am UTC](https://discuss.elastic.co/t/usage-of-elastic-search-in-production-environment/5049/5 "2011-08-04T06:13:29Z")

</div>

Can you (and other other people who responded) say something about your  
setup? Will be very interesting to know what configuration you're using to  
support this index.  
I'm particularly interested in:

- How many servers, and roughly what's their hardware configuration?
- How many shards?
- How many replicas?
- Are they all on the same index?
- How many documents per index?
- What kind of performance are you getting?
- Anything else interesting 🙂

Thanks.

---

<div class="post-metadata">

### Author: ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)
#### Post date: [August 4, 2011, 6:30am UTC](https://discuss.elastic.co/t/usage-of-elastic-search-in-production-environment/5049/6 "2011-08-04T06:30:16Z")

</div>

Some interesting information is also here : [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/users)

BTW, I'm going into production next week with a small ES cluster (0.16.5 with less than 200k docs) as a proof of concept.

David 😉

Le 4 août 2011 à 08:13, Eran Kutner [eran@gigya-inc.com](mailto:eran@gigya-inc.com) a écrit :

> Can you (and other other people who responded) say something about your setup? Will be very interesting to know what configuration you're using to support this index.  
> I'm particularly interested in:  
> How many servers, and roughly what's their hardware configuration?  
> How many shards?  
> How many replicas?  
> Are they all on the same index?  
> How many documents per index?  
> What kind of performance are you getting?  
> Anything else interesting 🙂
> 
> Thanks.

---

<div class="post-metadata">

### Author: ![danpolites](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danpolites/32/2136_2.png) [@danpolites](https://discuss.elastic.co/u/danpolites)
#### Post date: [August 4, 2011, 12:49pm UTC](https://discuss.elastic.co/t/usage-of-elastic-search-in-production-environment/5049/7 "2011-08-04T12:49:01Z")

</div>

We just went into production this week with 0.17.0 on Amazon EC2.

Two large instances running tomcat for with an Amazon elastic load  
balancer out in front.  
Two large instances that run Mongo (using replicasets) and  
Elasticsearch. We are using the default settings for shards/replicas  
at the moment for elasticsearch.  
We have a micro AMI with only elasticsearch on it so that we can bring  
up/take down one or more elasticsearch nodes at any given time.  
At this point we are only in the thousands as far as documents go, but  
our application records every request that a user makes so that will  
quickly grow.  
All I can say about performance is that it's been fast and it's been  
reliable. We had an issue with the trans log and too many open files  
(we had our limit set at 32k) the other day after we imported 10  
thousand users without using a bulk request in our Grails application.  
0.17.3 should fix the trans log issues.

Elasticsearch is amazing. It was so easy to set up and scale out. I  
think that MongoDB and Elasticsearch make a great couple in an  
environment like EC2 😉

On Aug 4, 2:13 am, Eran Kutner [e...@gigya-inc.com](mailto:e...@gigya-inc.com) wrote:

> Can you (and other other people who responded) say something about your  
> setup? Will be very interesting to know what configuration you're using to  
> support this index.  
> I'm particularly interested in:
> 
> - How many servers, and roughly what's their hardware configuration?
> - How many shards?
> - How many replicas?
> - Are they all on the same index?
> - How many documents per index?
> - What kind of performance are you getting?
> - Anything else interesting 🙂
> 
> Thanks.

---

<div class="post-metadata">

### Author: ![dbenson](https://avatars.discourse-cdn.com/v4/letter/d/958977/32.png) [@dbenson](https://discuss.elastic.co/u/dbenson)
#### Post date: [August 4, 2011, 4:30pm UTC](https://discuss.elastic.co/t/usage-of-elastic-search-in-production-environment/5049/8 "2011-08-04T16:30:27Z")

</div>

We have been in production since November.

- How many servers, and roughly what's their hardware configuration?
  - 4 servers. Each 24 core, 48G ram, 4x600G 10k rpm drives
  - 32M total documents

- How many shards?
  - Depends on the index size, typically 1-5. (indexes only using meta  
fields have 1, indexes using lots of full text extracted from PDFs get more)

- How many replicas?
  - 1, but we are adding more storage to increase this.

- Are they all on the same index?
  - We have ~50 indexes. Some for historical reasons, others for  
operational efficiency. We've found that when an index gets very large, the  
pain of handling incompatible field mappings, requiring a rebuild becomes a  
motivation fr capping indexes.

- How many documents per index?
  - thousands to few million (depending on application)

- What kind of performance are you getting?
  - 35-50ms for 200k req/min.

- Anything else interesting 🙂
  - ES is amazing and Shay's support totally rocks!

---

<div class="post-metadata">

### Author: ![jjasinek](https://avatars.discourse-cdn.com/v4/letter/j/b782af/32.png) [@jjasinek](https://discuss.elastic.co/u/jjasinek)
#### Post date: [August 4, 2011, 7:43pm UTC](https://discuss.elastic.co/t/usage-of-elastic-search-in-production-environment/5049/9 "2011-08-04T19:43:44Z")

</div>

Are you utilizing ES for just searching or for storage of the original  
document as well?

On Aug 4, 11:30 am, dbenson [dben...@dbenson.net](mailto:dben...@dbenson.net) wrote:

> We have been in production since November.
> 
> - How many servers, and roughly what's their hardware configuration?
> - 4 servers. Each 24 core, 48G ram, 4x600G 10k rpm drives
> - 32M total documents
> 
> - How many shards?
> - Depends on the index size, typically 1-5. (indexes only using meta  
> fields have 1, indexes using lots of full text extracted from PDFs get more)
> 
> - How many replicas?
> - 1, but we are adding more storage to increase this.
> 
> - Are they all on the same index?
> - We have ~50 indexes. Some for historical reasons, others for  
> operational efficiency. We've found that when an index gets very large, the  
> pain of handling incompatible field mappings, requiring a rebuild becomes a  
> motivation fr capping indexes.
> 
> - How many documents per index?
> - thousands to few million (depending on application)
> 
> - What kind of performance are you getting?
> - 35-50ms for 200k req/min.
> 
> - Anything else interesting 🙂
> - ES is amazing and Shay's support totally rocks!

---

<div class="post-metadata">

### Author: ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)
#### Post date: [August 4, 2011, 7:56pm UTC](https://discuss.elastic.co/t/usage-of-elastic-search-in-production-environment/5049/10 "2011-08-04T19:56:32Z")

</div>

Regarding the crashes, both Lucene itself and elasticsearch go through  
great effort to not corrupt the data in case of out of memory or other  
problems, like open file handles. I actually have several (non automated  
tests) that I run regularly that simulate those problems with no data loss.  
If you can help in trying to recreate what you saw, it would be great!

Regarding recovery of data, I have been thinking, and mentioned on the  
mailing list, of the ability to snapshot an index to a shared storage when  
using local gateway (basically, combine the shared gateway snapshot  
capabilities with local gateway). The main thing to note with this is the  
fact that it takes time to transfer large amount of data back to the nodes  
in case a full recovery is needed...

On Thu, Aug 4, 2011 at 2:21 AM, Tom Le [dottom@gmail.com](mailto:dottom@gmail.com) wrote:

> I am up to 200M documents inserted per day using HTTP REST API bulk  
> insertion with refresh settings from 5s to 20s.
> 
> The only problems I've had so far:
> 
> - Transaction logs take up a lot of space, which I believe 0.17.3 will  
> address that.
> - Some operational issues had catastrophic impact, such as not having  
> enough open files and the index getting corrupted.
> - Smaller java heap size that ran out of memory, when crashing, sometimes  
> corrupted the index. I tried to reproduce but corruption on crash doesn't  
> occur in all cases..
> - Index and search response can be slow for a few minutes when starting up  
> ES and you have really large indexes.
> 
> I would like a way to recover an index short of re-indexing, as 6 months  
> down the road, I do not want to re-index everything should something  
> unexpected happen.
> 
> On Wed, Aug 3, 2011 at 1:58 PM, Jay Taylor [outtatime@gmail.com](mailto:outtatime@gmail.com) wrote:
> 
> > I've been ingesting a few hundred GB per day now with ES without any major  
> > issues.

---

<div class="post-metadata">

### Author: ![dbenson](https://avatars.discourse-cdn.com/v4/letter/d/958977/32.png) [@dbenson](https://discuss.elastic.co/u/dbenson)
#### Post date: [August 5, 2011, 3:25pm UTC](https://discuss.elastic.co/t/usage-of-elastic-search-in-production-environment/5049/11 "2011-08-05T15:25:19Z")

</div>

We use ES just for searching and not as our document store. Our document  
store is home grown.

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 3:58am UTC](https://discuss.elastic.co/t/usage-of-elastic-search-in-production-environment/5049/12 "2017-07-06T03:58:10Z")

</div>


