# Configuration of elasticsearch to index 300 Million documents

**URL:** <https://discuss.elastic.co/t/configuration-of-elasticsearch-to-index-300-million-documents/21795>\
**Category:** Elasticsearch\
**Created:** [January 23, 2015, 11:37am UTC](https://discuss.elastic.co/t/configuration-of-elasticsearch-to-index-300-million-documents/21795 "2015-01-23T11:37:52Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![hitesh\_shekhada](https://avatars.discourse-cdn.com/v4/letter/h/3bc359/32.png) [@hitesh\_shekhada](https://discuss.elastic.co/u/hitesh_shekhada)\
**Post date:** [January 23, 2015, 11:37am UTC](https://discuss.elastic.co/t/configuration-of-elasticsearch-to-index-300-million-documents/21795/1 "2015-01-23T11:37:52Z")

</div>

Hi,

1. My goal is to index 300 million documents/products from MS SQL server  
database.
2. To get all documents I need to join 14 different tables.
3. Total data size of 300 million documents is 300GB
4. There 70 fields in one document
5. One document is of size 0.8 kb
6. Need to update value of 10 fields (out of 70) for almost 90 million  
documents every night.

I need to know ...

1. Does anybody has indexed such a large amount of data to elasticsearch  
server?
2. How many cluster/nodes I need for handling it.

Please let me know if someone has used elasticsearch for this amount of  
data.

Thanks,  
Hitesh

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/acc7bbd4-7abc-4451-bfd5-b21ab725e02e%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/acc7bbd4-7abc-4451-bfd5-b21ab725e02e%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 23, 2015, 10:24pm UTC](https://discuss.elastic.co/t/configuration-of-elasticsearch-to-index-300-million-documents/21795/2 "2015-01-23T22:24:55Z")

</div>

1 - Yes, that's not a lot of data for ES 🙂  
2 - Depends, you should be able to do that on a single node with ~16GB  
heap, but you should test yourself.

On 23 January 2015 at 22:37, hitesh shekhada [shekhada@gmail.com](mailto:shekhada@gmail.com) wrote:

> Hi,
> 
> 1. My goal is to index 300 million documents/products from MS SQL  
> server database.
> 2. To get all documents I need to join 14 different tables.
> 3. Total data size of 300 million documents is 300GB
> 4. There 70 fields in one document
> 5. One document is of size 0.8 kb
> 6. Need to update value of 10 fields (out of 70) for almost 90 million  
> documents every night.
> 
> I need to know ...
> 
> 1. Does anybody has indexed such a large amount of data to  
> elasticsearch server?
> 2. How many cluster/nodes I need for handling it.
> 
> Please let me know if someone has used elasticsearch for this amount of  
> data.
> 
> Thanks,  
> Hitesh
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/acc7bbd4-7abc-4451-bfd5-b21ab725e02e%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/acc7bbd4-7abc-4451-bfd5-b21ab725e02e%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/acc7bbd4-7abc-4451-bfd5-b21ab725e02e%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/acc7bbd4-7abc-4451-bfd5-b21ab725e02e%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAEYi1X-Y28Sf\_MnwDUb7iZp1MERxEWPehO4CRzmrSfBCHidumA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAEYi1X-Y28Sf_MnwDUb7iZp1MERxEWPehO4CRzmrSfBCHidumA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![vaniaravinda](https://avatars.discourse-cdn.com/v4/letter/v/7993a0/32.png) [@vaniaravinda](https://discuss.elastic.co/u/vaniaravinda)\
**Post date:** [July 7, 2015, 5:11am UTC](https://discuss.elastic.co/t/configuration-of-elasticsearch-to-index-300-million-documents/21795/3 "2015-07-07T05:11:25Z")

</div>

Hi @warkolm,

I want to do indexing on millions of records in SQL Server table. Can you please suggest me which approach I can follow and if there is any updates in the table how can I handle those records also.

Can you please suggest me as soon as possible. Your reply will be more helpful.

Thanks,  
Vani Aravinda

---

<div class="post-metadata">

**Author:** ![Jason\_Wee](https://avatars.discourse-cdn.com/v4/letter/j/7ea924/32.png) [@Jason\_Wee](https://discuss.elastic.co/u/Jason_Wee)\
**Post date:** [July 7, 2015, 7:35am UTC](https://discuss.elastic.co/t/configuration-of-elasticsearch-to-index-300-million-documents/21795/4 "2015-07-07T07:35:55Z")

</div>

concurred with warkolm statement, the requirements specified should be able to handle by elasticsearch. at least in the company i work for, we have five production nodes with 340m docs with index size of 863G.

hth

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:03am UTC](https://discuss.elastic.co/t/configuration-of-elasticsearch-to-index-300-million-documents/21795/5 "2017-07-06T00:03:18Z")

</div>


