# Time taken by ELK stack to create indices

**URL:** <https://discuss.elastic.co/t/time-taken-by-elk-stack-to-create-indices/307836>\
**Category:** Elasticsearch\
**Tags:** docker\
**Created:** [June 22, 2022, 6:48am UTC](https://discuss.elastic.co/t/time-taken-by-elk-stack-to-create-indices/307836 "2022-06-22T06:48:50Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Divyank\_Garg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/divyank_garg/32/99581_2.png) [@Divyank\_Garg](https://discuss.elastic.co/u/Divyank_Garg)\
**Post date:** [June 22, 2022, 6:48am UTC](https://discuss.elastic.co/t/time-taken-by-elk-stack-to-create-indices/307836/1 "2022-06-22T06:48:50Z")

</div>

How much time does it take to create indices using ELK services via docker-compose file?  
We process total 450 files of total size of 4.6GB to create 450 indices. All these files are basically the log files. It takes around 8 hours to ingest the data from directory to ELK stack indices creation.  
The server we are using is 16Gb RAM with 2T hard disk but still it is taking a lot of time to create indices.  
Is this figure correct? Does ELK takes time to create indices? How much time thus 1 GB takes to create indices? How it can be optimised?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 22, 2022, 7:14am UTC](https://discuss.elastic.co/t/time-taken-by-elk-stack-to-create-indices/307836/2 "2022-06-22T07:14:58Z")

</div>

> [@Divyank\_Garg](#):
>
> We process total 450 files of total size of 4.6GB to create 450 indices.

Creating an index per file does not seem optimal. Generally indices are create by file types instead as having lots of small indicxes and shards is inefficient and can cause problems down the line.

> [@Divyank\_Garg](#):
>
> It takes around 8 hours to ingest the data from directory to ELK stack indices creation.

How are you ingesting the data? Are you using Logstash or Filebeat?

In order to optimise indexing performance I would recommend looking into [this official guide](https://www.elastic.co/guide/en/elasticsearch/reference/8.2/tune-for-indexing-speed.html).

---

<div class="post-metadata">

**Author:** ![Divyank\_Garg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/divyank_garg/32/99581_2.png) [@Divyank\_Garg](https://discuss.elastic.co/u/Divyank_Garg)\
**Post date:** [June 24, 2022, 5:48am UTC](https://discuss.elastic.co/t/time-taken-by-elk-stack-to-create-indices/307836/3 "2022-06-24T05:48:24Z")

</div>

Hi Christion,  
Thanks for reply.  
450 files means 450 different directories having different type of log types in it and these log types has log files. We are making ingesting using Logstash.  
We allocated more memory to these docker containers but still total 15Gb of these 450 directories taking 6-7 hrs of time for data ingestion.  
Best way to optimise it?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 24, 2022, 5:59am UTC](https://discuss.elastic.co/t/time-taken-by-elk-stack-to-create-indices/307836/4 "2022-06-24T05:59:09Z")

</div>

Look at the link I provided in my previous response. This is a great starting point.

> [@Divyank\_Garg](#):
>
> 450 files means 450 different directories having different type of log types in it and these log types has log files.

I am not sure I understand this. Typically indices are designed to hold data of specific types and not different types of data from a certain location. The only reason I could see for the approach you have taken is if each directory belongs to a different user and you want to control access at the index level. Is that the case?

---

<div class="post-metadata">

**Author:** ![Divyank\_Garg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/divyank_garg/32/99581_2.png) [@Divyank\_Garg](https://discuss.elastic.co/u/Divyank_Garg)\
**Post date:** [June 24, 2022, 7:18am UTC](https://discuss.elastic.co/t/time-taken-by-elk-stack-to-create-indices/307836/5 "2022-06-24T07:18:12Z")

</div>

Yes I am looking to the link.  
450 directories are basically 450 different machines and each having similar kind of log files in it. Just all these machines are indexed separately to search query the files inside the log files and find some patterns present in each kind of machines.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 24, 2022, 7:51am UTC](https://discuss.elastic.co/t/time-taken-by-elk-stack-to-create-indices/307836/6 "2022-06-24T07:51:24Z")

</div>

> [@Divyank\_Garg](#):
>
> Just all these machines are indexed separately to search query the files inside the log files and find some patterns present in each kind of machines.

That does not make sense to me. I would recommend you instead index into a single index and add data about the host or directory onto the events so that you easily can filter on it in Kibana or queries. Indexing into a large number of indices will be inefficient and slow down ingestion. The thing that makes bulk indexing fast is that groups of documents are sent to each shard and indexed together, reducing overhead per document. If a bulk request contains documents for a very large number of shards the groups of documents sent to each shard is likely to be small and you get much more overhead per document.

I have seen users collect data from thousands of machines into a single index and then filter the way I explained. Your approach does not scale and may not perform very well either as it may result in a lot of small indices.

---

<div class="post-metadata">

**Author:** ![Divyank\_Garg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/divyank_garg/32/99581_2.png) [@Divyank\_Garg](https://discuss.elastic.co/u/Divyank_Garg)\
**Post date:** [June 24, 2022, 10:07am UTC](https://discuss.elastic.co/t/time-taken-by-elk-stack-to-create-indices/307836/7 "2022-06-24T10:07:26Z")

</div>

The shape of our data structure:  
Directories A, B ...116k with unrelated to one another

1. A  
a) RS log  
b) ME log  
c) CH log
2. B  
a) RS log  
b) ME log  
c) CH log
3. ...go on till 16k  
We are doing indexing for each directories and making indices as A, B, C, .. 16K

I believe you are telling to create one index only and put all log files into different events within that same index. Can you send some link for this kind of approach. Any example or link to look upon and implement.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 24, 2022, 10:36am UTC](https://discuss.elastic.co/t/time-taken-by-elk-stack-to-create-indices/307836/8 "2022-06-24T10:36:13Z")

</div>

That is the standard/default approach and would be the default configuration e.g. if you used filebeat to directly index into Elasticsearch. As it is the default I have not found anything in the standard docs that explicitly describe it. [This blog post](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster) contains some explicit guidelines that are largely still applicable even though the post is old and a lot of improvements have been made in recent versions, e.g.:

> _ **TIP:** In order to reduce the number of indices and avoid large and sprawling mappings, consider storing data with similar structure in the same index rather than splitting into separate indices based on where the data comes from._

[This section in the docs](https://www.elastic.co/guide/en/elasticsearch/reference/8.2/size-your-shards.html) also contains some guidelines.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 22, 2022, 10:36am UTC](https://discuss.elastic.co/t/time-taken-by-elk-stack-to-create-indices/307836/9 "2022-07-22T10:36:57Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
