# One vs Multiple indices on ES using LS's 'date' filter

**URL:** <https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939>\
**Category:** Logstash\
**Created:** [June 4, 2015, 2:29pm UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939 "2015-06-04T14:29:15Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![refaelos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/refaelos/32/516_2.png) [@refaelos](https://discuss.elastic.co/u/refaelos)\
**Post date:** [June 4, 2015, 2:29pm UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/1 "2015-06-04T14:29:15Z")

</div>

Hey Guys,

Love the elastic team products 👍

I thought if this question should go here or on the elasticsearch forum and only b/c it's a logstash configuration I decided to put it here.

What i'm doing is transferring s3 files' data containing a huge amount of JSON logs into ES. There's something I want to ask about the 'date' filter of LS and how it affects the number of indices and if it really matters.

When I'm doing this transfer of data (using the 's3' input) without the 'date' filter everything is done really fast and there's only one index created in ES for the current day: 'logstash-2015.06.04'.

The second time I decided to run with the 'date' filter and multiple indices were created and the load on the machine sky rocketed, which even causes a loss of data from time to time b/c ES became unresponsive.  
The 'date' filter i'm using is:

```auto
  date {
    match => ["server_time", "UNIX"]
  }

```

1. Is it expected that when I add the 'date' filter the load on the machine become this heavy and so many indices will be created ?
2. What's the difference between having one index and multiple indices created in ES ?

Thanks!!

Some info:

- I'm running on an EC2 r3.xlarge machine with 32g RAM.
- I have set 16g for the ES\_HEAP\_SIZE env variable

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [June 4, 2015, 4:33pm UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/2 "2015-06-04T16:33:38Z")

</div>

Are you creating hundreds of indices with the first few events that you're logging?

---

<div class="post-metadata">

**Author:** ![refaelos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/refaelos/32/516_2.png) [@refaelos](https://discuss.elastic.co/u/refaelos)\
**Post date:** [June 4, 2015, 5:30pm UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/3 "2015-06-04T17:30:03Z")

</div>

Actually i'm not creating anything manually. When I put the 'date' filter those indices are just created.  
And yes, hundreds of them are created for even less than 50K events. In Marvel they all appear in red and ES becomes unresponsive.

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [June 4, 2015, 6:24pm UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/4 "2015-06-04T18:24:49Z")

</div>

What does your Logstash output block look like? Specifically, the elasticsearch portion.

---

<div class="post-metadata">

**Author:** ![refaelos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/refaelos/32/516_2.png) [@refaelos](https://discuss.elastic.co/u/refaelos)\
**Post date:** [June 5, 2015, 5:07am UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/5 "2015-06-05T05:07:39Z")

</div>

The output block:

```auto
output {  
  elasticsearch {
    host => localhost
    protocol => http
    idle_flush_time => 100
    flush_size => 2000
  }   
} 

```

I tried different flush\_sizes. When I set a really small value (e.g.10) it works fine but the process is extremely slow. When I set higher values (\>= 500) the same scenario I described before happens.

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [June 5, 2015, 6:55pm UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/6 "2015-06-05T18:55:17Z")

</div>

What are some of the index names? If you're seeing lots of indices created, it implies that each one has a unique date, perhaps in the future or past. Is it possible that instead of UNIX (whole seconds since the epoch) that you have UNIX\_MS or nanoseconds or something like that? It would definitely make the date filter project indices into the distant future if you were sending higher precision timestamps to a "whole second" parser.

---

<div class="post-metadata">

**Author:** ![refaelos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/refaelos/32/516_2.png) [@refaelos](https://discuss.elastic.co/u/refaelos)\
**Post date:** [June 5, 2015, 10:04pm UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/7 "2015-06-05T22:04:44Z")

</div>

You're right. The indices i'm getting are weird.

I'll test and ket you know.

---

<div class="post-metadata">

**Author:** ![refaelos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/refaelos/32/516_2.png) [@refaelos](https://discuss.elastic.co/u/refaelos)\
**Post date:** [June 6, 2015, 9:22am UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/8 "2015-06-06T09:22:15Z")

</div>

Ok so running for around 10 hours now (I guess it's going to take a few more days), it looks better than before. Indices are generated for actual dates and the machine is better on responsiveness although it's still very highly loaded.

 ![](https://us1.discourse-cdn.com/elastic/original/1X/8de33ff0feb0066f3da287a6b73df57be0583075.png)

On my 4-cpu machine this load average is pretty heavy. When I ran the same think without the 'date' filter the load was a lot lower (probably b/c of just one index that was created).

Can you explain why the 'date' filter creates an index per day?  
Is it a better practice than having just one index created?  
If yes, how does it affect ES performance?

Thanks for your help so far.

UPDATE:  
Around the time I wrote the last message things started to break. The machine started becoming unresponsive and data insertion to ES became slow. Up until now it's like that and it's getting slower and slower...  
Ideas?

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [June 8, 2015, 5:15pm UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/9 "2015-06-08T17:15:45Z")

</div>

Time series data should be logically grouped into common blocks of time, e.g. days, weeks, months, etc. This simplifies searching and data retention.

As far as your performance issues go, how much memory have you allocated for heap in Elasticsearch? Logstash?

How many nodes do you have? How many events per second?

I can tell by looking at your "JVM Heap Usage" chart that your cluster is overwhelmed. That solid line pegged at 75% usage indicates that your system (which I am guessing only has one node) is constantly performing garbage collections and is never getting ahead. That line should look like a sawtooth pattern, but it's just a flat line pegged at 70%+. Your indexing load seems to indicate the need to grow your cluster to a few more nodes, and/or bigger heap set aside for Elasticsearch.

---

<div class="post-metadata">

**Author:** ![refaelos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/refaelos/32/516_2.png) [@refaelos](https://discuss.elastic.co/u/refaelos)\
**Post date:** [June 9, 2015, 11:09am UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/10 "2015-06-09T11:09:30Z")

</div>

Memory allocation:  
ES - 16g  
LS - 1g

1 node that runs both ES and LS.

Events per second - i'm working to load historical events from S3. There are loads of events in files as JSON. So I guess it's being delivered from s3 pretty fast.

I can increase the heap size but Marvel shows that it never really got to 16g. I seriously don't know what to do more ...

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [June 9, 2015, 3:43pm UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/11 "2015-06-09T15:43:04Z")

</div>

You need more nodes in your Elasticsearch cluster. A single node is not able to handle the load you're placing on it.

---

<div class="post-metadata">

**Author:** ![refaelos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/refaelos/32/516_2.png) [@refaelos](https://discuss.elastic.co/u/refaelos)\
**Post date:** [June 9, 2015, 3:58pm UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/12 "2015-06-09T15:58:54Z")

</div>

Now I see a lot of errors in logstash like this:  
{:timestamp=\>"2015-06-09T04:25:06.156000+0000", :message=\>"A plugin had an unrecoverable error. Will restart this plugin.\n Plugin: \<LogStash::Inputs::S3 bucket=\>"\*\*\*", access\_key\_id=\>"\*\*\*", secret\_access\_key=\>"\*\*\*", region=\>"us-east-1", prefix=\>"\*\*\*", sincedb\_path=\>"/var/lib/logstash/.sincedb\_s3", temporary\_directory=\>"/var/lib/logstash/logstash"\>\n Error: Invalid UTF-32 character 0x7b226e61(above 10ffff) at char #199, byte #799)", :level=\>:error}

Not sure if it's related ... it's probably what you said. I'll try adding another cluster. Is it possible to just add memory ?

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [June 9, 2015, 4:23pm UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/13 "2015-06-09T16:23:15Z")

</div>

It is possible to add memory. The maximum recommended heap for an Elasticsearch node is 30.5G. We recommend that this number be no more than 1/2 of the available system memory, which implies a 64G server to use a 30.5G heap.

Not sure what the character error is.

[EDIT] Changed 31G to 30.5G based on most recent information.

---

<div class="post-metadata">

**Author:** ![refaelos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/refaelos/32/516_2.png) [@refaelos](https://discuss.elastic.co/u/refaelos)\
**Post date:** [June 10, 2015, 12:15pm UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/14 "2015-06-10T12:15:01Z")

</div>

Thanks man!

Appreciate your help. I'll investigate and will let you know if I need anything else.

---

<div class="post-metadata">

**Author:** ![zkun](https://avatars.discourse-cdn.com/v4/letter/z/8797f3/32.png) [@zkun](https://discuss.elastic.co/u/zkun)\
**Post date:** [July 21, 2015, 10:10am UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/15 "2015-07-21T10:10:04Z")

</div>

Hey I saw the same problem with refaelos about "Invalid UTF-32" stuff.  
My scenario here is: there's a py script appending to a log file, while LS has been configured to get fresh lines from it. The file was placed on NFS mountpoint, which means both py script and LS are actually interacting with this file remotely.  
This error takes place randomly and I really appreciate you guys can tip about what's happening \>  
As for memory, LS runs on the machine without ES installed and there's still \>50% left. So I presume that would not cause trouble.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 5:34am UTC](https://discuss.elastic.co/t/one-vs-multiple-indices-on-es-using-lss-date-filter/1939/16 "2017-07-06T05:34:08Z")

</div>


