# Cluster configuration for log storage. 140Gb/day

**URL:** <https://discuss.elastic.co/t/cluster-configuration-for-log-storage-140gb-day/103886>\
**Category:** Elasticsearch\
**Created:** [October 13, 2017, 11:52am UTC](https://discuss.elastic.co/t/cluster-configuration-for-log-storage-140gb-day/103886 "2017-10-13T11:52:23Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 13, 2017, 12:34pm UTC](https://discuss.elastic.co/t/cluster-configuration-for-log-storage-140gb-day/103886/2 "2017-10-13T12:34:10Z")

</div>

Please format your code using `</>` icon as explained in [this guide](https://discuss.elastic.co/t/about-the-elasticsearch-category/21). It will make your post more readable.

Or use markdown style like:

````
```
CODE
```

````

Some thoughts:

- What is exactly the full query? I mean is the `query_string` part inside a `query` or a `filter`?
- `analyze_wildcard`: do you really intend to run queries like `foo*bar`? As per doc says it's super slow.
- Do you really want to compute a bucket for every 3 hours but for the full 6 days? Don't you want to add a filter by date and just look at the last 24 hours for example?

What are the index settings? How many shards per day?

Also using `_exists_:aggregate_final` is going to most likely in your use case give back all the documents. So you compute an aggregation on 3.5 billion docs most likely + the cost of running the query which could be faster with a `match_all`.

One thing you can do is to run a query filtered per day and compute the agg only for that day. Then use a multisearch query to run 5 of them in parallel.

> Can SSD help me?

Yes.

> Should I check mapping because index size is 5 times bigger in raw size?

Yes. Remove `_all`, remove non needed `keyword` fields, non needed `text` fields.

If you are planning to query often on the existence of `aggregate_final` field, may be you should simply index that value as a boolean and filter by that.

Just some thoughts.

---

_[View the full topic](https://discuss.elastic.co/t/cluster-configuration-for-log-storage-140gb-day/103886)._
