# Feedback on tuning Data Streams for my use case

**URL:** <https://discuss.elastic.co/t/feedback-on-tuning-data-streams-for-my-use-case/260848>\
**Category:** Elasticsearch\
**Tags:** datastreams\
**Created:** [January 12, 2021, 3:21pm UTC](https://discuss.elastic.co/t/feedback-on-tuning-data-streams-for-my-use-case/260848 "2021-01-12T15:21:37Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![ismarslomic](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ismarslomic/32/63870_2.png) [@ismarslomic](https://discuss.elastic.co/u/ismarslomic)\
**Post date:** [January 12, 2021, 3:21pm UTC](https://discuss.elastic.co/t/feedback-on-tuning-data-streams-for-my-use-case/260848/1 "2021-01-12T15:21:37Z")

</div>

I'm working on a use where we need to copy messages from 2 Kafka topics to Elasticsearch, in order to search and visualize data in Web App.

I have setup Elastic Cloud on Kubernetes with Elasticsearch and Kibana, and performed quick PoC by using ordinary index without aliases and ILM.

Now I want to take the PoC further and create production ready Elastic stack and since Im working with time series data, I thought Data Streams would be a good fit for my use case.

I have created an overview of my current cluster setup and characteristics for Document A and B in diagram below. Document B contains GPS position + some metadata for an vehicle at given point of time, Document A contains trip data for vehicle and will be used in search field. When trip is selected map is visualizing all GPS positions for vehicle by retrieving all related Document B (based on a ref value).

I would like feedback on following questions:

- Since charasteristics and payload of Document A and B are different Im thinking on separating those on two different Data Streams/Indices. I guess you agree?
- How many shards should I have for Data Streams for Document A and B?
- Given data retention for last 30 days, should I have rollover pattern per day (one index per day) for both Data Stream or use different approach for each one?
- Any input on other cluster related settings are also welcome!

The Web app will be used by limited amount of users (2-5 users). Replication is set to 0 since we can afford loosing data if one Node gets broken.

 ![Screenshot 2021-01-12 at 16.12.37](https://us1.discourse-cdn.com/elastic/original/3X/1/6/165a8f41cec1a7dd89fb5da0f44bc5417401f87e.png)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 9, 2021, 3:21pm UTC](https://discuss.elastic.co/t/feedback-on-tuning-data-streams-for-my-use-case/260848/2 "2021-02-09T15:21:42Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
