# Help for select the best solution for data analyse

**URL:** <https://discuss.elastic.co/t/help-for-select-the-best-solution-for-data-analyse/50948>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [May 25, 2016, 2:38pm UTC](https://discuss.elastic.co/t/help-for-select-the-best-solution-for-data-analyse/50948 "2016-05-25T14:38:50Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Vincent\_Vost](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vincent_vost/32/9946_2.png) [@Vincent\_Vost](https://discuss.elastic.co/u/Vincent_Vost)\
**Post date:** [May 25, 2016, 2:38pm UTC](https://discuss.elastic.co/t/help-for-select-the-best-solution-for-data-analyse/50948/1 "2016-05-25T14:38:51Z")

</div>

Hi 🙂

I have an open question.  
I'm a student and I work in a projet of anomalies detection we can say.

I have choose to store my data in Elasticsearch. Currently I have 5/6 weeks of data ready.  
But now I have a big problem, how analyse this data ?

My problem is connected with anomalies detection in time, like something happened during the night...  
For this, my idea is to learn a pattern of times of my log. So maybe machine learning is the best in my case. (I never used)

I found too much possibility in google.  
We have Spark with Mlib which provide machine learning,  
We have Hadoop, DeepDetect etc etc....

If you have any experience with this problem, maybe you can give me some advices.  
Which solution choose ?

I try to use the JAVA API of Elasticsearch, but for the moment I failled.  
thank for all reply who can help me a little

vincent

---

<div class="post-metadata">

**Author:** ![ramaguruprasad1](https://avatars.discourse-cdn.com/v4/letter/r/977dab/32.png) [@ramaguruprasad1](https://discuss.elastic.co/u/ramaguruprasad1)\
**Post date:** [June 22, 2016, 4:57am UTC](https://discuss.elastic.co/t/help-for-select-the-best-solution-for-data-analyse/50948/2 "2016-06-22T04:57:23Z")

</div>

Hello vincent,  
For the Anomoly detection problem you may want to look at the following link where a combinatorial approach for handling different native data is taken. The logical building blocks are here:

> **[Building a Streaming Data Hub with Elasticsearch, Kafka and Cassandra - The New...](https://thenewstack.io/building-streaming-data-hub-elasticsearch-kafka-cassandra/)**
>
> Over the past year or so, I’ve met a handful of software companies to discuss dealing with the data that pours out of their software (typically in the form of logs and metrics). In these discussions, I often hear frustration with having to get by...

Regards,  
Guruprasad

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:24pm UTC](https://discuss.elastic.co/t/help-for-select-the-best-solution-for-data-analyse/50948/3 "2017-07-06T13:24:14Z")

</div>


