# Supervised machine learning plugins or tools for elasticsearch?

**URL:** <https://discuss.elastic.co/t/supervised-machine-learning-plugins-or-tools-for-elasticsearch/59501>\
**Category:** Elasticsearch\
**Created:** [September 1, 2016, 7:37am UTC](https://discuss.elastic.co/t/supervised-machine-learning-plugins-or-tools-for-elasticsearch/59501 "2016-09-01T07:37:18Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Shrikanth\_N\_C](https://avatars.discourse-cdn.com/v4/letter/s/919ad9/32.png) [@Shrikanth\_N\_C](https://discuss.elastic.co/u/Shrikanth_N_C)\
**Post date:** [September 1, 2016, 7:37am UTC](https://discuss.elastic.co/t/supervised-machine-learning-plugins-or-tools-for-elasticsearch/59501/1 "2016-09-01T07:37:18Z")

</div>

Similar to carrot2 plugin are there any plugins or tools on top of elk, that can be used to train and execute supervised learning algorithms ? ( preferably java )

I want to classify search results based on training data.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [September 1, 2016, 9:22am UTC](https://discuss.elastic.co/t/supervised-machine-learning-plugins-or-tools-for-elasticsearch/59501/2 "2016-09-01T09:22:19Z")

</div>

The `significant_terms` aggregation can be used to extract features given a set of representative data - see [https://www.elastic.co/blog/significant-terms-aggregation#classifier](https://www.elastic.co/blog/significant-terms-aggregation#classifier)

---

<div class="post-metadata">

**Author:** ![mainec](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mainec/32/5557_2.png) [@mainec](https://discuss.elastic.co/u/mainec)\
**Post date:** [September 1, 2016, 9:22am UTC](https://discuss.elastic.co/t/supervised-machine-learning-plugins-or-tools-for-elasticsearch/59501/3 "2016-09-01T09:22:51Z")

</div>

Can you elaborate a bit on your use case? Why do you need to classify search results, would it also be ok to classify documents prior to ingestion?

What kind of classification algorithm would you like to apply to your data?

I'm not aware of any officially supported plugins. A quick search turned up the following results:

> **[pandastrike/bayzee](https://github.com/pandastrike/bayzee)**
>
> bayzee - Text classification using Naive Bayes and Elasticsearch

> **[Categorizing images with deep learning into Elasticsearch](https://www.elastic.co/blog/categorizing-images-with-deep-learning-into-elasticsearch)**
>
> Deepdetect, an open source deep-learning server working, created a range of machine learning applications, including an image search engine using Elasticsearch.

> **[polyfractal/ElasticBayes](https://github.com/polyfractal/ElasticBayes)**
>
> ElasticBayes - Naive Bayes Classifier implemented with Elasticsearch Aggregations

Hope this helps and looking forward to hear more on your use case,  
Isabel

---

<div class="post-metadata">

**Author:** ![Shrikanth\_N\_C](https://avatars.discourse-cdn.com/v4/letter/s/919ad9/32.png) [@Shrikanth\_N\_C](https://discuss.elastic.co/u/Shrikanth_N_C)\
**Post date:** [September 1, 2016, 9:42am UTC](https://discuss.elastic.co/t/supervised-machine-learning-plugins-or-tools-for-elasticsearch/59501/4 "2016-09-01T09:42:17Z")

</div>

Thanks for the quick reply, I am going through your nicely written post. Just want to clarify, I don't want to do semantic analysis. In other words, I will input/train to correlate apples with zebras. As in, if I query for 'apples' elasticsearch should return 'zebras' based on my training and statistics. Can you suggest ?

---

<div class="post-metadata">

**Author:** ![Shrikanth\_N\_C](https://avatars.discourse-cdn.com/v4/letter/s/919ad9/32.png) [@Shrikanth\_N\_C](https://discuss.elastic.co/u/Shrikanth_N_C)\
**Post date:** [September 1, 2016, 9:55am UTC](https://discuss.elastic.co/t/supervised-machine-learning-plugins-or-tools-for-elasticsearch/59501/5 "2016-09-01T09:55:25Z")

</div>

Thanks for the spontaneous reply. It would not be ok to classify prior ingestion, because the results may evolve while using machine learning based on history.

Any classification algorithm ( decision tree ) , essentially I am interested to see ML supervised learning support using elasticsearch.

Thank you for the links I came across them too!

---

<div class="post-metadata">

**Author:** ![mainec](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mainec/32/5557_2.png) [@mainec](https://discuss.elastic.co/u/mainec)\
**Post date:** [September 1, 2016, 10:24am UTC](https://discuss.elastic.co/t/supervised-machine-learning-plugins-or-tools-for-elasticsearch/59501/6 "2016-09-01T10:24:12Z")

</div>

> [@Shrikanth\_N\_C](#):
>
> It would not be ok to classify prior ingestion, because the results may evolve while using machine learning based on history.

What do you mean - the results may evolve while using machine learning based on history?

Do you mean the results may change as soon as you re-train your classification model? How often would you re-train your model then?

> [@Shrikanth\_N\_C](#):
>
> As in, if I query for 'apples' elasticsearch should return 'zebras' based on my training and statistics

This sounds like a typical use case for synonyms to me.

> **[Synonym token filter | Elasticsearch Guide \[8.11\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-synonym-tokenfilter.html)**

should have more details on this.

Hope this helps,  
Isabel

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [September 1, 2016, 10:37am UTC](https://discuss.elastic.co/t/supervised-machine-learning-plugins-or-tools-for-elasticsearch/59501/7 "2016-09-01T10:37:09Z")

</div>

> [@mainec](#):
>
> As in, if I query for 'apples' elasticsearch should return 'zebras' based on my training and statistics

Depending on your user volumes It might be worth considering on-the-fly analysis rather than pre-computing a limited set of trained responses.  
If the query was something you might not have predicted e.g. not just `apples` but `apples new AR headset` you can run significant terms on the best-matching docs and discover the term `iGizmo` (or whatever the product name might be). Doing the analysis on the fly tailors suggestions to the long tail of many and varied queries users provide rather than just the "head of the tail" ones you were able to predict in your training data.

---

<div class="post-metadata">

**Author:** ![Shrikanth\_N\_C](https://avatars.discourse-cdn.com/v4/letter/s/919ad9/32.png) [@Shrikanth\_N\_C](https://discuss.elastic.co/u/Shrikanth_N_C)\
**Post date:** [September 2, 2016, 5:10am UTC](https://discuss.elastic.co/t/supervised-machine-learning-plugins-or-tools-for-elasticsearch/59501/8 "2016-09-02T05:10:25Z")

</div>

Thanks Mark will think on those lines as well.

---

<div class="post-metadata">

**Author:** ![Shrikanth\_N\_C](https://avatars.discourse-cdn.com/v4/letter/s/919ad9/32.png) [@Shrikanth\_N\_C](https://discuss.elastic.co/u/Shrikanth_N_C)\
**Post date:** [September 2, 2016, 5:11am UTC](https://discuss.elastic.co/t/supervised-machine-learning-plugins-or-tools-for-elasticsearch/59501/9 "2016-09-02T05:11:08Z")

</div>

Synonym might work, will think on those lines. Thank you for your time!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:23pm UTC](https://discuss.elastic.co/t/supervised-machine-learning-plugins-or-tools-for-elasticsearch/59501/10 "2017-07-05T22:23:01Z")

</div>


