# Implementing random forest in ElasticSearch

**URL:** <https://discuss.elastic.co/t/implementing-random-forest-in-elasticsearch/46845>\
**Category:** Elasticsearch\
**Created:** [April 8, 2016, 8:08pm UTC](https://discuss.elastic.co/t/implementing-random-forest-in-elasticsearch/46845 "2016-04-08T20:08:20Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![paulshadepanda](https://avatars.discourse-cdn.com/v4/letter/p/bbe5ce/32.png) [@paulshadepanda](https://discuss.elastic.co/u/paulshadepanda)\
**Post date:** [April 8, 2016, 8:08pm UTC](https://discuss.elastic.co/t/implementing-random-forest-in-elasticsearch/46845/1 "2016-04-08T20:08:20Z")

</div>

Hi All,

I am new to ElasticSearch, so any input and suggestion is much appreciated. I have built a random forest classification model (binary class) through scikit learn, and I want to use the probability to predict class 1 to sort the documents (rfc.predict\_proba(X\_test) from sklearn) in ElasticSearch (version 1.7). The features I use are from both the documents and some user inputs (query). What I am currently doing is:

1. Use function score to filter out the documents that meet the query;
2. Run a script\_score (implemented the random forest classifier model in a Groovy script) as the score to sort the filtered documents from 1.  
The problem I am running into is that even when I am only implementing one tree my script already exceeds the java single method limit (64k). The tree has 50 layers and about 15000 nodes. Like I mentioned at the beginning, I am still a newbie. This probably is not the best way to implement things. I am wondering what is a better way, considering the feasibility and performance (speed).

The whole dataset is 25G in size and has 10M documents. After filtering, it's about 30K documents.

Thanks a lot,  
Wei

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [April 8, 2016, 10:13pm UTC](https://discuss.elastic.co/t/implementing-random-forest-in-elasticsearch/46845/2 "2016-04-08T22:13:48Z")

</div>

Forget about script\_score, this is not for classifiers. You have to write a plugin for Elasticsearch that can execute random forest classifier.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:01pm UTC](https://discuss.elastic.co/t/implementing-random-forest-in-elasticsearch/46845/3 "2017-07-05T23:01:04Z")

</div>


