# In Elasticsearch, is possible to cluster documents that share the most similar texts, without giving an initial query to compare to?

**URL:** <https://discuss.elastic.co/t/in-elasticsearch-is-possible-to-cluster-documents-that-share-the-most-similar-texts-without-giving-an-initial-query-to-compare-to/90962>\
**Category:** Elasticsearch\
**Created:** [June 27, 2017, 12:57pm UTC](https://discuss.elastic.co/t/in-elasticsearch-is-possible-to-cluster-documents-that-share-the-most-similar-texts-without-giving-an-initial-query-to-compare-to/90962 "2017-06-27T12:57:17Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![pachilo](https://avatars.discourse-cdn.com/v4/letter/p/b487fb/32.png) [@pachilo](https://discuss.elastic.co/u/pachilo)\
**Post date:** [June 27, 2017, 12:57pm UTC](https://discuss.elastic.co/t/in-elasticsearch-is-possible-to-cluster-documents-that-share-the-most-similar-texts-without-giving-an-initial-query-to-compare-to/90962/1 "2017-06-27T12:57:17Z")

</div>

In Elasticsearch, is possible to group documents that share the most similar texts, without giving an initial query to compare to?

I know is possible to query and get "more like this document" but, is possible to cluster documents within an index according to a field values?

For instance:

document 1: **The quick brown fox jumps over the lazy dog**

document 2: **Barcelona is a great city**

document 3: **The fast orange fox jumps over the lazy dog**

document 4: **Madrid is a great city**

document 5: **I do not like to eat fish**

Now, perform some kind of aggregation that, without giving a search query, it can group:

**Group 1** : document 1 and document 3

**Group 2** : document 2 and document 4

**Group 3** : document 5

I will really appreciate any clue!

---

<div class="post-metadata">

**Author:** ![colings86](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/colings86/32/44960_2.png) [@colings86](https://discuss.elastic.co/u/colings86)\
**Post date:** [June 27, 2017, 2:08pm UTC](https://discuss.elastic.co/t/in-elasticsearch-is-possible-to-cluster-documents-that-share-the-most-similar-texts-without-giving-an-initial-query-to-compare-to/90962/2 "2017-06-27T14:08:11Z")

</div>

There is not currently an aggregation which performs clustering. There is an issue for adding k-means clustering as an aggregation ([https://github.com/elastic/elasticsearch/issues/5512](https://github.com/elastic/elasticsearch/issues/5512)) and I played around with a prototype for this a while ago but there are some changes that would need to be made to the aggregations framework itself to support this kind of aggregation and that work is yet to be done.

---

<div class="post-metadata">

**Author:** ![pachilo](https://avatars.discourse-cdn.com/v4/letter/p/b487fb/32.png) [@pachilo](https://discuss.elastic.co/u/pachilo)\
**Post date:** [June 27, 2017, 2:18pm UTC](https://discuss.elastic.co/t/in-elasticsearch-is-possible-to-cluster-documents-that-share-the-most-similar-texts-without-giving-an-initial-query-to-compare-to/90962/3 "2017-06-27T14:18:15Z")

</div>

Ohhh, good and sad to know then, thanks a lot Colin.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 25, 2017, 2:18pm UTC](https://discuss.elastic.co/t/in-elasticsearch-is-possible-to-cluster-documents-that-share-the-most-similar-texts-without-giving-an-initial-query-to-compare-to/90962/4 "2017-07-25T14:18:20Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
