# Term vectors for computing document similarity

**URL:** <https://discuss.elastic.co/t/term-vectors-for-computing-document-similarity/10311>\
**Category:** Elasticsearch\
**Created:** [January 11, 2013, 7:37am UTC](https://discuss.elastic.co/t/term-vectors-for-computing-document-similarity/10311 "2013-01-11T07:37:28Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Aditya\_Rajgarhia](https://avatars.discourse-cdn.com/v4/letter/a/dec6dc/32.png) [@Aditya\_Rajgarhia](https://discuss.elastic.co/u/Aditya_Rajgarhia)\
**Post date:** [January 11, 2013, 7:37am UTC](https://discuss.elastic.co/t/term-vectors-for-computing-document-similarity/10311/1 "2013-01-11T07:37:28Z")

</div>

Hello,

I'm trying to build some search features for a website. I don't have any  
prior experience with search and decided to go with elasticsearch mostly  
because of the ease of use.

I am indexing two types of documents (each with it's own index) since I  
want to offer search functionality for either type of document. I have this  
part working.

Now, I also want to offer a feature for comparing documents from one index  
with those from the other. What I had in mind was that since ES uses  
Lucene, I could fetch the term vectors for a pair of documents and then  
compute the cosine similarity (as explained in  
[http://sujitpal.blogspot.in/2011/10/computing-document-similarity-using.html](http://sujitpal.blogspot.in/2011/10/computing-document-similarity-using.html)).

However, from what I can tell ES doesn't expose the term vectors. Is it  
still possible for me to use ES if I absolutely need the above feature? Is  
it possible to read the Lucene index generated by ES directly without too  
much trouble?

Of course, I could always generate the term vector dynamically for each  
document for the purpose of implementing this particular feature, but  
that's inefficient (I will be performing a large number of such  
comparisons) and I don't want to do that if there is an alternate -- solr  
seems to allow fetching the term vectors ☹

Any help would be appreciated!

Thanks,  
Aditya

--

---

<div class="post-metadata">

**Author:** ![Loic\_Bertron](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/loic_bertron/32/1794_2.png) [@Loic\_Bertron](https://discuss.elastic.co/u/Loic_Bertron)\
**Post date:** [January 11, 2013, 7:44pm UTC](https://discuss.elastic.co/t/term-vectors-for-computing-document-similarity/10311/2 "2013-01-11T19:44:26Z")

</div>

Hello,

You should have a look at this feature : Fuzzy Like this and More like  
this:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

If you compare your text field to all the text fields of all others  
documents, you can reduce results only to documents matching 95% and more.

Le vendredi 11 janvier 2013 02:37:28 UTC-5, [adi...@blobinfotech.com](mailto:adi...@blobinfotech.com) a  
écrit :

> Hello,
> 
> I'm trying to build some search features for a website. I don't have any  
> prior experience with search and decided to go with elasticsearch mostly  
> because of the ease of use.
> 
> I am indexing two types of documents (each with it's own index) since I  
> want to offer search functionality for either type of document. I have this  
> part working.
> 
> Now, I also want to offer a feature for comparing documents from one index  
> with those from the other. What I had in mind was that since ES uses  
> Lucene, I could fetch the term vectors for a pair of documents and then  
> compute the cosine similarity (as explained in  
> [Salmon Run: Computing Document Similarity using Lucene Term Vectors](http://sujitpal.blogspot.in/2011/10/computing-document-similarity-using.html)  
> ).
> 
> However, from what I can tell ES doesn't expose the term vectors. Is it  
> still possible for me to use ES if I absolutely need the above feature? Is  
> it possible to read the Lucene index generated by ES directly without too  
> much trouble?
> 
> Of course, I could always generate the term vector dynamically for each  
> document for the purpose of implementing this particular feature, but  
> that's inefficient (I will be performing a large number of such  
> comparisons) and I don't want to do that if there is an alternate -- solr  
> seems to allow fetching the term vectors ☹
> 
> Any help would be appreciated!
> 
> Thanks,  
> Aditya

--

---

<div class="post-metadata">

**Author:** ![Aditya\_Rajgarhia](https://avatars.discourse-cdn.com/v4/letter/a/dec6dc/32.png) [@Aditya\_Rajgarhia](https://discuss.elastic.co/u/Aditya_Rajgarhia)\
**Post date:** [January 12, 2013, 4:50am UTC](https://discuss.elastic.co/t/term-vectors-for-computing-document-similarity/10311/3 "2013-01-12T04:50:29Z")

</div>

Loïc, thanks for the response.

I am familiar with mlt and am already using it to produce similar documents  
from each of my indexes. However, for the particular feature that I  
described in the last post, I want to explicitly compare several specific  
documents from one index with a specific document from the second index and  
get the score for each pair. In other words, I don't want to run a  
comparison over every document in one or both indexes since there will be a  
large number of documents (millions) in each index. My understanding is  
that flt/mlt will do that, unfortunately.

If this is still achievable via flt/mlt, could you elaborate a bit on how?

Thanks,  
Aditya

On Saturday, January 12, 2013 1:14:26 AM UTC+5:30, Loïc Bertron wrote:

> Hello,
> 
> You should have a look at this feature : Fuzzy Like this and More like  
> this:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/flt-query.html)  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/mlt-query.html)
> 
> If you compare your text field to all the text fields of all others  
> documents, you can reduce results only to documents matching 95% and more.
> 
> Le vendredi 11 janvier 2013 02:37:28 UTC-5, [adi...@blobinfotech.com](mailto:adi...@blobinfotech.com) a  
> écrit :
> 
> > Hello,
> > 
> > I'm trying to build some search features for a website. I don't have any  
> > prior experience with search and decided to go with elasticsearch mostly  
> > because of the ease of use.
> > 
> > I am indexing two types of documents (each with it's own index) since I  
> > want to offer search functionality for either type of document. I have this  
> > part working.
> > 
> > Now, I also want to offer a feature for comparing documents from one  
> > index with those from the other. What I had in mind was that since ES uses  
> > Lucene, I could fetch the term vectors for a pair of documents and then  
> > compute the cosine similarity (as explained in  
> > [Salmon Run: Computing Document Similarity using Lucene Term Vectors](http://sujitpal.blogspot.in/2011/10/computing-document-similarity-using.html)  
> > ).
> > 
> > However, from what I can tell ES doesn't expose the term vectors. Is it  
> > still possible for me to use ES if I absolutely need the above feature? Is  
> > it possible to read the Lucene index generated by ES directly without too  
> > much trouble?
> > 
> > Of course, I could always generate the term vector dynamically for each  
> > document for the purpose of implementing this particular feature, but  
> > that's inefficient (I will be performing a large number of such  
> > comparisons) and I don't want to do that if there is an alternate -- solr  
> > seems to allow fetching the term vectors ☹
> > 
> > Any help would be appreciated!
> > 
> > Thanks,  
> > Aditya

--

---

<div class="post-metadata">

**Author:** ![Pratik\_Poddar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pratik_poddar/32/1529_2.png) [@Pratik\_Poddar](https://discuss.elastic.co/u/Pratik_Poddar)\
**Post date:** [April 24, 2014, 10:29am UTC](https://discuss.elastic.co/t/term-vectors-for-computing-document-similarity/10311/4 "2014-04-24T10:29:00Z")

</div>

Aditya, any luck here? Would appreciate if you could share your learning  
please? Thanks a ton

Regards,  
Pratik Poddar

On Saturday, January 12, 2013 10:20:29 AM UTC+5:30, Aditya Rajgarhia wrote:

> Loïc, thanks for the response.
> 
> I am familiar with mlt and am already using it to produce similar  
> documents from each of my indexes. However, for the particular feature that  
> I described in the last post, I want to explicitly compare several specific  
> documents from one index with a specific document from the second index and  
> get the score for each pair. In other words, I don't want to run a  
> comparison over every document in one or both indexes since there will be a  
> large number of documents (millions) in each index. My understanding is  
> that flt/mlt will do that, unfortunately.
> 
> If this is still achievable via flt/mlt, could you elaborate a bit on how?
> 
> Thanks,  
> Aditya
> 
> On Saturday, January 12, 2013 1:14:26 AM UTC+5:30, Loïc Bertron wrote:
> 
> > Hello,
> > 
> > You should have a look at this feature : Fuzzy Like this and More like  
> > this:  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/flt-query.html)  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/mlt-query.html)
> > 
> > If you compare your text field to all the text fields of all others  
> > documents, you can reduce results only to documents matching 95% and more.
> > 
> > Le vendredi 11 janvier 2013 02:37:28 UTC-5, [adi...@blobinfotech.com](mailto:adi...@blobinfotech.com) a  
> > écrit :
> > 
> > > Hello,
> > > 
> > > I'm trying to build some search features for a website. I don't have any  
> > > prior experience with search and decided to go with elasticsearch mostly  
> > > because of the ease of use.
> > > 
> > > I am indexing two types of documents (each with it's own index) since I  
> > > want to offer search functionality for either type of document. I have this  
> > > part working.
> > > 
> > > Now, I also want to offer a feature for comparing documents from one  
> > > index with those from the other. What I had in mind was that since ES uses  
> > > Lucene, I could fetch the term vectors for a pair of documents and then  
> > > compute the cosine similarity (as explained in  
> > > [Salmon Run: Computing Document Similarity using Lucene Term Vectors](http://sujitpal.blogspot.in/2011/10/computing-document-similarity-using.html)  
> > > ).
> > > 
> > > However, from what I can tell ES doesn't expose the term vectors. Is it  
> > > still possible for me to use ES if I absolutely need the above feature? Is  
> > > it possible to read the Lucene index generated by ES directly without too  
> > > much trouble?
> > > 
> > > Of course, I could always generate the term vector dynamically for each  
> > > document for the purpose of implementing this particular feature, but  
> > > that's inefficient (I will be performing a large number of such  
> > > comparisons) and I don't want to do that if there is an alternate -- solr  
> > > seems to allow fetching the term vectors ☹
> > > 
> > > Any help would be appreciated!
> > > 
> > > Thanks,  
> > > Aditya

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/662f4e65-5f78-4cb4-9759-c6976763fc02%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/662f4e65-5f78-4cb4-9759-c6976763fc02%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Aditya\_Rajgarhia](https://avatars.discourse-cdn.com/v4/letter/a/dec6dc/32.png) [@Aditya\_Rajgarhia](https://discuss.elastic.co/u/Aditya_Rajgarhia)\
**Post date:** [April 25, 2014, 12:24pm UTC](https://discuss.elastic.co/t/term-vectors-for-computing-document-similarity/10311/5 "2014-04-25T12:24:09Z")

</div>

For my purposes I was able to use mlt-field, which is slightly different  
from mlt-query and offers you more customizability. Combined with and/or  
queries, you can construct some really powerful queries.

For what it's worth, I believe they've recently added a term vectors API as  
well, which I didn't use since the above worked better and allowed me to  
operate at a higher level.

You can search for all of the above on their docs.

On Thursday, April 24, 2014 3:59:00 PM UTC+5:30, Pratik Poddar wrote:

> Aditya, any luck here? Would appreciate if you could share your learning  
> please? Thanks a ton
> 
> Regards,  
> Pratik Poddar
> 
> On Saturday, January 12, 2013 10:20:29 AM UTC+5:30, Aditya Rajgarhia wrote:
> 
> > Loïc, thanks for the response.
> > 
> > I am familiar with mlt and am already using it to produce similar  
> > documents from each of my indexes. However, for the particular feature that  
> > I described in the last post, I want to explicitly compare several specific  
> > documents from one index with a specific document from the second index and  
> > get the score for each pair. In other words, I don't want to run a  
> > comparison over every document in one or both indexes since there will be a  
> > large number of documents (millions) in each index. My understanding is  
> > that flt/mlt will do that, unfortunately.
> > 
> > If this is still achievable via flt/mlt, could you elaborate a bit on how?
> > 
> > Thanks,  
> > Aditya
> > 
> > On Saturday, January 12, 2013 1:14:26 AM UTC+5:30, Loïc Bertron wrote:
> > 
> > > Hello,
> > > 
> > > You should have a look at this feature : Fuzzy Like this and More like  
> > > this:  
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/flt-query.html)  
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/mlt-query.html)
> > > 
> > > If you compare your text field to all the text fields of all others  
> > > documents, you can reduce results only to documents matching 95% and more.
> > > 
> > > Le vendredi 11 janvier 2013 02:37:28 UTC-5, [adi...@blobinfotech.com](mailto:adi...@blobinfotech.com) a  
> > > écrit :
> > > 
> > > > Hello,
> > > > 
> > > > I'm trying to build some search features for a website. I don't have  
> > > > any prior experience with search and decided to go with elasticsearch  
> > > > mostly because of the ease of use.
> > > > 
> > > > I am indexing two types of documents (each with it's own index) since I  
> > > > want to offer search functionality for either type of document. I have this  
> > > > part working.
> > > > 
> > > > Now, I also want to offer a feature for comparing documents from one  
> > > > index with those from the other. What I had in mind was that since ES uses  
> > > > Lucene, I could fetch the term vectors for a pair of documents and then  
> > > > compute the cosine similarity (as explained in  
> > > > [Salmon Run: Computing Document Similarity using Lucene Term Vectors](http://sujitpal.blogspot.in/2011/10/computing-document-similarity-using.html)  
> > > > ).
> > > > 
> > > > However, from what I can tell ES doesn't expose the term vectors. Is it  
> > > > still possible for me to use ES if I absolutely need the above feature? Is  
> > > > it possible to read the Lucene index generated by ES directly without too  
> > > > much trouble?
> > > > 
> > > > Of course, I could always generate the term vector dynamically for each  
> > > > document for the purpose of implementing this particular feature, but  
> > > > that's inefficient (I will be performing a large number of such  
> > > > comparisons) and I don't want to do that if there is an alternate -- solr  
> > > > seems to allow fetching the term vectors ☹
> > > > 
> > > > Any help would be appreciated!
> > > > 
> > > > Thanks,  
> > > > Aditya

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/67fb94c5-0ba5-4630-873e-6dd7be1068f9%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/67fb94c5-0ba5-4630-873e-6dd7be1068f9%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Pratik\_Poddar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pratik_poddar/32/1529_2.png) [@Pratik\_Poddar](https://discuss.elastic.co/u/Pratik_Poddar)\
**Post date:** [April 25, 2014, 12:28pm UTC](https://discuss.elastic.co/t/term-vectors-for-computing-document-similarity/10311/6 "2014-04-25T12:28:45Z")

</div>

Aditya,  
Thanks for your reply. But even mlt\_field gives you close documents. How do  
we measure similarity between two documents? If you are able to solve this,  
do you mind sharing the snippet please? Thanks a ton. Really appreciate it.

Regards,  
Pratik

On Fri, Apr 25, 2014 at 5:54 PM, Aditya Rajgarhia  
[aditya@blobinfotech.com](mailto:aditya@blobinfotech.com)wrote:

> For my purposes I was able to use mlt-field, which is slightly different  
> from mlt-query and offers you more customizability. Combined with and/or  
> queries, you can construct some really powerful queries.
> 
> For what it's worth, I believe they've recently added a term vectors API  
> as well, which I didn't use since the above worked better and allowed me to  
> operate at a higher level.
> 
> You can search for all of the above on their docs.
> 
> On Thursday, April 24, 2014 3:59:00 PM UTC+5:30, Pratik Poddar wrote:
> 
> > Aditya, any luck here? Would appreciate if you could share your learning  
> > please? Thanks a ton
> > 
> > Regards,  
> > Pratik Poddar
> > 
> > On Saturday, January 12, 2013 10:20:29 AM UTC+5:30, Aditya Rajgarhia  
> > wrote:
> > 
> > > Loïc, thanks for the response.
> > > 
> > > I am familiar with mlt and am already using it to produce similar  
> > > documents from each of my indexes. However, for the particular feature that  
> > > I described in the last post, I want to explicitly compare several specific  
> > > documents from one index with a specific document from the second index and  
> > > get the score for each pair. In other words, I don't want to run a  
> > > comparison over every document in one or both indexes since there will be a  
> > > large number of documents (millions) in each index. My understanding is  
> > > that flt/mlt will do that, unfortunately.
> > > 
> > > If this is still achievable via flt/mlt, could you elaborate a bit on  
> > > how?
> > > 
> > > Thanks,  
> > > Aditya
> > > 
> > > On Saturday, January 12, 2013 1:14:26 AM UTC+5:30, Loïc Bertron wrote:
> > > 
> > > > Hello,
> > > > 
> > > > You should have a look at this feature : Fuzzy Like this and More like  
> > > > this:  
> > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/flt-query.html)  
> > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/mlt-query.html)
> > > > 
> > > > If you compare your text field to all the text fields of all others  
> > > > documents, you can reduce results only to documents matching 95% and more.
> > > > 
> > > > Le vendredi 11 janvier 2013 02:37:28 UTC-5, [adi...@blobinfotech.com](mailto:adi...@blobinfotech.com) a  
> > > > écrit :
> > > > 
> > > > > Hello,
> > > > > 
> > > > > I'm trying to build some search features for a website. I don't have  
> > > > > any prior experience with search and decided to go with elasticsearch  
> > > > > mostly because of the ease of use.
> > > > > 
> > > > > I am indexing two types of documents (each with it's own index) since  
> > > > > I want to offer search functionality for either type of document. I have  
> > > > > this part working.
> > > > > 
> > > > > Now, I also want to offer a feature for comparing documents from one  
> > > > > index with those from the other. What I had in mind was that since ES uses  
> > > > > Lucene, I could fetch the term vectors for a pair of documents and then  
> > > > > compute the cosine similarity (as explained in  
> > > > > [http://sujitpal.blogspot.in/2011/10/computing-document-](http://sujitpal.blogspot.in/2011/10/computing-document-)  
> > > > > similarity-using.html).
> > > > > 
> > > > > However, from what I can tell ES doesn't expose the term vectors. Is  
> > > > > it still possible for me to use ES if I absolutely need the above feature?  
> > > > > Is it possible to read the Lucene index generated by ES directly without  
> > > > > too much trouble?
> > > > > 
> > > > > Of course, I could always generate the term vector dynamically for  
> > > > > each document for the purpose of implementing this particular feature, but  
> > > > > that's inefficient (I will be performing a large number of such  
> > > > > comparisons) and I don't want to do that if there is an alternate -- solr  
> > > > > seems to allow fetching the term vectors ☹
> > > > > 
> > > > > Any help would be appreciated!
> > > > > 
> > > > > Thanks,  
> > > > > Aditya
> > > > 
> > > > --  
> > > > You received this message because you are subscribed to a topic in the  
> > > > Google Groups "elasticsearch" group.  
> > > > To unsubscribe from this topic, visit  
> > > > [https://groups.google.com/d/topic/elasticsearch/VExh3UhD5Yg/unsubscribe](https://groups.google.com/d/topic/elasticsearch/VExh3UhD5Yg/unsubscribe).  
> > > > To unsubscribe from this group and all its topics, send an email to  
> > > > [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > > > To view this discussion on the web visit  
> > > > [https://groups.google.com/d/msgid/elasticsearch/67fb94c5-0ba5-4630-873e-6dd7be1068f9%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/67fb94c5-0ba5-4630-873e-6dd7be1068f9%40googlegroups.com)[https://groups.google.com/d/msgid/elasticsearch/67fb94c5-0ba5-4630-873e-6dd7be1068f9%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/67fb94c5-0ba5-4630-873e-6dd7be1068f9%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > > .
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
Pratik Poddar  
[www.linkedin.com/in/pratikpoddar](http://www.linkedin.com/in/pratikpoddar)

> **[CseBlog.com is for sale | HugeDomains](https://www.hugedomains.com/domain_profile.cfm?d=cseblog.com)**
>
> Choosing the right domain name can be overwhelming. Our personalized customer service helps you get a great domain.

> **[Pratik Poddar's Web Space](https://pratikpoddar.wordpress.com/)**
>
> Visit the post for more.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAFiYsPc1kAS6Samqx6EqVYUyqL-DOAtg3Lev-rDjtE27rTnSoQ%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAFiYsPc1kAS6Samqx6EqVYUyqL-DOAtg3Lev-rDjtE27rTnSoQ%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Aditya\_Rajgarhia](https://avatars.discourse-cdn.com/v4/letter/a/dec6dc/32.png) [@Aditya\_Rajgarhia](https://discuss.elastic.co/u/Aditya_Rajgarhia)\
**Post date:** [April 25, 2014, 12:43pm UTC](https://discuss.elastic.co/t/term-vectors-for-computing-document-similarity/10311/7 "2014-04-25T12:43:54Z")

</div>

I didn't need to compute scores since chaining and nesting queries allowed  
me a much better solution for my needs than I would have been ever been  
able to get by writing the algorithm from scratch. Some of these query  
types were not available when I posted this thread.

As I said, they've added term vectors recently:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

Why can't you use this? Also, even before they added this API there was a  
way to get term vectors by writing low level code to get the lucene  
information from ES.

On Friday, April 25, 2014 5:58:45 PM UTC+5:30, Pratik Poddar wrote:

> Aditya,  
> Thanks for your reply. But even mlt\_field gives you close documents. How  
> do we measure similarity between two documents? If you are able to solve  
> this, do you mind sharing the snippet please? Thanks a ton. Really  
> appreciate it.
> 
> Regards,  
> Pratik
> 
> On Fri, Apr 25, 2014 at 5:54 PM, Aditya Rajgarhia \<[adi...@blobinfotech.com](mailto:adi...@blobinfotech.com)\<javascript:\>
> 
> > wrote:
> 
> > For my purposes I was able to use mlt-field, which is slightly different  
> > from mlt-query and offers you more customizability. Combined with and/or  
> > queries, you can construct some really powerful queries.
> > 
> > For what it's worth, I believe they've recently added a term vectors API  
> > as well, which I didn't use since the above worked better and allowed me to  
> > operate at a higher level.
> > 
> > You can search for all of the above on their docs.
> > 
> > On Thursday, April 24, 2014 3:59:00 PM UTC+5:30, Pratik Poddar wrote:
> > 
> > > Aditya, any luck here? Would appreciate if you could share your learning  
> > > please? Thanks a ton
> > > 
> > > Regards,  
> > > Pratik Poddar
> > > 
> > > On Saturday, January 12, 2013 10:20:29 AM UTC+5:30, Aditya Rajgarhia  
> > > wrote:
> > > 
> > > > Loïc, thanks for the response.
> > > > 
> > > > I am familiar with mlt and am already using it to produce similar  
> > > > documents from each of my indexes. However, for the particular feature that  
> > > > I described in the last post, I want to explicitly compare several specific  
> > > > documents from one index with a specific document from the second index and  
> > > > get the score for each pair. In other words, I don't want to run a  
> > > > comparison over every document in one or both indexes since there will be a  
> > > > large number of documents (millions) in each index. My understanding is  
> > > > that flt/mlt will do that, unfortunately.
> > > > 
> > > > If this is still achievable via flt/mlt, could you elaborate a bit on  
> > > > how?
> > > > 
> > > > Thanks,  
> > > > Aditya
> > > > 
> > > > On Saturday, January 12, 2013 1:14:26 AM UTC+5:30, Loïc Bertron wrote:
> > > > 
> > > > > Hello,
> > > > > 
> > > > > You should have a look at this feature : Fuzzy Like this and More like  
> > > > > this:  
> > > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/flt-query.html)  
> > > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/mlt-query.html)
> > > > > 
> > > > > If you compare your text field to all the text fields of all others  
> > > > > documents, you can reduce results only to documents matching 95% and more.
> > > > > 
> > > > > Le vendredi 11 janvier 2013 02:37:28 UTC-5, [adi...@blobinfotech.com](mailto:adi...@blobinfotech.com) a  
> > > > > écrit :
> > > > > 
> > > > > > Hello,
> > > > > > 
> > > > > > I'm trying to build some search features for a website. I don't have  
> > > > > > any prior experience with search and decided to go with elasticsearch  
> > > > > > mostly because of the ease of use.
> > > > > > 
> > > > > > I am indexing two types of documents (each with it's own index) since  
> > > > > > I want to offer search functionality for either type of document. I have  
> > > > > > this part working.
> > > > > > 
> > > > > > Now, I also want to offer a feature for comparing documents from one  
> > > > > > index with those from the other. What I had in mind was that since ES uses  
> > > > > > Lucene, I could fetch the term vectors for a pair of documents and then  
> > > > > > compute the cosine similarity (as explained in  
> > > > > > [http://sujitpal.blogspot.in/2011/10/computing-document-](http://sujitpal.blogspot.in/2011/10/computing-document-)  
> > > > > > similarity-using.html).
> > > > > > 
> > > > > > However, from what I can tell ES doesn't expose the term vectors. Is  
> > > > > > it still possible for me to use ES if I absolutely need the above feature?  
> > > > > > Is it possible to read the Lucene index generated by ES directly without  
> > > > > > too much trouble?
> > > > > > 
> > > > > > Of course, I could always generate the term vector dynamically for  
> > > > > > each document for the purpose of implementing this particular feature, but  
> > > > > > that's inefficient (I will be performing a large number of such  
> > > > > > comparisons) and I don't want to do that if there is an alternate -- solr  
> > > > > > seems to allow fetching the term vectors ☹
> > > > > > 
> > > > > > Any help would be appreciated!
> > > > > > 
> > > > > > Thanks,  
> > > > > > Aditya
> > > > > 
> > > > > --  
> > > > > You received this message because you are subscribed to a topic in the  
> > > > > Google Groups "elasticsearch" group.  
> > > > > To unsubscribe from this topic, visit  
> > > > > [https://groups.google.com/d/topic/elasticsearch/VExh3UhD5Yg/unsubscribe](https://groups.google.com/d/topic/elasticsearch/VExh3UhD5Yg/unsubscribe).  
> > > > > To unsubscribe from this group and all its topics, send an email to  
> > > > > [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > > > > To view this discussion on the web visit  
> > > > > [https://groups.google.com/d/msgid/elasticsearch/67fb94c5-0ba5-4630-873e-6dd7be1068f9%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/67fb94c5-0ba5-4630-873e-6dd7be1068f9%40googlegroups.com)[https://groups.google.com/d/msgid/elasticsearch/67fb94c5-0ba5-4630-873e-6dd7be1068f9%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/67fb94c5-0ba5-4630-873e-6dd7be1068f9%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > > > .
> > 
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> 
> --  
> Pratik Poddar  
> [Pratik Poddar - Nexus Venture Partners | LinkedIn](http://www.linkedin.com/in/pratikpoddar)  
> [http://www.cseblog.com](http://www.cseblog.com)  
> [http://pratikpoddar.wordpress.com/](http://pratikpoddar.wordpress.com/)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/315374ec-4236-4bd3-a175-f9f0f481259c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/315374ec-4236-4bd3-a175-f9f0f481259c%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:33am UTC](https://discuss.elastic.co/t/term-vectors-for-computing-document-similarity/10311/8 "2017-07-06T01:33:33Z")

</div>


