# Use Proper Nouns / Named Entities to improve MLT results?

**URL:** <https://discuss.elastic.co/t/use-proper-nouns-named-entities-to-improve-mlt-results/10611>\
**Category:** Elasticsearch\
**Created:** [February 4, 2013, 1:26am UTC](https://discuss.elastic.co/t/use-proper-nouns-named-entities-to-improve-mlt-results/10611 "2013-02-04T01:26:43Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![racedo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/racedo/32/2499_2.png) [@racedo](https://discuss.elastic.co/u/racedo)\
**Post date:** [February 4, 2013, 1:26am UTC](https://discuss.elastic.co/t/use-proper-nouns-named-entities-to-improve-mlt-results/10611/1 "2013-02-04T01:26:43Z")

</div>

Hi all,

I'm indexing news articles to basically relate them within a cluster of  
documents with "more like this". In this case, the results could be heavily  
improved if there was a way to give more weight to terms that are proper  
nouns / named entities. Is there any way to do this with elasticsearch?

Many thanks in advance.

Ramon

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [February 4, 2013, 9:27am UTC](https://discuss.elastic.co/t/use-proper-nouns-named-entities-to-improve-mlt-results/10611/2 "2013-02-04T09:27:12Z")

</div>

The detection of proper nouns / named entities is currently outside of  
the scope of Elasticsearch.

It is possible to integrate indexing programs with text mining  
algorithms, such as Stanford NER, Apache OpenNLP NER, or LingPipe, and  
create fields for the recognized entities, and a related JSON object /  
array that contains the "more like this" synonyms for a field.

Because the NER task is heavy and time consuming, I would not recommend  
to run it on the same machine where an ES data node is running. So, a  
plugin would be possible, but only for TransportClient side.

Jörg

Am 04.02.13 02:26, schrieb racedo:

> Hi all,
> 
> I'm indexing news articles to basically relate them within a cluster  
> of documents with "more like this". In this case, the results could be  
> heavily improved if there was a way to give more weight to terms that  
> are proper nouns / named entities. Is there any way to do this with  
> elasticsearch?
> 
> Many thanks in advance.
> 
> Ramon
> 
> --  
> You received this message because you are subscribed to the Google  
> Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send  
> an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Alex\_At\_Ikanow](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alex_at_ikanow/32/676_2.png) [@Alex\_At\_Ikanow](https://discuss.elastic.co/u/Alex_At_Ikanow)\
**Post date:** [February 5, 2013, 2:37pm UTC](https://discuss.elastic.co/t/use-proper-nouns-named-entities-to-improve-mlt-results/10611/3 "2013-02-05T14:37:22Z")

</div>

Ramon,

I have some experience in this sort of thing.

(note: I build an open source document analysis platform[https://github.com/IKANOW/Infinit.e](https://github.com/IKANOW/Infinit.e) built  
on Elasticsearch and including various support for entity extractors, that  
is probably too heavyweight for the purposes you describe; though some of  
the code may prove useful)

There are a few options, depending on the sort of article you are indexing:  
1] You can very cheaply run the text through the OpenNLP "POS tagger"[http://opennlp.apache.org/documentation/manual/opennlp.html#tools.postagger.tagging.cmdline](http://opennlp.apache.org/documentation/manual/opennlp.html#tools.postagger.tagging.cmdline) ...  
this does an OKish job of picking out proper nouns and other things you  
might want to run NLP over. You can then stick them in a multi field and  
use that with "mlt". Not sure which language your news articles are in, you  
can find POS models for many languages though.  
2] A slightly more complex but better solution would be to use TextRank[https://github.com/turian/textrank](https://github.com/turian/textrank).  
This is built on-top of OpenNLP POS, but then uses a snazzy bit of maths to  
pick out significant keyphrases. The trick would then be to tokenize the  
keyphrases back into keywords and put them in the mlt array.  
3] For mainstream news articles written in English/French/Spanish, OpenCalais  
[http://www.opencalais.com/](http://www.opencalais.com/)is a very good free SaaS named entity  
extractor, with a decent daily call allowance. (Disclaimer: I haven't  
looked at its performance on languages other than English)  
3a] (There are also many commercial alternatives of comparable quality,  
some of which have low volume free tiers, eg we have used AlchemyAPI[http://www.alchemyapi.com/](http://www.alchemyapi.com/)  
)  
3b] (For slightly less mainstream news articles you will find that the SaaS  
offerings like OpenCalais and AlchemyAPI will tend to be over-aggressive at  
resolving names to names of famous people - we had to build in some post  
processing to "unresolve" names)

(Finally, for geo-tagging, give Clavin[https://github.com/Berico-Technologies/CLAVIN](https://github.com/Berico-Technologies/CLAVIN)a look. We haven't integrated it into "Infinit.e" yet, but it looks pretty  
good.)

I will also say that from experience, you will still probably find the  
results of the "mlt" query disappointing (I'd love to hear about it if you  
don't!) - the "standard" way of clustering documents involves using Mahout[https://cwiki.apache.org/MAHOUT/quick-tour-of-text-analysis-using-the-mahout-command-line.html](https://cwiki.apache.org/MAHOUT/quick-tour-of-text-analysis-using-the-mahout-command-line.html),  
which integrates well with Lucene but not so much Elasticsearch.

Hope this helps!

Alex  
[www.ikanow.com](http://www.ikanow.com)

On Sunday, February 3, 2013 8:26:43 PM UTC-5, racedo wrote:

> Hi all,
> 
> I'm indexing news articles to basically relate them within a cluster of  
> documents with "more like this". In this case, the results could be heavily  
> improved if there was a way to give more weight to terms that are proper  
> nouns / named entities. Is there any way to do this with elasticsearch?
> 
> Many thanks in advance.
> 
> Ramon

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![racedo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/racedo/32/2499_2.png) [@racedo](https://discuss.elastic.co/u/racedo)\
**Post date:** [February 7, 2013, 1:17am UTC](https://discuss.elastic.co/t/use-proper-nouns-named-entities-to-improve-mlt-results/10611/4 "2013-02-07T01:17:57Z")

</div>

Alex, this is excellent feedback, much appreciated. I'd like to keep using  
elasticsearch for simplicity reasons and as Joerg suggested a plug-in might  
come handy too. Thanks to Joerg as well for his suggestions.

On 5 February 2013 07:37, Alex at Ikanow [apiggott@ikanow.com](mailto:apiggott@ikanow.com) wrote:

> Ramon,
> 
> I have some experience in this sort of thing.
> 
> (note: I build an open source document analysis platform[https://github.com/IKANOW/Infinit.e](https://github.com/IKANOW/Infinit.e) built  
> on Elasticsearch and including various support for entity extractors, that  
> is probably too heavyweight for the purposes you describe; though some of  
> the code may prove useful)
> 
> There are a few options, depending on the sort of article you are indexing:  
> 1] You can very cheaply run the text through the OpenNLP "POS tagger"[http://opennlp.apache.org/documentation/manual/opennlp.html#tools.postagger.tagging.cmdline](http://opennlp.apache.org/documentation/manual/opennlp.html#tools.postagger.tagging.cmdline) ...  
> this does an OKish job of picking out proper nouns and other things you  
> might want to run NLP over. You can then stick them in a multi field and  
> use that with "mlt". Not sure which language your news articles are in, you  
> can find POS models for many languages though.  
> 2] A slightly more complex but better solution would be to use TextRank[https://github.com/turian/textrank](https://github.com/turian/textrank).  
> This is built on-top of OpenNLP POS, but then uses a snazzy bit of maths to  
> pick out significant keyphrases. The trick would then be to tokenize the  
> keyphrases back into keywords and put them in the mlt array.  
> 3] For mainstream news articles written in English/French/Spanish, OpenCalais  
> [http://www.opencalais.com/](http://www.opencalais.com/)is a very good free SaaS named entity  
> extractor, with a decent daily call allowance. (Disclaimer: I haven't  
> looked at its performance on languages other than English)  
> 3a] (There are also many commercial alternatives of comparable quality,  
> some of which have low volume free tiers, eg we have used AlchemyAPI[http://www.alchemyapi.com/](http://www.alchemyapi.com/)  
> )  
> 3b] (For slightly less mainstream news articles you will find that the  
> SaaS offerings like OpenCalais and AlchemyAPI will tend to be  
> over-aggressive at resolving names to names of famous people - we had to  
> build in some post processing to "unresolve" names)
> 
> (Finally, for geo-tagging, give Clavin[https://github.com/Berico-Technologies/CLAVIN](https://github.com/Berico-Technologies/CLAVIN)a look. We haven't integrated it into "Infinit.e" yet, but it looks pretty  
> good.)
> 
> I will also say that from experience, you will still probably find the  
> results of the "mlt" query disappointing (I'd love to hear about it if you  
> don't!) - the "standard" way of clustering documents involves using Mahout[https://cwiki.apache.org/MAHOUT/quick-tour-of-text-analysis-using-the-mahout-command-line.html](https://cwiki.apache.org/MAHOUT/quick-tour-of-text-analysis-using-the-mahout-command-line.html),  
> which integrates well with Lucene but not so much Elasticsearch.
> 
> Hope this helps!
> 
> Alex  
> [www.ikanow.com](http://www.ikanow.com)
> 
> On Sunday, February 3, 2013 8:26:43 PM UTC-5, racedo wrote:
> 
> > Hi all,
> > 
> > I'm indexing news articles to basically relate them within a cluster of  
> > documents with "more like this". In this case, the results could be heavily  
> > improved if there was a way to give more weight to terms that are proper  
> > nouns / named entities. Is there any way to do this with elasticsearch?
> > 
> > Many thanks in advance.
> > 
> > Ramon
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
Ramon

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:52am UTC](https://discuss.elastic.co/t/use-proper-nouns-named-entities-to-improve-mlt-results/10611/5 "2017-07-06T02:52:40Z")

</div>


