# Classification with percolator

**URL:** <https://discuss.elastic.co/t/classification-with-percolator/15343>\
**Category:** Elasticsearch\
**Created:** [January 21, 2014, 3:01pm UTC](https://discuss.elastic.co/t/classification-with-percolator/15343 "2014-01-21T15:01:36Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Arthur\_Denning](https://avatars.discourse-cdn.com/v4/letter/a/5daacb/32.png) [@Arthur\_Denning](https://discuss.elastic.co/u/Arthur_Denning)\
**Post date:** [January 21, 2014, 3:01pm UTC](https://discuss.elastic.co/t/classification-with-percolator/15343/1 "2014-01-21T15:01:36Z")

</div>

I am considering using the percolator API to classify document, namely, by  
posting query like "football", "art" to the percolator, and then when  
adding new documents, percolator should return the right tags. My concerns  
is, suppose there is thousands of tag to be identified in this way, would  
it be a performance nightmare? Is there thousands of query that is  
implicitly running behind the scene?

And what would be the recommended way to tackle these kind of  
classification problem in Elasticsearch?

It seems that Lucene has a classification api. Is it already integrated  
elsewhere in Elasticsearch? Is there any roadmap concerning its  
implementation?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/8cd363be-5c9b-4b10-925c-fb4f1de4d4c3%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/8cd363be-5c9b-4b10-925c-fb4f1de4d4c3%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Binh\_Ly](https://avatars.discourse-cdn.com/v4/letter/b/ce7236/32.png) [@Binh\_Ly](https://discuss.elastic.co/u/Binh_Ly)\
**Post date:** [January 22, 2014, 12:27am UTC](https://discuss.elastic.co/t/classification-with-percolator/15343/2 "2014-01-22T00:27:03Z")

</div>

Arthur,

You should be able to use filters in your percolator queries so for example  
you can use a term/terms filter. Also, in ES 1.0 you can shard the  
percolator query index out so that percolation can distribute that load  
around for better scalability. The best way is to experiment with it:  
[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/downloads/1-0-0-RC1).

I actually worked for a company that did content classification this way,  
and the percolator was a perfect fit for that use-case.

On Tuesday, January 21, 2014 10:01:36 AM UTC-5, Arthur Denning wrote:

> I am considering using the percolator API to classify document, namely, by  
> posting query like "football", "art" to the percolator, and then when  
> adding new documents, percolator should return the right tags. My concerns  
> is, suppose there is thousands of tag to be identified in this way, would  
> it be a performance nightmare? Is there thousands of query that is  
> implicitly running behind the scene?
> 
> And what would be the recommended way to tackle these kind of  
> classification problem in Elasticsearch?
> 
> It seems that Lucene has a classification api. Is it already integrated  
> elsewhere in Elasticsearch? Is there any roadmap concerning its  
> implementation?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/a81c8c74-06a2-452c-8c82-3b0358d18380%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/a81c8c74-06a2-452c-8c82-3b0358d18380%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Arthur\_Denning](https://avatars.discourse-cdn.com/v4/letter/a/5daacb/32.png) [@Arthur\_Denning](https://discuss.elastic.co/u/Arthur_Denning)\
**Post date:** [January 22, 2014, 10:12am UTC](https://discuss.elastic.co/t/classification-with-percolator/15343/3 "2014-01-22T10:12:54Z")

</div>

Hey Binh, Thanks a lot and it is really nice to hear from someone with  
practical experience on this. Is it correct to say if I had a thousand  
tags, I would need to make thousands of

curl -XPUT 'localhost:9200/my-index1/.percolator/tagname1'

to register each tags? In your implementation is there any pitfalls or nice  
tricks that is worth noting?

On Wednesday, January 22, 2014 8:27:03 AM UTC+8, Binh Ly wrote:

> Arthur,
> 
> You should be able to use filters in your percolator queries so for  
> example you can use a term/terms filter. Also, in ES 1.0 you can shard the  
> percolator query index out so that percolation can distribute that load  
> around for better scalability. The best way is to experiment with it:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/downloads/1-0-0-RC1).
> 
> I actually worked for a company that did content classification this way,  
> and the percolator was a perfect fit for that use-case.
> 
> On Tuesday, January 21, 2014 10:01:36 AM UTC-5, Arthur Denning wrote:
> 
> > I am considering using the percolator API to classify document, namely,  
> > by posting query like "football", "art" to the percolator, and then when  
> > adding new documents, percolator should return the right tags. My concerns  
> > is, suppose there is thousands of tag to be identified in this way, would  
> > it be a performance nightmare? Is there thousands of query that is  
> > implicitly running behind the scene?
> > 
> > And what would be the recommended way to tackle these kind of  
> > classification problem in Elasticsearch?
> > 
> > It seems that Lucene has a classification api. Is it already integrated  
> > elsewhere in Elasticsearch? Is there any roadmap concerning its  
> > implementation?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/965b464c-1cf2-4ae5-83c1-5f18fe8d0228%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/965b464c-1cf2-4ae5-83c1-5f18fe8d0228%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Binh\_Ly](https://avatars.discourse-cdn.com/v4/letter/b/ce7236/32.png) [@Binh\_Ly](https://discuss.elastic.co/u/Binh_Ly)\
**Post date:** [January 22, 2014, 5:28pm UTC](https://discuss.elastic.co/t/classification-with-percolator/15343/4 "2014-01-22T17:28:50Z")

</div>

Arthur,

I am assuming that you will define a query/rule for each tag, so in your  
case yes, that would be the way to define the percolator queries.

Couple of things that you might want to be aware:

1. Percolation is CPU intensive
2. The lesser the queries you can percolate against, the better. So when  
you call the percolate API, see if you can also pass in a query criteria to  
limit the queries to percolate against.

On Wednesday, January 22, 2014 5:12:54 AM UTC-5, Arthur Denning wrote:

> Hey Binh, Thanks a lot and it is really nice to hear from someone with  
> practical experience on this. Is it correct to say if I had a thousand  
> tags, I would need to make thousands of
> 
> curl -XPUT 'localhost:9200/my-index1/.percolator/tagname1'
> 
> to register each tags? In your implementation is there any pitfalls or  
> nice tricks that is worth noting?
> 
> On Wednesday, January 22, 2014 8:27:03 AM UTC+8, Binh Ly wrote:
> 
> > Arthur,
> > 
> > You should be able to use filters in your percolator queries so for  
> > example you can use a term/terms filter. Also, in ES 1.0 you can shard the  
> > percolator query index out so that percolation can distribute that load  
> > around for better scalability. The best way is to experiment with it:  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/downloads/1-0-0-RC1).
> > 
> > I actually worked for a company that did content classification this way,  
> > and the percolator was a perfect fit for that use-case.
> > 
> > On Tuesday, January 21, 2014 10:01:36 AM UTC-5, Arthur Denning wrote:
> > 
> > > I am considering using the percolator API to classify document, namely,  
> > > by posting query like "football", "art" to the percolator, and then when  
> > > adding new documents, percolator should return the right tags. My concerns  
> > > is, suppose there is thousands of tag to be identified in this way, would  
> > > it be a performance nightmare? Is there thousands of query that is  
> > > implicitly running behind the scene?
> > > 
> > > And what would be the recommended way to tackle these kind of  
> > > classification problem in Elasticsearch?
> > > 
> > > It seems that Lucene has a classification api. Is it already integrated  
> > > elsewhere in Elasticsearch? Is there any roadmap concerning its  
> > > implementation?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/b6707b03-734a-4518-a12d-0e34e09e01f7%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/b6707b03-734a-4518-a12d-0e34e09e01f7%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:55am UTC](https://discuss.elastic.co/t/classification-with-percolator/15343/5 "2017-07-06T01:55:19Z")

</div>


