# Plugin development guidance

**URL:** <https://discuss.elastic.co/t/plugin-development-guidance/15737>\
**Category:** Elasticsearch\
**Created:** [February 11, 2014, 6:13pm UTC](https://discuss.elastic.co/t/plugin-development-guidance/15737 "2014-02-11T18:13:02Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Josh\_Harrison](https://avatars.discourse-cdn.com/v4/letter/j/0ea827/32.png) [@Josh\_Harrison](https://discuss.elastic.co/u/Josh_Harrison)\
**Post date:** [February 11, 2014, 6:13pm UTC](https://discuss.elastic.co/t/plugin-development-guidance/15737/1 "2014-02-11T18:13:02Z")

</div>

Hi all,  
We've got an internal Java library that allows us to do keyword extraction  
that seems like a great thing to turn into an integrated elasticsearch  
function.  
Ultimately, I want to be able to access the result of this library from  
search results/etc, but I wanted to do a sanity check to make sure my  
approach was right - or if I should be looking at doing a custom analyzer  
or something instead.

Given a string field, the type would become a multi-field, {name} and  
keywords/phrases as subfields. A plugin would be written to handle this  
keywords field, run the strings through the library and return a list of  
strings like:  
"my\_data":"Jack and Jill went up the hill, Jack fell down and bumped his  
crown, and Jill came tumbling after."  
"my\_data.keywords":["Jack", "Jack fell"]  
That's a trivial example, of course, and the algorithm is more complex than  
the standard stopword filtering.

Ultimately, I want to be able to expose the my\_data.keywords field as an  
actual list like above, so that we can use it in other things like facets  
down the line.

So is a custom type plugin the right way to go here, or should I be looking  
at developing a more complex analyzer/tokenizer/stopword combo?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/44d3ba64-f0b3-4727-9a49-745a2167d34d%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/44d3ba64-f0b3-4727-9a49-745a2167d34d%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [February 11, 2014, 9:18pm UTC](https://discuss.elastic.co/t/plugin-development-guidance/15737/2 "2014-02-11T21:18:40Z")

</div>

An analyzer plugin is the right thing. Adding the recognized/extracted  
terms needs access to ES mapping service. There are a few plugins out there  
which work in this manner, for example, the attachment mapper plugin.

Or the lang-detect plugin, it adds the recognized language(s) as a keyword  
code into a neighbor field for filtering or faceting:

> **[jprante/elasticsearch-langdetect](https://github.com/jprante/elasticsearch-langdetect)**
>
> elasticsearch-langdetect - A plugin for language detection in Elasticsearch using Nakatani Shuyo's language detector

Also, I developed a similar plugin that works with recognition techniques,  
it can recognize ISBN or other standard number in a text, and injects extra  
tokens into the token stream to identify these numbers:

> **[jprante/elasticsearch-analysis-standardnumber](https://github.com/jprante/elasticsearch-analysis-standardnumber)**
>
> elasticsearch-analysis-standardnumber - Analyze standard numbers like ARK, DOI, EAN, GTIN, IBAN, ISAN, ISBN, ISMN, ISNI, ISSN, ISTC, ISWC, ORCID, PPN, SICI, UPC, ZDB with Elasticsearch

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoEzwFzy4zGbcg1w66LQgcEq8L5O9tjj2ke\_6krw9nc%2B7A%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoEzwFzy4zGbcg1w66LQgcEq8L5O9tjj2ke_6krw9nc%2B7A%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Josh\_Harrison](https://avatars.discourse-cdn.com/v4/letter/j/0ea827/32.png) [@Josh\_Harrison](https://discuss.elastic.co/u/Josh_Harrison)\
**Post date:** [February 11, 2014, 9:26pm UTC](https://discuss.elastic.co/t/plugin-development-guidance/15737/3 "2014-02-11T21:26:24Z")

</div>

Great, thanks Jörg!  
I'll start fiddling around with the langdetect plugin to see if I can get  
it going with our library.

On Tue, Feb 11, 2014 at 1:18 PM, [joergprante@gmail.com](mailto:joergprante@gmail.com) \<  
[joergprante@gmail.com](mailto:joergprante@gmail.com)\> wrote:

> An analyzer plugin is the right thing. Adding the recognized/extracted  
> terms needs access to ES mapping service. There are a few plugins out there  
> which work in this manner, for example, the attachment mapper plugin.
> 
> Or the lang-detect plugin, it adds the recognized language(s) as a keyword  
> code into a neighbor field for filtering or faceting:  
> [GitHub - jprante/elasticsearch-langdetect: A plugin for language detection in Elasticsearch using Nakatani Shuyo's language detector](https://github.com/jprante/elasticsearch-langdetect)
> 
> Also, I developed a similar plugin that works with recognition techniques,  
> it can recognize ISBN or other standard number in a text, and injects extra  
> tokens into the token stream to identify these numbers:  
> [GitHub - jprante/elasticsearch-analysis-standardnumber: Analyze standard numbers like ARK, DOI, EAN, GTIN, IBAN, ISAN, ISBN, ISMN, ISNI, ISSN, ISTC, ISWC, ORCID, PPN, SICI, UPC, ZDB with Elasticsearch](https://github.com/jprante/elasticsearch-analysis-standardnumber)
> 
> Jörg
> 
> --  
> You received this message because you are subscribed to a topic in the  
> Google Groups "elasticsearch" group.  
> To unsubscribe from this topic, visit  
> [https://groups.google.com/d/topic/elasticsearch/1lexzKdBbP8/unsubscribe](https://groups.google.com/d/topic/elasticsearch/1lexzKdBbP8/unsubscribe).  
> To unsubscribe from this group and all its topics, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoEzwFzy4zGbcg1w66LQgcEq8L5O9tjj2ke\_6krw9nc%2B7A%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoEzwFzy4zGbcg1w66LQgcEq8L5O9tjj2ke_6krw9nc%2B7A%40mail.gmail.com)  
> .
> 
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALX3-AmgVoyoTy1nso0\_SW%3DaVNkZ3aSKXN8tbPTnmrOfqjnVDQ%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALX3-AmgVoyoTy1nso0_SW%3DaVNkZ3aSKXN8tbPTnmrOfqjnVDQ%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:50am UTC](https://discuss.elastic.co/t/plugin-development-guidance/15737/4 "2017-07-06T01:50:54Z")

</div>


