# Using ElasticSearch analyzers outside of ElasticSearch

**URL:** <https://discuss.elastic.co/t/using-elasticsearch-analyzers-outside-of-elasticsearch/8600>\
**Category:** Elasticsearch\
**Created:** [August 1, 2012, 6:07pm UTC](https://discuss.elastic.co/t/using-elasticsearch-analyzers-outside-of-elasticsearch/8600 "2012-08-01T18:07:18Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [August 1, 2012, 6:07pm UTC](https://discuss.elastic.co/t/using-elasticsearch-analyzers-outside-of-elasticsearch/8600/1 "2012-08-01T18:07:18Z")

</div>

Just thinking aloud, have not tried to implement anything, but I  
probably will soon...

I am currently using span queries for the bulk of my queries.  
Unfortunately, span queries only support term queries, which mean no  
analysis will happen on the query terms. My current approach utilizes  
a Lucene analyzer to analyze the terms used by the SpamTermQueries.

Using both a custom Lucene analyzer and an ElasticSearch analyzer (via  
elasticsearch.yml) has numerous issues: need to support two systems,  
potential mismatch, duplication of efforts, etc... The analysis API is  
useful, but the network hop to analyze each term would be too high.

My current thinking is to create a local embedded ElasticSearch  
instance who sole purpose would be fulfill analyze requests. The  
existing TransportClient would continue to communicate with the actual  
cluster, while this new NodeClient would only exist within the JVM.  
The embedded node would use the same analyzer definitions found in the  
cluster elasticsearch.yml file ("include" config files would be  
extremely useful).

Questions/thoughts:

1. I have never created an embedded ElasticSearch server, but I assume  
I can construct one that does not interfere with the existing cluster  
and/or the other middle boxes in the network. Correct?

2. Performance-wise. I expect that the performance of using the  
analyze API locally would be identical to using a Lucene analyzer.

3. How heavyweight is an embedded ElasticSearch instance?

Cheers,

Ivan

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [August 1, 2012, 11:39pm UTC](https://discuss.elastic.co/t/using-elasticsearch-analyzers-outside-of-elasticsearch/8600/2 "2012-08-01T23:39:06Z")

</div>

To answer my own question:

Tried created an embedded instance and quickly ran into the issue  
where a custom analyzer is not created unless it is tied to an index.  
I then simply followed the code used by the analysis unit tests  
(AnalysisModuleTests) and built an analyzer using the various  
Elasticsearch modules. Works like a charm.

Cheers,

Ivan

On Wed, Aug 1, 2012 at 11:07 AM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:

> Just thinking aloud, have not tried to implement anything, but I  
> probably will soon...
> 
> I am currently using span queries for the bulk of my queries.  
> Unfortunately, span queries only support term queries, which mean no  
> analysis will happen on the query terms. My current approach utilizes  
> a Lucene analyzer to analyze the terms used by the SpamTermQueries.
> 
> Using both a custom Lucene analyzer and an Elasticsearch analyzer (via  
> elasticsearch.yml) has numerous issues: need to support two systems,  
> potential mismatch, duplication of efforts, etc... The analysis API is  
> useful, but the network hop to analyze each term would be too high.
> 
> My current thinking is to create a local embedded Elasticsearch  
> instance who sole purpose would be fulfill analyze requests. The  
> existing TransportClient would continue to communicate with the actual  
> cluster, while this new NodeClient would only exist within the JVM.  
> The embedded node would use the same analyzer definitions found in the  
> cluster elasticsearch.yml file ("include" config files would be  
> extremely useful).
> 
> Questions/thoughts:
> 
> 1. I have never created an embedded Elasticsearch server, but I assume  
> I can construct one that does not interfere with the existing cluster  
> and/or the other middle boxes in the network. Correct?
> 
> 2. Performance-wise. I expect that the performance of using the  
> analyze API locally would be identical to using a Lucene analyzer.
> 
> 3. How heavyweight is an embedded Elasticsearch instance?
> 
> Cheers,
> 
> Ivan

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:18am UTC](https://discuss.elastic.co/t/using-elasticsearch-analyzers-outside-of-elasticsearch/8600/3 "2017-07-06T03:18:03Z")

</div>


