# Is there a concatenation filter?

**URL:** <https://discuss.elastic.co/t/is-there-a-concatenation-filter/6577>\
**Category:** Elasticsearch\
**Created:** [February 2, 2012, 8:25pm UTC](https://discuss.elastic.co/t/is-there-a-concatenation-filter/6577 "2012-02-02T20:25:54Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![cole](https://avatars.discourse-cdn.com/v4/letter/c/53a042/32.png) [@cole](https://discuss.elastic.co/u/cole)\
**Post date:** [February 2, 2012, 8:25pm UTC](https://discuss.elastic.co/t/is-there-a-concatenation-filter/6577/1 "2012-02-02T20:25:54Z")

</div>

Was there any progress in adding the concatenation filter [1] to  
Lucene (and ES) last summer? I can't find any evidence of built-in  
support for this type of filter.

Thanks,  
Cole

[1] [http://elasticsearch-users.115913.n3.nabble.com/Code-contribution-Concatenate-filter-td3137058.html#a3707818](http://elasticsearch-users.115913.n3.nabble.com/Code-contribution-Concatenate-filter-td3137058.html#a3707818)

---

<div class="post-metadata">

**Author:** ![Stephane\_Bastian](https://avatars.discourse-cdn.com/v4/letter/s/35a633/32.png) [@Stephane\_Bastian](https://discuss.elastic.co/u/Stephane_Bastian)\
**Post date:** [February 3, 2012, 1:42pm UTC](https://discuss.elastic.co/t/is-there-a-concatenation-filter/6577/2 "2012-02-03T13:42:59Z")

</div>

Hi Cole,

no, I got so busy right after my email last summer that I didn't follow up  
and dropped the ball.  
However if you are interested I can send you the code for the filter. just  
let me know

Stephane

---

<div class="post-metadata">

**Author:** ![cole](https://avatars.discourse-cdn.com/v4/letter/c/53a042/32.png) [@cole](https://discuss.elastic.co/u/cole)\
**Post date:** [February 3, 2012, 7:27pm UTC](https://discuss.elastic.co/t/is-there-a-concatenation-filter/6577/3 "2012-02-03T19:27:30Z")

</div>

Hi Stephane,

Thanks for the reply. I'd very much appreciate seeing your  
concatentation filter code. Do you have it up somewhere you can link  
to?

Thanks,  
Cole

On Feb 3, 5:42 am, Stephane Bastian [stephane.bastian....@gmail.com](mailto:stephane.bastian....@gmail.com)  
wrote:

> Hi Cole,
> 
> no, I got so busy right after my email last summer that I didn't follow up  
> and dropped the ball.  
> However if you are interested I can send you the code for the filter. just  
> let me know
> 
> Stephane

---

<div class="post-metadata">

**Author:** ![Stephane\_Bastian](https://avatars.discourse-cdn.com/v4/letter/s/35a633/32.png) [@Stephane\_Bastian](https://discuss.elastic.co/u/Stephane_Bastian)\
**Post date:** [February 6, 2012, 10:25am UTC](https://discuss.elastic.co/t/is-there-a-concatenation-filter/6577/4 "2012-02-06T10:25:33Z")

</div>

Hello Cole,

Here is the code for the concatenate filter. As you can see, it's very  
simple but does the job for me.

|public final class ConcatenateFilter extends TokenFilter {

```
 private final static String DEFAULT_TOKEN_SEPARATOR = " ";

 private final CharTermAttribute termAtt = 

```

addAttribute(CharTermAttribute.class);  
private String tokenSeparator = null;  
private StringBuilder builder = new StringBuilder();

```
 public ConcatenateFilter(Version matchVersion, TokenStream input, 

```

String tokenSeparator) {  
super(input);  
this.tokenSeparator = tokenSeparator!=null ? tokenSeparator :  
DEFAULT\_TOKEN\_SEPARATOR;  
}

```
 @Override
 public boolean incrementToken() throws IOException {
     boolean result = false;
     builder.setLength(0);
     while (input.incrementToken()) {
         if (builder.length()>0) {
             // append the token separator
             builder.append(tokenSeparator);
         }
         // append the term of the current token
         builder.append(termAtt.buffer(), 0, termAtt.length());
     }
     if (builder.length()>0) {
         termAtt.setEmpty().append(builder);
         result = true;
     }
     return result;
 }

```

}|

As you can see above the code is pure lucene (no ES code). In order to  
use the filter in ES you need to implement another class:  
|  
public class ConcatenateTokenFilterFactory extends  
AbstractTokenFilterFactory {

```
 private String tokenSeparator = null;

 @Inject
 public ConcatenateTokenFilterFactory(Index index, @IndexSettings 

```

Settings indexSettings, @Assisted String name, @Assisted Settings  
settings) {  
super(index, indexSettings, name, settings);  
// ||the token\_separator is defined in the ES configuration file|  
| tokenSeparator = settings.get("token\_separator");  
}

```
 @Override
 public TokenStream create(TokenStream tokenStream) {
     return new *ConcatenateFilter*(Version.LUCENE_CURRENT, 

```

tokenStream, tokenSeparator);  
}  
}|

and to glue things together you then need to declare the  
|ConcatenateTokenFilterFactory in ES config file:|

| "index": {  
"analysis": {  
"analyzer": {  
"myAnalyzer": {  
"tokenizer": "letter",  
"filter": ["lowercase", "asciifolding", "_filter-concatenate_"]  
}  
},  
"filter": {  
"_filter-concatenate_": {  
"type":  
"com.monpetitguide.elasticsearch.analysis._ConcatenateTokenFilterFactory_",  
"_token\_separator_": " "  
}  
}  
}  
} |

Cole, feel free to use any part of the code above. I'm glad if it helps

Al the best,

Stephane Bastian

On 02/03/2012 08:27 PM, cole wrote:

> Hi Stephane,
> 
> Thanks for the reply. I'd very much appreciate seeing your  
> concatentation filter code. Do you have it up somewhere you can link  
> to?
> 
> Thanks,  
> Cole
> 
> On Feb 3, 5:42 am, Stephane Bastian[stephane.bastian....@gmail.com](mailto:stephane.bastian....@gmail.com)  
> wrote:
> 
> > Hi Cole,
> > 
> > no, I got so busy right after my email last summer that I didn't follow up  
> > and dropped the ball.  
> > However if you are interested I can send you the code for the filter. just  
> > let me know
> > 
> > Stephane

---

<div class="post-metadata">

**Author:** ![cole](https://avatars.discourse-cdn.com/v4/letter/c/53a042/32.png) [@cole](https://discuss.elastic.co/u/cole)\
**Post date:** [February 6, 2012, 10:58pm UTC](https://discuss.elastic.co/t/is-there-a-concatenation-filter/6577/5 "2012-02-06T22:58:28Z")

</div>

Thanks, Stephane! I appreciate you explaining how everything is glued  
together. Very helpful!

Thanks,  
Cole

On Feb 6, 2:25 am, Stephane Bastian [stephane.bastian....@gmail.com](mailto:stephane.bastian....@gmail.com)  
wrote:

> Hello Cole,
> 
> Here is the code for the concatenate filter. As you can see, it's very  
> simple but does the job for me.
> 
> |public final class ConcatenateFilter extends TokenFilter {
> 
> ```
> private final static String DEFAULT_TOKEN_SEPARATOR = " ";
> 
> private final CharTermAttribute termAtt =
> 
> ```
> 
> addAttribute(CharTermAttribute.class);  
> private String tokenSeparator = null;  
> private StringBuilder builder = new StringBuilder();
> 
> ```
> public ConcatenateFilter(Version matchVersion, TokenStream input,
> 
> ```
> 
> String tokenSeparator) {  
> super(input);  
> this.tokenSeparator = tokenSeparator!=null ? tokenSeparator :  
> DEFAULT\_TOKEN\_SEPARATOR;  
> }
> 
> ```
> @Override
> public boolean incrementToken() throws IOException {
> boolean result = false;
> builder.setLength(0);
> while (input.incrementToken()) {
> if (builder.length()>0) {
> // append the token separator
> builder.append(tokenSeparator);
> }
> // append the term of the current token
> builder.append(termAtt.buffer(), 0, termAtt.length());
> }
> if (builder.length()>0) {
> termAtt.setEmpty().append(builder);
> result = true;
> }
> return result;
> }
> 
> ```
> 
> }|
> 
> As you can see above the code is pure lucene (no ES code). In order to  
> use the filter in ES you need to implement another class:  
> |  
> public class ConcatenateTokenFilterFactory extends  
> AbstractTokenFilterFactory {
> 
> ```
> private String tokenSeparator = null;
> 
> @Inject
> public ConcatenateTokenFilterFactory(Index index, @IndexSettings
> 
> ```
> 
> Settings indexSettings, @Assisted String name, @Assisted Settings  
> settings) {  
> super(index, indexSettings, name, settings);  
> // ||the token\_separator is defined in the ES configuration file|  
> | tokenSeparator = settings.get("token\_separator");  
> }
> 
> ```
> @Override
> public TokenStream create(TokenStream tokenStream) {
> return new *ConcatenateFilter*(Version.LUCENE_CURRENT,
> 
> ```
> 
> tokenStream, tokenSeparator);  
> }
> 
> }|
> 
> and to glue things together you then need to declare the  
> |ConcatenateTokenFilterFactory in ES config file:|
> 
> | "index": {  
> "analysis": {  
> "analyzer": {  
> "myAnalyzer": {  
> "tokenizer": "letter",  
> "filter": ["lowercase", "asciifolding", "_filter-concatenate_"]  
> }  
> },  
> "filter": {  
> "_filter-concatenate_": {  
> "type":  
> "com.monpetitguide.elasticsearch.analysis._ConcatenateTokenFilterFactory_",  
> "_token\_separator_": " "  
> }  
> }  
> }  
> } |
> 
> Cole, feel free to use any part of the code above. I'm glad if it helps
> 
> Al the best,
> 
> Stephane Bastian
> 
> On 02/03/2012 08:27 PM, cole wrote:
> 
> > Hi Stephane,
> 
> > Thanks for the reply. I'd very much appreciate seeing your  
> > concatentation filter code. Do you have it up somewhere you can link  
> > to?
> 
> > Thanks,  
> > Cole
> 
> > On Feb 3, 5:42 am, Stephane Bastian[stephane.bastian....@gmail.com](mailto:stephane.bastian....@gmail.com)  
> > wrote:
> > 
> > > Hi Cole,
> 
> > > no, I got so busy right after my email last summer that I didn't follow up  
> > > and dropped the ball.  
> > > However if you are interested I can send you the code for the filter. just  
> > > let me know
> 
> > > Stephane

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:40am UTC](https://discuss.elastic.co/t/is-there-a-concatenation-filter/6577/6 "2017-07-06T03:40:27Z")

</div>


