# Allow a dot at the beginning of a token, like '.Net'

**URL:** <https://discuss.elastic.co/t/allow-a-dot-at-the-beginning-of-a-token-like-net/4198>\
**Category:** Elasticsearch\
**Created:** [April 7, 2011, 2:13am UTC](https://discuss.elastic.co/t/allow-a-dot-at-the-beginning-of-a-token-like-net/4198 "2011-04-07T02:13:39Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![AGuereca](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aguereca/32/2357_2.png) [@AGuereca](https://discuss.elastic.co/u/AGuereca)\
**Post date:** [April 7, 2011, 2:13am UTC](https://discuss.elastic.co/t/allow-a-dot-at-the-beginning-of-a-token-like-net/4198/1 "2011-04-07T02:13:39Z")

</div>

I have all day trying to find a way to make the word “ **.Net** ” being considered as a token instead of just “ **Net** ” without the dot; Because I’m getting noise in my searches when the term .Net is required with dot (which mean something completely different without).

So far I haven’t been successful the closest result has been the withspace tokenizer, the problem with that is sentences like “This features: ” are being tokenized like “this”, “ **features:** ” and of course any search by term “features” will ignore the document because the “ **:** ”.

I’m hoping there is something else besides the “Pattern Analyzer”; which I don’t have a clear idea of how make it work properly.

Any ideas will be greatly appreciated since I’m out of them after all day knocking my head with this.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [April 9, 2011, 11:09am UTC](https://discuss.elastic.co/t/allow-a-dot-at-the-beginning-of-a-token-like-net/4198/2 "2011-04-09T11:09:58Z")

</div>

On Wed, 2011-04-06 at 19:13 -0700, AGuereca wrote:

> I have all day trying to find a way to make the word â.Netâ being considered  
> as a token instead of just âNetâ without the dot; Because Iâm getting noise  
> in my searches when the term .Net is required with dot (which mean something  
> completely different without).

This is tricky to do in the middle of a blob of text, because there are  
loads of other uses of '.' where you do want to ignore the '.' and other  
punctuation.

One thing you could do, short of writing a custom tokenizer in Java, is  
to preprocess both your document text and your query strings to replace  
occurrences of ".Net" with something like "dotNet"

It's a bit manual, but at least this way you won't lose the benefits of  
the standard analyzer.

hth

clint

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [April 9, 2011, 2:35pm UTC](https://discuss.elastic.co/t/allow-a-dot-at-the-beginning-of-a-token-like-net/4198/3 "2011-04-09T14:35:13Z")

</div>

I've did similar things for @user and #hashtag.

Therefor I've stolen the WordDelimiterFilter from solr:

> <https://github.com/karussell/Jetwick/blob/master/src/main/java/org/apache/solr/analysis/WordDelimiterFilter.java>

Then I've extended it:

> <https://github.com/karussell/Jetwick/blob/master/src/main/java/de/jetwick/es/JetwickFilterFactory.java>

Now you could then specify handleAsChar = ".";

I've used handleAsDigit = "@" and the solr setting  
splitOnNumerics=true so that @user is indexed as two terms: user and  
@user

Regards,  
Peter.

--  
[http://jetwick.com/](http://jetwick.com/) Personalized Twitter Search

On 7 Apr., 04:13, AGuereca [aguer...@gmail.com](mailto:aguer...@gmail.com) wrote:

> I have all day trying to find a way to make the word “.Net” being considered  
> as a token instead of just “Net” without the dot; Because I’m getting noise  
> in my searches when the term .Net is required with dot (which mean something  
> completely different without).
> 
> So far I haven’t been successful the closest result has been the withspace  
> tokenizer, the problem with that is sentences like “This features: ” are  
> being tokenized like “this”, “features:” and of course any search by term  
> “features” will ignore the document because the “:”.
> 
> I’m hoping there is something else besides the “Pattern Analyzer”; which I  
> don’t have a clear idea of how make it work properly.
> 
> Any ideas will be greatly appreciated since I’m out of them after all day  
> knocking my head with this.
> 
> --  
> View this message in context:[http://elasticsearch-users.115913.n3.nabble.com/Allow-a-dot-at-the-be](http://elasticsearch-users.115913.n3.nabble.com/Allow-a-dot-at-the-be)...  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

---

<div class="post-metadata">

**Author:** ![Gabriel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gabriel/32/2968_2.png) [@Gabriel](https://discuss.elastic.co/u/Gabriel)\
**Post date:** [March 6, 2012, 10:13pm UTC](https://discuss.elastic.co/t/allow-a-dot-at-the-beginning-of-a-token-like-net/4198/4 "2012-03-06T22:13:35Z")

</div>

Is there any way to do that without create a java class?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:36am UTC](https://discuss.elastic.co/t/allow-a-dot-at-the-beginning-of-a-token-like-net/4198/5 "2017-07-06T03:36:59Z")

</div>


