# 'dot' analyzer

**URL:** <https://discuss.elastic.co/t/dot-analyzer/3635>\
**Category:** Elasticsearch\
**Created:** [December 6, 2010, 11:57pm UTC](https://discuss.elastic.co/t/dot-analyzer/3635 "2010-12-06T23:57:51Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Rich\_Kroll](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rich_kroll/32/3229_2.png) [@Rich\_Kroll](https://discuss.elastic.co/u/Rich_Kroll)\
**Post date:** [December 6, 2010, 11:57pm UTC](https://discuss.elastic.co/t/dot-analyzer/3635/1 "2010-12-06T23:57:51Z")

</div>

I am indexing some log data, and just realized that the class/package names  
are not being indexed. For example "com.java.util.List" could only be  
searched using "_list_". Is there a way to modify the analyzer to tokenize  
on the 'dot' as well as whitespace?

Regards,  
Rich

--  
“We can't solve problems by using the same kind of thinking we used when we  
created them.” ~ Albert Einstein

---

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [December 7, 2010, 10:30am UTC](https://discuss.elastic.co/t/dot-analyzer/3635/2 "2010-12-07T10:30:13Z")

</div>

Hi,

this depends on the analyzer. By default the standard analyzer does not  
break text into tokens on dot if it is not followed by a whitespace.

Take the following text as an example:  
"one.two.three. four"

Using default analyzer:

curl -XGET '  
[http://localhost:9200/twitter/\_analyze?text=one.two.three.+four&pretty=1](http://localhost:9200/twitter/_analyze?text=one.two.three.+four&pretty=1)'

{  
"tokens" : [ {  
"token" : "one.two.three",  
"start\_offset" : 0,  
"end\_offset" : 14,  
"type" : "",  
"position" : 1  
}, {  
"token" : "four",  
"start\_offset" : 15,  
"end\_offset" : 19,  
"type" : "",  
"position" : 2  
} ]  
}

But if you use different analyzer (for example  
'simple[http://www.elasticsearch.com/docs/elasticsearch/index\_modules/analysis/analyzer/simple/](http://www.elasticsearch.com/docs/elasticsearch/index_modules/analysis/analyzer/simple/)'  
analyzer) then you can get what you want:

curl -XGET '  
[http://localhost:9200/twitter/\_analyze?analyzer=simple&text=one.two.three.+four&pretty=1](http://localhost:9200/twitter/_analyze?analyzer=simple&text=one.two.three.+four&pretty=1)  
'

{  
"tokens" : [ {  
"token" : "one",  
"start\_offset" : 0,  
"end\_offset" : 3,  
"type" : "word",  
"position" : 1  
}, {  
"token" : "two",  
"start\_offset" : 4,  
"end\_offset" : 7,  
"type" : "word",  
"position" : 2  
}, {  
"token" : "three",  
"start\_offset" : 8,  
"end\_offset" : 13,  
"type" : "word",  
"position" : 3  
}, {  
"token" : "four",  
"start\_offset" : 15,  
"end\_offset" : 19,  
"type" : "word",  
"position" : 4  
} ]  
}

Regards,  
Lukas

On Tue, Dec 7, 2010 at 12:57 AM, Rich Kroll [kroll.rich@gmail.com](mailto:kroll.rich@gmail.com) wrote:

> I am indexing some log data, and just realized that the class/package names  
> are not being indexed. For example "com.java.util.List" could only be  
> searched using "_list_". Is there a way to modify the analyzer to tokenize  
> on the 'dot' as well as whitespace?
> 
> Regards,  
> Rich
> 
> --  
> “We can't solve problems by using the same kind of thinking we used when we  
> created them.” ~ Albert Einstein

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:15am UTC](https://discuss.elastic.co/t/dot-analyzer/3635/3 "2017-07-06T04:15:37Z")

</div>


