# Most frequently occurring phrases?

**URL:** <https://discuss.elastic.co/t/most-frequently-occurring-phrases/16859>\
**Category:** Elasticsearch\
**Created:** [April 7, 2014, 11:26pm UTC](https://discuss.elastic.co/t/most-frequently-occurring-phrases/16859 "2014-04-07T23:26:59Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![John\_Stanford](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/john_stanford/32/1647_2.png) [@John\_Stanford](https://discuss.elastic.co/u/John_Stanford)\
**Post date:** [April 7, 2014, 11:26pm UTC](https://discuss.elastic.co/t/most-frequently-occurring-phrases/16859/1 "2014-04-07T23:26:59Z")

</div>

Hi,

I have a bunch of text events indexed as a message field, and in many  
cases, they are similar but not exactly the same. Is there a way to return  
the top n most frequently occurring similar phrases, and if so, how would I  
control the definition of similar?

Thanks,  
John

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/861cd17c-3897-4fd0-8a66-847f7cabdb8a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/861cd17c-3897-4fd0-8a66-847f7cabdb8a%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![John\_Stanford](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/john_stanford/32/1647_2.png) [@John\_Stanford](https://discuss.elastic.co/u/John_Stanford)\
**Post date:** [April 9, 2014, 5:34pm UTC](https://discuss.elastic.co/t/most-frequently-occurring-phrases/16859/2 "2014-04-09T17:34:26Z")

</div>

Here's an example. If I use aggregations to search for the top 10 most  
frequent messages:

POST \_search  
{  
"query": {  
"match": {  
"loglevel": "error"  
}  
},  
"aggs": {  
"freqent\_msgs": {  
"terms": {  
"field": "message.raw",  
"size": 10  
}  
}  
}  
}

I end up with a list that exhibit two undesirable characteristics. The top  
3 entries are the same type of message, but have different instances. The  
remaining messages are a few different types, but each of them has a  
repetitive counter. Is there a way to overlook these differences so the  
result would be closer to the 4 message types?

"aggregations": {  
"freqent\_msgs": {  
"buckets": [  
{  
"key": "Getting disk size of instance-0000bcbb: [Errno 2] No  
such file or directory:  
'/var/lib/nova/instances/9b173949-c34d-401e-a214-8e3d8ddefd46/disk'",  
"doc\_count": 22599  
},  
{  
"key": "Getting disk size of instance-0000bd08: [Errno 2] No  
such file or directory:  
'/var/lib/nova/instances/a4e2c7b5-093a-494f-bdef-5b6997e7c3bb/disk'",  
"doc\_count": 13447  
},  
{  
"key": "Getting disk size of instance-0000bd09: [Errno 2] No  
such file or directory:  
'/var/lib/nova/instances/ca680c42-f7c8-49ea-b46e-8864051c860c/disk'",  
"doc\_count": 13447  
},  
{  
"key": "Unable to connect to AMQP server: [Errno 113]  
EHOSTUNREACH. Sleeping 60 seconds",  
"doc\_count": 32  
},  
{  
"key": "Unable to connect to AMQP server: [Errno 113]  
EHOSTUNREACH. Sleeping 32 seconds",  
"doc\_count": 15  
},  
{  
"key": "Unable to connect to AMQP server: [Errno 111]  
ECONNREFUSED. Sleeping 2 seconds",  
"doc\_count": 12  
},  
{  
"key": "Unable to connect to AMQP server: [Errno 111]  
ECONNREFUSED. Sleeping 4 seconds",  
"doc\_count": 10  
},  
{  
"key": "Unable to connect to AMQP server: [Errno 111]  
ECONNREFUSED. Sleeping 8 seconds",  
"doc\_count": 9  
},  
{  
"key": "Unable to connect to AMQP server: [Errno 110]  
ETIMEDOUT. Sleeping 16 seconds",  
"doc\_count": 7  
},  
{  
"key": "Unable to connect to AMQP server: [Errno 111]  
ECONNREFUSED. Sleeping 1 seconds",  
"doc\_count": 7  
}  
]  
}  
}

Thanks,  
John

On Monday, April 7, 2014 4:26:59 PM UTC-7, John Stanford wrote:

> Hi,
> 
> I have a bunch of text events indexed as a message field, and in many  
> cases, they are similar but not exactly the same. Is there a way to return  
> the top n most frequently occurring similar phrases, and if so, how would I  
> control the definition of similar?
> 
> Thanks,  
> John

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/803575fb-fae1-43d0-9085-2e7fdc21f321%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/803575fb-fae1-43d0-9085-2e7fdc21f321%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [April 11, 2014, 6:05am UTC](https://discuss.elastic.co/t/most-frequently-occurring-phrases/16859/3 "2014-04-11T06:05:49Z")

</div>

Hey,

as these two sample messages a very different in nature, it is hard to use  
something like scripting to cut those messages off after a certain length  
as a workaround. I would go with some sort of preprocessing (maybe using  
logstash), where you give each message a certain type/identifier and facet  
on that one.

--Alex

On Wed, Apr 9, 2014 at 7:34 PM, John Stanford [jxstanford@gmail.com](mailto:jxstanford@gmail.com) wrote:

> Here's an example. If I use aggregations to search for the top 10 most  
> frequent messages:
> 
> POST \_search  
> {  
> "query": {  
> "match": {  
> "loglevel": "error"  
> }  
> },  
> "aggs": {  
> "freqent\_msgs": {  
> "terms": {  
> "field": "message.raw",  
> "size": 10  
> }  
> }  
> }  
> }
> 
> I end up with a list that exhibit two undesirable characteristics. The  
> top 3 entries are the same type of message, but have different instances.  
> The remaining messages are a few different types, but each of them has a  
> repetitive counter. Is there a way to overlook these differences so the  
> result would be closer to the 4 message types?
> 
> "aggregations": {  
> "freqent\_msgs": {  
> "buckets": [  
> {  
> "key": "Getting disk size of instance-0000bcbb: [Errno 2]  
> No such file or directory:  
> '/var/lib/nova/instances/9b173949-c34d-401e-a214-8e3d8ddefd46/disk'",  
> "doc\_count": 22599  
> },  
> {  
> "key": "Getting disk size of instance-0000bd08: [Errno 2]  
> No such file or directory:  
> '/var/lib/nova/instances/a4e2c7b5-093a-494f-bdef-5b6997e7c3bb/disk'",  
> "doc\_count": 13447  
> },  
> {  
> "key": "Getting disk size of instance-0000bd09: [Errno 2]  
> No such file or directory:  
> '/var/lib/nova/instances/ca680c42-f7c8-49ea-b46e-8864051c860c/disk'",  
> "doc\_count": 13447  
> },  
> {  
> "key": "Unable to connect to AMQP server: [Errno 113]  
> EHOSTUNREACH. Sleeping 60 seconds",  
> "doc\_count": 32  
> },  
> {  
> "key": "Unable to connect to AMQP server: [Errno 113]  
> EHOSTUNREACH. Sleeping 32 seconds",  
> "doc\_count": 15  
> },  
> {  
> "key": "Unable to connect to AMQP server: [Errno 111]  
> ECONNREFUSED. Sleeping 2 seconds",  
> "doc\_count": 12  
> },  
> {  
> "key": "Unable to connect to AMQP server: [Errno 111]  
> ECONNREFUSED. Sleeping 4 seconds",  
> "doc\_count": 10  
> },  
> {  
> "key": "Unable to connect to AMQP server: [Errno 111]  
> ECONNREFUSED. Sleeping 8 seconds",  
> "doc\_count": 9  
> },  
> {  
> "key": "Unable to connect to AMQP server: [Errno 110]  
> ETIMEDOUT. Sleeping 16 seconds",  
> "doc\_count": 7  
> },  
> {  
> "key": "Unable to connect to AMQP server: [Errno 111]  
> ECONNREFUSED. Sleeping 1 seconds",  
> "doc\_count": 7  
> }  
> ]  
> }  
> }
> 
> Thanks,  
> John
> 
> On Monday, April 7, 2014 4:26:59 PM UTC-7, John Stanford wrote:
> 
> > Hi,
> > 
> > I have a bunch of text events indexed as a message field, and in many  
> > cases, they are similar but not exactly the same. Is there a way to return  
> > the top n most frequently occurring similar phrases, and if so, how would I  
> > control the definition of similar?
> > 
> > Thanks,  
> > John
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/803575fb-fae1-43d0-9085-2e7fdc21f321%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/803575fb-fae1-43d0-9085-2e7fdc21f321%40googlegroups.com)[https://groups.google.com/d/msgid/elasticsearch/803575fb-fae1-43d0-9085-2e7fdc21f321%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/803575fb-fae1-43d0-9085-2e7fdc21f321%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAGCwEM\_OoWWp1nBVdwkWriSk4zFftEr2hRX%3DTAsx8vMT2StfQA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAGCwEM_OoWWp1nBVdwkWriSk4zFftEr2hRX%3DTAsx8vMT2StfQA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![John\_Stanford](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/john_stanford/32/1647_2.png) [@John\_Stanford](https://discuss.elastic.co/u/John_Stanford)\
**Post date:** [April 11, 2014, 1:58pm UTC](https://discuss.elastic.co/t/most-frequently-occurring-phrases/16859/4 "2014-04-11T13:58:34Z")

</div>

Hi Alex,

Yeah, I'm doing that with some other message types, but was hoping to keep that to select messages with metrics in them. I may look into some post processing strategies, and will keep searching for a reasonable solution within elasticsearch.

Thanks,

John

> On Apr 10, 2014, at 11:05 PM, Alexander Reelsen [alr@spinscale.de](mailto:alr@spinscale.de) wrote:
> 
> Hey,
> 
> as these two sample messages a very different in nature, it is hard to use something like scripting to cut those messages off after a certain length as a workaround. I would go with some sort of preprocessing (maybe using logstash), where you give each message a certain type/identifier and facet on that one.
> 
> --Alex
> 
> > On Wed, Apr 9, 2014 at 7:34 PM, John Stanford [jxstanford@gmail.com](mailto:jxstanford@gmail.com) wrote:  
> > Here's an example. If I use aggregations to search for the top 10 most frequent messages:
> > 
> > POST \_search  
> > {  
> > "query": {  
> > "match": {  
> > "loglevel": "error"  
> > }  
> > },  
> > "aggs": {  
> > "freqent\_msgs": {  
> > "terms": {  
> > "field": "message.raw",  
> > "size": 10  
> > }  
> > }  
> > }  
> > }
> > 
> > I end up with a list that exhibit two undesirable characteristics. The top 3 entries are the same type of message, but have different instances. The remaining messages are a few different types, but each of them has a repetitive counter. Is there a way to overlook these differences so the result would be closer to the 4 message types?
> > 
> > "aggregations": {  
> > "freqent\_msgs": {  
> > "buckets": [  
> > {  
> > "key": "Getting disk size of instance-0000bcbb: [Errno 2] No such file or directory: '/var/lib/nova/instances/9b173949-c34d-401e-a214-8e3d8ddefd46/disk'",  
> > "doc\_count": 22599  
> > },  
> > {  
> > "key": "Getting disk size of instance-0000bd08: [Errno 2] No such file or directory: '/var/lib/nova/instances/a4e2c7b5-093a-494f-bdef-5b6997e7c3bb/disk'",  
> > "doc\_count": 13447  
> > },  
> > {  
> > "key": "Getting disk size of instance-0000bd09: [Errno 2] No such file or directory: '/var/lib/nova/instances/ca680c42-f7c8-49ea-b46e-8864051c860c/disk'",  
> > "doc\_count": 13447  
> > },  
> > {  
> > "key": "Unable to connect to AMQP server: [Errno 113] EHOSTUNREACH. Sleeping 60 seconds",  
> > "doc\_count": 32  
> > },  
> > {  
> > "key": "Unable to connect to AMQP server: [Errno 113] EHOSTUNREACH. Sleeping 32 seconds",  
> > "doc\_count": 15  
> > },  
> > {  
> > "key": "Unable to connect to AMQP server: [Errno 111] ECONNREFUSED. Sleeping 2 seconds",  
> > "doc\_count": 12  
> > },  
> > {  
> > "key": "Unable to connect to AMQP server: [Errno 111] ECONNREFUSED. Sleeping 4 seconds",  
> > "doc\_count": 10  
> > },  
> > {  
> > "key": "Unable to connect to AMQP server: [Errno 111] ECONNREFUSED. Sleeping 8 seconds",  
> > "doc\_count": 9  
> > },  
> > {  
> > "key": "Unable to connect to AMQP server: [Errno 110] ETIMEDOUT. Sleeping 16 seconds",  
> > "doc\_count": 7  
> > },  
> > {  
> > "key": "Unable to connect to AMQP server: [Errno 111] ECONNREFUSED. Sleeping 1 seconds",  
> > "doc\_count": 7  
> > }  
> > ]  
> > }  
> > }
> > 
> > Thanks,  
> > John
> > 
> > > On Monday, April 7, 2014 4:26:59 PM UTC-7, John Stanford wrote:  
> > > Hi,
> > > 
> > > I have a bunch of text events indexed as a message field, and in many cases, they are similar but not exactly the same. Is there a way to return the top n most frequently occurring similar phrases, and if so, how would I control the definition of similar?
> > > 
> > > Thanks,  
> > > John
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/803575fb-fae1-43d0-9085-2e7fdc21f321%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/803575fb-fae1-43d0-9085-2e7fdc21f321%40googlegroups.com).
> > 
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> 
> --  
> You received this message because you are subscribed to a topic in the Google Groups "elasticsearch" group.  
> To unsubscribe from this topic, visit [https://groups.google.com/d/topic/elasticsearch/9bQdUgTQqgU/unsubscribe](https://groups.google.com/d/topic/elasticsearch/9bQdUgTQqgU/unsubscribe).  
> To unsubscribe from this group and all its topics, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAGCwEM\_OoWWp1nBVdwkWriSk4zFftEr2hRX%3DTAsx8vMT2StfQA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAGCwEM_OoWWp1nBVdwkWriSk4zFftEr2hRX%3DTAsx8vMT2StfQA%40mail.gmail.com).  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/14301933-4556-4F89-BB5E-B4E9A3F79D3E%40gmail.com](https://groups.google.com/d/msgid/elasticsearch/14301933-4556-4F89-BB5E-B4E9A3F79D3E%40gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:36am UTC](https://discuss.elastic.co/t/most-frequently-occurring-phrases/16859/5 "2017-07-06T01:36:30Z")

</div>


