# Storing and analyzing user agent strings, general approach

**URL:** <https://discuss.elastic.co/t/storing-and-analyzing-user-agent-strings-general-approach/18346>\
**Category:** Elasticsearch\
**Created:** [June 26, 2014, 7:09am UTC](https://discuss.elastic.co/t/storing-and-analyzing-user-agent-strings-general-approach/18346 "2014-06-26T07:09:28Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mark\_Dodwell](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_dodwell/32/1446_2.png) [@Mark\_Dodwell](https://discuss.elastic.co/u/Mark_Dodwell)\
**Post date:** [June 26, 2014, 7:09am UTC](https://discuss.elastic.co/t/storing-and-analyzing-user-agent-strings-general-approach/18346/1 "2014-06-26T07:09:28Z")

</div>

I want to store a bunch of documents in elasticsearch (which represent a  
hit to a website) including the user agent of the client that made the  
original HTTP request.

Since user agent strings have a lot of variance, and the useful parts need  
parsing out (OS, browser, version etc.) I would like to be able to perform  
aggregations on those extracted features.

The simplest way I can think to do this would be to analyze the user agent  
string before indexing the document. The downside to this approach is as  
new/different user agent strings emerge (which is not unlikely) you would  
have to proactively update the parser.

This may be impossibly/undesirable for a number of reasons, but what I'd  
really like to do is index the raw user agent string and then perform the  
analysis/feature extraction post-hoc at query time. Any ideas/pointers on  
how to do this?

Aggregators? Custom analyzers? (How would you handle an update to the  
analyzer, would you need to re-run against all existing stored data?)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/ed9bf030-f9bf-480a-88b1-a80421b9e79e%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ed9bf030-f9bf-480a-88b1-a80421b9e79e%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Patrick\_Proniewski](https://avatars.discourse-cdn.com/v4/letter/p/96bed5/32.png) [@Patrick\_Proniewski](https://discuss.elastic.co/u/Patrick_Proniewski)\
**Post date:** [June 26, 2014, 7:34am UTC](https://discuss.elastic.co/t/storing-and-analyzing-user-agent-strings-general-approach/18346/2 "2014-06-26T07:34:11Z")

</div>

Hi,

You should give [Useragent filter plugin | Logstash Reference [8.11] | Elastic](http://logstash.net/docs/1.4.2/filters/useragent) a try before anything else.

Here is the relevant part of logstash.conf I'm using:

filter {  
if [type] == "apache" {  
if [user-agent] != "-" and [user-agent] != "" {  
useragent {  
add\_tag =\> ["UA"]  
source =\> "user-agent"  
}  
}  
if "UA" in [tags] {  
if [device] == "Other" { mutate { remove\_field =\> "device" } }  
if [name] == "Other" { mutate { remove\_field =\> "name" } }  
if [os] == "Other" { mutate { remove\_field =\> "os" } }  
}  
}  
}

It retains the full user-agent field, and add nice fileds like "device", "major" and "minor" version, "name", "os", "os\_name", "os\_major" and "os\_minor".

sample:  
"user-agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10\_8\_5) AppleWebKit/537.76.4 (KHTML, like Gecko) Version/6.1.4 Safari/537.76.4",  
"name": "Safari",  
"os": "Mac OS X 10.8.5",  
"os\_name": "Mac OS X",  
"os\_major": "10",  
"os\_minor": "8",  
"major": "6",  
"minor": "1",  
"patch": "4",

On 26 juin 2014, at 09:09, Mark Dodwell [mark@mkdynamic.co.uk](mailto:mark@mkdynamic.co.uk) wrote:

> I want to store a bunch of documents in elasticsearch (which represent a  
> hit to a website) including the user agent of the client that made the  
> original HTTP request.
> 
> Since user agent strings have a lot of variance, and the useful parts need  
> parsing out (OS, browser, version etc.) I would like to be able to perform  
> aggregations on those extracted features.
> 
> The simplest way I can think to do this would be to analyze the user agent  
> string before indexing the document. The downside to this approach is as  
> new/different user agent strings emerge (which is not unlikely) you would  
> have to proactively update the parser.
> 
> This may be impossibly/undesirable for a number of reasons, but what I'd  
> really like to do is index the raw user agent string and then perform the  
> analysis/feature extraction post-hoc at query time. Any ideas/pointers on  
> how to do this?
> 
> Aggregators? Custom analyzers? (How would you handle an update to the  
> analyzer, would you need to re-run against all existing stored data?)
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/ed9bf030-f9bf-480a-88b1-a80421b9e79e%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ed9bf030-f9bf-480a-88b1-a80421b9e79e%40googlegroups.com).  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/48F3FC4E-D872-43B9-A60D-D2755094AC85%40patpro.net](https://groups.google.com/d/msgid/elasticsearch/48F3FC4E-D872-43B9-A60D-D2755094AC85%40patpro.net).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Mark\_Dodwell](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_dodwell/32/1446_2.png) [@Mark\_Dodwell](https://discuss.elastic.co/u/Mark_Dodwell)\
**Post date:** [June 27, 2014, 8:03pm UTC](https://discuss.elastic.co/t/storing-and-analyzing-user-agent-strings-general-approach/18346/3 "2014-06-27T20:03:56Z")

</div>

Thanks, lots of useful stuff there.

--

Sent from Mailbox for iPhone

On Thu, Jun 26, 2014 at 12:34 AM, Patrick Proniewski  
[elasticsearch@patpro.net](mailto:elasticsearch@patpro.net) wrote:

> Hi,  
> You should give [Useragent filter plugin | Logstash Reference [8.11] | Elastic](http://logstash.net/docs/1.4.2/filters/useragent) a try before anything else.  
> Here is the relevant part of logstash.conf I'm using:  
> filter {  
> if [type] == "apache" {  
> if [user-agent] != "-" and [user-agent] != "" {  
> useragent {  
> add\_tag =\> ["UA"]  
> source =\> "user-agent"  
> }  
> }  
> if "UA" in [tags] {  
> if [device] == "Other" { mutate { remove\_field =\> "device" } }  
> if [name] == "Other" { mutate { remove\_field =\> "name" } }  
> if [os] == "Other" { mutate { remove\_field =\> "os" } }  
> }  
> }  
> }  
> It retains the full user-agent field, and add nice fileds like "device", "major" and "minor" version, "name", "os", "os\_name", "os\_major" and "os\_minor".  
> sample:  
> "user-agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10\_8\_5) AppleWebKit/537.76.4 (KHTML, like Gecko) Version/6.1.4 Safari/537.76.4",  
> "name": "Safari",  
> "os": "Mac OS X 10.8.5",  
> "os\_name": "Mac OS X",  
> "os\_major": "10",  
> "os\_minor": "8",  
> "major": "6",  
> "minor": "1",  
> "patch": "4",  
> On 26 juin 2014, at 09:09, Mark Dodwell [mark@mkdynamic.co.uk](mailto:mark@mkdynamic.co.uk) wrote:
> 
> > I want to store a bunch of documents in elasticsearch (which represent a  
> > hit to a website) including the user agent of the client that made the  
> > original HTTP request.
> > 
> > Since user agent strings have a lot of variance, and the useful parts need  
> > parsing out (OS, browser, version etc.) I would like to be able to perform  
> > aggregations on those extracted features.
> > 
> > The simplest way I can think to do this would be to analyze the user agent  
> > string before indexing the document. The downside to this approach is as  
> > new/different user agent strings emerge (which is not unlikely) you would  
> > have to proactively update the parser.
> > 
> > This may be impossibly/undesirable for a number of reasons, but what I'd  
> > really like to do is index the raw user agent string and then perform the  
> > analysis/feature extraction post-hoc at query time. Any ideas/pointers on  
> > how to do this?
> > 
> > Aggregators? Custom analyzers? (How would you handle an update to the  
> > analyzer, would you need to re-run against all existing stored data?)
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/ed9bf030-f9bf-480a-88b1-a80421b9e79e%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ed9bf030-f9bf-480a-88b1-a80421b9e79e%40googlegroups.com).  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).  
> > --  
> > You received this message because you are subscribed to a topic in the Google Groups "elasticsearch" group.  
> > To unsubscribe from this topic, visit [https://groups.google.com/d/topic/elasticsearch/H-sUPppQMp8/unsubscribe](https://groups.google.com/d/topic/elasticsearch/H-sUPppQMp8/unsubscribe).  
> > To unsubscribe from this group and all its topics, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/48F3FC4E-D872-43B9-A60D-D2755094AC85%40patpro.net](https://groups.google.com/d/msgid/elasticsearch/48F3FC4E-D872-43B9-A60D-D2755094AC85%40patpro.net).  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/1403899436158.1403ba21%40Nodemailer](https://groups.google.com/d/msgid/elasticsearch/1403899436158.1403ba21%40Nodemailer).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Patrick\_Proniewski](https://avatars.discourse-cdn.com/v4/letter/p/96bed5/32.png) [@Patrick\_Proniewski](https://discuss.elastic.co/u/Patrick_Proniewski)\
**Post date:** [June 28, 2014, 5:22am UTC](https://discuss.elastic.co/t/storing-and-analyzing-user-agent-strings-general-approach/18346/4 "2014-06-28T05:22:44Z")

</div>

I just realize that the "user-agent" field comes from my Apache config, where I define a JSON logging format:

LogFormat "{ "@timestamp": "%{%Y-%m-%dT%H:%M:%S%z}t", "message": "%r", "host": "%v", "user-agent": "%{User-agent}i", "client": "%a", "duration\_usec": %D, "duration\_sec": %T, "status": %s, "size": %B, "request\_path": "%U", "request": "%U%q", "method": "%m", "referrer": "%{Referer}i" }" logstash\_ext\_json

everything else is in my first email.

On 27 juin 2014, at 22:03, Mark Dodwell wrote:

> Thanks, lots of useful stuff there.
> 
> --
> 
> Sent from Mailbox for iPhone
> 
> On Thu, Jun 26, 2014 at 12:34 AM, Patrick Proniewski [elasticsearch@patpro.net](mailto:elasticsearch@patpro.net) wrote:
> 
> Hi,
> 
> You should give [Useragent filter plugin | Logstash Reference [8.11] | Elastic](http://logstash.net/docs/1.4.2/filters/useragent) a try before anything else.
> 
> Here is the relevant part of logstash.conf I'm using:
> 
> filter {  
> if [type] == "apache" {  
> if [user-agent] != "-" and [user-agent] != "" {  
> useragent {  
> add\_tag =\> ["UA"]  
> source =\> "user-agent"  
> }  
> }  
> if "UA" in [tags] {  
> if [device] == "Other" { mutate { remove\_field =\> "device" } }  
> if [name] == "Other" { mutate { remove\_field =\> "name" } }  
> if [os] == "Other" { mutate { remove\_field =\> "os" } }  
> }  
> }  
> }
> 
> It retains the full user-agent field, and add nice fileds like "device", "major" and "minor" version, "name", "os", "os\_name", "os\_major" and "os\_minor".
> 
> sample:  
> "user-agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10\_8\_5) AppleWebKit/537.76.4 (KHTML, like Gecko) Version/6.1.4 Safari/537.76.4",  
> "name": "Safari",  
> "os": "Mac OS X 10.8.5",  
> "os\_name": "Mac OS X",  
> "os\_major": "10",  
> "os\_minor": "8",  
> "major": "6",  
> "minor": "1",  
> "patch": "4",
> 
> On 26 juin 2014, at 09:09, Mark Dodwell [mark@mkdynamic.co.uk](mailto:mark@mkdynamic.co.uk) wrote:
> 
> > I want to store a bunch of documents in elasticsearch (which represent a  
> > hit to a website) including the user agent of the client that made the  
> > original HTTP request.
> > 
> > Since user agent strings have a lot of variance, and the useful parts need  
> > parsing out (OS, browser, version etc.) I would like to be able to perform  
> > aggregations on those extracted features.
> > 
> > The simplest way I can think to do this would be to analyze the user agent  
> > string before indexing the document. The downside to this approach is as  
> > new/different user agent strings emerge (which is not unlikely) you would  
> > have to proactively update the parser.
> > 
> > This may be impossibly/undesirable for a number of reasons, but what I'd  
> > really like to do is index the raw user agent string and then perform the  
> > analysis/feature extraction post-hoc at query time. Any ideas/pointers on  
> > how to do this?
> > 
> > Aggregators? Custom analyzers? (How would you handle an update to the  
> > analyzer, would you need to re-run against all existing stored data?)
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/ed9bf030-f9bf-480a-88b1-a80421b9e79e%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ed9bf030-f9bf-480a-88b1-a80421b9e79e%40googlegroups.com).  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> 
> --  
> You received this message because you are subscribed to a topic in the Google Groups "elasticsearch" group.  
> To unsubscribe from this topic, visit [https://groups.google.com/d/topic/elasticsearch/H-sUPppQMp8/unsubscribe](https://groups.google.com/d/topic/elasticsearch/H-sUPppQMp8/unsubscribe).  
> To unsubscribe from this group and all its topics, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/48F3FC4E-D872-43B9-A60D-D2755094AC85%40patpro.net](https://groups.google.com/d/msgid/elasticsearch/48F3FC4E-D872-43B9-A60D-D2755094AC85%40patpro.net).  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/1403899436158.1403ba21%40Nodemailer](https://groups.google.com/d/msgid/elasticsearch/1403899436158.1403ba21%40Nodemailer).  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/DF64C8A3-2285-4B8C-9B56-2201C4649EF1%40patpro.net](https://groups.google.com/d/msgid/elasticsearch/DF64C8A3-2285-4B8C-9B56-2201C4649EF1%40patpro.net).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:19am UTC](https://discuss.elastic.co/t/storing-and-analyzing-user-agent-strings-general-approach/18346/5 "2017-07-06T01:19:13Z")

</div>


