# Extracting from Bind 9 log files

**URL:** <https://discuss.elastic.co/t/extracting-from-bind-9-log-files/75744>\
**Category:** Logstash\
**Created:** [February 20, 2017, 1:09pm UTC](https://discuss.elastic.co/t/extracting-from-bind-9-log-files/75744 "2017-02-20T13:09:35Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![HeadScratcher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/headscratcher/32/98085_2.png) [@HeadScratcher](https://discuss.elastic.co/u/HeadScratcher)\
**Post date:** [February 20, 2017, 1:09pm UTC](https://discuss.elastic.co/t/extracting-from-bind-9-log-files/75744/1 "2017-02-20T13:09:35Z")

</div>

So I've been basically testing and experimenting with the whole ELK and i'm impressed with its versatility and already have a few things being displayed from various data sources which pleases my boss and my team quite nicely, to the point where i get coffee made for me in the mornings LOL

I've now got something I'm just not able to get around and I've been trying for about a week to make it work.

message:client, 127.0.0.1#56692, ([microsoft.com](http://microsoft.com)):, query:, [microsoft.com](http://microsoft.com), IN, A, +, (127.0.0.1) @version:1 @timestamp:February 20th 2017, 12:46:47.533 path:/var/log/bind9/query.log host:ip-127.0.0.1 tags:\_grokparsefailure _id:AVpbj2TWl5ZobaQBDUM_ \_type:logs \_index:logstash-2017.02.20 \_score:1

I am trying to extract as a base metric either ([microsoft.com](http://microsoft.com)) or [microsoft.com](http://microsoft.com) to i can see the most queried domains on a DNS server. I've tried basic file filters, which show the data but do not allow me to query the 'message' field beyond showing the full field data.

I've tried reg ex, but i get inconsistent results, and i think thats due to the amount of time its taking to run on a larger log file ( note to self, must try a smaller log file to check that ) using /^(query?://)?([\da-z.-]+).([a-z.]{2,6})([/\w .-]_)_/?$/ but i must admit, my developer who gave me this was somewhat distracted by pokemon go at the time.

I've tried building the query using grok to parse out the fields so i can at least see what i need to see ...

```
 grok {
match => ["message", "%{WORD:query:,} %{WORD:query} %{WORD:IN}"]
}

```

But this doesn't seem to drop new fields into my kibana discovery console.

I peeled back to basics and started on a copy to the Apache access log i had and tried to build up on that, but i think i might be running in circles as this pretty much does the same thing as my existing conf file below

input {  
file {  
path =\> "/var/log/bind9/query.log"  
start\_position =\> beginning  
}  
}

filter {  
grok {  
match =\> ["message", "%{WORD:query:,} %{WORD:query} %{WORD:IN}"]  
}  
mutate {  
split =\> ["message", " "]  
}  
kv{  
field\_split =\> " "  
}  
}

output {  
elasticsearch {  
hosts =\> ["127.0.0.1:9200"]  
}  
}

Has anyone else ever had to extract a single bit of datum from the message field before that doesn't have any identifiable markers to work from ?

---

<div class="post-metadata">

**Author:** ![HeadScratcher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/headscratcher/32/98085_2.png) [@HeadScratcher](https://discuss.elastic.co/u/HeadScratcher)\
**Post date:** [February 20, 2017, 5:05pm UTC](https://discuss.elastic.co/t/extracting-from-bind-9-log-files/75744/2 "2017-02-20T17:05:27Z")

</div>

so narrowing it down by using this  
\s\*([-a-z-].[a-z]):\s\*

i get [microsoft.com](http://microsoft.com), IN, A, +, (127.0.0.1)... still working on removing or screening out the IN, A, +, (127.0.0.1) part

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [February 20, 2017, 9:33pm UTC](https://discuss.elastic.co/t/extracting-from-bind-9-log-files/75744/3 "2017-02-20T21:33:34Z")

</div>

I'd use a bunch of `%{NOTSPACE}`, cleaner than the regexp.

---

<div class="post-metadata">

**Author:** ![HeadScratcher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/headscratcher/32/98085_2.png) [@HeadScratcher](https://discuss.elastic.co/u/HeadScratcher)\
**Post date:** [February 21, 2017, 8:13am UTC](https://discuss.elastic.co/t/extracting-from-bind-9-log-files/75744/4 "2017-02-21T08:13:47Z")

</div>

tried a bunch of them last night including trying to mutate the output of my code, im not an expert by any imagination, but i know im doing something wrong, each time i read the grok parsing instructions and the grok patterns page ( [https://github.com/elastic/logstash/blob/v1.4.2/patterns/grok-patterns](https://github.com/elastic/logstash/blob/v1.4.2/patterns/grok-patterns) ) i just slid further and further into confusion.

I think i have to use a drop filter, but everything im reading says drop filters drop whole lines so im not sure... otherwise, i need to find a way to insert the results into a string and then query the string somehow.  
i can't believe no ones come across the need to extract and report on a single piece of datum before and not hit this wall.

---

<div class="post-metadata">

**Author:** ![HeadScratcher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/headscratcher/32/98085_2.png) [@HeadScratcher](https://discuss.elastic.co/u/HeadScratcher)\
**Post date:** [February 21, 2017, 8:15am UTC](https://discuss.elastic.co/t/extracting-from-bind-9-log-files/75744/5 "2017-02-21T08:15:56Z")

</div>

by the way

[http://grokconstructor.appspot.com](http://grokconstructor.appspot.com) FTW!!!!

ive gained a better understanding of grok patterns from testing with this and reading endless blogs and posts that only seem to be concerned with numbers and apache logs

---

<div class="post-metadata">

**Author:** ![HeadScratcher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/headscratcher/32/98085_2.png) [@HeadScratcher](https://discuss.elastic.co/u/HeadScratcher)\
**Post date:** [February 21, 2017, 10:28pm UTC](https://discuss.elastic.co/t/extracting-from-bind-9-log-files/75744/6 "2017-02-21T22:28:28Z")

</div>

\A%{NOTSPACE}%{SPACE}%{NOTSPACE}%{JAVALOGMESSAGE} gives me the same output as \s\*([-a-z-].[a-z]):\s\*

im definitely running in circles now

---

<div class="post-metadata">

**Author:** ![HeadScratcher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/headscratcher/32/98085_2.png) [@HeadScratcher](https://discuss.elastic.co/u/HeadScratcher)\
**Post date:** [February 23, 2017, 9:20am UTC](https://discuss.elastic.co/t/extracting-from-bind-9-log-files/75744/7 "2017-02-23T09:20:44Z")

</div>

ok, i need some help i just can't figure out what the correct pattern would be can anyone lend a hand ?

---

<div class="post-metadata">

**Author:** ![HeadScratcher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/headscratcher/32/98085_2.png) [@HeadScratcher](https://discuss.elastic.co/u/HeadScratcher)\
**Post date:** [February 23, 2017, 11:35am UTC](https://discuss.elastic.co/t/extracting-from-bind-9-log-files/75744/8 "2017-02-23T11:35:24Z")

</div>

finally managed to figure this out

input {  
file {  
path =\> "/var/log/bind9/query.log"  
start\_position =\> beginning  
}  
}

filter {  
grok {  
match =\> {"message" =\> "client %{IP:clientip}#%{POSINT:clientport} (%{GREEDYDATA:query}): query: %{GREEDYDATA:Target} IN %{GREEDYDATA:querytype} (%{IP:dns})"}  
}  
}

output {  
elasticsearch {  
hosts =\> ["127.0.0.1:9200"]  
}  
}

this is my conf file for extracting the DNS names of DNS queries into Kibana, it works, so now im mining data 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 23, 2017, 11:36am UTC](https://discuss.elastic.co/t/extracting-from-bind-9-log-files/75744/9 "2017-03-23T11:36:04Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
