# Google Compute Engine systemd-resolved errors

**URL:** <https://discuss.elastic.co/t/google-compute-engine-systemd-resolved-errors/185757>\
**Category:** Beats\
**Tags:** packetbeat\
**Created:** [June 14, 2019, 4:40am UTC](https://discuss.elastic.co/t/google-compute-engine-systemd-resolved-errors/185757 "2019-06-14T04:40:40Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![baerrach](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/baerrach/32/48119_2.png) [@baerrach](https://discuss.elastic.co/u/baerrach)\
**Post date:** [June 14, 2019, 4:40am UTC](https://discuss.elastic.co/t/google-compute-engine-systemd-resolved-errors/185757/1 "2019-06-14T04:40:40Z")

</div>

I've just installed packetbeats on our Google Compute Engine, and I have been watching a days worth of visualizations.

The visualization for "Errors vs successful transactions [Packetbeat] ECS" has a much higher error rate than I expected, around 40-60% errors. The majority of which are from **process.name: systemd-resolved** with two **error.messages**

- Another query with the same DNS ID from this client was received so this query was closed without receiving a response
- Response: received without an associated Query

My google-fu is failing me and I can find nothing with either of these two error messages.

Has anyone else noticed similar problems?  
Is this something wrong with my configuration of packetbeats?  
Or is this something I need to investigate systemd-resolved for?

Cheers  
Barrie

---

<div class="post-metadata">

**Author:** ![andrewkroh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrewkroh/32/3784_2.png) [@andrewkroh](https://discuss.elastic.co/u/andrewkroh)\
**Post date:** [June 19, 2019, 11:56am UTC](https://discuss.elastic.co/t/google-compute-engine-systemd-resolved-errors/185757/2 "2019-06-19T11:56:29Z")

</div>

Based on your description of the two errors I suspect that the the systemd-resolved is re-sending DNS requests (maybe if it doesn't receive a DNS response within some timeout). Packetbeat's DNS protocol analyzer has a pretty simple state machine. It receives a request and expects a response. In this case it gets a request, another request (that overwrites the first based on request ID), a DNS response, then another response to which Packetbeat no longer has a request to match it with since the previous response closed the state for that ID.

If you capture a PCAP trace of the DNS traffic on port 53 (`sudo tcpdump -i <interface> -w dns-capture.pcap port 53`) you'll be able to confirm this by looking at the request and response traffic in a tool like Wireshark. If the PCAP data does not prove this theory then it could some other issue like dropped packets in which case we can look at optimizing your configuration (like using af\_packet as shown [here](https://www.elastic.co/guide/en/beats/packetbeat/7.x/configuration-interfaces.html#_sniffing_configuration_options) and a few other things).

I think it would be possible to improve the state machine in Packetbeat's DNS protocol analyzer to avoid these errors by adding a counter to track the number of pending requests.

---

<div class="post-metadata">

**Author:** ![baerrach](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/baerrach/32/48119_2.png) [@baerrach](https://discuss.elastic.co/u/baerrach)\
**Post date:** [June 20, 2019, 2:00am UTC](https://discuss.elastic.co/t/google-compute-engine-systemd-resolved-errors/185757/3 "2019-06-20T02:00:46Z")

</div>

Thanks.

Its been an eon since I've done pcap debugging.

So I've grabbed everything from any interface `sudo tcpdump -i any -w dns-capture.pcap port 53` over a 10 minute window.

Here is a small sample of that file

 ![Wireshark_heNKHd62sW](https://us1.discourse-cdn.com/elastic/original/3X/9/f/9fa127724c070365b8ffd7f9b3d6a88a969dfd87.png)

It doesn't look like its resending the query, I can see the response match up for every request.  
It is however sequentially asking for resolution of the same name quite quickly, why it isn't caching these queries is another question.

Now, I think I have the same information displayed in Kibana.

 ![firefox_2bJnHFo3Gy](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1aa77c96150a30788b3d85b64f458d24d228ca8d.png)

You can see that packetbeats isn't able to group these correctly.

I am already using these optimization values

```auto
packetbeat.interfaces.type: af_packet
packetbeat.interfaces.buffer_size_mb: 100 

```

How do I look for dropped packets in packetbeats?  
Or what should I investigate next.

Cheers

---

<div class="post-metadata">

**Author:** ![andrewkroh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrewkroh/32/3784_2.png) [@andrewkroh](https://discuss.elastic.co/u/andrewkroh)\
**Post date:** [June 20, 2019, 4:05am UTC](https://discuss.elastic.co/t/google-compute-engine-systemd-resolved-errors/185757/4 "2019-06-20T04:05:56Z")

</div>

What version of Packetbeat? What OS version and kernel version is it?

Can you run test using `pcap` instead of `af_packet`?

---

<div class="post-metadata">

**Author:** ![baerrach](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/baerrach/32/48119_2.png) [@baerrach](https://discuss.elastic.co/u/baerrach)\
**Post date:** [June 20, 2019, 4:55am UTC](https://discuss.elastic.co/t/google-compute-engine-systemd-resolved-errors/185757/5 "2019-06-20T04:55:36Z")

</div>

$ sudo packetbeat version  
packetbeat version 7.1.1 (amd64), libbeat 7.1.1 [3358d9a5a09e3c6709a2d3aaafde628ea34e8419 built 2019-05-23 13:15:09 +0000 UTC]

$ uname -a  
Linux my-machine 4.15.0-1034-gcp #36-Ubuntu SMP Thu Jun 6 14:47:38 UTC 2019 x86\_64 x86\_64 x86\_64 GNU/Linux

I've switched to `packetbeat.interfaces.type: pcap` and restarted packetbeat, will get back to you about errors.

---

<div class="post-metadata">

**Author:** ![baerrach](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/baerrach/32/48119_2.png) [@baerrach](https://discuss.elastic.co/u/baerrach)\
**Post date:** [June 20, 2019, 5:16am UTC](https://discuss.elastic.co/t/google-compute-engine-systemd-resolved-errors/185757/6 "2019-06-20T05:16:36Z")

</div>

The change has been in for almost 20 minutes and the visualization `[Packetbeat] DNS Overview ECS` show 0 errors.

This would indicate it works with `pcap` but not with `af_packet`.

---

<div class="post-metadata">

**Author:** ![andrewkroh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrewkroh/32/3784_2.png) [@andrewkroh](https://discuss.elastic.co/u/andrewkroh)\
**Post date:** [June 20, 2019, 6:46pm UTC](https://discuss.elastic.co/t/google-compute-engine-systemd-resolved-errors/185757/7 "2019-06-20T18:46:21Z")

</div>

This sounds like [https://github.com/elastic/beats/issues/621](https://github.com/elastic/beats/issues/621).

---

<div class="post-metadata">

**Author:** ![baerrach](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/baerrach/32/48119_2.png) [@baerrach](https://discuss.elastic.co/u/baerrach)\
**Post date:** [June 21, 2019, 2:03am UTC](https://discuss.elastic.co/t/google-compute-engine-systemd-resolved-errors/185757/8 "2019-06-21T02:03:10Z")

</div>

Yes, it looks like it might be the same issue.

Unfortunately I don't have the technical background to do anything about it.

I'm happy to try running debug versions of packetbeats to check whether it fixes it.

Our setup on GCP is relatively new, we are running in zone australia-southeast1-b on an n1-standard-2 (2 vCPUs, 7.5 GB memory) so I expect that spinning up a compute engine with the right OS will be able to reproduce the problem.

As I've got a workaround, running pcap, I'll live with that for now.

Thanks for your help.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 19, 2019, 2:03am UTC](https://discuss.elastic.co/t/google-compute-engine-systemd-resolved-errors/185757/9 "2019-07-19T02:03:12Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
