# What is the best way to collect yarn application logs from hdfs?

**URL:** <https://discuss.elastic.co/t/what-is-the-best-way-to-collect-yarn-application-logs-from-hdfs/173641>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [March 24, 2019, 2:09pm UTC](https://discuss.elastic.co/t/what-is-the-best-way-to-collect-yarn-application-logs-from-hdfs/173641 "2019-03-24T14:09:33Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Drahkar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/drahkar/32/42648_2.png) [@Drahkar](https://discuss.elastic.co/u/Drahkar)\
**Post date:** [March 24, 2019, 2:09pm UTC](https://discuss.elastic.co/t/what-is-the-best-way-to-collect-yarn-application-logs-from-hdfs/173641/1 "2019-03-24T14:09:33Z")

</div>

We have a dynamic infrastructure for Hadoop, which means the yarn application logs only exist for a limited period of time.

Currently we are using Splunk HadoopConnect to ingest those logs as a live feed into Splunk. However this requires the installation of the full Splunk server on the cluster to accomplish, which is not only resource intensive, but not the most efficient thing to do every time we spin up a dynamic Hadoop cluster.

Does Elasticsearch, Logstash, etc have an alternative to HadoopConnect that could be used to collect the Yarn application logs out of HDFS and feed them into the ELK stack?

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [April 5, 2019, 8:37pm UTC](https://discuss.elastic.co/t/what-is-the-best-way-to-collect-yarn-application-logs-from-hdfs/173641/2 "2019-04-05T20:37:43Z")

</div>

I don't have much experience with Splunk's HadoopConnect or know much about what it even is, but in terms of collecting log files, I would suggest something like starting a [Filebeat](https://www.elastic.co/guide/en/beats/filebeat/current/index.html) along side your NodeManager instances as they come online and tearing it down as they come offline.

Granted, this means running a data shipper process on all the nodes that you would want to collect data from (my recollection is that you can get the YARN application logs from the local directories on the NodeManagers/ResourceManagers that launch them, but your setup may be different than what I'm used to).

Additionally, I know of a tool that one of the engineers at Elastic has made public called [FSCrawler](https://github.com/dadoonet/fscrawler), which has a blurb about indexing data through an [HDFS NFS Gateway](https://fscrawler.readthedocs.io/en/fscrawler-2.6/user/tips.html#indexing-from-hdfs-drive). Maybe that would be easier to implement instead of using a sidecar datashipper on dynamic infrastructure?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 3, 2019, 8:37pm UTC](https://discuss.elastic.co/t/what-is-the-best-way-to-collect-yarn-application-logs-from-hdfs/173641/3 "2019-05-03T20:37:49Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
