# How to index HDFS data

**URL:** <https://discuss.elastic.co/t/how-to-index-hdfs-data/59129>\
**Category:** Elasticsearch\
**Created:** [August 28, 2016, 6:02pm UTC](https://discuss.elastic.co/t/how-to-index-hdfs-data/59129 "2016-08-28T18:02:47Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![johan313](https://avatars.discourse-cdn.com/v4/letter/j/b9bd4f/32.png) [@johan313](https://discuss.elastic.co/u/johan313)\
**Post date:** [August 28, 2016, 6:02pm UTC](https://discuss.elastic.co/t/how-to-index-hdfs-data/59129/1 "2016-08-28T18:02:47Z")

</div>

Hello,

I'm prototyping the use of ES and Hadoop for a project but I cannot figure out the most obvious.

I have a hadoop cluster that contains some log data on HDFS. I installed ES Yarn according to this guide: [https://www.elastic.co/guide/en/elasticsearch/hadoop/current/ey-usage.html](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/ey-usage.html). Elasticsearch seems to work properly, however it stores all data locally. This was mentioned as the default storage solution in the guide, so ok.

Question one: ES created an index called "Titan" during the installation. What is this? Looking at the content is has nothing to do with any data I have put into HDFS.

Question two: What is the proper way to read the HDFS data into en ES index? I feel really stupid, but besides writting an application that pushes it through REST I could not figure this out. Is there any out-of-the-box support for populating ES Indexes?

R  
Johan

---

<div class="post-metadata">

**Author:** ![jasontedor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jasontedor/32/66992_2.png) [@jasontedor](https://discuss.elastic.co/u/jasontedor)\
**Post date:** [August 28, 2016, 8:37pm UTC](https://discuss.elastic.co/t/how-to-index-hdfs-data/59129/2 "2016-08-28T20:37:14Z")

</div>

> [@johan313](#):
>
> ES created an index called "Titan" during the installation. What is this?

That's not possible, names of indexes must be lowercase. Maybe you meant `titan`? Still, that's on your end. Perhaps you sent a post request to `/titan` and have auto index creation enabled?

> [@johan313](#):
>
> What is the proper way to read the HDFS data into en ES index?

You can use the [elasticsearch-hadoop](https://github.com/elastic/elasticsearch-hadoop) framework.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:24pm UTC](https://discuss.elastic.co/t/how-to-index-hdfs-data/59129/3 "2017-07-05T22:24:34Z")

</div>


