# Index dataset in text file

**URL:** <https://discuss.elastic.co/t/index-dataset-in-text-file/60870>\
**Category:** Elasticsearch\
**Created:** [September 19, 2016, 11:04am UTC](https://discuss.elastic.co/t/index-dataset-in-text-file/60870 "2016-09-19T11:04:35Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![maniut](https://avatars.discourse-cdn.com/v4/letter/m/bcef8e/32.png) [@maniut](https://discuss.elastic.co/u/maniut)\
**Post date:** [September 19, 2016, 11:04am UTC](https://discuss.elastic.co/t/index-dataset-in-text-file/60870/1 "2016-09-19T11:04:35Z")

</div>

Hi,  
I am new to elastic search .i have to index data(half million entries) available in text file. i don't understand how i have to start doing it

---

<div class="post-metadata">

**Author:** ![mainec](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mainec/32/5557_2.png) [@mainec](https://discuss.elastic.co/u/mainec)\
**Post date:** [September 19, 2016, 11:07am UTC](https://discuss.elastic.co/t/index-dataset-in-text-file/60870/2 "2016-09-19T11:07:41Z")

</div>

Did you look at Logstash or Filebeat already?

[https://www.elastic.co/guide/en/logstash/current/plugins-inputs-file.html](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-file.html)

[https://www.elastic.co/products/beats/filebeat](https://www.elastic.co/products/beats/filebeat)

Other options would include writing code to parse the file and send separate index requests/ bulk index requests to ES. However I've found using one of the two options above much faster in terms of development time needed in the past.

Hope this helps,  
Isabel

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [September 19, 2016, 11:31am UTC](https://discuss.elastic.co/t/index-dataset-in-text-file/60870/3 "2016-09-19T11:31:06Z")

</div>

Adding to Isabel's answer that if it's non structured data (well, just text), you can have a look at [FSCrawler](https://github.com/dadoonet/fscrawler) project.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:19pm UTC](https://discuss.elastic.co/t/index-dataset-in-text-file/60870/4 "2017-07-05T22:19:17Z")

</div>


