# Getting data into Elasticsearch?

**URL:** <https://discuss.elastic.co/t/getting-data-into-elasticsearch/25666>\
**Category:** Elasticsearch\
**Created:** [July 15, 2015, 6:38pm UTC](https://discuss.elastic.co/t/getting-data-into-elasticsearch/25666 "2015-07-15T18:38:09Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![ziocho](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ziocho/32/3734_2.png) [@ziocho](https://discuss.elastic.co/u/ziocho)\
**Post date:** [July 15, 2015, 6:38pm UTC](https://discuss.elastic.co/t/getting-data-into-elasticsearch/25666/1 "2015-07-15T18:38:09Z")

</div>

Hello! This is my first large project in working with aggregating and searching data. I understand how to use Elasticsearch querying, however I'm unsure of how to actually give it the data I want.

I'm looking to index data from various sources (Confluence, GitHub, Google Docs, box, etc.). So far, I learned that I can get data in JSON format through their respective APIs, but how would I go about actually scripting the indexing through http(s)? While testing how Elasticsearch would work on the server, I had to access the API call in-browser to retrieve the JSON data, then copy that into the -XPOST call for indexing in Elasticsearch. There's definitely a better way to do this, haha

I'm proficient in Java and Ruby, as well as HTML, JavaScript and PHP. My end-goal is to be able to search for documents from those various services!

---

<div class="post-metadata">

**Author:** ![colings86](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/colings86/32/44960_2.png) [@colings86](https://discuss.elastic.co/u/colings86)\
**Post date:** [July 15, 2015, 8:34pm UTC](https://discuss.elastic.co/t/getting-data-into-elasticsearch/25666/2 "2015-07-15T20:34:26Z")

</div>

There are many official language clients for Elasticsearch includes clients for Ruby, JS, PHP and Java. Take a look at the documentation for your preferred language clients in the following link (under the clients section): [https://www.elastic.co/guide/index.html](https://www.elastic.co/guide/index.html)

Hope that helps

---

<div class="post-metadata">

**Author:** ![ziocho](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ziocho/32/3734_2.png) [@ziocho](https://discuss.elastic.co/u/ziocho)\
**Post date:** [July 15, 2015, 8:45pm UTC](https://discuss.elastic.co/t/getting-data-into-elasticsearch/25666/3 "2015-07-15T20:45:34Z")

</div>

Thanks for the reply!

I've looked through some of the client documentation, and I guess I'm looking for a recommendation for starting this project. My first goal is to be able to index documents using Confluence's API (with HTTP(s) calls) but I'm not sure whether it's easier to do with Logstash+plugins or with the Java client for this particular step.

---

<div class="post-metadata">

**Author:** ![colings86](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/colings86/32/44960_2.png) [@colings86](https://discuss.elastic.co/u/colings86)\
**Post date:** [July 16, 2015, 3:49pm UTC](https://discuss.elastic.co/t/getting-data-into-elasticsearch/25666/4 "2015-07-16T15:49:46Z")

</div>

I think there is a certain amount of personal preference here. I think you could do it with both approaches but I haven't personally used the http input for Logstash so I don't have an opinion on which way would be easier. Sorry.

Maybe other people will be able to comment

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [July 17, 2015, 2:41am UTC](https://discuss.elastic.co/t/getting-data-into-elasticsearch/25666/5 "2015-07-17T02:41:59Z")

</div>

I'd try both and see what works best for you 😄

---

<div class="post-metadata">

**Author:** ![ziocho](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ziocho/32/3734_2.png) [@ziocho](https://discuss.elastic.co/u/ziocho)\
**Post date:** [July 17, 2015, 5:11pm UTC](https://discuss.elastic.co/t/getting-data-into-elasticsearch/25666/6 "2015-07-17T17:11:45Z")

</div>

Thanks everyone! I've tried both the PHP and Ruby clients and found that documentation really pushed me towards the former. It's working out with multi\_match, and the conversion from JSON array to PHP code (or reverse) is simple enough! I can send queries and put the hits on the web page!

My next step is to aggregate data from those various sites now. Has anyone tried the Elasticsearch River App to crawl for content? I've looked into [some threads here](https://discuss.elastic.co/t/crawling-web-sites-and-indexing-the-extracted-content/3545/4) that discusses web crawlers, but the last comment was from four years ago and the suggestions there are probably not the best or current.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [July 18, 2015, 4:26am UTC](https://discuss.elastic.co/t/getting-data-into-elasticsearch/25666/7 "2015-07-18T04:26:35Z")

</div>

Rivers are deprecated anyway, so it's not worth using them.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:00am UTC](https://discuss.elastic.co/t/getting-data-into-elasticsearch/25666/8 "2017-07-06T00:00:40Z")

</div>


