# Loading JSON to ElasticSearch

**URL:** <https://discuss.elastic.co/t/loading-json-to-elasticsearch/15451>\
**Category:** Elasticsearch\
**Created:** [January 28, 2014, 3:45pm UTC](https://discuss.elastic.co/t/loading-json-to-elasticsearch/15451 "2014-01-28T15:45:56Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![IronMike](https://avatars.discourse-cdn.com/v4/letter/i/e9a140/32.png) [@IronMike](https://discuss.elastic.co/u/IronMike)\
**Post date:** [January 28, 2014, 3:45pm UTC](https://discuss.elastic.co/t/loading-json-to-elasticsearch/15451/1 "2014-01-28T15:45:56Z")

</div>

I would like to get your perspective on how to load json to index server in  
my scenario.  
We have about 15 million documents in html/pdf/... on Server 1  
I would like to process the data and convert to json on server 2  
I would like the indexer to index json n a separate machine/server server 3

Ideally I thought on Server 2, as I prepare json and have it ready in  
memory, I can feed it to indexer. But since data processing is cpu  
intensive, I want indexing to be done on a separate machines/server.  
How do you guys deal with this since I can no longer feed in-memory json to  
the indexer on separate machine? Do I just grab files from server 2 and  
index them then?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/05b977ac-00d0-45c0-9e58-8df523e6978c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/05b977ac-00d0-45c0-9e58-8df523e6978c%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 28, 2014, 6:05pm UTC](https://discuss.elastic.co/t/loading-json-to-elasticsearch/15451/2 "2014-01-28T18:05:32Z")

</div>

Did you try [https://github.com/dadoonet/fsriver](https://github.com/dadoonet/fsriver)?  
Never tested it with so many docs but may be it could help you here?

If you have already generated json files on a server, then I would recommend trying logstash to send them into elasticsearch.

My 2 cents

--  
David Pilato | Technical Advocate | [Elasticsearch.com](http://Elasticsearch.com)  
@dadoonet | @elasticsearchfr

Le 28 janvier 2014 at 16:46:06, ZenMaster80 ([sabdalla80@gmail.com](mailto:sabdalla80@gmail.com)) a écrit:

I would like to get your perspective on how to load json to index server in my scenario.  
We have about 15 million documents in html/pdf/... on Server 1  
I would like to process the data and convert to json on server 2  
I would like the indexer to index json n a separate machine/server server 3

## Ideally I thought on Server 2, as I prepare json and have it ready in memory, I can feed it to indexer. But since data processing is cpu intensive, I want indexing to be done on a separate machines/server. How do you guys deal with this since I can no longer feed in-memory json to the indexer on separate machine? Do I just grab files from server 2 and index them then?

You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/05b977ac-00d0-45c0-9e58-8df523e6978c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/05b977ac-00d0-45c0-9e58-8df523e6978c%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/etPan.52e7f16c.74b0dc51.ec%40MacBook-Air-de-David.local](https://groups.google.com/d/msgid/elasticsearch/etPan.52e7f16c.74b0dc51.ec%40MacBook-Air-de-David.local).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![IronMike](https://avatars.discourse-cdn.com/v4/letter/i/e9a140/32.png) [@IronMike](https://discuss.elastic.co/u/IronMike)\
**Post date:** [January 28, 2014, 6:56pm UTC](https://discuss.elastic.co/t/loading-json-to-elasticsearch/15451/3 "2014-01-28T18:56:10Z")

</div>

Thanks David, I will certainly look into hashtag. Do you think it is a good  
idea to separate data analysis and indexing into 2 different machines since  
both require lots of cpu time.  
If I use hashtag to send files over to ES, will I be able to use native  
Java API or http, and is there any preference to the API? I have noticed  
there are somethings that aren't very easy and may be don't even work in  
the native API?  
Thanks again.

On Tuesday, January 28, 2014 1:05:32 PM UTC-5, David Pilato wrote:

> Did you try [GitHub - dadoonet/fscrawler: Elasticsearch File System Crawler (FS Crawler)](https://github.com/dadoonet/fsriver)?  
> Never tested it with so many docs but may be it could help you here?
> 
> If you have already generated json files on a server, then I would  
> recommend trying logstash to send them into elasticsearch.
> 
> My 2 cents
> 
> --  
> _David Pilato_ | _Technical Advocate_ | _[Elasticsearch.com](http://Elasticsearch.com)_  
> @dadoonet [https://twitter.com/dadoonet](https://twitter.com/dadoonet) | @elasticsearchfr[https://twitter.com/elasticsearchfr](https://twitter.com/elasticsearchfr)
> 
> Le 28 janvier 2014 at 16:46:06, ZenMaster80 ([sabda...@gmail.com](mailto:sabda...@gmail.com)\<javascript:\>)  
> a écrit:
> 
> I would like to get your perspective on how to load json to index server  
> in my scenario.  
> We have about 15 million documents in html/pdf/... on Server 1  
> I would like to process the data and convert to json on server 2  
> I would like the indexer to index json n a separate machine/server server 3
> 
> ## Ideally I thought on Server 2, as I prepare json and have it ready in memory, I can feed it to indexer. But since data processing is cpu intensive, I want indexing to be done on a separate machines/server. How do you guys deal with this since I can no longer feed in-memory json to the indexer on separate machine? Do I just grab files from server 2 and index them then?
> 
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/05b977ac-00d0-45c0-9e58-8df523e6978c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/05b977ac-00d0-45c0-9e58-8df523e6978c%40googlegroups.com)  
> .  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/a02427ec-a3d8-484f-9cfb-2ba7628192b1%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/a02427ec-a3d8-484f-9cfb-2ba7628192b1%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![IronMike](https://avatars.discourse-cdn.com/v4/letter/i/e9a140/32.png) [@IronMike](https://discuss.elastic.co/u/IronMike)\
**Post date:** [January 28, 2014, 7:13pm UTC](https://discuss.elastic.co/t/loading-json-to-elasticsearch/15451/4 "2014-01-28T19:13:01Z")

</div>

Thanks David, I will certainly look into logstash. Do you think it is a  
good idea to separate data analysis and indexing into 2 different machines  
since both require lots of cpu time.  
If I use logstash to send files over to ES, will I be able to use native  
Java API or http, and is there any preference to the API? I have noticed  
there are somethings that aren't very easy and may be don't even work in  
the native API?  
Thanks again

On Tuesday, January 28, 2014 1:05:32 PM UTC-5, David Pilato wrote:

> Did you try [GitHub - dadoonet/fscrawler: Elasticsearch File System Crawler (FS Crawler)](https://github.com/dadoonet/fsriver)?  
> Never tested it with so many docs but may be it could help you here?
> 
> If you have already generated json files on a server, then I would  
> recommend trying logstash to send them into elasticsearch.
> 
> My 2 cents
> 
> --  
> _David Pilato_ | _Technical Advocate_ | _[Elasticsearch.com](http://Elasticsearch.com)_  
> @dadoonet [https://twitter.com/dadoonet](https://twitter.com/dadoonet) | @elasticsearchfr[https://twitter.com/elasticsearchfr](https://twitter.com/elasticsearchfr)
> 
> Le 28 janvier 2014 at 16:46:06, ZenMaster80 ([sabda...@gmail.com](mailto:sabda...@gmail.com)\<javascript:\>)  
> a écrit:
> 
> I would like to get your perspective on how to load json to index server  
> in my scenario.  
> We have about 15 million documents in html/pdf/... on Server 1  
> I would like to process the data and convert to json on server 2  
> I would like the indexer to index json n a separate machine/server server 3
> 
> ## Ideally I thought on Server 2, as I prepare json and have it ready in memory, I can feed it to indexer. But since data processing is cpu intensive, I want indexing to be done on a separate machines/server. How do you guys deal with this since I can no longer feed in-memory json to the indexer on separate machine? Do I just grab files from server 2 and index them then?
> 
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/05b977ac-00d0-45c0-9e58-8df523e6978c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/05b977ac-00d0-45c0-9e58-8df523e6978c%40googlegroups.com)  
> .  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/f536d58c-89ab-4609-b5ca-cef44e2b879a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f536d58c-89ab-4609-b5ca-cef44e2b879a%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 30, 2014, 8:01am UTC](https://discuss.elastic.co/t/loading-json-to-elasticsearch/15451/5 "2014-01-30T08:01:59Z")

</div>

Logstash uses native API when you choose elasticsearch output: [http://logstash.net/docs/1.3.3/outputs/elasticsearch](http://logstash.net/docs/1.3.3/outputs/elasticsearch)  
About machines separation, I would say that you should test it. If your nodes are not really intensively used (CPU / IO), you can probably use the same machine for extracting content and produce JSON docs.

HTH

--  
David Pilato | Technical Advocate | [Elasticsearch.com](http://Elasticsearch.com)  
@dadoonet | @elasticsearchfr

Le 28 janvier 2014 at 20:14:07, ZenMaster80 ([sabdalla80@gmail.com](mailto:sabdalla80@gmail.com)) a écrit:

Thanks David, I will certainly look into logstash. Do you think it is a good idea to separate data analysis and indexing into 2 different machines since both require lots of cpu time.  
If I use logstash to send files over to ES, will I be able to use native Java API or http, and is there any preference to the API? I have noticed there are somethings that aren't very easy and may be don't even work in the native API?  
Thanks again

On Tuesday, January 28, 2014 1:05:32 PM UTC-5, David Pilato wrote:  
Did you try [https://github.com/dadoonet/fsriver](https://github.com/dadoonet/fsriver)?  
Never tested it with so many docs but may be it could help you here?

If you have already generated json files on a server, then I would recommend trying logstash to send them into elasticsearch.

My 2 cents

--  
David Pilato | Technical Advocate | [Elasticsearch.com](http://Elasticsearch.com)  
@dadoonet | @elasticsearchfr

Le 28 janvier 2014 at 16:46:06, ZenMaster80 ([sabda...@gmail.com](mailto:sabda...@gmail.com)) a écrit:

I would like to get your perspective on how to load json to index server in my scenario.  
We have about 15 million documents in html/pdf/... on Server 1  
I would like to process the data and convert to json on server 2  
I would like the indexer to index json n a separate machine/server server 3

## Ideally I thought on Server 2, as I prepare json and have it ready in memory, I can feed it to indexer. But since data processing is cpu intensive, I want indexing to be done on a separate machines/server. How do you guys deal with this since I can no longer feed in-memory json to the indexer on separate machine? Do I just grab files from server 2 and index them then?

## You received this message because you are subscribed to the Google Groups "elasticsearch" group. To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com). To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/05b977ac-00d0-45c0-9e58-8df523e6978c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/05b977ac-00d0-45c0-9e58-8df523e6978c%40googlegroups.com). For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/f536d58c-89ab-4609-b5ca-cef44e2b879a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f536d58c-89ab-4609-b5ca-cef44e2b879a%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/etPan.52ea06f7.41b71efb.45fa%40MacBook-Air-de-David.local](https://groups.google.com/d/msgid/elasticsearch/etPan.52ea06f7.41b71efb.45fa%40MacBook-Air-de-David.local).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:54am UTC](https://discuss.elastic.co/t/loading-json-to-elasticsearch/15451/6 "2017-07-06T01:54:02Z")

</div>


