# Issue while using Elastic Search (ESTap) for processing data

**URL:** <https://discuss.elastic.co/t/issue-while-using-elastic-search-estap-for-processing-data/92791>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [July 12, 2017, 10:25am UTC](https://discuss.elastic.co/t/issue-while-using-elastic-search-estap-for-processing-data/92791 "2017-07-12T10:25:02Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Kunal\_Ghosh](https://avatars.discourse-cdn.com/v4/letter/k/838e76/32.png) [@Kunal\_Ghosh](https://discuss.elastic.co/u/Kunal_Ghosh)\
**Post date:** [July 12, 2017, 10:25am UTC](https://discuss.elastic.co/t/issue-while-using-elastic-search-estap-for-processing-data/92791/1 "2017-07-12T10:25:02Z")

</div>

Hi, I am new to the Elastic Search and I am stuck with an issue. I am developing an application with cascading api.  
I am processing 10 million rows of data with 43 columns. Now my issue is when i am dumping data to sink tap ( using default hfs sink tap) it takes me 1 -2 minutes to completely dump the data but when i use ESTap instead of hfs tap it takes 1 hour. While configuring the elastic search we configured 3 nodes with all nodes acting as master as well as data nodes and "bootstrap.memory\_lock: true" , rest is kept as default settings.  
Do I need to change the configuration so that the process takes less time? Please help. Thanks in Advance.

String inputPath = args[0]+File.separator+"10\_million\_rows.csv";  
Tap inputTap = new Hfs( new TextDelimited( new Fields( "empid","gender","title","nameset","surname","city","statefull","zipcode","header1","header2","header3","header4","header5","header6","header7","header8","header9","header10","header11","header12","header13","header14","header15","header16","header17","header18","header19","header20","header21","header22","header23","header24","header25","header26","header27","header28","header29","header30","header31","header32","expr1","expr2","expr3","expr4" ) ,true , "," ), inputPath );  
Pipe pipe = new Pipe("pipe");

/\*  
**Tap sinkTap = new Hfs( new TextDelimited( new Fields( "empid","gender","title","nameset","surname","city","statefull","zipcode","header1","header2","header3","header4","header5","header6","header7","header8","header9","header10","header11","header12","header13","header14","header15","header16","header17","header18","header19","header20","header21","header22","header23","header24","header25","header26","header27","header28","header29","header30","header31","header32","expr1","expr2","expr3","expr4" ) ,true , "," ), "/hdfsdata/output" );**  
\*/

_Tap sinkTap = new EsTap("master-host",9200,"index1/type1", new Fields( "empid","gender","title","nameset","surname","city","statefull","zipcode","header1","header2","header3","header4","header5","header6","header7","header8","header9","header10","header11","header12","header13","header14","header15","header16","header17","header18","header19","header20","header21","header22","header23","header24","header25","header26","header27","header28","header29","header30","header31","header32","expr1","expr2","expr3","expr4" ));_

FlowDef flowDef = FlowDef.flowDef()  
.addSource(pipe,inputTap)  
.addTailSink(pipe,sinkTap);

Properties properties = new Properties();  
Flow flow = new Hadoop2MR1FlowConnector(properties).connect(flowDef);  
flow.complete();

# ---------------------------------- Node1 -----------------------------------

[cluster.name](http://cluster.name): electrik-io  
[node.name](http://node.name): master  
node.master: true  
node.data: true  
path.data: "/secondary/elasticsearch/data"  
path.logs: "/secondary/elasticsearch/logs"  
bootstrap.memory\_lock: true  
bootstrap.system\_call\_filter: false  
network.host: ["master", "localhost"]  
http.port: 9200  
transport.tcp.port: 9300  
http.enabled: true  
discovery.zen.ping.unicast.hosts: ["master", "slave1", "slave2"]  
discovery.zen.minimum\_master\_nodes: 3

# ---------------------------------- Node2 -----------------------------------

[cluster.name](http://cluster.name): electrik-io  
[node.name](http://node.name): slave1  
node.master: true  
node.data: true  
path.data: "/secondary/elasticsearch/data"  
path.logs: "/secondary/elasticsearch/logs"  
bootstrap.memory\_lock: true  
bootstrap.system\_call\_filter: false  
network.host: ["slave1", "localhost"]  
http.port: 9200  
transport.tcp.port: 9300  
http.enabled: true  
discovery.zen.ping.unicast.hosts: ["master", "slave1", "slave2"]  
discovery.zen.minimum\_master\_nodes: 3

# ---------------------------------- Node3 -----------------------------------

[cluster.name](http://cluster.name): electrik-io  
[node.name](http://node.name): slave2  
node.master: true  
node.data: true  
path.data: "/secondary/elasticsearch/data"  
path.logs: "/secondary/elasticsearch/logs"  
bootstrap.memory\_lock: true  
bootstrap.system\_call\_filter: false  
network.host: ["slave2", "localhost"]  
http.port: 9200  
transport.tcp.port: 9300  
http.enabled: true  
discovery.zen.ping.unicast.hosts: ["master", "slave1", "slave2"]  
discovery.zen.minimum\_master\_nodes: 3

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 12, 2017, 10:56am UTC](https://discuss.elastic.co/t/issue-while-using-elastic-search-estap-for-processing-data/92791/2 "2017-07-12T10:56:31Z")

</div>

I moved the question to #elasticsearch-and-hadoop

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [July 12, 2017, 6:09pm UTC](https://discuss.elastic.co/t/issue-while-using-elastic-search-estap-for-processing-data/92791/3 "2017-07-12T18:09:34Z")

</div>

@Kunal_Ghosh One option for optimizing your job is to tune the batch output sizes for Elasticsearch on the EsTap configuration. Feel free to peruse our [documentation](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/performance.html) on performance for more information.

---

<div class="post-metadata">

**Author:** ![Kunal\_Ghosh](https://avatars.discourse-cdn.com/v4/letter/k/838e76/32.png) [@Kunal\_Ghosh](https://discuss.elastic.co/u/Kunal_Ghosh)\
**Post date:** [July 13, 2017, 1:20pm UTC](https://discuss.elastic.co/t/issue-while-using-elastic-search-estap-for-processing-data/92791/4 "2017-07-13T13:20:51Z")

</div>

**@james.baiera Thanks for the prompt response !**  
I am using Elastic Search Version 5.5.0, the default figure for number of shards is 5 so in my case i have 3 nodes = 15 shards. Will this have adverse effect on performance??  
Also how do I configure number of shards per node in Elastic Search 5.5.0 ? In earlier versions it was in elasticsearch.yml file but now I could not find where to configure it.

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [July 17, 2017, 2:32am UTC](https://discuss.elastic.co/t/issue-while-using-elastic-search-estap-for-processing-data/92791/5 "2017-07-17T02:32:04Z")

</div>

I'm not sure I understand how your math pans out: Shard counts are per index and are distributed across the cluster. Shard counts are not based on the number of nodes you have, unless you mean that you have an index with 15 shards and it happens to have 5 shards on each of your 3 nodes?

Shards are configurable at index creation time via the index's settings. If no settings are provided, the default number of shards and replicas for the cluster are used.

In terms of execution time for your job, have you modified the batch sizes for the job at all? You can tune the maximum batch request sizes by using `es.batch.size.bytes` (default 1mb), `es.batch.size.entries` (default 1000), and `es.batch.write.refresh` (default true). These are often good starting numbers but tend to be fairly limiting for larger datasets.

Another thing that might be worth tuning is the job's parallelism. You want to make sure you have a reasonable rate of indexing with enough writing clients to push your data, but not so many clients that you start seeing rejections from Elasticsearch. Remember a good rule of thumb is that each node can handle about 50 outstanding bulk requests in its queue at a time, so you want to be below this mark for concurrent writers.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 14, 2017, 2:35am UTC](https://discuss.elastic.co/t/issue-while-using-elastic-search-estap-for-processing-data/92791/6 "2017-08-14T02:35:43Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
