# Help Needed in improving the data ingestion time

**URL:** <https://discuss.elastic.co/t/help-needed-in-improving-the-data-ingestion-time/107767>\
**Category:** Logstash\
**Created:** [November 15, 2017, 2:25pm UTC](https://discuss.elastic.co/t/help-needed-in-improving-the-data-ingestion-time/107767 "2017-11-15T14:25:10Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![arun\_s](https://avatars.discourse-cdn.com/v4/letter/a/e9a140/32.png) [@arun\_s](https://discuss.elastic.co/u/arun_s)\
**Post date:** [November 15, 2017, 2:25pm UTC](https://discuss.elastic.co/t/help-needed-in-improving-the-data-ingestion-time/107767/1 "2017-11-15T14:25:10Z")

</div>

Hi

I am having csv file of size 1.6 GB and i used default configuration of logstash and elasticsearch to upload the data via logstash , it took more than 15hrs and when i tried with LS\_HEAP\_SIZE=2gb the data uploaded in 4 hrs. I am doing some POC with system configuration 64 bit ubuntu machine with dual core and 4 gb ram.

Can some one help me to improve the data uploading rate as my real time data will be more than 1GB and we cant wait for long time for each uploading.

---

<div class="post-metadata">

**Author:** ![arun\_s](https://avatars.discourse-cdn.com/v4/letter/a/e9a140/32.png) [@arun\_s](https://discuss.elastic.co/u/arun_s)\
**Post date:** [November 15, 2017, 2:27pm UTC](https://discuss.elastic.co/t/help-needed-in-improving-the-data-ingestion-time/107767/2 "2017-11-15T14:27:09Z")

</div>

actually i used LS\_HEAP\_SIZE= 2gb along with -b 1000 and -w 2

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [November 15, 2017, 2:30pm UTC](https://discuss.elastic.co/t/help-needed-in-improving-the-data-ingestion-time/107767/3 "2017-11-15T14:30:45Z")

</div>

Are the CPUs saturated during the hours Logstash is working on the data? What kind of filters do you have? What's the event rate? Have you looked into using the montoring APIs introduced in Logstash 5 to analyze where the bottlenecks are?

---

<div class="post-metadata">

**Author:** ![arun\_s](https://avatars.discourse-cdn.com/v4/letter/a/e9a140/32.png) [@arun\_s](https://discuss.elastic.co/u/arun_s)\
**Post date:** [November 15, 2017, 2:34pm UTC](https://discuss.elastic.co/t/help-needed-in-improving-the-data-ingestion-time/107767/4 "2017-11-15T14:34:44Z")

</div>

Hi Magnus

i am using the following filter

input {  
file {  
path =\> "\home\Projects\kibana\DataSet\test.csv"  
start\_position =\> "beginning"  
sincedb\_path =\> "/dev/null"  
}  
}

filter {  
csv {  
separator =\> ","  
columns =\> ["time","id\_geo","gw","id","server\_id","good","responses"]

```
} 

date{
 match => ["time","UNIX"]
 target => "unixtime"
}
mutate {convert =>["id_geo","string"]}
mutate {convert =>["gw","string"]}
mutate {convert =>["id","integer"]}
mutate {convert =>["server_id","integer"]}
mutate {convert =>["good","integer"]}
mutate {convert =>["responses","integer"]}

```

}

output {  
elasticsearch {  
hosts =\> "[http://localhost:9200](http://localhost:9200)"  
index =\> 'waittest-point'  
}  
stdout{}  
}

And i am not sure about monitoring API , can you give me pointers how to use that and also how to find CPU is getting saturated or not

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [November 15, 2017, 2:50pm UTC](https://discuss.elastic.co/t/help-needed-in-improving-the-data-ingestion-time/107767/5 "2017-11-15T14:50:58Z")

</div>

Is Logstash even the bottleneck here? Or is it Elasticsearch? Maybe they're competing for the same machine resources.

> And i am not sure about monitoring API , can you give me pointers how to use that

Did you look at the documentation?

> how to find CPU is getting saturated or not

Use `top`? Are the CPUs running at full load? If not increasing the parallelism could be an option.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 13, 2017, 2:51pm UTC](https://discuss.elastic.co/t/help-needed-in-improving-the-data-ingestion-time/107767/6 "2017-12-13T14:51:01Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
