# Performance issues while importing CSV files into Elasticsearch

**URL:** <https://discuss.elastic.co/t/performance-issues-while-importing-csv-files-into-elasticsearch/143723>\
**Category:** Logstash\
**Created:** [August 9, 2018, 2:52pm UTC](https://discuss.elastic.co/t/performance-issues-while-importing-csv-files-into-elasticsearch/143723 "2018-08-09T14:52:50Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![jnci](https://avatars.discourse-cdn.com/v4/letter/j/c77e96/32.png) [@jnci](https://discuss.elastic.co/u/jnci)\
**Post date:** [August 9, 2018, 2:52pm UTC](https://discuss.elastic.co/t/performance-issues-while-importing-csv-files-into-elasticsearch/143723/1 "2018-08-09T14:52:50Z")

</div>

I'm trying to import some gigabytes of CSV files (approximately 100+ million rows with 5 columns) but the throughput is very low (~1mb/s). I'm not quite sure yet what the issue may be, but maybe someone here has some leads.

What throughput could I expect on an 8GB, i7 (octocore) box with an SSD for the stack and an external HDD from which the data is imported (USB 3.0, known to be readable at 200mb/s+), using the default Security Onion stack (Evaluation mode)? Are there any known throughput issues with importing CSV files using the CSV filter? Likewise for using the Date filter (to extract timestamp values from the CSV)?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [August 9, 2018, 5:28pm UTC](https://discuss.elastic.co/t/performance-issues-while-importing-csv-files-into-elasticsearch/143723/2 "2018-08-09T17:28:21Z")

</div>

When logstash starts it will log something like

```
[logstash.agent] Successfully started Logstash API endpoint {:port=>9600}

```

You can query that using the [APIs](https://www.elastic.co/guide/en/logstash/current/monitoring.html). For example this will tell you the time spent in each part of the pipeline.

```
curl -XGET 'localhost:9600/_node/stats/pipelines?pretty'

```

On one of my servers I see a csv filter processing about 7,000 rows per second with a single worker thread. That would scale with the number of CPUs. A simple date filter is cheaper than a 5 column csv.

csv appears to be quite expensive. A stripped down regex that does not handle quoted fields gets about 3 times the throughput.

```
    ruby { code => '
        m = event.get("message").scan(/([^,]+)(,|$)/)
        m.each_index { |i|
            event.set("column#{i}", m[i][0])
        }
    ' }

```

See [this](https://www.elastic.co/blog/logstash-configuration-tuning) blog post also (but note that you have to use nested notation not dot notation now, so [documents][rate\_1m]).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 6, 2018, 5:28pm UTC](https://discuss.elastic.co/t/performance-issues-while-importing-csv-files-into-elasticsearch/143723/3 "2018-09-06T17:28:22Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
