# Taking too much time to index

**URL:** <https://discuss.elastic.co/t/taking-too-much-time-to-index/87918>\
**Category:** Logstash\
**Created:** [June 1, 2017, 2:06pm UTC](https://discuss.elastic.co/t/taking-too-much-time-to-index/87918 "2017-06-01T14:06:07Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![munotshubham](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/munotshubham/32/22663_2.png) [@munotshubham](https://discuss.elastic.co/u/munotshubham)\
**Post date:** [June 1, 2017, 2:06pm UTC](https://discuss.elastic.co/t/taking-too-much-time-to-index/87918/1 "2017-06-01T14:06:07Z")

</div>

So I'm trying to index a DB with 91 columns and around 300,000 rows. It is in a CSV file and I'm using logstash to load it in ES.

It's been running since 20 hours, it still hasn't been indexed.  
For test purposes I had taken just 10 rows and indexed it. It was working fine.

I have ran logstash on debug mode

```
10:00:34.221 [Ruby-0-Thread-11: /usr/share/logstash/logstash-core/lib/logstash/pipeline.rb:532] DEBUG logstash.pipeline - Pushing flush onto pipeline
10:00:34.965 [[main]<file] DEBUG logstash.inputs.file - each: file grew: /home/patagonia/Documents/testserver-patients-May31.csv: old size 0, new size 356553232
10:00:35.966 [[main]<file] DEBUG logstash.inputs.file - each: file grew: /home/patagonia/Documents/testserver-patients-May31.csv: old size 0, new size 356553232
10:00:36.967 [[main]<file] DEBUG logstash.inputs.file - each: file grew: /home/patagonia/Documents/testserver-patients-May31.csv: old size 0, new size 356553232
10:00:37.968 [[main]<file] DEBUG logstash.inputs.file - each: file grew: /home/patagonia/Documents/testserver-patients-May31.csv: old size 0, new size 356553232
10:00:38.969 [[main]<file] DEBUG logstash.inputs.file - each: file grew: /home/patagonia/Documents/testserver-patients-May31.csv: old size 0, new size 356553232
10:00:39.220 [Ruby-0-Thread-11: /usr/share/logstash/logstash-core/lib/logstash/pipeline.rb:532] DEBUG logstash.pipeline - Pushing flush onto pipeline
10:00:39.970 [[main]<file] DEBUG logstash.inputs.file - each: file grew: /home/patagonia/Documents/testserver-patients-May31.csv: old size 0, new size 356553232
10:00:40.972 [[main]<file] DEBUG logstash.inputs.file - each: file grew: /home/patagonia/Documents/testserver-patients-May31.csv: old size 0, new size 356553232
10:00:41.973 [[main]<file] DEBUG logstash.inputs.file - each: file grew: /home/patagonia/Documents/testserver-patients-May31.csv: old size 0, new size 356553232
10:00:42.975 [[main]<file] DEBUG logstash.inputs.file - each: file grew: /home/patagonia/Documents/testserver-patients-May31.csv: old size 0, new size 356553232
10:00:43.976 [[main]<file] DEBUG logstash.inputs.file - each: file grew: /home/patagonia/Documents/testserver-patients-May31.csv: old size 0, new size 356553232
10:00:43.977 [[main]<file] DEBUG logstash.inputs.file - _globbed_files: /home/patagonia/Documents/testserver-patients-May31.csv: glob is: ["/home/patagonia/Documents/testserver-patients-May31.csv"]

```

This is repeating since 20 hours. I've no clue how much has been indexed till now.  
There is no new folder in /var/lib/elasticsearch/nodes/0/indices apart from the indices that are already there on localhost:9200/\_cat/indices

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 1, 2017, 2:25pm UTC](https://discuss.elastic.co/t/taking-too-much-time-to-index/87918/2 "2017-06-01T14:25:14Z")

</div>

What does you Logstash config look like?

---

<div class="post-metadata">

**Author:** ![munotshubham](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/munotshubham/32/22663_2.png) [@munotshubham](https://discuss.elastic.co/u/munotshubham)\
**Post date:** [June 1, 2017, 2:40pm UTC](https://discuss.elastic.co/t/taking-too-much-time-to-index/87918/3 "2017-06-01T14:40:04Z")

</div>

```
input {
  file {
    path => "/home/patagonia/Documents/testserver-patients*.csv"
  }
}
filter {
  csv {
    columns => ["PatientID", "PaperChartNumber", "PMRecorndNumber", "SubscriberID", "PracticeID", "UserID", "PatientType", "StartDate", "EndDate", "IsActive", "IsDeleted", "IsReportable", "Title", "FirstName", "MiddleName", "LastName", "Suffix", "PreferredName", "GuardianName", "MaidenName", "DateOfBirth", "DateOfDeath", "Gender", "MaritalStatus", "RaceCode", "LanguageCode", "InsuranceType", "PharmacyName", "PreferredContact", "InactiveReason", "BloodType", "AddressLine1", "AddressLine2", "AddressLine3", "City", "State", "Zipcode", "HomePhoneNumber", "WorkPhoneNumber", "MobileNumber", "EmailAddress", "PatientPhotoLocation", "MergedPatientID", "InsertedBy", "InsertDate", "LastEditedBy", "LastEditDate", "EthnicGroup", "ReferringPhysician", "ReferringPhysicianPhone", "ReferringPhysicianFax", "PharmacyPhone", "PharmacyFax", "ReferringPhysicianCity", "PharmacyCity", "PCPName", "PCPPhone", "PCPFax", "PCPCity", "AltPhone1", "AltPhone2", "InsuranceID", "PMPatientID", "PMCaseName", "PatInsProfileID", "PatientInsuranceID", "PatientInsuranceName", "Employer", "Employment", "LocationID", "Comments", "PatientState", "PatientStateDate", "County", "RefPhysicianID", "CNDSID", "SSN", "Nosnailmail", "NeedsInterpreter", "CountryCode", "Veteranstatus", "WCServedIn", "Driverlicense", "Cl"]
  }
}
output {
  elasticsearch {
    hosts => ["localhost:9200"]
    index => "patient"
  }
}
```

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 1, 2017, 2:46pm UTC](https://discuss.elastic.co/t/taking-too-much-time-to-index/87918/4 "2017-06-01T14:46:01Z")

</div>

What does resource utilisation, particularly CPU and disk I/O, look like on the host while indexing? How many CPU cores do you have available?

---

<div class="post-metadata">

**Author:** ![munotshubham](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/munotshubham/32/22663_2.png) [@munotshubham](https://discuss.elastic.co/u/munotshubham)\
**Post date:** [June 1, 2017, 3:16pm UTC](https://discuss.elastic.co/t/taking-too-much-time-to-index/87918/5 "2017-06-01T15:16:20Z")

</div>

How do I check all that?

I cannot see all these information using top command on terminal.

---

<div class="post-metadata">

**Author:** ![munotshubham](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/munotshubham/32/22663_2.png) [@munotshubham](https://discuss.elastic.co/u/munotshubham)\
**Post date:** [June 2, 2017, 1:33am UTC](https://discuss.elastic.co/t/taking-too-much-time-to-index/87918/6 "2017-06-02T01:33:27Z")

</div>

Hey,  
I figured out the solution.  
Basically logstash works similar to beats agent, Harvestor had already indexed the data once, so it wasn't indexing the same file. So I had to create a new file with the same content to get it indexed.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 30, 2017, 1:33am UTC](https://discuss.elastic.co/t/taking-too-much-time-to-index/87918/7 "2017-06-30T01:33:34Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
