# How to split CSV file and filter CSV Files

**URL:** <https://discuss.elastic.co/t/how-to-split-csv-file-and-filter-csv-files/28710>\
**Category:** Logstash\
**Created:** [September 4, 2015, 11:03pm UTC](https://discuss.elastic.co/t/how-to-split-csv-file-and-filter-csv-files/28710 "2015-09-04T23:03:18Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Hans](https://avatars.discourse-cdn.com/v4/letter/h/e19b73/32.png) [@Hans](https://discuss.elastic.co/u/Hans)\
**Post date:** [September 4, 2015, 11:03pm UTC](https://discuss.elastic.co/t/how-to-split-csv-file-and-filter-csv-files/28710/1 "2015-09-04T23:03:18Z")

</div>

Hi,  
I have a challenge where I am not able to filter a CSV file in logstash and pass this a separate files to elasticsearch.

**Original files:**  
CLIENT\_IP,ISP,TEST\_DATE,SERVER\_NAME,DOWNLOAD\_KBPS,UPLOAD\_KBPS,LATENCY,LATITUDE,LONGITUDE,CONNECTION\_TYPE  
41.151._._,Telkom Internet,8/14/2012 01:40:43 GMT,Windhoek,3894,2401,194,-26.0947,28.2161,Cell  
41.243._._,Airtel DRC,8/14/2012 04:51:21 GMT,Windhoek,871,1170,146,-26.0749,28.2503,WiFi  
41.243._._,Airtel DRC,8/14/2012 04:52:09 GMT,Windhoek,878,1086,156,-26.0749,28.2503,WiFi

**Configuration file:**  
input {  
file {  
type =\> "speedtest"  
path =\> ["/data/speedtest3.txt"]  
start\_position =\> "beginning"  
}  
}  
filter {  
grok {  
patterns\_dir =\> "/opt/logstash/vendor/bundle/jruby/1.9/gems/logstash-patterns-core-0.3.0/patterns"  
match =\> ["messages", "%{URIHOST}_._,%{WORD:ISP},%{QS},%{WORD:SERVER\_NAME},%{NUMBER:DOWNLOAD\_KBPS},%{NUMBER:UPLOAD\_KBPS},%{NUMBER:LATENCY},%{JAVACLASS},%{JAVACLASS},%{WORD:CONNECTION\_TYPE}"]  
tag\_on\_failure =\> []  
add\_tag =\> "Android-Speedtest"  
}  
mutate {  
split =\> ["messages", ","]  
}  
kv{  
field\_split =\> ","  
value\_split =\> ","  
source =\> "kvdata"  
remove\_field =\> "kvdata"  
}  
}  
output {  
elasticsearch {  
protocol =\> "node"  
host =\> "localhost"  
cluster =\> "elasticsearch"  
}  
}

**Output information:**  
{  
"\_index": ".kibana",  
"\_type": "index-pattern",  
"\_id": "logstash-_",  
"\_score": 1,  
"\_source": {  
"title": "logstash-_",  
"timeFieldName": "@timestamp",  
"customFormats": "{}",  
"fields": "[{"type":"string","indexed":false,"analyzed":false,"name":"\_index","count":0,"scripted":false},{"type":"string","indexed":true,"analyzed":false,"name":"\_type","count":0,"scripted":false},{"type":"geo\_point","indexed":true,"analyzed":false,"doc\_values":false,"name":"geoip.location","count":0,"scripted":false},{"type":"string","indexed":true,"analyzed":false,"doc\_values":false,"name":"@version","count":0,"scripted":false},{"type":"string","indexed":false,"analyzed":false,"name":"\_source","count":0,"scripted":false},{"type":"string","indexed":false,"analyzed":false,"name":"\_id","count":1,"scripted":false},{"type":"string","indexed":true,"analyzed":false,"doc\_values":false,"name":"host.raw","count":0,"scripted":false},{"type":"string","indexed":true,"analyzed":false,"doc\_values":false,"name":"type.raw","count":0,"scripted":false},{"type":"string","indexed":true,"analyzed":true,"doc\_values":false,"name":"message","count":0,"scripted":false},{"type":"string","indexed":true,"analyzed":true,"doc\_values":false,"name":"type","count":0,"scripted":false},{"type":"string","indexed":true,"analyzed":true,"doc\_values":false,"name":"path","count":0,"scripted":false},{"type":"date","indexed":true,"analyzed":false,"doc\_values":false,"name":"@timestamp","count":2,"scripted":false},{"type":"string","indexed":true,"analyzed":true,"doc\_values":false,"name":"host","count":0,"scripted":false},{"type":"string","indexed":true,"analyzed":false,"doc\_values":false,"name":"path.raw","count":0,"scripted":false}]"  
},  
"fields": {}  
}  
The output does not separate the different fields. When trying to use CSV the @timestamp could not be collected. I would appreciate any kind of assistance to get this challenge resolved as there is Latitude and longitude information also that I would like to use in Kibana 4.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 5, 2015, 10:35am UTC](https://discuss.elastic.co/t/how-to-split-csv-file-and-filter-csv-files/28710/2 "2015-09-05T10:35:55Z")

</div>

For CSV style input it might be worthwhile looking into using the [CSV filter](https://www.elastic.co/guide/en/logstash/current/plugins-filters-csv.html) instead of grok. This filter extracts all fields as strings, so you may need to also use the mutate filter to convert field types where necessary.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 5:29am UTC](https://discuss.elastic.co/t/how-to-split-csv-file-and-filter-csv-files/28710/3 "2017-07-06T05:29:59Z")

</div>


