# Indexing through logstash - varying columns in the input

**URL:** <https://discuss.elastic.co/t/indexing-through-logstash-varying-columns-in-the-input/66154>\
**Category:** Elasticsearch\
**Created:** [November 15, 2016, 7:00pm UTC](https://discuss.elastic.co/t/indexing-through-logstash-varying-columns-in-the-input/66154 "2016-11-15T19:00:37Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![deepu.sundar](https://avatars.discourse-cdn.com/v4/letter/d/e0b2c6/32.png) [@deepu.sundar](https://discuss.elastic.co/u/deepu.sundar)\
**Post date:** [November 15, 2016, 7:00pm UTC](https://discuss.elastic.co/t/indexing-through-logstash-varying-columns-in-the-input/66154/1 "2016-11-15T19:00:37Z")

</div>

Hi Team,

We built a solution leveraging the ELK to parse and generate reports out of AWS billing files. Currently we are indexing all the columns available in the log file as it is. But we found that the number of columns can be changed in the future like more columns can be added or dropped based on how user managing the resource tags names in aws.

How we handle this in ES or even at Logstash level so that the indexing process does not break even if the input file columns changes ?

Appreciate your help if anybody came across any similar cases and resolved it.

---

<div class="post-metadata">

**Author:** ![suanmeiguo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/suanmeiguo/32/11758_2.png) [@suanmeiguo](https://discuss.elastic.co/u/suanmeiguo)\
**Post date:** [November 15, 2016, 7:48pm UTC](https://discuss.elastic.co/t/indexing-through-logstash-varying-columns-in-the-input/66154/2 "2016-11-15T19:48:31Z")

</div>

I think ES is supporting that. Did you have any specific issue?

---

<div class="post-metadata">

**Author:** ![deepu.sundar](https://avatars.discourse-cdn.com/v4/letter/d/e0b2c6/32.png) [@deepu.sundar](https://discuss.elastic.co/u/deepu.sundar)\
**Post date:** [November 15, 2016, 8:02pm UTC](https://discuss.elastic.co/t/indexing-through-logstash-varying-columns-in-the-input/66154/3 "2016-11-15T20:02:48Z")

</div>

Hi, Thank you for your reply. I have imported a template in ES first with the exact list of columns available in the input file before indexing the files. And in the logstash configuration I have given same exact list of columns in the filter.

filter {  
csv {  
columns =\> ["InvoiceID","PayerAccountId","....  
separator =\> ","  
}

I see that when the columns list if different than in the configurations, some records are not getting indexed properly. Am I missing anything ?

---

<div class="post-metadata">

**Author:** ![deepu.sundar](https://avatars.discourse-cdn.com/v4/letter/d/e0b2c6/32.png) [@deepu.sundar](https://discuss.elastic.co/u/deepu.sundar)\
**Post date:** [November 15, 2016, 11:17pm UTC](https://discuss.elastic.co/t/indexing-through-logstash-varying-columns-in-the-input/66154/4 "2016-11-15T23:17:03Z")

</div>

Basically what I am trying to figure out is way in ES or through Logstash wherein we can map the actual column names with the values dynamically. (assuming the first row of the file will be the column header)

In this way we don't mess up the order by dropping or adding columns in the input files and indexing should process file correctly

---

<div class="post-metadata">

**Author:** ![suanmeiguo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/suanmeiguo/32/11758_2.png) [@suanmeiguo](https://discuss.elastic.co/u/suanmeiguo)\
**Post date:** [November 17, 2016, 6:59pm UTC](https://discuss.elastic.co/t/indexing-through-logstash-varying-columns-in-the-input/66154/5 "2016-11-17T18:59:39Z")

</div>

Hi Deepu,

I have very limited experience on logstash. But from my elasticsearch experience, this is doable.

For example you can have a python script to read the file parse each line into dictionary and just send that to elasticsearch. Elasticsearch will handle columns for you.

So if there's any missing column in your data, elasticsearch will not put any value for it, but other columns in the row can still have values (so it's just like how other no-sql handles it). If there's extra column in the data elasticsearch can create dynamic mapping for it, meaning it'll try to identify the field type and create the column for you.

I know logstash support csv but not sure if it has any feature like this. Another question is why your csv column changes? AWS billing csv should have fixed format.

---

<div class="post-metadata">

**Author:** ![deepu.sundar](https://avatars.discourse-cdn.com/v4/letter/d/e0b2c6/32.png) [@deepu.sundar](https://discuss.elastic.co/u/deepu.sundar)\
**Post date:** [November 18, 2016, 7:21pm UTC](https://discuss.elastic.co/t/indexing-through-logstash-varying-columns-in-the-input/66154/6 "2016-11-18T19:21:07Z")

</div>

Hi Vincent,

Thank you so much for the reply. The columns in the billing file can be customized. For example, we can add user tags also in the DBR file. But we will know when the order of columns or count altered.

---

<div class="post-metadata">

**Author:** ![suanmeiguo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/suanmeiguo/32/11758_2.png) [@suanmeiguo](https://discuss.elastic.co/u/suanmeiguo)\
**Post date:** [November 18, 2016, 8:59pm UTC](https://discuss.elastic.co/t/indexing-through-logstash-varying-columns-in-the-input/66154/7 "2016-11-18T20:59:46Z")

</div>

Got it. Yeah then I think it's not a problem for elasticsearch or logstash. Good luck!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 16, 2016, 8:59pm UTC](https://discuss.elastic.co/t/indexing-through-logstash-varying-columns-in-the-input/66154/8 "2016-12-16T20:59:51Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
