# Dynamic Data Type

**URL:** <https://discuss.elastic.co/t/dynamic-data-type/206639>\
**Category:** Logstash\
**Created:** [November 5, 2019, 4:05pm UTC](https://discuss.elastic.co/t/dynamic-data-type/206639 "2019-11-05T16:05:42Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![zmink-pxc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zmink-pxc/32/57138_2.png) [@zmink-pxc](https://discuss.elastic.co/u/zmink-pxc)\
**Post date:** [November 5, 2019, 4:05pm UTC](https://discuss.elastic.co/t/dynamic-data-type/206639/1 "2019-11-05T16:05:42Z")

</div>

Hi all,

` I have a data format as shown in the attached image -`

![image](https://us1.discourse-cdn.com/elastic/original/3X/6/0/6011ef789a9fa744a4219e01d42cc5a8a99a113f.png)

I'm able to import the data via the CSV input plugin which works great. However, I'm stuck on the next step which is to map the rows to numbers so that the data can be analyzed in kibana. The issue is that the values should generally be numbers as in the second two columns, however, if there is an error with the data point at some point in time, an error code will be generated as shown in the last two columns. Is there some way to map the columns to number datatype while also handling the occasional case where the value will be a string?

For reference, below is my current logstash config which needs to be expanded upon

input {  
file {  
path =\> "C:/Users/zach/Downloads/pdr\*.csv"  
start\_position =\> "beginning"  
sincedb\_path =\> "NUL"  
}  
}

filter {  
csv {  
separator =\> ","  
autodetect\_column\_names =\> true  
autogenerate\_column\_names =\> true  
}  
}

output {  
stdout { codec =\> rubydebug }

elasticsearch {  
hosts =\> ["localhost:9200"]  
index =\> "pdr-data"  
}  
}

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [November 5, 2019, 4:50pm UTC](https://discuss.elastic.co/t/dynamic-data-type/206639/2 "2019-11-05T16:50:14Z")

</div>

> [@zmink-pxc](#):
>
> Is there some way to map the columns to number datatype while also handling the occasional case where the value will be a string?

In elasticsearch, if you have a template, then if a field is expected to be an integer I think (I have not tested) that you would get a mapping exception if you try to ingest a document where it is a string that cannot be parsed as a number.

If you do not have a template then you run the risk that the first document indexed contains a string in that field and the field type gets set to text.

You could record the fact that the field contained an error in another field, and then remove it. Something like

```
    ruby {
        code => '
            errors = []
            event.to_hash.each { |k, v|
                if k =~ /column[0-9]+/
                    unless v.to_f.to_s == v.to_s
                        event.remove(k)
                        errors << k
                    end
                end
            if errors != []
                event.set("errorFields", errors)
            end
            }
        '
    }

```

---

<div class="post-metadata">

**Author:** ![zmink-pxc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zmink-pxc/32/57138_2.png) [@zmink-pxc](https://discuss.elastic.co/u/zmink-pxc)\
**Post date:** [November 12, 2019, 4:13pm UTC](https://discuss.elastic.co/t/dynamic-data-type/206639/3 "2019-11-12T16:13:45Z")

</div>

Thanks @Badger! I used a slightly different method but the solution was spot on. For reference,

```
 ruby {
    code => "
        event.to_hash.each { |k, v|
            if !['@version','@timestamp','message','path','Timestamp','host'].include?(k)
                if v.include? 'ERR'
                    event.set(k+'-ERR',v)
                    event.remove(k)
                else
                    event.set(k,v.to_f)
                end
            end
        }
        path = event.get('[path]')
        unitExists = path.include? 'pdr'
        if unitExists
            filename = event.get('[path]').split('/').last
            pdrID = filename[0...-19]
            event.set('unit',pdrID)
        end
    "
}
```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 10, 2019, 4:19pm UTC](https://discuss.elastic.co/t/dynamic-data-type/206639/4 "2019-12-10T16:19:56Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
