# Logstash json to multiple documents

**URL:** <https://discuss.elastic.co/t/logstash-json-to-multiple-documents/207007>\
**Category:** Logstash\
**Created:** [November 7, 2019, 8:09pm UTC](https://discuss.elastic.co/t/logstash-json-to-multiple-documents/207007 "2019-11-07T20:09:52Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![vikasp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vikasp/32/90872_2.png) [@vikasp](https://discuss.elastic.co/u/vikasp)\
**Post date:** [November 7, 2019, 8:09pm UTC](https://discuss.elastic.co/t/logstash-json-to-multiple-documents/207007/1 "2019-11-07T20:09:53Z")

</div>

I have an input in the below format:  
`<data><te0><id>1</id><text>this is first event</text></te0><te1><id>2</id><text>this is second event</text></te1><te2><id>3</id><text>this is third event></text></te2><te0><id>4</id><text>this is fourth event</text></te0></data>`

The above is a single input event to logstash. I want to convert the above single event to multiple events and push it to elasticsearch only for the te0 attribute value as a document identified by it's id. So, for the above example, result should be:  
say index: xml\_test  
xml\_test/doc/1/: { "\_id": 1, "text": "this is first event" }  
xml\_test/doc/4/: { "\_id": 4, "text": this is fourth event" }

Below is the logstash config I am trying to use:  
grok {  
match =\> ["message", "%{GREEDYDATA:inxml}"]  
}  
xml {  
source =\> "inxml"  
target =\> "xmldata"  
force\_array =\> false  
}  
json {  
source =\> "xmldata"  
}  
split {  
field =\> "xmldata"  
}  
}

I am getting \_json\_parse\_failure, I see that it's already a json in elasticsearch and also \_split\_failure since it can only happen on string or array but says xmldata is a hash.  
something like this:  
"xmldata": {  
"te0": {  
"text": "this is a second doc",  
"id": "1"  
},  
"te2": {  
"text": " this is supposed to be 3rdor4th doc",  
"id": "2"  
}  
},  
and  
"tags": [  
"\_jsonparsefailure",  
"\_split\_type\_failure"  
],

How can I convert the above input to multiple docs for the values in attributes only inside

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [November 7, 2019, 8:42pm UTC](https://discuss.elastic.co/t/logstash-json-to-multiple-documents/207007/2 "2019-11-07T20:42:47Z")

</div>

What does your data look like?

---

<div class="post-metadata">

**Author:** ![vikasp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vikasp/32/90872_2.png) [@vikasp](https://discuss.elastic.co/u/vikasp)\
**Post date:** [November 7, 2019, 8:52pm UTC](https://discuss.elastic.co/t/logstash-json-to-multiple-documents/207007/3 "2019-11-07T20:52:53Z")

</div>

Hi,  
I did put a sample event, not sure what happened to it: here it is:  
`<data><te0><id>1</id><text>this is first event</text></te0><te1><id>2</id><text>this is second event</text></te1><te0><id>3</id><text>this is third event</text></te0><te2><id>4</id><text>this is fourth event</text></te2></data>`

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [November 7, 2019, 11:35pm UTC](https://discuss.elastic.co/t/logstash-json-to-multiple-documents/207007/4 "2019-11-07T23:35:40Z")

</div>

Parsing your sample xml results in

```
   "xmldata" => {
    "te0" => [
        [0] {
            "text" => "this is first event",
              "id" => "1"
        },
        [1] {
            "text" => "this is third event",
              "id" => "3"
        }
    ],
    "te1" => {
        "text" => "this is second event",
          "id" => "2"
    },
    "te2" => {
        "text" => "this is fourth event",
          "id" => "4"
    }
},

```

You say you want ids 1 and 4. What test do you use to drop ids 2 and 3?

---

<div class="post-metadata">

**Author:** ![vikasp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vikasp/32/90872_2.png) [@vikasp](https://discuss.elastic.co/u/vikasp)\
**Post date:** [November 8, 2019, 12:34am UTC](https://discuss.elastic.co/t/logstash-json-to-multiple-documents/207007/5 "2019-11-08T00:34:17Z")

</div>

All I want is to push to elasticsearch as seprate documents for id 1 and id 3 and remove anything else other than te0.  
so kind out output should index two documents to elasticsearch. and below are the two docs:  
doc 1 with id: 1 and doc as { "id":1, "text": "this is first event", "@timestamp":".....".......all metadata...}  
doc 2 with id: 3 and doc as { "id":3, "text": "this is third event", "@timestamp":".....".......all metadata...}

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [November 8, 2019, 1:10am UTC](https://discuss.elastic.co/t/logstash-json-to-multiple-documents/207007/6 "2019-11-08T01:10:53Z")

</div>

So if you iterate over the members of xmldata you want to ignore any that are not arrays? And if they are arrays then split them?

---

<div class="post-metadata">

**Author:** ![vikasp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vikasp/32/90872_2.png) [@vikasp](https://discuss.elastic.co/u/vikasp)\
**Post date:** [November 9, 2019, 6:31am UTC](https://discuss.elastic.co/t/logstash-json-to-multiple-documents/207007/7 "2019-11-09T06:31:33Z")

</div>

I want only the members under xmldata with attribute te0, and ignore everything (can be multiple te1, te2 ,te3 and so on). and most of them are arrays (te\*)

---

<div class="post-metadata">

**Author:** ![vikasp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vikasp/32/90872_2.png) [@vikasp](https://discuss.elastic.co/u/vikasp)\
**Post date:** [November 9, 2019, 6:34am UTC](https://discuss.elastic.co/t/logstash-json-to-multiple-documents/207007/8 "2019-11-09T06:34:15Z")

</div>

xml data will merge all te0s into one array and will have a single te0 key and first occurence in the xml of te0 will be 1st element in array and so on....Now, my split is creating documents with te0 [0] and the rest of the vars (te1, te2....), te0[1] and the rest of the vars(te1, te2....) and so no....  
But I want is only te0[0] as one document without any other te\*s, te[1] as another document.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [November 9, 2019, 2:18pm UTC](https://discuss.elastic.co/t/logstash-json-to-multiple-documents/207007/9 "2019-11-09T14:18:49Z")

</div>

Try

```
    ruby {
        code => '
            event.get("xmldata").each { |k, v|
                unless k == "te0"
                    event.remove("[xmldata][#{k}]")
                end
            }
        '
    }
    split { field => "[xmldata][te0]" }
```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 7, 2019, 2:31pm UTC](https://discuss.elastic.co/t/logstash-json-to-multiple-documents/207007/10 "2019-12-07T14:31:54Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
