# XML filter

**URL:** <https://discuss.elastic.co/t/xml-filter/48133>\
**Category:** Logstash\
**Created:** [April 22, 2016, 7:52am UTC](https://discuss.elastic.co/t/xml-filter/48133 "2016-04-22T07:52:52Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![meons](https://avatars.discourse-cdn.com/v4/letter/m/b5a626/32.png) [@meons](https://discuss.elastic.co/u/meons)\
**Post date:** [April 22, 2016, 7:52am UTC](https://discuss.elastic.co/t/xml-filter/48133/1 "2016-04-22T07:52:52Z")

</div>

Hello,

I have to parse some xml files (with xml filter) and store each child item (car) as a document in Elasticsearch but I don't see how I can do that.

My xml file could containt n child items (around 100-200 per xml file) and I need 1 record for each child. There is an example :

`<cars> <car> <model>BMW</model> <color>Red</color> </car> <car> <model>Audi</model> <color>Yellow</color> </car> <car> <model>Mercedes</model> <color>Black</color> </car> </cars>`

So I want to have 1 record for each car contained in my xml document. So the xpath param will be something like that  
`xpath => ["/cars/car/model", "Model", "cars/car/color", "Color"]`  
But I don't get the logic of how can I tell Logstash to store each item in a new record =/

Could you explain me please?

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [April 24, 2016, 4:07pm UTC](https://discuss.elastic.co/t/xml-filter/48133/2 "2016-04-24T16:07:03Z")

</div>

You should be able to use the [split filter](https://www.elastic.co/guide/en/logstash/current/plugins-filters-split.html) as long as you can convert the XML into something like this:

```auto
{
  "cars": {
    [
      {
        "model": "BMW",
        "color": "Red"
      },
      {
        "model": "Audi",
        "color": "Yellow"
      },
      ...
    ]
  }
}

```

I wonder, is the `xpath` option really the best way forward? Can't you parse the whole XML document, delete any unwanted fields, and use the split filter on the result?

---

<div class="post-metadata">

**Author:** ![meons](https://avatars.discourse-cdn.com/v4/letter/m/b5a626/32.png) [@meons](https://discuss.elastic.co/u/meons)\
**Post date:** [April 26, 2016, 8:56am UTC](https://discuss.elastic.co/t/xml-filter/48133/3 "2016-04-26T08:56:25Z")

</div>

Yop @magnusbaeck thanks for your response.

Finally I just used the following filter

`xml { source => "cars" target => "cars" }`

So in elastic I have one document for each car with following message field and it's exatcly what I want.  
`<car> <model>xxx</model> <color>yyy</color> </car>`  
Once parsed like this, I can easly add fields model or color to have a more clean document (instead of a message field with all data inside).

👌

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 5:00am UTC](https://discuss.elastic.co/t/xml-filter/48133/4 "2017-07-06T05:00:33Z")

</div>


