# Parse LDIF format, aggregate several lines into one ES index entry

**URL:** https://discuss.elastic.co/t/parse-ldif-format-aggregate-several-lines-into-one-es-index-entry/139760
**Category:** Logstash
**Created:** [July 12, 2018, 1:14pm UTC](https://discuss.elastic.co/t/parse-ldif-format-aggregate-several-lines-into-one-es-index-entry/139760 "2018-07-12T13:14:09Z")
**Posts on this page:** 4
**Page:** 1

<div class="post-metadata">

### Author: ![izinovik](https://avatars.discourse-cdn.com/v4/letter/i/e95f7d/32.png) [@izinovik](https://discuss.elastic.co/u/izinovik)
#### Post date: [July 12, 2018, 1:14pm UTC](https://discuss.elastic.co/t/parse-ldif-format-aggregate-several-lines-into-one-es-index-entry/139760/1 "2018-07-12T13:14:09Z")

</div>

Hello.

I'm trying to solve following task: parse LDIF (LDAP data interchange) with Logstash using 'aggregate' plugin, but currently with no success.

Here is an example of LDIF:  
time: 20180708032751  
dn: fqdn=host.acme.local,cn=computers,cn=accounts,dc=acme,dc=local  
result: 0  
changetype: modify  
replace: krbLastSuccessfulAuth  
krbLastSuccessfulAuth: 20180708002752Z  
-  
replace: modifiersname  
modifiersname: cn=Directory Manager  
-  
replace: modifytimestamp  
modifytimestamp: 20180708002751Z  
-  
replace: entryusn  
entryusn: 252729871

What I want to get is something like this:  
{  
"host" =\> "host.acme.local",  
"@timestamp" =\> 2018-07-12T12:06:14.980Z,  
"dn" =\> "fqdn=host.acme.local,cn=computers,cn=accounts,dc=acme,dc=local"  
"result" =\> "0"  
"changetype" =\> "modify"  
"replace" =\> [  
"krbLastSuccessfulAuth" =\> "20180708002752Z"  
"modifiersname" =\> "cn=Directory Manager"  
"modifytimestamp" =\> "20180708002752Z"  
]  
"message" =\> "entryusn: 252730169",  
"@version" =\> "1",  
"entryusn" =\> "252730169"  
}

In my case LDIF entry begins with 'time' attribute and ends with 'entryusn'. As is can be seen there is  
no any kind of ID on each line that will help to aggregate them into one elasticsearch entry.

First I thought that 'time' can be used as 'task\_id' for 'aggregate' plugin, but several LDIF entries can  
have same time, so only unique id is entryusn (update sequence number), but it is written as last line so I cannot access it with 'aggregate { task\_id =\> "%{entryusn} map\_action =\> "create" }'

---

<div class="post-metadata">

### Author: ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)
#### Post date: [July 12, 2018, 3:18pm UTC](https://discuss.elastic.co/t/parse-ldif-format-aggregate-several-lines-into-one-es-index-entry/139760/2 "2018-07-12T15:18:27Z")

</div>

If there is no unique id then aggregate is not going to work. If all these lines are written as a unit then perhaps you can use a multiline codec on the input?

---

<div class="post-metadata">

### Author: ![izinovik](https://avatars.discourse-cdn.com/v4/letter/i/e95f7d/32.png) [@izinovik](https://discuss.elastic.co/u/izinovik)
#### Post date: [July 13, 2018, 7:05am UTC](https://discuss.elastic.co/t/parse-ldif-format-aggregate-several-lines-into-one-es-index-entry/139760/3 "2018-07-13T07:05:04Z")

</div>

I think it is not possible in my case since my log flow comes into central rsyslog collector which stores logs in redis from which logstash reads log entries. Can I somehow selectively apply multiline codec to entries that Logstash reads from single input?

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [August 10, 2018, 7:05am UTC](https://discuss.elastic.co/t/parse-ldif-format-aggregate-several-lines-into-one-es-index-entry/139760/4 "2018-08-10T07:05:04Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
