# Logstsash is running pipelines in parallel

**URL:** <https://discuss.elastic.co/t/logstsash-is-running-pipelines-in-parallel/289614>\
**Category:** Logstash\
**Created:** [November 18, 2021, 3:23pm UTC](https://discuss.elastic.co/t/logstsash-is-running-pipelines-in-parallel/289614 "2021-11-18T15:23:16Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![rodri.gz](https://avatars.discourse-cdn.com/v4/letter/r/aca169/32.png) [@rodri.gz](https://discuss.elastic.co/u/rodri.gz)\
**Post date:** [November 18, 2021, 3:23pm UTC](https://discuss.elastic.co/t/logstsash-is-running-pipelines-in-parallel/289614/1 "2021-11-18T15:23:16Z")

</div>

Hello!

I have a divided pipeline that generates 2 indices that enrich each other.

I would like to know how I can make it do one etl first and when it finishes the other.

my logstash filter wich enrichs has this filter:

```auto
  elasticsearch {
    hosts => "https://xxxxxxxxx:9200"
    index => "prueba-agentes"
    query => "alias.keyword:%{[alias]}"
    fields => {
      "estado_agente" => "estado_agente_enriquecido"
    }
    user => "xxx"
    password => "xxxx"
    ca_file => 'xxx.pem'
  }

```

my pipelines.yml :

```auto
- pipeline.id: age-mod-pand
  path.config: "/apps/elastic/logstash/config/custom/Pruebas-mod-age/*"
  pipeline.workers: 1
  pipeline.order: true

```

the 2 filters (with input and output) :

```auto
[root@host]# ll /apps/elastic/logstash/config/custom/Pruebas-mod-age/
total 8
-rw-r--r-- 1 root root 2058 nov 18 16:07 A__agentes.conf
-rw-r--r-- 1 root root 2343 nov 18 16:07 B__modules.conf

```

If I first launch A and stop it and then launch B it works perfectly but when I put it in pipelines it shows the error indicating the field with which I want to enrich it does not exist ye  
therefore I understand that logstash is executing the filters in parallel omitting the alphabetical order.

Thank you in advanced!

---

<div class="post-metadata">

**Author:** ![AquaX](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aquax/32/92006_2.png) [@AquaX](https://discuss.elastic.co/u/AquaX)\
**Post date:** [November 18, 2021, 9:18pm UTC](https://discuss.elastic.co/t/logstsash-is-running-pipelines-in-parallel/289614/2 "2021-11-18T21:18:21Z")

</div>

If you want to truly preserve the order of operations then use pipeline to pipeline communication

> **[Pipeline-to-Pipeline Communication | Logstash Reference \[7.15\] | Elastic](https://www.elastic.co/guide/en/logstash/current/pipeline-to-pipeline.html)**

config/pipelines.yml

```auto
- pipeline.id: age-mod-pand-input
  #THIS EXECUTES FIRST
  path.config: "/apps/elastic/logstash/config/custom/Pruebas-mod-age/A__agentes.conf"
  pipeline.workers: 1
  pipeline.order: true
  

- pipeline.id: age-mod-pand-enrich
 #THIS EXECUTES WHENEVER AN EVENT IS SENT FROM THE age-mod-pand-input
  path.config: "/apps/elastic/logstash/config/custom/Pruebas-mod-age/B__modules.conf"
  pipeline.workers: 1
  pipeline.order: true

```

A\_\_agentes.conf

```auto
input {
...
}
filter {
  elasticsearch {
    hosts => "https://xxxxxxxxx:9200"
    index => "prueba-agentes"
    query => "alias.keyword:%{[alias]}"
    fields => {
      "estado_agente" => "estado_agente_enriquecido"
    }
    user => "xxx"
    password => "xxxx"
    ca_file => 'xxx.pem'
  }
}
output{
     pipeline { send_to => age-mod-pand-enrich}
}

```

B\_\_modules.conf

```auto
input {
     pipeline { address => age-mod-pand-enrich}
}
filter {
...
}
output {
...
}

```

---

<div class="post-metadata">

**Author:** ![rodri.gz](https://avatars.discourse-cdn.com/v4/letter/r/aca169/32.png) [@rodri.gz](https://discuss.elastic.co/u/rodri.gz)\
**Post date:** [November 19, 2021, 8:59am UTC](https://discuss.elastic.co/t/logstsash-is-running-pipelines-in-parallel/289614/3 "2021-11-19T08:59:15Z")

</div>

First of all Thanks for replying.

This is very interesting that you tell me because I did not know that the interconnection of pipelines existed.

but I don't see how to apply it well.

I currently have 2 jdbc inputs that run every x time.

2 independent filters and in the second an enrichment to bring me the fields of an index

and 2 outputs to 2 indices.

If I use your model, how does it enrich since the index is not previously uploaded? How could the Elasticsearch filter know that such data exists?

```auto

filter {
  elasticsearch {
    hosts => "https://xxxxxxxxx:9200"
    # ------------> index => "prueba-agentes" <-----------------
    query => "alias.keyword:%{[alias]}"
    fields => {
      "estado_agente" => "estado_agente_enriquecido"
    }
    user => "xxx"
    password => "xxxx"
    ca_file => 'xxx.pem'
  }
}

```

---

<div class="post-metadata">

**Author:** ![AquaX](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aquax/32/92006_2.png) [@AquaX](https://discuss.elastic.co/u/AquaX)\
**Post date:** [November 19, 2021, 2:17pm UTC](https://discuss.elastic.co/t/logstsash-is-running-pipelines-in-parallel/289614/4 "2021-11-19T14:17:16Z")

</div>

Could you give me an example of what you want this whole workflow to look like?

---

<div class="post-metadata">

**Author:** ![rodri.gz](https://avatars.discourse-cdn.com/v4/letter/r/aca169/32.png) [@rodri.gz](https://discuss.elastic.co/u/rodri.gz)\
**Post date:** [November 22, 2021, 8:15am UTC](https://discuss.elastic.co/t/logstsash-is-running-pipelines-in-parallel/289614/5 "2021-11-22T08:15:38Z")

</div>

yes.

I have 2 indices.

Each of them has a jdbc oracle entry that fetches the data every 5 minutes.

I am trying to enrich one with fields from the other using the Elasticsearch filter. The problem is the following:

When I put it in pipelines they both run at the same time and therefore the enrichment does not complete.  
This I have already seen that with workers 1 order etc it can be solved. The most important problem that I do not see how to solve is that logstash tries to enrich with everything and really should only do it with the last data that is originally indexed.

I only get it to work if I run the first ETL by hand and then the second.

Is there a way to do this enrichment with dynamic indexes taking only the last indexed documents?

This would be of great help to me because it is being repeated in different use cases, thank you.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 20, 2021, 8:16am UTC](https://discuss.elastic.co/t/logstsash-is-running-pipelines-in-parallel/289614/6 "2021-12-20T08:16:26Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
