# Logstash or Spark -- Elasticsearch

**URL:** <https://discuss.elastic.co/t/logstash-or-spark-elasticsearch/269313>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [April 6, 2021, 9:33am UTC](https://discuss.elastic.co/t/logstash-or-spark-elasticsearch/269313 "2021-04-06T09:33:42Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![alvgoro](https://avatars.discourse-cdn.com/v4/letter/a/aeb1de/32.png) [@alvgoro](https://discuss.elastic.co/u/alvgoro)\
**Post date:** [April 6, 2021, 9:33am UTC](https://discuss.elastic.co/t/logstash-or-spark-elasticsearch/269313/1 "2021-04-06T09:33:43Z")

</div>

Hi there!

I was searching on Google and this forum but I am still undecided.

I would like to develop a system which collects logs from several servers and then analyze them.

I thinked about Beats to collect logs (and maybe Logstash to parse logs or whatever) and Elastic as a centralized store. Then, I would use Spark for reading from ES and processing data to create some machine learning models.

But, during the previous search, some people use Spark to preprocess the data before to write into ES. Why do not they use Logstash for that? Is Spark better than Logstash?

Is there any way to know a possible bottleneck in this kind of steps?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 6, 2021, 9:52am UTC](https://discuss.elastic.co/t/logstash-or-spark-elasticsearch/269313/2 "2021-04-06T09:52:00Z")

</div>

Welcome!

> [@alvgoro](#):
>
> I thinked about Beats to collect logs (and maybe Logstash to parse logs or whatever)

Very good. Elasticsearch has built-in ingest pipelines I'd recommend instead of adding Logstash.

> [@alvgoro](#):
>
> Then, I would use Spark for reading from ES and processing data to create some machine learning models.

Just note that we do have a built-in machine learning feature (available on cloud or with a commercial license) which does that out of the box to perform things like anomaly detection for example. So may be you don't need to reinvent the wheel.

> [@alvgoro](#):
>
> Why do not they use Logstash for that?

I guess that Spark is more for computing things where Logstash is built to parse, enrich and load single events one by one. Logstash is an ETL. Spark is an analytics engine. Not the same purpose...

---

<div class="post-metadata">

**Author:** ![alvgoro](https://avatars.discourse-cdn.com/v4/letter/a/aeb1de/32.png) [@alvgoro](https://discuss.elastic.co/u/alvgoro)\
**Post date:** [April 6, 2021, 10:24am UTC](https://discuss.elastic.co/t/logstash-or-spark-elasticsearch/269313/3 "2021-04-06T10:24:50Z")

</div>

Thanks for your early reply,

> Just note that we do have a built-in machine learning feature (available on cloud or with a commercial license) which does that out of the box to perform things like anomaly detection for example. So may be you don't need to reinvent the wheel.

Yes, I used it during the 30 days trial license, but I would like to use Spark (or others) for more flexibility. I mean, I would like to have the possibility to get the data out from ES.

> I guess that Spark is more for computing things where Logstash is built to parse, enrich and load single events one by one. Logstash is an ETL. Spark is an analytics engine. Not the same purpose...

Ummm, I am overwhelmed... It is not clear for me what I have to use. I know there is a Spark-ES connector and I tried it yesterday. Is not it the best solution? ☹

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 6, 2021, 4:05pm UTC](https://discuss.elastic.co/t/logstash-or-spark-elasticsearch/269313/4 "2021-04-06T16:05:13Z")

</div>

> [@alvgoro](#):
>
> I know there is a Spark-ES connector and I tried it yesterday. Is not it the best solution? ☹

I don't know it. Probably that if you want to run a "Spark job", that's the best option.  
I'd probably start from [Apache Spark support | Elasticsearch for Apache Hadoop [8.11] | Elastic](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/spark.html#spark-read)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 4, 2021, 4:05pm UTC](https://discuss.elastic.co/t/logstash-or-spark-elasticsearch/269313/5 "2021-05-04T16:05:43Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
