# How to construct Spark DStream from continued RDD？

**URL:** <https://discuss.elastic.co/t/how-to-construct-spark-dstream-from-continued-rdd/44388>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [March 15, 2016, 2:54am UTC](https://discuss.elastic.co/t/how-to-construct-spark-dstream-from-continued-rdd/44388 "2016-03-15T02:54:16Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Kramer\_Li](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kramer_li/32/46872_2.png) [@Kramer\_Li](https://discuss.elastic.co/u/Kramer_Li)\
**Post date:** [March 15, 2016, 2:54am UTC](https://discuss.elastic.co/t/how-to-construct-spark-dstream-from-continued-rdd/44388/1 "2016-03-15T02:54:16Z")

</div>

I`m reading data from ElasticSearch to Spark every 5min. So there will be a RDD every 5 minutes.

I hope to construct a DStream based on these RDDs, so that I can get report for data within last 1 day, last 1 hour , last 5 minutes and so on.

To construct the DStream, I was thinking about create my own receiver, but the official documents of spark only give information using scala or java to do so. And I use python.

So do you know any way to do it? I know we can. After all the DStream is a series of RDDs, of course we should be about create DStream from continued RDDs. I just do not know how. Please give some advice

apache elasticsearch

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [March 21, 2016, 10:17am UTC](https://discuss.elastic.co/t/how-to-construct-spark-dstream-from-continued-rdd/44388/2 "2016-03-21T10:17:59Z")

</div>

I'm afraid I can't help when it comes to Python. ES-Hadoop is based on the JVM - Spark python integration can leverage the `InputFormat`/`OutputFormat` but that's not enough when it comes to an RDD in terms of efficiency.  
There is however a community wrapper in python around ES-Hadoop that is available on github; maybe that one will address your problem.

> [@Kramer\_Li](#):
>
> apache elasticsearch

Pardon?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:25pm UTC](https://discuss.elastic.co/t/how-to-construct-spark-dstream-from-continued-rdd/44388/3 "2017-07-06T13:25:41Z")

</div>


