# Processing data in ES (in sequential approach). What would be the best approach?

**URL:** <https://discuss.elastic.co/t/processing-data-in-es-in-sequential-approach-what-would-be-the-best-approach/6347>\
**Category:** Elasticsearch\
**Created:** [January 11, 2012, 2:55pm UTC](https://discuss.elastic.co/t/processing-data-in-es-in-sequential-approach-what-would-be-the-best-approach/6347 "2012-01-11T14:55:47Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [January 11, 2012, 2:55pm UTC](https://discuss.elastic.co/t/processing-data-in-es-in-sequential-approach-what-would-be-the-best-approach/6347/1 "2012-01-11T14:55:47Z")

</div>

Hi,

I need to implement some document enhancing functionality and I am looking  
for the best practices/examples/references about how to do it. Basically I  
need the following:  
1/ the code would be scheduled to start every minute or so  
2/ it would pull data from one index, process it and insert (or update)  
into other index  
3/ I need to make sure that at certain points there is only a single  
processing unit for whole cluster (parallel execution could lead to  
inaccurate results)

Given the above points I think I can implement it as a river, especially  
due to #3 but I have also concerns about it:

- if the code has some bugs (for example memory leaks), what would be the  
impact on cluster if it runs as a river? And is there anything I can do to  
minimize risk that crappy river hurts ES cluster?
- although river can execute any general code, its origin is to allow for  
pull/push data from external sources. Isn't it serious misuse to use river  
to get data from one index and index it into another index within the same  
cluster?
- as for the scheduling, I know it is possible start Java Timer inside the  
river but isn't there any built in scheduling API in ES that I could use  
instead?

Regards,  
Lukas

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [January 12, 2012, 10:29am UTC](https://discuss.elastic.co/t/processing-data-in-es-in-sequential-approach-what-would-be-the-best-approach/6347/2 "2012-01-12T10:29:46Z")

</div>

On Wed, Jan 11, 2012 at 4:55 PM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com) wrote:

> Hi,
> 
> I need to implement some document enhancing functionality and I am looking  
> for the best practices/examples/references about how to do it. Basically I  
> need the following:  
> 1/ the code would be scheduled to start every minute or so  
> 2/ it would pull data from one index, process it and insert (or update)  
> into other index  
> 3/ I need to make sure that at certain points there is only a single  
> processing unit for whole cluster (parallel execution could lead to  
> inaccurate results)
> 
> Given the above points I think I can implement it as a river, especially  
> due to #3 but I have also concerns about it:
> 
> - if the code has some bugs (for example memory leaks), what would be the  
> impact on cluster if it runs as a river? And is there anything I can do to  
> minimize risk that crappy river hurts ES cluster?

Not much, if it leaks memory then it will cause OOM on that node.

> - although river can execute any general code, its origin is to allow for  
> pull/push data from external sources. Isn't it serious misuse to use river  
> to get data from one index and index it into another index within the same  
> cluster?

It can be done, don't think its a misuse.

> - as for the scheduling, I know it is possible start Java Timer inside the  
> river but isn't there any built in scheduling API in ES that I could use  
> instead?

I suggest you use your own, thats fine. Just make sure to close it when the  
river closes.

> Regards,  
> Lukas

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:42am UTC](https://discuss.elastic.co/t/processing-data-in-es-in-sequential-approach-what-would-be-the-best-approach/6347/3 "2017-07-06T03:42:59Z")

</div>


