# Use case

**URL:** <https://discuss.elastic.co/t/use-case/13006>\
**Category:** Elasticsearch\
**Created:** [July 30, 2013, 8:49pm UTC](https://discuss.elastic.co/t/use-case/13006 "2013-07-30T20:49:04Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Michael\_Andrews](https://avatars.discourse-cdn.com/v4/letter/m/b5ac83/32.png) [@Michael\_Andrews](https://discuss.elastic.co/u/Michael_Andrews)\
**Post date:** [July 30, 2013, 8:49pm UTC](https://discuss.elastic.co/t/use-case/13006/1 "2013-07-30T20:49:04Z")

</div>

Hello all, I was curious about your input on a potential use case for  
Elasticsearch. We collect massive amounts of netflow data (which is  
essentially a description of a TCP conversation that occurred between two  
endpoints through a router/switch/etc) which we save and analyze using  
basic netflow tools. Our current software solution is deployed on a single  
machine and can only store very limit quantities of netflow records, and  
does basic functions to read that data. We would like to use a scalable  
database of some kind to store the information that gives us the ability to  
do interesting queries over millions or even billions of records. We have  
experimented with Cassandra in this regard, however Cassandra really does  
not allow one to ask questions about the data as it's not capable of any  
aggregation other than a simple row count (we did create multiple data  
models that could answer a number of our queries, but being able to do ad  
hoc aggregation is much more appealing).

Our use case involves storing these netflow records which are comprised of  
a number of fields of data such as source IP address, destination IP  
address, source port, destination port, total bytes in the conversation,  
and a few others. We would like to be able to do interesting queries of  
the data such as asking the question "How many destination IP address's did  
the source IP address of w.x.y.z connect to between 4am and 7am on Tuesday  
last week?", or maybe ask the question "How many bytes of data were sent to  
the destination port 443 during the month of May?".

The biggest consideration is the size of the dataset, which we estimate  
would be adding tens of millions of records a day. Considering this, would  
Elasticsearch be capable of handling such a large volume of data or be able  
to do efficient searches over said data?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Christian\_Th](https://avatars.discourse-cdn.com/v4/letter/c/f14d63/32.png) [@Christian\_Th](https://discuss.elastic.co/u/Christian_Th)\
**Post date:** [July 30, 2013, 10:24pm UTC](https://discuss.elastic.co/t/use-case/13006/2 "2013-07-30T22:24:16Z")

</div>

10 millions of records in one hour? Yes, or did you mean per day ? Answer  
is still yes.  
The size of the dataset can be managed with ES, from my perspective. Your  
Client should be a little bit smart.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Norberto\_Meijome](https://avatars.discourse-cdn.com/v4/letter/n/ed655f/32.png) [@Norberto\_Meijome](https://discuss.elastic.co/u/Norberto_Meijome)\
**Post date:** [July 31, 2013, 9:21am UTC](https://discuss.elastic.co/t/use-case/13006/3 "2013-07-31T09:21:15Z")

</div>

Not trying to shoot down the idea (or maybe i'll hear some interesting  
counter opinions! ) - but at the moment there isn't a clear, well tested  
and reliable way to snapshot and backup an ES cluster's data - specially on  
a large, and active cluster (one where you simply cant stop sending updates  
/ queries to).  
AFAICT , ES is not to be considered your source of data. ( yes, to us it  
means keep the source of data somewhere 'stable' as a SQL DB, and have some  
trusted, tested process to rebuild the index).

Again, maybe (Hopefully), I'm wrong and I have missed some announcement wrt  
support for native snapshots / log streaming a'la RDBMS replication to a  
separate cluster for backups / snapshots / geographical  
federation/distribution..

cheers,  
B

On Wed, Jul 31, 2013 at 8:24 AM, Christian Th. [chth.exensio@gmail.com](mailto:chth.exensio@gmail.com)wrote:

> 10 millions of records in one hour? Yes, or did you mean per day ? Answer  
> is still yes.  
> The size of the dataset can be managed with ES, from my perspective. Your  
> Client should be a little bit smart.
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
Norberto 'Beto' Meijome

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Michael\_Andrews](https://avatars.discourse-cdn.com/v4/letter/m/b5ac83/32.png) [@Michael\_Andrews](https://discuss.elastic.co/u/Michael_Andrews)\
**Post date:** [July 31, 2013, 1:26pm UTC](https://discuss.elastic.co/t/use-case/13006/4 "2013-07-31T13:26:05Z")

</div>

Norberto,

Thanks for the insight, I am very new to Elasticsearch and I had wondered  
whether this system billed itself as a true data warehouse or simply a  
ephemeral search index. It seems it could take on a new life as a database  
considering it's capabilities working with such large data. We are now  
looking at Datastax Enterprise Solr integration which adds search on top of  
Cassandra, and possibly using Cassandra and Elasticsearch in conjunction.

I found a thread from 2010 where users were asking whether Elasticsearch  
would add multi-datacenter capabilities and it seems the idea never came to  
fruition  
([http://elasticsearch-users.115913.n3.nabble.com/ES-and-multiple-datacenters-td551665.html](http://elasticsearch-users.115913.n3.nabble.com/ES-and-multiple-datacenters-td551665.html)),  
which would be ideal for us as we operate two geo-redundant serving  
centers.

On Wednesday, July 31, 2013 4:21:15 AM UTC-5, Norberto Meijome wrote:

> Again, maybe (Hopefully), I'm wrong and I have missed some announcement  
> wrt support for native snapshots / log streaming a'la RDBMS replication to  
> a separate cluster for backups / snapshots / geographical  
> federation/distribution.
> 
> Norberto 'Beto' Meijome

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [July 31, 2013, 10:44pm UTC](https://discuss.elastic.co/t/use-case/13006/5 "2013-07-31T22:44:12Z")

</div>

Michael, afaik cross-datacenter syncing and snapshots for backup/restore  
are one important feature that may find the way into the 1.0 release.

Elasticsearch is not a ephemeral container. The concept of gateways  
[http://www.elasticsearch.org/guide/reference/modules/gateway/](http://www.elasticsearch.org/guide/reference/modules/gateway/) and the  
recovery from redundant sharded indexes allow reliable persistency of data.

With the new aggregation module  
[https://github.com/elasticsearch/elasticsearch/issues/3300](https://github.com/elasticsearch/elasticsearch/issues/3300) ES will become a  
real advanced big data analysis & search platform, also with map-reduce  
capabilities.

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Michael\_Andrews](https://avatars.discourse-cdn.com/v4/letter/m/b5ac83/32.png) [@Michael\_Andrews](https://discuss.elastic.co/u/Michael_Andrews)\
**Post date:** [August 1, 2013, 1:58pm UTC](https://discuss.elastic.co/t/use-case/13006/6 "2013-08-01T13:58:51Z")

</div>

Jörg,

Thanks for clarifying, I did go back and read more of the documentation and  
have become familiar with replicas and persistence. Can you speak to the  
use case I illustrated above in any fashion? We are very curious to hear  
about operational experiences scaling Elasticsearch and performance writing  
to and reading from massive data sets.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [August 1, 2013, 5:07pm UTC](https://discuss.elastic.co/t/use-case/13006/7 "2013-08-01T17:07:52Z")

</div>

Jorg,

Aggregations is really noble feature, but strictly speaking, I do not think  
it is really a map-reduce. Or is it?  
I might be wrong, but AFAICT there is not shuffle phase (it would be  
expensive in real-time), so instead of doing aggregations for all key  
values in single reducer, there are probably done a lot of particular  
segment (or node) level aggregations and then all these results are  
combined (so this IMO relies on the fact that you can get the final result  
by additions).

But as I said, I would be happy to be proven wrong.

Regards,  
Lukas

On Thu, Aug 1, 2013 at 12:44 AM, [joergprante@gmail.com](mailto:joergprante@gmail.com) \<  
[joergprante@gmail.com](mailto:joergprante@gmail.com)\> wrote:

> Michael, afaik cross-datacenter syncing and snapshots for backup/restore  
> are one important feature that may find the way into the 1.0 release.
> 
> Elasticsearch is not a ephemeral container. The concept of gateways  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/modules/gateway/) and the  
> recovery from redundant sharded indexes allow reliable persistency of data.
> 
> With the new aggregation module  
> [Aggregation Module - Phase 1 - Functional Design · Issue #3300 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/3300) ES will become  
> a real advanced big data analysis & search platform, also with map-reduce  
> capabilities.
> 
> Jörg
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [August 1, 2013, 7:56pm UTC](https://discuss.elastic.co/t/use-case/13006/8 "2013-08-01T19:56:06Z")

</div>

Lukáš,

aggregations are one piece for certain map-reduce-like algorithms, it could  
be accompanied by a modified bulk indexing action for presorting data so  
they arrive in-place at the shards to get them effectively processed by the  
aggregation framework.

I agree that ES will never be Hadoop but I see no reason why the ES  
distributed architecture should not be extensible to run some sort of  
simple map-reduce analysis on indexed documents. For example, creating  
ordered lists of all the values of all fields, for statistics. For example,  
in bibliographic data, librarians tend to ask how many occurrences of  
values are in what fields and what fields are less used.

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:23am UTC](https://discuss.elastic.co/t/use-case/13006/9 "2017-07-06T02:23:18Z")

</div>


