# Elastic Fan Out Read

**URL:** <https://discuss.elastic.co/t/elastic-fan-out-read/78757>\
**Category:** Elasticsearch\
**Created:** [March 15, 2017, 7:22pm UTC](https://discuss.elastic.co/t/elastic-fan-out-read/78757 "2017-03-15T19:22:57Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Nicholas\_Ventimiglia](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nicholas_ventimiglia/32/16383_2.png) [@Nicholas\_Ventimiglia](https://discuss.elastic.co/u/Nicholas_Ventimiglia)\
**Post date:** [March 15, 2017, 7:22pm UTC](https://discuss.elastic.co/t/elastic-fan-out-read/78757/1 "2017-03-15T19:22:57Z")

</div>

I am a first time user with elastic search. I am tasked with implementing a news feed which uses a fan-out-read to scan news from my friends and return the news in chronological order.

My first hypothesis is that I could create an index which is tagged with the publisher's user id. The consumer will get the news, passing a giant array of their friends user Ids to include in the request. Is this a realistic solution?

Additionally, a friend suggested that I could leverage shards (each publisher has their own shard, and the consumer will scan all their friends shards).

Am I on the right path ? Could you recommend any related best practices or blogs ?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [March 16, 2017, 2:58am UTC](https://discuss.elastic.co/t/elastic-fan-out-read/78757/2 "2017-03-16T02:58:27Z")

</div>

> [@Nicholas\_Ventimiglia](#):
>
> Additionally, a friend suggested that I could leverage shards (each publisher has their own shard, and the consumer will scan all their friends shards).

That'll hit limits eventually, each shard can only hold ~2^32 docs.

Just go with time based indices and tag them with the various IDs, then you can filter with good efficiency.

---

<div class="post-metadata">

**Author:** ![Nicholas\_Ventimiglia](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nicholas_ventimiglia/32/16383_2.png) [@Nicholas\_Ventimiglia](https://discuss.elastic.co/u/Nicholas_Ventimiglia)\
**Post date:** [March 16, 2017, 4:44pm UTC](https://discuss.elastic.co/t/elastic-fan-out-read/78757/3 "2017-03-16T16:44:55Z")

</div>

> That'll hit limits eventually, each shard can only hold ~2^32 docs.

What if I were to include a TTL on the document ?

> Just go with time based indices and tag them with the various IDs

I`m not sure what you mean by this. Like, create a new monthly index or something like that ? This seems strange to me.

That said, I came up with a minimal example yesterday. I have a single index and tagged each document with a publisherId. I could request documents with associated publishers using a Terms query pretty easily.

I had 400k documents in my single index and had no special shard logic at first. On a second iteration I split my documents into two shards. When I did this I saw no performance increases and was wondering if this extra step was worth the risk. Perhaps I was doing it wrong ?

Bulk insert Logic

```
            for (int p = 0; p < publisherCount; p++)
            {
                for (int r = 0; r < recordCount; r++)
                {
                    var ops = new BulkCreateDescriptor<NewsModel>();
                    ops.Routing(route); // "1" or "2"
                    ops.Document(new NewsModel
                    {
                        id = Guid.NewGuid().ToString(),
                        created = DateTime.UtcNow.Subtract(TimeSpan.FromHours(rand.Next(1, publisherCount * recordCount))),
                        publisherId = p + publisherStart
                    });

                    descriptor.AddOperation(ops);
                }
            }

            var response = await client.BulkAsync(descriptor);

```

Query Logic

```
        List<int> friends = new List<int>();
        for (int i = 0; i < friendCount; i++)
        {
            friends.Add(i);
        }

        var search = new SearchDescriptor<NewsModel>();
        
        search.Routing("1", "2");
        search.Sort(so => so.Descending(a => a.created));
        search.Size(size);
        search.Query(q => q.Terms(t => t.Field(f => f.publisherId).Terms<int>(friends)));

        var result = client.Search<NewsModel>(search);

```

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 16, 2017, 5:23pm UTC](https://discuss.elastic.co/t/elastic-fan-out-read/78757/4 "2017-03-16T17:23:52Z")

</div>

> What if I were to include a TTL on the document ?

TTL feature have been removed now. [Mapping changes | Elasticsearch Guide [5.2] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/5.2/breaking_50_mapping_changes.html#_literal __timestamp_literal_and_literal__ ttl_literal)

Removing docs in elasticsearch with something like a TTL field will cost you a lot of IOs. Removing a doc is actually adding somewhere an indicator that a doc has been removed. It does not really remove physically the doc one the disk until a Lucene merge operation happens.

Removing docs will generate a lot of merges, so a lot of IOs.

That's why @warkolm recommended:

> Just go with time based indices and tag them with the various IDs

That's the way to go.

> I had 400k documents in my single index and had no special shard logic at first. On a second iteration I split my documents into two shards. When I did this I saw no performance increases and was wondering if this extra step was worth the risk. Perhaps I was doing it wrong ?

No. You will see a lot of difference at scale. If you go to time based indices, I'd recommend to test if a single index, single shard can hold all the data you need per timeframe (day, month, whatever). And increase the number of shards if needed only.  
Having more shards will also help to spread the index load on more writers I'd say. So find the right balance for you.

My 2 cents.

---

<div class="post-metadata">

**Author:** ![Nicholas\_Ventimiglia](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nicholas_ventimiglia/32/16383_2.png) [@Nicholas\_Ventimiglia](https://discuss.elastic.co/u/Nicholas_Ventimiglia)\
**Post date:** [March 16, 2017, 5:45pm UTC](https://discuss.elastic.co/t/elastic-fan-out-read/78757/5 "2017-03-16T17:45:34Z")

</div>

> Just go with time based indices and tag them with the various IDs

So, maybe an index each week and then search for the last 4 weeks ? Then my cron can just nuke older indicies ? This sounds good.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 16, 2017, 5:56pm UTC](https://discuss.elastic.co/t/elastic-fan-out-read/78757/6 "2017-03-16T17:56:07Z")

</div>

Exactly!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 13, 2017, 5:56pm UTC](https://discuss.elastic.co/t/elastic-fan-out-read/78757/7 "2017-04-13T17:56:25Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
