# Social search

**URL:** <https://discuss.elastic.co/t/social-search/12023>\
**Category:** Elasticsearch\
**Created:** [May 18, 2013, 9:52pm UTC](https://discuss.elastic.co/t/social-search/12023 "2013-05-18T21:52:51Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mike\_Kaplinskiy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mike_kaplinskiy/32/1776_2.png) [@Mike\_Kaplinskiy](https://discuss.elastic.co/u/Mike_Kaplinskiy)\
**Post date:** [May 18, 2013, 9:52pm UTC](https://discuss.elastic.co/t/social-search/12023/1 "2013-05-18T21:52:51Z")

</div>

Hey folks,

I was reading the new features list in 0.90 and saw social search. The  
terms lookup mechanism seems to have some promise, but I have a few  
questions/issues:

- It doesn't seem to work for the \_id field (I.e. {"\_id": {"terms":{ ... }  
} })
- The design means that you need to store the entire set of followers in a  
single doc array. Would that mean reindexing the entire list (which for  
us can be 300K+ longs) whenever the list changes?
- if I wanted to denormalize the data instead and use a has\_child filter to  
check the relationship, do you have any hints on how to create the minimal  
possible child doc so 100M+ of these don't kill the index size? I would be  
fine with losing the ability to do any other type of query (well except for  
having a stable id for these docs). Here is what I have so far:

{"mapping": {"follower": {  
"\_parent": {"type": "user"},  
"\_source": {"enabled": false},  
"\_all": {"enabled": false},  
"properties": {  
"followerId": { "type": "long", "precision\_step": 0 },  
},  
} }

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [May 18, 2013, 10:02pm UTC](https://discuss.elastic.co/t/social-search/12023/2 "2013-05-18T22:02:19Z")

</div>

Hi Mike

I was reading the new features list in 0.90 and saw social search. The

> terms lookup mechanism seems to have some promise, but I have a few  
> questions/issues:
> 
> - It doesn't seem to work for the \_id field (I.e. {"\_id": {"terms":{ ... }  
> } })

you want:

{ terms: { \_id: { index... etc }}}

> - The design means that you need to store the entire set of followers in a  
> single doc array. Would that mean reindexing the entire list (which for  
> us can be 300K+ longs) whenever the list changes?

Yes, although you could break them down into smaller chunks and use a bool  
filter to combine them

> - if I wanted to denormalize the data instead and use a has\_child filter  
> to check the relationship, do you have any hints on how to create the  
> minimal possible child doc so 100M+ of these don't kill the index size? I  
> would be fine with losing the ability to do any other type of query (well  
> except for having a stable id for these docs). Here is what I have so far:
> 
> {"mapping": {"follower": {  
> "\_parent": {"type": "user"},  
> "\_source": {"enabled": false},  
> "\_all": {"enabled": false},  
> "properties": {  
> "followerId": { "type": "long", "precision\_step": 0 },  
> },  
> } }

I wouldn't disable the \_source field - you'll regret it later on, eg when  
you want to rebuild your index, or debug why a particular query isn't  
working as expected. And I wouldn't worry about the precision\_step either.

Also, in master, there is a big memory improvement on parent/child queries.  
Now only parent IDs are loaded into memory. Previously it used to load  
child IDs too

clint

> 

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Mike\_Kaplinskiy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mike_kaplinskiy/32/1776_2.png) [@Mike\_Kaplinskiy](https://discuss.elastic.co/u/Mike_Kaplinskiy)\
**Post date:** [May 20, 2013, 3:59am UTC](https://discuss.elastic.co/t/social-search/12023/3 "2013-05-20T03:59:21Z")

</div>

Hey Clinton,

Thanks for the quick reply.

On Saturday, May 18, 2013 6:02:19 PM UTC-4, Clinton Gormley wrote:

> Hi Mike
> 
> I was reading the new features list in 0.90 and saw social search. The
> 
> > terms lookup mechanism seems to have some promise, but I have a few  
> > questions/issues:
> > 
> > - It doesn't seem to work for the \_id field (I.e. {"\_id": {"terms":{ ...  
> > } } })
> 
> you want:
> 
> { terms: { \_id: { index... etc }}}

Sorry I that wasn't a valid test case. Here's one that doesn't work:

$ curl -XPUT [http://localhost:9200/index1/t1/123](http://localhost:9200/index1/t1/123) -d '{ "name": "123" }'  
{"ok":true,"\_index":"index1","\_type":"t1","\_id":"123","\_version":1}  
$ curl -XPUT [http://localhost:9200/index1/t1/456](http://localhost:9200/index1/t1/456) -d '{ "name": "456" }'  
{"ok":true,"\_index":"index1","\_type":"t1","\_id":"456","\_version":1}  
$ curl -XPUT [http://localhost:9200/index1/t2/1](http://localhost:9200/index1/t2/1) -d '{ "ids": ["123", "456"]  
}'  
{"ok":true,"\_index":"index1","\_type":"t2","\_id":"1","\_version":1}  
$ curl [http://localhost:9200/index1/t1/\_search](http://localhost:9200/index1/t1/_search) -d '{ "query": { "filtered":  
{ "filter": { "terms": { "\_id": { "index": "index1", "type": "t2", "id":  
"1", "path": "ids" } } } } } }'  
{"took":48,"timed\_out":false,"\_shards":{"total":5,"successful":5,"failed":0},"hits":{"total":0,"max\_score":null,"hits":}}  
$ curl [http://localhost:9200/index1/t1/\_search](http://localhost:9200/index1/t1/_search) -d '{ "query": { "filtered":  
{ "filter": { "terms": { "\_id": ["123", "456"] } } } } }'  
{"took":14,"timed\_out":false,"\_shards":{"total":5,"successful":5,"failed":0},"hits":{"total":2,"max\_score":1.0,"hits":[{"\_index":"index1","\_type":"t1","\_id":"456","\_score":1.0,  
"\_source" : { "name": "456"  
}},{"\_index":"index1","\_type":"t1","\_id":"123","\_score":1.0, "\_source" : {  
"name": "123" }}]}}

> > - The design means that you need to store the entire set of followers in  
> > a single doc array. Would that mean reindexing the entire list (which for  
> > us can be 300K+ longs) whenever the list changes?
> 
> Yes, although you could break them down into smaller chunks and use a bool  
> filter to combine them

Hmm good point.

> > - if I wanted to denormalize the data instead and use a has\_child filter  
> > to check the relationship, do you have any hints on how to create the  
> > minimal possible child doc so 100M+ of these don't kill the index size? I  
> > would be fine with losing the ability to do any other type of query (well  
> > except for having a stable id for these docs). Here is what I have so far:
> > 
> > {"mapping": {"follower": {  
> > "\_parent": {"type": "user"},  
> > "\_source": {"enabled": false},  
> > "\_all": {"enabled": false},  
> > "properties": {  
> > "followerId": { "type": "long", "precision\_step": 0 },  
> > },  
> > } }
> 
> I wouldn't disable the \_source field - you'll regret it later on, eg when  
> you want to rebuild your index, or debug why a particular query isn't  
> working as expected. And I wouldn't worry about the precision\_step either.

ES isn't the main datastore here, so reindexing from the database isn't an  
issue. I ran into an issue when doing this with the above mapping - the  
index got too big for the FS cache and query & indexing performance went  
through the floor. This was with 3 nodes with 15G ram and an EBS RAID0.  
Before adding the children the index was ~ 8G in size; afterwards it was  
80G which is ~ 680 bytes for a doc that's 2 ints.

> Also, in master, there is a big memory improvement on parent/child  
> queries. Now only parent IDs are loaded into memory. Previously it used to  
> load child IDs too

I saw that. I'm quite looking forward 0.90.1 - mostly because of the bulk  
update support. 🙂

> clint
> 
> >

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [May 20, 2013, 11:32am UTC](https://discuss.elastic.co/t/social-search/12023/4 "2013-05-20T11:32:50Z")

</div>

On 20 May 2013 05:59, Mike Kaplinskiy [mike.kaplinskiy@gmail.com](mailto:mike.kaplinskiy@gmail.com) wrote:

> $ curl -XPUT [http://localhost:9200/index1/t1/123](http://localhost:9200/index1/t1/123) -d '{ "name": "123" }'  
> {"ok":true,"\_index":"index1","\_type":"t1","\_id":"123","\_version":1}  
> $ curl -XPUT [http://localhost:9200/index1/t1/456](http://localhost:9200/index1/t1/456) -d '{ "name": "456" }'  
> {"ok":true,"\_index":"index1","\_type":"t1","\_id":"456","\_version":1}  
> $ curl -XPUT [http://localhost:9200/index1/t2/1](http://localhost:9200/index1/t2/1) -d '{ "ids": ["123",  
> "456"] }'  
> {"ok":true,"\_index":"index1","\_type":"t2","\_id":"1","\_version":1}  
> $ curl [http://localhost:9200/index1/t1/\_search](http://localhost:9200/index1/t1/_search) -d '{ "query": {  
> "filtered": { "filter": { "terms": { "\_id": { "index": "index1", "type":  
> "t2", "id": "1", "path": "ids" } } } } } }'
> 
> {"took":48,"timed\_out":false,"\_shards":{"total":5,"successful":5,"failed":0},"hits":{"total":0,"max\_score":null,"hits":}}  
> $ curl [http://localhost:9200/index1/t1/\_search](http://localhost:9200/index1/t1/_search) -d '{ "query": {  
> "filtered": { "filter": { "terms": { "\_id": ["123", "456"] } } } } }'  
> {"took":14,"timed\_out":false,"\_shards":{"total":5,"successful":5,"failed":0},"hits":{"total":2,"max\_score":1.0,"hits":[{"\_index":"index1","\_type":"t1","\_id":"456","\_score":1.0,  
> "\_source" : { "name": "456"  
> }},{"\_index":"index1","\_type":"t1","\_id":"123","\_score":1.0, "\_source" : {  
> "name": "123" }}]}}

You're right, this doesn't work. I've opened this issue:

> <https://github.com/elastic/elasticsearch/issues/3063>
>
> \`\`\`
> curl -XPUT http://localhost:9200/index1/t1/123 -d '{ "name": "123" }'
> curl -…XPUT http://localhost:9200/index1/t1/456 -d '{ "name": "456" }'
> curl -XPUT http://localhost:9200/index1/t2/1 -d '{ "ids": \["123", "456"\] }'
> \`\`\`
> 
> Query with external terms returns no results:
> 
> \`\`\`
> curl http://localhost:9200/index1/t1/\_search?pretty -d '{ "query": { "filtered": { "filter": { "terms": { "\_id": { "index": "index1", "type": "t2", "id": "1", "path": "ids" } } } } } }'
> \`\`\`
> 
> Query with listed terms works:
> 
> \`\`\`
> curl http://localhost:9200/index1/t1/\_search ?pretty -d '{ "query": { "filtered": { "filter": { "terms": { "\_id": \["123", "456"\] } } } } }'
> \`\`\`
> 
> External terms on \`name\` field works:
> 
> \`\`\`
> curl http://localhost:9200/index1/t1/\_search?pretty -d '{ "query": { "filtered": { "filter": { "terms": { "name": { "index": "index1", "type": "t2", "id": "1", "path": "ids" } } } } } }'
> \`\`\`
> 
> Side issue: unmapped field throws NPE:
> 
> \`\`\`
> curl http://localhost:9200/index1/t1/\_search?pretty -d '{ "query": { "filtered": { "filter": { "terms": { "XXX": { "index": "index1", "type": "t2", "id": "1", "path": "ids" } } } } } }'
> \`\`\`

clint

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:35am UTC](https://discuss.elastic.co/t/social-search/12023/5 "2017-07-06T02:35:44Z")

</div>


