# Indexing around 140 million addresses - need some performance tips

**URL:** <https://discuss.elastic.co/t/indexing-around-140-million-addresses-need-some-performance-tips/13462>\
**Category:** Elasticsearch\
**Created:** [September 4, 2013, 7:30pm UTC](https://discuss.elastic.co/t/indexing-around-140-million-addresses-need-some-performance-tips/13462 "2013-09-04T19:30:30Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![Anthony\_Campagna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anthony_campagna/32/2079_2.png) [@Anthony\_Campagna](https://discuss.elastic.co/u/Anthony_Campagna)\
**Post date:** [September 4, 2013, 7:30pm UTC](https://discuss.elastic.co/t/indexing-around-140-million-addresses-need-some-performance-tips/13462/1 "2013-09-04T19:30:30Z")

</div>

I am about to begin a project to index 140 million documents with street  
addresses, city, state, and zip. I will need to do searches against the  
entire index everytime a user types in a letter in our search. I was  
wondering if there was any way to organize street addresses in my index in  
a smarter way than just dumping them all in a single index and type. Maybe  
a different type for each number? Or maybe it's a clever way of using  
nested or parent/child objects/documents. Or maybe it's not necessary at  
all and we can just rely on a good use of filters. If that's the case, any  
suggestions on how to filter this while doing searches?

Example of a street address (which is the field that we will be searching  
against): 243 Broadway, New York, NY 10060

Example document:

{  
street\_address: "243 broadway",  
street\_number: 243,  
street\_name: "broadway",  
city: "New York",  
state: "NY",  
zip: "10060"  
}

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Max\_Seleznev](https://avatars.discourse-cdn.com/v4/letter/m/d6d6ee/32.png) [@Max\_Seleznev](https://discuss.elastic.co/u/Max_Seleznev)\
**Post date:** [September 4, 2013, 7:48pm UTC](https://discuss.elastic.co/t/indexing-around-140-million-addresses-need-some-performance-tips/13462/2 "2013-09-04T19:48:29Z")

</div>

I think you should try to experiment  
if you do that just search for a string in a string?

{  
_full\_address: "243 Broadway, New York, NY 10060",_  
street\_address: "243 broadway",  
street\_number: 243,  
street\_name: "broadway",  
city: "New York",  
state: "NY",  
zip: "10060"  
}

On Wednesday, September 4, 2013 3:30:30 PM UTC-4, Anthony Campagna wrote:

> I am about to begin a project to index 140 million documents with street  
> addresses, city, state, and zip. I will need to do searches against the  
> entire index everytime a user types in a letter in our search. I was  
> wondering if there was any way to organize street addresses in my index in  
> a smarter way than just dumping them all in a single index and type. Maybe  
> a different type for each number? Or maybe it's a clever way of using  
> nested or parent/child objects/documents. Or maybe it's not necessary at  
> all and we can just rely on a good use of filters. If that's the case, any  
> suggestions on how to filter this while doing searches?
> 
> Example of a street address (which is the field that we will be searching  
> against): 243 Broadway, New York, NY 10060
> 
> Example document:
> 
> {  
> street\_address: "243 broadway",  
> street\_number: 243,  
> street\_name: "broadway",  
> city: "New York",  
> state: "NY",  
> zip: "10060"  
> }

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Anthony\_Campagna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anthony_campagna/32/2079_2.png) [@Anthony\_Campagna](https://discuss.elastic.co/u/Anthony_Campagna)\
**Post date:** [September 4, 2013, 8:04pm UTC](https://discuss.elastic.co/t/indexing-around-140-million-addresses-need-some-performance-tips/13462/3 "2013-09-04T20:04:27Z")

</div>

What exactly do you mean by searching for a string within a string? A fuzzy  
search? If so, I was under the impression that fuzzy searches are  
incredibly slow and resource intensive compared to other methods of  
searching.

I do plan on doing plenty of exprimenting but I have two barriers to that:

1. There are just so many options available to me, I was curious what  
others thing or have found to be successful
2. I'm not exactly sure how to quantify how resource-intensive/taxing a  
single query is on an elasticsearch cluster

On Wednesday, September 4, 2013 3:48:29 PM UTC-4, Max Seleznev wrote:

> I think you should try to experiment  
> if you do that just search for a string in a string?
> 
> {  
> _full\_address: "243 Broadway, New York, NY 10060",_  
> street\_address: "243 broadway",  
> street\_number: 243,  
> street\_name: "broadway",  
> city: "New York",  
> state: "NY",  
> zip: "10060"  
> }
> 
> On Wednesday, September 4, 2013 3:30:30 PM UTC-4, Anthony Campagna wrote:
> 
> > I am about to begin a project to index 140 million documents with street  
> > addresses, city, state, and zip. I will need to do searches against the  
> > entire index everytime a user types in a letter in our search. I was  
> > wondering if there was any way to organize street addresses in my index in  
> > a smarter way than just dumping them all in a single index and type. Maybe  
> > a different type for each number? Or maybe it's a clever way of using  
> > nested or parent/child objects/documents. Or maybe it's not necessary at  
> > all and we can just rely on a good use of filters. If that's the case, any  
> > suggestions on how to filter this while doing searches?
> > 
> > Example of a street address (which is the field that we will be searching  
> > against): 243 Broadway, New York, NY 10060
> > 
> > Example document:
> > 
> > {  
> > street\_address: "243 broadway",  
> > street\_number: 243,  
> > street\_name: "broadway",  
> > city: "New York",  
> > state: "NY",  
> > zip: "10060"  
> > }

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Max\_Seleznev](https://avatars.discourse-cdn.com/v4/letter/m/d6d6ee/32.png) [@Max\_Seleznev](https://discuss.elastic.co/u/Max_Seleznev)\
**Post date:** [September 4, 2013, 8:13pm UTC](https://discuss.elastic.co/t/indexing-around-140-million-addresses-need-some-performance-tips/13462/4 "2013-09-04T20:13:56Z")

</div>

No, I did not mean fuzzy search.

what are you use the scheme? what types and analyzers?  
can you show full scheme?

On Wednesday, September 4, 2013 4:04:27 PM UTC-4, Anthony Campagna wrote:

> What exactly do you mean by searching for a string within a string? A  
> fuzzy search? If so, I was under the impression that fuzzy searches are  
> incredibly slow and resource intensive compared to other methods of  
> searching.
> 
> I do plan on doing plenty of exprimenting but I have two barriers to that:
> 
> 1. There are just so many options available to me, I was curious what  
> others thing or have found to be successful
> 2. I'm not exactly sure how to quantify how resource-intensive/taxing a  
> single query is on an elasticsearch cluster
> 
> On Wednesday, September 4, 2013 3:48:29 PM UTC-4, Max Seleznev wrote:
> 
> > I think you should try to experiment  
> > if you do that just search for a string in a string?
> > 
> > {  
> > _full\_address: "243 Broadway, New York, NY 10060",_  
> > street\_address: "243 broadway",  
> > street\_number: 243,  
> > street\_name: "broadway",  
> > city: "New York",  
> > state: "NY",  
> > zip: "10060"  
> > }
> > 
> > On Wednesday, September 4, 2013 3:30:30 PM UTC-4, Anthony Campagna wrote:
> > 
> > > I am about to begin a project to index 140 million documents with street  
> > > addresses, city, state, and zip. I will need to do searches against the  
> > > entire index everytime a user types in a letter in our search. I was  
> > > wondering if there was any way to organize street addresses in my index in  
> > > a smarter way than just dumping them all in a single index and type. Maybe  
> > > a different type for each number? Or maybe it's a clever way of using  
> > > nested or parent/child objects/documents. Or maybe it's not necessary at  
> > > all and we can just rely on a good use of filters. If that's the case, any  
> > > suggestions on how to filter this while doing searches?
> > > 
> > > Example of a street address (which is the field that we will be  
> > > searching against): 243 Broadway, New York, NY 10060
> > > 
> > > Example document:
> > > 
> > > {  
> > > street\_address: "243 broadway",  
> > > street\_number: 243,  
> > > street\_name: "broadway",  
> > > city: "New York",  
> > > state: "NY",  
> > > zip: "10060"  
> > > }

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Anthony\_Campagna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anthony_campagna/32/2079_2.png) [@Anthony\_Campagna](https://discuss.elastic.co/u/Anthony_Campagna)\
**Post date:** [September 4, 2013, 8:33pm UTC](https://discuss.elastic.co/t/indexing-around-140-million-addresses-need-some-performance-tips/13462/5 "2013-09-04T20:33:03Z")

</div>

I have not gotten that far yet. The whole point of this thread is to see  
what others suggest or have done. Scheme and Analyzers would be completely  
determined on what route I decide to go, which might be heavily influenced  
by this thread.

On Wednesday, September 4, 2013 4:13:56 PM UTC-4, Max Seleznev wrote:

> No, I did not mean fuzzy search.
> 
> what are you use the scheme? what types and analyzers?  
> can you show full scheme?
> 
> On Wednesday, September 4, 2013 4:04:27 PM UTC-4, Anthony Campagna wrote:
> 
> > What exactly do you mean by searching for a string within a string? A  
> > fuzzy search? If so, I was under the impression that fuzzy searches are  
> > incredibly slow and resource intensive compared to other methods of  
> > searching.
> > 
> > I do plan on doing plenty of exprimenting but I have two barriers to that:
> > 
> > 1. There are just so many options available to me, I was curious what  
> > others thing or have found to be successful
> > 2. I'm not exactly sure how to quantify how resource-intensive/taxing a  
> > single query is on an elasticsearch cluster
> > 
> > On Wednesday, September 4, 2013 3:48:29 PM UTC-4, Max Seleznev wrote:
> > 
> > > I think you should try to experiment  
> > > if you do that just search for a string in a string?
> > > 
> > > {  
> > > _full\_address: "243 Broadway, New York, NY 10060",_  
> > > street\_address: "243 broadway",  
> > > street\_number: 243,  
> > > street\_name: "broadway",  
> > > city: "New York",  
> > > state: "NY",  
> > > zip: "10060"  
> > > }
> > > 
> > > On Wednesday, September 4, 2013 3:30:30 PM UTC-4, Anthony Campagna wrote:
> > > 
> > > > I am about to begin a project to index 140 million documents with  
> > > > street addresses, city, state, and zip. I will need to do searches against  
> > > > the entire index everytime a user types in a letter in our search. I was  
> > > > wondering if there was any way to organize street addresses in my index in  
> > > > a smarter way than just dumping them all in a single index and type. Maybe  
> > > > a different type for each number? Or maybe it's a clever way of using  
> > > > nested or parent/child objects/documents. Or maybe it's not necessary at  
> > > > all and we can just rely on a good use of filters. If that's the case, any  
> > > > suggestions on how to filter this while doing searches?
> > > > 
> > > > Example of a street address (which is the field that we will be  
> > > > searching against): 243 Broadway, New York, NY 10060
> > > > 
> > > > Example document:
> > > > 
> > > > {  
> > > > street\_address: "243 broadway",  
> > > > street\_number: 243,  
> > > > street\_name: "broadway",  
> > > > city: "New York",  
> > > > state: "NY",  
> > > > zip: "10060"  
> > > > }

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Israel\_Ekpo](https://avatars.discourse-cdn.com/v4/letter/i/49beb7/32.png) [@Israel\_Ekpo](https://discuss.elastic.co/u/Israel_Ekpo)\
**Post date:** [September 4, 2013, 8:33pm UTC](https://discuss.elastic.co/t/indexing-around-140-million-addresses-need-some-performance-tips/13462/6 "2013-09-04T20:33:14Z")

</div>

One thing you should consider is how to handle synonyms in the street name.

So for example:

123 North Main Street could be equivalent to  
123 N Main Street  
123 N Main St  
123 North Main St

500 Second Street  
500 2nd St  
500 2nd Street

Also consider stripping out commas and other delimiters if you store the  
entire address as one field

So users may search with commas and some may search without commas.

> **[Street name](https://en.wikipedia.org/wiki/Street_or_road_name)**
>
> A street name is an identifying name given to a street or road. In toponymic terminology, names of streets and roads are referred to as odonyms or hodonyms (from Ancient Greek ὁδός hodós 'road', and ὄνυμα ónuma 'name', i.e., the Doric and Aeolic form of ὄνομα ónoma 'name'). The street name usually forms part of the address (though addresses in some parts of the world, notably most of Japan, make no reference to street names). Buildings are often given numbers along the street to further help i Na...

_Author and Instructor for the Upcoming Book and Lecture Series_  
_Massive Log Data Aggregation, Processing, Searching and Visualization with  
Open Source Software_  
_[http://massivelogdata.com](http://massivelogdata.com)_

On Wed, Sep 4, 2013 at 3:30 PM, Anthony Campagna [gucommander@gmail.com](mailto:gucommander@gmail.com)wrote:

> I am about to begin a project to index 140 million documents with street  
> addresses, city, state, and zip. I will need to do searches against the  
> entire index everytime a user types in a letter in our search. I was  
> wondering if there was any way to organize street addresses in my index in  
> a smarter way than just dumping them all in a single index and type. Maybe  
> a different type for each number? Or maybe it's a clever way of using  
> nested or parent/child objects/documents. Or maybe it's not necessary at  
> all and we can just rely on a good use of filters. If that's the case, any  
> suggestions on how to filter this while doing searches?
> 
> Example of a street address (which is the field that we will be searching  
> against): 243 Broadway, New York, NY 10060
> 
> Example document:
> 
> {  
> street\_address: "243 broadway",  
> street\_number: 243,  
> street\_name: "broadway",  
> city: "New York",  
> state: "NY",  
> zip: "10060"  
> }
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Anthony\_Campagna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anthony_campagna/32/2079_2.png) [@Anthony\_Campagna](https://discuss.elastic.co/u/Anthony_Campagna)\
**Post date:** [September 4, 2013, 8:45pm UTC](https://discuss.elastic.co/t/indexing-around-140-million-addresses-need-some-performance-tips/13462/7 "2013-09-04T20:45:12Z")

</div>

Yea, i'm not 100% sure how i'm going to handle synonyms yet but it's on my  
radar.

On Wednesday, September 4, 2013 4:33:14 PM UTC-4, Israel Ekpo wrote:

> One thing you should consider is how to handle synonyms in the street name.
> 
> So for example:
> 
> 123 North Main Street could be equivalent to  
> 123 N Main Street  
> 123 N Main St  
> 123 North Main St
> 
> 500 Second Street  
> 500 2nd St  
> 500 2nd Street
> 
> Also consider stripping out commas and other delimiters if you store the  
> entire address as one field
> 
> So users may search with commas and some may search without commas.
> 
> [Street name - Wikipedia](http://en.wikipedia.org/wiki/Street_or_road_name)
> 
> _Author and Instructor for the Upcoming Book and Lecture Series_  
> _Massive Log Data Aggregation, Processing, Searching and Visualization  
> with Open Source Software_  
> _[http://massivelogdata.com](http://massivelogdata.com)_
> 
> On Wed, Sep 4, 2013 at 3:30 PM, Anthony Campagna \<[gucom...@gmail.com](mailto:gucom...@gmail.com)\<javascript:\>
> 
> > wrote:
> 
> > I am about to begin a project to index 140 million documents with street  
> > addresses, city, state, and zip. I will need to do searches against the  
> > entire index everytime a user types in a letter in our search. I was  
> > wondering if there was any way to organize street addresses in my index in  
> > a smarter way than just dumping them all in a single index and type. Maybe  
> > a different type for each number? Or maybe it's a clever way of using  
> > nested or parent/child objects/documents. Or maybe it's not necessary at  
> > all and we can just rely on a good use of filters. If that's the case, any  
> > suggestions on how to filter this while doing searches?
> > 
> > Example of a street address (which is the field that we will be searching  
> > against): 243 Broadway, New York, NY 10060
> > 
> > Example document:
> > 
> > {  
> > street\_address: "243 broadway",  
> > street\_number: 243,  
> > street\_name: "broadway",  
> > city: "New York",  
> > state: "NY",  
> > zip: "10060"  
> > }
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Max\_Seleznev](https://avatars.discourse-cdn.com/v4/letter/m/d6d6ee/32.png) [@Max\_Seleznev](https://discuss.elastic.co/u/Max_Seleznev)\
**Post date:** [September 4, 2013, 8:45pm UTC](https://discuss.elastic.co/t/indexing-around-140-million-addresses-need-some-performance-tips/13462/8 "2013-09-04T20:45:32Z")

</div>

yes, you right

https://maps.google.com/maps?q=second%2Bave%2BNew%2BYork,%2BNY&hl=en&ll=40.757561,-73.966235&spn=0.001764,0.003954&sll=40.758816,-73.974703&sspn=0.007054,0.015814&hnear=Second%2BAve,%2BNew%2BYork&t=m&z=19&iwloc=A

google maps all this can do

i search: second ave New York, NY and Google maps show me correct address

Did you mean:  
\*2nd Ave, New York, NY 10022[https://maps.google.com/maps?q=2nd+Ave,+New+York,+10022&hl=en&sll=40.758816,-73.974703&sspn=0.007054,0.015814&hnear=Second+Ave,+New+York&t=m&ie=UTF8&oi=georefine&ct=clnk&cd=2&geocode=FSDubQIdUTyX-w&split=0](https://maps.google.com/maps?q=2nd+Ave,+New+York,+10022&hl=en&sll=40.758816,-73.974703&sspn=0.007054,0.015814&hnear=Second+Ave,+New+York&t=m&ie=UTF8&oi=georefine&ct=clnk&cd=2&geocode=FSDubQIdUTyX-w&split=0)  
\*

On Wednesday, September 4, 2013 4:33:14 PM UTC-4, Israel Ekpo wrote:

> One thing you should consider is how to handle synonyms in the street name.
> 
> So for example:
> 
> 123 North Main Street could be equivalent to  
> 123 N Main Street  
> 123 N Main St  
> 123 North Main St
> 
> 500 Second Street  
> 500 2nd St  
> 500 2nd Street
> 
> Also consider stripping out commas and other delimiters if you store the  
> entire address as one field
> 
> So users may search with commas and some may search without commas.
> 
> [Street name - Wikipedia](http://en.wikipedia.org/wiki/Street_or_road_name)
> 
> _Author and Instructor for the Upcoming Book and Lecture Series_  
> _Massive Log Data Aggregation, Processing, Searching and Visualization  
> with Open Source Software_  
> _[http://massivelogdata.com](http://massivelogdata.com)_
> 
> On Wed, Sep 4, 2013 at 3:30 PM, Anthony Campagna \<[gucom...@gmail.com](mailto:gucom...@gmail.com)\<javascript:\>
> 
> > wrote:
> 
> > I am about to begin a project to index 140 million documents with street  
> > addresses, city, state, and zip. I will need to do searches against the  
> > entire index everytime a user types in a letter in our search. I was  
> > wondering if there was any way to organize street addresses in my index in  
> > a smarter way than just dumping them all in a single index and type. Maybe  
> > a different type for each number? Or maybe it's a clever way of using  
> > nested or parent/child objects/documents. Or maybe it's not necessary at  
> > all and we can just rely on a good use of filters. If that's the case, any  
> > suggestions on how to filter this while doing searches?
> > 
> > Example of a street address (which is the field that we will be searching  
> > against): 243 Broadway, New York, NY 10060
> > 
> > Example document:
> > 
> > {  
> > street\_address: "243 broadway",  
> > street\_number: 243,  
> > street\_name: "broadway",  
> > city: "New York",  
> > state: "NY",  
> > zip: "10060"  
> > }
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Max\_Seleznev](https://avatars.discourse-cdn.com/v4/letter/m/d6d6ee/32.png) [@Max\_Seleznev](https://discuss.elastic.co/u/Max_Seleznev)\
**Post date:** [September 4, 2013, 8:47pm UTC](https://discuss.elastic.co/t/indexing-around-140-million-addresses-need-some-performance-tips/13462/9 "2013-09-04T20:47:21Z")

</div>

then your search results will not be as good  
you need use synonyms

On Wednesday, September 4, 2013 4:45:12 PM UTC-4, Anthony Campagna wrote:

> Yea, i'm not 100% sure how i'm going to handle synonyms yet but it's on my  
> radar.
> 
> On Wednesday, September 4, 2013 4:33:14 PM UTC-4, Israel Ekpo wrote:
> 
> > One thing you should consider is how to handle synonyms in the street  
> > name.
> > 
> > So for example:
> > 
> > 123 North Main Street could be equivalent to  
> > 123 N Main Street  
> > 123 N Main St  
> > 123 North Main St
> > 
> > 500 Second Street  
> > 500 2nd St  
> > 500 2nd Street
> > 
> > Also consider stripping out commas and other delimiters if you store the  
> > entire address as one field
> > 
> > So users may search with commas and some may search without commas.
> > 
> > [Street name - Wikipedia](http://en.wikipedia.org/wiki/Street_or_road_name)
> > 
> > _Author and Instructor for the Upcoming Book and Lecture Series_  
> > _Massive Log Data Aggregation, Processing, Searching and Visualization  
> > with Open Source Software_  
> > _[http://massivelogdata.com](http://massivelogdata.com)_
> > 
> > On Wed, Sep 4, 2013 at 3:30 PM, Anthony Campagna [gucom...@gmail.com](mailto:gucom...@gmail.com)wrote:
> > 
> > > I am about to begin a project to index 140 million documents with street  
> > > addresses, city, state, and zip. I will need to do searches against the  
> > > entire index everytime a user types in a letter in our search. I was  
> > > wondering if there was any way to organize street addresses in my index in  
> > > a smarter way than just dumping them all in a single index and type. Maybe  
> > > a different type for each number? Or maybe it's a clever way of using  
> > > nested or parent/child objects/documents. Or maybe it's not necessary at  
> > > all and we can just rely on a good use of filters. If that's the case, any  
> > > suggestions on how to filter this while doing searches?
> > > 
> > > Example of a street address (which is the field that we will be  
> > > searching against): 243 Broadway, New York, NY 10060
> > > 
> > > Example document:
> > > 
> > > {  
> > > street\_address: "243 broadway",  
> > > street\_number: 243,  
> > > street\_name: "broadway",  
> > > city: "New York",  
> > > state: "NY",  
> > > zip: "10060"  
> > > }
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google  
> > > Groups "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send  
> > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).  
> > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Anthony\_Campagna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anthony_campagna/32/2079_2.png) [@Anthony\_Campagna](https://discuss.elastic.co/u/Anthony_Campagna)\
**Post date:** [September 4, 2013, 8:54pm UTC](https://discuss.elastic.co/t/indexing-around-140-million-addresses-need-some-performance-tips/13462/10 "2013-09-04T20:54:01Z")

</div>

I know. It's on my radar. I'm just not 100% sure how i'm going to handle it  
yet. Having synonyms for N,S,E,W,NE,NW,SE,SW is fine. But to create a  
synonym dictionary of every number up to 200th is a lot to do. There might  
be a better way of doing it than a dictionary for numbers.

On Wednesday, September 4, 2013 4:47:21 PM UTC-4, Max Seleznev wrote:

> then your search results will not be as good  
> you need use synonyms
> 
> On Wednesday, September 4, 2013 4:45:12 PM UTC-4, Anthony Campagna wrote:
> 
> > Yea, i'm not 100% sure how i'm going to handle synonyms yet but it's on  
> > my radar.
> > 
> > On Wednesday, September 4, 2013 4:33:14 PM UTC-4, Israel Ekpo wrote:
> > 
> > > One thing you should consider is how to handle synonyms in the street  
> > > name.
> > > 
> > > So for example:
> > > 
> > > 123 North Main Street could be equivalent to  
> > > 123 N Main Street  
> > > 123 N Main St  
> > > 123 North Main St
> > > 
> > > 500 Second Street  
> > > 500 2nd St  
> > > 500 2nd Street
> > > 
> > > Also consider stripping out commas and other delimiters if you store the  
> > > entire address as one field
> > > 
> > > So users may search with commas and some may search without commas.
> > > 
> > > [Street name - Wikipedia](http://en.wikipedia.org/wiki/Street_or_road_name)
> > > 
> > > _Author and Instructor for the Upcoming Book and Lecture Series_  
> > > _Massive Log Data Aggregation, Processing, Searching and Visualization  
> > > with Open Source Software_  
> > > _[http://massivelogdata.com](http://massivelogdata.com)_
> > > 
> > > On Wed, Sep 4, 2013 at 3:30 PM, Anthony Campagna [gucom...@gmail.com](mailto:gucom...@gmail.com)wrote:
> > > 
> > > > I am about to begin a project to index 140 million documents with  
> > > > street addresses, city, state, and zip. I will need to do searches against  
> > > > the entire index everytime a user types in a letter in our search. I was  
> > > > wondering if there was any way to organize street addresses in my index in  
> > > > a smarter way than just dumping them all in a single index and type. Maybe  
> > > > a different type for each number? Or maybe it's a clever way of using  
> > > > nested or parent/child objects/documents. Or maybe it's not necessary at  
> > > > all and we can just rely on a good use of filters. If that's the case, any  
> > > > suggestions on how to filter this while doing searches?
> > > > 
> > > > Example of a street address (which is the field that we will be  
> > > > searching against): 243 Broadway, New York, NY 10060
> > > > 
> > > > Example document:
> > > > 
> > > > {  
> > > > street\_address: "243 broadway",  
> > > > street\_number: 243,  
> > > > street\_name: "broadway",  
> > > > city: "New York",  
> > > > state: "NY",  
> > > > zip: "10060"  
> > > > }
> > > > 
> > > > --  
> > > > You received this message because you are subscribed to the Google  
> > > > Groups "elasticsearch" group.  
> > > > To unsubscribe from this group and stop receiving emails from it, send  
> > > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).  
> > > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Israel\_Ekpo](https://avatars.discourse-cdn.com/v4/letter/i/49beb7/32.png) [@Israel\_Ekpo](https://discuss.elastic.co/u/Israel_Ekpo)\
**Post date:** [September 4, 2013, 8:54pm UTC](https://discuss.elastic.co/t/indexing-around-140-million-addresses-need-some-performance-tips/13462/11 "2013-09-04T20:54:07Z")

</div>

If you can share more information about the dataset and how you plan to use  
the the data, you may get better recommendations.

If you plan to do local search that is limited to specific geographical  
areas, it might be helpful to have the geo point type added as one of the  
fields.

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

This could simplify the process of narrowing down searches to a specific  
radius across multiple states, if necessary.

Are all these addresses within the United States? Or do you have other  
countries and territories?

How are they distributed in terms of number of documents per state?

Your responses to these questions will influence how the architecture is  
designed and configured.

_Author and Instructor for the Upcoming Book and Lecture Series_  
_Massive Log Data Aggregation, Processing, Searching and Visualization with  
Open Source Software_  
_[http://massivelogdata.com](http://massivelogdata.com)_

On Wed, Sep 4, 2013 at 3:30 PM, Anthony Campagna [gucommander@gmail.com](mailto:gucommander@gmail.com)wrote:

> I am about to begin a project to index 140 million documents with street  
> addresses, city, state, and zip. I will need to do searches against the  
> entire index everytime a user types in a letter in our search. I was  
> wondering if there was any way to organize street addresses in my index in  
> a smarter way than just dumping them all in a single index and type. Maybe  
> a different type for each number? Or maybe it's a clever way of using  
> nested or parent/child objects/documents. Or maybe it's not necessary at  
> all and we can just rely on a good use of filters. If that's the case, any  
> suggestions on how to filter this while doing searches?
> 
> Example of a street address (which is the field that we will be searching  
> against): 243 Broadway, New York, NY 10060
> 
> Example document:
> 
> {  
> street\_address: "243 broadway",  
> street\_number: 243,  
> street\_name: "broadway",  
> city: "New York",  
> state: "NY",  
> zip: "10060"  
> }
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Anthony\_Campagna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anthony_campagna/32/2079_2.png) [@Anthony\_Campagna](https://discuss.elastic.co/u/Anthony_Campagna)\
**Post date:** [September 4, 2013, 9:12pm UTC](https://discuss.elastic.co/t/indexing-around-140-million-addresses-need-some-performance-tips/13462/12 "2013-09-04T21:12:32Z")

</div>

I am working with just a small sample of a dataset that I am negotiating a  
purchase for. Assume the dataset just has what I posted in my original post  
plus lat/lon information. The dataset has 90%+ of all addresses in the  
United States, and only within the United States. Number of documents per  
state are unknown at this time, but will ultimately be fairly similar to  
the distribution of populations across the states. While we will have  
geopoints for all addresses, they will only be used for scoring. The  
address search must be done across the entire dataset. If I type in 100 8th  
st it should give me the top 10 closest results if there are more than 10,  
if there are less than 10 then it must be able to give me as many addresses  
as there are in the index reguardless of location.

On Wednesday, September 4, 2013 4:54:07 PM UTC-4, Israel Ekpo wrote:

> If you can share more information about the dataset and how you plan to  
> use the the data, you may get better recommendations.
> 
> If you plan to do local search that is limited to specific geographical  
> areas, it might be helpful to have the geo point type added as one of the  
> fields.
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/geo-point-type/)
> 
> This could simplify the process of narrowing down searches to a specific  
> radius across multiple states, if necessary.
> 
> Are all these addresses within the United States? Or do you have other  
> countries and territories?
> 
> How are they distributed in terms of number of documents per state?
> 
> Your responses to these questions will influence how the architecture is  
> designed and configured.
> 
> _Author and Instructor for the Upcoming Book and Lecture Series_  
> _Massive Log Data Aggregation, Processing, Searching and Visualization  
> with Open Source Software_  
> _[http://massivelogdata.com](http://massivelogdata.com)_
> 
> On Wed, Sep 4, 2013 at 3:30 PM, Anthony Campagna \<[gucom...@gmail.com](mailto:gucom...@gmail.com)\<javascript:\>
> 
> > wrote:
> 
> > I am about to begin a project to index 140 million documents with street  
> > addresses, city, state, and zip. I will need to do searches against the  
> > entire index everytime a user types in a letter in our search. I was  
> > wondering if there was any way to organize street addresses in my index in  
> > a smarter way than just dumping them all in a single index and type. Maybe  
> > a different type for each number? Or maybe it's a clever way of using  
> > nested or parent/child objects/documents. Or maybe it's not necessary at  
> > all and we can just rely on a good use of filters. If that's the case, any  
> > suggestions on how to filter this while doing searches?
> > 
> > Example of a street address (which is the field that we will be searching  
> > against): 243 Broadway, New York, NY 10060
> > 
> > Example document:
> > 
> > {  
> > street\_address: "243 broadway",  
> > street\_number: 243,  
> > street\_name: "broadway",  
> > city: "New York",  
> > state: "NY",  
> > zip: "10060"  
> > }
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:18am UTC](https://discuss.elastic.co/t/indexing-around-140-million-addresses-need-some-performance-tips/13462/13 "2017-07-06T02:18:10Z")

</div>


