# 2 Node cluster hanging while bulk indexing and adding types

**URL:** <https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211>\
**Category:** Elasticsearch\
**Created:** [December 21, 2011, 3:10pm UTC](https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211 "2011-12-21T15:10:33Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Greg\_Brown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greg_brown/32/1373_2.png) [@Greg\_Brown](https://discuss.elastic.co/u/Greg_Brown)\
**Post date:** [December 21, 2011, 3:10pm UTC](https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211/1 "2011-12-21T15:10:33Z")

</div>

Hi,

We're trying to index a few million responses to many (10000+) forms.  
We're creating a separate type for each form since they can have  
different fields. So each form is associated with a type named form-  
.

In bulk indexing the existing data, we are creating the new types and  
their mappings on the fly, interspersed with bulk indexing the  
documents. After running through about 200k responses the server  
stopped responding and appeared to be running out of memory. The  
cluster node stats are pasted below.

This two node cluster (8 GB machines) has indices for other  
applications (of similar size) already running successfully on it, and  
we've can bulk index them without a problem. So I assume the large  
number of types is causing the issue.

So some questions:

1. I'm planning to try creating all of the types first and then bulk  
index. Is intermixing type creation and adding docs expected to run  
into performance problems?

2. It seems from the cluster stats that there is a surprising amount  
of data going in one direction. Could this be the performance problem?

3. We're planning to do term facet searches (for word cloud  
generation) for each of the above types. I think I read that the  
entire index gets loaded when a term facet is done. If I do a term  
facet on a particular type, will only the portion of the index for  
that type be loaded, or will it still be the whole thing? If the whole  
thing, any way I can move the data into separate indices without  
having 10k indices show up when I go to look at the size of my  
indices. There is no need for them to be in the same index, I just  
don't want the 10k tables making it harder for me to examine the  
indices for other applications.

Thanks for the help  
-Greg

The index in question:  
"index" : {  
"primary\_size" : "55.8mb",  
"primary\_size\_in\_bytes" : 58586978,  
"size" : "111.7mb",  
"size\_in\_bytes" : 117177113  
},  
"translog" : {  
"operations" : 0  
},  
"docs" : {  
"num\_docs" : 214308,  
"max\_doc" : 214314,  
"deleted\_docs" : 6  
},  
"merges" : {  
"current" : 0,  
"current\_docs" : 0,  
"current\_size" : "0b",  
"current\_size\_in\_bytes" : 0,  
"total" : 0,  
"total\_time" : "0s",  
"total\_time\_in\_millis" : 0,  
"total\_docs" : 0,  
"total\_size" : "0b",  
"total\_size\_in\_bytes" : 0  
},  
"refresh" : {  
"total" : 6426,  
"total\_time" : "4.8m",  
"total\_time\_in\_millis" : 289292  
},  
"flush" : {  
"total" : 2702,  
"total\_time" : "1.6m",  
"total\_time\_in\_millis" : 97459  
},

/\_cluster/node/stats  
{  
"cluster\_name" : "mgs2",  
"nodes" : {  
"x25SzznQQZ2uPGQvgGjs4Q" : {  
"name" : "Nut",  
"indices" : {  
"store" : {  
"size" : "1.6gb",  
"size\_in\_bytes" : 1806371221  
},  
"docs" : {  
"count" : 557654,  
"deleted" : 114387  
},  
"indexing" : {  
"index\_total" : 18577042,  
"index\_time" : "40m",  
"index\_time\_in\_millis" : 2403162,  
"index\_current" : 0,  
"delete\_total" : 0,  
"delete\_time" : "0s",  
"delete\_time\_in\_millis" : 0,  
"delete\_current" : 0  
},  
"get" : {  
"total" : 86,  
"time" : "60ms",  
"time\_in\_millis" : 60,  
"exists\_total" : 24,  
"exists\_time" : "40ms",  
"exists\_time\_in\_millis" : 40,  
"missing\_total" : 62,  
"missing\_time" : "20ms",  
"missing\_time\_in\_millis" : 20,  
"current" : 0  
},  
"search" : {  
"query\_total" : 20121,  
"query\_time" : "4.8m",  
"query\_time\_in\_millis" : 288251,  
"query\_current" : 0,  
"fetch\_total" : 15337,  
"fetch\_time" : "57.4s",  
"fetch\_time\_in\_millis" : 57470,  
"fetch\_current" : 0  
},  
"cache" : {  
"field\_evictions" : 0,  
"field\_size" : "4.4mb",  
"field\_size\_in\_bytes" : 4675324,  
"filter\_count" : 3,  
"filter\_evictions" : 0,  
"filter\_size" : "92.9kb",  
"filter\_size\_in\_bytes" : 95144  
},  
"merges" : {  
"current" : 0,  
"current\_docs" : 0,  
"current\_size" : "0b",  
"current\_size\_in\_bytes" : 0,  
"total" : 56,  
"total\_time" : "15.4s",  
"total\_time\_in\_millis" : 15492,  
"total\_docs" : 1887347,  
"total\_size" : "495.7mb",  
"total\_size\_in\_bytes" : 519823123  
},  
"refresh" : {  
"total" : 6748,  
"total\_time" : "4.8m",  
"total\_time\_in\_millis" : 290519  
},  
"flush" : {  
"total" : 2972,  
"total\_time" : "1.6m",  
"total\_time\_in\_millis" : 101684  
}  
},  
"os" : {  
"timestamp" : 1324472103660,  
"uptime" : "-1 seconds",  
"uptime\_in\_millis" : -1000,  
"load\_average" : []  
},  
"process" : {  
"timestamp" : 1324472103660,  
"open\_file\_descriptors" : 1257  
},  
"jvm" : {  
"timestamp" : 1324472103660,  
"uptime" : "15 hours, 58 minutes, 35 seconds and 124  
milliseconds",  
"uptime\_in\_millis" : 57515124,  
"mem" : {  
"heap\_used" : "2.3gb",  
"heap\_used\_in\_bytes" : 2524820736,  
"heap\_committed" : "4.9gb",  
"heap\_committed\_in\_bytes" : 5340397568,  
"non\_heap\_used" : "46mb",  
"non\_heap\_used\_in\_bytes" : 48276152,  
"non\_heap\_committed" : "69.8mb",  
"non\_heap\_committed\_in\_bytes" : 73248768  
},  
"threads" : {  
"count" : 72,  
"peak\_count" : 86  
},  
"gc" : {  
"collection\_count" : 16609,  
"collection\_time" : "4 minutes, 25 seconds and 568  
milliseconds",  
"collection\_time\_in\_millis" : 265568,  
"collectors" : {  
"ParNew" : {  
"collection\_count" : 16600,  
"collection\_time" : "4 minutes, 24 seconds and 807  
milliseconds",  
"collection\_time\_in\_millis" : 264807  
},  
"ConcurrentMarkSweep" : {  
"collection\_count" : 9,  
"collection\_time" : "761 milliseconds",  
"collection\_time\_in\_millis" : 761  
}  
}  
}  
},  
"network" : {  
},  
"transport" : {  
"server\_open" : 14,  
"rx\_count" : 243057,  
"rx\_size" : "452.2mb",  
"rx\_size\_in\_bytes" : 474228725,  
"tx\_count" : 321262,  
"tx\_size" : "4.1gb",  
"tx\_size\_in\_bytes" : 4459909980  
},  
"http" : {  
"current\_open" : 8,  
"total\_opened" : 7208  
}  
},  
"fxJdAyuKSZeLRQ0nrBybGg" : {  
"name" : "Bantam",  
"indices" : {  
"store" : {  
"size" : "1.7gb",  
"size\_in\_bytes" : 1889182815  
},  
"docs" : {  
"count" : 557472,  
"deleted" : 114386  
},  
"indexing" : {  
"index\_total" : 18576863,  
"index\_time" : "1.1h",  
"index\_time\_in\_millis" : 4074696,  
"index\_current" : 5,  
"delete\_total" : 0,  
"delete\_time" : "0s",  
"delete\_time\_in\_millis" : 0,  
"delete\_current" : 0  
},  
"get" : {  
"total" : 91,  
"time" : "34ms",  
"time\_in\_millis" : 34,  
"exists\_total" : 43,  
"exists\_time" : "16ms",  
"exists\_time\_in\_millis" : 16,  
"missing\_total" : 48,  
"missing\_time" : "18ms",  
"missing\_time\_in\_millis" : 18,  
"current" : 0  
},  
"search" : {  
"query\_total" : 20060,  
"query\_time" : "11.4m",  
"query\_time\_in\_millis" : 686600,  
"query\_current" : 5,  
"fetch\_total" : 15146,  
"fetch\_time" : "2.4m",  
"fetch\_time\_in\_millis" : 148005,  
"fetch\_current" : 1  
},  
"cache" : {  
"field\_evictions" : 0,  
"field\_size" : "4.4mb",  
"field\_size\_in\_bytes" : 4674888,  
"filter\_count" : 2,  
"filter\_evictions" : 0,  
"filter\_size" : "92.5kb",  
"filter\_size\_in\_bytes" : 94760  
},  
"merges" : {  
"current" : 5,  
"current\_docs" : 214061,  
"current\_size" : "55.8mb",  
"current\_size\_in\_bytes" : 58566804,  
"total" : 55,  
"total\_time" : "5.6m",  
"total\_time\_in\_millis" : 341870,  
"total\_docs" : 2010928,  
"total\_size" : "527.5mb",  
"total\_size\_in\_bytes" : 553224978  
},  
"refresh" : {  
"total" : 7365,  
"total\_time" : "7.4m",  
"total\_time\_in\_millis" : 447627  
},  
"flush" : {  
"total" : 2977,  
"total\_time" : "13.4m",  
"total\_time\_in\_millis" : 804327  
}  
},  
"os" : {  
"timestamp" : 1324472120587,  
"uptime" : "-1 seconds",  
"uptime\_in\_millis" : -1000,  
"load\_average" : []  
},  
"process" : {  
"timestamp" : 1324472120587,  
"open\_file\_descriptors" : 1438  
},  
"jvm" : {  
"timestamp" : 1324472120589,  
"uptime" : "15 hours, 58 minutes, 25 seconds and 155  
milliseconds",  
"uptime\_in\_millis" : 57505155,  
"mem" : {  
"heap\_used" : "4.9gb",  
"heap\_used\_in\_bytes" : 5317828824,  
"heap\_committed" : "4.9gb",  
"heap\_committed\_in\_bytes" : 5340397568,  
"non\_heap\_used" : "44.4mb",  
"non\_heap\_used\_in\_bytes" : 46654016,  
"non\_heap\_committed" : "68.6mb",  
"non\_heap\_committed\_in\_bytes" : 71991296  
},  
"threads" : {  
"count" : 94,  
"peak\_count" : 96  
},  
"gc" : {  
"collection\_count" : 15361,  
"collection\_time" : "7 minutes, 38 seconds and 876  
milliseconds",  
"collection\_time\_in\_millis" : 458876,  
"collectors" : {  
"ParNew" : {  
"collection\_count" : 15245,  
"collection\_time" : "6 minutes, 6 seconds and 39  
milliseconds",  
"collection\_time\_in\_millis" : 366039  
},  
"ConcurrentMarkSweep" : {  
"collection\_count" : 116,  
"collection\_time" : "1 minute, 32 seconds and 837  
milliseconds",  
"collection\_time\_in\_millis" : 92837  
}  
}  
}  
},  
"network" : {  
},  
"transport" : {  
"server\_open" : 14,  
"rx\_count" : 242790,  
"rx\_size" : "4.1gb",  
"rx\_size\_in\_bytes" : 4459575446,  
"tx\_count" : 250201,  
"tx\_size" : "452mb",  
"tx\_size\_in\_bytes" : 474044111  
},  
"http" : {  
"current\_open" : 3,  
"total\_opened" : 7074  
}  
}  
}  
}

---

<div class="post-metadata">

**Author:** ![Craig\_Brown](https://avatars.discourse-cdn.com/v4/letter/c/ce7236/32.png) [@Craig\_Brown](https://discuss.elastic.co/u/Craig_Brown)\
**Post date:** [December 21, 2011, 6:47pm UTC](https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211/2 "2011-12-21T18:47:19Z")

</div>

Greg, have you checked/increased open file handle limits for your machine?  
ES/Lucene tend to require lots of file handles. I believe that more indices  
require more file handles. I had a case last night where I was trying to  
index 1m records and it hung at about 340K because it ran out of file  
handles. Once I increased the open file handle limit, I was able to index  
1m without problems.

- Craig

On Wed, Dec 21, 2011 at 8:10 AM, Greg Brown [gbrown5878@gmail.com](mailto:gbrown5878@gmail.com) wrote:

> Hi,
> 
> We're trying to index a few million responses to many (10000+) forms.  
> We're creating a separate type for each form since they can have  
> different fields. So each form is associated with a type named form-  
> .
> 
> In bulk indexing the existing data, we are creating the new types and  
> their mappings on the fly, interspersed with bulk indexing the  
> documents. After running through about 200k responses the server  
> stopped responding and appeared to be running out of memory. The  
> cluster node stats are pasted below.
> 
> This two node cluster (8 GB machines) has indices for other  
> applications (of similar size) already running successfully on it, and  
> we've can bulk index them without a problem. So I assume the large  
> number of types is causing the issue.
> 
> So some questions:
> 
> 1. I'm planning to try creating all of the types first and then bulk  
> index. Is intermixing type creation and adding docs expected to run  
> into performance problems?
> 
> 2. It seems from the cluster stats that there is a surprising amount  
> of data going in one direction. Could this be the performance problem?
> 
> 3. We're planning to do term facet searches (for word cloud  
> generation) for each of the above types. I think I read that the  
> entire index gets loaded when a term facet is done. If I do a term  
> facet on a particular type, will only the portion of the index for  
> that type be loaded, or will it still be the whole thing? If the whole  
> thing, any way I can move the data into separate indices without  
> having 10k indices show up when I go to look at the size of my  
> indices. There is no need for them to be in the same index, I just  
> don't want the 10k tables making it harder for me to examine the  
> indices for other applications.
> 
> Thanks for the help  
> -Greg
> 
> The index in question:  
> "index" : {  
> "primary\_size" : "55.8mb",  
> "primary\_size\_in\_bytes" : 58586978,  
> "size" : "111.7mb",  
> "size\_in\_bytes" : 117177113  
> },  
> "translog" : {  
> "operations" : 0  
> },  
> "docs" : {  
> "num\_docs" : 214308,  
> "max\_doc" : 214314,  
> "deleted\_docs" : 6  
> },  
> "merges" : {  
> "current" : 0,  
> "current\_docs" : 0,  
> "current\_size" : "0b",  
> "current\_size\_in\_bytes" : 0,  
> "total" : 0,  
> "total\_time" : "0s",  
> "total\_time\_in\_millis" : 0,  
> "total\_docs" : 0,  
> "total\_size" : "0b",  
> "total\_size\_in\_bytes" : 0  
> },  
> "refresh" : {  
> "total" : 6426,  
> "total\_time" : "4.8m",  
> "total\_time\_in\_millis" : 289292  
> },  
> "flush" : {  
> "total" : 2702,  
> "total\_time" : "1.6m",  
> "total\_time\_in\_millis" : 97459  
> },
> 
> /\_cluster/node/stats  
> {  
> "cluster\_name" : "mgs2",  
> "nodes" : {  
> "x25SzznQQZ2uPGQvgGjs4Q" : {  
> "name" : "Nut",  
> "indices" : {  
> "store" : {  
> "size" : "1.6gb",  
> "size\_in\_bytes" : 1806371221  
> },  
> "docs" : {  
> "count" : 557654,  
> "deleted" : 114387  
> },  
> "indexing" : {  
> "index\_total" : 18577042,  
> "index\_time" : "40m",  
> "index\_time\_in\_millis" : 2403162,  
> "index\_current" : 0,  
> "delete\_total" : 0,  
> "delete\_time" : "0s",  
> "delete\_time\_in\_millis" : 0,  
> "delete\_current" : 0  
> },  
> "get" : {  
> "total" : 86,  
> "time" : "60ms",  
> "time\_in\_millis" : 60,  
> "exists\_total" : 24,  
> "exists\_time" : "40ms",  
> "exists\_time\_in\_millis" : 40,  
> "missing\_total" : 62,  
> "missing\_time" : "20ms",  
> "missing\_time\_in\_millis" : 20,  
> "current" : 0  
> },  
> "search" : {  
> "query\_total" : 20121,  
> "query\_time" : "4.8m",  
> "query\_time\_in\_millis" : 288251,  
> "query\_current" : 0,  
> "fetch\_total" : 15337,  
> "fetch\_time" : "57.4s",  
> "fetch\_time\_in\_millis" : 57470,  
> "fetch\_current" : 0  
> },  
> "cache" : {  
> "field\_evictions" : 0,  
> "field\_size" : "4.4mb",  
> "field\_size\_in\_bytes" : 4675324,  
> "filter\_count" : 3,  
> "filter\_evictions" : 0,  
> "filter\_size" : "92.9kb",  
> "filter\_size\_in\_bytes" : 95144  
> },  
> "merges" : {  
> "current" : 0,  
> "current\_docs" : 0,  
> "current\_size" : "0b",  
> "current\_size\_in\_bytes" : 0,  
> "total" : 56,  
> "total\_time" : "15.4s",  
> "total\_time\_in\_millis" : 15492,  
> "total\_docs" : 1887347,  
> "total\_size" : "495.7mb",  
> "total\_size\_in\_bytes" : 519823123  
> },  
> "refresh" : {  
> "total" : 6748,  
> "total\_time" : "4.8m",  
> "total\_time\_in\_millis" : 290519  
> },  
> "flush" : {  
> "total" : 2972,  
> "total\_time" : "1.6m",  
> "total\_time\_in\_millis" : 101684  
> }  
> },  
> "os" : {  
> "timestamp" : 1324472103660,  
> "uptime" : "-1 seconds",  
> "uptime\_in\_millis" : -1000,  
> "load\_average" :   
> },  
> "process" : {  
> "timestamp" : 1324472103660,  
> "open\_file\_descriptors" : 1257  
> },  
> "jvm" : {  
> "timestamp" : 1324472103660,  
> "uptime" : "15 hours, 58 minutes, 35 seconds and 124  
> milliseconds",  
> "uptime\_in\_millis" : 57515124,  
> "mem" : {  
> "heap\_used" : "2.3gb",  
> "heap\_used\_in\_bytes" : 2524820736,  
> "heap\_committed" : "4.9gb",  
> "heap\_committed\_in\_bytes" : 5340397568,  
> "non\_heap\_used" : "46mb",  
> "non\_heap\_used\_in\_bytes" : 48276152,  
> "non\_heap\_committed" : "69.8mb",  
> "non\_heap\_committed\_in\_bytes" : 73248768  
> },  
> "threads" : {  
> "count" : 72,  
> "peak\_count" : 86  
> },  
> "gc" : {  
> "collection\_count" : 16609,  
> "collection\_time" : "4 minutes, 25 seconds and 568  
> milliseconds",  
> "collection\_time\_in\_millis" : 265568,  
> "collectors" : {  
> "ParNew" : {  
> "collection\_count" : 16600,  
> "collection\_time" : "4 minutes, 24 seconds and 807  
> milliseconds",  
> "collection\_time\_in\_millis" : 264807  
> },  
> "ConcurrentMarkSweep" : {  
> "collection\_count" : 9,  
> "collection\_time" : "761 milliseconds",  
> "collection\_time\_in\_millis" : 761  
> }  
> }  
> }  
> },  
> "network" : {  
> },  
> "transport" : {  
> "server\_open" : 14,  
> "rx\_count" : 243057,  
> "rx\_size" : "452.2mb",  
> "rx\_size\_in\_bytes" : 474228725,  
> "tx\_count" : 321262,  
> "tx\_size" : "4.1gb",  
> "tx\_size\_in\_bytes" : 4459909980  
> },  
> "http" : {  
> "current\_open" : 8,  
> "total\_opened" : 7208  
> }  
> },  
> "fxJdAyuKSZeLRQ0nrBybGg" : {  
> "name" : "Bantam",  
> "indices" : {  
> "store" : {  
> "size" : "1.7gb",  
> "size\_in\_bytes" : 1889182815  
> },  
> "docs" : {  
> "count" : 557472,  
> "deleted" : 114386  
> },  
> "indexing" : {  
> "index\_total" : 18576863,  
> "index\_time" : "1.1h",  
> "index\_time\_in\_millis" : 4074696,  
> "index\_current" : 5,  
> "delete\_total" : 0,  
> "delete\_time" : "0s",  
> "delete\_time\_in\_millis" : 0,  
> "delete\_current" : 0  
> },  
> "get" : {  
> "total" : 91,  
> "time" : "34ms",  
> "time\_in\_millis" : 34,  
> "exists\_total" : 43,  
> "exists\_time" : "16ms",  
> "exists\_time\_in\_millis" : 16,  
> "missing\_total" : 48,  
> "missing\_time" : "18ms",  
> "missing\_time\_in\_millis" : 18,  
> "current" : 0  
> },  
> "search" : {  
> "query\_total" : 20060,  
> "query\_time" : "11.4m",  
> "query\_time\_in\_millis" : 686600,  
> "query\_current" : 5,  
> "fetch\_total" : 15146,  
> "fetch\_time" : "2.4m",  
> "fetch\_time\_in\_millis" : 148005,  
> "fetch\_current" : 1  
> },  
> "cache" : {  
> "field\_evictions" : 0,  
> "field\_size" : "4.4mb",  
> "field\_size\_in\_bytes" : 4674888,  
> "filter\_count" : 2,  
> "filter\_evictions" : 0,  
> "filter\_size" : "92.5kb",  
> "filter\_size\_in\_bytes" : 94760  
> },  
> "merges" : {  
> "current" : 5,  
> "current\_docs" : 214061,  
> "current\_size" : "55.8mb",  
> "current\_size\_in\_bytes" : 58566804,  
> "total" : 55,  
> "total\_time" : "5.6m",  
> "total\_time\_in\_millis" : 341870,  
> "total\_docs" : 2010928,  
> "total\_size" : "527.5mb",  
> "total\_size\_in\_bytes" : 553224978  
> },  
> "refresh" : {  
> "total" : 7365,  
> "total\_time" : "7.4m",  
> "total\_time\_in\_millis" : 447627  
> },  
> "flush" : {  
> "total" : 2977,  
> "total\_time" : "13.4m",  
> "total\_time\_in\_millis" : 804327  
> }  
> },  
> "os" : {  
> "timestamp" : 1324472120587,  
> "uptime" : "-1 seconds",  
> "uptime\_in\_millis" : -1000,  
> "load\_average" :   
> },  
> "process" : {  
> "timestamp" : 1324472120587,  
> "open\_file\_descriptors" : 1438  
> },  
> "jvm" : {  
> "timestamp" : 1324472120589,  
> "uptime" : "15 hours, 58 minutes, 25 seconds and 155  
> milliseconds",  
> "uptime\_in\_millis" : 57505155,  
> "mem" : {  
> "heap\_used" : "4.9gb",  
> "heap\_used\_in\_bytes" : 5317828824,  
> "heap\_committed" : "4.9gb",  
> "heap\_committed\_in\_bytes" : 5340397568,  
> "non\_heap\_used" : "44.4mb",  
> "non\_heap\_used\_in\_bytes" : 46654016,  
> "non\_heap\_committed" : "68.6mb",  
> "non\_heap\_committed\_in\_bytes" : 71991296  
> },  
> "threads" : {  
> "count" : 94,  
> "peak\_count" : 96  
> },  
> "gc" : {  
> "collection\_count" : 15361,  
> "collection\_time" : "7 minutes, 38 seconds and 876  
> milliseconds",  
> "collection\_time\_in\_millis" : 458876,  
> "collectors" : {  
> "ParNew" : {  
> "collection\_count" : 15245,  
> "collection\_time" : "6 minutes, 6 seconds and 39  
> milliseconds",  
> "collection\_time\_in\_millis" : 366039  
> },  
> "ConcurrentMarkSweep" : {  
> "collection\_count" : 116,  
> "collection\_time" : "1 minute, 32 seconds and 837  
> milliseconds",  
> "collection\_time\_in\_millis" : 92837  
> }  
> }  
> }  
> },  
> "network" : {  
> },  
> "transport" : {  
> "server\_open" : 14,  
> "rx\_count" : 242790,  
> "rx\_size" : "4.1gb",  
> "rx\_size\_in\_bytes" : 4459575446,  
> "tx\_count" : 250201,  
> "tx\_size" : "452mb",  
> "tx\_size\_in\_bytes" : 474044111  
> },  
> "http" : {  
> "current\_open" : 3,  
> "total\_opened" : 7074  
> }  
> }  
> }  
> }

--  
…  
CRAIG BROWN  
chief architect  
youwho, Inc.

_[www.youwho.com](http://www.youwho.com)_ [http://www.youwho.com/](http://www.youwho.com/)

T: 801.855. 0921  
M: 801.913. 0939

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [December 21, 2011, 7:08pm UTC](https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211/3 "2011-12-21T19:08:19Z")

</div>

> Greg, have you checked/increased open file handle limits for your machine?

First, check/post your logs. If too many files open ES would log that.

Peter.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [December 22, 2011, 12:54am UTC](https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211/4 "2011-12-22T00:54:16Z")

</div>

My guess is that the problem is with creating so many types, which ends up  
being a large overhead in the system. Each time a type is introduced, it  
needs to be broadcasted to the rest of the nodes and persisted as part of  
the cluster meta data. Can you try just indexing into the same type as a  
test and see if it still happens?

On Wed, Dec 21, 2011 at 9:08 PM, Karussell [tableyourtime@googlemail.com](mailto:tableyourtime@googlemail.com)wrote:

> > Greg, have you checked/increased open file handle limits for your  
> > machine?
> 
> First, check/post your logs. If too many files open ES would log that.
> 
> Peter.

---

<div class="post-metadata">

**Author:** ![Greg\_Brown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greg_brown/32/1373_2.png) [@Greg\_Brown](https://discuss.elastic.co/u/Greg_Brown)\
**Post date:** [December 22, 2011, 2:37pm UTC](https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211/5 "2011-12-22T14:37:45Z")

</div>

Checking through the logs, there isn't any mention of there not being  
enough file handles, the errors I am running into are out of memory on  
the heap space errors.

Shay,

Thanks, will give that a try and let you know. Will have to wait until  
after the weekend so I can set up a development cluster. I've brought  
down the production cluster a few too many times this week, and its  
time to be more careful. 🙂

Thanks for the fast responses.  
-Greg

On Dec 21, 6:54 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> My guess is that the problem is with creating so many types, which ends up  
> being a large overhead in the system. Each time a type is introduced, it  
> needs to be broadcasted to the rest of the nodes and persisted as part of  
> the cluster meta data. Can you try just indexing into the same type as a  
> test and see if it still happens?
> 
> On Wed, Dec 21, 2011 at 9:08 PM, Karussell [tableyourt...@googlemail.com](mailto:tableyourt...@googlemail.com)wrote:
> 
> > > Greg, have you checked/increased open file handle limits for your  
> > > machine?
> 
> > First, check/post your logs. If too many files open ES would log that.
> 
> > Peter.

---

<div class="post-metadata">

**Author:** ![Greg\_Brown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greg_brown/32/1373_2.png) [@Greg\_Brown](https://discuss.elastic.co/u/Greg_Brown)\
**Post date:** [January 10, 2012, 3:45am UTC](https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211/6 "2012-01-10T03:45:08Z")

</div>

Indexing all data to a single type did work fine (3.3mil docs) as  
expected.

I submitted a bug ([elasticsearch · GitHub](https://github.com/elasticsearch/)  
[elasticsearch.github.com/issues/134](http://elasticsearch.github.com/issues/134)) on the large number of types  
because I was able to get the server to become unresponsive even when  
there was only a single server and I tried to add many types.

For the moment I am going ahead with using all of the documents in a  
single index. However, this significantly reduces query performance  
compared to having a separate type for each set of documents. I looped  
and profiled the following queries on the larger sets of documents  
(10k-70k): [Profiling queries for building word clouds · GitHub](https://gist.github.com/1586723) This was all run on a  
single server, and the query from a different machine.

The first query has each set of docs in its own Type. On average it  
took about 11 ms to complete.

The second has all of the docs in one index with a field pd\_id to  
distinguish the sets. The query uses the facet\_filter and average ~190  
ms.

The third uses the same index as the second, but uses a query to do  
the "filtering" of the docs. ~140 ms. I was surprised that this was  
faster than the facet\_filter.

Any suggestions on how to improve the last two queries?

Any ideas on how to create multiple types without creating 10k  
separate indices. In this case all I am using the Type for is a  
partitioning/grouping of multiple separate indices, since the Mapping  
of each Type is identical.

Thanks for the help.  
-Greg

On Dec 22 2011, 7:37 am, Greg Brown [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com) wrote:

> Checking through the logs, there isn't any mention of there not being  
> enough file handles, the errors I am running into are out of memory on  
> the heap space errors.
> 
> Shay,
> 
> Thanks, will give that a try and let you know. Will have to wait until  
> after the weekend so I can set up a development cluster. I've brought  
> down the production cluster a few too many times this week, and its  
> time to be more careful. 🙂
> 
> Thanks for the fast responses.  
> -Greg
> 
> On Dec 21, 6:54 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > My guess is that the problem is with creating so many types, which ends up  
> > being a large overhead in the system. Each time a type is introduced, it  
> > needs to be broadcasted to the rest of the nodes and persisted as part of  
> > the cluster meta data. Can you try just indexing into the same type as a  
> > test and see if it still happens?
> 
> > On Wed, Dec 21, 2011 at 9:08 PM, Karussell [tableyourt...@googlemail.com](mailto:tableyourt...@googlemail.com)wrote:
> 
> > > > Greg, have you checked/increased open file handle limits for your  
> > > > machine?
> 
> > > First, check/post your logs. If too many files open ES would log that.
> 
> > > Peter.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [January 10, 2012, 9:40am UTC](https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211/7 "2012-01-10T09:40:41Z")

</div>

Just add the "type" as a field to the doc, and filter by it.

On Tue, Jan 10, 2012 at 5:45 AM, Greg Brown [gbrown5878@gmail.com](mailto:gbrown5878@gmail.com) wrote:

> Indexing all data to a single type did work fine (3.3mil docs) as  
> expected.
> 
> I submitted a bug ([elasticsearch · GitHub](https://github.com/elasticsearch/)  
> [elasticsearch.github.com/issues/134](http://elasticsearch.github.com/issues/134)) on the large number of types  
> because I was able to get the server to become unresponsive even when  
> there was only a single server and I tried to add many types.
> 
> For the moment I am going ahead with using all of the documents in a  
> single index. However, this significantly reduces query performance  
> compared to having a separate type for each set of documents. I looped  
> and profiled the following queries on the larger sets of documents  
> (10k-70k): [Profiling queries for building word clouds · GitHub](https://gist.github.com/1586723) This was all run on a  
> single server, and the query from a different machine.
> 
> The first query has each set of docs in its own Type. On average it  
> took about 11 ms to complete.
> 
> The second has all of the docs in one index with a field pd\_id to  
> distinguish the sets. The query uses the facet\_filter and average ~190  
> ms.
> 
> The third uses the same index as the second, but uses a query to do  
> the "filtering" of the docs. ~140 ms. I was surprised that this was  
> faster than the facet\_filter.
> 
> Any suggestions on how to improve the last two queries?
> 
> Any ideas on how to create multiple types without creating 10k  
> separate indices. In this case all I am using the Type for is a  
> partitioning/grouping of multiple separate indices, since the Mapping  
> of each Type is identical.
> 
> Thanks for the help.  
> -Greg
> 
> On Dec 22 2011, 7:37 am, Greg Brown [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com) wrote:
> 
> > Checking through the logs, there isn't any mention of there not being  
> > enough file handles, the errors I am running into are out of memory on  
> > the heap space errors.
> > 
> > Shay,
> > 
> > Thanks, will give that a try and let you know. Will have to wait until  
> > after the weekend so I can set up a development cluster. I've brought  
> > down the production cluster a few too many times this week, and its  
> > time to be more careful. 🙂
> > 
> > Thanks for the fast responses.  
> > -Greg
> > 
> > On Dec 21, 6:54 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> > 
> > > My guess is that the problem is with creating so many types, which  
> > > ends up  
> > > being a large overhead in the system. Each time a type is introduced,  
> > > it  
> > > needs to be broadcasted to the rest of the nodes and persisted as part  
> > > of  
> > > the cluster meta data. Can you try just indexing into the same type as  
> > > a  
> > > test and see if it still happens?
> > 
> > > On Wed, Dec 21, 2011 at 9:08 PM, Karussell \<  
> > > [tableyourt...@googlemail.com](mailto:tableyourt...@googlemail.com)\>wrote:
> > 
> > > > > Greg, have you checked/increased open file handle limits for your  
> > > > > machine?
> > 
> > > > First, check/post your logs. If too many files open ES would log  
> > > > that.
> > 
> > > > Peter.

---

<div class="post-metadata">

**Author:** ![Greg\_Brown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greg_brown/32/1373_2.png) [@Greg\_Brown](https://discuss.elastic.co/u/Greg_Brown)\
**Post date:** [January 10, 2012, 3:48pm UTC](https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211/8 "2012-01-10T15:48:59Z")

</div>

That's what I did. Functionally works, but it is 10x slower to query  
using either a query to filter or a facet\_filter. Is there another  
way? According to the docs: "search filters restrict only returned  
documents — but not facet counts"

On Jan 10, 2:40 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> Just add the "type" as a field to the doc, and filter by it.
> 
> On Tue, Jan 10, 2012 at 5:45 AM, Greg Brown [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com) wrote:
> 
> > Indexing all data to a single type did work fine (3.3mil docs) as  
> > expected.
> 
> > I submitted a bug ([elasticsearch · GitHub](https://github.com/elasticsearch/)  
> > [elasticsearch.github.com/issues/134](http://elasticsearch.github.com/issues/134)) on the large number of types  
> > because I was able to get the server to become unresponsive even when  
> > there was only a single server and I tried to add many types.
> 
> > For the moment I am going ahead with using all of the documents in a  
> > single index. However, this significantly reduces query performance  
> > compared to having a separate type for each set of documents. I looped  
> > and profiled the following queries on the larger sets of documents  
> > (10k-70k):[https://gist.github.com/1586723This](https://gist.github.com/1586723This) was all run on a  
> > single server, and the query from a different machine.
> 
> > The first query has each set of docs in its own Type. On average it  
> > took about 11 ms to complete.
> 
> > The second has all of the docs in one index with a field pd\_id to  
> > distinguish the sets. The query uses the facet\_filter and average ~190  
> > ms.
> 
> > The third uses the same index as the second, but uses a query to do  
> > the "filtering" of the docs. ~140 ms. I was surprised that this was  
> > faster than the facet\_filter.
> 
> > Any suggestions on how to improve the last two queries?
> 
> > Any ideas on how to create multiple types without creating 10k  
> > separate indices. In this case all I am using the Type for is a  
> > partitioning/grouping of multiple separate indices, since the Mapping  
> > of each Type is identical.
> 
> > Thanks for the help.  
> > -Greg
> 
> > On Dec 22 2011, 7:37 am, Greg Brown [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com) wrote:
> > 
> > > Checking through the logs, there isn't any mention of there not being  
> > > enough file handles, the errors I am running into are out of memory on  
> > > the heap space errors.
> 
> > > Shay,
> 
> > > Thanks, will give that a try and let you know. Will have to wait until  
> > > after the weekend so I can set up a development cluster. I've brought  
> > > down the production cluster a few too many times this week, and its  
> > > time to be more careful. 🙂
> 
> > > Thanks for the fast responses.  
> > > -Greg
> 
> > > On Dec 21, 6:54 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > > > My guess is that the problem is with creating so many types, which  
> > > > ends up  
> > > > being a large overhead in the system. Each time a type is introduced,  
> > > > it  
> > > > needs to be broadcasted to the rest of the nodes and persisted as part  
> > > > of  
> > > > the cluster meta data. Can you try just indexing into the same type as  
> > > > a  
> > > > test and see if it still happens?
> 
> > > > On Wed, Dec 21, 2011 at 9:08 PM, Karussell \<  
> > > > [tableyourt...@googlemail.com](mailto:tableyourt...@googlemail.com)\>wrote:
> 
> > > > > > Greg, have you checked/increased open file handle limits for your  
> > > > > > machine?
> 
> > > > > First, check/post your logs. If too many files open ES would log  
> > > > > that.
> 
> > > > > Peter.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [January 10, 2012, 5:50pm UTC](https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211/9 "2012-01-10T17:50:18Z")

</div>

10x slower than types? It makes little sense since types, at teh end of the  
day, is just a field called \_type in a document, and when you search within  
a type, your query provided is simply wrapped in a filtered query a filter  
on the type. So, you can do it yourself, just wrap your query in a filtered  
query with a filter on your "type".

On Tue, Jan 10, 2012 at 5:48 PM, Greg Ichneumon Brown  
[gbrown5878@gmail.com](mailto:gbrown5878@gmail.com)wrote:

> That's what I did. Functionally works, but it is 10x slower to query  
> using either a query to filter or a facet\_filter. Is there another  
> way? According to the docs: "search filters restrict only returned  
> documents — but not facet counts"
> 
> On Jan 10, 2:40 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > Just add the "type" as a field to the doc, and filter by it.
> > 
> > On Tue, Jan 10, 2012 at 5:45 AM, Greg Brown [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com)  
> > wrote:
> > 
> > > Indexing all data to a single type did work fine (3.3mil docs) as  
> > > expected.
> > 
> > > I submitted a bug ([elasticsearch · GitHub](https://github.com/elasticsearch/)  
> > > [elasticsearch.github.com/issues/134](http://elasticsearch.github.com/issues/134)) on the large number of types  
> > > because I was able to get the server to become unresponsive even when  
> > > there was only a single server and I tried to add many types.
> > 
> > > For the moment I am going ahead with using all of the documents in a  
> > > single index. However, this significantly reduces query performance  
> > > compared to having a separate type for each set of documents. I looped  
> > > and profiled the following queries on the larger sets of documents  
> > > (10k-70k):[https://gist.github.com/1586723This](https://gist.github.com/1586723This) was all run on a  
> > > single server, and the query from a different machine.
> > 
> > > The first query has each set of docs in its own Type. On average it  
> > > took about 11 ms to complete.
> > 
> > > The second has all of the docs in one index with a field pd\_id to  
> > > distinguish the sets. The query uses the facet\_filter and average ~190  
> > > ms.
> > 
> > > The third uses the same index as the second, but uses a query to do  
> > > the "filtering" of the docs. ~140 ms. I was surprised that this was  
> > > faster than the facet\_filter.
> > 
> > > Any suggestions on how to improve the last two queries?
> > 
> > > Any ideas on how to create multiple types without creating 10k  
> > > separate indices. In this case all I am using the Type for is a  
> > > partitioning/grouping of multiple separate indices, since the Mapping  
> > > of each Type is identical.
> > 
> > > Thanks for the help.  
> > > -Greg
> > 
> > > On Dec 22 2011, 7:37 am, Greg Brown [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com) wrote:
> > > 
> > > > Checking through the logs, there isn't any mention of there not being  
> > > > enough file handles, the errors I am running into are out of memory  
> > > > on  
> > > > the heap space errors.
> > 
> > > > Shay,
> > 
> > > > Thanks, will give that a try and let you know. Will have to wait  
> > > > until  
> > > > after the weekend so I can set up a development cluster. I've brought  
> > > > down the production cluster a few too many times this week, and its  
> > > > time to be more careful. 🙂
> > 
> > > > Thanks for the fast responses.  
> > > > -Greg
> > 
> > > > On Dec 21, 6:54 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> > 
> > > > > My guess is that the problem is with creating so many types, which  
> > > > > ends up  
> > > > > being a large overhead in the system. Each time a type is  
> > > > > introduced,  
> > > > > it  
> > > > > needs to be broadcasted to the rest of the nodes and persisted as  
> > > > > part  
> > > > > of  
> > > > > the cluster meta data. Can you try just indexing into the same  
> > > > > type as  
> > > > > a  
> > > > > test and see if it still happens?
> > 
> > > > > On Wed, Dec 21, 2011 at 9:08 PM, Karussell \<  
> > > > > [tableyourt...@googlemail.com](mailto:tableyourt...@googlemail.com)\>wrote:
> > 
> > > > > > > Greg, have you checked/increased open file handle limits for  
> > > > > > > your  
> > > > > > > machine?
> > 
> > > > > > First, check/post your logs. If too many files open ES would log  
> > > > > > that.
> > 
> > > > > > Peter.

---

<div class="post-metadata">

**Author:** ![Greg\_Brown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greg_brown/32/1373_2.png) [@Greg\_Brown](https://discuss.elastic.co/u/Greg_Brown)\
**Post date:** [January 11, 2012, 4:35pm UTC](https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211/10 "2012-01-11T16:35:26Z")

</div>

Ah, I see what you are saying. But I am totally flummoxed as to how to  
formulate the query to get a filtered query that matches the  
performance of the default filtered query used for the type.

I think what you are suggesting is:  
curl -XGET "${SERVER}/pd-test-0/\_search?pretty"   
-d '{  
"size" : 0,  
"query" : {  
"filtered" : {  
"query" : { "match\_all" : { } },  
"filter" : {  
"term" : { "pd\_id" : "$ID" }  
}  
} },  
"facets" : {  
"q1" : {  
"terms" : {  
"field" : "q1",  
"size" : 100  
}  
}  
}  
}'

Is that query match\_all correct? This query takes about 125 ms vs 13  
ms for a query on the type (curl -XGET "${SERVER}/pd-0/${ID}/\_search).

From what I can gather from the Java code this should mostly match  
that query, but don't know the code well enough. Is there an easy way  
to enable logging that would let me compare the structure of the  
parsed queries for debugging this?

Thanks for all the help, Shay! Much appreciated.  
-Greg

On Jan 10, 10:50 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> 10x slower than types? It makes little sense since types, at teh end of the  
> day, is just a field called \_type in a document, and when you search within  
> a type, your query provided is simply wrapped in a filtered query a filter  
> on the type. So, you can do it yourself, just wrap your query in a filtered  
> query with a filter on your "type".
> 
> On Tue, Jan 10, 2012 at 5:48 PM, Greg Ichneumon Brown  
> [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com)wrote:
> 
> > That's what I did. Functionally works, but it is 10x slower to query  
> > using either a query to filter or a facet\_filter. Is there another  
> > way? According to the docs: "search filters restrict only returned  
> > documents — but not facet counts"
> 
> > On Jan 10, 2:40 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> > 
> > > Just add the "type" as a field to the doc, and filter by it.
> 
> > > On Tue, Jan 10, 2012 at 5:45 AM, Greg Brown [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com)  
> > > wrote:
> > > 
> > > > Indexing all data to a single type did work fine (3.3mil docs) as  
> > > > expected.
> 
> > > > I submitted a bug ([elasticsearch · GitHub](https://github.com/elasticsearch/)  
> > > > [elasticsearch.github.com/issues/134](http://elasticsearch.github.com/issues/134)) on the large number of types  
> > > > because I was able to get the server to become unresponsive even when  
> > > > there was only a single server and I tried to add many types.
> 
> > > > For the moment I am going ahead with using all of the documents in a  
> > > > single index. However, this significantly reduces query performance  
> > > > compared to having a separate type for each set of documents. I looped  
> > > > and profiled the following queries on the larger sets of documents  
> > > > (10k-70k):[https://gist.github.com/1586723Thiswas](https://gist.github.com/1586723Thiswas) all run on a  
> > > > single server, and the query from a different machine.
> 
> > > > The first query has each set of docs in its own Type. On average it  
> > > > took about 11 ms to complete.
> 
> > > > The second has all of the docs in one index with a field pd\_id to  
> > > > distinguish the sets. The query uses the facet\_filter and average ~190  
> > > > ms.
> 
> > > > The third uses the same index as the second, but uses a query to do  
> > > > the "filtering" of the docs. ~140 ms. I was surprised that this was  
> > > > faster than the facet\_filter.
> 
> > > > Any suggestions on how to improve the last two queries?
> 
> > > > Any ideas on how to create multiple types without creating 10k  
> > > > separate indices. In this case all I am using the Type for is a  
> > > > partitioning/grouping of multiple separate indices, since the Mapping  
> > > > of each Type is identical.
> 
> > > > Thanks for the help.  
> > > > -Greg
> 
> > > > On Dec 22 2011, 7:37 am, Greg Brown [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com) wrote:
> > > > 
> > > > > Checking through the logs, there isn't any mention of there not being  
> > > > > enough file handles, the errors I am running into are out of memory  
> > > > > on  
> > > > > the heap space errors.
> 
> > > > > Shay,
> 
> > > > > Thanks, will give that a try and let you know. Will have to wait  
> > > > > until  
> > > > > after the weekend so I can set up a development cluster. I've brought  
> > > > > down the production cluster a few too many times this week, and its  
> > > > > time to be more careful. 🙂
> 
> > > > > Thanks for the fast responses.  
> > > > > -Greg
> 
> > > > > On Dec 21, 6:54 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > > > > > My guess is that the problem is with creating so many types, which  
> > > > > > ends up  
> > > > > > being a large overhead in the system. Each time a type is  
> > > > > > introduced,  
> > > > > > it  
> > > > > > needs to be broadcasted to the rest of the nodes and persisted as  
> > > > > > part  
> > > > > > of  
> > > > > > the cluster meta data. Can you try just indexing into the same  
> > > > > > type as  
> > > > > > a  
> > > > > > test and see if it still happens?
> 
> > > > > > On Wed, Dec 21, 2011 at 9:08 PM, Karussell \<  
> > > > > > [tableyourt...@googlemail.com](mailto:tableyourt...@googlemail.com)\>wrote:
> 
> > > > > > > > Greg, have you checked/increased open file handle limits for  
> > > > > > > > your  
> > > > > > > > machine?
> 
> > > > > > > First, check/post your logs. If too many files open ES would log  
> > > > > > > that.
> 
> > > > > > > Peter.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [January 11, 2012, 6:00pm UTC](https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211/11 "2012-01-11T18:00:13Z")

</div>

Yes, what you posted is exactly the query that you will end up with when  
searching against a type. Maybe you just didn't let it do its caching bit?  
How fast is the 2-3rd execution using the same pid (the term filter result  
is cached).

On Wed, Jan 11, 2012 at 6:35 PM, Greg Ichneumon Brown  
[gbrown5878@gmail.com](mailto:gbrown5878@gmail.com)wrote:

> Ah, I see what you are saying. But I am totally flummoxed as to how to  
> formulate the query to get a filtered query that matches the  
> performance of the default filtered query used for the type.
> 
> I think what you are suggesting is:  
> curl -XGET "${SERVER}/pd-test-0/\_search?pretty"   
> -d '{  
> "size" : 0,  
> "query" : {  
> "filtered" : {  
> "query" : { "match\_all" : { } },  
> "filter" : {  
> "term" : { "pd\_id" : "$ID" }  
> }  
> } },  
> "facets" : {  
> "q1" : {  
> "terms" : {  
> "field" : "q1",  
> "size" : 100  
> }  
> }  
> }  
> }'
> 
> Is that query match\_all correct? This query takes about 125 ms vs 13  
> ms for a query on the type (curl -XGET "${SERVER}/pd-0/${ID}/\_search).
> 
> From what I can gather from the Java code this should mostly match  
> that query, but don't know the code well enough. Is there an easy way  
> to enable logging that would let me compare the structure of the  
> parsed queries for debugging this?
> 
> Thanks for all the help, Shay! Much appreciated.  
> -Greg
> 
> On Jan 10, 10:50 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > 10x slower than types? It makes little sense since types, at teh end of  
> > the  
> > day, is just a field called \_type in a document, and when you search  
> > within  
> > a type, your query provided is simply wrapped in a filtered query a  
> > filter  
> > on the type. So, you can do it yourself, just wrap your query in a  
> > filtered  
> > query with a filter on your "type".
> > 
> > On Tue, Jan 10, 2012 at 5:48 PM, Greg Ichneumon Brown  
> > [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com)wrote:
> > 
> > > That's what I did. Functionally works, but it is 10x slower to query  
> > > using either a query to filter or a facet\_filter. Is there another  
> > > way? According to the docs: "search filters restrict only returned  
> > > documents — but not facet counts"
> > 
> > > On Jan 10, 2:40 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> > > 
> > > > Just add the "type" as a field to the doc, and filter by it.
> > 
> > > > On Tue, Jan 10, 2012 at 5:45 AM, Greg Brown [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com)  
> > > > wrote:
> > > > 
> > > > > Indexing all data to a single type did work fine (3.3mil docs) as  
> > > > > expected.
> > 
> > > > > I submitted a bug ([elasticsearch · GitHub](https://github.com/elasticsearch/)  
> > > > > [elasticsearch.github.com/issues/134](http://elasticsearch.github.com/issues/134)) on the large number of types  
> > > > > because I was able to get the server to become unresponsive even  
> > > > > when  
> > > > > there was only a single server and I tried to add many types.
> > 
> > > > > For the moment I am going ahead with using all of the documents in  
> > > > > a  
> > > > > single index. However, this significantly reduces query performance  
> > > > > compared to having a separate type for each set of documents. I  
> > > > > looped  
> > > > > and profiled the following queries on the larger sets of documents  
> > > > > (10k-70k):[https://gist.github.com/1586723Thiswas](https://gist.github.com/1586723Thiswas) all run on a  
> > > > > single server, and the query from a different machine.
> > 
> > > > > The first query has each set of docs in its own Type. On average it  
> > > > > took about 11 ms to complete.
> > 
> > > > > The second has all of the docs in one index with a field pd\_id to  
> > > > > distinguish the sets. The query uses the facet\_filter and average  
> > > > > ~190  
> > > > > ms.
> > 
> > > > > The third uses the same index as the second, but uses a query to do  
> > > > > the "filtering" of the docs. ~140 ms. I was surprised that this was  
> > > > > faster than the facet\_filter.
> > 
> > > > > Any suggestions on how to improve the last two queries?
> > 
> > > > > Any ideas on how to create multiple types without creating 10k  
> > > > > separate indices. In this case all I am using the Type for is a  
> > > > > partitioning/grouping of multiple separate indices, since the  
> > > > > Mapping  
> > > > > of each Type is identical.
> > 
> > > > > Thanks for the help.  
> > > > > -Greg
> > 
> > > > > On Dec 22 2011, 7:37 am, Greg Brown [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com) wrote:
> > > > > 
> > > > > > Checking through the logs, there isn't any mention of there not  
> > > > > > being  
> > > > > > enough file handles, the errors I am running into are out of  
> > > > > > memory  
> > > > > > on  
> > > > > > the heap space errors.
> > 
> > > > > > Shay,
> > 
> > > > > > Thanks, will give that a try and let you know. Will have to wait  
> > > > > > until  
> > > > > > after the weekend so I can set up a development cluster. I've  
> > > > > > brought  
> > > > > > down the production cluster a few too many times this week, and  
> > > > > > its  
> > > > > > time to be more careful. 🙂
> > 
> > > > > > Thanks for the fast responses.  
> > > > > > -Greg
> > 
> > > > > > On Dec 21, 6:54 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> > 
> > > > > > > My guess is that the problem is with creating so many types,  
> > > > > > > which  
> > > > > > > ends up  
> > > > > > > being a large overhead in the system. Each time a type is  
> > > > > > > introduced,  
> > > > > > > it  
> > > > > > > needs to be broadcasted to the rest of the nodes and persisted  
> > > > > > > as  
> > > > > > > part  
> > > > > > > of  
> > > > > > > the cluster meta data. Can you try just indexing into the same  
> > > > > > > type as  
> > > > > > > a  
> > > > > > > test and see if it still happens?
> > 
> > > > > > > On Wed, Dec 21, 2011 at 9:08 PM, Karussell \<  
> > > > > > > [tableyourt...@googlemail.com](mailto:tableyourt...@googlemail.com)\>wrote:
> > 
> > > > > > > > > Greg, have you checked/increased open file handle limits for  
> > > > > > > > > your  
> > > > > > > > > machine?
> > 
> > > > > > > > First, check/post your logs. If too many files open ES would  
> > > > > > > > log  
> > > > > > > > that.
> > 
> > > > > > > > Peter.

---

<div class="post-metadata">

**Author:** ![Greg\_Brown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greg_brown/32/1373_2.png) [@Greg\_Brown](https://discuss.elastic.co/u/Greg_Brown)\
**Post date:** [January 11, 2012, 8:14pm UTC](https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211/12 "2012-01-11T20:14:19Z")

</div>

No matter how many times I repeat that query I am getting "took" :  
125, whereas the type query gives "took" : 13. I also built a filtered  
query using \_type and that completes in 13 ms also.

Could my mapping for this field be the culprit? I am doing:

```
	'pd_id' => array( 'type' => 'long', 'store' => 'yes', 'index' =>

```

'not\_analyzed' )

Each document has an integer as an id, I used long as the storage type  
to reduce memory, but does this not work with the term filter?

I tried setting \_cache to true in the filter to see if that forced the  
caching, but then I get the correct number of total hits, but no facet  
results:  
{  
"took" : 1,  
"timed\_out" : false,  
"\_shards" : {  
"total" : 5,  
"successful" : 5,  
"failed" : 0  
},  
"hits" : {  
"total" : 11588,  
"max\_score" : 1.0,  
"hits" :   
}  
}

Should be:  
{  
"took" : 13,  
"timed\_out" : false,  
"\_shards" : {  
"total" : 5,  
"successful" : 5,  
"failed" : 0  
},  
"hits" : {  
"total" : 11588,  
"max\_score" : 1.0,  
"hits" :   
},  
"facets" : {  
"q1" : {  
"\_type" : "terms",  
"missing" : 0,  
"total" : 11720,  
"other" : 56,  
"terms" : [ {  
"term" : "adopt",  
"count" : 11475  
}, {  
"term" : "adoption",  
"count" : 39  
}, {  
etc...

On Jan 11, 11:00 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> Yes, what you posted is exactly the query that you will end up with when  
> searching against a type. Maybe you just didn't let it do its caching bit?  
> How fast is the 2-3rd execution using the same pid (the term filter result  
> is cached).
> 
> On Wed, Jan 11, 2012 at 6:35 PM, Greg Ichneumon Brown  
> [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com)wrote:
> 
> > Ah, I see what you are saying. But I am totally flummoxed as to how to  
> > formulate the query to get a filtered query that matches the  
> > performance of the default filtered query used for the type.
> 
> > I think what you are suggesting is:  
> > curl -XGET "${SERVER}/pd-test-0/\_search?pretty"   
> > -d '{  
> > "size" : 0,  
> > "query" : {  
> > "filtered" : {  
> > "query" : { "match\_all" : { } },  
> > "filter" : {  
> > "term" : { "pd\_id" : "$ID" }  
> > }  
> > } },  
> > "facets" : {  
> > "q1" : {  
> > "terms" : {  
> > "field" : "q1",  
> > "size" : 100  
> > }  
> > }  
> > }  
> > }'
> 
> > Is that query match\_all correct? This query takes about 125 ms vs 13  
> > ms for a query on the type (curl -XGET "${SERVER}/pd-0/${ID}/\_search).
> 
> > From what I can gather from the Java code this should mostly match  
> > that query, but don't know the code well enough. Is there an easy way  
> > to enable logging that would let me compare the structure of the  
> > parsed queries for debugging this?
> 
> > Thanks for all the help, Shay! Much appreciated.  
> > -Greg
> 
> > On Jan 10, 10:50 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> > 
> > > 10x slower than types? It makes little sense since types, at teh end of  
> > > the  
> > > day, is just a field called \_type in a document, and when you search  
> > > within  
> > > a type, your query provided is simply wrapped in a filtered query a  
> > > filter  
> > > on the type. So, you can do it yourself, just wrap your query in a  
> > > filtered  
> > > query with a filter on your "type".
> 
> > > On Tue, Jan 10, 2012 at 5:48 PM, Greg Ichneumon Brown  
> > > [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com)wrote:
> 
> > > > That's what I did. Functionally works, but it is 10x slower to query  
> > > > using either a query to filter or a facet\_filter. Is there another  
> > > > way? According to the docs: "search filters restrict only returned  
> > > > documents — but not facet counts"
> 
> > > > On Jan 10, 2:40 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> > > > 
> > > > > Just add the "type" as a field to the doc, and filter by it.
> 
> > > > > On Tue, Jan 10, 2012 at 5:45 AM, Greg Brown [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com)  
> > > > > wrote:
> > > > > 
> > > > > > Indexing all data to a single type did work fine (3.3mil docs) as  
> > > > > > expected.
> 
> > > > > > I submitted a bug ([elasticsearch · GitHub](https://github.com/elasticsearch/)  
> > > > > > [elasticsearch.github.com/issues/134](http://elasticsearch.github.com/issues/134)) on the large number of types  
> > > > > > because I was able to get the server to become unresponsive even  
> > > > > > when  
> > > > > > there was only a single server and I tried to add many types.
> 
> > > > > > For the moment I am going ahead with using all of the documents in  
> > > > > > a  
> > > > > > single index. However, this significantly reduces query performance  
> > > > > > compared to having a separate type for each set of documents. I  
> > > > > > looped  
> > > > > > and profiled the following queries on the larger sets of documents  
> > > > > > (10k-70k):[https://gist.github.com/1586723Thiswasall](https://gist.github.com/1586723Thiswasall) run on a  
> > > > > > single server, and the query from a different machine.
> 
> > > > > > The first query has each set of docs in its own Type. On average it  
> > > > > > took about 11 ms to complete.
> 
> > > > > > The second has all of the docs in one index with a field pd\_id to  
> > > > > > distinguish the sets. The query uses the facet\_filter and average  
> > > > > > ~190  
> > > > > > ms.
> 
> > > > > > The third uses the same index as the second, but uses a query to do  
> > > > > > the "filtering" of the docs. ~140 ms. I was surprised that this was  
> > > > > > faster than the facet\_filter.
> 
> > > > > > Any suggestions on how to improve the last two queries?
> 
> > > > > > Any ideas on how to create multiple types without creating 10k  
> > > > > > separate indices. In this case all I am using the Type for is a  
> > > > > > partitioning/grouping of multiple separate indices, since the  
> > > > > > Mapping  
> > > > > > of each Type is identical.
> 
> > > > > > Thanks for the help.  
> > > > > > -Greg
> 
> > > > > > On Dec 22 2011, 7:37 am, Greg Brown [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com) wrote:
> > > > > > 
> > > > > > > Checking through the logs, there isn't any mention of there not  
> > > > > > > being  
> > > > > > > enough file handles, the errors I am running into are out of  
> > > > > > > memory  
> > > > > > > on  
> > > > > > > the heap space errors.
> 
> > > > > > > Shay,
> 
> > > > > > > Thanks, will give that a try and let you know. Will have to wait  
> > > > > > > until  
> > > > > > > after the weekend so I can set up a development cluster. I've  
> > > > > > > brought  
> > > > > > > down the production cluster a few too many times this week, and  
> > > > > > > its  
> > > > > > > time to be more careful. 🙂
> 
> > > > > > > Thanks for the fast responses.  
> > > > > > > -Greg
> 
> > > > > > > On Dec 21, 6:54 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > > > > > > > My guess is that the problem is with creating so many types,  
> > > > > > > > which  
> > > > > > > > ends up  
> > > > > > > > being a large overhead in the system. Each time a type is  
> > > > > > > > introduced,  
> > > > > > > > it  
> > > > > > > > needs to be broadcasted to the rest of the nodes and persisted  
> > > > > > > > as  
> > > > > > > > part  
> > > > > > > > of  
> > > > > > > > the cluster meta data. Can you try just indexing into the same  
> > > > > > > > type as  
> > > > > > > > a  
> > > > > > > > test and see if it still happens?
> 
> > > > > > > > On Wed, Dec 21, 2011 at 9:08 PM, Karussell \<  
> > > > > > > > [tableyourt...@googlemail.com](mailto:tableyourt...@googlemail.com)\>wrote:
> 
> > > > > > > > > > Greg, have you checked/increased open file handle limits for  
> > > > > > > > > > your  
> > > > > > > > > > machine?
> 
> > > > > > > > > First, check/post your logs. If too many files open ES would  
> > > > > > > > > log  
> > > > > > > > > that.
> 
> > > > > > > > > Peter.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [January 11, 2012, 8:38pm UTC](https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211/13 "2012-01-11T20:38:17Z")

</div>

Regarding the caching, you did not show an example of how you try and set  
it, so I can't help, but the \_cache is probably set in the wrong place. In  
any cas, you don't need to set it, as term filter is cached by default.

Regarding the field type of pd\_id, you define it as numeric, which is fine.  
Note that \_type is a String, but it does not really matter that much since  
we are caching the filters results. Note though, long is 64bit signed , and  
integer is 32bit signed.

I don't really understand where this change is coming from, it makes very  
little sense. You can dropbox the ES data directory you are working with,  
and two sample curl queries, one against type and one against pd\_id, and I  
can have a look.

On Wed, Jan 11, 2012 at 10:14 PM, Greg Ichneumon Brown \<[gbrown5878@gmail.com](mailto:gbrown5878@gmail.com)

> wrote:

> No matter how many times I repeat that query I am getting "took" :  
> 125, whereas the type query gives "took" : 13. I also built a filtered  
> query using \_type and that completes in 13 ms also.
> 
> Could my mapping for this field be the culprit? I am doing:
> 
> ```
> 'pd_id' => array( 'type' => 'long', 'store' => 'yes',
> 
> ```
> 
> 'index' =\>  
> 'not\_analyzed' )
> 
> Each document has an integer as an id, I used long as the storage type  
> to reduce memory, but does this not work with the term filter?
> 
> I tried setting \_cache to true in the filter to see if that forced the  
> caching, but then I get the correct number of total hits, but no facet  
> results:  
> {  
> "took" : 1,  
> "timed\_out" : false,  
> "\_shards" : {  
> "total" : 5,  
> "successful" : 5,  
> "failed" : 0  
> },  
> "hits" : {  
> "total" : 11588,  
> "max\_score" : 1.0,  
> "hits" :   
> }  
> }
> 
> Should be:  
> {  
> "took" : 13,  
> "timed\_out" : false,  
> "\_shards" : {  
> "total" : 5,  
> "successful" : 5,  
> "failed" : 0  
> },  
> "hits" : {  
> "total" : 11588,  
> "max\_score" : 1.0,  
> "hits" :   
> },  
> "facets" : {  
> "q1" : {  
> "\_type" : "terms",  
> "missing" : 0,  
> "total" : 11720,  
> "other" : 56,  
> "terms" : [ {  
> "term" : "adopt",  
> "count" : 11475  
> }, {  
> "term" : "adoption",  
> "count" : 39  
> }, {  
> etc...
> 
> On Jan 11, 11:00 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > Yes, what you posted is exactly the query that you will end up with when  
> > searching against a type. Maybe you just didn't let it do its caching  
> > bit?  
> > How fast is the 2-3rd execution using the same pid (the term filter  
> > result  
> > is cached).
> > 
> > On Wed, Jan 11, 2012 at 6:35 PM, Greg Ichneumon Brown  
> > [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com)wrote:
> > 
> > > Ah, I see what you are saying. But I am totally flummoxed as to how to  
> > > formulate the query to get a filtered query that matches the  
> > > performance of the default filtered query used for the type.
> > 
> > > I think what you are suggesting is:  
> > > curl -XGET "${SERVER}/pd-test-0/\_search?pretty"   
> > > -d '{  
> > > "size" : 0,  
> > > "query" : {  
> > > "filtered" : {  
> > > "query" : { "match\_all" : { } },  
> > > "filter" : {  
> > > "term" : { "pd\_id" : "$ID" }  
> > > }  
> > > } },  
> > > "facets" : {  
> > > "q1" : {  
> > > "terms" : {  
> > > "field" : "q1",  
> > > "size" : 100  
> > > }  
> > > }  
> > > }  
> > > }'
> > 
> > > Is that query match\_all correct? This query takes about 125 ms vs 13  
> > > ms for a query on the type (curl -XGET "${SERVER}/pd-0/${ID}/\_search).
> > 
> > > From what I can gather from the Java code this should mostly match  
> > > that query, but don't know the code well enough. Is there an easy way  
> > > to enable logging that would let me compare the structure of the  
> > > parsed queries for debugging this?
> > 
> > > Thanks for all the help, Shay! Much appreciated.  
> > > -Greg
> > 
> > > On Jan 10, 10:50 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> > > 
> > > > 10x slower than types? It makes little sense since types, at teh end  
> > > > of  
> > > > the  
> > > > day, is just a field called \_type in a document, and when you search  
> > > > within  
> > > > a type, your query provided is simply wrapped in a filtered query a  
> > > > filter  
> > > > on the type. So, you can do it yourself, just wrap your query in a  
> > > > filtered  
> > > > query with a filter on your "type".
> > 
> > > > On Tue, Jan 10, 2012 at 5:48 PM, Greg Ichneumon Brown  
> > > > [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com)wrote:
> > 
> > > > > That's what I did. Functionally works, but it is 10x slower to  
> > > > > query  
> > > > > using either a query to filter or a facet\_filter. Is there another  
> > > > > way? According to the docs: "search filters restrict only returned  
> > > > > documents — but not facet counts"
> > 
> > > > > On Jan 10, 2:40 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> > > > > 
> > > > > > Just add the "type" as a field to the doc, and filter by it.
> > 
> > > > > > On Tue, Jan 10, 2012 at 5:45 AM, Greg Brown \<  
> > > > > > [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com)\>  
> > > > > > wrote:
> > > > > > 
> > > > > > > Indexing all data to a single type did work fine (3.3mil docs)  
> > > > > > > as  
> > > > > > > expected.
> > 
> > > > > > > I submitted a bug ([elasticsearch · GitHub](https://github.com/elasticsearch/)  
> > > > > > > [elasticsearch.github.com/issues/134](http://elasticsearch.github.com/issues/134)) on the large number of  
> > > > > > > types  
> > > > > > > because I was able to get the server to become unresponsive  
> > > > > > > even  
> > > > > > > when  
> > > > > > > there was only a single server and I tried to add many types.
> > 
> > > > > > > For the moment I am going ahead with using all of the  
> > > > > > > documents in  
> > > > > > > a  
> > > > > > > single index. However, this significantly reduces query  
> > > > > > > performance  
> > > > > > > compared to having a separate type for each set of documents. I  
> > > > > > > looped  
> > > > > > > and profiled the following queries on the larger sets of  
> > > > > > > documents  
> > > > > > > (10k-70k):[https://gist.github.com/1586723Thiswasall](https://gist.github.com/1586723Thiswasall) run on a  
> > > > > > > single server, and the query from a different machine.
> > 
> > > > > > > The first query has each set of docs in its own Type. On  
> > > > > > > average it  
> > > > > > > took about 11 ms to complete.
> > 
> > > > > > > The second has all of the docs in one index with a field pd\_id  
> > > > > > > to  
> > > > > > > distinguish the sets. The query uses the facet\_filter and  
> > > > > > > average  
> > > > > > > ~190  
> > > > > > > ms.
> > 
> > > > > > > The third uses the same index as the second, but uses a query  
> > > > > > > to do  
> > > > > > > the "filtering" of the docs. ~140 ms. I was surprised that  
> > > > > > > this was  
> > > > > > > faster than the facet\_filter.
> > 
> > > > > > > Any suggestions on how to improve the last two queries?
> > 
> > > > > > > Any ideas on how to create multiple types without creating 10k  
> > > > > > > separate indices. In this case all I am using the Type for is a  
> > > > > > > partitioning/grouping of multiple separate indices, since the  
> > > > > > > Mapping  
> > > > > > > of each Type is identical.
> > 
> > > > > > > Thanks for the help.  
> > > > > > > -Greg
> > 
> > > > > > > On Dec 22 2011, 7:37 am, Greg Brown [gbrown5...@gmail.com](mailto:gbrown5...@gmail.com)  
> > > > > > > wrote:
> > > > > > > 
> > > > > > > > Checking through the logs, there isn't any mention of there  
> > > > > > > > not  
> > > > > > > > being  
> > > > > > > > enough file handles, the errors I am running into are out of  
> > > > > > > > memory  
> > > > > > > > on  
> > > > > > > > the heap space errors.
> > 
> > > > > > > > Shay,
> > 
> > > > > > > > Thanks, will give that a try and let you know. Will have to  
> > > > > > > > wait  
> > > > > > > > until  
> > > > > > > > after the weekend so I can set up a development cluster. I've  
> > > > > > > > brought  
> > > > > > > > down the production cluster a few too many times this week,  
> > > > > > > > and  
> > > > > > > > its  
> > > > > > > > time to be more careful. 🙂
> > 
> > > > > > > > Thanks for the fast responses.  
> > > > > > > > -Greg
> > 
> > > > > > > > On Dec 21, 6:54 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> > 
> > > > > > > > > My guess is that the problem is with creating so many  
> > > > > > > > > types,  
> > > > > > > > > which  
> > > > > > > > > ends up  
> > > > > > > > > being a large overhead in the system. Each time a type is  
> > > > > > > > > introduced,  
> > > > > > > > > it  
> > > > > > > > > needs to be broadcasted to the rest of the nodes and  
> > > > > > > > > persisted  
> > > > > > > > > as  
> > > > > > > > > part  
> > > > > > > > > of  
> > > > > > > > > the cluster meta data. Can you try just indexing into the  
> > > > > > > > > same  
> > > > > > > > > type as  
> > > > > > > > > a  
> > > > > > > > > test and see if it still happens?
> > 
> > > > > > > > > On Wed, Dec 21, 2011 at 9:08 PM, Karussell \<  
> > > > > > > > > [tableyourt...@googlemail.com](mailto:tableyourt...@googlemail.com)\>wrote:
> > 
> > > > > > > > > > > Greg, have you checked/increased open file handle  
> > > > > > > > > > > limits for  
> > > > > > > > > > > your  
> > > > > > > > > > > machine?
> > 
> > > > > > > > > > First, check/post your logs. If too many files open ES  
> > > > > > > > > > would  
> > > > > > > > > > log  
> > > > > > > > > > that.
> > 
> > > > > > > > > > Peter.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:43am UTC](https://discuss.elastic.co/t/2-node-cluster-hanging-while-bulk-indexing-and-adding-types/6211/14 "2017-07-06T03:43:01Z")

</div>


