# After upgrade to .16, problems

**URL:** <https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298>\
**Category:** Elasticsearch\
**Created:** [April 26, 2011, 4:00pm UTC](https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298 "2011-04-26T16:00:26Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![electic](https://avatars.discourse-cdn.com/v4/letter/e/ecccb3/32.png) [@electic](https://discuss.elastic.co/u/electic)\
**Post date:** [April 26, 2011, 4:00pm UTC](https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298/1 "2011-04-26T16:00:26Z")

</div>

Hi,

I upgraded to .16 and I am having a multitude of issues. The biggest  
seems to be during inserts:

org.elasticsearch.index.engine.EngineClosedException: [documents][3]  
CurrentState[CLOSED]  
at  
org.elasticsearch.index.engine.robin.RobinEngine.create(RobinEngine.java:  
264)  
at  
org.elasticsearch.index.shard.service.InternalIndexShard.create(InternalIndexShard.java:  
272)  
at  
org.elasticsearch.action.bulk.TransportShardBulkAction.shardOperationOnPrimary(TransportShardBulkAction.java:  
136)  
at  
org.elasticsearch.action.support.replication.TransportShardReplicationOperationAction  
$AsyncShardOperationAction.performOnPrimary(TransportShardReplicationOperationAction.java:  
418)  
at  
org.elasticsearch.action.support.replication.TransportShardReplicationOperationAction  
$AsyncShardOperationAction.access  
$100(TransportShardReplicationOperationAction.java:233)  
at  
org.elasticsearch.action.support.replication.TransportShardReplicationOperationAction  
$AsyncShardOperationAction  
$1.run(TransportShardReplicationOperationAction.java:331)  
at  
java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:  
1110)  
at java.util.concurrent.ThreadPoolExecutor  
$Worker.run(ThreadPoolExecutor.java:603)  
at java.lang.Thread.run(Thread.java:636)  
Caused by: java.lang.OutOfMemoryError: Java heap space  
at  
org.elasticsearch.common.compress.lzf.BufferRecycler.allocEncodingBuffer(BufferRecycler.java:  
76)  
at  
org.elasticsearch.common.compress.lzf.ChunkEncoder.(ChunkEncoder.java:  
69)  
at  
org.elasticsearch.common.io.stream.LZFStreamOutput.(LZFStreamOutput.java:  
46)  
at org.elasticsearch.common.io.stream.CachedStreamOutput  
$1.initialValue(CachedStreamOutput.java:47)  
at org.elasticsearch.common.io.stream.CachedStreamOutput  
$1.initialValue(CachedStreamOutput.java:43)  
at java.lang.ThreadLocal.setInitialValue(ThreadLocal.java:160)  
at java.lang.ThreadLocal.get(ThreadLocal.java:150)  
at  
org.elasticsearch.common.io.stream.CachedStreamOutput.cachedBytes(CachedStreamOutput.java:  
56)  
at org.elasticsearch.index.translog.fs.FsTranslog.add(FsTranslog.java:  
159)  
at  
org.elasticsearch.index.engine.robin.RobinEngine.innerCreate(RobinEngine.java:  
361)  
at  
org.elasticsearch.index.engine.robin.RobinEngine.create(RobinEngine.java:  
266)  
... 8 more

We are also noticing really high CPU usage, 400% (there are 4 cores)  
on the Ubuntu machine. Is there a way to query the cluster and find  
out what operations it is working on? Here is the cluster health (the  
index we use is called documents):

{"cluster\_name":"viralheat","master\_node":"mjQhyNydTZycnLZseq8lHQ","blocks":  
{},"nodes":{"mjQhyNydTZycnLZseq8lHQ":  
{"name":"Alchemy","transport\_address":"inet[/  
192.168.8.230:9300]","attributes":{}}},"metadata":{"templates":  
{},"indices":{"twitter":{"state":"open","settings":  
{"index.number\_of\_shards":"5","index.number\_of\_replicas":"1"},"mappings":  
{},"aliases":[]},"documents":{"state":"open","settings":  
{"index.number\_of\_replicas":"0","index.number\_of\_shards":"5"},"mappings":  
{"document":{"properties":{"tags":{"type":"string"},"platform":  
{"type":"string"},"utimestamp":{"type":"string"},"document":  
{"type":"string"},"created\_at":  
{"format":"dateOptionalTime","type":"date"},"record\_id":  
{"type":"string"}}},"documents":{"properties":{"tags":  
{"type":"string"},"platform":{"type":"string"},"utimestamp":  
{"type":"long"},"document":{"type":"string"},"created\_at":  
{"format":"dateOptionalTime","type":"date"},"record\_id":  
{"type":"string"}}}},"aliases":[]}}},"routing\_table":{"indices":  
{"twitter":{"shards":{"0":  
[{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
0,"index":"twitter"},  
{"state":"UNASSIGNED","primary":false,"node":null,"relocating\_node":null,"shard":  
0,"index":"twitter"}],"1":  
[{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
1,"index":"twitter"},  
{"state":"UNASSIGNED","primary":false,"node":null,"relocating\_node":null,"shard":  
1,"index":"twitter"}],"2":  
[{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
2,"index":"twitter"},  
{"state":"UNASSIGNED","primary":false,"node":null,"relocating\_node":null,"shard":  
2,"index":"twitter"}],"3":  
[{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
3,"index":"twitter"},  
{"state":"UNASSIGNED","primary":false,"node":null,"relocating\_node":null,"shard":  
3,"index":"twitter"}],"4":  
[{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
4,"index":"twitter"},  
{"state":"UNASSIGNED","primary":false,"node":null,"relocating\_node":null,"shard":  
4,"index":"twitter"}]}},"documents":{"shards":{"0":  
[{"state":"UNASSIGNED","primary":true,"node":null,"relocating\_node":null,"shard":  
0,"index":"documents"}],"1":  
[{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
1,"index":"documents"}],"2":  
[{"state":"UNASSIGNED","primary":true,"node":null,"relocating\_node":null,"shard":  
2,"index":"documents"}],"3":  
[{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
3,"index":"documents"}],"4":  
[{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
4,"index":"documents"}]}}}},"routing\_nodes":{"unassigned":  
[{"state":"UNASSIGNED","primary":false,"node":null,"relocating\_node":null,"shard":  
0,"index":"twitter"},  
{"state":"UNASSIGNED","primary":false,"node":null,"relocating\_node":null,"shard":  
1,"index":"twitter"},  
{"state":"UNASSIGNED","primary":false,"node":null,"relocating\_node":null,"shard":  
2,"index":"twitter"},  
{"state":"UNASSIGNED","primary":false,"node":null,"relocating\_node":null,"shard":  
3,"index":"twitter"},  
{"state":"UNASSIGNED","primary":false,"node":null,"relocating\_node":null,"shard":  
4,"index":"twitter"},  
{"state":"UNASSIGNED","primary":true,"node":null,"relocating\_node":null,"shard":  
0,"index":"documents"},  
{"state":"UNASSIGNED","primary":true,"node":null,"relocating\_node":null,"shard":  
2,"index":"documents"}],"nodes":{"mjQhyNydTZycnLZseq8lHQ":  
[{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
0,"index":"twitter"},  
{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
1,"index":"twitter"},  
{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
2,"index":"twitter"},  
{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
3,"index":"twitter"},  
{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
4,"index":"twitter"},  
{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
1,"index":"documents"},  
{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
3,"index":"documents"},  
{"state":"STARTED","primary":true,"node":"mjQhyNydTZycnLZseq8lHQ","relocating\_node":null,"shard":  
4,"index":"documents"}]}},"allocations":[]}

---

<div class="post-metadata">

**Author:** ![electic](https://avatars.discourse-cdn.com/v4/letter/e/ecccb3/32.png) [@electic](https://discuss.elastic.co/u/electic)\
**Post date:** [April 26, 2011, 4:12pm UTC](https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298/2 "2011-04-26T16:12:56Z")

</div>

One more thing, I also noticed one of my nodes is gone. I had two  
machines in the cluster. For some odd reason, it's not discovering the  
231 box:

{"cluster\_name":"x","nodes":{"mjQhyNydTZycnLZseq8lHQ":  
{"name":"Alchemy","transport\_address":"inet[/  
192.168.8.230:9300]","attributes":{},"http\_address":"inet[/  
192.168.8.230:9200]","os":{"refresh\_interval":5000,"cpu":  
{"vendor":"AMD","model":"Dual Core AMD Opteron(tm) Processor  
275","mhz":2209,"total\_cores":4,"total\_sockets":2,"cores\_per\_socket":  
2,"cache\_size":"1kb","cache\_size\_in\_bytes":1024},"mem":  
{"total":"14.9gb","total\_in\_bytes":16070246400},"swap":  
{"total":"37.6gb","total\_in\_bytes":40470831104}},"process":  
{"refresh\_interval":5000,"id":2298},"jvm":{"pid":  
2298,"version":"1.6.0\_20","vm\_name":"OpenJDK 64-Bit Server  
VM","vm\_version":"19.0-b09","vm\_vendor":"Sun Microsystems  
Inc.","start\_time":1303806081608,"mem":  
{"heap\_init":"256mb","heap\_init\_in\_bytes":  
268435456,"heap\_max":"1015.6mb","heap\_max\_in\_bytes":  
1065025536,"non\_heap\_init":"23.1mb","non\_heap\_init\_in\_bytes":  
24313856,"non\_heap\_max":"214mb","non\_heap\_max\_in\_bytes":  
224395264}},"network":{"refresh\_interval":5000,"primary\_interface":  
{"address":"192.168.8.230","name":"eth1","mac\_address":"00:E0:81:41:AF:  
7F"}},"transport":{"bound\_address":"inet[/  
192.168.8.230:9300]","publish\_address":"inet[/192.168.8.230:9300]"}}}}

---

<div class="post-metadata">

**Author:** ![Igor\_Motov1](https://avatars.discourse-cdn.com/v4/letter/i/c5a1d2/32.png) [@Igor\_Motov1](https://discuss.elastic.co/u/Igor_Motov1)\
**Post date:** [April 26, 2011, 4:17pm UTC](https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298/3 "2011-04-26T16:17:10Z")

</div>

It looks like one of your nodes ran out of memory:

java.lang.OutOfMemoryError: Java heap space

On Tue, Apr 26, 2011 at 12:12 PM, electic [electic@gmail.com](mailto:electic@gmail.com) wrote:

> One more thing, I also noticed one of my nodes is gone. I had two  
> machines in the cluster. For some odd reason, it's not discovering the  
> 231 box:
> 
> {"cluster\_name":"x","nodes":{"mjQhyNydTZycnLZseq8lHQ":  
> {"name":"Alchemy","transport\_address":"inet[/  
> 192.168.8.230:9300]","attributes":{},"http\_address":"inet[/  
> 192.168.8.230:9200]","os":{"refresh\_interval":5000,"cpu":  
> {"vendor":"AMD","model":"Dual Core AMD Opteron(tm) Processor  
> 275","mhz":2209,"total\_cores":4,"total\_sockets":2,"cores\_per\_socket":  
> 2,"cache\_size":"1kb","cache\_size\_in\_bytes":1024},"mem":  
> {"total":"14.9gb","total\_in\_bytes":16070246400},"swap":  
> {"total":"37.6gb","total\_in\_bytes":40470831104}},"process":  
> {"refresh\_interval":5000,"id":2298},"jvm":{"pid":  
> 2298,"version":"1.6.0\_20","vm\_name":"OpenJDK 64-Bit Server  
> VM","vm\_version":"19.0-b09","vm\_vendor":"Sun Microsystems  
> Inc.","start\_time":1303806081608,"mem":  
> {"heap\_init":"256mb","heap\_init\_in\_bytes":  
> 268435456,"heap\_max":"1015.6mb","heap\_max\_in\_bytes":  
> 1065025536,"non\_heap\_init":"23.1mb","non\_heap\_init\_in\_bytes":  
> 24313856,"non\_heap\_max":"214mb","non\_heap\_max\_in\_bytes":  
> 224395264}},"network":{"refresh\_interval":5000,"primary\_interface":  
> {"address":"192.168.8.230","name":"eth1","mac\_address":"00:E0:81:41:AF:  
> 7F"}},"transport":{"bound\_address":"inet[/  
> 192.168.8.230:9300]","publish\_address":"inet[/192.168.8.230:9300]"}}}}

---

<div class="post-metadata">

**Author:** ![electic](https://avatars.discourse-cdn.com/v4/letter/e/ecccb3/32.png) [@electic](https://discuss.elastic.co/u/electic)\
**Post date:** [April 26, 2011, 4:21pm UTC](https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298/4 "2011-04-26T16:21:05Z")

</div>

That's interesting. As I just turned it on for about an hour after I  
upgraded it. Here is the box memory usage:

Mem: 15693600k total, 6643536k used, 9050064k free, 73648k  
buffers

is there anyway to solve this?

On Apr 26, 9:17 am, Igor Motov [igor.mo...@sonian.net](mailto:igor.mo...@sonian.net) wrote:

> It looks like one of your nodes ran out of memory:
> 
> java.lang.OutOfMemoryError: Java heap space
> 
> On Tue, Apr 26, 2011 at 12:12 PM, electic [elec...@gmail.com](mailto:elec...@gmail.com) wrote:
> 
> > One more thing, I also noticed one of my nodes is gone. I had two  
> > machines in the cluster. For some odd reason, it's not discovering the  
> > 231 box:
> 
> > {"cluster\_name":"x","nodes":{"mjQhyNydTZycnLZseq8lHQ":  
> > {"name":"Alchemy","transport\_address":"inet[/  
> > 192.168.8.230:9300]","attributes":{},"http\_address":"inet[/  
> > 192.168.8.230:9200]","os":{"refresh\_interval":5000,"cpu":  
> > {"vendor":"AMD","model":"Dual Core AMD Opteron(tm) Processor  
> > 275","mhz":2209,"total\_cores":4,"total\_sockets":2,"cores\_per\_socket":  
> > 2,"cache\_size":"1kb","cache\_size\_in\_bytes":1024},"mem":  
> > {"total":"14.9gb","total\_in\_bytes":16070246400},"swap":  
> > {"total":"37.6gb","total\_in\_bytes":40470831104}},"process":  
> > {"refresh\_interval":5000,"id":2298},"jvm":{"pid":  
> > 2298,"version":"1.6.0\_20","vm\_name":"OpenJDK 64-Bit Server  
> > VM","vm\_version":"19.0-b09","vm\_vendor":"Sun Microsystems  
> > Inc.","start\_time":1303806081608,"mem":  
> > {"heap\_init":"256mb","heap\_init\_in\_bytes":  
> > 268435456,"heap\_max":"1015.6mb","heap\_max\_in\_bytes":  
> > 1065025536,"non\_heap\_init":"23.1mb","non\_heap\_init\_in\_bytes":  
> > 24313856,"non\_heap\_max":"214mb","non\_heap\_max\_in\_bytes":  
> > 224395264}},"network":{"refresh\_interval":5000,"primary\_interface":  
> > {"address":"192.168.8.230","name":"eth1","mac\_address":"00:E0:81:41:AF:  
> > 7F"}},"transport":{"bound\_address":"inet[/  
> > 192.168.8.230:9300]","publish\_address":"inet[/192.168.8.230:9300]"}}}}

---

<div class="post-metadata">

**Author:** ![Igor\_Motov1](https://avatars.discourse-cdn.com/v4/letter/i/c5a1d2/32.png) [@Igor\_Motov1](https://discuss.elastic.co/u/Igor_Motov1)\
**Post date:** [April 26, 2011, 4:37pm UTC](https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298/5 "2011-04-26T16:37:33Z")

</div>

What about jvm memory usage? What does jvm section of curl -XGET '  
[http://localhost:9200/\_cluster/nodes/stats?pretty=true](http://localhost:9200/_cluster/nodes/stats?pretty=true)' contain and how does  
it compare to what is specified in -Xmx parameter of elasticsearch process?

On Tue, Apr 26, 2011 at 12:21 PM, electic [electic@gmail.com](mailto:electic@gmail.com) wrote:

> That's interesting. As I just turned it on for about an hour after I  
> upgraded it. Here is the box memory usage:
> 
> Mem: 15693600k total, 6643536k used, 9050064k free, 73648k  
> buffers
> 
> is there anyway to solve this?
> 
> On Apr 26, 9:17 am, Igor Motov [igor.mo...@sonian.net](mailto:igor.mo...@sonian.net) wrote:
> 
> > It looks like one of your nodes ran out of memory:
> > 
> > java.lang.OutOfMemoryError: Java heap space
> > 
> > On Tue, Apr 26, 2011 at 12:12 PM, electic [elec...@gmail.com](mailto:elec...@gmail.com) wrote:
> > 
> > > One more thing, I also noticed one of my nodes is gone. I had two  
> > > machines in the cluster. For some odd reason, it's not discovering the  
> > > 231 box:
> > 
> > > {"cluster\_name":"x","nodes":{"mjQhyNydTZycnLZseq8lHQ":  
> > > {"name":"Alchemy","transport\_address":"inet[/  
> > > 192.168.8.230:9300]","attributes":{},"http\_address":"inet[/  
> > > 192.168.8.230:9200]","os":{"refresh\_interval":5000,"cpu":  
> > > {"vendor":"AMD","model":"Dual Core AMD Opteron(tm) Processor  
> > > 275","mhz":2209,"total\_cores":4,"total\_sockets":2,"cores\_per\_socket":  
> > > 2,"cache\_size":"1kb","cache\_size\_in\_bytes":1024},"mem":  
> > > {"total":"14.9gb","total\_in\_bytes":16070246400},"swap":  
> > > {"total":"37.6gb","total\_in\_bytes":40470831104}},"process":  
> > > {"refresh\_interval":5000,"id":2298},"jvm":{"pid":  
> > > 2298,"version":"1.6.0\_20","vm\_name":"OpenJDK 64-Bit Server  
> > > VM","vm\_version":"19.0-b09","vm\_vendor":"Sun Microsystems  
> > > Inc.","start\_time":1303806081608,"mem":  
> > > {"heap\_init":"256mb","heap\_init\_in\_bytes":  
> > > 268435456,"heap\_max":"1015.6mb","heap\_max\_in\_bytes":  
> > > 1065025536,"non\_heap\_init":"23.1mb","non\_heap\_init\_in\_bytes":  
> > > 24313856,"non\_heap\_max":"214mb","non\_heap\_max\_in\_bytes":  
> > > 224395264}},"network":{"refresh\_interval":5000,"primary\_interface":  
> > > {"address":"192.168.8.230","name":"eth1","mac\_address":"00:E0:81:41:AF:  
> > > 7F"}},"transport":{"bound\_address":"inet[/  
> > > 192.168.8.230:9300]","publish\_address":"inet[/192.168.8.230:9300]"}}}}

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [April 26, 2011, 4:41pm UTC](https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298/6 "2011-04-26T16:41:03Z")

</div>

> We are also noticing really high CPU usage,  
> 400% (there are 4 cores) on the Ubuntu machine.

When does your CPU use 400%? Directly after the start or some hours  
after indexing?

I had hard problems (0.16.0 snapshot) discribed here:

[http://groups.google.com/a/elasticsearch.com/group/users/browse\_thread/thread/beb060bf58a2d1df](http://groups.google.com/a/elasticsearch.com/group/users/browse_thread/thread/beb060bf58a2d1df)

which is now partially solved: the OS did not crash anymore, thread  
count is now normal (~70),  
the node is responsive, but still high CPU usage (in our case 700%)!  
After a restart of ES the node takes \<= 50% of CPU and it somehow goes  
up after some hours.  
Now I need to find out when (my feeling was that it missed out some  
optimizations and goes up due to missing optimization ... but I'm  
still investigating)

Regards,  
Peter.

On Apr 26, 6:21 pm, electic [elec...@gmail.com](mailto:elec...@gmail.com) wrote:

> That's interesting. As I just turned it on for about an hour after I  
> upgraded it. Here is the box memory usage:
> 
> Mem: 15693600k total, 6643536k used, 9050064k free, 73648k  
> buffers
> 
> is there anyway to solve this?
> 
> On Apr 26, 9:17 am, Igor Motov [igor.mo...@sonian.net](mailto:igor.mo...@sonian.net) wrote:
> 
> > It looks like one of your nodes ran out of memory:
> 
> > java.lang.OutOfMemoryError: Java heap space
> 
> > On Tue, Apr 26, 2011 at 12:12 PM, electic [elec...@gmail.com](mailto:elec...@gmail.com) wrote:
> > 
> > > One more thing, I also noticed one of my nodes is gone. I had two  
> > > machines in the cluster. For some odd reason, it's not discovering the  
> > > 231 box:
> 
> > > {"cluster\_name":"x","nodes":{"mjQhyNydTZycnLZseq8lHQ":  
> > > {"name":"Alchemy","transport\_address":"inet[/  
> > > 192.168.8.230:9300]","attributes":{},"http\_address":"inet[/  
> > > 192.168.8.230:9200]","os":{"refresh\_interval":5000,"cpu":  
> > > {"vendor":"AMD","model":"Dual Core AMD Opteron(tm) Processor  
> > > 275","mhz":2209,"total\_cores":4,"total\_sockets":2,"cores\_per\_socket":  
> > > 2,"cache\_size":"1kb","cache\_size\_in\_bytes":1024},"mem":  
> > > {"total":"14.9gb","total\_in\_bytes":16070246400},"swap":  
> > > {"total":"37.6gb","total\_in\_bytes":40470831104}},"process":  
> > > {"refresh\_interval":5000,"id":2298},"jvm":{"pid":  
> > > 2298,"version":"1.6.0\_20","vm\_name":"OpenJDK 64-Bit Server  
> > > VM","vm\_version":"19.0-b09","vm\_vendor":"Sun Microsystems  
> > > Inc.","start\_time":1303806081608,"mem":  
> > > {"heap\_init":"256mb","heap\_init\_in\_bytes":  
> > > 268435456,"heap\_max":"1015.6mb","heap\_max\_in\_bytes":  
> > > 1065025536,"non\_heap\_init":"23.1mb","non\_heap\_init\_in\_bytes":  
> > > 24313856,"non\_heap\_max":"214mb","non\_heap\_max\_in\_bytes":  
> > > 224395264}},"network":{"refresh\_interval":5000,"primary\_interface":  
> > > {"address":"192.168.8.230","name":"eth1","mac\_address":"00:E0:81:41:AF:  
> > > 7F"}},"transport":{"bound\_address":"inet[/  
> > > 192.168.8.230:9300]","publish\_address":"inet[/192.168.8.230:9300]"}}}}

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [April 26, 2011, 4:43pm UTC](https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298/7 "2011-04-26T16:43:26Z")

</div>

But check RAM usage first and how long the GC takes  
or if you have problems with merging.

e.g. use jvisualvm, node info, thread dumps (kill -3  
) etc

On Apr 26, 6:21 pm, electic [elec...@gmail.com](mailto:elec...@gmail.com) wrote:

> That's interesting. As I just turned it on for about an hour after I  
> upgraded it. Here is the box memory usage:
> 
> Mem: 15693600k total, 6643536k used, 9050064k free, 73648k  
> buffers
> 
> is there anyway to solve this?
> 
> On Apr 26, 9:17 am, Igor Motov [igor.mo...@sonian.net](mailto:igor.mo...@sonian.net) wrote:
> 
> > It looks like one of your nodes ran out of memory:
> 
> > java.lang.OutOfMemoryError: Java heap space
> 
> > On Tue, Apr 26, 2011 at 12:12 PM, electic [elec...@gmail.com](mailto:elec...@gmail.com) wrote:
> > 
> > > One more thing, I also noticed one of my nodes is gone. I had two  
> > > machines in the cluster. For some odd reason, it's not discovering the  
> > > 231 box:
> 
> > > {"cluster\_name":"x","nodes":{"mjQhyNydTZycnLZseq8lHQ":  
> > > {"name":"Alchemy","transport\_address":"inet[/  
> > > 192.168.8.230:9300]","attributes":{},"http\_address":"inet[/  
> > > 192.168.8.230:9200]","os":{"refresh\_interval":5000,"cpu":  
> > > {"vendor":"AMD","model":"Dual Core AMD Opteron(tm) Processor  
> > > 275","mhz":2209,"total\_cores":4,"total\_sockets":2,"cores\_per\_socket":  
> > > 2,"cache\_size":"1kb","cache\_size\_in\_bytes":1024},"mem":  
> > > {"total":"14.9gb","total\_in\_bytes":16070246400},"swap":  
> > > {"total":"37.6gb","total\_in\_bytes":40470831104}},"process":  
> > > {"refresh\_interval":5000,"id":2298},"jvm":{"pid":  
> > > 2298,"version":"1.6.0\_20","vm\_name":"OpenJDK 64-Bit Server  
> > > VM","vm\_version":"19.0-b09","vm\_vendor":"Sun Microsystems  
> > > Inc.","start\_time":1303806081608,"mem":  
> > > {"heap\_init":"256mb","heap\_init\_in\_bytes":  
> > > 268435456,"heap\_max":"1015.6mb","heap\_max\_in\_bytes":  
> > > 1065025536,"non\_heap\_init":"23.1mb","non\_heap\_init\_in\_bytes":  
> > > 24313856,"non\_heap\_max":"214mb","non\_heap\_max\_in\_bytes":  
> > > 224395264}},"network":{"refresh\_interval":5000,"primary\_interface":  
> > > {"address":"192.168.8.230","name":"eth1","mac\_address":"00:E0:81:41:AF:  
> > > 7F"}},"transport":{"bound\_address":"inet[/  
> > > 192.168.8.230:9300]","publish\_address":"inet[/192.168.8.230:9300]"}}}}

---

<div class="post-metadata">

**Author:** ![electic](https://avatars.discourse-cdn.com/v4/letter/e/ecccb3/32.png) [@electic](https://discuss.elastic.co/u/electic)\
**Post date:** [April 26, 2011, 4:44pm UTC](https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298/8 "2011-04-26T16:44:27Z")

</div>

My guess is that its a few hours after restart. That is why I somewhat  
wanted to know what it was doing.

@igor, I had to reboot the box but I will paste that if it happens  
again. Anyway to increate the -Xmx variable for the java heap size?

---

<div class="post-metadata">

**Author:** ![electic](https://avatars.discourse-cdn.com/v4/letter/e/ecccb3/32.png) [@electic](https://discuss.elastic.co/u/electic)\
**Post date:** [April 26, 2011, 4:46pm UTC](https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298/9 "2011-04-26T16:46:11Z")

</div>

Okay, just restarted it. Its at about 105 percent CPU. Here is the  
stats:

{  
"cluster\_name" : "xx",  
"nodes" : {  
"Mk1V3glvRxinJr\_bEZ8DDQ" : {  
"name" : "Shadowcat",  
"indices" : {  
"size" : "60gb",  
"size\_in\_bytes" : 64434244345,  
"docs" : {  
"num\_docs" : 29663747  
},  
"cache" : {  
"field\_evictions" : 0,  
"field\_size" : "0b",  
"field\_size\_in\_bytes" : 0,  
"filter\_count" : 0,  
"filter\_evictions" : 0,  
"filter\_mem\_evictions" : 0,  
"filter\_size" : "0b",  
"filter\_size\_in\_bytes" : 0  
},  
"merges" : {  
"current" : 1,  
"total" : 2,  
"total\_time" : "35.7s",  
"total\_time\_in\_millis" : 35783  
}  
},  
"os" : {  
"timestamp" : 1303836311658,  
"uptime" : "8 hours, 40 minutes and 15 seconds",  
"uptime\_in\_millis" : 31215000,  
"load\_average" : [13.65, 70.19, 94.85],  
"cpu" : {  
"sys" : 7,  
"user" : 37,  
"idle" : 47  
},  
"mem" : {  
"free" : "6.5gb",  
"free\_in\_bytes" : 7065710592,  
"used" : "8.3gb",  
"used\_in\_bytes" : 9004535808,  
"free\_percent" : 90,  
"used\_percent" : 9,  
"actual\_free" : "13.5gb",  
"actual\_free\_in\_bytes" : 14502236160,  
"actual\_used" : "1.4gb",  
"actual\_used\_in\_bytes" : 1568010240  
},  
"swap" : {  
"used" : "0b",  
"used\_in\_bytes" : 0,  
"free" : "37.6gb",  
"free\_in\_bytes" : 40470831104  
}  
},  
"process" : {  
"timestamp" : 1303836311658,  
"cpu" : {  
"percent" : 183,  
"sys" : "31 seconds and 700 milliseconds",  
"sys\_in\_millis" : 31700,  
"user" : "3 minutes and 600 milliseconds",  
"user\_in\_millis" : 180600,  
"total" : "-1 milliseconds",  
"total\_in\_millis" : -1  
},  
"mem" : {  
"resident" : "980.3mb",  
"resident\_in\_bytes" : 1027993600,  
"share" : "10.7mb",  
"share\_in\_bytes" : 11259904,  
"total\_virtual" : "1.5gb",  
"total\_virtual\_in\_bytes" : 1668825088  
},  
"fd" : {  
"total" : 651  
}  
},  
"jvm" : {  
"timestamp" : 1303836311660,  
"uptime" : "1 minute, 56 seconds and 642 milliseconds",  
"uptime\_in\_millis" : 116642,  
"mem" : {  
"heap\_used" : "718.7mb",  
"heap\_used\_in\_bytes" : 753624584,  
"heap\_committed" : "971.1mb",  
"heap\_committed\_in\_bytes" : 1018298368,  
"non\_heap\_used" : "33.8mb",  
"non\_heap\_used\_in\_bytes" : 35487376,  
"non\_heap\_committed" : "54.7mb",  
"non\_heap\_committed\_in\_bytes" : 57389056  
},  
"threads" : {  
"count" : 50,  
"peak\_count" : 82  
},  
"gc" : {  
"collection\_count" : 763,  
"collection\_time" : "6 seconds and 401 milliseconds",  
"collection\_time\_in\_millis" : 6401,  
"collectors" : {  
"ParNew" : {  
"collection\_count" : 757,  
"collection\_time" : "6 seconds and 319 milliseconds",  
"collection\_time\_in\_millis" : 6319  
},  
"ConcurrentMarkSweep" : {  
"collection\_count" : 6,  
"collection\_time" : "82 milliseconds",  
"collection\_time\_in\_millis" : 82  
}  
}  
}  
},  
"network" : {  
"tcp" : {  
"active\_opens" : 37,  
"passive\_opens" : 2678,  
"curr\_estab" : 36,  
"in\_segs" : 574496,  
"out\_segs" : 186058,  
"retrans\_segs" : 20,  
"estab\_resets" : 414,  
"attempt\_fails" : 0,  
"in\_errs" : 0,  
"out\_rsts" : 3799  
}  
},  
"transport" : {  
"rx\_count" : 2,  
"rx\_size" : "278b",  
"rx\_size\_in\_bytes" : 278,  
"tx\_count" : 2,  
"tx\_size" : "26b",  
"tx\_size\_in\_bytes" : 26  
}  
}  
}  
}

---

<div class="post-metadata">

**Author:** ![Igor\_Motov1](https://avatars.discourse-cdn.com/v4/letter/i/c5a1d2/32.png) [@Igor\_Motov1](https://discuss.elastic.co/u/Igor_Motov1)\
**Post date:** [April 26, 2011, 4:49pm UTC](https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298/10 "2011-04-26T16:49:15Z")

</div>

If it's linux box, you can run

ps -aef | grep elasticsearch

On Tue, Apr 26, 2011 at 12:46 PM, electic [electic@gmail.com](mailto:electic@gmail.com) wrote:

> Okay, just restarted it. Its at about 105 percent CPU. Here is the  
> stats:
> 
> {  
> "cluster\_name" : "xx",  
> "nodes" : {  
> "Mk1V3glvRxinJr\_bEZ8DDQ" : {  
> "name" : "Shadowcat",  
> "indices" : {  
> "size" : "60gb",  
> "size\_in\_bytes" : 64434244345,  
> "docs" : {  
> "num\_docs" : 29663747  
> },  
> "cache" : {  
> "field\_evictions" : 0,  
> "field\_size" : "0b",  
> "field\_size\_in\_bytes" : 0,  
> "filter\_count" : 0,  
> "filter\_evictions" : 0,  
> "filter\_mem\_evictions" : 0,  
> "filter\_size" : "0b",  
> "filter\_size\_in\_bytes" : 0  
> },  
> "merges" : {  
> "current" : 1,  
> "total" : 2,  
> "total\_time" : "35.7s",  
> "total\_time\_in\_millis" : 35783  
> }  
> },  
> "os" : {  
> "timestamp" : 1303836311658,  
> "uptime" : "8 hours, 40 minutes and 15 seconds",  
> "uptime\_in\_millis" : 31215000,  
> "load\_average" : [13.65, 70.19, 94.85],  
> "cpu" : {  
> "sys" : 7,  
> "user" : 37,  
> "idle" : 47  
> },  
> "mem" : {  
> "free" : "6.5gb",  
> "free\_in\_bytes" : 7065710592,  
> "used" : "8.3gb",  
> "used\_in\_bytes" : 9004535808,  
> "free\_percent" : 90,  
> "used\_percent" : 9,  
> "actual\_free" : "13.5gb",  
> "actual\_free\_in\_bytes" : 14502236160,  
> "actual\_used" : "1.4gb",  
> "actual\_used\_in\_bytes" : 1568010240  
> },  
> "swap" : {  
> "used" : "0b",  
> "used\_in\_bytes" : 0,  
> "free" : "37.6gb",  
> "free\_in\_bytes" : 40470831104  
> }  
> },  
> "process" : {  
> "timestamp" : 1303836311658,  
> "cpu" : {  
> "percent" : 183,  
> "sys" : "31 seconds and 700 milliseconds",  
> "sys\_in\_millis" : 31700,  
> "user" : "3 minutes and 600 milliseconds",  
> "user\_in\_millis" : 180600,  
> "total" : "-1 milliseconds",  
> "total\_in\_millis" : -1  
> },  
> "mem" : {  
> "resident" : "980.3mb",  
> "resident\_in\_bytes" : 1027993600,  
> "share" : "10.7mb",  
> "share\_in\_bytes" : 11259904,  
> "total\_virtual" : "1.5gb",  
> "total\_virtual\_in\_bytes" : 1668825088  
> },  
> "fd" : {  
> "total" : 651  
> }  
> },  
> "jvm" : {  
> "timestamp" : 1303836311660,  
> "uptime" : "1 minute, 56 seconds and 642 milliseconds",  
> "uptime\_in\_millis" : 116642,  
> "mem" : {  
> "heap\_used" : "718.7mb",  
> "heap\_used\_in\_bytes" : 753624584,  
> "heap\_committed" : "971.1mb",  
> "heap\_committed\_in\_bytes" : 1018298368,  
> "non\_heap\_used" : "33.8mb",  
> "non\_heap\_used\_in\_bytes" : 35487376,  
> "non\_heap\_committed" : "54.7mb",  
> "non\_heap\_committed\_in\_bytes" : 57389056  
> },  
> "threads" : {  
> "count" : 50,  
> "peak\_count" : 82  
> },  
> "gc" : {  
> "collection\_count" : 763,  
> "collection\_time" : "6 seconds and 401 milliseconds",  
> "collection\_time\_in\_millis" : 6401,  
> "collectors" : {  
> "ParNew" : {  
> "collection\_count" : 757,  
> "collection\_time" : "6 seconds and 319 milliseconds",  
> "collection\_time\_in\_millis" : 6319  
> },  
> "ConcurrentMarkSweep" : {  
> "collection\_count" : 6,  
> "collection\_time" : "82 milliseconds",  
> "collection\_time\_in\_millis" : 82  
> }  
> }  
> }  
> },  
> "network" : {  
> "tcp" : {  
> "active\_opens" : 37,  
> "passive\_opens" : 2678,  
> "curr\_estab" : 36,  
> "in\_segs" : 574496,  
> "out\_segs" : 186058,  
> "retrans\_segs" : 20,  
> "estab\_resets" : 414,  
> "attempt\_fails" : 0,  
> "in\_errs" : 0,  
> "out\_rsts" : 3799  
> }  
> },  
> "transport" : {  
> "rx\_count" : 2,  
> "rx\_size" : "278b",  
> "rx\_size\_in\_bytes" : 278,  
> "tx\_count" : 2,  
> "tx\_size" : "26b",  
> "tx\_size\_in\_bytes" : 26  
> }  
> }  
> }  
> }

---

<div class="post-metadata">

**Author:** ![electic](https://avatars.discourse-cdn.com/v4/letter/e/ecccb3/32.png) [@electic](https://discuss.elastic.co/u/electic)\
**Post date:** [April 26, 2011, 4:51pm UTC](https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298/11 "2011-04-26T16:51:09Z")

</div>

root 4671 1 64 09:43 pts/0 00:04:32 /usr/bin/java -Xms256m  
-Xmx1g -Xss128k -Djline.enabled=true -XX:+UseParNewGC -XX:  
+UseConcMarkSweepGC -XX:+CMSParallelRemarkEnabled -XX:SurvivorRatio=8 -  
XX:MaxTenuringThreshold=1 -XX:CMSInitiatingOccupancyFraction=75 -XX:  
+UseCMSInitiatingOccupancyOnly -XX:+HeapDumpOnOutOfMemoryError -  
Delasticsearch -Des.path.home=/usr/local/elasticsearch -Des-pidfile=/  
var/run/elasticsearch.pid -cp :/usr/local/elasticsearch/lib/  
elasticsearch-0.16.0.jar:/usr/local/elasticsearch/lib/_:/usr/local/  
elasticsearch/lib/sigar/_ -Des.config=/etc/elasticsearch/  
elasticsearch.yml -Des.path.home=/usr/local/elasticsearch -  
Des.path.logs=/var/log/elasticsearch -Des.path.data=/var/lib/  
elasticsearch -Des.path.work=/tmp/elasticsearch  
org.elasticsearch.bootstrap.ElasticSearch  
electic 4824 4511 0 09:50 pts/0 00:00:00 grep --color=auto  
elasticsearch

The CPU died down. Not sure why last time it stayed at 300 percent.  
The only remaining issue is its not discovery it's nodes for some odd  
reason. Both boxes are fine and running, not sure after upgrade why I  
can't get them to talk.

---

<div class="post-metadata">

**Author:** ![Igor\_Motov1](https://avatars.discourse-cdn.com/v4/letter/i/c5a1d2/32.png) [@Igor\_Motov1](https://discuss.elastic.co/u/Igor_Motov1)\
**Post date:** [April 26, 2011, 4:54pm UTC](https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298/12 "2011-04-26T16:54:52Z")

</div>

There are several ways to increase -Xmx size. You find them in comments at  
the beginning of bin/elasticsearch file. For example, you can set ES\_MAX\_MEM  
environment variable or specify it in ES\_JAVA\_OPTS.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [April 26, 2011, 8:45pm UTC](https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298/13 "2011-04-26T20:45:16Z")

</div>

One thing that might explain it is the change to allow for 4 concurrent primaries recoveries on a node, which might strain the node too much with the current memory allocation. It used to be 2. You can control that by setting node\_initial\_primaries\_recoveries to 2.

If your node fails when trying to recover 4 shards concurrently, in any case I would say that its in a bad place memory wise, so, I would suggest increasing the memory to something like 2gb MIN and MAX.  
On Tuesday, April 26, 2011 at 7:54 PM, Igor Motov wrote:

> There are several ways to increase -Xmx size. You find them in comments at the beginning of bin/elasticsearch file. For example, you can set ES\_MAX\_MEM environment variable or specify it in ES\_JAVA\_OPTS.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:07am UTC](https://discuss.elastic.co/t/after-upgrade-to-16-problems/4298/14 "2017-07-06T04:07:36Z")

</div>


