not vendor specific. Impacting Inovelli, 3rd reality, leviton devices
I believe it is zigbee mesh related because the storm will start if I introduce a lot of traffic such as a zigbee broadcast or put hub in pairing mode. A switch that is not part of the zigbee mesh isn't effected.
I've powered off all of my devices room by room with no effect
I've pulled batteries from all battery powered devices with no effect
However, sometimes, powering off a device immediately stops the boot storm temporarily
I've left the hub powered off for 30 minutes to force a network re-build with no effect
Problem was intermittent, began ~1-2 months ago but has now become almost constant.
Changing from channel 20 - 25 appeared to correct the problem for about a month
Lights don't work! I can't even reset my light switches because they are constantly rebooting.
Currently running on zigbee channel 25. WiFi analyzer shows channel clear
I wish I had a zigbee sniffer...
Got any troubleshooting tips or has anybody even heard of this sort of thing happening? Need to figure this out, I can't get any lights to work at night and since I can't reset devices, there's no way to correct this!
Can you get into the Zigbee logs when this is happening? It seems most likely that a single Zigbee device has gone rogue. I had this happen years ago with a Peanut plug that was flooding the network. I removed it and all was well. It had sent around 75,000 messages in just a few hours.
This rules out this being a hub issue, i would focus on the devices and not the hub for this issue.
When you say “all of your zigbee devices are rebooting” considering you can detect this with the hub powered off… what are the symptoms that lead you to your zigbee devices are rebooting?
This is a good symptom, keep it powered off. Keep powering off devices until the symptoms stop for 7 days - a month. This could be a multi device issue..
Once you have stability, reintroduce 1 device at a time. give it a few days to a few weeks in between adding devices back. when the issues return, power off the last device added and confirm stability before moving on to the next one.
"all zigbee devices are rebooting, how do you know this?"
With Inovelli Switches, I see them all doing the colorbar startup sequence. With leviton and Sinope devices, I see the green light extinguish then light up. With all devices, the load extinguishes then lights back up. This does not happen to devices that are not part of the zigbee mesh so it rules out an actual power related issue
More troubleshooting data:
I've been using the breaker box to power off regions of the house and this problem appears to be related to the number of mesh connected zigbee devices powered up. There is no one region that definitivley makes the problem go away or re-appear. However, the more devices I have powered up, the more frequently the boot storm occurs.
There is nothing unusual that I can see in the zigbee logs. Its as if the system becomes sensitive to elevated traffic levels once the mesh grows large enough. So things like doing a switch LED color change broadcast or putting the mesh into pairing mode will initiate the boot storm if I have enough circuit breakers closed.
Here's the most maddening part: It seems like I can isolate the problem to a given circuit; turning on the breaker starts up the problem and vice versa. But then the problem shifts to another circuit and I can stop the storm by de-energizing its breaker then start it up again once the circuit is re-energized.
I have 76 zigbee devices. Does that seem like an excessive number for one hub?
More interesting data:
I switched back to zigbee channel 20 and the mesh was stable until about 60 of my 76 devices had found the new channel. Then the boot storm started up again. So supporting the hypothesis that it is somehow the number of devices in the mesh that is behind the problem.
When I had a similar issue, I was able to isolate it to an Inovelli On/Off Zigbee switch.
I think that I reset the switch to resolve the issue. So I would say that if you think you have the culprit, reset the device, then join it again, while leaving it on the hub. It will take back it’s old slot.
I have 207 real Zigbee devices, over 70 virtual devices, and a handful of Z-wave devices on my main hub. Has the WiFi network environment you are in (yours and your neighbors') changed recently?
Not that I can see based on the last analysis I did back when the problem originally occurred. That's what convinced me to move to channel 25 from channel 20. Everything looks about the same now from a WiFi perspective. Nevertheless, I'm moving to channel 26 now. It looks like my devices all support it; seeing reports from all the manufacturers that I use. I have little hope of success here though.
You would think that a device would show up in the circuit breaker based troubleshooting I was doing though. Meaning that the same breaker would always control the problem. In my case, the problem isolates itself to different breakers meaning that when I think I have the problem isolated (toggling a given breaker toggles the problem), the problem itself shifts to another circuit during the time I'm spending isolating individual devices on the original circuit.
I sympathize! Several years ago I had an Aqara battery operated device that apparently went insane and crashed my Zigbee network. Took days to track it down and kill it.
So is this "boot storm" effect typical of a crashed network? It is crippling the house! Not even any manual control over light switches available with lights flashing on and off randomly everywhere.
Take all battery-powered sensors offline for 48 hours by removing their batteries (disabling them alone will not cause their radios to stop broadcasting).
Pick a Zigbee channel and stick with it; changing channels makes all devices start to seek the new channel.
For everything except the Inovelli switches/dimmers, disconnect from power.
For the Inovelli switches/dimmers (you said Zigbee, so they should be Blue), pull all of the local power interrupt tabs and then slowly add them back. You should be able to regain local control of your lights.
Then add the mains powered devices one by one.
Then add the battery-powered devices one by one, but only after the 48 hr period expires.
Well, it looks like its down to one of my four Sinope baseboard thermostats. With the heater circuit breaker off, the mesh has been completely stable. I must say that zigbee troubleshooting is certainly hit or miss. The "divide and conquor" troubleshooting method does not work very well if at all. Also not a fan of what a "crashed network" looks like... Would a zigbee sniffer make this simpler? Trying to see if its worth the somewhat considerable effort to set one up and then actually use it to get unencrypted information off the mesh.
In retrospect...
Most important to me is the fact that my zigbee mesh doesn't "fail safe". Every other protocol I've worked with still provided manual switch control when the protocol was in a failed state. Not the case here. If something were to happen to me and this sort of failure occurred, the house and its occupants would be completely crippled with no ability to use any electrical lighting and (even worse) with the lighting system itself exhibiting completely random and unstable light on/off behaviors. The re-boot on network congestion condition is also not vendor specific as I observed empirically. Ironically, the only devices that didn't exhibit this condition were the zigbee thermostats, one of which was itself the source of the problem.
Ultimately it looks like the solution may be more hubs creating smaller meshes as @Ken_Fraleigh suggested. It certainly looks like the boot condition didn't occur until ~50 devices were active.
Yes - A sniffer would definitely help to see large broadcasts of traffic from a single device. - It's always good to be able to see what's on the wire, when troubleshooting. - And if the root case is a battery device, it can clearly take days to isolate it.
As for "fail safe", ZB isn't the only protocol to behave that way. - ZW and TCP/IP also can behave that way, specifically with broadcast storms - For OT devices with poor IP stacks, broadcast storms on larger networks with a single bad device, can easily crash an OS, and effect whatever that device was responsible for (I've personally taken down a production line in a LCD factory, due to running slow and gentle IP vulnerability scans (nmap) on a manufacturing network and "bumping/crashing" an VERY old PLC). In the TCP/IP case, VLANs (or more accurately network segmentation) is your friend, to lower the blast radius, and lower the likely-hood of a broadcast storm in the first place.
In HE land (or in your case ZB), you don't have many tools to work with (aka VLANs), hence the recommendation to break the mesh up. This carries it's own risk, in terms of interdependencies across hubs, and a reliance on more points of failure (hubs). If you have ZB or ZW on critical circuits, running critical devices (lighting, HVAC, food storage, etc.), then those items are obviously at risk.
This is why I'm presonally a fan of Matter, as much of it is IPV6 based, and you have more "real networking tools" to work with. That all said, segementing and bridging OT networks (which use alot of mDNS) brings it's own levels of complexity to troubleshoot.
You "pays your money, and you takes your choice"... But there are tradeoffs on all of these approaches. If/when I die before my wife, I expected all the Home Automation devices to be likely removed by a paid electrician (good luck to them finding all the devices).
Yeah, literally thinking the same thing here. Or if reactive situation happens leave behind instructions to power off breakers (I essentially have 5 house lighting "zones" and one heating zone controlled by breakers) then go through reset process for devices still powered up. So power off 5, do the one that's on, power on another and do that etc until all four zones are re-energized/devices reset. Or choose proactive which would be to reset the devices one by one through the entire house the day after I get planted... Actually wondering if I can write a "doomsday script" to automate this... edit: nope there isn't a way. Inovelli for one requires physical interaction for a hard reset.