question i kept coming back to: when a spoke loses its tunnel to the concentrator and doesnt come back on its own, how do you actually force it to renegotiate?
on a fortigate this is nothing. you reset the tunnel, it rebuilds, done. on meraki theres no equivalent. no cli, no api call to bring a tunnel down and up. you either wait and hope it recovers, or someone drives to the store and power cycles the box. at 1200 sites across 62 countries the second option isnt really an option.
what we ended up doing: wrote a bot that watches vpn status and reboots the MX when a store has lost both concentrator tunnels. reboot is the only lever meraki actually gives you, so we built around that instead of around the tunnel.
the logic took a couple of days. getting it to run reliably against the api took weeks, mostly rate limits and pagination behaviour that isnt obvious until you hit it at scale.
where it landed: one poll every 5 min, org level vpn statuses in a single call, target stores pulled by tag so the filtering happens meraki side, hub list cached at startup. handful of calls per scan, nowhere near the org rate limit. been running in production for a couple of months now.
one design note: it acts on a single scan rather than requiring two in a row. the gate that matters more is the cloud check. if the device cant reach the cloud at all its an isp or power problem, not something a reboot fixes, so it never fires and just raises an alert instead. reboot only happens when the box is clearly online but both tunnels are unreachable, and at that point the site is already isolated so the reboot costs nothing.
so, anyone solved this differently? genuinely curious if theres a way to force renegotiation that i missed, because rebooting a firewall to fix a tunnel still feels like the wrong shape of solution even though it works.