Building a VPN Kill Switch That Fails Closed

A VPN is only as private as its worst second. If the tunnel drops for a moment and your traffic quietly falls back to the normal connection, the VPN has failed exactly when it mattered. That’s why a kill switch exists, and why building one for Umbra, my Windows VPN client, taught me more than the tunnel itself did.
Fail closed, not open
The core rule: if anything goes wrong, traffic stops. It never “just works” over the unprotected connection.
Umbra enforces that with the Windows Filtering Platform (WFP), the same filtering layer the Windows firewall is built on:
- Block everything in a dedicated Umbra sublayer.
- Open a single hole to the VPN server, just for the handshake.
- Once the tunnel is up, permit only the tunnel adapter, plus loopback, DHCP and, if you choose, your local network.
- Block IPv6 unless the tunnel carries it, a classic leak path.
The filters live in a static WFP session, so they don’t depend on the app staying alive. If the service crashes, the block stays in place and there’s no leak window.
There’s also a lockdown mode: keep blocking even after you disconnect, until you explicitly turn it off.
The day it locked me out
Failing closed has a dark side, and I found it the hard way.
A bad key or an unreachable server left the tunnel half-open: the adapter was up, but no handshake ever completed. The kill switch did exactly what it was built to do and blocked everything. The only way back online was to stop the service by hand.
The root cause was two design decisions interacting:
- The WFP filters intentionally outlive the process. That’s the no-leak-on-crash property.
- But nothing cleared stale filters on restart, and connect reported success before a handshake had happened.
The fixes
- Startup filter reconcile: when the service starts, it checks for leftover filters and reconciles them, so a crash can no longer strand you.
- Handshake verification at connect: connect only succeeds after a real handshake. A bad config now fails cleanly and restores your internet instead of trapping you.
- Service auto-restart: if the service crashes, it comes back and reconciles.
The lesson stuck with me: a fail-closed system needs an equally reliable way to recover, or it just becomes a very secure outage.
Other things the audit caught
Building a privileged Windows service meant auditing it like one:
- Imported OpenVPN configs could run programs as SYSTEM through
up,down,route-up,pluginandscript-security. Umbra now strips every program-executing directive and forcesscript-security 0. - The service’s named pipe accepted any local caller. It now requires medium integrity (sandboxed processes are rejected), prevents pipe squatting, and in release builds only accepts the installed Umbra UI.
- DLL planting: release builds load
wireguard.dllandopenvpn.exeonly from their absolute install paths, with no fallback to the current directory. - Secrets at rest are DPAPI-encrypted and bound to app-specific entropy.
Takeaways
- Decide your failure mode up front. For privacy tools, closed beats open.
- Make state outlive crashes, and then make startup clean up that state.
- Don’t report success early. Verify the thing actually worked, like a real handshake.
- Treat every input to a privileged process as hostile, even a config file the user imported themselves.