Sudden shutdown of all interfaces

Started by uhillebrand, October 01, 2026, 12:31:31 PM

Previous topic - Next topic
Hi all, we had a severe outage at a customer's site a few days ago (OPNsense stopped responding on all interfaces). After going through the logs I am baffled how this could happen. Maybe someone can help finding the cause for this.

The firewall in question is running 26.7.1_1 on a Protectli VP2440.

There is a multi WAN-Setup:
- igc0 has a PPPoE connection to the primary ISP
- igc1 has an ethernet connection to the backup ISP

On the LAN side there is a lagg0 with 2 10G Interfaces (ixl0, ixl1); all internal traffic is routed via a number of VLANs on this bond.

This is how the problem developed:

1) At first, dpinger alerted that both ISPs had packet loss:

<165>1 2026-09-29T17:18:18+02:00 hwk-gw01.c.cibex.net dpinger 10740 - [meta sequenceId="1"] ALERT: WAN1_PPPOE (Addr: 8.8.8.8 Alarm: none -> loss RTT: 17.9 ms RTTd: 0.1 ms Loss: 12.0 %)
<165>1 2026-09-29T17:18:18+02:00 hwk-gw01.c.cibex.net dpinger 10740 - [meta sequenceId="2"] ALERT: WAN2GW (Addr: 1.1.1.1 Alarm: none -> loss RTT: 22.2 ms RTTd: 0.4 ms Loss: 12.0 %)
<165>1 2026-09-29T17:18:30+02:00 hwk-gw01.c.cibex.net dpinger 10740 - [meta sequenceId="3"] ALERT: WAN1_PPPOE (Addr: 8.8.8.8 Alarm: loss -> down RTT: 17.9 ms RTTd: 0.2 ms Loss: 33.0 %)
<165>1 2026-09-29T17:18:30+02:00 hwk-gw01.c.cibex.net dpinger 10740 - [meta sequenceId="4"] ALERT: WAN2GW (Addr: 1.1.1.1 Alarm: loss -> down RTT: 22.1 ms RTTd: 0.4 ms Loss: 33.0 %)


2) In system.log error messages regarding PPPoE started to show:

<29>1 2026-09-29T17:18:34+02:00 hwk-gw01.c.cibex.net ppp 35846 - [meta sequenceId="17"] [opt15_link0] LCP: no reply to 1 echo request(s)
[...]
<29>1 2026-09-29T17:19:04+02:00 hwk-gw01.c.cibex.net ppp 35846 - [meta sequenceId="20"] [opt15_link0] LCP: no reply to 4 echo request(s)
<29>1 2026-09-29T17:19:14+02:00 hwk-gw01.c.cibex.net ppp 35846 - [meta sequenceId="21"] [opt15_link0] LCP: no reply to 5 echo request(s)
<29>1 2026-09-29T17:19:14+02:00 hwk-gw01.c.cibex.net ppp 35846 - [meta sequenceId="22"] [opt15_link0] LCP: peer not respond ing to echo requests
<29>1 2026-09-29T17:19:14+02:00 hwk-gw01.c.cibex.net ppp 35846 - [meta sequenceId="23"] [opt15_link0] LCP: state change Opened --> Stopping


3) And now starts the part that is baffling me: a few seconds later, system.log reported that the LAN interface was detatched, and shut down both interfaces:

<13>1 2026-09-29T17:19:29+02:00 hwk-gw01.c.cibex.net opnsense 80063 - [meta sequenceId="57"] /usr/local/etc/rc.linkup: DEVD: Ethernet detached event for opt2(lagg0)
<13>1 2026-09-29T17:19:30+02:00 hwk-gw01.c.cibex.net kernel - - [meta sequenceId="58"] [3363137] ixl1: Interface stopped DISTRIBUTING, possible flapping
<13>1 2026-09-29T17:19:30+02:00 hwk-gw01.c.cibex.net kernel - - [meta sequenceId="59"] [3363137] ixl0: Interface stopped DISTRIBUTING, possible flapping
<13>1 2026-09-29T17:19:30+02:00 hwk-gw01.c.cibex.net kernel - - [meta sequenceId="60"] <6>[3363137] lagg0: link state changed to DOWN


-


Some time later the firewall was physically reset, and started working without problems again.

What kind of problem can cause a shutdown of *all* available interfaces in the matter of less than a minute? I checked CPU temp, RAM usage, state table, nothing out of the ordinary there. If anyone has an explanation or could tell me how to further debug this issue any input would be very welcome.

Thanks
Urban

Unpatched production firewall internet facing ? The cause could have been external

Could you give an example of what kind of missing bugfix could cause something like this? There is not a single service exposed to the internet, and only NTP and DNS from selected internal IPs. I do read the release notes of every new release. I don't tend to apply every update immediately, as these are rather frequent, and we did experience some issues in the past. So we developed a slower staging process.

If you feel this is not a safe way to run OPNsense and I am missing something important here, please let me know.

Thanks!
Urban

Today at 12:00:12 AM #3 Last Edit: Today at 12:44:13 AM by (MARLOO)
Hi Urban,

What you're seeing looks very much like a driver/hardware issue on the 10G interfaces (ixl) that drags down the whole LAN side and, as a side effect, breaks WAN monitoring and PPPoE.

Key points from your logs:

ixl0/ixl1: Interface stopped DISTRIBUTING, possible flapping

lagg0: link state changed to DOWN

DEVD: Ethernet detached event for opt2(lagg0)

Those "possible flapping" messages are typical of the ixl driver when it detects abnormal conditions (errors, watchdog, DMA issues, etc.) and stops the interfaces. When lagg0 goes down, all VLANs on top of it disappear, dpinger loses its route, and PPPoE stops getting LCP echoes – so it looks like "everything died", even if the WAN PHYs are still up

----------------------------------------Things I'd check-----------------------------------------------------------------------------------------------------

Full system.log around the incident for ixl0/ixl1 errors: watchdog, TX/RX hang, DMA, no buffers.

ifconfig ixl0, ifconfig ixl1, ifconfig lagg0 and netstat -i for growing error counters.

Lagg type vs switch config (LACP/failover, active/passive, speed/duplex). If you don't need aggregation, try failover or a single 10G temporarily.

OPNsense forum for 26.7.x + ixl/lagg + "interface down" / "flapping"; in some cases, rolling back version fixed similar issues.

Basic hardware checks: different SFP+/DAC ports, cables, and (if available) disable PCIe power saving in BIOS..

If you can share a larger log excerpt (anonymized) and your lagg/VLAN config, it's easier to narrow down whether this is driver, switch negotiation, or hardware


-------------------------------------------------------------------------------------------------------------------------------------------------------------
These kinds of issues often show up only in specific hardware/driver/topology combinations (Protectli + Intel 10G + LACP + certain switches), so they may not be obvious in general testing and only surface in production setups like yours.

That doesn't mean you must run the absolute latest release everywhere. A pragmatic approach that many use:

Keep your current version on staging and monitor for similar symptoms.

If you see any 10G/ixl/lagg‑related fixes in later release notes, consider testing those on staging before deciding whether to move production.

For critical sites, maintain a known‑good version and only move after checking forums/issues for regressions on your hardware.
-------------------------------------------------------------------------------------------------------------------------------------------------------------

Check interface status and errors

ifconfig ixl0
ifconfig ixl1
ifconfig lagg0

Look for:

status: (active / no carrier)
****************************************************
Counters like input errors, output errors, collisions, dropped.

Then:

netstat -i

Check if ixl0/ixl1 show increasing error counters compared to other interfaces
****************************************************************
grep -i "ixl" /var/log/system.log | less
dmesg | grep -i "ixl"

Search for patterns like:

watchdog timeout

TX hang / RX hang

DMA error

no buffers

descriptor errors
****************************************

ifconfig lagg0

*****************************************
During normal operation

top -P
systat -if 1
*****************************************
maybe this will help your troubleshooting
Hardware: N5105 Intel Celeron  
                       OPNsense | Home Lab | Linux & Home Automation
                               "Secure the network, automate the rest."