Sudden shutdown of all interfaces

Started by uhillebrand, Today at 12:31:31 PM

Previous topic - Next topic
Hi all, we had a severe outage at a customer's site a few days ago (OPNsense stopped responding on all interfaces). After going through the logs I am baffled how this could happen. Maybe someone can help finding the cause for this.

The firewall in question is running 26.7.1_1 on a Protectli VP2440.

There is a multi WAN-Setup:
- igc0 has a PPPoE connection to the primary ISP
- igc1 has an ethernet connection to the backup ISP

On the LAN side there is a lagg0 with 2 10G Interfaces (ixl0, ixl1); all internal traffic is routed via a number of VLANs on this bond.

This is how the problem developed:

1) At first, dpinger alerted that both ISPs had packet loss:

<165>1 2026-09-29T17:18:18+02:00 hwk-gw01.c.cibex.net dpinger 10740 - [meta sequenceId="1"] ALERT: WAN1_PPPOE (Addr: 8.8.8.8 Alarm: none -> loss RTT: 17.9 ms RTTd: 0.1 ms Loss: 12.0 %)
<165>1 2026-09-29T17:18:18+02:00 hwk-gw01.c.cibex.net dpinger 10740 - [meta sequenceId="2"] ALERT: WAN2GW (Addr: 1.1.1.1 Alarm: none -> loss RTT: 22.2 ms RTTd: 0.4 ms Loss: 12.0 %)
<165>1 2026-09-29T17:18:30+02:00 hwk-gw01.c.cibex.net dpinger 10740 - [meta sequenceId="3"] ALERT: WAN1_PPPOE (Addr: 8.8.8.8 Alarm: loss -> down RTT: 17.9 ms RTTd: 0.2 ms Loss: 33.0 %)
<165>1 2026-09-29T17:18:30+02:00 hwk-gw01.c.cibex.net dpinger 10740 - [meta sequenceId="4"] ALERT: WAN2GW (Addr: 1.1.1.1 Alarm: loss -> down RTT: 22.1 ms RTTd: 0.4 ms Loss: 33.0 %)


2) In system.log error messages regarding PPPoE started to show:

<29>1 2026-09-29T17:18:34+02:00 hwk-gw01.c.cibex.net ppp 35846 - [meta sequenceId="17"] [opt15_link0] LCP: no reply to 1 echo request(s)
[...]
<29>1 2026-09-29T17:19:04+02:00 hwk-gw01.c.cibex.net ppp 35846 - [meta sequenceId="20"] [opt15_link0] LCP: no reply to 4 echo request(s)
<29>1 2026-09-29T17:19:14+02:00 hwk-gw01.c.cibex.net ppp 35846 - [meta sequenceId="21"] [opt15_link0] LCP: no reply to 5 echo request(s)
<29>1 2026-09-29T17:19:14+02:00 hwk-gw01.c.cibex.net ppp 35846 - [meta sequenceId="22"] [opt15_link0] LCP: peer not respond ing to echo requests
<29>1 2026-09-29T17:19:14+02:00 hwk-gw01.c.cibex.net ppp 35846 - [meta sequenceId="23"] [opt15_link0] LCP: state change Opened --> Stopping


3) And now starts the part that is baffling me: a few seconds later, system.log reported that the LAN interface was detatched, and shut down both interfaces:

<13>1 2026-09-29T17:19:29+02:00 hwk-gw01.c.cibex.net opnsense 80063 - [meta sequenceId="57"] /usr/local/etc/rc.linkup: DEVD: Ethernet detached event for opt2(lagg0)
<13>1 2026-09-29T17:19:30+02:00 hwk-gw01.c.cibex.net kernel - - [meta sequenceId="58"] [3363137] ixl1: Interface stopped DISTRIBUTING, possible flapping
<13>1 2026-09-29T17:19:30+02:00 hwk-gw01.c.cibex.net kernel - - [meta sequenceId="59"] [3363137] ixl0: Interface stopped DISTRIBUTING, possible flapping
<13>1 2026-09-29T17:19:30+02:00 hwk-gw01.c.cibex.net kernel - - [meta sequenceId="60"] <6>[3363137] lagg0: link state changed to DOWN


-


Some time later the firewall was physically reset, and started working without problems again.

What kind of problem can cause a shutdown of *all* available interfaces in the matter of less than a minute? I checked CPU temp, RAM usage, state table, nothing out of the ordinary there. If anyone has an explanation or could tell me how to further debug this issue any input would be very welcome.

Thanks
Urban

Unpatched production firewall internet facing ? The cause could have been external

Could you give an example of what kind of missing bugfix could cause something like this? There is not a single service exposed to the internet, and only NTP and DNS from selected internal IPs. I do read the release notes of every new release. I don't tend to apply every update immediately, as these are rather frequent, and we did experience some issues in the past. So we developed a slower staging process.

If you feel this is not a safe way to run OPNsense and I am missing something important here, please let me know.

Thanks!
Urban