26.7.6: unbound (and syslog-ng) segfault with signal 11 after upgrade (Hyper-V)

Started by Nex VII, October 08, 2026, 08:38:22 PM

Previous topic - Next topic
Hi,
after upgrading from 26.7.5 to 26.7.6 today, unbound keeps crashing with signal 11, roughly every 6-15 minutes, both at startup and at runtime.

Environment:
- OPNsense 26.7.6 (amd64), VM on Hyper-V
- unbound 1.26.1, unchanged since 18/09 and stable until today
- Unbound with DNSSEC, prefetch, serve-expired, aggressive NSEC, forwarding (DoT); DHCP lease registration off; DNSBL off

Observed (kernel log):
- unbound: 14x "exited on signal 11" since the upgrade, e.g. 19:29, 19:35, 19:41, 19:53, 19:59, 20:10, 20:18
- syslog-ng: 2x signal 11 right after boot
- python3.13: 2x signal 6 (abort)
- No "exited on signal" entries at all in the logs of the previous 4 weeks (since 10/09)
- Unbound logs nothing before dying: the last lines are normal resolution activity. A reboot doesn't help.

Since several unrelated binaries started segfaulting right after the update, this looks like a regression in base/libs rather than in unbound itself. Is anyone else seeing this on 26.7.6? I can enable kern.coredump and provide a backtrace if useful.

Thanks

I have the feeling this is about https://github.com/opnsense/src/issues/297 which was supposed to be solved with 26.7.6, not cause more regressions on previously unaffected installs.

Can you try to confirm by downgrading the kernel?

# opnsense-update -zkr 26.7.4
(reboot)


Cheers,
Franco

Thanks Franco.

Host: Intel Xeon E-2236, Hyper-V Gen 2 VM with 4 vCPUs and 8 GB static RAM.

What I did:
- opnsense-revert -r 26.7.5 opnsense
- opnsense-update -bk -r 26.7.4
- reboot at 20:58

Current versions: core 26.7.5, base 26.7.4, kernel 26.7.4.

Result: since the reboot (about 1 hour so far) there have been no "exited on signal" entries at all. Before, on 26.7.6, unbound crashed repeatedly, the last time at 20:41.

I'll keep monitoring and will report back here if any crash happens again.

Cheers

Same problem on Hyper-V.
Other physical machines are ok, no Unbound crashes.

I let AI run over this, take with a grain of salt (could, would...)

Two potential issues in sys/dev/hyperv/vmbus/hyperv_mmu.c:

* In hv_vm_tlb_flush(), (end == 0 || ...) > max_gvas compares a Boolean against max_gvas, making the condition always false and preventing full TLB flushes.
* In hv_flush_tlb_others_ex(), fill_gva_list() receives end, start instead of start, end, potentially invalidating the wrong page.

Both could leave stale TLB entries and explain the crashes.
Hardware:
DEC740

@Nex VII

Can you lock the kernel at 26.7.4 and update the rest back to 26.7.6? It would help narrowing this down before building a kernel that has the relevant commit reverted.


Thanks,
Franco


Hi,

I'm seeing the same thing on Hyper-V, though less frequently than described here.

Environment:

OPNsense 26.7.6-amd64, FreeBSD 15.1-RELEASE-p4, OpenSSL 3.5.9
Virtualized on Hyper-V (host CPU: AMD EPYC 4584PX), VM with 8 GB RAM
Unbound 1.26.1 (since 18/09), Python module with dnsbl_module loaded
[Upgraded to 26.7.6 on: ...]

What I observe:

Today, 2026-10-09 04:15:53, kernel log:
pid 19272 (unbound), jid 0, uid 59: exited on signal 11 (no core dump - denied by kern.coredump)
The VM had rebooted at about 04:10 and Unbound started cleanly at 04:10:55. It crashed roughly 5 minutes later.
The resolver log shows nothing unusual before the crash. The last line is dnsbl_module: successfully opened pipe at 04:11:05, then silence until "Closing logger".
Earlier unclean death on 2026-10-04 at 03:01:31: the resolver log shows "Closing logger" without the usual "service stopped", and no restart afterwards. Unbound was apparently down until I rebooted on 2026-10-09 01:16, so about 5 days without DNS resolution from OPNsense. I can't confirm a signal 11 for that one, since my kernel log excerpt has no matching entry.
I restarted Unbound manually after noticing it was stopped on the dashboard.

I have not downgraded the kernel yet. Reverting to 26.7.5...

Thanks

I've built a test kernel with the mentioned commit:

# opnsense-update -zkr 26.7.6-hyperv


Cheers,
Franco

Updated to 26.7.6 again.
Run opnsense-update -zkr 26.7.6-hyperv

After 5min it seams ok (older kernel already crashed few processes after 5min)
Will keep eye on it little bit longer.

Good news!  Thanks for testing.  :)

We may have to move to 26.7.7 early next week to ship this fix for everyone.


Cheers,
Franco

That might have happened to me during the update as well.
Here is my post: https://forum.opnsense.org/index.php?topic=53121.0

I managed to spot an error entry in the Unbound logs.
Also in the logs via the console.

Do I understand this correctly: You updated everything to the latest version except Unbound?
Was it locked and left at the previous version?

This isn't about Unbound. It's about a Hyper-V patch in 26.7.6 that fixes a boot issue on som AMD architecture but apparently causes these userspace segfaults.  Reverting the kernel to 26.7.4 or trying the hyperv kernel mentioned above should help.... it would help more if you could also confirm the hyperv kernel working.


Cheers,
Franco

Quote from: franco on Today at 08:42:50 AMI've built a test kernel with the mentioned commit:

# opnsense-update -zkr 26.7.6-hyperv


I'm now using the kernel you updated!
It's running beautifully!
Many thanks for the quick fix!

Unbound is now remaining stable (for the time being); everything is running as it should!


Is there any way I can help out with logs or anything else?