Menu

Show posts

This section allows you to view all posts made by this member. Note that you can only see posts made in areas you currently have access to.

Show posts Menu

Messages - meikel

#1
@nero355

It's an SAMSUNG EVO 850 250gb - it's used but when this issue happens again I'm fine with downtime. I may move to a raid1 at some point but it's not worth the hustle right now as there is no dedicated space for another ssd.

Once I knew what the issue was (very easy to find out if you don't blindfold yourself) it's quite easy to fix with a backup. Except the dhcp migration part.
#2
The memtest was fine:




So I think I found the culprit:


Well I guess it was the SSD all along. I didn't think a system drive failing would have such symptoms.

Sadly the story didn't end here:

* I tried to clone the drive onto a new drive which took way to long and didn't work at the end. The OS booted but couldn't find some files (don't ask me what exactly).
* So I took another road: Backup opnsense via UI and import in a freshly installed version. Should be straightforward, right? Wrong
* I didn't ever bothering migrating from "deprecated" features as long as they worked or I get the message "This will no longer work in the next version". It's just my homelab after all, if it's working it's good enough.
* Today I learned ISC DCHP is so hard deprecated that it is not even backed up when you create a backup. It backups the config values but does not backup (now a plugin) ISC from what I could tell.
* I booted the original drive (and hoped it would survive as long as necessary (spoiler: it did not - had to restart multiple times) and migrated everything to Kea according to this guide
* Now everything is working as before and I hope the device survives longer this time

Afterthought: Shouldn't opnsense warn me about bad smart values? Is there any way to enable this?

Thanks for everyone helping me, bringing ideas to the table and for your valuable time.
#3
Quote from: Patrick M. Hausen on June 03, 2026, 04:36:36 PMDoes your device have a serial console? If yes, connect a PC with a terminal program and let that run until the next crash. The serial console output will not be cleared on reboot unlike VGA/HDMI.

I just ordered a serial cable and will try that. I noticed in the flashy screen that there were a lot of hex codes so I ran me test.

Sadly both ram sticks had no issues after 30min.

I tried running a live debian13 and the device hung after some seconds on the desktop.

I improved cooling (airflow) even though I don't think this is the issue (the issue persists even after the cooling body feels way cooler now).

I also replaced the CMOS battery as a friend suggested this can be the root cause of such issues but no luck here either.

I will wait for the serial cable and try to debug that way. In the meantime I might try to update the bios of the mainboard.

I also have to correct myself: the SSD is not running in raid1. It's a single SSD but as a live OS is also crashing I assume it's not the SSD. I also doubt that a faulty SSD crashes the whole machine.
#4
So I conneted a display and withnessed a crash today. It actually crashes. When the crash happens opnsense beeps first and floods the UI with log messages (maybe kernel messages). After that the device reboots itself and is stuck in:

"Reboot and Select proper Boot device or Insert Boot Media in selected Boot device and press a key"

Is this an indication that the drive is faulty or doesn't this proof anything?
As this is my gateway to the internet a longly test should be my last resort right now.
#5
I doubt it's the SSD(s) as it's a raid1 (zfs)
#6
26.1, 26,4 Series / Opnsense randomly (?) crashes
June 02, 2026, 09:06:59 AM
So I've been using Opnsense for quite some time now and am very pleased with it.

However just Yesterday morning I went into my home office and noticed that I have no internet. After a short troubleshoot I found out that OPNsense is powered up and running but I get no IP or anything from it, I can't ping, ssh into it or get to the Web UI. I just quickly hard rebooted it and the issue was solved. Until today where this issue appeared again. I solved it quickly the same way as Yesterday however the issue just reappeared just about an hour later after the first reboot.

I'm unable to diagnose this issue. The logs give no information about what could have happened:

<13>1 2026-06-02T07:39:05+02:00 OPNsense.intern opnsense 61182 - [meta sequenceId="70"] /usr/local/etc/rc.linkup: plugins_configure ipsec (,lan)
<13>1 2026-06-02T07:39:05+02:00 OPNsense.intern opnsense 61182 - [meta sequenceId="71"] /usr/local/etc/rc.linkup: plugins_configure ipsec (execute task : ipsec_configure_do(,lan))
<13>1 2026-06-02T07:39:05+02:00 OPNsense.intern opnsense 61182 - [meta sequenceId="72"] /usr/local/etc/rc.linkup: plugins_configure dhcp ()
<13>1 2026-06-02T07:39:05+02:00 OPNsense.intern opnsense 61182 - [meta sequenceId="73"] /usr/local/etc/rc.linkup: plugins_configure dhcp (execute task : dhcpd_dhcp_configure())
<13>1 2026-06-02T07:39:06+02:00 OPNsense.intern opnsense 61182 - [meta sequenceId="74"] /usr/local/etc/rc.linkup: plugins_configure dhcp (execute task : radvd_configure_dhcp())
<13>1 2026-06-02T07:39:06+02:00 OPNsense.intern opnsense 61182 - [meta sequenceId="75"] /usr/local/etc/rc.linkup: plugins_configure dns ()
<13>1 2026-06-02T07:39:06+02:00 OPNsense.intern opnsense 61182 - [meta sequenceId="76"] /usr/local/etc/rc.linkup: plugins_configure dns (execute task : dnsmasq_configure_do())
<13>1 2026-06-02T07:39:06+02:00 OPNsense.intern opnsense 61182 - [meta sequenceId="77"] /usr/local/etc/rc.linkup: plugins_configure dns (execute task : unbound_configure_do())
<12>1 2026-06-02T07:39:06+02:00 OPNsense.intern opnsense 61182 - [meta sequenceId="78"] /usr/local/etc/rc.linkup: warning: ignoring missing default tunable request: vm.pmap.pti
<12>1 2026-06-02T07:39:06+02:00 OPNsense.intern opnsense 61182 - [meta sequenceId="79"] /usr/local/etc/rc.linkup: warning: ignoring missing default tunable request: hw.ibrs_disable
<13>1 2026-06-02T07:39:07+02:00 OPNsense.intern opnsense 61182 - [meta sequenceId="80"] /usr/local/etc/rc.linkup: plugins_configure newwanip:rfc2136 (,[lan])
<13>1 2026-06-02T07:39:29+02:00 OPNsense.intern kernel - - [meta sequenceId="81"] <6>[102] igb1: promiscuous mode enabled
<45>1 2026-06-02T08:39:09+02:00 OPNsense.intern syslog-ng 20363 - [meta sequenceId="1"] syslog-ng starting up; version='4.11.0'
<13>1 2026-06-02T08:39:09+02:00 OPNsense.intern kernel - - [meta sequenceId="2"] ---<<BOOT>>---
<13>1 2026-06-02T08:39:09+02:00 OPNsense.intern kernel - - [meta sequenceId="3"] Copyright (c) 1992-2023 The FreeBSD Project.
<13>1 2026-06-02T08:39:09+02:00 OPNsense.intern kernel - - [meta sequenceId="4"] Copyright (c) 1979, 1980, 1983, 1986, 1988, 1989, 1991, 1992, 1993, 1994
<13>1 2026-06-02T08:39:09+02:00 OPNsense.intern kernel - - [meta sequenceId="5"]        The Regents of the University of California. All rights reserved.
<13>1 2026-06-02T08:39:09+02:00 OPNsense.intern kernel - - [meta sequenceId="6"] FreeBSD is a registered trademark of The FreeBSD Foundation.
<13>1 2026-06-02T08:39:09+02:00 OPNsense.intern kernel - - [meta sequenceId="7"] FreeBSD 14.3-RELEASE-p12 stable/26.1-n272089-81f87c4d694c SMP amd64

I rebooted the device at about 07:39 and 08:39 so the logs from 07:xx are just the boot logs, no more logs after that.

Are there any other logs I can look into? I used opnsense-log to look at these logs.

I also was in the room once the device became faulty (It threw me out of my remote connection to work) and I noticed that Opnsense also did a beep (I don't know which kind of beep it is).

I saw some users suggesting the DIMM could be overheating but before I replace that I'd like to verify that these are in fact the issue. Even though it's somewhat summer here in Germany right now it's not really as hot in the room where the hardware is and it hasn't been a problem in the last year.

My Opnsense Version: OPNsense 26.1.8_5-amd64
My Hardware: Sophos SG 310 Rev.1