OPNsense Forum

English Forums => Hardware and Performance => Topic started by: aru.persia on July 09, 2026, 11:15:04 PM

Title: Severe Boot Delays & Recurring Network Outages Due to NVMe Timeouts (Heat-Relate
Post by: aru.persia on July 09, 2026, 11:15:04 PM


OPNsense Support Request – Severe Boot Delays & Recurring Network Outages Due to NVMe Timeouts (Heat‑Related)



PROBLEM SUMMARY

My OPNsense firewall (version 26.1) takes approximately 20 minutes to complete the boot process.

While the system does eventually become fully operational, this delay occurs reliably every two weeks. When the issue strikes, the firewall becomes unresponsive during this period, and internal DHCP clients (devices at home) are unable to obtain IP addresses – effectively taking the entire local network offline until the boot sequence fully recovers.

Critical physical observation: The device chassis becomes extremely hot – hot enough to "cook eggs on it." If I shut down the system, wait a short while for it to physically cool down, and then power it back on, the 20‑minute boot delay reliably manifests. The root cause appears to be a thermal‑related lockup of the NVMe storage controller during the cold‑start initialization phase.



SYSTEM INFORMATION

Operating System
OPNsense 26.1
FreeBSD 14.3-RELEASE-p15
stable/26.1-n272132-0101c21cac77
amd64

Hardware 



DETAILED SYMPTOM TIMELINE & RECURRENCE PATTERN




STORAGE DEVICE




ZFS HEALTH CHECK (performed while cool/stable)

zpool status -v
pool: zroot
state: ONLINE
config:
  NAME        STATE     READ WRITE CKSUM
  zroot       ONLINE       0     0     0
    nda0p3    ONLINE       0     0     0
errors: No known data errors

A scrub completed successfully with 0 errors. The filesystem integrity is sound.



SMART / NVMe HEALTH INFORMATION (Critical Thermal Data)

smartctl -x /dev/nvme0
SMART overall-health self-assessment test result: PASSED
Critical Warning:                   0x00
Available Spare:                    100%
Percentage Used:                    26%
Data Units Written:                 32.0 TB
Media and Data Integrity Errors:    0
Error Information Log Entries:      0

Temperature Readings (Excerpt)
Initial reading (idle/operational):

Later reading (stabilized):

QuoteNote: The sustained high temperature (Sensor 1 hitting 79°C) strongly suggests the controller is experiencing thermal throttling or outright thermal lockups during initialization.



BOOT PROBLEM DETAILS & KERNEL LOG EVIDENCE

During a cold boot following a heat‑soak period, the system pauses extensively. The relevant logs show a relentless NVMe reset loop:

nvme0: Resetting controller due to a timeout.
nvme0: resetting controller
nvme0: Waiting for reset to complete
nvme0: Waiting for reset to complete
nvme0: Waiting for reset to complete
...
# Eventually, after ~20 minutes:
nvme0: resubmitting queued i/o
nvme0: READ sqid:1 cid:0 nsid:1
nvme0: WRITE sqid:3 cid:0 nsid:1

Logs confirming recurrence:
/var/log/system/latest.log:
2026-07-09T17:29:35 nvme0: Resetting controller due to a timeout.
2026-07-09T17:38:03 nvme0: Resetting controller due to a timeout.
2026-07-09T20:10:22 nvme0: Resetting controller due to a timeout.

Additional Warnings (Cosmetic)
BOOT LOADER IS TOO OLD. PLEASE UPGRADE.
Please set hw.efifb.address and hw.efifb.stride.
These are noted but considered cosmetic/unrelated to the 20‑minute delay.



ACTIONS ALREADY PERFORMED




QUESTIONS FOR OPNsense SUPPORT

Given the clear thermal component and the bi‑weekly failure cycle, we need targeted guidance:




CURRENT ASSESSMENT


QuoteRecommendation needed: Please advise whether this is a software‑configuration fix (power management/PCIe linkspeed tuning) or if we are looking at a critical hardware failure requiring immediate SSD replacement and improved thermal management.


Title: Re: Severe Boot Delays & Recurring Network Outages Due to NVMe Timeouts (Heat-Relate
Post by: nero355 on July 09, 2026, 11:40:58 PM
Quote from: aru.persia on July 09, 2026, 11:15:04 PMCritical physical observation: The device chassis becomes extremely hot – hot enough to "cook eggs on it." If I shut down the system, wait a short while for it to physically cool down, and then power it back on, the 20‑minute boot delay reliably manifests. The root cause appears to be a thermal‑related lockup of the NVMe storage controller during the cold‑start initialization phase.

Temperature Readings (Excerpt)
Initial reading (idle/operational):
  • Temperature: 66°C
  • Temperature Sensor 1: 79°C ⚠️ (This is dangerously high for NVMe and likely the source of the timeout.)

Later reading (stabilized):
  • Temperature: 56°C
  • Sensor 1: 67°C
Have a look at this topic : https://forum.opnsense.org/index.php?topic=51263.0
You might have to do the same mod in order to keep the NVMe a bit cooler :)

Quotesmartctl -x /dev/nvme0
SMART overall-health self-assessment test result: PASSED
That's nice and all, but please post the full output for all S.M.A.R.T. data so we can double check for you ;)
Title: Re: Severe Boot Delays & Recurring Network Outages Due to NVMe Timeouts (Heat-Relate
Post by: pfry on July 10, 2026, 12:31:12 AM
Quote from: aru.persia on July 09, 2026, 11:15:04 PM[...]Temperature Sensor 1: 79°C ⚠️ (This is dangerously high for NVMe and likely the source of the timeout.)[...]

It's high, but not unusual - most SSD controllers have limits in the 75-85C range. They draw 2-20W, with M.2 devices generally limited to 6-8W. Most NVME devices will be right up at that limit. They have (generally) very effective thermal throttling, which you appear to be seeing. Yours could just be a hot runner (which would arguably be a hardware fault). Out of curiosity, what's the ambient temperature?

I'm not defending use of M.2 SSDs with no/minimal thermal management, by the way.
Title: Re: Severe Boot Delays & Recurring Network Outages Due to NVMe Timeouts (Heat-Relate
Post by: aru.persia on July 10, 2026, 09:48:53 AM
Thanks for taking a look nero355!

root@OPNsense:~ # date
Fri Jul 10 09:45:45 CEST 2026
root@OPNsense:~ # smartctl -x /dev/nvme0
smartctl 7.5 2025-04-30 r5714 [FreeBSD 14.3-RELEASE-p15 amd64] (local build)
Copyright (C) 2002-25, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Model Number:                       TS256GMTE710T
Serial Number:                      H532690002
Firmware Version:                   82B0U9MP
PCI Vendor/Subsystem ID:            0x1d79
IEEE OUI Identifier:                0x48357c
Controller ID:                      0
NVMe Version:                       1.4
Number of Namespaces:               1
Namespace 1 Size/Capacity:          256,060,514,304 [256 GB]
Namespace 1 Utilization:            246,970,736,640 [246 GB]
Namespace 1 Formatted LBA Size:     512
Namespace 1 IEEE EUI-64:            7c3548 51fc968452
Local Time is:                      Fri Jul 10 09:45:47 2026 CEST
Firmware Updates (0x14):            2 Slots, no Reset required
Optional Admin Commands (0x0017):   Security Format Frmw_DL Self_Test
Optional NVM Commands (0x005f):     Comp Wr_Unc DS_Mngmt Wr_Zero Sav/Sel_Feat Timestmp
Log Page Attributes (0x0f):         S/H_per_NS Cmd_Eff_Lg Ext_Get_Lg Telmtry_Lg
Maximum Data Transfer Size:         64 Pages
Warning  Comp. Temp. Threshold:     85 Celsius
Critical Comp. Temp. Threshold:     90 Celsius

Supported Power States
St Op     Max   Active     Idle   RL RT WL WT  Ent_Lat  Ex_Lat
 0 +     9.00W       -        -    0  0  0  0        0       0

Supported LBA Sizes (NSID 0x1)
Id Fmt  Data  Metadt  Rel_Perf
 0 +     512       0         0

=== START OF SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

SMART/Health Information (NVMe Log 0x02, NSID 0xffffffff)
Critical Warning:                   0x00
Temperature:                        54 Celsius
Available Spare:                    100%
Available Spare Threshold:          10%
Percentage Used:                    26%
Data Units Read:                    111,316 [56.9 GB]
Data Units Written:                 62,633,447 [32.0 TB]
Host Read Commands:                 2,665,357
Host Write Commands:                1,007,349,067
Controller Busy Time:               3,068
Power Cycles:                       16
Power On Hours:                     28,706
Unsafe Shutdowns:                   11
Media and Data Integrity Errors:    0
Error Information Log Entries:      0
Warning  Comp. Temperature Time:    6
Critical Comp. Temperature Time:    0
Temperature Sensor 1:               65 Celsius
Temperature Sensor 2:               41 Celsius
Temperature Sensor 3:               54 Celsius
Thermal Temp. 1 Transition Count:   1844
Thermal Temp. 2 Transition Count:   537
Thermal Temp. 1 Total Time:         3954648
Thermal Temp. 2 Total Time:         14109868

Error Information (NVMe Log 0x01, 16 of 256 entries)
No Errors Logged

Self-test Log (NVMe Log 0x06, NSID 0xffffffff)
Self-test status: No self-test in progress
No Self-tests Logged


Thanks again for double-checking this for me—I appreciate the help!
Title: Re: Severe Boot Delays & Recurring Network Outages Due to NVMe Timeouts (Heat-Relate
Post by: Patrick M. Hausen on July 10, 2026, 10:30:43 AM
I have a DEC750, too, same brand and model of SSD. Since the unit is passively cooled and there is not much convection in the cupboard in my study where I keep my network infrastructure, I placed one of these on top of the firewall:

https://www.amazon.de/dp/B08QYY87XW (the "80mm-1" style)

Really quiet, the rubber feet keep it in place nicely, powered from the DEC750's own USB port. Blowing upward so (I hope) air is getting sucked in and over the upper part of the case from the sides.
Title: Re: Severe Boot Delays & Recurring Network Outages Due to NVMe Timeouts (Heat-Relate
Post by: aru.persia on July 10, 2026, 12:54:40 PM
Quote from: pfry on July 10, 2026, 12:31:12 AM
Quote from: aru.persia on July 09, 2026, 11:15:04 PM[...]Temperature Sensor 1: 79°C ⚠️ (This is dangerously high for NVMe and likely the source of the timeout.)[...]

It's high, but not unusual - most SSD controllers have limits in the 75-85C range. They draw 2-20W, with M.2 devices generally limited to 6-8W. Most NVME devices will be right up at that limit. They have (generally) very effective thermal throttling, which you appear to be seeing. Yours could just be a hot runner (which would arguably be a hardware fault). Out of curiosity, what's the ambient temperature?

I'm not defending use of M.2 SSDs with no/minimal thermal management, by the way.

Ambient temperature is around 25–30°C. The SSD is a Transcend MTE710T 256GB NVMe (TS256GMTE710T). It is installed inside a rack.

The drive has reached over 90°C before, which is why I was looking into the thermal behavior. Current SMART data shows 54°C composite temperature and 65°C on sensor 1.
Title: Re: Severe Boot Delays & Recurring Network Outages Due to NVMe Timeouts (Heat-Relate
Post by: nero355 on July 10, 2026, 04:49:24 PM
Quote from: aru.persia on July 10, 2026, 09:48:53 AMThanks for taking a look nero355!

root@OPNsense:~ # smartctl -x /dev/nvme0Thanks again for double-checking this for me—I appreciate the help!
Why the -x option ?!

I usually just use -a to show as much data as possible and in case of some RAID controllers there is another option that's needed to look into the HDD/SSD via the controller, but that's not the case here luckily.

Maybe try also the 'nvmecontrol' command to see if there is any difference, but so far it seems to be mostly a heat thing...[/code]
Title: Re: Severe Boot Delays & Recurring Network Outages Due to NVMe Timeouts (Heat-Relate
Post by: pfry on July 10, 2026, 05:49:21 PM
Quote from: aru.persia on July 10, 2026, 09:48:53 AM[...]Warning  Comp. Temp. Threshold:    85 Celsius[...]

Hm. If the SSD controller is not running at or very near this temperature it should not throttle. Given the tiny thermal mass of the chip, the temperature change will be rapid - I don't know if you can get a diagnostic reading while it is... having difficulties. Deliberately testing the thing with a write benchmark ablates its endurance; the 256GB MTE710T is pretty good in that respect - 400GB/d (over the 3y warranty period). I wouldn't expect it to throttle under a read load, but then booting shouldn't be an issue.
Title: Re: Severe Boot Delays & Recurring Network Outages Due to NVMe Timeouts (Heat-Relate
Post by: BrandyWine on July 10, 2026, 08:17:45 PM
When the system is booting, is this from room ambient, or is the thing already toasty when you reboot it?


This however only works when the OS is running. So depends where in boot stage it is.
nvmecontrol logpage -p 2 nvme0



    Poll it repeatedly (example: once per second, modify as needed as print command only prints the value, etc):

sh

while true; do
  nvmecontrol logpage -p 2 nvme0 | awk -F': ' '/Temperature:/{print $2; exit}'
  [ do some checks and actions here ]
  sleep 1
done
Title: Re: Severe Boot Delays & Recurring Network Outages Due to NVMe Timeouts (Heat-Relate
Post by: aru.persia on July 16, 2026, 08:36:35 PM
Quote from: Patrick M. Hausen on July 10, 2026, 10:30:43 AMI have a DEC750, too, same brand and model of SSD. Since the unit is passively cooled and there is not much convection in the cupboard in my study where I keep my network infrastructure, I placed one of these on top of the firewall:

https://www.amazon.de/dp/B08QYY87XW (the "80mm-1" style)

Really quiet, the rubber feet keep it in place nicely, powered from the DEC750's own USB port. Blowing upward so (I hope) air is getting sucked in and over the upper part of the case from the sides.

Hi Patrick,

Just wanted to let you know that your solution worked for me as well! I picked up the same fan, powered it from the DEC750's USB port as you suggested, and it's made a noticeable difference.

Thank you so much for sharing your setup and the recommendation — I really appreciate it!



Before:
Temperature Sensor 1:               88 Celsius
Temperature Sensor 2:               61 Celsius
Temperature Sensor 3:               75 Celsius

After:
Temperature Sensor 1:               66 Celsius
Temperature Sensor 2:               41 Celsius
Temperature Sensor 3:               55 Celsius