Menu

Show posts

This section allows you to view all posts made by this member. Note that you can only see posts made in areas you currently have access to.

Show posts Menu

Messages - meyergru

#1
Your latest tcpdump provides an important clue: Even in the broken state, DNS queries from 10.127.127.2 reach 10.127.127.1, and valid DNS responses are sent back through wg4.

This suggests that Unbound is working. The problem is more likely related to packet delivery, firewall rules, or routing/gateway selection during startup.

Since your setup involves multiple gateways and policy-based routing (PBR), I particularly suspect a routing or firewall state issue during boot, possibly involving incorrect gateway selection or asymmetric routing.

Could you perform the following tests before restarting WireGuard, while simultaneously running `tcpdump -ni wg4 'host 10.127.127.2'` on OPNsense?

1. Explicit DNS lookup: Query `10.127.127.1` directly from the affected client (e.g. `nslookup example.com 10.127.127.1`). If this fails despite DNS responses appearing in the capture, the problem is likely on the return path to the client or in the client's DNS stack. If it succeeds, DNS transport itself is working.

2. Direct IP connection: Try accessing a known reachable Internet server by IP address, bypassing DNS entirely. If this fails while DNS works, suspect forwarding, NAT, firewall rules, or gateway selection. The capture will show whether the connection attempt reaches wg4.

3. Connection by hostname: Access the same server by hostname. If direct IP access works but hostname access fails, investigate the client's DNS configuration or resolver behavior rather than general IP routing.

Given your multiple gateways and PBR rules, I would also compare the firewall states and gateway selection for the affected client traffic before and after restarting WireGuard.

These tests should distinguish DNS transport issues from forwarding/routing problems and incorrect client-side name resolution.

I would avoid introducing startup delays until we have identified the failing component.
#2
Generally, there are usually one of these problems (not all with all DHCP servers, but hey):

1. With some servers (at least ISC), the creation of a static reservation does not get active unless the service is restarted.
2. Some servers (also ISC), you should not have static reservations within dynamic ranges.
3. When you let a client get a dynamic reservation and convert it to a static one, the dynamic reservation is not deleted, such that a reboot still gives the old IP. You have to manually clean the dynamic reservation from the lease file (and restart the service to recognize that fact).
4. With newer services, a DUID ist being used instead of a MAC (unless you change that setting) - which may give the impression that static reservations (for the MAC) do not work.

That being said, I know that is not an exact answer, just because I happen to use Kea and not DNSmasq.
#3
German - Deutsch / Re: [Hardware] Temperaturen
October 08, 2026, 10:32:27 AM
NVME SSDs unterscheiden sich ziemlich in den Controllern. Bei Geizhals kannst Du sehen, wieviel die Modelle im Betrieb brauchen. Es gibt welche, die bei >7 Watt liegen, das ist mehr als ein N100. Bei der minimalen thermischen Trägheit und winzigen Abstrahlfläche müssen die also heiß werden, zudem die Hitze sich nicht gleichmäßig auf die Fläche verteilt.

Es gibt passive Kühlkörper, die man aufsetzen kann, die oft >10°C Abkühlung bringen.

Ich suche gute NVME-SSDs für ZFS danach aus, ob sie geringe Leistung brauchen, hohen TBW-Wert und möglichst DRAM-Cache haben. Solche Modelle sind dann leider auch etwas teurer.

Eine weitere Möglichkeit ist, die SSD mittel eines Kommandos in einen geringeren Power-State zu bringen - sie agiert dann langsamer, aber sparsamer. Man sieht die möglichen States mit "nvmecontrol power -l nvme0".
#4
Your time might be better spent by creating a PR for this: https://github.com/opnsense/core/issues/8239

While I generally like the idea, I personally refrain from applying non-vetted patches or packages to a security-relevant appliance like OPNsense.
#5
Aha. Found it:

https://github.com/flaviuvlaicu/opnsense-topo-map

I had a look at the code. The plugin does not actually discover the network topology. It merely collects clients from ARP/Kea; the relationships between clients, switches and APs are then defined manually by the user via drag & drop and stored in topology.json.

There is no LLDP, SNMP, switch MAC-table or controller data being queried. Therefore, the plugin cannot determine which switch/port a client is actually connected through, nor can it detect when a wireless client roams from one AP to another.

So essentially, it is a graphical network diagram editor with an automatically populated client list, not an automatic topology discovery tool.

That said, I would be rather cautious about installing software from arbitrary GitHub repositories on an OPNsense firewall.
#7
OpnSense does not know anything about this other than on which interface a client is connected to, so what you want is not feasible using it.

There is network equipment available that can show you things like that, however. Take a look at Unifi switches and APs. My personal opinion is that they are quite good at making those, less good at making routers (I use OpnSense for that).

You do not need a dream box or any of their routers, though. The network controller can be run on x64 hardware or as a VM (the software is called Unifi OS server and is free).
#8
That is a collision of interest of the purest kind.

You can block DoT/DoH, but obviously, you cannot intercept and modify that traffic, because it is encrypted and protected via certificates. Normally, such an approach would be used in conjunction with a traffic interception of normal, unencrypted DNS traffic to be able to block certain DNS domains. If anyone wants to use your network, he must fall back to standard DNS in such a scenario, i.e. abide by your imposed rules.

If you want to give a device internet access that does not play by those rules, you can, but then you cannot block DoT/DoH, as you have seen.
In that case, I would put that device into a separate VLAN (probably also a separate WLAN), such that it can do whatever it chooses, but has no possibility to access any of your own LAN ressources. Of course, you can allow specific services, like accessing a printer.

Or to put it short: You cannot have the cake and eat it, too.

BTW: If there was a canary domain, you would be able to see a DNS request (and, obviously, it would be via normal DNS).

Since Chromium deliberately chose not to implement a Firefox-style canary domain, there is no DNS-side mechanism to force such a fallback; and if the Chromebook is managed by the school, the DoH-only setting is most likely policy-locked anyway.

#9
26.7 Series / Re: Pppoe Interfaces Linkage
September 26, 2026, 01:20:46 PM
...because the documentation knows better than you if your ISP needs a VLAN and if your modem supplies it in that case?

There are two cases:

1. your ISP does not need a VLAN - then your use the NIC device without one.
2. your ISP needs a VLAN - then you can do one of two things:
   a. if your modem supports it, you can set the VLAN in the modem and your NIC device without VLAN.
   b. if your modem does not support it or you do not want to lose modem access via the untagged VLAN, you can use the NIC device with VLAN.


I prefer 2b if needed for the reason given.
#10
26.1, 26,4 Series / Re: Legacy FTP ON 26.X BE
September 25, 2026, 10:53:04 PM
You only need that for active FTP, which I assumed why the question was asked. These days, most FTP servers can do passive FTP.
#11
26.1, 26,4 Series / Re: Legacy FTP ON 26.X BE
September 25, 2026, 06:17:23 PM
The official documentation currently does not describe the complete setup of os-ftp-proxy and the old how-to explains the concept, but it is about 10 years old and the screenshots are no longer available.

The current documentation only mentions the basic principle.

The important missing step is that merely enabling the FTP proxy service is not enough. For a transparent forward proxy you also need a NAT port-forward rule for the FTP control connection, for example:

Interface:        LAN
Protocol:         TCP
Source:           LAN net
Destination:      any
Destination port: 21

Redirect target:  127.0.0.1
Redirect port:    8021

With the proxy enabled on 127.0.0.1:8021, but without that redirect, active FTP failed here exactly as expected:

> EPRT |1|192.168.10.97|22827|
< 425 Connections to other hosts not allowed.

After adding the redirect rule, the same active FTP test worked immediately.

The proxy handles PORT/EPRT, substitutes an externally reachable endpoint and dynamically creates the required PF rules for the incoming FTP data connection. Therefore no permanent WAN rule for the FTP data ports is required.

This can easily be tested against the public Rebex FTP test server:

curl.exe -v --ftp-port - --user demo:password ftp://test.rebex.net/pub/example/readme.txt

Important limitation: this only works for plain, unencrypted FTP.

It cannot work for FTPS/TLS, because once the FTP control connection is encrypted, ftp-proxy can no longer inspect or rewrite commands such as PORT, EPRT, PASV or EPSV, nor derive the required dynamic firewall rules from them.

The underlying ftp-proxy behaviour is documented here.


So I think the current OPNsense documentation should explicitly include

  • the required NAT redirect to 127.0.0.1:8021,
  • an explanation that ftp-proxy dynamically opens and redirects the active-mode data connection, and
  • the limitation that this cannot work with encrypted FTP control connections (FTPS/TLS).
#12
Hardware and Performance / Re: Upgrade from J6413 Question
September 24, 2026, 06:34:27 PM
The thing is, such N1x0 units were like 300€ complete with 16 GByte RAM and 256 GByte SSD a year ago. Now they are more like 600€ (both assuming good brands, not el cheapo no-name SSDs that fail after one year of heavy ZFS use).
#13
Thanks. This does not look like the LAPIC calibration problem from the thread I linked.

Your LAPIC frequency of about 500 MHz looks sane and is almost exactly what was reported for the working Proxmox case there. The broken case had a frequency off by roughly three orders of magnitude and around 65k timer interrupts/sec on every vCPU.

The `vmstat -i` numbers are interesting, but note that the displayed rate is averaged since boot, so it may hide what happens during one of the short stalls.

Since you are using `kvmclock`, I think one simple A/B test would still be worthwhile:

sysctl kern.timecounter.hardware=ACPI-fast

Leave everything else unchanged and see whether the stalls still occur. With several events per hour it should not take too long to get a useful result.

You can switch back with:

sysctl kern.timecounter.hardware=kvmclock

I would not conclude from the DTrace samples yet that the TCP retransmission timers are the cause. They may also be a consequence of the actual stall: if packet processing stops briefly, retransmission timers expire and `softclock_thread` subsequently has a lot of work to do.

The snapshot-related KVM clock problem I mentioned earlier also looks less likely in your case, since you have many events which clearly do not coincide with snapshots or backups.
#14
That was only a boot issue and does not explain these short outages. I would try disabling multiqueue on the NICs of the OPNsense VM.

There have also been reports of a current issue that may be related:
https://forum.opnsense.org/index.php?topic=52420.0

Can you give the output of:

sysctl kern.timecounter.hardware
sysctl kern.timecounter.choice
sysctl kern.timecounter.fast_gettime

sysctl kern.eventtimer.timer
sysctl kern.eventtimer.choice
sysctl kern.eventtimer.periodic

sysctl kern.eventtimer.et.LAPIC.frequency
sysctl kern.eventtimer.et.LAPIC.quality

Also, do the outages coincide with VM snapshots or backups?

#15
Just applied and rebooted that - came up fine just as usual. PPPoE over VLAN setup.