Your latest tcpdump provides an important clue: Even in the broken state, DNS queries from 10.127.127.2 reach 10.127.127.1, and valid DNS responses are sent back through wg4.
This suggests that Unbound is working. The problem is more likely related to packet delivery, firewall rules, or routing/gateway selection during startup.
Since your setup involves multiple gateways and policy-based routing (PBR), I particularly suspect a routing or firewall state issue during boot, possibly involving incorrect gateway selection or asymmetric routing.
Could you perform the following tests before restarting WireGuard, while simultaneously running `tcpdump -ni wg4 'host 10.127.127.2'` on OPNsense?
1. Explicit DNS lookup: Query `10.127.127.1` directly from the affected client (e.g. `nslookup example.com 10.127.127.1`). If this fails despite DNS responses appearing in the capture, the problem is likely on the return path to the client or in the client's DNS stack. If it succeeds, DNS transport itself is working.
2. Direct IP connection: Try accessing a known reachable Internet server by IP address, bypassing DNS entirely. If this fails while DNS works, suspect forwarding, NAT, firewall rules, or gateway selection. The capture will show whether the connection attempt reaches wg4.
3. Connection by hostname: Access the same server by hostname. If direct IP access works but hostname access fails, investigate the client's DNS configuration or resolver behavior rather than general IP routing.
Given your multiple gateways and PBR rules, I would also compare the firewall states and gateway selection for the affected client traffic before and after restarting WireGuard.
These tests should distinguish DNS transport issues from forwarding/routing problems and incorrect client-side name resolution.
I would avoid introducing startup delays until we have identified the failing component.
This suggests that Unbound is working. The problem is more likely related to packet delivery, firewall rules, or routing/gateway selection during startup.
Since your setup involves multiple gateways and policy-based routing (PBR), I particularly suspect a routing or firewall state issue during boot, possibly involving incorrect gateway selection or asymmetric routing.
Could you perform the following tests before restarting WireGuard, while simultaneously running `tcpdump -ni wg4 'host 10.127.127.2'` on OPNsense?
1. Explicit DNS lookup: Query `10.127.127.1` directly from the affected client (e.g. `nslookup example.com 10.127.127.1`). If this fails despite DNS responses appearing in the capture, the problem is likely on the return path to the client or in the client's DNS stack. If it succeeds, DNS transport itself is working.
2. Direct IP connection: Try accessing a known reachable Internet server by IP address, bypassing DNS entirely. If this fails while DNS works, suspect forwarding, NAT, firewall rules, or gateway selection. The capture will show whether the connection attempt reaches wg4.
3. Connection by hostname: Access the same server by hostname. If direct IP access works but hostname access fails, investigate the client's DNS configuration or resolver behavior rather than general IP routing.
Given your multiple gateways and PBR rules, I would also compare the firewall states and gateway selection for the affected client traffic before and after restarting WireGuard.
These tests should distinguish DNS transport issues from forwarding/routing problems and incorrect client-side name resolution.
I would avoid introducing startup delays until we have identified the failing component.
"