Menu

Show posts

This section allows you to view all posts made by this member. Note that you can only see posts made in areas you currently have access to.

Show posts Menu

Messages - Andreas L.

#31
Sorry for the delay, I had to provoke it first.
I can now clearly destroy it with this approach:

  • Reboot OPNsense --> Everything is fine
  • Let a ping go constanly from my network over into the other net of the wireguard tunnel --> Everything is fine
  • Stop my ping  --> Everything is fine
  • Let other wireguard side ping one device of my local network --> MBUF error appears nearly immediatly with the first ping

So when everything is working fine it looks like this:

root@OPNsense:~ # ifconfig bce0
bce0: flags=8843<UP,BROADCAST,RUNNING,SIMPLEX,MULTICAST> metric 0 mtu 1500
   options=80028<VLAN_MTU,JUMBO_MTU,LINKSTATE>
   ether 00:21:5e:c8:be:88
   inet6 fe80::221:5eff:fec8:be88%bce0 prefixlen 64 scopeid 0x1
   inet6 2a02:2f4:xxxx:xxx0:221:5eff:fec8:be88 prefixlen 64 autoconf
   inet6 fd00:0:cafe:affe:221:5eff:fec8:be88 prefixlen 64 autoconf
   inet 192.168.0.100 netmask 0xffffff00 broadcast 192.168.0.255
   media: Ethernet autoselect (1000baseT <full-duplex>)
   status: active
   nd6 options=23<PERFORMNUD,ACCEPT_RTADV,AUTO_LINKLOCAL>


After wireguard ping trigger from other side partner:

root@OPNsense:~ # ifconfig bce0
bce0: flags=8843<UP,BROADCAST,RUNNING,SIMPLEX,MULTICAST> metric 0 mtu 1500
   options=80028<VLAN_MTU,JUMBO_MTU,LINKSTATE>
   ether 00:21:5e:c8:be:88
   inet6 fe80::221:5eff:fec8:be88%bce0 prefixlen 64 scopeid 0x1
   inet6 2a02:2f4:xxxx:xxx0:221:5eff:fec8:be88 prefixlen 64 autoconf
   inet6 fd00:0:cafe:affe:221:5eff:fec8:be88 prefixlen 64 autoconf
   inet 192.168.0.100 netmask 0xffffff00 broadcast 192.168.0.255
   media: Ethernet autoselect (1000baseT <full-duplex>)
   status: active
   nd6 options=23<PERFORMNUD,ACCEPT_RTADV,AUTO_LINKLOCAL>


I cannot see an immediate difference here ???.
#32
Some bad news :(, router was running 5 days, nearly 6 days now, without this issue, but today - out of a sudden - the error messages returned:
[4006] netmap_transmit bce0 drop mbuf that needs checksum offload

MBUF usage is slightly higher than normal, but far (!) away from the ciritcal maximum:

MBUF Usage  0% (10432/1271498)

So I'm not sure, if I can consider this still as solved, but at least as remarkably better.

And I said "out of a sudden" but I'm afraid the trigger might be somehow in relation to my wireguard side2side VPN, it started briefly after the counter part send a test ping after a quite long period of data silence between the two locations. I'm not sure, if this might be related, it could be coincidence, but I think, I should mention it.
#33
Hmm, Ok understandable and good to know, so hoping, I will not expect these errors again then, but then I know at least why :). Thanks for making me aware.
#34
Just to return some feedback here, I'm testing now for 24h under "full load" incl. IDS/IPS and Sensei and the messages did not appear anymore. So I consider this issue as solved with the new kernel "kernel-20.7.2-netmap-amd64.txz"!

Thank you very much :)!

PS: I assume, my preloading to test before next official update is not an issue for the upcoming release aka official update or will I get in troube with this kernel now?
#35
Awesome @mb, thank you! I have done that and rebooted:

root@OPNsense:~ # opnsense-update -kr 20.7.2-netmap
Fetching kernel-20.7.2-netmap-amd64.txz: ....... done
!!!!!!!!!!!! ATTENTION !!!!!!!!!!!!!!!
! A critical upgrade is in progress. !
! Please do not turn off the system. !
!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
Installing kernel-20.7.2-netmap-amd64.txz... done
Please reboot.

I've also activated IDS/IPS again to monitor it now. 5 mins later no problems yet, so still monitoring.
I keep you posted!

PS: And as requested, all offloadings and vlan hardware filtering were already set to disabled.
#36
I'm running OPNsense (20.7.2-amd64) with one Broadcom NetXtreme II BCM5709 for WAN (bce0) and one for LAN (bce1), further on I have 4x Intel 82580, which I use for other LANs like IoT (igb1) and Guests (igb0) etc.

I have "some" traffic on WAN with quite constantly 60 to 100MBit (mainly due to IP cam streams), which I consider as handeable with my setup. I also have IDS/IPS up and running as well as Sensei.

After "a while" (usually only minutes after reboot) of traffic I get the following error in the log, multiple times per second:

2020-09-10T00:28:10   kernel   490.690419 [4006] netmap_transmit bce0 drop mbuf that needs checksum offload
2020-09-10T00:28:05   kernel   485.572543 [4006] netmap_transmit bce0 drop mbuf that needs checksum offload
2020-09-10T00:28:00   kernel   480.194945 [4006] netmap_transmit bce0 drop mbuf that needs checksum offload
2020-09-10T00:28:00   kernel   479.940436 [4006] netmap_transmit bce0 drop mbuf that needs checksum offload
2020-09-10T00:27:54   kernel   474.761838 [4006] netmap_transmit bce0 drop mbuf that needs checksum offload
2020-09-10T00:27:49   kernel   469.475112 [4006] netmap_transmit bce0 drop mbuf that needs checksum offload
2020-09-10T00:27:44   kernel   464.324372 [4006] netmap_transmit bce0 drop mbuf that needs checksum offload
2020-09-10T00:27:39   kernel   459.205033 [4006] netmap_transmit bce0 drop mbuf that needs checksum offload
2020-09-10T00:27:33   kernel   453.830080 [4006] netmap_transmit bce0 drop mbuf that needs checksum offload
2020-09-10T00:27:28   kernel   448.126626 [4006] netmap_transmit bce0 drop mbuf that needs checksum offload
2020-09-10T00:27:23   kernel   443.431391 [ 320] generic_netmap_register Emulated adapter for bce0 activated
2020-09-10T00:27:23   kernel   443.431259 [1130] generic_netmap_attach Emulated adapter for bce0 created (prev was NULL)
2020-09-10T00:27:23   kernel   bce0: permanently promiscuous mode enabled
2020-09-10T00:27:23   kernel   443.407436 [1035] generic_netmap_dtor Emulated netmap adapter for bce0 destroyed
2020-09-10T00:27:23   kernel   443.407409 [1130] generic_netmap_attach Emulated adapter for bce0 created (prev was NULL)

As you can see on the attached screenshot, the MBUF usage is at 0% and with ~9720 way below the limit of 1.271.626, so there should be plenty of MBUF available.

So what triggers this error?

I can get rid of it, when deactivating IDS/IPS, and since I'm testing it, the error did not show up again. So is it somehow IPS throughput related? Nonetheless, I would like to turn IDS/IPS on again :).

How can I tune my system, so the "netmap_transmit" can handle the load? (BTW: What process/step ist it, what does it do here?)
And whay does the mbuf "need checksum offload"? What does that exactly mean?

Some more config details:

I have all three hooks set, so all of these three are disabled:
- Hardware CRC
- Hardware TSO
- Hardware LRO


root@OPNsense:~ # sysctl -a | grep nmbclusters
kern.ipc.nmbclusters: 1271626

root@OPNsense:~ # sysctl -a | grep msi
hw.sdhci.enable_msi: 1
hw.puc.msi_disable: 0
hw.pci.honor_msi_blacklist: 1
hw.pci.msix_rewrite_table: 0
hw.pci.enable_msix: 1
hw.pci.enable_msi: 1
hw.mfi.msi: 1
hw.malo.pci.msi_disable: 0
hw.ix.enable_msix: 1
hw.bce.msi_enable: 1
hw.aac.enable_msi: 1
machdep.disable_msix_migration: 0
machdep.num_msi_irqs: 512
dev.igb.3.iflib.disable_msix: 0
dev.igb.2.iflib.disable_msix: 0
dev.igb.1.iflib.disable_msix: 0
dev.igb.0.iflib.disable_msix: 0


BTW: I also experimented with following values, which did not bring any change:

kern.ipc.nmbclusters="2543660"
hw.bce.tso_enable="0"
hw.pci.enable_msix="0"
#37
Any news here? Could anyone somehow prove that direct ping6 of gateway from OPNsense via IPv6 link local works?
Looking for someone running OPNsense behind a FritzBox with a delegated sub net prefix to compare - maybe in a working setup.

I simply don't get, why ICMPv6 does not find a route (as shown before), when routing table clearly states, it's the default route to go :o and traceroute also works as it should ???. I consider this as a bug or at least - if somewhere hidden - some in-transparent setup somewhere. I'm open to test anything to move on here.

Seems IPv6 support is not that sophisticated yet as IPv4 within OPNsense. :-\
#38
Thanks bobm, very good hint, but this I already worked through (also a very recommended requirement for "Intrusion Detection" to be started/used).
So I have all three hooks set, so all of these three are disabled:
- Hardware CRC
- Hardware TSO
- Hardware LRO
#39
20.7 Legacy Series / Re: IPv6 / PPP dropout troubles
September 08, 2020, 09:39:11 PM
Sounds briefly like my story. I'm behind a CGNAT net now as well since some days. My ISP is "komflat" in Germany, if this is known? It's not a very professional provider and I'm considering to move out there as well asap, we can discuss this at the side :).

None-the-less, I like the dual stack approach now but I also expect IPv6 to work with my OPNsense. I get a /56 and "forward" a /60 further on to my OPNsense internally from my FritzBox. This seems to work fine, but I cannot ping my FB, which is currently a huge blocker for me.
#40
Problem is still not solved :-[. Today I removed "netdata" as this was causing thousand of these error log entries after a short while:

netmap_transmit bce0 drop mbuf that needs checksum offload

So the log get's cleaner, but no idea how to progress and solve or analyse this IPv6 issue :-\.
#41
20.7 Legacy Series / Re: IPv6 / PPP dropout troubles
September 08, 2020, 08:49:21 PM
I think I can second it, I posted a very detailed report here some days ago: https://forum.opnsense.org/index.php?topic=18923.0

Until today, I still have issues with IPv6, while IPv4 in parallel (and since ever before) is running fine.

I also cannot ping6 my IPv6 gateway, but traceroute6 works fine as I described here: https://forum.opnsense.org/index.php?topic=19018.0

Maybe, if you read through my threads, you might confirm something in relation to your setup and behaviour?
#42
Hehe, agreed, you started with language switch, I'm flexible I just adapted :).

But good to know, that this seems to be common in BSD with the interface being part of the link local IPv6 gateway IP and not being a bug.
A /64 I haven't tried as this would destroy the approach of separated sub nets. Just for testing I can give it a try.

And all your other questions are answered in my former post(s). :)
netstat output is visible in my example before, ICMPv6 firewall rule for WAN is in place (ping is also visible and green in FW log).
Bogon and RFC1918 is deactivated as WAN being in a common private network.

I also activated in "Firewall: Settings: Advanced" the option "Disable force gateway" as I read somewhere this might influence usage of routing table.

So the key question is left, what is needed to ensure OPNsense uses the route as announced in the routing table? Why does traceroute6 work, but ping6 cannot determine a route? What else is needed?

Would be nice to find someone with a compareable setup. ::)


And second, this is something appearing from time to time in the log, could this influence this behaviour and is the root cause known? Curenly I guess this problem is independent, but I habe not done any more research yet.

error in configd communication Traceback (most recent call last): File "/usr/local/opnsense/service/configd_ctl.py", line 68, in exec_config_cmd line = sock.recv(65536).decode() socket.timeout: timed out
#43
Ich habe noch eine Überlegung, könnte das ein Bug sein, dass das Interface fest bei der Link Local mit hinterlegt wird in der Routing Tabelle?

Warum "fe80::c225:6ff:feff:820d%bce0"? Das zu verwendende Interface steht ja am Ende der Routing Tabelle schon, warum auch in der Adresse? In der Routing Tabelle macht es doch auch als Link Local keinen Sinn.

Ich habe auch mal mit einem Linux Rechner hinter der FB vergleichen, da steht als default IPv6-Gateway auch die Link Local von der FB drin, allerdings ohne das Interface - sonst ist alles gleich.

Frage ob das wirklich ein Problem ist bei BSD? Und würde es dann auch zu dem "No route to host" error kommen?
#44
Nein, natürlich nicht. Alles wurde automatisch vergeben, also wie es sich gehört und wie man es erwartet.
#45
And one more thing here to add. There is allegedly no route to the default gateway, what I cannot understand as the routing table clearly states, there is the route:

root@OPNsense:~ # traceroute6 fe80::c225:6ff:feff:820d (It does not matter, if I add "%bce0" or not.)
traceroute6 to fe80::c225:6ff:feff:820d (fe80::c225:6ff:feff:820d) from fe80::221:5eff:fec8:be88%bce0, 64 hops max, 20 byte packets
sendto: No route to host
1 traceroute6: wrote fe80::c225:6ff:feff:820d 12 chars, ret=-1
*sendto: No route to host
traceroute6: wrote fe80::c225:6ff:feff:820d 12 chars, ret=-1
*sendto: No route to host
traceroute6: wrote fe80::c225:6ff:feff:820d 12 chars, ret=-1


But when I check the routing table, this is exactly how it should be and what I would expect :o

root@OPNsense:~ # netstat -nr
Routing tables

Internet:
[...]

Internet6:
Destination                       Gateway                       Flags     Netif Expire
default                           fe80::c225:6ff:feff:820d%bce0 UG         bce0
::1                               link#8                        UH          lo0
2a02:2f4:xxxx:xxxx::/64           link#1                        U          bce0
2a02:2f4:xxxx:xxxx:221:5eff:fec8:be88 link#1                    UHS         lo0
fd00:0:cafe:affe::/64             link#1                        U          bce0
fd00:0:cafe:affe:221:5eff:fec8:be88 link#1                      UHS         lo0
fe80::%bce0/64                    link#1                        U          bce0
fe80::221:5eff:fec8:be88%bce0     link#1                        UHS         lo0
fe80::%bce1/64                    link#2                        U          bce1
fe80::221:5eff:fec8:be8a%bce1     link#2                        UHS         lo0
fe80::%igb0/64                    link#3                        U          igb0
fe80::92e2:baff:fe68:cd74%igb0    link#3                        UHS         lo0
fe80::%igb1/64                    link#4                        U          igb1
fe80::92e2:baff:fe68:cd75%igb1    link#4                        UHS         lo0
fe80::%lo0/64                     link#8                        U           lo0
fe80::1%lo0                       link#8                        UHS         lo0