Limitations

This section records the limits of ASAP² Direct as validated on MOSK: which OpenStack networking features stay off the hardware path, what to watch in day-2 operations and node life cycle, and how narrowly hardware and software compatibility was tested. It is not a vendor support or compatibility list.

Functional and feature limitations

ASAP² Direct accelerates eligible tenant flows in the NIC eSwitch. Traffic that Open vSwitch cannot program into hardware stays on the host software path, and several OpenStack networking features never enter the offload path at all.

The following OpenStack and datapath features could not be hardware-offloaded in the validated configuration. Treat them as out of scope for an accelerated plane for the given combination of software and hardware.

Neutron routers and NAT

Any traffic that traverses a Neutron logical router—centralized or distributed (DVR), with or without SNAT or DNAT — does not offload. It falls back to host CPU switching. To keep acceleration, connect instances directly to the accelerated VXLAN (or to a provider VLAN) without a Neutron router hop. North-south routing and stateful policy that you still need belong in a guest virtual router or firewall that bridges the accelerated tenant network and the external network.

Neutron floating IPs

Associating a floating IP with an accelerated switchdev SR-IOV port (vnic-type=direct with capabilities: ["switchdev"]) failed in the validated DVR environment. With distributed virtual routing, the compute node that hosts the instance also performs floating-IP NAT, and that binding did not succeed on the switchdev port.

Where floating-IP traffic is handled by Neutron routers on dedicated gateway nodes (centralized FIP or SNAT, not DVR on the compute node), associating an SR-IOV port with a floating IP should theoretically work. That topology was not validated. Even if the association succeeds, north-south traffic still traverses a Neutron router and NAT, so it stays on the software path and does not gain ASAP² acceleration. See Neutron routers and NAT above.

For accelerated north-south access, attach the instance directly to a provider VLAN or an external VXLAN through a switchdev port, without a floating IP.

The lab used Distributed Virtual Routing (DVR). That is not a requirement for ASAP² Direct. The validated overlay path should work with centralized (non-DVR) routing as well.

Security groups and port security

Standard Neutron security groups and port security are incompatible with the OVS hardware-offload path used here. Accelerated ports must be created with --disable-port-security. Do not expect conntrack security-group rules to be programmed into the eSwitch.

Neutron QoS

Disable Neutron QoS for the whole cloud. That includes omitting the qos service plugin and unloading the Open vSwitch agent extension. Bandwidth-limit and minimum-bandwidth policies are incompatible with this offload datapath. Leaving policies off the accelerated ports is not enough. The OpenStackDeployment fragment is in Configuration and activation.

ARP, neighbor discovery, and first packets

ARP and IPv6 Neighbor Discovery stay on the host. The first packet of each new flow is trapped to the VF representor so OVS can install a TC flower rule. Subsequent matching packets of an eligible flow can run in the eSwitch until the idle timeout expires. IPv6 tenant unicast was not validated explicitly; it should follow the same path. See Architecture.

Live migration

Moving a running instance that uses a direct ASAP² port is generally limited. The guest is tied to a physical NIC function on that host, so live migration is not a straightforward OpenStack operation. It was not validated here and remains a future requirement, not a procedure you can rely on from this blueprint.

According to NVIDIA documentation, newer adapters (ConnectX-6 Dx and later) can be tuned so the NIC matches flows in a way that allows SR-IOV live migration; that tuning was not used on the ConnectX-6 Lx hardware in this work.

A longer-term approach is vDPA: the guest sees a normal virtio-net NIC while the host still offloads traffic, which keeps live migration possible. vDPA was not part of this validation.

The following matrix summarizes the validated tenant and provider bindings. Details of physnets and bonds are in Topologies and use cases.

Network types and offload status

Network

Encapsulation

Port binding

Offload

Tenant overlay

VXLAN

vnic-type=direct with capabilities: ["switchdev"] and port security disabled

Validated. Eligible flows run in the eSwitch.

Tenant overlay

VXLAN

virtio (vnic-type=normal)

Attachment validated. Host OVS can use the same switchdev PF. This is not the direct VF fast path used for line-rate tests.

Provider network

VLAN

Direct switchdev SR-IOV

Validated for hardware VLAN push and pop.

Floating IP / external

VLAN or VXLAN via br-fip / br-ex

Floating IP on a switchdev port

Association failed under DVR. Centralized gateway FIP was not validated; north-south would not offload in any case.

Operational and life-cycle caveats

Boot order versus running pods

Physical Functions must enter switchdev mode, representors must exist, hw-tc-offload must be on, and VFs must be rebound before containerd and kubelet start. If those services are already up, you cannot insert the sequence into the current boot. On a provisioned compute node, persist the automation, drain the node, and reboot. See Configuration and activation.

The experimental Host OS configuration module used in the lab is not a productized MOSK procedure. Use it only as a reference for ordering around host networking.

Node replacement

A replacement compute node needs the same BIOS options (SR-IOV, IOMMU, ARI, MMIO above 4G) and the same non-volatile adapter profile (mlxconfig: SR-IOV enabled, Ethernet link type, VF count within the VF-LAG maximum). Factory firmware does not match the cluster profile automatically.

OpenStackDeployment scope

Declare the accelerated physnet, SR-IOV NIC name, and PCI device specification only on node types that have the SmartNIC. A cluster-wide mapping to a named PF causes the OVS or SR-IOV agent to fail on compute nodes that lack that device. Node-specific overrides also lengthen OSDPL apply time. Allow a longer reconciliation window than for a default Helm apply.

Bonding the PF into OVS too early

Do not enslave an active VF-LAG bond into br-ex or an SR-IOV bridge from generic OVS startup if the kernel already owns the PFs. That pattern produced device or resource busy and OVS pidfile errors during validation.

The validated path keeps the bond in the kernel (Netplan from the compute L2Template), puts the VXLAN underlay on a VLAN of that bond, and lets Neutron attach VF representors. The experimental switchdev Host OS module only prepares the NIC around that host networking; it does not add the bond to OVS. See Configuration and activation.

Firmware and drivers

Validation used the in-tree Ubuntu mlx5_core / mlx5e_rep driver and a single firmware revision listed in Requirements. Unlisted driver or firmware combinations were not characterized. Firmware reset during driver health recovery can take longer than a naive timeout; treat live firmware flash as a maintenance event, not a rolling in-place change under tenant load.

Compatibility constraints

  • NIC and firmware

    Claims apply only to the ConnectX-6 Lx dual-port 25GbE adapter, PCI IDs, firmware, and kernel driver in Requirements. ConnectX-5, ConnectX-6 Dx, ConnectX-7, and BlueField DPUs appear in vendor literature but were not part of this MOSK validation.

  • VF-LAG on one ASIC

    When you use VF-LAG, both bond members must be ports of the same dual-port adapter. Cross-card bonding across two PCIe NICs is unsupported. Both PFs must be in switchdev before the bond comes up.

  • VF count

    VF-LAG initialization fails if NUM_OF_VFS exceeds 64.

  • Dedicated tenant NIC

    Life-cycle, storage, and Kubernetes underlay traffic must not share the ASAP² adapter. Mixing infrastructure and tenant traffic on one SmartNIC is not a supported MOSK LCM layout. Isolate the VF-LAG bond as described in Topologies and use cases.

  • Networking backend

    Only ML2/OVS with the kernel datapath and sriovnicswitch was validated. OVN, OpenSDN, Tungsten Fabric, and OVS-DPDK are out of scope.