Limitations
This section records the limits of ASAP² Direct as validated on MOSK: which OpenStack networking features stay off the hardware path, what to watch in day-2 operations and node life cycle, and how narrowly hardware and software compatibility was tested. It is not a vendor support or compatibility list.
Functional and feature limitations
ASAP² Direct accelerates eligible tenant flows in the NIC eSwitch. Traffic that Open vSwitch cannot program into hardware stays on the host software path, and several OpenStack networking features never enter the offload path at all.
The following OpenStack and datapath features could not be hardware-offloaded in the validated configuration. Treat them as out of scope for an accelerated plane for the given combination of software and hardware.
Neutron routers and NAT
Any traffic that traverses a Neutron logical router—centralized or distributed (DVR), with or without SNAT or DNAT — does not offload. It falls back to host CPU switching. To keep acceleration, connect instances directly to the accelerated VXLAN (or to a provider VLAN) without a Neutron router hop. North-south routing and stateful policy that you still need belong in a guest virtual router or firewall that bridges the accelerated tenant network and the external network.
Neutron floating IPs
Associating a floating IP with an accelerated switchdev SR-IOV port
(vnic-type=direct with capabilities: ["switchdev"]) failed in the
validated DVR environment. With distributed virtual routing, the compute node
that hosts the instance also performs floating-IP NAT, and that binding did not
succeed on the switchdev port.
Where floating-IP traffic is handled by Neutron routers on dedicated gateway nodes (centralized FIP or SNAT, not DVR on the compute node), associating an SR-IOV port with a floating IP should theoretically work. That topology was not validated. Even if the association succeeds, north-south traffic still traverses a Neutron router and NAT, so it stays on the software path and does not gain ASAP² acceleration. See Neutron routers and NAT above.
For accelerated north-south access, attach the instance directly to a provider
VLAN or an external VXLAN through a switchdev port, without a floating IP.
The lab used Distributed Virtual Routing (DVR). That is not a requirement for ASAP² Direct. The validated overlay path should work with centralized (non-DVR) routing as well.
Security groups and port security
Standard Neutron security groups and port security are incompatible
with the OVS hardware-offload path used here. Accelerated ports must
be created with --disable-port-security. Do not expect conntrack
security-group rules to be programmed into the eSwitch.
Neutron QoS
Disable Neutron QoS for the whole cloud. That includes omitting the qos
service plugin and unloading the Open vSwitch agent extension. Bandwidth-limit
and minimum-bandwidth policies are incompatible with this offload datapath.
Leaving policies off the accelerated ports is not enough. The
OpenStackDeployment fragment is in
Configuration and activation.
ARP, neighbor discovery, and first packets
ARP and IPv6 Neighbor Discovery stay on the host. The first packet of each new flow is trapped to the VF representor so OVS can install a TC flower rule. Subsequent matching packets of an eligible flow can run in the eSwitch until the idle timeout expires. IPv6 tenant unicast was not validated explicitly; it should follow the same path. See Architecture.
Live migration
Moving a running instance that uses a direct ASAP² port is generally limited. The guest is tied to a physical NIC function on that host, so live migration is not a straightforward OpenStack operation. It was not validated here and remains a future requirement, not a procedure you can rely on from this blueprint.
According to NVIDIA documentation, newer adapters (ConnectX-6 Dx and later) can be tuned so the NIC matches flows in a way that allows SR-IOV live migration; that tuning was not used on the ConnectX-6 Lx hardware in this work.
A longer-term approach is vDPA: the guest sees a normal virtio-net NIC
while the host still offloads traffic, which keeps live migration possible.
vDPA was not part of this validation.
The following matrix summarizes the validated tenant and provider bindings. Details of physnets and bonds are in Topologies and use cases.
Operational and life-cycle caveats
Boot order versus running pods
Physical Functions must enter switchdev mode, representors must exist,
hw-tc-offload must be on, and VFs must be rebound before containerd
and kubelet start. If those services are already up, you cannot insert the
sequence into the current boot. On a provisioned compute node, persist the
automation, drain the node, and reboot. See
Configuration and activation.
The experimental Host OS configuration module used in the lab is not a productized MOSK procedure. Use it only as a reference for ordering around host networking.
Node replacement
A replacement compute node needs the same BIOS options (SR-IOV, IOMMU, ARI,
MMIO above 4G) and the same non-volatile adapter profile (mlxconfig: SR-IOV
enabled, Ethernet link type, VF count within the VF-LAG maximum). Factory
firmware does not match the cluster profile automatically.
OpenStackDeployment scope
Declare the accelerated physnet, SR-IOV NIC name, and PCI device specification only on node types that have the SmartNIC. A cluster-wide mapping to a named PF causes the OVS or SR-IOV agent to fail on compute nodes that lack that device. Node-specific overrides also lengthen OSDPL apply time. Allow a longer reconciliation window than for a default Helm apply.
Bonding the PF into OVS too early
Do not enslave an active VF-LAG bond into br-ex or an SR-IOV bridge from
generic OVS startup if the kernel already owns the PFs. That pattern produced
device or resource busy and OVS pidfile errors during validation.
The validated path keeps the bond in the kernel (Netplan from the compute
L2Template), puts the VXLAN underlay on a VLAN of that bond, and lets
Neutron attach VF representors. The experimental switchdev Host OS module
only prepares the NIC around that host networking; it does not add the bond to
OVS. See
Configuration and activation.
Firmware and drivers
Validation used the in-tree Ubuntu mlx5_core / mlx5e_rep driver and a
single firmware revision listed in
Requirements. Unlisted driver or firmware
combinations were not characterized. Firmware reset during driver health
recovery can take longer than a naive timeout; treat live firmware flash as a
maintenance event, not a rolling in-place change under tenant load.
Compatibility constraints
- NIC and firmware
Claims apply only to the ConnectX-6 Lx dual-port 25GbE adapter, PCI IDs, firmware, and kernel driver in Requirements. ConnectX-5, ConnectX-6 Dx, ConnectX-7, and BlueField DPUs appear in vendor literature but were not part of this MOSK validation.
- VF-LAG on one ASIC
When you use VF-LAG, both bond members must be ports of the same dual-port adapter. Cross-card bonding across two PCIe NICs is unsupported. Both PFs must be in
switchdevbefore the bond comes up.
- VF count
VF-LAG initialization fails if
NUM_OF_VFSexceeds 64.
- Dedicated tenant NIC
Life-cycle, storage, and Kubernetes underlay traffic must not share the ASAP² adapter. Mixing infrastructure and tenant traffic on one SmartNIC is not a supported MOSK LCM layout. Isolate the VF-LAG bond as described in Topologies and use cases.
- Networking backend
Only ML2/OVS with the kernel datapath and
sriovnicswitchwas validated. OVN, OpenSDN, Tungsten Fabric, and OVS-DPDK are out of scope.