Requirements

Warning

Compatibility is limited to the hardware, firmware, and software versions listed after lab validation. Do not assume that other ConnectX models or driver combinations work the same way.

This section records the hardware, firmware, and software combination used to validate NVIDIA ASAP² Direct on MOSK. It is a tested matrix, not a vendor compatibility list or a MOSK support commitment.

Hardware prerequisites

Validated ASAP² adapter

Item

Validated value

NIC

NVIDIA ConnectX-6 Lx, dual-port 25 GbE

PCI vendor ID

15b3

PCI product ID (VF)

101e

Firmware

26.35.2000

Kernel drivers

In-tree mlx5_core and mlx5e_rep

NVIDIA documentation for ConnectX-6 Lx may list a newer firmware minimum, for example 26.42.1000. This validation used 26.35.2000. Do not assume that untested firmware revisions work the same way.

Caution

ConnectX-5, ConnectX-6 Dx, ConnectX-7, and BlueField DPUs are described in vendor documentation but were not part of this MOSK validation.

VF-LAG

VF-LAG is optional for ASAP² Direct. Production environments are recommended to use it, so validation used the following path: both ports of the same dual-port ASIC bonded as IEEE 802.3ad (LACP) with transmit hash layer3+4 and a jumbo MTU on the ASAP² bond and underlay VLAN. When you use VF-LAG, cross-card bonding of two PCIe NICs is unsupported. Both Physical Functions (PFs) must be in switchdev before the bond comes up. Firmware NUM_OF_VFS must not exceed 64, VF-LAG initialization fails above that limit.

BIOS and firmware programming

Enable SR-IOV, IOMMU (Intel VT-d or AMD-Vi), SR-IOV global / Alternative Routing-ID Interpretation (ARI), and memory-mapped I/O above 4G decoding on each ASAP² compute node. Program the adapter with mlxconfig as described in Configuration and activation.

This requires NVIDIA MFT or the Ubuntu mstflint package on the host as described in Host software stack. Replacement adapters typically ship with defaults that do not match this profile.

Server profile

Validation used a dense compute node (256 CPU threads, AMD EPYC referenced in the design). That size is informational, not a requirement. Host hugepage allocation was not recorded for this path. Some performance guests used large pages in the flavor only.

Dedicated adapter

Life-cycle management, storage, and Kubernetes underlay traffic must use other NICs. Mixing infrastructure and tenant traffic on one SmartNIC is not a supported MOSK LCM layout, see Topologies and use cases.

Host software stack

Operating system and kernel

Validation used the Ubuntu host image managed by MOSK with in-tree Linux kernel 6.8.0-52-generic. The ConnectX driver and representor subsystem were mlx5_core and mlx5e_rep at that kernel version.

NVIDIA documentation cites kernel 4.14 or later for ASAP² Direct switchdev, and 5.6 or later for connection-tracking offload. This validation did not rely on conntrack or security-group offload: accelerated ports must disable port security. See Limitations.

Firmware tools

Each ASAP² compute node must provide mlxconfig and mlxfwreset on the MOSK host image so you can program adapter firmware (SRIOV_EN, NUM_OF_VFS, LINK_TYPE) before switchdev bring-up.

Caution

These tools come from NVIDIA Mellanox Firmware Tools (MFT) or the Ubuntu mstflint package. They are required for this blueprint and are not part of the default host image. Install the package on the ASAP² compute nodes, or bake it into a custom image. Commands are in Configuration and activation. A specific MFT or mstflint version was not recorded.

NVIDIA OFED and DOCA

No MLNX_OFED or DOCA host suite version was validated. The environment used the in-tree Ubuntu mlx5 stack, not an out-of-tree OFED or DOCA package. MFT / mstflint for mlxconfig is separate from OFED and is required as described above.

Open vSwitch

The MOSK 25.1 (cluster release 17.4.0) Caracal artifacts ship Open vSwitch 2.17 (openvswitch:2.17-jammy-*). The validated other_config settings were hw-offload true and max-idle 30000 (milliseconds).

Apply these through OpenStackDeployment as described in Configuration and activation.

Switchdev bring-up

Host-side switchdev bring-up is not a productized MOSK LCM procedure. Validation used an experimental Host OS configuration module that discovers PFs, creates VFs, switches the eSwitch to switchdev, enables hw-tc-offload, then rebinds VFs before containerd and kubelet. Early steps must finish before host networking starts. VF rebind must run after the network is online and before these container runtimes.

Note

The module is not part of the product. You can use it as a reference, see switchdev Host OS configuration module.

For the ordered enablement sequence, see Configuration and activation.

MOSK and OpenStack matrix

Validated platform combination

Layer

Validated value

MOSK cluster release

17.4.0+25.1

MOSK controller

1.0.7

OpenStack release

Caracal

Neutron

ML2 with openvswitch and sriovnicswitch

Open vSwitch

2.17 (MOSK 25.1 / 17.4.0 Caracal image)

OVS datapath

Kernel datapath with TC Flower / switchdev

Neutron QoS

Disabled cluster-wide (qos service plugin omitted; Open vSwitch agent extension unloaded)

Validation window

Staging enablement in mid-2025 (applied state July 2025)

Other MOSK or OpenStack combinations were not characterized.

The following networking backends are out of scope:

  • OVN

  • OpenSDN (Tungsten Fabric)

  • OVS-DPDK

  • DPU-hosted OVS or DOCA (for example on BlueField). Only ML2/OVS with the kernel datapath and sriovnicswitch was validated.

Cluster assumptions

  • Disable Neutron QoS for the whole cloud. That includes omitting the qos service plugin and unloading the Open vSwitch agent extension. Leaving policies off accelerated ports is not enough. See Limitations and Configuration and activation.

  • Enable ASAP² only on compute nodes that have the SmartNIC. Do not declare the accelerated physnet, SR-IOV NIC name, or PCI device specification cluster-wide. The OVS or SR-IOV agent fails on compute nodes that lack that device. See Limitations.

  • Isolate three bond roles: host management and Kubernetes LCM, storage and cluster infrastructure, and a dedicated ASAP² VF-LAG bond.

  • Keep the VF-LAG bond in the kernel (Netplan from L2Template). Place the VXLAN underlay on a VLAN of the ASAP² bond. Set tunnel_interface in OpenStackDeployment to the VLAN interface so OVS sends overlay traffic out the ASAP² underlay.

  • Use a dedicated ASAP² physnet on those compute nodes, separate from the infrastructure physnet on the non-ASAP² NICs. Declare the SR-IOV NIC and PCI device specification only on nodes that have the SmartNIC. Provider VLAN mapping in OpenStackDeployment is in Configuration and activation.

  • Configure the ToR with a matching LACP port-channel (typically MLAG) and trunk the underlay VLAN on that channel.

Bond, VLAN, and physnet layout is described in Topologies and use cases. For the enablement steps, refer to Configuration and activation.