Configuration and activation

This section describes how to enable NVIDIA ASAP² Direct on MOSK compute nodes that use the Open vSwitch (OVS) kernel datapath. The procedure builds on classic SR-IOV with the Neutron OVS backend.

ASAP² additionally requires the Physical Functions (PFs) to run in switchdev mode, VF representors for OVS, and hardware offload in the OVS daemon.

Host-side switchdev bring-up is not yet a productized MOSK procedure. The following steps outline the general approach to configuring the host OS and NIC for ASAP² Direct offloading. Apply these steps with any durable host automation you operate, for example, first-boot scripts, image customization, or a custom module through the host OS configuration API.

Starting point

The procedure assumes a MOSK compute node has already been provisioned:

  • The host is a compute node that has already fully joined the cluster: Kubernetes underlay components (containerd and kubelet) are running, the node networking stack is up, and the Open vSwitch and Compute (Nova) pods are already scheduled on the node.

  • An NVIDIA SmartNIC that meets Requirements is installed and cabled for tenant traffic. The adapter is not yet enabled for ASAP²: Physical Functions remain in legacy (non-switchdev) mode, VFs and representors are not prepared for offload, and OpenStackDeployment does not yet declare SR-IOV hardware offload on that node.

  • Life-cycle management and storage traffic already use other NICs. The SmartNIC is reserved for the ASAP² path described in Topologies and use cases.

Approach

Host OS configuration for switchdev mode, VF representors, hw-tc-offload, and VF rebind must be in place before containerd and kubelet start, and therefore before the Open vSwitch and OpenStack pods consume representors and VF PCI devices. If those services are already up, you cannot insert that sequence into the current boot.

On a provisioned compute node, treat host enablement as persist, drain, and reboot—not a live change under running pods. Firmware, kernel IOMMU, and L2Template updates can each trigger a reboot. Install the switchdev boot automation before the reboot that brings up the ASAP² bond so the order on that boot is: early switchdev, then Netplan bond and underlay, then VF rebind, then containerd and kubelet.

Host and LCM

  1. Evacuate or otherwise remove instances from the compute node so no guest depends on the SmartNIC.

  2. Enable server BIOS/UEFI options, then program the adapter with mlxconfig. See Firmware and BIOS.

  3. Enable kernel IOMMU with the grub_settings module. See Host IOMMU.

  4. Install durable automation for the Switchdev bring-up sequence (VF create, unbind, devlink switchdev, hw-tc-offload, VF rebind). It must run on every boot, around host networking, and finish before containerd.service and kubelet.service. Do not restart these services in place. See Global recommendations for implementation of custom modules.

  5. Declare the VF-LAG bond and VXLAN underlay VLAN in a new compute L2Template and apply it as described in Bond and underlay. IPAM renders Netplan. Do not create the bond with ip link.

  6. After the host is back, confirm Host readiness (eSwitch mode, offload, rebound VFs, representors, bond and underlay).

Note

During validation, Mirantis developed an experimental host OS configuration module that automates host OS and NIC steps for switchdev bring-up. The module uses systemd units to run that configuration in the required order at node startup.

The module is not yet part of the product. Mirantis does not guarantee that it works for every software and hardware combination. You can use it as a reference for the operations described in switchdev Host OS configuration module.

OpenStack and workloads

After the host is in switchdev and the underlay is up:

  1. Configure Networking, Compute, and OVS in the OpenStackDeployment custom resource.

  2. Create accelerated networks and switchdev ports, then launch instances.

  3. Verify that eligible flows offload to the NIC.

Hardware, firmware, and software bounds are listed in Requirements. For bond, VLAN, and physnet layout, see Topologies and use cases. Functional constraints, such as disabled port security, Neutron QoS disabled globally, and no offload through Neutron routers or floating IPs, are listed in Limitations.

Host and NIC preparation

Firmware and BIOS

Enable the following options in the server BIOS/UEFI of each ASAP² compute node (boot-time firmware setup or the out-of-band BMC equivalent). Menu names vary by vendor. Complete this before the operating system creates Virtual Functions (VFs) or changes eSwitch mode, and before you program the adapter with mlxconfig:

  • SR-IOV support

  • IOMMU (Intel VT-d or AMD-Vi)

  • SR-IOV global / Alternative Routing-ID Interpretation (ARI), required for high VF counts

  • Memory-mapped I/O above 4G decoding, required for 64-bit PCI BAR allocation across dense VF sets

Then configure adapter firmware on the ASAP² NIC with mlxconfig (from mstflint / MFT). LINK_TYPE value 2 selects Ethernet. Set NUM_OF_VFS to the maximum VF count you will use. The validated VF-LAG path does not accept a firmware VF count above 64. The num_vfs value you later set in OpenStackDeployment may be lower than the firmware maximum.

mlxconfig -d <PCI_BDF> set SRIOV_EN=1 NUM_OF_VFS=<VF_COUNT> \
  LINK_TYPE_P1=2 LINK_TYPE_P2=2
mlxfwreset -d <PCI_BDF> reset

If a firmware reset is not available, cold-reboot the host so the new registers take effect. Use the NIC model and firmware versions from Requirements. Replacement adapters typically ship with defaults that do not match this profile; configure them before you enable switchdev on the node.

Host IOMMU

After IOMMU is enabled in server BIOS/UEFI, enable it in the compute-node kernel so the hypervisor can assign VFs to guests. Add the flag that matches the CPU vendor:

  • Intel: intel_iommu=on

  • AMD: amd_iommu=on

Validation used AMD EPYC hosts, so use amd_iommu=on on that profile.

On a compute node that is already provisioned, use the grub_settings module provided by MOSK. See Modules provided by MOSK and the grub_settings module documentation. Add the module to a HostOSConfiguration object that selects the ASAP² compute nodes. The drop-in file the module writes takes precedence over /etc/default/grub for GRUB_CMDLINE_LINUX. Include any other command-line options that the drop-in must own.

Intel hosts:

spec:
  configs:
  - module: grub_settings
    moduleVersion: <MODULE_VERSION>
    values:
      options:
        grub_cmdline_linux:
        - intel_iommu=on

AMD hosts (validated EPYC profile):

spec:
  configs:
  - module: grub_settings
    moduleVersion: <MODULE_VERSION>
    values:
      options:
        grub_cmdline_linux:
        - amd_iommu=on

The module requests a host reboot. Once all the changes are applied, perform a graceful reboot using GracefulRebootRequest resource.

Switchdev bring-up

ASAP² requires the PFs to operate in switchdev mode so the mlx5e_rep driver creates one representor netdev per VF. Open vSwitch attaches those representors and programs hardware flows through Linux TC Flower. See Architecture.

The order of operations matters. The eSwitch must be in switchdev mode, representors must exist, hw-tc-offload must be on, and VFs must be rebound to mlx5_core before containerd and kubelet start, so that Kubernetes can schedule the Open vSwitch and Compute pods against a ready NIC. If those pods start first, representor attachment and VF PCI passthrough fail. On a node that is already provisioned, install this sequence so it runs at the next boot, then drain and reboot as described in Approach. Do not apply it live under running pods.

When using VF-LAG, switch both PFs of the same dual-port adapter into switchdev mode before you form the bond. Cross-NIC bonding is not supported.

Apply the following sequence on each ASAP² compute node, for every PF that will participate. These commands are illustrative. Persist them with the host automation you operate so they survive reboot.

  1. Create VFs on the PF:

    echo <VF_COUNT> > /sys/class/net/<PF_IFACE>/device/sriov_numvfs
    
  2. Unbind each newly created VF from mlx5_core so the eSwitch can change mode. VF PCI addresses appear as virtfn* under the PF sysfs device:

    echo <VF_PCI_BDF> > /sys/bus/pci/drivers/mlx5_core/unbind
    
  3. Set eSwitch mode to switchdev:

    devlink dev eswitch set pci/<PF_PCI_BDF> mode switchdev
    
  4. Enable hardware TC offload on the PF:

    ethtool -K <PF_IFACE> hw-tc-offload on
    
  5. Bring up the host bond and underlay VLAN as described in Bond and underlay.

  6. Rebind the VFs to mlx5_core after the underlay is online and before containerd and kubelet start:

    echo <VF_PCI_BDF> > /sys/bus/pci/drivers/mlx5_core/bind
    

Caution

The setup sequence must be repeated on startup. A one-time interactive configuration will not survive a node replacement or a standard host reboot unless you encode it in durable automation.

Note

Changing eSwitch mode on a NIC that already has leftover state can fail in the mlx5 driver. That happens during lab trials when you hop between legacy and switchdev, or between VF counts and bond layouts, without a reboot. Typical leftovers are an existing kernel bond on the PFs, VFs still bound to mlx5_core, or a non-zero sriov_numvfs from the previous attempt. devlink dev eswitch set ... mode switchdev then returns a driver error.

For those trials, return both PFs to a clean legacy state first: take the ASAP² bond down, unbind or destroy VFs (sriov_numvfs to 0), and reload the PF driver if the eSwitch does not reset. Then run the sequence above from the start, including hw-tc-offload. That cleanup is not a substitute for the durable boot order in Approach. On a provisioned compute node that will keep ASAP², persist the sequence, drain, and reboot. Do not use a live reset under running pods.

Bond and underlay

Reserve the ASAP² adapter for tenant offload traffic. Do not share that NIC with life-cycle management or storage networks. The validated layout is described in Topologies and use cases. VF-LAG is optional for ASAP² Direct, production deployments are highly likely to use it.

For VF-LAG, enslave both PFs of the same dual-port card into an IEEE 802.3ad (LACP) bond. Place the VXLAN underlay on a VLAN subinterface of that bond: that tagged physical network carries encapsulated tenant packets between compute nodes. Use a jumbo MTU on the bond and the underlay VLAN so encapsulated frames fit. The validated value was 9050 (<JUMBO_MTU> in the L2Template fragment below). Apply the same MTU on both PFs, the ASAP² bond, and the tenant underlay VLAN.

Declare the bond and underlay in the compute L2Template so MOSK LCM renders the Netplan into static configuration in the host. Do not form the bond with one-shot ip link commands, and do not add the bond or VF representors to Open vSwitch with ovs-vsctl. Neutron attaches representors. Set tunnel_interface in OpenStackDeployment to <TENANT_UNDERLAY_IFACE> as described in OpenStackDeployment configuration.

To order relative to switchdev:

Both PFs must be in switchdev mode before the bond comes up. Cross-NIC bonding is not supported. After the underlay is online, rebind VFs as described in Switchdev bring-up, before containerd and kubelet start.

On a node that is already provisioned, applying an L2Template drains the machine and re-runs LCM. Persist the switchdev sequence so it still runs in that order on the reboot LCM triggers. Do not rely on a live ip link set master change under running pods.

To declare the bond in L2Template:

Merge the following fragment into the compute L2Template npTemplate (Netplan) and add the underlay subnet to spec.l3Layout. Map N and M to the two PFs of the ASAP² NIC in the host interface mapping. The bond and VLAN names below are placeholders. Use the same names in OpenStackDeployment (tunnel_interface). See L2Template, L2 template example with bonds and bridges, and Create an L2 template for a MOSK compute node.

Configure the ToR as a matching LACP port-channel and trunk the underlay VLAN on that channel:

spec:
  l3Layout:
  - subnetName: <TENANT_UNDERLAY_SUBNET_ALIAS>
    scope: namespace
  npTemplate: |
    version: 2
    ethernets:
      {{nic N}}:
        dhcp4: false
        dhcp6: false
        match:
          macaddress: {{mac N}}
        set-name: {{nic N}}
        mtu: <JUMBO_MTU>
      {{nic M}}:
        dhcp4: false
        dhcp6: false
        match:
          macaddress: {{mac M}}
        set-name: {{nic M}}
        mtu: <JUMBO_MTU>
    bonds:
      <ASAP_BOND_IFACE>:
        dhcp4: false
        dhcp6: false
        interfaces:
        - {{nic N}}
        - {{nic M}}
        mtu: <JUMBO_MTU>
        parameters:
          mode: 802.3ad
          transmit-hash-policy: layer3+4
          mii-monitor-interval: 100
    vlans:
      <TENANT_UNDERLAY_IFACE>:
        id: <UNDERLAY_VLAN_ID>
        link: <ASAP_BOND_IFACE>
        mtu: <JUMBO_MTU>
        dhcp4: false
        dhcp6: false
        addresses:
        - {{ip "<TENANT_UNDERLAY_IFACE>:<TENANT_UNDERLAY_SUBNET_ALIAS>"}}

Where:

  • mode is required

  • mii-monitor-interval is recommended for MOSK bonds

  • transmit-hash-policy: layer3+4 matches the validated VF-LAG profile

To apply on a provisioned compute node:

For a machine that is already in the cluster, create a new L2Template that includes the fragment above (do not silently edit a template that is already in use), assign it, inspect the rendered candidate, then approve. For full procedure, see Modify network configuration on an existing machine.

Caution

Applying a new L2 template drains the node and re-runs LCM, the same class of disruption as a cluster update. Netplan cannot apply every arbitrary change. See the warnings on Modify network configuration on an existing machine.

If you can still set the host profile before first boot, put the same stanza in the compute L2Template used at provision time. See Create an L2 template for a new cluster.

Verify the bond:

On the compute node, confirm the bond is UP, both PFs are enslaved, LACP is active, and the underlay VLAN has the tunnel address. See also Host readiness.

ip link show <ASAP_BOND_IFACE>
ip link show <PF0_IFACE>
ip link show <PF1_IFACE>
cat /proc/net/bonding/<ASAP_BOND_IFACE>
ip -br addr show <TENANT_UNDERLAY_IFACE>

Host readiness

Confirm the following on the compute node before you rely on OVS or Nova to consume the devices:

  • devlink dev eswitch show pci/<PF_PCI_BDF> reports mode switchdev on each bonded PF.

  • ethtool -k <PF_IFACE> reports hw-tc-offload: on.

  • /sys/class/net/<PF_IFACE>/device/sriov_numvfs matches the VF count you configured.

  • Each VF /sys/bus/pci/devices/<VF_PCI_BDF>/driver symlink points to mlx5_core.

  • Representor netdevs exist and use the mlx5e_rep driver. Names vary by udev policy. Typical forms are a kernel name such as eth<N> with a predictable altname of the form enp<BUS>s<SLOT>f<FUNC>npf<PF>vf<VF>.

  • The ASAP² bond is UP and the underlay VLAN has the host tunnel address.

Example checks:

devlink dev eswitch show pci/<PF_PCI_BDF>
ethtool -k <PF_IFACE> | grep hw-tc-offload
ethtool -i <REPRESENTOR_IFACE>
ip -d link show <REPRESENTOR_IFACE>
ip link show <PF_IFACE>
ip link show <ASAP_BOND_IFACE>
ip -br addr show <TENANT_UNDERLAY_IFACE>

OpenStackDeployment configuration

Configure Networking SR-IOV, the Nova PCI device specification, and OVS hardware offload on the ASAP² compute nodes through OpenStackDeployment. Use a node selector that matches only those compute nodes. Merge the SR-IOV, PCI, and OVS fragments under the same spec.nodes override. The override style matches Enable SR-IOV with OVS/OVN and Configure PCI passthrough for guests. Disable the Neutron QoS service plugin and the Open vSwitch QoS agent extension at the cluster spec.services level, not on that node selector.

Enable the SR-IOV mechanism driver and declare the PFs that host switchdev VFs. For VF-LAG, list both PFs of the same dual-port adapter. Set tunnel_interface to <TENANT_UNDERLAY_IFACE>, the underlay VLAN that carries VXLAN. That mapping is the difference from classic SR-IOV: OVS must send overlay traffic out the ASAP² underlay so the eSwitch can offload encapsulation.

spec:
  nodes:
    <NODE-LABEL>::<NODE-LABEL-VALUE>:
      features:
        neutron:
          sriov:
            enabled: true
            nics:
            - device: <PF0_IFACE>
              num_vfs: <VF_COUNT>
              physnet: <PHYSNET>
              tunnel_interface: <TENANT_UNDERLAY_IFACE>
            - device: <PF1_IFACE>
              num_vfs: <VF_COUNT>
              physnet: <PHYSNET>

Allow Nova to assign ConnectX VFs to guests. Use pci.device_spec (the current nova.conf option; older snippets used the deprecated passthrough_whitelist name). Match the VF vendor and product IDs for the NIC you validated. The ConnectX-6 Lx Virtual Function IDs used in the test environment were vendor 15b3 and product 101e. Confirm IDs for your adapters in Requirements.

spec:
  nodes:
    <NODE-LABEL>::<NODE-LABEL-VALUE>:
      services:
        compute:
          nova:
            nova_compute:
              values:
                conf:
                  nova:
                    pci:
                      device_spec: |
                        [{"vendor_id": "15b3", "product_id": "101e"}]

Enable TC Flower hardware offload in the containerized OVS daemon and set idle aging for offloaded flows. max-idle is in milliseconds. The validated value 30000 ages idle hardware flows after 30 seconds, which matches the datapath description in Architecture. Apply these keys through OpenStackDeployment. Do not treat a host Netplan openvswitch.other-config stanza as a substitute.

spec:
  nodes:
    <NODE-LABEL>::<NODE-LABEL-VALUE>:
      openvswitch:
        values:
          conf:
            ovs_other_config:
              hw-offload: true
              max-idle: 30000

Disable Neutron QoS globally for the whole cluster. That includes omitting the qos service plugin and unloading the Open vSwitch agent extension. Disabling QoS only on ASAP² ports or ASAP² compute nodes is not enough. MOSK enables QoS by default. See Limitations.

Apply the following fragment at the cluster level spec.services. service_plugins is a single comma-separated string: copy the list from your OpenStack deployment and omit qos. The example below matches the MOSK ML2 default set without qos (including trunk, which is on by default). Keep any extra plugins you enabled, such as VPNaaS. Set agent.extensions to empty to unload the agent extension.

To obtain the current service_plugins and Open vSwitch agent extensions values from the Neutron Helm release:

kubectl -n osh-system exec -t deployment/rockoon -- \
  helm3 -n openstack get values --all openstack-neutron -o json | \
  jq -r .conf.neutron.DEFAULT.service_plugins

kubectl -n osh-system exec -t deployment/rockoon -- \
  helm3 -n openstack get values --all openstack-neutron -o json | \
  jq -r .conf.plugins.openvswitch_agent.agent.extensions
spec:
  services:
    networking:
      neutron:
        values:
          conf:
            neutron:
              DEFAULT:
                # Values from your OpenStack deployment (without qos)
                service_plugins: router,metering,trunk
            plugins:
              openvswitch_agent:
                agent:
                  # Values from your deployment (without qos)
                  extensions: ""

This spec.services override is a low-level Helm values merge. See OpenStackDeployment spec:services.

If you also attach direct ports to provider VLANs, map that physnet at the cluster spec.features.neutron.external_networks level. Use the same <PHYSNET> as in sriov.nics. interface is the ASAP² kernel bond from the compute L2Template, not the underlay VLAN. Do not enslave that bond into Open vSwitch from generic OVS startup; see Limitations. VLAN ranges are operator-chosen. <MTU> is the provider-network MTU. The lab used 9000, which is independent of <JUMBO_MTU> on the bond and underlay. Size both so encapsulated overlay traffic still fits on the underlay.

spec:
  features:
    neutron:
      external_networks:
      - physnet: <PHYSNET>
        interface: <ASAP_BOND_IFACE>
        bridge: <PROVIDER_BRIDGE>
        network_types:
        - vlan
        vlan_ranges: <VLAN_RANGES>
        mtu: <MTU>

Provider VLAN layout is in Topologies and use cases. The schema matches Enable SR-IOV with OVS/OVN.

Instance and port attachment

Create a VXLAN tenant network for accelerated workloads. Use addressing that fits your project; the values below are examples.

openstack network create <NETWORK> --provider-network-type vxlan

openstack subnet create <SUBNET> \
  --network <NETWORK> \
  --subnet-range <CIDR>

Accelerated ports must be SR-IOV direct ports with the switchdev capability. Disable port security. Neutron QoS must already be disabled for the cluster in OpenStackDeployment as shown in OpenStackDeployment configuration. Omitting a QoS policy on the port is not enough. Floating IPs in Distributed Virtual Routing (DVR) are not compatible with ASAP², associating a floating IP with a switchdev port fails. See Limitations.

openstack port create <PORT> \
  --network <NETWORK> \
  --vnic-type direct \
  --binding-profile '{"capabilities": ["switchdev"]}' \
  --disable-port-security

To exercise hardware offload across the fabric, create two such ports and launch one instance per port on different ASAP² compute hosts. The guest must run the NVIDIA/Mellanox driver; the VF appears as a PCI network device, not virtio-net. See Architecture.

openstack server create <INSTANCE> \
  --flavor <FLAVOR> \
  --image <IMAGE> \
  --port <PORT> \
  --availability-zone nova:<COMPUTE_HOST>

For provider VLAN testing, create a VLAN network on the physnet you mapped to the ASAP² adapter and attach a direct switchdev port the same way:

openstack network create <PROVIDER_NETWORK> \
  --provider-network-type vlan \
  --provider-physical-network <PHYSNET> \
  --provider-segment <VLAN_ID>

Validation checks

After instances obtain addresses, confirm tenant connectivity first, then prove that matching flows leave the host CPU.

From one guest, ping the peer and, if you want a sustained flow, generate traffic with iperf3 or a flood ping. Quantitative throughput results belong in Performance. This section only needs enough packets to install and hit hardware flows.

Dump offloaded datapath flows inside the openvswitch-vswitchd container on the compute node that hosts the instance. Use kubectl against the MOSK cluster. The Open vSwitch daemon runs as a DaemonSet in the openstack namespace, for example, openvswitch-openvswitch-vswitchd-default.

kubectl -n openstack get pods -o wide | grep vswitchd

kubectl -n openstack exec -it <OVS_VSWITCHD_POD> \
  -c openvswitch-vswitchd -- \
  ovs-appctl dpctl/dump-flows type=offloaded

kubectl -n openstack exec -it <OVS_VSWITCHD_POD> \
  -c openvswitch-vswitchd -- \
  ovs-appctl dpctl/offload-stats-show

A successful VXLAN offload shows type=offloaded (or the dedicated offloaded dump), a tunnel(...) match for the underlay, inner Ethernet or IP match fields, and non-zero packets / bytes counters while traffic runs.

To confirm host CPU bypass, run tcpdump on the VF representor that belongs to the instance:

tcpdump -i <REPRESENTOR_IFACE> -nn

The first packet of a new flow (the eSwitch miss) appears on the representor while OVS programs the hardware rule. Subsequent packets of that offloaded flow do not. ARP and other exception traffic can still be visible.

Re-check PF offload state and, optionally, TC Flower hardware installation on the representor. The in_hw flag indicates that the filter is in the eSwitch:

ethtool -k <PF_IFACE> | grep hw-tc-offload
tc -s filter show dev <REPRESENTOR_IFACE> ingress

Useful documentation