Configuration and activation
This section describes how to enable NVIDIA ASAP² Direct on MOSK compute nodes that use the Open vSwitch (OVS) kernel datapath. The procedure builds on classic SR-IOV with the Neutron OVS backend.
ASAP² additionally requires the Physical Functions (PFs) to run in
switchdev mode, VF representors for OVS, and hardware offload in the OVS
daemon.
Host-side switchdev bring-up is not yet a productized MOSK procedure. The
following steps outline the general approach to configuring the host OS and
NIC for ASAP² Direct offloading. Apply these steps with any durable host
automation you operate, for example, first-boot scripts, image customization,
or a custom module through the host OS configuration API.
See also
Starting point
The procedure assumes a MOSK compute node has already been provisioned:
The host is a compute node that has already fully joined the cluster: Kubernetes underlay components (
containerdandkubelet) are running, the node networking stack is up, and the Open vSwitch and Compute (Nova) pods are already scheduled on the node.An NVIDIA SmartNIC that meets Requirements is installed and cabled for tenant traffic. The adapter is not yet enabled for ASAP²: Physical Functions remain in legacy (non-
switchdev) mode, VFs and representors are not prepared for offload, andOpenStackDeploymentdoes not yet declare SR-IOV hardware offload on that node.Life-cycle management and storage traffic already use other NICs. The SmartNIC is reserved for the ASAP² path described in Topologies and use cases.
Approach
Host OS configuration for switchdev mode, VF representors,
hw-tc-offload, and VF rebind must be in place before containerd and
kubelet start, and therefore before the Open vSwitch and OpenStack pods
consume representors and VF PCI devices. If those services are already up, you
cannot insert that sequence into the current boot.
On a provisioned compute node, treat host enablement as persist, drain, and
reboot—not a live change under running pods. Firmware, kernel IOMMU, and
L2Template updates can each trigger a reboot. Install the switchdev
boot automation before the reboot that brings up the ASAP² bond so the order on
that boot is: early switchdev, then Netplan bond and underlay, then VF
rebind, then containerd and kubelet.
Host and LCM
Evacuate or otherwise remove instances from the compute node so no guest depends on the SmartNIC.
Enable server BIOS/UEFI options, then program the adapter with
mlxconfig. See Firmware and BIOS.Enable kernel IOMMU with the
grub_settingsmodule. See Host IOMMU.Install durable automation for the Switchdev bring-up sequence (VF create, unbind,
devlinkswitchdev,hw-tc-offload, VF rebind). It must run on every boot, around host networking, and finish beforecontainerd.serviceandkubelet.service. Do not restart these services in place. See Global recommendations for implementation of custom modules.Declare the VF-LAG bond and VXLAN underlay VLAN in a new compute
L2Templateand apply it as described in Bond and underlay. IPAM renders Netplan. Do not create the bond withip link.After the host is back, confirm Host readiness (eSwitch mode, offload, rebound VFs, representors, bond and underlay).
Note
During validation, Mirantis developed an experimental host OS
configuration module that automates host OS and NIC steps for switchdev
bring-up. The module uses systemd units to run that configuration in the
required order at node startup.
The module is not yet part of the product. Mirantis does not guarantee that it works for every software and hardware combination. You can use it as a reference for the operations described in switchdev Host OS configuration module.
OpenStack and workloads
After the host is in switchdev and the underlay is up:
Configure Networking, Compute, and OVS in the
OpenStackDeploymentcustom resource.Create accelerated networks and
switchdevports, then launch instances.Verify that eligible flows offload to the NIC.
Hardware, firmware, and software bounds are listed in Requirements. For bond, VLAN, and physnet layout, see Topologies and use cases. Functional constraints, such as disabled port security, Neutron QoS disabled globally, and no offload through Neutron routers or floating IPs, are listed in Limitations.
Host and NIC preparation
Firmware and BIOS
Enable the following options in the server BIOS/UEFI of each ASAP² compute
node (boot-time firmware setup or the out-of-band BMC equivalent). Menu names
vary by vendor. Complete this before the operating system creates Virtual
Functions (VFs) or changes eSwitch mode, and before you program the adapter
with mlxconfig:
SR-IOV support
IOMMU (Intel VT-d or AMD-Vi)
SR-IOV global / Alternative Routing-ID Interpretation (ARI), required for high VF counts
Memory-mapped I/O above 4G decoding, required for 64-bit PCI BAR allocation across dense VF sets
Then configure adapter firmware on the ASAP² NIC with mlxconfig (from
mstflint / MFT). LINK_TYPE value 2 selects Ethernet. Set
NUM_OF_VFS to the maximum VF count you will use. The validated VF-LAG path
does not accept a firmware VF count above 64. The num_vfs value you
later set in OpenStackDeployment may be lower than the firmware maximum.
mlxconfig -d <PCI_BDF> set SRIOV_EN=1 NUM_OF_VFS=<VF_COUNT> \
LINK_TYPE_P1=2 LINK_TYPE_P2=2
mlxfwreset -d <PCI_BDF> reset
If a firmware reset is not available, cold-reboot the host so the new registers
take effect. Use the NIC model and firmware versions from
Requirements. Replacement adapters typically ship
with defaults that do not match this profile; configure them before you enable
switchdev on the node.
Host IOMMU
After IOMMU is enabled in server BIOS/UEFI, enable it in the compute-node kernel so the hypervisor can assign VFs to guests. Add the flag that matches the CPU vendor:
Intel:
intel_iommu=onAMD:
amd_iommu=on
Validation used AMD EPYC hosts, so use amd_iommu=on on that profile.
On a compute node that is already provisioned, use the grub_settings module
provided by MOSK. See Modules provided by MOSK and the grub_settings module
documentation.
Add the module to a HostOSConfiguration object that selects the ASAP²
compute nodes. The drop-in file the module writes takes precedence over
/etc/default/grub for GRUB_CMDLINE_LINUX. Include any other
command-line options that the drop-in must own.
Intel hosts:
spec:
configs:
- module: grub_settings
moduleVersion: <MODULE_VERSION>
values:
options:
grub_cmdline_linux:
- intel_iommu=on
AMD hosts (validated EPYC profile):
spec:
configs:
- module: grub_settings
moduleVersion: <MODULE_VERSION>
values:
options:
grub_cmdline_linux:
- amd_iommu=on
The module requests a host reboot. Once all the changes are applied, perform a graceful reboot using GracefulRebootRequest resource.
Switchdev bring-up
ASAP² requires the PFs to operate in switchdev mode so the mlx5e_rep
driver creates one representor netdev per VF. Open vSwitch attaches those
representors and programs hardware flows through Linux TC Flower. See
Architecture.
The order of operations matters. The eSwitch must be in switchdev mode,
representors must exist, hw-tc-offload must be on, and VFs must be rebound
to mlx5_core before containerd and kubelet start, so that
Kubernetes can schedule the Open vSwitch and Compute pods against a ready NIC.
If those pods start first, representor attachment and VF PCI passthrough fail.
On a node that is already provisioned, install this sequence so it runs at the
next boot, then drain and reboot as described in Approach.
Do not apply it live under running pods.
When using VF-LAG, switch both PFs of the same dual-port adapter into
switchdev mode before you form the bond. Cross-NIC bonding is not
supported.
Apply the following sequence on each ASAP² compute node, for every PF that will participate. These commands are illustrative. Persist them with the host automation you operate so they survive reboot.
Create VFs on the PF:
echo <VF_COUNT> > /sys/class/net/<PF_IFACE>/device/sriov_numvfs
Unbind each newly created VF from
mlx5_coreso the eSwitch can change mode. VF PCI addresses appear asvirtfn*under the PF sysfs device:echo <VF_PCI_BDF> > /sys/bus/pci/drivers/mlx5_core/unbind
Set eSwitch mode to
switchdev:devlink dev eswitch set pci/<PF_PCI_BDF> mode switchdev
Enable hardware TC offload on the PF:
ethtool -K <PF_IFACE> hw-tc-offload on
Bring up the host bond and underlay VLAN as described in Bond and underlay.
Rebind the VFs to
mlx5_coreafter the underlay is online and beforecontainerdandkubeletstart:echo <VF_PCI_BDF> > /sys/bus/pci/drivers/mlx5_core/bind
Caution
The setup sequence must be repeated on startup. A one-time interactive configuration will not survive a node replacement or a standard host reboot unless you encode it in durable automation.
Note
Changing eSwitch mode on a NIC that already has leftover state can
fail in the mlx5 driver. That happens during lab trials when you hop
between legacy and switchdev, or between VF counts and bond layouts,
without a reboot. Typical leftovers are an existing kernel bond on the PFs,
VFs still bound to mlx5_core, or a non-zero sriov_numvfs from the
previous attempt. devlink dev eswitch set ... mode switchdev then
returns a driver error.
For those trials, return both PFs to a clean legacy state first: take
the ASAP² bond down, unbind or destroy VFs (sriov_numvfs to 0), and
reload the PF driver if the eSwitch does not reset. Then run the sequence
above from the start, including hw-tc-offload. That cleanup is not a
substitute for the durable boot order in Approach. On a
provisioned compute node that will keep ASAP², persist the sequence, drain,
and reboot. Do not use a live reset under running pods.
Bond and underlay
Reserve the ASAP² adapter for tenant offload traffic. Do not share that NIC with life-cycle management or storage networks. The validated layout is described in Topologies and use cases. VF-LAG is optional for ASAP² Direct, production deployments are highly likely to use it.
For VF-LAG, enslave both PFs of the same dual-port card into an IEEE
802.3ad (LACP) bond. Place the VXLAN underlay on a VLAN subinterface of that
bond: that tagged physical network carries encapsulated tenant packets between
compute nodes. Use a jumbo MTU on the bond and the underlay VLAN so
encapsulated frames fit. The validated value was 9050 (<JUMBO_MTU> in
the L2Template fragment below). Apply the same MTU on both PFs, the ASAP² bond,
and the tenant underlay VLAN.
Declare the bond and underlay in the compute L2Template so MOSK LCM renders
the Netplan into static configuration in the host. Do not form the bond with
one-shot ip link commands, and do not add the bond or VF representors to
Open vSwitch with ovs-vsctl. Neutron attaches representors. Set
tunnel_interface in OpenStackDeployment to <TENANT_UNDERLAY_IFACE>
as described in
OpenStackDeployment configuration.
To order relative to switchdev:
Both PFs must be in switchdev mode before the bond comes up. Cross-NIC
bonding is not supported. After the underlay is online, rebind VFs as described
in Switchdev bring-up, before containerd and
kubelet start.
On a node that is already provisioned, applying an L2Template drains the
machine and re-runs LCM. Persist the switchdev sequence so it still runs in
that order on the reboot LCM triggers. Do not rely on a live ip link set
master change under running pods.
To declare the bond in L2Template:
Merge the following fragment into the compute L2Template npTemplate
(Netplan) and add the underlay subnet to spec.l3Layout. Map N and M
to the two PFs of the ASAP² NIC in the host interface mapping. The bond and
VLAN names below are placeholders. Use the same names in
OpenStackDeployment (tunnel_interface). See L2Template,
L2 template example with bonds and bridges, and
Create an L2 template for a MOSK compute node.
Configure the ToR as a matching LACP port-channel and trunk the underlay VLAN on that channel:
spec:
l3Layout:
- subnetName: <TENANT_UNDERLAY_SUBNET_ALIAS>
scope: namespace
npTemplate: |
version: 2
ethernets:
{{nic N}}:
dhcp4: false
dhcp6: false
match:
macaddress: {{mac N}}
set-name: {{nic N}}
mtu: <JUMBO_MTU>
{{nic M}}:
dhcp4: false
dhcp6: false
match:
macaddress: {{mac M}}
set-name: {{nic M}}
mtu: <JUMBO_MTU>
bonds:
<ASAP_BOND_IFACE>:
dhcp4: false
dhcp6: false
interfaces:
- {{nic N}}
- {{nic M}}
mtu: <JUMBO_MTU>
parameters:
mode: 802.3ad
transmit-hash-policy: layer3+4
mii-monitor-interval: 100
vlans:
<TENANT_UNDERLAY_IFACE>:
id: <UNDERLAY_VLAN_ID>
link: <ASAP_BOND_IFACE>
mtu: <JUMBO_MTU>
dhcp4: false
dhcp6: false
addresses:
- {{ip "<TENANT_UNDERLAY_IFACE>:<TENANT_UNDERLAY_SUBNET_ALIAS>"}}
Where:
modeis requiredmii-monitor-intervalis recommended for MOSK bondstransmit-hash-policy: layer3+4matches the validated VF-LAG profile
To apply on a provisioned compute node:
For a machine that is already in the cluster, create a new L2Template
that includes the fragment above (do not silently edit a template that is
already in use), assign it, inspect the rendered candidate, then approve. For
full procedure, see
Modify network configuration on an existing machine.
Caution
Applying a new L2 template drains the node and re-runs LCM, the same class of disruption as a cluster update. Netplan cannot apply every arbitrary change. See the warnings on Modify network configuration on an existing machine.
If you can still set the host profile before first boot, put the same stanza in
the compute L2Template used at provision time. See
Create an L2 template for a new cluster.
Verify the bond:
On the compute node, confirm the bond is UP, both PFs are enslaved, LACP is
active, and the underlay VLAN has the tunnel address. See also
Host readiness.
ip link show <ASAP_BOND_IFACE>
ip link show <PF0_IFACE>
ip link show <PF1_IFACE>
cat /proc/net/bonding/<ASAP_BOND_IFACE>
ip -br addr show <TENANT_UNDERLAY_IFACE>
Host readiness
Confirm the following on the compute node before you rely on OVS or Nova to consume the devices:
devlink dev eswitch show pci/<PF_PCI_BDF>reportsmode switchdevon each bonded PF.ethtool -k <PF_IFACE>reportshw-tc-offload: on./sys/class/net/<PF_IFACE>/device/sriov_numvfsmatches the VF count you configured.Each VF
/sys/bus/pci/devices/<VF_PCI_BDF>/driversymlink points tomlx5_core.Representor netdevs exist and use the
mlx5e_repdriver. Names vary by udev policy. Typical forms are a kernel name such aseth<N>with a predictablealtnameof the formenp<BUS>s<SLOT>f<FUNC>npf<PF>vf<VF>.The ASAP² bond is
UPand the underlay VLAN has the host tunnel address.
Example checks:
devlink dev eswitch show pci/<PF_PCI_BDF>
ethtool -k <PF_IFACE> | grep hw-tc-offload
ethtool -i <REPRESENTOR_IFACE>
ip -d link show <REPRESENTOR_IFACE>
ip link show <PF_IFACE>
ip link show <ASAP_BOND_IFACE>
ip -br addr show <TENANT_UNDERLAY_IFACE>
OpenStackDeployment configuration
Configure Networking SR-IOV, the Nova PCI device specification, and OVS
hardware offload on the ASAP² compute nodes through OpenStackDeployment.
Use a node selector that matches only those compute nodes. Merge the SR-IOV,
PCI, and OVS fragments under the same spec.nodes override. The override
style matches
Enable SR-IOV with OVS/OVN and Configure PCI passthrough for guests. Disable the
Neutron QoS service plugin and the Open vSwitch QoS agent extension at the
cluster spec.services level, not on that node selector.
Enable the SR-IOV mechanism driver and declare the PFs that host switchdev VFs.
For VF-LAG, list both PFs of the same dual-port adapter. Set
tunnel_interface to <TENANT_UNDERLAY_IFACE>, the underlay VLAN that
carries VXLAN. That mapping is the difference from classic SR-IOV: OVS must
send overlay traffic out the ASAP² underlay so the eSwitch can offload
encapsulation.
spec:
nodes:
<NODE-LABEL>::<NODE-LABEL-VALUE>:
features:
neutron:
sriov:
enabled: true
nics:
- device: <PF0_IFACE>
num_vfs: <VF_COUNT>
physnet: <PHYSNET>
tunnel_interface: <TENANT_UNDERLAY_IFACE>
- device: <PF1_IFACE>
num_vfs: <VF_COUNT>
physnet: <PHYSNET>
Allow Nova to assign ConnectX VFs to guests. Use pci.device_spec (the
current nova.conf option; older snippets used the deprecated
passthrough_whitelist name). Match the VF vendor and product IDs for the
NIC you validated. The ConnectX-6 Lx Virtual Function IDs used in the test
environment were vendor 15b3 and product 101e. Confirm IDs for your
adapters in Requirements.
spec:
nodes:
<NODE-LABEL>::<NODE-LABEL-VALUE>:
services:
compute:
nova:
nova_compute:
values:
conf:
nova:
pci:
device_spec: |
[{"vendor_id": "15b3", "product_id": "101e"}]
Enable TC Flower hardware offload in the containerized OVS daemon and set idle
aging for offloaded flows. max-idle is in milliseconds. The validated value
30000 ages idle hardware flows after 30 seconds, which matches the datapath
description in
Architecture. Apply these keys through
OpenStackDeployment. Do not treat a host Netplan
openvswitch.other-config stanza as a substitute.
spec:
nodes:
<NODE-LABEL>::<NODE-LABEL-VALUE>:
openvswitch:
values:
conf:
ovs_other_config:
hw-offload: true
max-idle: 30000
Disable Neutron QoS globally for the whole cluster. That includes omitting the
qos service plugin and unloading the Open vSwitch agent extension.
Disabling QoS only on ASAP² ports or ASAP² compute nodes is not enough. MOSK
enables QoS by default. See
Limitations.
Apply the following fragment at the cluster level spec.services.
service_plugins is a single comma-separated string: copy the list from your
OpenStack deployment and omit qos. The example below matches the MOSK ML2
default set without qos (including trunk, which is on by default). Keep
any extra plugins you enabled, such as VPNaaS. Set agent.extensions to
empty to unload the agent extension.
To obtain the current service_plugins and Open vSwitch agent extensions
values from the Neutron Helm release:
kubectl -n osh-system exec -t deployment/rockoon -- \
helm3 -n openstack get values --all openstack-neutron -o json | \
jq -r .conf.neutron.DEFAULT.service_plugins
kubectl -n osh-system exec -t deployment/rockoon -- \
helm3 -n openstack get values --all openstack-neutron -o json | \
jq -r .conf.plugins.openvswitch_agent.agent.extensions
spec:
services:
networking:
neutron:
values:
conf:
neutron:
DEFAULT:
# Values from your OpenStack deployment (without qos)
service_plugins: router,metering,trunk
plugins:
openvswitch_agent:
agent:
# Values from your deployment (without qos)
extensions: ""
This spec.services override is a low-level Helm values merge. See
OpenStackDeployment spec:services.
If you also attach direct ports to provider VLANs, map that physnet at the
cluster spec.features.neutron.external_networks level. Use the same
<PHYSNET> as in sriov.nics. interface is the ASAP² kernel bond from
the compute L2Template, not the underlay VLAN. Do not enslave that bond
into Open vSwitch from generic OVS startup; see
Limitations. VLAN ranges are operator-chosen. <MTU> is
the provider-network MTU. The lab used 9000, which is independent of
<JUMBO_MTU> on the bond and underlay. Size both so encapsulated overlay
traffic still fits on the underlay.
spec:
features:
neutron:
external_networks:
- physnet: <PHYSNET>
interface: <ASAP_BOND_IFACE>
bridge: <PROVIDER_BRIDGE>
network_types:
- vlan
vlan_ranges: <VLAN_RANGES>
mtu: <MTU>
Provider VLAN layout is in Topologies and use cases. The schema matches Enable SR-IOV with OVS/OVN.
Instance and port attachment
Create a VXLAN tenant network for accelerated workloads. Use addressing that fits your project; the values below are examples.
openstack network create <NETWORK> --provider-network-type vxlan
openstack subnet create <SUBNET> \
--network <NETWORK> \
--subnet-range <CIDR>
Accelerated ports must be SR-IOV direct ports with the switchdev
capability. Disable port security. Neutron QoS must already be disabled for the
cluster in OpenStackDeployment as shown in
OpenStackDeployment configuration. Omitting a QoS policy on the
port is not enough. Floating IPs in Distributed Virtual Routing (DVR) are not
compatible with ASAP², associating a floating IP with a switchdev port
fails. See Limitations.
openstack port create <PORT> \
--network <NETWORK> \
--vnic-type direct \
--binding-profile '{"capabilities": ["switchdev"]}' \
--disable-port-security
To exercise hardware offload across the fabric, create two such ports and
launch one instance per port on different ASAP² compute hosts. The guest must
run the NVIDIA/Mellanox driver; the VF appears as a PCI network device, not
virtio-net. See
Architecture.
openstack server create <INSTANCE> \
--flavor <FLAVOR> \
--image <IMAGE> \
--port <PORT> \
--availability-zone nova:<COMPUTE_HOST>
For provider VLAN testing, create a VLAN network on the physnet you mapped to
the ASAP² adapter and attach a direct switchdev port the same way:
openstack network create <PROVIDER_NETWORK> \
--provider-network-type vlan \
--provider-physical-network <PHYSNET> \
--provider-segment <VLAN_ID>
Validation checks
After instances obtain addresses, confirm tenant connectivity first, then prove that matching flows leave the host CPU.
From one guest, ping the peer and, if you want a sustained flow, generate
traffic with iperf3 or a flood ping. Quantitative throughput results belong
in Performance. This section only needs enough packets to
install and hit hardware flows.
Dump offloaded datapath flows inside the openvswitch-vswitchd container on
the compute node that hosts the instance. Use kubectl against the MOSK
cluster. The Open vSwitch daemon runs as a DaemonSet in the openstack
namespace, for example, openvswitch-openvswitch-vswitchd-default.
kubectl -n openstack get pods -o wide | grep vswitchd
kubectl -n openstack exec -it <OVS_VSWITCHD_POD> \
-c openvswitch-vswitchd -- \
ovs-appctl dpctl/dump-flows type=offloaded
kubectl -n openstack exec -it <OVS_VSWITCHD_POD> \
-c openvswitch-vswitchd -- \
ovs-appctl dpctl/offload-stats-show
A successful VXLAN offload shows type=offloaded (or the dedicated offloaded
dump), a tunnel(...) match for the underlay, inner Ethernet or IP match
fields, and non-zero packets / bytes counters while traffic runs.
To confirm host CPU bypass, run tcpdump on the VF representor that belongs
to the instance:
tcpdump -i <REPRESENTOR_IFACE> -nn
The first packet of a new flow (the eSwitch miss) appears on the representor while OVS programs the hardware rule. Subsequent packets of that offloaded flow do not. ARP and other exception traffic can still be visible.
Re-check PF offload state and, optionally, TC Flower hardware installation on
the representor. The in_hw flag indicates that the filter is in the
eSwitch:
ethtool -k <PF_IFACE> | grep hw-tc-offload
tc -s filter show dev <REPRESENTOR_IFACE> ingress