Performance

This section records throughput measured in the MOSK lab that validated NVIDIA ASAP² Direct. The numbers are indicative for that hardware, topology, and traffic profile. They are not product performance targets, NVIDIA SLAs, or a MOSK support commitment.

Test methodology

Traffic was east-west between two instances on different ASAP² compute nodes, on one VXLAN tenant overlay. Each instance used a switchdev direct port with port security disabled. The physical path was dual-port 25 GbE ConnectX-6 Lx VF-LAG (IEEE 802.3ad, transmit hash layer3+4) with jumbo MTU on the ASAP² bond and underlay.

For the host layout and the validated overlay path, see Topologies and use cases. For hardware and firmware bounds, see Requirements.

Test methodology parameters

Parameter

Details

Tool and duration

Guests ran iperf3 for 30 seconds per profile.

TCP profiles

Parallel stream counts -P1, -P15, and -P30.

UDP profiles

The same stream counts with unlimited offered load (-u -b0) and with a per-stream cap (-b 500M). A subset of UDP runs used 700-byte payloads (-l700).

Guests

Ubuntu cloud images on a 16 vCPU / 32 GB flavor.

Baseline

Software Open vSwitch in the host kernel on non-offloaded Intel 25 GbE NICs, without hardware flow offload (labeled NIC standard in the tables). The ASAP² column is the ConnectX-6 Lx switchdev path (NIC asap2).

Complementary checks

These confirm that measured traffic is on the hardware path, they are not extra throughput tables. Flood ping and a sustained iperf3 flow install TC Flower rules. After the first packet of a matching flow, representor tcpdump on the host shows no further packets of that flow. ovs-appctl dpctl/dump-flows type=offloaded shows non-zero hardware packet and byte counters. Commands are documented in Configuration and activation.

Results

TCP throughput is the iperf3 reported goodput. Retransmission hides drops, so a TCP row is a single rate.

UDP rows list offered load, delivered bandwidth, and loss:

  • Offered is the rate the sender generated: how much UDP iperf3 tried to put on the wire, not how much the receiver got. With -u -b0 that is as fast as the guest CPU and NIC allow. With -b 500M it is the configured per-stream target (500 Mbit/s times the stream count).

  • Delivered is the rate that arrived at the peer.

  • Loss is the share that did not arrive at the peer. UDP has no congestion control, so the path (eSwitch, VF-LAG, ToR, peer guest) can drop packets and offered can exceed delivered.

Compare UDP paths on delivered bandwidth. A higher offered rate alone is not a win if most of it is lost.

TCP throughput, 30-second runs, ASAP² versus kernel OVS baseline

Profile

ASAP² (NIC asap2)

Baseline (NIC standard)

Approximate gain

1 stream (-P1)

16.5 Gbit/s

2.76 Gbit/s

~6×

15 streams (-P15)

22.7 Gbit/s

11.5 Gbit/s

~2×

30 streams (-P30)

22.8 Gbit/s

14.9 Gbit/s

ASAP² near 25 GbE line rate

UDP unlimited offered load (-u -b0), 30-second runs, offered versus delivered

Profile

ASAP² offered / delivered / loss

Baseline offered / delivered / loss

Approximate delivered gain

1 stream (-P1)

3.18 Gbit/s / 2.34 Gbit/s / 27%

891 Mbit/s / 890 Mbit/s / 0.046%

~2.6×

15 streams (-P15)

23.1 Gbit/s / 19.2 Gbit/s / 17%

1.01 Gbit/s / 980 Mbit/s / 2.7%

~20×

30 streams (-P30)

23.1 Gbit/s / 18.4 Gbit/s / 20%

2.87 Gbit/s / 2.70 Gbit/s / 5.7%

~6.8×

15 streams, 700 B (-P15 -l700)

16.6 Gbit/s / 10.8 Gbit/s / 35%

903 Mbit/s / 302 Mbit/s / 0.00017%

~36×

30 streams, 700 B (-P30 -l700)

16.6 Gbit/s / 10.6 Gbit/s / 33%

1.28 Gbit/s / 1.27 Gbit/s / 0.3%

~8.3×

UDP, per-stream cap -b 500M, 30-second runs. Aggregate caps are 0.5 Gbit/s (1 stream), 7.5 Gbit/s (15 streams), and 15 Gbit/s (30 streams). Hitting those caps is not 25 GbE line rate.

UDP capped load (500 Mbit/s per stream)

Profile

ASAP² offered / delivered / loss

Baseline offered / delivered / loss

Approximate delivered gain

1 stream

500 Mbit/s / 499 Mbit/s / 0.12%

500 Mbit/s / 500 Mbit/s / 0%

Parity at this cap

15 streams

7.50 Gbit/s / 7.50 Gbit/s / 0.022%

2.31 Gbit/s / 2.30 Gbit/s / 0.012%

~3.3× (ASAP² at the 7.5 Gbit/s cap)

30 streams

7.50 Gbit/s / 7.49 Gbit/s / 0.058%

2.85 Gbit/s / 2.72 Gbit/s / 4.5%

~2.8×

1 stream, 700 B

500 Mbit/s / 500 Mbit/s / 0.00011%

446 Mbit/s / 445 Mbit/s / 0.14%

~1.1×

15 streams, 700 B

3.56 Gbit/s / 3.56 Gbit/s / 0.0094%

895 Mbit/s / 895 Mbit/s / 0.0061%

~4.0×

30 streams, 700 B

11.7 Gbit/s / 9.61 Gbit/s / 18%

1.37 Gbit/s / 1.34 Gbit/s / 1.6%

~7.2×

Latency and host CPU

This validation did not tabulate sockperf latency or hypervisor CPU percentages. Qualitatively, after OVS programs the hardware rule, matching packets bypass the host datapath: they do not appear on the VF representor. Only the first packet of a new flow (and other exception traffic) hits the host CPU. See Architecture.

Interpretation

Note

Treat the tables as a lab snapshot, not a capacity plan.

Interpretation considerations

Consideration

Details

Hardware generation

Results apply to dual-port ConnectX-6 Lx 25 GbE as tested. Do not assume the same rates on ConnectX-6 Dx, ConnectX-7, BlueField DPUs, or other vendors.

Flow concurrency

Profiles used 1, 15, and 30 synthetic iperf3 streams. They do not characterize tens of thousands of production flows or eSwitch forwarding-database pressure.

Packet size

Near line-rate TCP on large or jumbo frames does not imply lossless small-packet UDP. Unlimited 700-byte UDP at 15 streams delivered 10.8 Gbit/s with 35% loss.

Encapsulation and features

Acceleration was validated on VXLAN east-west between direct switchdev ports. Traffic that traverses Neutron routers, floating-IP NAT, or security groups stays on the software path. See Limitations.

Single-stream TCP

A single stream (16.5 Gbit/s vs 2.76 Gbit/s) can be limited by guest vCPU and queueing rather than by the ASIC. Multi-stream runs (-P15 / -P30) are the fairer comparison of offload versus kernel OVS.

Reproducibility

Use the same overlay, jumbo MTU, VF-LAG, disabled port security, and complementary offload checks. Different NICs, packet sizes, or OpenStack features will not reproduce these numbers.