Performance
This section records throughput measured in the MOSK lab that validated NVIDIA ASAP² Direct. The numbers are indicative for that hardware, topology, and traffic profile. They are not product performance targets, NVIDIA SLAs, or a MOSK support commitment.
Test methodology
Traffic was east-west between two instances on different ASAP² compute nodes,
on one VXLAN tenant overlay. Each instance used a switchdev direct port
with port security disabled. The physical path was dual-port 25 GbE ConnectX-6
Lx VF-LAG (IEEE 802.3ad, transmit hash layer3+4) with jumbo MTU on the
ASAP² bond and underlay.
For the host layout and the validated overlay path, see Topologies and use cases. For hardware and firmware bounds, see Requirements.
Parameter |
Details |
|---|---|
Tool and duration |
Guests ran |
TCP profiles |
Parallel stream counts |
UDP profiles |
The same stream counts with unlimited offered load ( |
Guests |
Ubuntu cloud images on a 16 vCPU / 32 GB flavor. |
Baseline |
Software Open vSwitch in the host kernel on non-offloaded
Intel 25 GbE NICs, without hardware flow offload (labeled
|
Complementary checks |
These confirm that measured traffic is on the hardware path,
they are not extra throughput tables. Flood ping and a
sustained |
Results
TCP throughput is the iperf3 reported goodput. Retransmission hides drops,
so a TCP row is a single rate.
UDP rows list offered load, delivered bandwidth, and loss:
Offered is the rate the sender generated: how much UDP
iperf3tried to put on the wire, not how much the receiver got. With-u -b0that is as fast as the guest CPU and NIC allow. With-b 500Mit is the configured per-stream target (500 Mbit/s times the stream count).Delivered is the rate that arrived at the peer.
Loss is the share that did not arrive at the peer. UDP has no congestion control, so the path (eSwitch, VF-LAG, ToR, peer guest) can drop packets and offered can exceed delivered.
Compare UDP paths on delivered bandwidth. A higher offered rate alone is not a win if most of it is lost.
UDP, per-stream cap -b 500M, 30-second runs. Aggregate caps are 0.5 Gbit/s
(1 stream), 7.5 Gbit/s (15 streams), and 15 Gbit/s (30 streams). Hitting those
caps is not 25 GbE line rate.
Latency and host CPU
This validation did not tabulate sockperf
latency or hypervisor CPU percentages. Qualitatively, after OVS programs
the hardware rule, matching packets bypass the host datapath: they do
not appear on the VF representor. Only the first packet of a new flow
(and other exception traffic) hits the host CPU. See
Architecture.
Interpretation
Note
Treat the tables as a lab snapshot, not a capacity plan.
Consideration |
Details |
|---|---|
Hardware generation |
Results apply to dual-port ConnectX-6 Lx 25 GbE as tested. Do not assume the same rates on ConnectX-6 Dx, ConnectX-7, BlueField DPUs, or other vendors. |
Flow concurrency |
Profiles used 1, 15, and 30 synthetic |
Packet size |
Near line-rate TCP on large or jumbo frames does not imply lossless small-packet UDP. Unlimited 700-byte UDP at 15 streams delivered 10.8 Gbit/s with 35% loss. |
Encapsulation and features |
Acceleration was validated on VXLAN east-west between direct
|
Single-stream TCP |
A single stream (16.5 Gbit/s vs 2.76 Gbit/s) can be limited by guest
vCPU and queueing rather than by the ASIC. Multi-stream runs ( |
Reproducibility |
Use the same overlay, jumbo MTU, VF-LAG, disabled port security, and complementary offload checks. Different NICs, packet sizes, or OpenStack features will not reproduce these numbers. |