Help with ConnectX 3 SR-IOV with Linux host and windows guest via kvm

Notice: Page may contain affiliate links for which we may earn a small commission through services like Amazon Affiliates or Skimlinks.

okrasit

Member
Jun 28, 2019
40
32
18
I did that and they both came back exactly as they were before.
I've tried with default Windows driver and with MLNX_VPI_WinOF-5_50_53000_All_Win2019_x64, in which case there was a driver version mismatch between VPI and ETH. So what's the correct process here?
The microsoft driver should work, so it points in the direction of linux ofed :(
I did get that error 10, when I had winof driver on the VPI and microsoft one on the ethernet device.
 

okrasit

Member
Jun 28, 2019
40
32
18
Here's one important thing to notice:
1623276568604.png
If the "fast start-up" (fast boot) is enabled in windows, the mlx vpi device will get the error 43, requiring reboot or disable/enable toggling. That is, with a patched kernel driver or a working linux ofed (if such a thing even exists anymore).
 

dwalme

New Member
Jan 16, 2018
6
0
1
47
It's been a couple years and I am running into this same issue. Proxmox 8.3 running 6.8 or 6.11 kernels that have the broken 4.0-0 mlx4 kernel driver.

My Linux skills are limited. Before I dive into figuring out how to use this patch does anyone know offhand if this will still work with newer kernels?
 

Basriram

New Member
Sep 16, 2021
3
7
3
It's been a couple years and I am running into this same issue. Proxmox 8.3 running 6.8 or 6.11 kernels that have the broken 4.0-0 mlx4 kernel driver.

My Linux skills are limited. Before I dive into figuring out how to use this patch does anyone know offhand if this will still work with newer kernels?
I have created a dirty dkms patch script of the above mentioned patch for proxmox running 6.8.12-5 here

Just make sure you have linux-headers installed for your kernel version. You can ignore the "patch unexpectedly ends in middle of line" remarks but to be safe you may want to do a dry run first.
 

phyerbarte

New Member
Dec 30, 2024
2
0
1
I have created a dirty dkms patch script of the above mentioned patch for proxmox running 6.8.12-5 here

Just make sure you have linux-headers installed for your kernel version. You can ignore the "patch unexpectedly ends in middle of line" remarks but to be safe you may want to do a dry run first.
Hi, I tried your patch with PVE, which kernal is 6.8.12-5, but seems the error 43 still there, is that means it not support ? How can I check if the patch correctly ? The scripts seems running correctly.

Bash:
modinfo  mlx4_core
filename:       /lib/modules/6.8.12-5-pve/updates/dkms/mlx4_core.ko
version:        99.4.0-0
license:        Dual BSD/GPL
description:    Mellanox ConnectX HCA low-level driver
author:         Roland Dreier
srcversion:     4EC12DB41AD0C1ADF790DF4
alias:          pci:v000015B3d00001010sv*sd*bc*sc*i*
alias:          pci:v000015B3d0000100Fsv*sd*bc*sc*i*
alias:          pci:v000015B3d0000100Esv*sd*bc*sc*i*
alias:          pci:v000015B3d0000100Dsv*sd*bc*sc*i*
alias:          pci:v000015B3d0000100Csv*sd*bc*sc*i*
alias:          pci:v000015B3d0000100Bsv*sd*bc*sc*i*
alias:          pci:v000015B3d0000100Asv*sd*bc*sc*i*
alias:          pci:v000015B3d00001009sv*sd*bc*sc*i*
alias:          pci:v000015B3d00001008sv*sd*bc*sc*i*
alias:          pci:v000015B3d00001007sv*sd*bc*sc*i*
alias:          pci:v000015B3d00001006sv*sd*bc*sc*i*
alias:          pci:v000015B3d00001005sv*sd*bc*sc*i*
alias:          pci:v000015B3d00001004sv*sd*bc*sc*i*
alias:          pci:v000015B3d00001003sv*sd*bc*sc*i*
alias:          pci:v000015B3d00001002sv*sd*bc*sc*i*
alias:          pci:v000015B3d0000676Esv*sd*bc*sc*i*
alias:          pci:v000015B3d00006746sv*sd*bc*sc*i*
alias:          pci:v000015B3d00006764sv*sd*bc*sc*i*
alias:          pci:v000015B3d0000675Asv*sd*bc*sc*i*
alias:          pci:v000015B3d00006372sv*sd*bc*sc*i*
alias:          pci:v000015B3d00006750sv*sd*bc*sc*i*
alias:          pci:v000015B3d00006368sv*sd*bc*sc*i*
alias:          pci:v000015B3d0000673Csv*sd*bc*sc*i*
alias:          pci:v000015B3d00006732sv*sd*bc*sc*i*
alias:          pci:v000015B3d00006354sv*sd*bc*sc*i*
alias:          pci:v000015B3d0000634Asv*sd*bc*sc*i*
alias:          pci:v000015B3d00006340sv*sd*bc*sc*i*
depends:
retpoline:      Y
name:           mlx4_core
vermagic:       6.8.12-5-pve SMP preempt mod_unload modversions
sig_id:         PKCS#7
signer:         DKMS module signing key
sig_key:        2D:F8:AE:B4:09:A5:69:3A:96:9A:15:62:C1:0C:F8:7F:E0:FD:B4:A1
sig_hashalgo:   sha512
signature:      58:66:CA:6E:97:EB:0A:09:49:BA:E1:AE:A5:6D:33:8E:87:EB:59:B4:
                E0:A7:81:68:EB:82:FC:29:25:41:06:88:E0:D6:33:C0:FD:68:AB:89:
                89:23:71:20:C2:6B:9A:EE:CE:E5:76:8E:A8:D8:B4:72:E5:28:BF:BA:
                1E:81:94:A6:F1:80:95:DF:F0:B1:12:57:2E:90:B6:D0:54:37:D3:EE:
                19:BA:54:D7:13:62:66:0E:5A:56:E2:94:B6:4E:0F:AD:A5:A2:37:6E:
                99:78:D0:A3:FA:CB:09:99:86:BA:5A:59:F4:A9:C2:EC:89:0F:6A:C2:
                DB:A8:3D:B0:F8:1F:BB:A8:D6:D4:37:B1:3A:2B:3A:03:5E:91:C7:9B:
                B5:10:C3:52:68:71:04:14:12:FB:37:C2:EE:03:5A:C9:CB:4A:4A:E2:
                1F:41:9F:0B:9E:8B:98:C9:EE:20:46:EF:C8:9D:79:71:67:F3:13:99:
                43:2F:53:06:48:3C:AB:E5:0C:12:DF:DC:18:6A:6D:12:03:05:0F:D7:
                0D:EC:88:62:1A:CD:65:38:6A:E6:90:15:EE:D5:90:EB:41:12:66:01:
                D7:F9:3A:38:DB:82:8C:AB:BB:4F:AC:1A:CE:D8:1E:B7:A8:46:F2:C9:
                F7:11:0B:47:84:09:1C:CC:A3:3B:2A:15:01:48:61:2B
parm:           debug_level:Enable debug tracing if > 0 (int)
parm:           msi_x:0 - don't use MSI-X, 1 - use MSI-X, >1 - limit number of MSI-X irqs to msi_x (int)
parm:           num_vfs:enable #num_vfs functions if num_vfs > 0
num_vfs=port1,port2,port1+2 (array of byte)
parm:           probe_vf:number of vfs to probe by pf driver (num_vfs > 0)
probe_vf=port1,port2,port1+2 (array of byte)
parm:           log_num_mgm_entry_size:log mgm size, that defines the num of qp per mcg, for example: 10 gives 248.range: 7 <= log_num_mgm_entry_size <= 12. To activate device managed flow steering when available, set to -1 (int)
parm:           enable_64b_cqe_eqe:Enable 64 byte CQEs/EQEs when the FW supports this (default: True) (bool)
parm:           enable_4k_uar:Enable using 4K UAR. Should not be enabled if have VFs which do not support 4K UARs (default: false) (bool)
parm:           log_num_mac:Log2 max number of MACs per ETH port (1-7) (int)
parm:           log_num_vlan:Log2 max number of VLANs per ETH port (0-7) (int)
parm:           use_prio:Enable steering by VLAN priority on ETH ports (deprecated) (bool)
parm:           log_mtts_per_seg:Log2 number of MTT entries per segment (0-7) (default: 0) (int)
parm:           port_type_array:Array of port types: HW_DEFAULT (0) is default 1 for IB, 2 for Ethernet (array of int)
parm:           enable_qos:Enable Enhanced QoS support (default: off) (bool)
parm:           internal_err_reset:Reset device on internal errors if non-zero (default 1) (int)
 

Basriram

New Member
Sep 16, 2021
3
7
3
Hi, I tried your patch with PVE, which kernal is 6.8.12-5, but seems the error 43 still there, is that means it not support ? How can I check if the patch correctly ? The scripts seems running correctly.

Bash:
modinfo  mlx4_core
filename:       /lib/modules/6.8.12-5-pve/updates/dkms/mlx4_core.ko
version:        99.4.0-0
license:        Dual BSD/GPL
description:    Mellanox ConnectX HCA low-level driver
author:         Roland Dreier
srcversion:     4EC12DB41AD0C1ADF790DF4
alias:          pci:v000015B3d00001010sv*sd*bc*sc*i*
alias:          pci:v000015B3d0000100Fsv*sd*bc*sc*i*
alias:          pci:v000015B3d0000100Esv*sd*bc*sc*i*
alias:          pci:v000015B3d0000100Dsv*sd*bc*sc*i*
alias:          pci:v000015B3d0000100Csv*sd*bc*sc*i*
alias:          pci:v000015B3d0000100Bsv*sd*bc*sc*i*
alias:          pci:v000015B3d0000100Asv*sd*bc*sc*i*
alias:          pci:v000015B3d00001009sv*sd*bc*sc*i*
alias:          pci:v000015B3d00001008sv*sd*bc*sc*i*
alias:          pci:v000015B3d00001007sv*sd*bc*sc*i*
alias:          pci:v000015B3d00001006sv*sd*bc*sc*i*
alias:          pci:v000015B3d00001005sv*sd*bc*sc*i*
alias:          pci:v000015B3d00001004sv*sd*bc*sc*i*
alias:          pci:v000015B3d00001003sv*sd*bc*sc*i*
alias:          pci:v000015B3d00001002sv*sd*bc*sc*i*
alias:          pci:v000015B3d0000676Esv*sd*bc*sc*i*
alias:          pci:v000015B3d00006746sv*sd*bc*sc*i*
alias:          pci:v000015B3d00006764sv*sd*bc*sc*i*
alias:          pci:v000015B3d0000675Asv*sd*bc*sc*i*
alias:          pci:v000015B3d00006372sv*sd*bc*sc*i*
alias:          pci:v000015B3d00006750sv*sd*bc*sc*i*
alias:          pci:v000015B3d00006368sv*sd*bc*sc*i*
alias:          pci:v000015B3d0000673Csv*sd*bc*sc*i*
alias:          pci:v000015B3d00006732sv*sd*bc*sc*i*
alias:          pci:v000015B3d00006354sv*sd*bc*sc*i*
alias:          pci:v000015B3d0000634Asv*sd*bc*sc*i*
alias:          pci:v000015B3d00006340sv*sd*bc*sc*i*
depends:
retpoline:      Y
name:           mlx4_core
vermagic:       6.8.12-5-pve SMP preempt mod_unload modversions
sig_id:         PKCS#7
signer:         DKMS module signing key
sig_key:        2D:F8:AE:B4:09:A5:69:3A:96:9A:15:62:C1:0C:F8:7F:E0:FD:B4:A1
sig_hashalgo:   sha512
signature:      58:66:CA:6E:97:EB:0A:09:49:BA:E1:AE:A5:6D:33:8E:87:EB:59:B4:
                E0:A7:81:68:EB:82:FC:29:25:41:06:88:E0:D6:33:C0:FD:68:AB:89:
                89:23:71:20:C2:6B:9A:EE:CE:E5:76:8E:A8:D8:B4:72:E5:28:BF:BA:
                1E:81:94:A6:F1:80:95:DF:F0:B1:12:57:2E:90:B6:D0:54:37:D3:EE:
                19:BA:54:D7:13:62:66:0E:5A:56:E2:94:B6:4E:0F:AD:A5:A2:37:6E:
                99:78:D0:A3:FA:CB:09:99:86:BA:5A:59:F4:A9:C2:EC:89:0F:6A:C2:
                DB:A8:3D:B0:F8:1F:BB:A8:D6:D4:37:B1:3A:2B:3A:03:5E:91:C7:9B:
                B5:10:C3:52:68:71:04:14:12:FB:37:C2:EE:03:5A:C9:CB:4A:4A:E2:
                1F:41:9F:0B:9E:8B:98:C9:EE:20:46:EF:C8:9D:79:71:67:F3:13:99:
                43:2F:53:06:48:3C:AB:E5:0C:12:DF:DC:18:6A:6D:12:03:05:0F:D7:
                0D:EC:88:62:1A:CD:65:38:6A:E6:90:15:EE:D5:90:EB:41:12:66:01:
                D7:F9:3A:38:DB:82:8C:AB:BB:4F:AC:1A:CE:D8:1E:B7:A8:46:F2:C9:
                F7:11:0B:47:84:09:1C:CC:A3:3B:2A:15:01:48:61:2B
parm:           debug_level:Enable debug tracing if > 0 (int)
parm:           msi_x:0 - don't use MSI-X, 1 - use MSI-X, >1 - limit number of MSI-X irqs to msi_x (int)
parm:           num_vfs:enable #num_vfs functions if num_vfs > 0
num_vfs=port1,port2,port1+2 (array of byte)
parm:           probe_vf:number of vfs to probe by pf driver (num_vfs > 0)
probe_vf=port1,port2,port1+2 (array of byte)
parm:           log_num_mgm_entry_size:log mgm size, that defines the num of qp per mcg, for example: 10 gives 248.range: 7 <= log_num_mgm_entry_size <= 12. To activate device managed flow steering when available, set to -1 (int)
parm:           enable_64b_cqe_eqe:Enable 64 byte CQEs/EQEs when the FW supports this (default: True) (bool)
parm:           enable_4k_uar:Enable using 4K UAR. Should not be enabled if have VFs which do not support 4K UARs (default: false) (bool)
parm:           log_num_mac:Log2 max number of MACs per ETH port (1-7) (int)
parm:           log_num_vlan:Log2 max number of VLANs per ETH port (0-7) (int)
parm:           use_prio:Enable steering by VLAN priority on ETH ports (deprecated) (bool)
parm:           log_mtts_per_seg:Log2 number of MTT entries per segment (0-7) (default: 0) (int)
parm:           port_type_array:Array of port types: HW_DEFAULT (0) is default 1 for IB, 2 for Ethernet (array of int)
parm:           enable_qos:Enable Enhanced QoS support (default: off) (bool)
parm:           internal_err_reset:Reset device on internal errors if non-zero (default 1) (int)
This will not solve the error 43 problem. Just disable the fast start up option that was mentioned by @okrasit and make sure both the bus driver and network interface drivers are from mellanox not a mix of Microsoft and mellanox.
 

phyerbarte

New Member
Dec 30, 2024
2
0
1
This will not solve the error 43 problem. Just disable the fast start up option that was mentioned by @okrasit and make sure both the bus driver and network interface drivers are from mellanox not a mix of Microsoft and mellanox.
Thanks for the explaination, it's sad, seems there is no way for the fixing, looks like the best chance may doing the 4.9 version driver build by our own with new Linux kernal, but it may impossible without necessary code changing, even we have the source code...ah..too bad. dead end for the connectx-3 with windows guest under SRIOV...
 

Basriram

New Member
Sep 16, 2021
3
7
3
Thanks for the explaination, it's sad, seems there is no way for the fixing, looks like the best chance may doing the 4.9 version driver build by our own with new Linux kernal, but it may impossible without necessary code changing, even we have the source code...ah..too bad. dead end for the connectx-3 with windows guest under SRIOV...
I should have been a little more clear. With the dkms patch for the proxmox 6.8 kernel you will be able to use the virtual network interface under SRIOV in windows. Here are some steps -
1. Get the dkms scripts installed on the host and build mlx4_core.ko module.
2. Enable vf ports by creating /etc/modprobe.d/mlx4_core.conf file with the required number of virtual functions. In my case below I have 2 vfs
cat /etc/modprobe.d/mlx4_core.conf
options mlx4_core port_type_array=2,2 num_vfs=2 probe_vf=0 enable_64b_cqe_eqe=0 log_num_mgm_entry_size=-1 debug_level=64 msi_x=1 enable_4k_uar=1 enable_qos=1 log_num_mac=7 log_mttr_per_seg=4
3. updateinitramfs -u -k `uname -r`, modprobe mlx4_en etc., and you should see the additional virtual adapters with lspci -nn
09:00.0 Ethernet controller [0200]: Mellanox Technologies MT27500 Family [ConnectX-3] [15b3:1003]
09:00.1 Ethernet controller [0200]: Mellanox Technologies MT27500/MT27520 Family [ConnectX-3/ConnectX-3 Pro Virtual Function] [15b3:1004]
09:00.2 Ethernet controller [0200]: Mellanox Technologies MT27500/MT27520 Family [ConnectX-3/ConnectX-3 Pro Virtual Function] [15b3:1004]
4. Create a snippet file in /var/lib/vz/snippets/ folder for any tweaks you need to apply on the specific virtual function. In my example I use a specific mac address and vlan tag
cat /var/lib/vz/snippets/windows.sh
#!/bin/sh
/usr/bin/ip link set enp9s0 vf 0 mac f4:52:14:68:b5:c0
/usr/bin/ip link set enp9s0 vf 0 vlan 30
5. Add this hookscript to your windows guest conf as well as add the pci node
cat /etc/pve/qemu-server/100.conf
hookscript: local:snippets/windows.sh
hostpci1: 0000:09:00.1
6. Install the Mellanox drivers for both the system bus and the network adapter
1736448908311.png
7. Ensure Turn on fast start-up is unchecked in settings
1736448959634.png
With this you should be able to use your virtual adapter in windows. Hope this helps.
 

kapone

Well-Known Member
May 23, 2015
2,063
1,395
113
Hi guys - paging @Basriram - Apologies for bringing up an old thread, but I'm in the middle of playing with PVE (planning to move off Hyper-V since I need native container support) and I'm running into issue with the CX3 cards and SR-IOV.

I've gone through the steps above to patch the mlx driver in PVE, but...I'm using the latest PVE 8.4.1 and..when I do modinfo mlx4_core, I'm on 6.8.12-11-pve vs what was just above:

name: mlx4_core
vermagic: 6.8.12-5-pve SMP preempt mod_unload modversions
Does this patch still work with the latest PVE 8.4.1?

Here's what I'm facing.

Background Info:
- CX3 dual port 40gb cards in the PVE host, all flashed to the latest firmware (2.42.5000), connected to an SX6036 switch.
- PVE 8.4.1 setup and running successfully.
- A test Windows 11 VM with Mellanox drivers installed (5.50.54000).
- PVE hookscript for this VM configured to use vlan 20

The VM starts fine, no yellow bangs in Device Manager, both Mellanox devices show up correctly and use the right driver versions. The ethernet adapter pulls a correct IP address based on vlan 20 from my DHCP server, and...then nothing. Stays at "Unidentified Network" and no traffic seems to pass. I can't even ping the gateway for that VLAN.

If I simply pass a virtio Nic from the same bond/bridge, everything works fine, so I know there's no issues with the VLAN routing/LACP/Switch etc.

What am I doing wrong/missing?

Edit: I'm seeing this in dmesg (The device vfio-pci 0000:81:00.1 is the VF adapter passed to the VM)

Code:
[  828.231590] mlx4_core 0000:81:00.0: default mac on vf 0 port 1 to 248A07DA1753 will take effect only after vf restart
[  828.233364] mlx4_core 0000:81:00.0: updating vf 0 port 1 config will take effect on next VF restart
[  829.172536] tap100i0: entered promiscuous mode
[  829.208563] fwbr100i0: port 1(tap100i0) entered blocking state
[  829.208567] fwbr100i0: port 1(tap100i0) entered disabled state
[  829.208586] tap100i0: entered allmulticast mode
[  829.208730] fwbr100i0: port 1(tap100i0) entered blocking state
[  829.208733] fwbr100i0: port 1(tap100i0) entered forwarding state
[  829.225425] fwln100o0: entered promiscuous mode
[  829.241318] fwbr100i0: port 2(fwln100o0) entered blocking state
[  829.241322] fwbr100i0: port 2(fwln100o0) entered disabled state
[  829.241339] fwln100o0: entered allmulticast mode
[  829.241426] fwbr100i0: port 2(fwln100o0) entered blocking state
[  829.241428] fwbr100i0: port 2(fwln100o0) entered forwarding state
[  831.564475] vfio-pci 0000:81:00.1: enabling device (0000 -> 0002)
[  831.898305] mlx4_core 0000:81:00.0: default mac on vf 0 port 1 to 248A07DA1753 will take effect only after vf restart
[  831.899954] mlx4_core 0000:81:00.0: updating vf 0 port 1 config will take effect on next VF restart
[  840.834036] mlx4_core 0000:81:00.0: Received reset from slave:1
[  841.046605] mlx4_core 0000:81:00.0: denying Global Pause change for slave:1
[ 1531.841656] vmbr0: the hash_elasticity option has been deprecated and is always 16
[ 1543.321475] mlx4_core 0000:81:00.0: denying Global Pause change for slave:1
 
Last edited:

kapone

Well-Known Member
May 23, 2015
2,063
1,395
113
Update: Well, it seems like I was getting a lil too ambitious. I was using OVS bonds/bridges because...well, they're newer! And it didn't work with them. Reverting back to Linux bonds/bridges makes everything work.

It must be something in my OVS config...but glad to report that the patch seems to work with the latest PVE 8.4.1.
 

kapone

Well-Known Member
May 23, 2015
2,063
1,395
113
Update 2: The OVS issue turned out to be a blonde moment on my part and not getting the config right. SR-IOV works just fine with an OVS bond/bridge with the CX3 adapters.

Oh, and RDMA on the VF adapters (Linux guests only) works great as well with PVE.
 

naptastic

Member
Jan 27, 2023
40
5
8
[casts "Revive Thread"]

Edit: I'm seeing this in dmesg (The device vfio-pci 0000:81:00.1 is the VF adapter passed to the VM)

Code:
[  828.231590] mlx4_core 0000:81:00.0: default mac on vf 0 port 1 to 248A07DA1753 will take effect only after vf restart
[  828.233364] mlx4_core 0000:81:00.0: updating vf 0 port 1 config will take effect on next VF restart
[  831.898305] mlx4_core 0000:81:00.0: default mac on vf 0 port 1 to 248A07DA1753 will take effect only after vf restart
[  831.899954] mlx4_core 0000:81:00.0: updating vf 0 port 1 config will take effect on next VF restart
[  840.834036] mlx4_core 0000:81:00.0: Received reset from slave:1
[  841.046605] mlx4_core 0000:81:00.0: denying Global Pause change for slave:1
How do you reset a vf so your changes are effective? I get the "...will take effect on next VF restart" message in dmesg but I have no idea what to do next and every oracle I've consulted doesn't know either.

Thanks,
Nap
 

klui

༺༻
Feb 3, 2019
1,061
626
113
Did you try
Code:
echo $pcie_add >/sys/bus/pci/drivers/mlx4_core/unbind
echo $pcie_add >/sys/bus/pci/drivers/mlx4_core/bind
From gist.github.com/koallen/32709a244d77a2c0f8e17ed79a4092ed#:~:text=reload%20driver
 

dsrhdev

Member
May 28, 2024
43
14
8
hello,
@klui @naptastic
on modern kernels (if im not mistaken - 6+) this way still works, but there is a little bit different (and recommended) approach;
- do not directly bind VFs to driver, otherwise use `driver_override` to use particular driver (mlx4_core, vfio-pci, etc)
Code:
echo mlx5_core > /sys/bus/pci/devices/0000:11:22.3/driver_override
- trigger driver reload for the device
Code:
echo '0000:11:22.3' > /sys/bus/pci/drivers_probe
 

klui

༺༻
Feb 3, 2019
1,061
626
113
I think what you're suggesting is to use driver_override for a PCIe address to explicitly use a different driver instead of going to drivers/$driver/{bind | unbind}.

Several sites still recommend using bind instead of pci/drivers_probe. Are they the same?
 

dsrhdev

Member
May 28, 2024
43
14
8
Several sites still recommend using bind instead of pci/drivers_probe. Are they the same?
i just wanted to say:
1) vfio-pci (ids, new_id) and driver bind/unbind are not idempotent
2) use driver_override instead of vfio-pci ids/new_id and driver_probe instead of bind/unbind

option `driver_override` is recommended to use since 3.16 for pci devices, and in 6.6 this mechanism is finally works for apci, pci, usb buses + reworked into one kernel core subsystem
or you can use `driverctl` utility
 

naptastic

Member
Jan 27, 2023
40
5
8
i just wanted to say:
1) vfio-pci (ids, new_id) and driver bind/unbind are not idempotent
2) use driver_override instead of vfio-pci ids/new_id and driver_probe instead of bind/unbind

option `driver_override` is recommended to use since 3.16 for pci devices, and in 6.6 this mechanism is finally works for apci, pci, usb buses + reworked into one kernel core subsystem
or you can use `driverctl` utility
Here's what I got:

Code:
# driverctl list-devices
(...snip...)
0000:0e:00.0 mlx4_core
0000:0e:00.1 (none)

# driverctl set-override 0000:0e:00.1 mlx4_core
# driverctl: failed to bind device 0000:0e:00.1 to driver mlx4_core
But, in `driverctl list-devices`, the entry for 0000:0e:00.1 has gone from "(none)" to "(none) [*]".

I can use `driverctl unset-override` and the entry changes back. When I do, this appears in dmesg:

Code:
[203151.972661] mlx4_core: Initializing 0000:0e:00.1
[203151.972689] mlx4_core 0000:0e:00.1: enabling device (0000 -> 0002)
[203151.973318] mlx4_core 0000:0e:00.1: Skipping virtual function:1
In my Linux CLI, I have mlx4.probe_vf=0,0,0. Is that the reason? (The parameter is read-only in sysfs, or I would try changing it.)
 

dsrhdev

Member
May 28, 2024
43
14
8
hello,
i've taken a look to code and found two issues regarding to override functionality::
1) at line 176 - it tries to resolve /sys/bus/pci/devices/0000:11:22:33.4/driver, but there is not particular attribute `driver` for mlx4_core driver (the attribute appears after driver bound)
2) at lines 167-171 - it tries to [de]register VF in vfio-pci - it is not necessary and just throw an error on sequential calls, and it is not correct if you have more than one card (CX-3 in your case) installed in your hardware
Code:
152  function set_override()
   153  {
   154      if [ ! -f "$syspath/driver_override" ]; then
   155          error "device does not support driver override: $dev"
   156      fi
   157      if [ -n "$drv" ] && [ "$drv" != "none" ]; then
   158          debug "setting driver override for $dev: $drv"
   159          if [ ! -d "/sys/module/$drv" ]; then
   160              debug "loading driver $drv"
   161              /sbin/modprobe -q "$drv" || error "no such module: $drv"
   162          fi
   163      else
   164          debug "unsetting driver override for $dev"
   165      fi
   166      unbind
   167      if [ "$drv" = "vfio-pci" ]; then
   168        echo -n "$(< "$syspath/vendor") $(< "$syspath/device")" > "/sys/module/vfio_pci/drivers/pci:vfio-pci/new_id" 2>/dev/null
   169      elif [ "$drv" = "none" ] || [ -z "$drv" ]; then
   170        echo -n "$(< "$syspath/vendor") $(< "$syspath/device")" > "/sys/module/vfio_pci/drivers/pci:vfio-pci/remove_id" 2>/dev/null
   171      fi
   172      echo "$drv" > "$syspath/driver_override"
   173
   174      if [ "$drv" != "none" ] && [ $probe -ne 0 ]; then
   175          probe_driver
   176          if [ ! -L "$syspath/driver" ]; then
   177              error "failed to bind device $dev to driver $drv"
   178          fi
   179      fi
   180  }
but it successfully overrides and probes driver (in my case), but with side effect - all not bound VF of mlx4 are going to bind to vfio-pci due to 167-171 code lines
you can override driver manually via `echo "driver_name' > /sys/bus/pci/devices/0000:11:22:33.4
and trigger probe via `echo 0000:11:22:33.4 > /sys/bus/pci/drivers_probe
in your case you see the `skipping virtual function: 1` due to VF not probed by driver itself since you passed `probe_vf=0,0,0`

sorry i incorrectly recommend the `driverctl` utility
upd: i guess mlx4_core driver is not fully compatible with modern kernel's SR-IOV (i.e. management of sriov_numvfs etc)
upd2: if you passed probe_vf=x,0,0 you are able to bind/unbind `probed` VFs to mlx4_core
 
Last edited:
  • Like
Reactions: naptastic