40Gb is dead... long live 40Gb in my proxmox cluster. :)

Notice: Page may contain affiliate links for which we may earn a small commission through services like Amazon Affiliates or Skimlinks.

willo

Member
Apr 26, 2024
41
17
8
I've been having some fun.
I have a pair of Dell c6400 cluster chassis that I was allowed to save from e-waste this week. Best perk ever.
Each has 4 sleds with dual Intel Silver Xeons and a nice amount of ram. They use U.2 NVME drives so that's been fun to sort out.
Second annoyance was a lack of mezzanine ethernet cards. I had to share the gig nics with idrac. The issue with having 8 nodes is buying 8 of every friggin thing. Thanks to ebay I was able to source brand new QSFP 40 GB copper DACs for $7.50 each. Ethernet cards were a hair over $10 each.

Anyway, I built them all with Debian, then ansible'd them into proxmox nodes with CEPH and a Mellanox dual 40Gb nic each and a 32 port Cisco Nexus.
The end result is a cluster of 8 nodes, each with one drive and with an 80Gb/s CEPH cluster fabric.

For backups, I'm leveraging an old 3U Supermicro with a bunch of lightly used SAS drives. That'll be my PBS server, which is getting it's own 40Gb NIC. The backplane on that thing is only good for 6Gb/s so it'll never have the chance of overrunning the 80Gb cluster network. The card was $15 shipped and the cables were another $20.

Ultimately my point is - my people. 40Gb hardware is DIRT cheap. Take advantage of this fantastic time!
 

ano

Well-Known Member
Nov 7, 2022
802
353
63
I still got some 40gbps clusters, its not only cheap, but low power usage vs cx6's
 
  • Like
Reactions: Stephan

willo

Member
Apr 26, 2024
41
17
8
CO2 bomb, power must too cheap in your area
Yeah that extra 100w to run that switch is a killer. Also I keep my gear in a DC where there's a power commit.


switch# show environment power
Power Supply:
Voltage: 12 Volts
Power Actual Actual Total
Supply Model Output Input Capacity Status
(Watts ) (Watts ) (Watts )
------- ------------------- ---------- ---------- ---------- --------------
1 N9K-PAC-650W-B 47 W 71 W 650 W Ok
2 N9K-PAC-650W-B 49 W 79 W 650 W Ok
 

kapone

Well-Known Member
May 23, 2015
2,067
1,402
113
If you had gone with a Mellanox SX6036 and QSFP+ FDR14 compliant interconnects…

1. power consumption goes down - the SX6036 sips power
2. your title could have been…”long live 56gbps….”

:)
 
  • Like
Reactions: itronin

Greg_E

Active Member
Oct 10, 2024
547
173
43
:( I'm still only 10g in my lab, with a bunch of 1g and a single 2.5g (moca 2.5 adapter).

That said, I have two 40g ports, but I broke both of them out to quad 10g as all the host are only 10g cards and not strong enough to use 40g (PCIe 3.0 x8 single slot).
 

clcorbin

Member
Feb 15, 2014
87
15
8
I have a C6400. It just pulled it a month or so ago in favor of some R640's (cluster) and a R740 (backup). ONE node on that C6400 was twice as loud as all four R servers. At least twice as loud... But a LOT of hardware in a small box. If you need a third, I have one (no CPU and no RAM. Guess were THAT is?). It does have 4x10Gb messanene cards and 2x40 ConnectX-3 cards in each node.
 

willo

Member
Apr 26, 2024
41
17
8
@willo What model nexus switch are you using?
Eh It's a nexus C9332PQ. It's good if you don't need changes. I found that I have to literally wipe it and reconfigure to get port channels to work properly if I need to re-config it. I plan on swapping it for a QFX sometime.
 

macrules34

Active Member
Mar 18, 2016
529
44
28
42
I have bought 9 HP T740’s what I’m going to install proxmox on and was origionally going to use 10gb Ethernet but found out that the 40gb cards are cheaper, so I will be using the 40gb network for ceph and vm migration traffic.
 

willo

Member
Apr 26, 2024
41
17
8
Just remember - 40gb is just 4x10gb. But honestly unless you have MASSIVE data transfers, that's fine. Scaling windows matter to TCP. In my experience 40GB is great for CEPH.
 

macrules34

Active Member
Mar 18, 2016
529
44
28
42
@willo I was looking at using qsfp+ to fiber adapters one each end (host and switch), just not sure if I need mellanox qsfp+ on the card and Cisco qsfp+ on the switch or can I just use one brand on all ends?
 

BackupProphet

Well-Known Member
Jul 2, 2014
1,426
1,070
113
Stavanger, Norway
intellistream.ai
Deep buffer switch, comes with multiple GB of buffers. They are made specially for storage solutions/video streaming as they handle microbursts much more effective. But if you use RDMA/RoCE, this is not a problem as long you enable DCQCN (ECN + PFC) end-to-end across your NICs and switches. If you use Infiniband, then the protocols credit system will save you.
 

macrules34

Active Member
Mar 18, 2016
529
44
28
42
Wow those switches that I could find on eBay were like $800-$900. What’s the ramifications if get a Cisco C9332PQ?
 

BackupProphet

Well-Known Member
Jul 2, 2014
1,426
1,070
113
Stavanger, Norway
intellistream.ai
For a homelab, you will probably be fine. But if CRUSH has to run, and rebalance the data, you will get issues. Nothing is slower than packets getting dropped and TCP goes into full crisis-management mode. There are multiple failure modes here, but the worst is handshake drops, SYN-ACK loss. The default timeout for a SYN/SYN-ACK is seconds...

ECN can help you though. And Arista 7050QX support ECN, I think the Cisco C9332PQ supports it too. Read its documentation.
 

Greg_E

Active Member
Oct 10, 2024
547
173
43
Doesn't flow control help with the lack of ram?

I'd have to look deeper, but even my Mikrotik switches should have enough ram to buffer (store and forward) until flow control tells the sender to wait. I know my Extreme switches have enough, they also offer cut through switching, but I was told there are some downsides to this. Cut through has been used to help audio and video over DANTE and NDI work in congested networks, even places where QOS policies fall short.