Minisforum ms-01 i9-13900H 96GB with Proxmox

Notice: Page may contain affiliate links for which we may earn a small commission through services like Amazon Affiliates or Skimlinks.

Toormser

New Member
Jul 15, 2024
7
1
3
This problem is hard to traking. Currently we recommand using kingston RAM if you want to run 5200Mhz.
There is no any report of memory issue from kingston RAM.
What did you think @JaxJiang ; is this Memory supported with maximum MTS? 2Rx8 KVR56S46BD8K2-96 (Kingston 96GB Kit)
 
Last edited:
  • Like
Reactions: jsunjones

hellohenri

New Member
Jun 10, 2024
3
0
1
I still get random hangs while using Bios v1.24, micocode 0x4121, and crucial DIMM.
The most recent was this morning after 5 days. Memory speed has been lowered to 4400.
The last thing I did was disable ASPM for the 2.5G network ports.

Usually, when the random hangs happen, the fans run at full speed and the computer becomes unresponsive, and I have to unplug the electricity and restart again.

For those people running Bios v1.24, do you disable ASPM for the 2.5G network ports? Also, do your fans also run at full speed when you get random hangs or reboots?

I would also like to know which Kingston DIMMs are recommended. I have 2 days to return these crucial DIMMs.
 

Toormser

New Member
Jul 15, 2024
7
1
3
I would also like to know which Kingston DIMMs are recommended. I have 2 days to return these crucial DIMMs.
In your case I would return your crucial memory directly and buy the Kingston from my latest posts. You don’t have to lose anything.
 

Lix

Member
Aug 6, 2017
46
13
8
41
I still get random hangs while using Bios v1.24, micocode 0x4121, and crucial DIMM.
The most recent was this morning after 5 days. Memory speed has been lowered to 4400.
The last thing I did was disable ASPM for the 2.5G network ports.

Usually, when the random hangs happen, the fans run at full speed and the computer becomes unresponsive, and I have to unplug the electricity and restart again.

For those people running Bios v1.24, do you disable ASPM for the 2.5G network ports? Also, do your fans also run at full speed when you get random hangs or reboots?

I would also like to know which Kingston DIMMs are recommended. I have 2 days to return these crucial DIMMs.
I have two units, same HW-setup in both, no changes to BIOS settings except secure boot off (13900H CPU / 1x NVMe / 64GB RAM)

Still running BIOS v1.17 + microcode, changed grub settings:

GRUB_CMDLINE_LINUX_DEFAULT="quiet pcie_port_pm=off pcie_aspm.policy=performance"

I also had lockups, with fans speed ramp up, but looking at the logs in the running VM's they would log the crash at the same time that I powered off the host. Did the GRUB mod to see if it was some sort of power saving that made the 2.5 nics go to sleep and not be able to wake up again. So far so good. No issues the last 13 days..
 

alex501

New Member
May 22, 2024
8
0
1
I still get random hangs while using Bios v1.24, micocode 0x4121, and crucial DIMM.
The most recent was this morning after 5 days. Memory speed has been lowered to 4400.
The last thing I did was disable ASPM for the 2.5G network ports.

Usually, when the random hangs happen, the fans run at full speed and the computer becomes unresponsive, and I have to unplug the electricity and restart again.

For those people running Bios v1.24, do you disable ASPM for the 2.5G network ports? Also, do your fans also run at full speed when you get random hangs or reboots?

I would also like to know which Kingston DIMMs are recommended. I have 2 days to return these crucial DIMMs.
Can you try to remove one dimm for test ? I have stable work for month after this on 5200 memory speed, slot closer to cpu is stable, far from cpu have random hangs.
 

JaxJiang

Member
Jan 10, 2023
94
84
18
I still get random hangs while using Bios v1.24, micocode 0x4121, and crucial DIMM.
The most recent was this morning after 5 days. Memory speed has been lowered to 4400.
The last thing I did was disable ASPM for the 2.5G network ports.

Usually, when the random hangs happen, the fans run at full speed and the computer becomes unresponsive, and I have to unplug the electricity and restart again.

For those people running Bios v1.24, do you disable ASPM for the 2.5G network ports? Also, do your fans also run at full speed when you get random hangs or reboots?

I would also like to know which Kingston DIMMs are recommended. I have 2 days to return these crucial DIMMs.
Hi. Thanks for your report.
Our BIOS enable ASPM on I226-V/LM default.
If downgrade memory speed and problem still. I strongly suggest you contact after-sales service to exchange or return the product.
Our tested kingston ram part number is: CBD56S46BD8HA-32. I am sorry currently there is no any kingston 48G RAM data.
Your Crucial 48G RAM part. We are investigating the issue.
 

JaxJiang

Member
Jan 10, 2023
94
84
18
I have two units, same HW-setup in both, no changes to BIOS settings except secure boot off (13900H CPU / 1x NVMe / 64GB RAM)

Still running BIOS v1.17 + microcode, changed grub settings:

GRUB_CMDLINE_LINUX_DEFAULT="quiet pcie_port_pm=off pcie_aspm.policy=performance"

I also had lockups, with fans speed ramp up, but looking at the logs in the running VM's they would log the crash at the same time that I powered off the host. Did the GRUB mod to see if it was some sort of power saving that made the 2.5 nics go to sleep and not be able to wake up again. So far so good. No issues the last 13 days..
Thanks for your report.We will keep an eye on this
 

jsunjones

New Member
May 15, 2023
22
10
3
I did disabled ASPM on all nics. I am using all 4 of them. The 2.5’s are running at 1Gb. I don’t recall my fans running high when lockup occurs but it’s located where i have some other devices making noise. So far with the new bios and 4400 I’m about 13/14 days without issue. Longest ever so far. Also my next resort was to try one dimm like suggested above. I haven’t modified the performance profile/policy but it was also something I’ve considered
 
  • Like
Reactions: B-C

hellohenri

New Member
Jun 10, 2024
3
0
1
Can you try to remove one dimm for test ? I have stable work for month after this on 5200 memory speed, slot closer to cpu is stable, far from cpu have random hangs.
Unfortunately I can't try this anymore with the 48GB DIMM. I've just returned it. I'm using the crucial 2x16GB DIMM back at 5200 now that comes with MS-01. I wanted to test if the issue persist.

Hi. Thanks for your report.
Our BIOS enable ASPM on I226-V/LM default.
If downgrade memory speed and problem still. I strongly suggest you contact after-sales service to exchange or return the product.
Our tested kingston ram part number is: CBD56S46BD8HA-32. I am sorry currently there is no any kingston 48G RAM data.
Your Crucial 48G RAM part. We are investigating the issue.
Thanks. I'll try the following then:
- use the stock Crucial 2x16G at 5200
- use the stock Crucial 2x16G at 4400
- use 1 slot as per suggested above

If none of this works, then I'll contact the after-sales team.

This message comes every time in the log. Does anyone know what this is?

[Thu Jul 18 15:39:36 2024] EDAC MC0: Giving out device to module igen6_edac controller Intel_client_SoC MC#0: DEV 0000:00:00.0 (INTERRUPT)
[Thu Jul 18 15:39:36 2024] EDAC MC1: Giving out device to module igen6_edac controller Intel_client_SoC MC#1: DEV 0000:00:00.0 (INTERRUPT)
[Thu Jul 18 15:39:36 2024] EDAC igen6 MC1: HANDLING IBECC MEMORY ERROR
[Thu Jul 18 15:39:36 2024] EDAC igen6 MC1: ADDR 0x1ffffffffff
[Thu Jul 18 15:39:36 2024] EDAC igen6 MC0: HANDLING IBECC MEMORY ERROR
[Thu Jul 18 15:39:36 2024] EDAC igen6 MC0: ADDR 0x1ffffffffff
[Thu Jul 18 15:39:36 2024] EDAC igen6: v2.5.1

Thanks
 
Last edited:

B-C

New Member
Jun 2, 2024
7
0
1
Have had a couple of failure events now...


Node 2 gets fenced - VMs move to another node and all works as expected so not major.

All three have good resources and not overloaded -

do see a CPU sustained load overnight into AM hours - possibly its just CPU hitting a no go point....
Screenshot 2024-07-20 170409.jpg


96g ram - issue is I don't have a Switched UPS... and couldn't get the RMM working on the MS01s before I deployed them
so 1000+ miles away physically - have to reboot the UPS and take the network down to get the failed host back online.

Seems to be the same node every month now.

Pretty sure I have the microcode installed on all three nodes

humm... used the promox helper scripts when I setup but possibly during upgrades it got modified again.


Affected nodes before update:
# journalctl -k | grep -E "microcode" | head -n 1

Jul 20 15:13:21 pve02 kernel: microcode: Current revision: 0x00004121

Ran their script:

✓ GenuineIntel was detected
✓ Intel iucode-tool is already installed
- Downloading the Intel Processor Microcode Package intel-microcode_3.20240514.1~deb12u1_amd64.
✓ Downloaded the Intel Processor Microcode Package intel-microcode_3.20240514.1~deb12u1_amd64.deb
✓ Installed intel-microcode_3.20240514.1~deb12u1_amd64.deb
✓ Cleaned

In order to apply the changes, a system reboot will be necessary.
# reboot

# journalctl -k | grep -E "microcode" | head -n 1

Jul 20 16:13:33 pve03 kernel: microcode: Current revision: 0x00004121

same result so expect thats not the issue I'm hitting.

lots of these - I'm betting my thunderbolt mesh is jacked up...
fabricd[892]: [S3GXJ-RJEC1] No supported protocols TLV in P2P IIH

Ideas?
 

B-C

New Member
Jun 2, 2024
7
0
1
I did find an oddity with my #2 node...
After yet another crash event 20240721 -
Moved moved VMs off that and moved some others onto it to test with.

Bios is 1.17 vs 1.22 - thought I checked that previously -

Seems I need to update the bios to the current version....
is there anyway to update that remotely via proxmox / debian 12?
vpro setup enabled but never linked so the IPs setup for that never took effect so no access directly to that either currently.

See @reneil1337 has the 1.24 bios and slowed the RAM speed - but my other nodes on 1.22 are stable so really just getting this node up to 1.22 would probably resolve my issue.


Node 1
# dmidecode | less
Getting SMBIOS data from sysfs.
SMBIOS 3.5.0 present.
....
BIOS Information
Vendor: American Megatrends International, LLC.
Version: AHWSA.1.22
Release Date: 03/12/2024

Node 2
# dmidecode | less
Getting SMBIOS data from sysfs.
SMBIOS 3.5.0 present.
....
BIOS Information
Vendor: American Megatrends International, LLC.

Version: AHWSA.1.17
Release Date: 12/14/2023


Node 3
# dmidecode | less
Getting SMBIOS data from sysfs.
SMBIOS 3.5.0 present.
....
BIOS Information
Vendor: American Megatrends International, LLC.
Version: AHWSA.1.22
Release Date: 03/12/2024
 

JaxJiang

Member
Jan 10, 2023
94
84
18
Have had a couple of failure events now...


Node 2 gets fenced - VMs move to another node and all works as expected so not major.

All three have good resources and not overloaded -

do see a CPU sustained load overnight into AM hours - possibly its just CPU hitting a no go point....
View attachment 38004


96g ram - issue is I don't have a Switched UPS... and couldn't get the RMM working on the MS01s before I deployed them
so 1000+ miles away physically - have to reboot the UPS and take the network down to get the failed host back online.

Seems to be the same node every month now.

Pretty sure I have the microcode installed on all three nodes

humm... used the promox helper scripts when I setup but possibly during upgrades it got modified again.


Affected nodes before update:
# journalctl -k | grep -E "microcode" | head -n 1

Jul 20 15:13:21 pve02 kernel: microcode: Current revision: 0x00004121

Ran their script:

✓ GenuineIntel was detected
✓ Intel iucode-tool is already installed
- Downloading the Intel Processor Microcode Package intel-microcode_3.20240514.1~deb12u1_amd64.
✓ Downloaded the Intel Processor Microcode Package intel-microcode_3.20240514.1~deb12u1_amd64.deb
✓ Installed intel-microcode_3.20240514.1~deb12u1_amd64.deb
✓ Cleaned

In order to apply the changes, a system reboot will be necessary.
# reboot

# journalctl -k | grep -E "microcode" | head -n 1

Jul 20 16:13:33 pve03 kernel: microcode: Current revision: 0x00004121

same result so expect thats not the issue I'm hitting.

lots of these - I'm betting my thunderbolt mesh is jacked up...
fabricd[892]: [S3GXJ-RJEC1] No supported protocols TLV in P2P IIH

Ideas?
Could you help try 1.24 TEST BIOS for RAM SPEED downgrade?
 

B-C

New Member
Jun 2, 2024
7
0
1
Sure - but again - Remote without access locally via a KVM / VPro

So stuck on 1.17 currently can't even stabilize under 1.22 like the others.

Unless there is a way to update the bios via the host directly in cli?
 

anewsome

Active Member
Mar 15, 2024
130
135
43
Here's my update with an uptime report on the 5 node MS01 Proxmox Cluster. I've been running BIOS 1.22 since it was released and the cluster definitely has experienced less crashing. Before 1.22, I didn't have a single node with more than a few days of uptime. Here's the current uptime on the 5 nodes:


Code:
ms01:  09:44:05 up 30 days, 14:59,  1 user,  load average: 0.76, 1.08, 1.09
ms02:  09:44:05 up 51 days, 18:42,  0 user,  load average: 0.22, 0.35, 0.66
ms03:  09:44:06 up 16 days, 20:25,  0 user,  load average: 0.64, 0.58, 0.68
ms04:  09:44:06 up 52 days,  1:02,  0 user,  load average: 1.18, 0.89, 0.84
ms05:  09:44:06 up 45 days, 17:42,  0 user,  load average: 0.74, 0.68, 0.79

Node 3 has seemed to crash a bit more than others, but every other node as 30+ days of uptime. Node 1 did have any issue where I couldn't get the thunderbolt ports activated, so it has less uptime than the other "stable" nodes.

@JaxJiang , I would like to try the 1.24 BIOS. It sounds like memory speed can be decreased to provide more stability. I'll read back into the thread to see if I can find the download link.
 

wishbone65

New Member
Jul 24, 2024
1
1
3
After reading this entire thread, it sounds a lot like the Intel 13th/14th gen stability and oxidization CPU issues confirmed by intel. Even the memory down clock resolution and random crashes. I stumbled on it from watching gamers nexus.

 
  • Like
Reactions: SloothNZ

P00r

New Member
May 29, 2024
26
3
3
This problem is hard to traking. Currently we recommand using kingston RAM if you want to run 5200Mhz.
There is no any report of memory issue from kingston RAM.
Not fully related but could confirm for the Kingston.

I have Kingston Fury in my i9-13900H 2x32GB DDR5-5600
I also have i9-12900H Running SABRENT Rocket DDR5 64GB SO-DIMM 4800MHz

The sabrent was not working in the i9-13900H (not booting) but work fine in the I9-12900H, both unit are rock stable, proxmox is running on the 12900H

Only issue I have is that in the I9-12900H with 3 x NVME SSD, I have to disable the 10G port to enable the 2.5G network, it turn itself off when booting.

The I9-13900H has nothing disabled with same 3x NVME SSD, so waiting patiently for new bios release to see if I can re-enable the 10G port. Both unit have a NVIDIA GX1650 but not the same brand.

Would like to test this new bios, but I can't report on 48G :)
 
Last edited: