Hello guys, looking for some help here. First please kindly forgive my stupidness.
I am using dual 8368es(QWAT). I have been getting MCE for a week now. The error log in Linux is like:
Code:
Aug 10 05:26:41 astraTZ kernel: mce: [Hardware Error]: CPU 38: Machine Check: 0 Bank 26: 80000040020000b1
Aug 10 05:26:41 astraTZ kernel: mce: [Hardware Error]: TSC 29abee5a339 PPIN 5d4fa8cb73c77cc0
I asked Gemini what this means, Gemini told me that Bank26 is related to memory issue. I am not sure if Gemini is correct, but I do have some memory issues on this platform. Please forgive my ignorance here. The issue I came across is quite complex:
1. When I was first installing memory to the slots, I got distracted and accidentally broke one of the slot. Then I left that slot empty and populated every other slot.
2. After that, I was able to get 14 memory slots working. It was supposed to be 15 since only one slot is missing memory module. I used a tool called cockpit to inspected memory info and found out that 15 out of 16 slots is presented with memory, but the memory in one of the slots has an unknown size, which explains why my server only detects 14.
3. I was still unable to locate the real issue until then. But as I am getting MCE over the last week, two more of my memory modules can’t be detected. I am starting to suspect that the issue is with the CPU (since the OS is reporting MCE about the CPU at the same time)
Below is an image of my memory info captured by cockpit.
Also there is a separate incident that might be related to this. Due to my stupidness, I bent ONE of the socket pins of the socket which the problem CPU is installed (this happened before I install the whole system). I paid a guy to repair the problem socket for me, he used tools to bend the pin back, the angle of the pin is not exactly identical to its original state. But I think it seems fine since it’s “almost” perfect and definitely not touching any other pins in the socket. Could this one imperfect pin cause my problem?
Any help is appreciated
