school me on NVME namespace best practice

Notice: Page may contain affiliate links for which we may earn a small commission through services like Amazon Affiliates or Skimlinks.

dragonme

Active Member
Apr 12, 2016
372
34
28
I picked up 2 Samsung / Dell pm1725A .. latest firmware 1.6TB pic-e NVME cards... eBay of course.. 70 bucks a pop...

well they came today and each had almost identical ~1.6 PB of data read/written and 48000+hours uptime.. not exactly spring chickens but heath still looked good from what I can make of NMVE tools version of smart output.

The use case will be for use in a ESXI server with 1 VM (nappit or truenas) having a 3008 HBA passed to it to manage the large 5x8tb data store pool
my previous servers I had this VM feeding internal AIO NFS volumes back to ZFS for VM storage but in this case I think I will keep the VMs native on a namespace

so ..

first.. ZFS / ESXI what is the best format for the namespaces .. its currently at 512.. I am thinking 4k would make more sense as pools get set up ashift=12 ?

Second... I want to partition a couple namespaces for ZIL and ARC.. the machine will have 128gb of ram.. and honestly I think regular l2arc in ram might be good enough. my current server is running just fine with no ZIL or ARC on the data pool .. but the VMs are on a pool of 2 stripped S3500SSDs

1 machine will have a 10gbe Nic.. but honestly the network is gigabit..

I dont thinK I will be running any active VM workloads on the pool requiring sync write like NFS.. but there will likely be times where NFS might serve out data to other machines.. and of course there will be SMB shares.. like to a VM for plex back end media.. again ... right now that pool on a 5 wide stripe of spinning iron seems to keep up..

I don't think ESXI can namespace the NVME natively.. so what.. should I pass it though to Ubuntu to manage setting up namespaces, then un-attach

each namespace should be treated like a LUN in ESXI.. so I was figuring I would pass though the namespace to Nappit or truenas.. to attach to the pool... rather than pass the Pcie device in total .. as I want to keep a namespace or 2 for native VMFS storage..

anywho.. school me on this .. its all new to to me... I go back far enough were storage was on punchcards so be gentle


haha


thanks in advance ..
 

mr44er

Active Member
Feb 22, 2020
168
51
28
First thing to do...check out if newer firmware exists and update it.
On linux, you want nvme-cli for the tasks.

Code:
nvme list

nvme fw-download /dev/nvme0 -f firmware.bin
nvme fw-commit -s 2 -a 1

*reboot*
I am thinking 4k would make more sense as pools get set up ashift=12
Yes. Everything you need to know and better than I could explain: Setting 4k sector size on NVMe SSDs: does performance actually change?

I don't think ESXI can namespace the NVME natively.. so what.. should I pass it though to Ubuntu to manage setting up namespaces, then un-attach
If ESXi can't do it, then yes. -> nvme-cli

regular l2arc in ram
You mean l1arc, RAM is level1 ;)
 

dragonme

Active Member
Apr 12, 2016
372
34
28
First thing to do...check out if newer firmware exists and update it.
On linux, you want nvme-cli for the tasks.

Code:
nvme list

nvme fw-download /dev/nvme0 -f firmware.bin
nvme fw-commit -s 2 -a 1

*reboot*

Yes. Everything you need to know and better than I could explain: Setting 4k sector size on NVMe SSDs: does performance actually change?


If ESXi can't do it, then yes. -> nvme-cli


You mean l1arc, RAM is level1 ;)

ok some things I have learned since starting the thread.. which I thought was dead so thanks for weighing in

this is a Samsung/Dell 1725A so at least it can do namespaces.. but no fabrics.. not sure if it even has multiple controllers

2 - ESXI will not recognize 4k sectors on NVME period.. has to be formatted at 512 .. several Broadcom/Vmware tech articles on it.. so there is that.

namespacing I believe cant be done either by GUI or esxcli commands... format, attach, etc can.. but the NVME has to be namespaced/partitioned elsewhere so I used a live ubuntu and used nvme-cli

one namespace and format to 512/0 esxi sees the controller AND the namespaces.. but only the controller is in the PCIe passthrough list.. not the namespaces... so there is no way of passing just the namespace to a VM .. at least not native.

What I found was that .. for example.. if you want a 10gig namespace as a passthrough device to a VM for cache..

you can give the VM a NVME controller
and using esxcli vmkfstools, build a fake vmfs disk on that namespace, and pass it though like an RDM disk.. seems to work ok.. and fairly decent iops.. not sure how 'native' the vm has control over that NVME directly ...

hope some other folks weigh it..

also I dont think the card is bootable.. saw some articles on that as well..

so for now, I have to boot ESXI with USB.. I have passed both sata controllers to the ZFS vm .. this way I can run the VMs native off NVME VMFS storage, or off the ZFS pool with NVME cache, NFS back to the ESXI all-in-one style..
 

mr44er

Active Member
Feb 22, 2020
168
51
28
ESXI will not recognize 4k sectors on NVME period.. has to be formatted at 512 .. several Broadcom/Vmware tech articles on it.. so there is that.
Meh...this alone would be enough reason for me to kick ESXi, really. :)
Overall it does not sound really good.

hope some other folks weigh it..
Yes, I turned my back on ESXi a long time ago because even back then there were too many things you just couldn't do with it and they weren't exotic at all. If no one here knows more or doesn't write anything about it... I've become a bit of a Proxmox fanboy. *cough cough* :cool:
 
  • Like
Reactions: name stolen

dragonme

Active Member
Apr 12, 2016
372
34
28
Meh...this alone would be enough reason for me to kick ESXi, really. :)
Overall it does not sound really good.


Yes, I turned my back on ESXi a long time ago because even back then there were too many things you just couldn't do with it and they weren't exotic at all. If no one here knows more or doesn't write anything about it... I've become a bit of a Proxmox fanboy. *cough cough* :cool:
yeah.. Broadcom is running VMware into the ground.. and I am defiantly playing with Proxmox.. but spent the last 10 years teaching myself ESXI enough to be dangerous.. hahah .. but moving it all to proxmox.. not quite there yet. dont have it all figured out.. but I think its an inevitability at this point ...

ESXI 8 seemed like it dropped drivers for everything and I think Broadcoms new model is .. bill them annually on subscription for software.. and hit them with new hardware requirements every build.. oh.. yeah.. we sell all those networking and SAS/Raid cards.. how convenient...
 
  • Like
Reactions: name stolen

mr44er

Active Member
Feb 22, 2020
168
51
28
but I think its an inevitability at this point
Yup, definitely can recommend that. I read their forum, and it's striking to read posts from new users beginning the sentence with "I come from ESXi..."

bill them annually on subscription for software.. and hit them with new hardware requirements every build.
I read that too, greed knows no bounds...if only the software were feature-rich.
 

dragonme

Active Member
Apr 12, 2016
372
34
28
like almost all software.. ESXI was build on the hard work, suggestions and many time contributions of the open source community and the users that were home labbing with free licenses...

where there is enthusiasm.. there is a crowd.. willing to work for free and contribute to something someone else sells... that is now GONE completely ...

Proxmox will see 90% of that base flock to them this year... and proxmox will begin advancing as a product at the same rate of growth VMware saw 2 decades ago... as they being to loose market share as these converts take the corporations they work for over to Proxmox as well..

what they use at home and are enthusiastic about.. is what they will buy at work...
 

mr44er

Active Member
Feb 22, 2020
168
51
28
If Proxmox will be shitty tomorrow, there is always my solution to switch back to FreeBSD's bhyve (for now not that comfortable to use and no HA), but it works. :)
 

dragonme

Active Member
Apr 12, 2016
372
34
28
If Proxmox will be shitty tomorrow, there is always my solution to switch back to FreeBSD's bhyve (for now not that comfortable to use and no HA), but it works. :)
XCP-NG is built on ZEN if you are into BSD more than linux.. some drawbacks with it as well if I recall.. either passthrough or vGPU or some other niche.. but XCP-NG has a bigger corporate install base than Proxmox for sure.. and I think is #2 to vmware in corporate install base.. IIRC
 

jsingh04

New Member
Mar 4, 2024
11
0
1
I am on the 1.2.1 firmware on my dell pm1725a 1.6 tb hhhl , does it actually support namespaces. I get this error

nvme id-ctrl /dev/nvme0n1 | grep nn
nn : 1

nvme delete-ns /dev/nvme0n1 -n 1
NVMe status: Invalid Command Opcode: A reserved coded value or an unsupported value in the command opcode field(0x1)
 

BackupProphet

Well-Known Member
Jul 2, 2014
1,426
1,069
113
Stavanger, Norway
intellistream.ai
These drives has a firmware bug(also present in the latest firmware) if you run blocks at 512B and not 4KB where you can corrupt the namespace. If the drive lose power or experiences a hang during one of these complex 512B Read-Modify-Write cycles, the translation tables are more likely to become inconsistent. A 4KB native write is "atomic" to the controller's mapping logic, whereas 512B writes are "sub-page" writes that require more complex tracking in the metadata. Also 4KB-byte sectors align perfectly with the controller's internal ECC.
I had one drive crash last week and did a lot of research, no mirror and had to restore from backup.
 
  • Like
Reactions: jsingh04