Available on Google Cloud: Intel Optane DC Persistent Memory(cloud.google.com)
cloud.google.com
Available on Google Cloud: Intel Optane DC Persistent Memory
https://cloud.google.com/blog/topics/partners/available-first-on-google-cloud-intel-optane-dc-persistent-memory
70 comments
I'm curious how endurance is handled in this case. Is there a wear level guarantee when provisioning it? I guess this is applicable to any consumable though this seems to be a pretty expensive one.
That’s part of the design of Optane (née 3DXpoint): it’s supposed to be lower latency than traditional SSD and higher durability. I don’t know enough about the how nor the details to say more, sorry.
This article from Anandtech [1] covers the SSD variant though which states:
> The endurance rating for both capacities is 200 GB/day for the five-year warranty period. Given the small capacity of the drives, this works out to 1.7 or 3.4 drive writes per day, which is considerably higher than normal for consumer SSDs.
[1] https://www.anandtech.com/show/12512/the-intel-optane-ssd-80...
This article from Anandtech [1] covers the SSD variant though which states:
> The endurance rating for both capacities is 200 GB/day for the five-year warranty period. Given the small capacity of the drives, this works out to 1.7 or 3.4 drive writes per day, which is considerably higher than normal for consumer SSDs.
[1] https://www.anandtech.com/show/12512/the-intel-optane-ssd-80...
> Optane (née 3DXpoint)
Optane is the brand name for Intel products that use 3D XPoint memory. Micron products that use 3D XPoint memory will be sold under Micron's QuantX trademark, if they ever ship. There's also a possibility that Micron could eventually start selling 3D XPoint memory to other SSD manufacturers, and the resulting products wouldn't bear either Optane or QuantX branding.
The more relevant comparison for endurance may be the enterprise Optane SSD DC P4800X, which was initially released with a 30 DWPD rating but has since been increased to 60 DWPD. Intel hasn't described what kind of error correction that drive is doing under the hood, but it's probably more robust than what the entry-level consumer Optane drives do, and possibly more robust than what the DIMMs can do within their latency budget.
Optane is the brand name for Intel products that use 3D XPoint memory. Micron products that use 3D XPoint memory will be sold under Micron's QuantX trademark, if they ever ship. There's also a possibility that Micron could eventually start selling 3D XPoint memory to other SSD manufacturers, and the resulting products wouldn't bear either Optane or QuantX branding.
The more relevant comparison for endurance may be the enterprise Optane SSD DC P4800X, which was initially released with a 30 DWPD rating but has since been increased to 60 DWPD. Intel hasn't described what kind of error correction that drive is doing under the hood, but it's probably more robust than what the entry-level consumer Optane drives do, and possibly more robust than what the DIMMs can do within their latency budget.
> The persistent part is that it's persistent across reboots.
So does that mean it can be considered persistent like a regular SSD/HDD?
So does that mean it can be considered persistent like a regular SSD/HDD?
Yes. It's durable storage with a RAM interface. Byte-addressable through conventional CPU load/store instructions. It follows the usual cache coherence protocols for memory, but once it gets flushed, it's stable—just like conventional storage devices.
How much of a problem is the gap for SSD's? Don't those have buffers with enough battery to write out if they lose power?
Sorry, I should have been more clear: the performance gap between DRAM and SSD. This sits in between those.
Yes, don't SSD's have DRAM buffers/cache that make the performance gap less important?
Still a big gap. Random read time from NVMe Flash is on the order of 100µs, whereas service time for reads from this medium are lower. You can get an idea of how much lower based on the QD1 numbers of the SSDs based on it at https://www.anandtech.com/show/11953/the-intel-optane-ssd-90....
On top of that, the DRAM-like form is supposed to much further lower latency by 1) getting rid of NVMe and software overhead and 2) letting you read/write as little as a (64b) cache line at a time. We still don't know what the exact remaining latency is, and I bet the few with access through this program will still be NDAd from talking about it. We also don't know about the price/GB.
Those numbers, esp. price, will have a lot to do with how broad an impact this has at first: you can imagine a scenario where most folks currently building large DBs would find it worthwhile to replace 2/3 of their RAM with 5x as much of this stuff for the same price, though that seems kind of unlikely to be imminent. Or we may find out it's only going to be worth it in some niche applications and the rest of us have to wait some generations before it reaches us (if it ever does).
NVMe Flash is, obviously, still sufficient for a lot of workloads and today is nothing like, say, 2004 when I was writing DB-backed apps and HDD seeks were my bane. Notably, if your OLTP DB can have a lot of transactions in flight at once, or your reporting workload is doing sequential reads, you can take advantage of the very high throughput of modern Flash SSDs when working on lots of I/Os in parallel.
On top of that, the DRAM-like form is supposed to much further lower latency by 1) getting rid of NVMe and software overhead and 2) letting you read/write as little as a (64b) cache line at a time. We still don't know what the exact remaining latency is, and I bet the few with access through this program will still be NDAd from talking about it. We also don't know about the price/GB.
Those numbers, esp. price, will have a lot to do with how broad an impact this has at first: you can imagine a scenario where most folks currently building large DBs would find it worthwhile to replace 2/3 of their RAM with 5x as much of this stuff for the same price, though that seems kind of unlikely to be imminent. Or we may find out it's only going to be worth it in some niche applications and the rest of us have to wait some generations before it reaches us (if it ever does).
NVMe Flash is, obviously, still sufficient for a lot of workloads and today is nothing like, say, 2004 when I was writing DB-backed apps and HDD seeks were my bane. Notably, if your OLTP DB can have a lot of transactions in flight at once, or your reporting workload is doing sequential reads, you can take advantage of the very high throughput of modern Flash SSDs when working on lots of I/Os in parallel.
I still do not completely understand what SAP HANA really is:
- an in-memory database technology?
- the name for SAP’s cloud platform?
- an on-premise DB to run SAP ERP (to replace Oracle)?
- a full-stack proprietary web development platform?
- a marketing term to solve all problems with SAP products?
Does anyone have an hands on experience with HANA, beyond the usual marketing BS? Is is that revolutionary? If it runs on such specialized hardware, is the speed increase that impressive?
- an in-memory database technology?
- the name for SAP’s cloud platform?
- an on-premise DB to run SAP ERP (to replace Oracle)?
- a full-stack proprietary web development platform?
- a marketing term to solve all problems with SAP products?
Does anyone have an hands on experience with HANA, beyond the usual marketing BS? Is is that revolutionary? If it runs on such specialized hardware, is the speed increase that impressive?
HANA was in the first instance a in memory database - at least partly invented by SAP. Because SAP marketing pushed "in memory" so much, all of the SAP employees believed that HANA was the solution to all problems with SAP software. Therefore SAP named all of its products SAP HANA .... Like SAP HANA Cloud Platform, SAP HANA Cloud Integration, etc. Those products were not even related to the in memory database. It's all marketing bullshit and SAP has even taken a few steps back as too many customers were disappointed.
Ultimately I think the HANA database is an overhyped in-memory database. The reason they make money with HANA is that they are successfully replacing Oracle Databases for their SAP on-premise customers.
Ultimately I think the HANA database is an overhyped in-memory database. The reason they make money with HANA is that they are successfully replacing Oracle Databases for their SAP on-premise customers.
You forgot about the HanaHaus in Palo Alto, which unfortunately makes coffee at the same speed of all other Blue Bottle locations.
This comment is at once so endearing and so alarmingly indicative of the bubble we live.
I have no practical experience with SAP HANA. But I’m a sysadmin and I’ve been to conferences where SAP has had stands and I have a tendency to ask too many questions.
SAP HANA is an in memory database. That’s all it is. It’s similar to Qlikview if you’ve used that. It’s licensed per socket, so scaling up is better than scaling out.
SAP will also host HANA for you, for a fee.
See. The thing is: SAP is a company for HR departments. The kinds of departments that aren’t technical, do not have technical staff to develop thing for them. (It’s an example) but your HR department talks to SAP, they deliver some service, sometimes something that the barely tech-literate members of HR is able to build on and then suddenly your sysadmins have to support those things in perpetuity.
Well. They market themselves as this anyway. And I have experience of our HR department buying all kinds of stuff from SAP and their child companies. (Concur, for example).
So I understand your incredulous-ness. This company is not made for us. And most of their tech isn’t either.
SAP HANA is an in memory database. That’s all it is. It’s similar to Qlikview if you’ve used that. It’s licensed per socket, so scaling up is better than scaling out.
SAP will also host HANA for you, for a fee.
See. The thing is: SAP is a company for HR departments. The kinds of departments that aren’t technical, do not have technical staff to develop thing for them. (It’s an example) but your HR department talks to SAP, they deliver some service, sometimes something that the barely tech-literate members of HR is able to build on and then suddenly your sysadmins have to support those things in perpetuity.
Well. They market themselves as this anyway. And I have experience of our HR department buying all kinds of stuff from SAP and their child companies. (Concur, for example).
So I understand your incredulous-ness. This company is not made for us. And most of their tech isn’t either.
> SAP is a company for HR departments.
And finance. And operations. And compliance, and ...
> The kinds of departments that aren’t technical
Which, in most businesses, are almost all of them. That's why they are so entrenched and so profitable, despite selling an utterly unfashionable and over-complicated stack.
And finance. And operations. And compliance, and ...
> The kinds of departments that aren’t technical
Which, in most businesses, are almost all of them. That's why they are so entrenched and so profitable, despite selling an utterly unfashionable and over-complicated stack.
We use SAP for finance in many countries world wide, many countries require you to use SAP now so they can give you the tax & legal information for each country in a SAP format to make sure you are up to their countries codes.
I suspect they also want information back from you, and most likely in a SAP compatible version.
This was all told to me at our IT quarterly meeting, with our new SAP deployments, as IT does not actually input the data, just run the servers.
Funny how we run the stuff, but sometimes dont even know how its used.
I suspect they also want information back from you, and most likely in a SAP compatible version.
This was all told to me at our IT quarterly meeting, with our new SAP deployments, as IT does not actually input the data, just run the servers.
Funny how we run the stuff, but sometimes dont even know how its used.
The other critical attribute is that it isn’t Oracle. It’s really not in SAP’s interest to have the Oracle guy in there for annual renewals.
Don't know specifically about HANA. But, any database with high throughput would benefit from high bandwidth PCIe non-volatile storage. Databases with significant random accesses in their workload would see a large improvement in performance due to Optane. Nothing else apart from main memory matches Optane with its insane random access performance.
ERP applications which use HANA typically do have significant random accesses in their workload. It is due to this reason that it was initially marketed as an in-memory database. The random access performance of RAM was necessary for high performance. If HANA was just being used in in-memory mode, the speed increase is probably not that impressive. But, if there is non-volatile storage involved, then yes the speed increase would be impressive.
ERP applications which use HANA typically do have significant random accesses in their workload. It is due to this reason that it was initially marketed as an in-memory database. The random access performance of RAM was necessary for high performance. If HANA was just being used in in-memory mode, the speed increase is probably not that impressive. But, if there is non-volatile storage involved, then yes the speed increase would be impressive.
Couldn't tell you what it does, I only know that the advertising of it in airports is ubiquitous. In the dozen plus airports I've been in this year, they've all been plastered in adds for HANA.
I've used SAP HANA for about 5 months earlier in the year. Between the overselling marketing, the expensive cost of the thing, and the bugs + growing view that SAP are using customers as guinea pigs; I'd say it's like IBM Watson for data.
You can get good performance out of it, no medals there as any IMDB would give you that.
My pain points were:
- it is sold as the virtualisation layer to rule them all, until you find crazy breakages. SAP's SQL dialect is case sensitive. That'll frustrate a lot of downstream databases that you deal with, or the users thereof.
- virtual warehouse. I'm not a Kimball expert, but our resident experts came to the same conclusion as I did, that you can't really build useful VIRTUAL data marts directly from disparate downstream data sources without persisting that somewhere. That then defeats the purpose, in my quasi-professional opinion.
- I spent over a month looking at the SDA/SDI to "Big Data" parts. Very immature, found interesting instances where push-down of predicates were inconsistent and poor. It's one thing to say you've got Apache Spark integration, but it's another thing when your system can't figure out that data comes from the same system, and this you could do some expensive computations downstream instead of saturating the network pulling a lot of data. I did my work early this year, so I would expect improvement here. The problem though was that many of us started feeling like SAP are releasing a product that should spend more time in development before coming to us. I'm entitled to say so because they charge an arm and a leg.
- data modelling/development. SQL is Turing complete, but some things are still hard to do in it. Instead of giving us more capabilities when modelling our data, they're getting "distracted" by wanting to tuck marketing boxes like R support, NodeJS support, etc. I didn't spend enough time on those to feel comfortable with producing a fair opinion, so I'll reserve one.
- SDLC, security. SAP has solid products there, and I learnt quite a lot around security integration. There were some quirks, but those were largely because of our requirements, not SAP. I enjoyed working on data access and security patterns.
- IDE. They initially used a modified NetBeans, which was also available as a plugin. I'm not a NB person, but this was solid. They decided to move to a web-based IDE. I don't know why, but they deprecated certain things while their web thingy was not even feature complete. It was a growing mess when you'd find that you can only do something in X but not Y.
To recap, a high cost for performance that you could get from a competitor, or even a well baby-sat Apache Ignite cluster (for the features they support, at least they aren't lying in their marketing). I think of it as a lame caching layer, that's useful if you're already invested in the SAP ecosystem, think of SAP CRMs or even their warehousing as a stretch.
So, no; it's not revolutionary in my view. The "specialised hardware" is marketing bait, our IT (I too was) were convinced that one could build themselves the same hardware as was speced. When I first heard of the special hardware thing, I thought of a custom Intel with features that only SAP would have, DDR4 RAM at crazy clock speed, etc. Only to find that it's a sly way of controlling how much they charge you. "Want to add moar RAMs? Pay us so you don't vid warranty, and we'll charge you more of course for the extra RAM".
To conclude my rant, I hope you've found my view useful. My conclusion when I rolled off from my SAP HANA (notice how little I've used this acronym, I despise it) project feeling like it was making everyone in the room bad. People were forsaking their Oracle DBs and Hadoop clusters to make design compromises for an over-marketed database thingy.
You can get good performance out of it, no medals there as any IMDB would give you that.
My pain points were:
- it is sold as the virtualisation layer to rule them all, until you find crazy breakages. SAP's SQL dialect is case sensitive. That'll frustrate a lot of downstream databases that you deal with, or the users thereof.
- virtual warehouse. I'm not a Kimball expert, but our resident experts came to the same conclusion as I did, that you can't really build useful VIRTUAL data marts directly from disparate downstream data sources without persisting that somewhere. That then defeats the purpose, in my quasi-professional opinion.
- I spent over a month looking at the SDA/SDI to "Big Data" parts. Very immature, found interesting instances where push-down of predicates were inconsistent and poor. It's one thing to say you've got Apache Spark integration, but it's another thing when your system can't figure out that data comes from the same system, and this you could do some expensive computations downstream instead of saturating the network pulling a lot of data. I did my work early this year, so I would expect improvement here. The problem though was that many of us started feeling like SAP are releasing a product that should spend more time in development before coming to us. I'm entitled to say so because they charge an arm and a leg.
- data modelling/development. SQL is Turing complete, but some things are still hard to do in it. Instead of giving us more capabilities when modelling our data, they're getting "distracted" by wanting to tuck marketing boxes like R support, NodeJS support, etc. I didn't spend enough time on those to feel comfortable with producing a fair opinion, so I'll reserve one.
- SDLC, security. SAP has solid products there, and I learnt quite a lot around security integration. There were some quirks, but those were largely because of our requirements, not SAP. I enjoyed working on data access and security patterns.
- IDE. They initially used a modified NetBeans, which was also available as a plugin. I'm not a NB person, but this was solid. They decided to move to a web-based IDE. I don't know why, but they deprecated certain things while their web thingy was not even feature complete. It was a growing mess when you'd find that you can only do something in X but not Y.
To recap, a high cost for performance that you could get from a competitor, or even a well baby-sat Apache Ignite cluster (for the features they support, at least they aren't lying in their marketing). I think of it as a lame caching layer, that's useful if you're already invested in the SAP ecosystem, think of SAP CRMs or even their warehousing as a stretch.
So, no; it's not revolutionary in my view. The "specialised hardware" is marketing bait, our IT (I too was) were convinced that one could build themselves the same hardware as was speced. When I first heard of the special hardware thing, I thought of a custom Intel with features that only SAP would have, DDR4 RAM at crazy clock speed, etc. Only to find that it's a sly way of controlling how much they charge you. "Want to add moar RAMs? Pay us so you don't vid warranty, and we'll charge you more of course for the extra RAM".
To conclude my rant, I hope you've found my view useful. My conclusion when I rolled off from my SAP HANA (notice how little I've used this acronym, I despise it) project feeling like it was making everyone in the room bad. People were forsaking their Oracle DBs and Hadoop clusters to make design compromises for an over-marketed database thingy.
The hardware sale is usually a scam to yield fatter renewals.
EMC did something similar with avamar... you’d buy spray painted SuperMicro boxes runnding RHEL and pay for software as well. Come node end of life, you have to buy the software and hardware again, even after paying the 20-25% maintenance.
EMC did something similar with avamar... you’d buy spray painted SuperMicro boxes runnding RHEL and pay for software as well. Come node end of life, you have to buy the software and hardware again, even after paying the 20-25% maintenance.
Netapp too. A quote for a new filer and a 5th year of support on a 4 year old one are usually about the same.
Thank you and all the others for your sharing. Nice to have some honest opinions.
One of the most interesting applications of NVM is databases. The design of most existing databases is predicated on needing to always persist writes to disk; the existence of non-volatile memory allows very different, potentially much faster designs. There's a great explanation of this in https://db.cs.cmu.edu/papers/2017/p1753-arulraj.pdf
Does anyone know much about the tech on this? I assume when they say persistent they mean across VM restarts, but are they actually doing some sort of disk persistence too?
This is byte-addressable persistent memory. They look like DRAM DIMMs and they plug into DIMM slots. You access them using your memory controller and not your storage controller. People sometimes refer to them as non-volatile memory (NVM). Intel used to call it Apache Pass.
They're a nightmare to program because OSes do not have a good abstraction for them (at least not yet). Accessing them through the file-system seems sub-optimal (this is byte-addressable memory and not a block device). Accessing them through virtual memory is also pretty bad because they're much slower than DRAM.
They're a nightmare to program because OSes do not have a good abstraction for them (at least not yet). Accessing them through the file-system seems sub-optimal (this is byte-addressable memory and not a block device). Accessing them through virtual memory is also pretty bad because they're much slower than DRAM.
Disclaimer: I work at Intel on PMDK (pmem.io)
Both Windows and Linux implement DAX, which, as @the8472 explained, allows bypassing page cache in memory mapped I/O. Additionally, DAX optionally allows you to flush your data directly from user-space instead of calling msync.
And that's the gist of NVM programming model [0], its entire point is to allow applications to avoid the now hugely excessive abstraction layer of traditional storage.
And I will freely admit that programming to raw memory mapped files can be difficult, but there is ongoing work on making it easier. An example of that is, excuse the shameless plug, Persistent Memory Development Kit [1], which makes writing new software for this new type of memory much simpler.
Performance of an NVDIMM is obviously hardware dependent, but the now widely accepted programming model works with the assumption that persistent memory is fast enough so that it is reasonable to stall a CPU while an instruction is accessing it. I'm not sure on what hardware evaluations you are basing your claims on, but let me assure you that the HW solution being described in the blog post does not violate that assumption.
[0] - https://www.snia.org/tech_activities/standards/curr_standard...
[1] - http://pmem.io/
Both Windows and Linux implement DAX, which, as @the8472 explained, allows bypassing page cache in memory mapped I/O. Additionally, DAX optionally allows you to flush your data directly from user-space instead of calling msync.
And that's the gist of NVM programming model [0], its entire point is to allow applications to avoid the now hugely excessive abstraction layer of traditional storage.
And I will freely admit that programming to raw memory mapped files can be difficult, but there is ongoing work on making it easier. An example of that is, excuse the shameless plug, Persistent Memory Development Kit [1], which makes writing new software for this new type of memory much simpler.
Performance of an NVDIMM is obviously hardware dependent, but the now widely accepted programming model works with the assumption that persistent memory is fast enough so that it is reasonable to stall a CPU while an instruction is accessing it. I'm not sure on what hardware evaluations you are basing your claims on, but let me assure you that the HW solution being described in the blog post does not violate that assumption.
[0] - https://www.snia.org/tech_activities/standards/curr_standard...
[1] - http://pmem.io/
Do you know how well tools like Cap'n Proto and Protocol Buffers help for dealing with this kind of scenario? I'd imagine that some kind of low latency/cost serialization system would help significantly in using the device. Cap'n Proto I'd imagine would work nicely for reading data off since it should be able to read and use the structure with no extra copying or decoding, but I have no idea how the situation with writing would win out.
That's an excellent question.
The answer is that those type of libraries will work just fine for read-only workloads since you cannot mutate a data structure once you have serialized and written it out to a file.
The best part is, this will work without any (or very little) modifications, as long as your application is suited for using mmap. All you have to do is to use a persistent memory resident file on a DAX file system.
If you need dynamic mutable state however, as great as these libraries are, you will need a more complex solution with memory allocation and transactions.
If you need dynamic mutable state however, as great as these libraries are, you will need a more complex solution with memory allocation and transactions.
The simplest way of using this is to not do any serialization at all, just store any information you want persisted in memory allocated from the region you mmaped to the Optane DIMMs instead of the DRAM DIMMs.
The main reason I've been thinking about tools like that is more because the persistent structure should probably work regardless of compiler settings/flags and code updates. Directly mmaping the structures you're going to have to worry about how things are packed, and if a new compiler optimization causes things to go differently (something gets eliminated in one version and not the other).
I would say that this is still a mentality of thinking of something as "data on disk", where the data should be in an "ABI-stable" format.
Think of persistent-memory data as more like data resident in the memory of a runtime which can experience a "hot code upgrade", like the Erlang runtime.
In Erlang, when you hot-upgrade your running code, you usually do so through a managed system of "relups" (RELease UPdates), which are sort of a cross between an RDBMS migration, and a traditional installer-package full of newer versions of code and assets.
The Erlang runtime takes this package, unpacks it, and then runs a master relup script, which can been authored to do arbitrary things (including, if ultimately necessary, fully rebooting the node, throwing away all that in-memory state.) Mostly, though, a relup script calls into individual "appup" scripts for each Erlang application. Those applications then specify how their corresponding running processes are to be updated—which can sometimes be fraught (if e.g. the new release requires that you add new service-processes or remove old ones, migrating in-memory state into a new architecture), but usually just means calling a "code_change" callback on all the service-processes.
This "code_change" callback is the thing that's most like an RDBMS migration: it is called from the event-loop running in the old version of the code of the service-process, and passes in the old in-memory state; and when it returns, it's returning the new in-memory state, to resume the event loop in the new version of the code of the service-process.
This is basically how I'd picture dealing with code updates (including ones due to build-setting changes) in software that deals with pmem: you'd architect your code such that the library that touches the pmem can have multiple versions of it dynamically loaded (though not running) at the same time; and then you'd stage a migration from the old code's pmem state encoding, to the new version's, by
1. dlopen(2)ing the new version of the lib;
2. telling the old version of the lib to stop any ongoing work;
3. handing off the toplevel pmem state-handle that the old version of the lib was using, to a "migrate" function in the new version of the lib;
4. replacing the old version's pmem state-handle with a dummy one;
5. telling the old version of the lib to terminate (and so do the trivial cleanup to the world it sees through the dummy handle);
6. tell the new version of the lib to initialize, using the handle to the now-migrated-in-format pmem;
7. dlclose(2) the old version of the lib.
Basically, picture what something like Photoshop would have to do to enable you to upgrade its plugins without restarting it or closing your working document, and you'll have the right architecture.
Think of persistent-memory data as more like data resident in the memory of a runtime which can experience a "hot code upgrade", like the Erlang runtime.
In Erlang, when you hot-upgrade your running code, you usually do so through a managed system of "relups" (RELease UPdates), which are sort of a cross between an RDBMS migration, and a traditional installer-package full of newer versions of code and assets.
The Erlang runtime takes this package, unpacks it, and then runs a master relup script, which can been authored to do arbitrary things (including, if ultimately necessary, fully rebooting the node, throwing away all that in-memory state.) Mostly, though, a relup script calls into individual "appup" scripts for each Erlang application. Those applications then specify how their corresponding running processes are to be updated—which can sometimes be fraught (if e.g. the new release requires that you add new service-processes or remove old ones, migrating in-memory state into a new architecture), but usually just means calling a "code_change" callback on all the service-processes.
This "code_change" callback is the thing that's most like an RDBMS migration: it is called from the event-loop running in the old version of the code of the service-process, and passes in the old in-memory state; and when it returns, it's returning the new in-memory state, to resume the event loop in the new version of the code of the service-process.
This is basically how I'd picture dealing with code updates (including ones due to build-setting changes) in software that deals with pmem: you'd architect your code such that the library that touches the pmem can have multiple versions of it dynamically loaded (though not running) at the same time; and then you'd stage a migration from the old code's pmem state encoding, to the new version's, by
1. dlopen(2)ing the new version of the lib;
2. telling the old version of the lib to stop any ongoing work;
3. handing off the toplevel pmem state-handle that the old version of the lib was using, to a "migrate" function in the new version of the lib;
4. replacing the old version's pmem state-handle with a dummy one;
5. telling the old version of the lib to terminate (and so do the trivial cleanup to the world it sees through the dummy handle);
6. tell the new version of the lib to initialize, using the handle to the now-migrated-in-format pmem;
7. dlclose(2) the old version of the lib.
Basically, picture what something like Photoshop would have to do to enable you to upgrade its plugins without restarting it or closing your working document, and you'll have the right architecture.
So does that update process rewrite all 6 TB of data in the new format? Because I can imagine why people would rather not do that.
I mean, it depends on whether your pmem data is a bunch of copy-on-write persistent data structures like HAMTs; or maybe packed data structures like Vector<Foo>s where you can't easily rewrite one Foo to be a different size without rewriting the whole vector; etc.
If 1. the new format is just like the old format except for one little difference to one struct, and 2. structs point to other structs, rather than containing them; then it's just a matter of calling your within-mmap(2)ed-arena malloc(2)-equivalent function to get a new chunk of the pmem arena of the right size for the new version of the struct; and then rewriting the pointer in the other struct to point to it; and then calling your free(2)-equivalent on the old version of the struct.
If you change the structure of some fundamental primitive type like how strings are represented, then you're probably going to have to rewrite your whole pmem arena.
Though, also, you can just make your code deal with both old and new versions of the struct, and only migrate structs when they're getting modified anyway. (This is equivalent to the way you'd avoid an RDBMS migration rewrite an entire table, by instead adding a trigger that makes the migration happen to a row on UPDATE, and then ensuring that your business-layer can deal with both migrated and un-migrated versions of the row.)
If 1. the new format is just like the old format except for one little difference to one struct, and 2. structs point to other structs, rather than containing them; then it's just a matter of calling your within-mmap(2)ed-arena malloc(2)-equivalent function to get a new chunk of the pmem arena of the right size for the new version of the struct; and then rewriting the pointer in the other struct to point to it; and then calling your free(2)-equivalent on the old version of the struct.
If you change the structure of some fundamental primitive type like how strings are represented, then you're probably going to have to rewrite your whole pmem arena.
Though, also, you can just make your code deal with both old and new versions of the struct, and only migrate structs when they're getting modified anyway. (This is equivalent to the way you'd avoid an RDBMS migration rewrite an entire table, by instead adding a trigger that makes the migration happen to a row on UPDATE, and then ensuring that your business-layer can deal with both migrated and un-migrated versions of the row.)
> If you change the structure of some fundamental primitive type like how strings are represented, then you're probably going to have to rewrite your whole pmem arena.
That's part of the reason why I was thinking something like Cap'n Proto or Protocol Buffers might make sense for a lot of structures. You pay a bit of cost for writing but get to gracefully handle upgrades to the structure if you do it right. I'd imagine you want to use something higher level just above them to organize the records. But this is all a really new area of thinking about this so I'm probably being a bit obtuse about it.
That's part of the reason why I was thinking something like Cap'n Proto or Protocol Buffers might make sense for a lot of structures. You pay a bit of cost for writing but get to gracefully handle upgrades to the structure if you do it right. I'd imagine you want to use something higher level just above them to organize the records. But this is all a really new area of thinking about this so I'm probably being a bit obtuse about it.
Probably something like Flatbuffers, with which you can skip the marshalling and unmarshalling. It was designed for games, where you would map files into memory.
The ultimate zero-cost serialization system is just mmapping the Optane and casting pointers, but if you want something less fragile you could probably use one of those libraries on top of mmap.
> They're a nightmare to program because OSes do not have a good abstraction for them (at least not yet).
With DAX[0] linux already has the ability to put a filesystem (currently ext4 and xfs) on NVDIMMS and then let userspace address them through mmap while skipping the page cache indirection. I.e. you're directly byte-addressing them through the memory controller via standard memory-mapped file abstractions. Direct block device mapping of nvdimms without filesystem is also possible.
[0] https://www.kernel.org/doc/Documentation/filesystems/dax.txt
With DAX[0] linux already has the ability to put a filesystem (currently ext4 and xfs) on NVDIMMS and then let userspace address them through mmap while skipping the page cache indirection. I.e. you're directly byte-addressing them through the memory controller via standard memory-mapped file abstractions. Direct block device mapping of nvdimms without filesystem is also possible.
[0] https://www.kernel.org/doc/Documentation/filesystems/dax.txt
The Persistent Memory Development Kit (PMDK) offers high-level abstractions over DAX. Looks like most mortals would use libpmemobj (or its C++ bindings).
http://pmem.io/pmdk/
http://pmem.io/pmdk/
It seems like Single-level store[1] would be a really good fit for this. I was going to make a crack about bringing back Multics but apparently IBM has an OS using this.
[1]https://en.wikipedia.org/wiki/Single-level_store
[1]https://en.wikipedia.org/wiki/Single-level_store
Is there are a reason to not just use DRAM along with a battery to achieve the same persistence but as fast as DRAM?
Probably because systems that passively do their job tend to be preferred.
Also a battery would only last so long. IIRC DRAM needs to be constantly refreshed, so, it would be a trade-off between capacity and duration.
Optane seems[1] to be 20~30X slower than DRAM but 4~10X faster than server SSDs
[1]: https://superuser.com/a/1195674/187732
Also a battery would only last so long. IIRC DRAM needs to be constantly refreshed, so, it would be a trade-off between capacity and duration.
Optane seems[1] to be 20~30X slower than DRAM but 4~10X faster than server SSDs
[1]: https://superuser.com/a/1195674/187732
There are NVDIMMs that have DRAM and a matching quantity of NAND flash memory to save the contents to in the event of a power failure. They require an external capacitor module and are limited in data capacity by how much DRAM you can fit on the module. You can fit far more 3D XPoint memory on a module than DRAM, and it doesn't require the external capacitors to achieve persistence, and it should be significantly cheaper on a per-GB basis.
Just to reiterate the point... this is an instance with terabytes of near-memory-speed storage.
If persistent memory pans out as a technology it will completely upend the way we think about building software and the cost tradeoffs of hardware. (as much as or more so than the transition from spinning disks to ssds)
If persistent memory pans out as a technology it will completely upend the way we think about building software and the cost tradeoffs of hardware. (as much as or more so than the transition from spinning disks to ssds)
How do you define "pans out"? What performance and price differences between it, flash, and DRAM do you have in mind?
Because I'll keep reminding people that putting a DRAM cache in front of some flash can very closely approximate a large persistent memory. If people wanted to build software for that kind of system, they could do it today. The hardware is not the blocker.
Because I'll keep reminding people that putting a DRAM cache in front of some flash can very closely approximate a large persistent memory. If people wanted to build software for that kind of system, they could do it today. The hardware is not the blocker.
Diablo Memory1 used flash DIMMs with DRAM DIMMs as cache. When we tried it, it worked OK for some workloads and poorly for others, but it was also buggy and then the company went out of business. So a large part of "pans out" is simply a production-quality implementation that you can buy.
Note that Optane DIMMs have been delayed by around two years at this point and we still don't know what they will cost.
Note that Optane DIMMs have been delayed by around two years at this point and we still don't know what they will cost.
Cache will always stay just cache, unless it is the same size that an underlying persistent storage. You cannot read or write larger-than-cache chunks without performance degradation. Also you have a start-up cache-warming problem.
Most workloads don't need the entire storage to be at maximum speed all the time. In other words, in most situations cache is plenty for speed purposes. And the mapping layer can hide the chunk sizes. But we still don't see people writing software based around persistence. Maybe a lot of people are simply stuck in their ways, or maybe the benefits aren't actually that big.
As for cache-warming, that's also a configuration issue. When you reboot, leave the 'cache' portion of DRAM alone. Then as soon as the service resumes, the cache is already hot. When you shut down a node for an extended period, consider spending five minutes writing the cache to disc. And the article is about cloud servers anyway, where a shutdown typically implies losing all local storage whether it's persistent or not.
As for cache-warming, that's also a configuration issue. When you reboot, leave the 'cache' portion of DRAM alone. Then as soon as the service resumes, the cache is already hot. When you shut down a node for an extended period, consider spending five minutes writing the cache to disc. And the article is about cloud servers anyway, where a shutdown typically implies losing all local storage whether it's persistent or not.
The reason is the battery. These devices can be powered off and save state, like an SSD.
But in the context of the OP, presumably devices in a data center would never be powered down on purpose to save energy? In which case it seems that battery-backup DRAM would work just as well for this use case, and be both cheaper and faster.
Optane is in between DRAM and flash in terms of cost (and performance) - it's also denser, so you can fit much more storage on each DIMM slot.
DRAM can not be simply naively battery-backed; it needs active refreshing. And as memory controllers reside in CPUs these days, that would mean keeping the CPU powered up.
It uses Intel Optane.
It's basically super fast SSD used as RAM.
How does this compare to NVMe based AWS EC2 instances like m5d, c5d, r5d?
It's between RAM and SSD. Much closer to RAM speeds but data persists after turning off power. Biggest advantage over just being a faster disk is that it's byte-addressable random access like memory, so you don't have to deal with writing pages/blocks to drives.
The throughput and latency is fast enough that many apps can skip using memory and work straight from this. We already see this in some mobile devices that only have solid state and instant boot times because the data is always ready, with no need to shift from disk to ram.
It'll be awhile before OS and applications take advantage but it has potential to be a major shift in how storage works.
The throughput and latency is fast enough that many apps can skip using memory and work straight from this. We already see this in some mobile devices that only have solid state and instant boot times because the data is always ready, with no need to shift from disk to ram.
It'll be awhile before OS and applications take advantage but it has potential to be a major shift in how storage works.
Disclosure: I work on Google Cloud.
NVMe is (now) a more general purpose "speak to flash like things". This Optane memory stuff is made of "flash", but unlike our Local SSD offering (or AWS's i3) it's at a latency and throughput closer to DRAM.
Despite it being marketing material, I find the pyramid diagram [1] helpful. This blog post is about "Optane Persistent Memory".
[1] https://newsroom.intel.com/wp-content/uploads/sites/11/2018/...
NVMe is (now) a more general purpose "speak to flash like things". This Optane memory stuff is made of "flash", but unlike our Local SSD offering (or AWS's i3) it's at a latency and throughput closer to DRAM.
Despite it being marketing material, I find the pyramid diagram [1] helpful. This blog post is about "Optane Persistent Memory".
[1] https://newsroom.intel.com/wp-content/uploads/sites/11/2018/...
What is its latency compared to RAM, 10X ?
Most latency comparisons mentioned in other comments compare RAM to pcie based optane memory, not the DRAM optane memory.
Edit: Article[1] here says the latency was 40µs with 13M IOPS, if you consider RAM latency to be 100ns[2], then it looks like 40X
[1] - https://blogs.technet.microsoft.com/filecab/2018/10/30/windo... [2] - https://gist.github.com/jboner/2841832
Edit: Article[1] here says the latency was 40µs with 13M IOPS, if you consider RAM latency to be 100ns[2], then it looks like 40X
[1] - https://blogs.technet.microsoft.com/filecab/2018/10/30/windo... [2] - https://gist.github.com/jboner/2841832
A better comparison would be EC2 x1 instances which are high memory instances designed for SAP HANA like these ones: https://aws.amazon.com/ec2/instance-types/x1/
Redis on Optane [1]. The GraphBLAS stars continue to align [2]. GPUs/TPUs next. Distributed to come.
[1] Redis on Optane https://redislabs.com/blog/redis-enterprise-flash-intel-opta...
[2] GraphBLAS & RedisGraph https://news.ycombinator.com/item?id=18099520
[1] Redis on Optane https://redislabs.com/blog/redis-enterprise-flash-intel-opta...
[2] GraphBLAS & RedisGraph https://news.ycombinator.com/item?id=18099520
Would this be a good fit for Smalltalk, as it is image based? It seems it can remove a level of complexity, if you can treat your image, and so all your objects, as constantly persisted?
Can't you already do this with an OODB? I'm asking sincerely, since I have never quite understood why OODBs aren't used more often.
What is the benefit ?
Why are we constantly creating pets out of what should be cattle ?
Why are we constantly creating pets out of what should be cattle ?
The benefits of persistent memory, as I understand it, are principally improved performance (higher throughput, lower latency) and secondarily ease of programming. (Edit: another obvious benefit is cost per gigabyte; this stuff is cheaper than RAM.)
The benefits of having access to that technology on GCP (or another cloud) are the usual: reduced operational burden, increased availability, flexible pricing structure, elastic scalability, etc.
Your question about pets vs cattle is a non sequitur. Nothing about this announcement or the underlying technologies are suggestive of how they should be used (or misused). It's just a new tool in the toolbox—and as a distributed systems engineer specializing in database technology, having easy access to this hardware at scale is extremely compelling.
The benefits of having access to that technology on GCP (or another cloud) are the usual: reduced operational burden, increased availability, flexible pricing structure, elastic scalability, etc.
Your question about pets vs cattle is a non sequitur. Nothing about this announcement or the underlying technologies are suggestive of how they should be used (or misused). It's just a new tool in the toolbox—and as a distributed systems engineer specializing in database technology, having easy access to this hardware at scale is extremely compelling.
too bad you got downvoted, i would imagine in-memory DB and large cache.
> too bad you got downvoted
It's pretty weird to call out people for making 'pet' servers because they're daring to attach storage to them.
It's pretty weird to call out people for making 'pet' servers because they're daring to attach storage to them.
His native language isn't English, and people probably misunderstood his literally-translated idiom. Pet servers?
Pet's vs cattle is an analogy that originated with English-speakers.
http://cloudscaling.com/blog/cloud-computing/the-history-of-...
http://cloudscaling.com/blog/cloud-computing/the-history-of-...
As I replied in a sub-thread, I think Intel's marketing diagram [1] is probably useful to help separate the Optane flavors. This is about the "near DRAM" variant.
While the blog post highlights running SAP HANA (SAP's in-memory focused database), you can use them for whatever you want. The persistent part is that it's persistent across reboots. The hope is that this might make it easier to have tiered database/caching systems, since the gap between DRAM and this new "memory" is much closer than say DRAM and SSD.
[1] https://newsroom.intel.com/wp-content/uploads/sites/11/2018/...