2021 Media Shipments
|
Exabytes |
Revenue |
$/GB |
| Flash | 598 | $68.6B | $0.115 |
| Hard Disk | 1418 | $28.0B | $0.020 |
| LTO Tape | 59.2 | $0.51B | $0.003 |
I summed up my big picture view of archival media eight years ago in
Archival Media: Not a Good Business. Whatever your choice of technology, the economics are brutal. This is a version of the table in that post, updated to 2021 and again based upon
IBM data. Note that archival media shiped 4% as many bytes as hard disk and generated 1.8% or the revenue. It remains a tiny market.
I have written several times about Microsoft Research's
Project Silica, most recently in last year's
Archival Storage:
I'm skeptical of "commoditizing the technology". Archival systems are a niche in the IT market, and one on which companies are loath to spend money. Realistically, there aren't going to be a vast number of Silica write heads. The only customers for systems like Silica are the large cloud providers, who will be reluctant to commit their archives to technology owned by a competitor. Unless a mass-market application for femtosecond lasers emerges, the scope for cost reduction is limited.
But the more I think about this technology, which is still in the lab, the more I think it probably has the best chance of impacting the market among all the rival archival storage technologies. Not great, but better than its competitors:
I followed this with a list of eight major reasons for my opinion.
Below the fold an update on the project and an assessment.
At the 2026
Library of Congress' Designing Storage Architectures meeting Richard Black of Microsoft Research gave a presentation on their
Project Silica. He announced that the "research phase is complete", and pointed to three significant papers laying out their achievements:
- Project Silica: Towards Sustainable Cloud Archival Storage in Glass 23rd October 2023.
- RASCAL: A Scalable, High-redundancy Robot for Automated Storage and Retrieval Systems, 8th August 2024.
- Laser writing in glass for dense, fast and efficient archival data storage 17th February 2026.
The team agreed with my earlier assessment that a pre-condition for commercialization is cost-reducing the femtosecond laser,
writing:
Viable path towards commercialization gated only by the laser
My view is that for this to happen there needs to be a big market for these lasers, and that archival media isn't big enough, leaving the technology in a chicken-and-egg situation.
Paper 1
I discussed this paper in 2024's
Microsoft's Archival Storage Research, pointing out that the idea of data in silica dated back
at least to 2009, and writing:
But in the last few years Microsoft Research has taken this idea and run with it, as they report in a 68-author paper at SOSP entitled Project Silica: Towards Sustainable Cloud Archival Storage in Glass. It is a fascinating paper that should be read by anyone interest in archival storage.
I compared Project Silica with Facebook's 2013 development of two cold storage systems, one based on spun-down hard drives and the other on robots holding Blu-Ray disks. Facebook established two important attributes of such systems:
- The key performance criterion was write bandwidth, because reads were extremely rare. They expected the major reason for a read would be a subpoena.
- The economics depended upon operating at cloud scale, and thus being able to house the systems in normal warehouse space with no special air conditioning or power supplies, instead of expensive data center space. The key criterion was worst-case power draw.
Project Silica's design was based upon an analysis of traffic at Azure's tape-based archival layer. This was both interesting in itself and in contrast to the traffic to Facebook's cloud storage. I
wrote:
Note these important differences between Microsoft's and Facebook's storage hierarchies:
- Microsoft stores generic data and depends upon user action to migrate it down the hierarchy to the archive layer, whereas Facebook stores 9 specific types of application data which is migrated automatically based upon detailed knowledge of the workload for each of the types.
- Because Facebook can migrate data automatically, it can interpose a warm layer above the archive layer of the hierarchy, and because it has detailed knowledge about the behavior of each of the data types it can make good decision about when to move each type down the hierarchy.
- Because the warm layer responds to the vast majority of the read requests and schedules the downward migrations, Facebook's archive layer's IOPS are a steady flow of large writes with very few reads, making efficient use of the hardware.
Contrast Facebook's consistent scheduled ingest flow with the bursty ingest rate shown in Figure 2 of the
Silica paper. The analysis of their archive's workload in Section 2
shows that:
on average for every MB read there are 47 MBs written, and for every read operation there are 174 writes. We can see some variation across months, but writes always dominate by over an order of magnitude.
...
Small files dominate the workload, with 58.7% of the reads for files of 4 MiB or smaller. However, these reads only contribute 1.2% of the volume of data read. Files larger than 256 MiB comprise around 85% of bytes read but less than 2% of total read requests. Additionally, there is a long tail of request sizes: there is ∼ 10 orders of magnitude between the smallest and largest requested file sizes.
...
We observe a variability in the workload within data centers, with up to 7 orders of magnitude difference between the median and the tail, as well as large variability across different data centers.
...
At the granularity of a day, the peak daily [ingress] rate is ∼16x higher than the mean daily rate. As the aggregation time increases beyond 30 days, the peak over mean ratio decreases significantly down to only ∼2, indicating that the average write rate is similar across different 30-day windows.
...
To summarize, as expected for archival storage, the workload is heavily write-dominated. However, unexpectedly, the IO operations are dominated by small file accesses.
There is a great deal of useful information like this in the paper.
Paper 2
This paper covers Project Silica's
innovative robotics:
we present RASCAL, a novel ASRS robot for small payload items in structured environments, with a focus on system-level scalability and redundancy. We describe the design objectives of RASCAL and how they address some of the limitations of existing robotic systems in this area, such as scalability and redundancy. We then demonstrate the viability of our design with a proof-of-concept implementation of a data centre storage media robot, and show through a series of experiments that its design, speed, accuracy, and energy efficiency are appropriate for this application.
As shown in their Fig. 1, RASCAL consists of an array of shelves with a set of small robots that can access anywhere in the array by moving horizontally along rails at the edge of each shelf, and vertically by unclipping from one rail and clipping to the one two above or below. With my comments, their
design goals were:
- Serviceability: Robots should be easy to add and remove from the system, without requiring specialist tools or expertise.
The operator just clips or unclips the robot from the pair of rails to which it is attached by the "wings" that carry the wheels that contact them. Servicability is a significant problem with tape robots.
- Addressability: Any robot should be able to access any item stored in the shelving.
Because each robot blocks only a small fraction of the shelf to which it is attached, other robots can navigate around it to access items on the same or other shelves.
- Scalability: A deployment should be able to scale its retrieval throughput by adding and removing robots. Likewise, it should be able to easily extend or reduce its storage capacity by adding or removing shelving without significant downtime.
Adding or removing relves of robots requires updating the central control computer, but should not interrupt operations.
- Availability: Failure of a robot should have a limited effect on the items that can be accessed, should not have a significant impact on routing, and should not obstruct other robots from continuing their operations.
A failed robot blocks access to the items immediately in front of it, but these are only a tiny fraction of the total.
The team is very clear that what they are describing is a proof-of-concept demonstrating the feasibiliy of this kind of robot, using the test case of the
silica tablets:
The proof-of-concept robot is around 240 mm wide, and sits 300 mm tall and 90 mm deep when mounted on the rails. The picker adds an additional 70 mm to the width, and increases the overall depth to around 150 mm. When flipping, the robot extends a maximum of 270 mm from the structure, and 250 mm from the wing pivot point. The total mass of the robot (including picker and battery) is around 3.5 kg.
 |
| Rascal Fig. 2 |
The way the robot moves horizontally is obvious, but the way it moves
vertically isn't, as shown in their Fig. 2:
The climbing manoeuvre ... consists of unlatching one wing from its current rail while remaining firmly attached with the other; rotating the robot outwards from the storage rack around the attached rail and wing; and latching the free wing onto a new rail, two rails either above or below its original position.
The result is
that:
These motion systems equip our robot with two key properties: independence and flexibility.
In this context, independence refers to the fact that a RASCAL does not depend on the state of any other robot or external motion system (such as an elevator) to perform its operations.
In a deployed Silica system reads would be rare, so the job of the robots would almost exclusively be to shuttle blank tablets to the write head(s) and written tablets to their resting place on the shelves. Because "the peak daily [ingress] rate is ∼16x higher than the mean daily rate" the robots must be over-provisioned, with most idle much of the time. This means there would typically be ample time for recharging their batteries.
Paper 3
This paper covers the physics of recording and reading
data in glass. Their abstract reads:
Here we report an optical archival storage technology based on femtosecond laser direct writing in glass that addresses the practical demands of archival storage, which we call Silica. We achieve a data density of 1.59 Gbit mm−3 in 301 layers for a capacity of 4.8 TB in a 120 mm square, 2 mm thick piece of glass. The demonstrated write regimes enable a write throughput of 25.6 Mbit s−1 per beam, limited by the laser repetition rate, with an energy efficiency of 10.1 nJ per bit. Moreover, we extend the storage ability to borosilicate glass, offering a lower-cost medium and reduced writing and reading complexity. Accelerated ageing tests on written voxels in borosilicate suggest data lifetimes exceeding 10,000 years.
The paper claims advances in four main areas. First,
Writing data:
Two efficient regimes of volume pixel (voxel) writing in glass: we use phase voxels relying on isotropic refractive index (RI) changes and birefringent voxels based on anisotropic changes ... We demonstrate high-quality voxels, each storing more than one bit, using a minimum number of pulses.
They describe phase voxel writing
thus:
Phase voxels are femtosecond-laser-induced isotropic modifications with locally altered [refractive index] and minimal optical scattering. The pulse energy is modulated to encode the symbol. ... Each voxel is written with a single pulse, so voxels are written at the laser repetition rate of 10 MHz. We modulate the beam energy using an acousto-optic modulator ... Different symbols correspond to distinct [refractive index] changes that can be read using Zernike phase-contrast microscopy
Furthermore, we demonstrate a throughput of 65.9 Mbit s−1 by splitting the laser into 4 independently modulated beams. We scan all beams with the same scanner and objective .... The written, read and decoded results show that throughput can be scaled in this way without damaging the media.
The reason the 4-beam throughput is more than 4 times the 10MHz laser pulse rate is that each pulse writes more than one bit, up to 1.8, per voxel.
They describe birefringent voxel writing
thus:
Birefringent voxels are composed of optically anisotropic sub-diffraction modifications, the in-plane orientation of which is determined by the polarization of the writing pulse. Varying this orientation encodes different data symbols, which we read using polarization-resolved imaging.
...
Our new pseudo-single-pulse regime shows the formation of elongated nanovoids with just two pulses, improving on previous work. We split each pulse into two: one that forms a void (seed pulse) and the other that elongates a previously formed void (data pulse). ... In this way, a single laser pulse simultaneously initiates the formation of a new seed structure and converts an existing seed structure into a data voxel, so voxels are written at the laser repetition rate of 10 MHz
Second,
Emissions-based control of voxel writing:
High-throughput, stable writing: we demonstrate writing at high throughput using multiple beams per laser ... We use a closed-loop feedback system to actively monitor and optimize the laser power, providing precise energy stability during writing and enabling predictability and reliability across different writers at scale ...
Using closed-loop feedback to control paramteres of the writing process is important because the process requires very high precision to achieve its very high volumetric density.
Third, as regards
Lifetime they address the question of whether the fact that borosilicate class is very stable means that data recorded in it is very stable:
To assess the thermal stability of phase voxels, we perform accelerated ageing experiments based on the Arrhenius law, using visible light diffraction measurements to track the decay of written structures. ... Extrapolation from the measurement points at elevated temperatures suggests exceptional long-term stability, indicating a modification lifetime that exceeds 10,000 years at 290 °C and therefore even longer at room temperature. This lifetime reflects the thermal stability of phase voxels under isolated conditions and does not account for external influences, such as mechanical stress or chemical corrosion, which are beyond the scope of this study.
As is normal in accelerated aging tests, there is a very large extrapolation from the measurements they made at temperatures from 500-440C (see graph) to the likely storage temperatures. But the result of the extrapolation is such a long lfe that the media are effectively immortal given "benign neglect".
Fourth, they describe their approach to the often overlooked complexity of
Reading and decoding data:
Machine learning decode: building on our previous work, here we apply machine-learning-based decode ... to account for noise and inter-voxel cross-talk.
This is an excellent way of implementing the pattern-matching that is needed to extract bits from the images of the nanopores from the camera. The most overlooked aspect of storage is the problem of converting the actual noisy analog signal from the media into bits.
They provide performance numbers for both types of media. First,
birefringent voxels:
Using birefringent voxels, in fused silica glass, we achieve 1.59 Gbit mm−3 data density (usable capacity of 4.84 TB per platter, 0.500 μm × 0.485 μm voxel pitch and 6 μm layer spacing, 301 layers, 8 azimuth levels at 0.85 quality factor), a write throughput of 25.6 Mbit s−1, and a write efficiency of 10.1 nJ per bit.
Second,
phase voxels:
Using phase voxels, in borosilicate glass we achieve 0.678 Gbit mm−3 data density (usable capacity 2.02 TB per platter, 0.5 μm × 0.7 μm voxel pitch, 7 μm layer spacing, and 258 layers, 4 energy levels at 0.92 quality factor), a write throughput of 18.4 Mbit s−1, and a write efficiency of 8.85 nJ per bit. Furthermore, our multibeam system achieves a throughput of 65.9 Mbit s−1 through parallel writing with four beams without inducing thermal damage. Thermal simulations indicate that writing with 16 or more beams should be possible
For comparison, a 4TB M.2 2280 SSD has 1.34TB mm
-3, comparable to the density in fused silica. But, of course, most of the SSD's volume is the PCB, not the actual medium.
So which type of
media is preferred?
We have shown that birefringent voxels achieve higher key metrics than phase voxels. However, efficient formation of birefringent voxels can be achieved only in high-purity silica glasses, whereas phase voxels can be written in potentially any durable transparent media, for example, borosilicate glass as demonstrated here. For phase voxels, the writing and reading hardware are simpler, requiring only one modulator per beamline and only one camera per reader, respectively. Both regimes can match the maximum laser repetition rate of 10 MHz or higher.
Assessment
In my view Project Silica was was not just excellent research but also, like
Facebook's earlier systems using spun-down hard drives and optical media robots, a really praiseworthy attempt to craft a technological solution to the extremely difficult economics of the market for archival media. Their technology had many important attributes
- The media is very cheap and very dense, so the effect of Kryder's Law economics driving media replacement and thus its economic rather than technical lifetime is minimal.
- The media is quasi-immortal and survives benign neglect, so opex once written is minimal.
- The media is write-once, and the write and read heads are physically separate, so the data cannot be encrypted or erased by malware. The long read latency makes exfiltrating large amounts of data hard.
- The robotics are simple and highly redundant. Any of the shuttles can reach any of the platters. They should be much less troublesome than tape library robotics because, unlike tape, a robot failure only renders a small fraction of the library inaccessible and is easily repaired, simply by removing and replacing the failed shuttle.
- All the technologies needed are in the market now, the only breakthroughs needed are economic, not technological.
- The team has worked on improving the write bandwidth which is a critical issue for archival storage at scale. They can currently write hundreds of megabytes a second.
- Like Facebook's archival storage technologies, Project Silica enjoys the synergies of data center scale without needing full data center environmental and power resources.
- As Facebook's technologies had, Project Silica has an in-house customer, Azure's archival storage, with a need for a product in this space.
Alas, my prediction is that this excellent technology will fail in the market. The market is too small to cost-reduce the lasers. The incumbent, LTO tape, is well established and has a credible road-map. It is built into the processes of the likely customers, making a technology transition risky. And with non-zero interest rates it is hard to justify spending more capex now to reduce, or in this case almost completely eliminate, future opex.
No comments:
Post a Comment