IT Hub · MyWorkLog Back
Deep Dive · Storage technology

How an SSD remembers

Nothing spins inside an SSD. Data still survives when the power is gone. The trick is a tiny charge trapped inside insulated cells. From there the path runs through NAND pages, blocks and a translator to the file — and a direct comparison with a hard disk shows why the two drives behave so differently.

CHAPTER 1The cell holds charge

A flash cell is a transistor with one extra, fully insulated gate. This floating gate sits between the control gate and the channel, and it can hold on to electrons. Every electron up there screens part of the gate field, so the transistor needs more voltage before it conducts. That shifted threshold is exactly what the controller reads back as a data value.

During programming, a high voltage on the control gate pulls electrons out of the channel and through the thin oxide layer onto the floating gate — Fowler-Nordheim tunnelling. Erasing reverses the field and the charge travels back. That oxide layer is the insulator and at the same time the most fragile part of the cell: every cycle leaves a few charges stuck inside it.

Figure 1Charge on the floating gate shifts the threshold
Charge state
Electrons
432 e⁻
Threshold
1,8 V
State
00
Bits per cell
2 MLC

An MLC cell stores two bits across four distinguishable voltage windows; TLC uses eight, QLC sixteen. The erased state 11 sits below zero volts — the cell already conducts without any gate voltage. The levels are coded so that neighbours differ in exactly one bit (Gray code): misreading a neighbouring level therefore costs at most one bit instead of two.

The electron count here is derived from the threshold shift (Q = C · ΔVTH, roughly 120 electrons per volt). In the smallest cells only a few dozen are left — there, practically every single electron counts.

So the cell does not store “a file”, and it does not directly store a zero or a one either. It stores an analogue amount of charge that the controller sorts into a voltage window. Many such windows make a page, pages make a block, and blocks make a drive.

There is no floating gate in there any more

Current drives stack their cells more than 200 layers deep (3D NAND) and usually no longer trap the charge in a conducting gate but in an insulating nitride layer — charge-trap flash. The principle is unchanged: trapped charge shifts the threshold. Because stacking removes the pressure to shrink in area, the individual cells actually grew again, which noticeably improved endurance and error rates.

CHAPTER 2Many cells become NAND

NAND flash wires cells into long chains. A page is the smallest unit the controller can program — often 4 to 16 KiB of payload plus room for error correction. Erasing, on the other hand, only works per block: a block holds hundreds to thousands of pages and is reset to the erased state in one go by a high voltage on the substrate.

That gives the rule everything else follows from: programming happens page by page, erasing only block by block — and an occupied page cannot simply be overwritten. Change a page and you get a new one.

Figure 2Writing fills pages, erasing clears the whole block
Valid pages
2 / 8
Stale pages
0
Free pages
6
Erase cycles
0

“Change page” shows the core of it: the new version lands in a free page, the old one is merely marked stale. The space only comes back when the whole block is erased — garbage collection takes care of that later. Once no free page is left, erasing is the only option.

The block here has eight pages so you can count them; real blocks hold hundreds to thousands, and therefore several megabytes. The timings beside it are orders of magnitude for TLC NAND: reading is cheap, programming costs several times as much, erasing considerably more again.

CHAPTER 3The controller translates

The operating system thinks in logical blocks: LBA 42 is always LBA 42. The NAND pages underneath have a limited lifetime, cannot be overwritten and are erased only in whole blocks. So the SSD controller slots a translation table between the two worlds: the Flash Translation Layer, or FTL.

Figure 3The same logical address, a different page every time
Fill level
Logical
LBA 42
Physical
B3 · P07
Write amplification
1,72×
Free pages
56 %

The logical address stays put, the physical one moves with every write — preferably into a different block, so the ageing spreads out (wear levelling). The old page turns red: occupied, but worthless.

Write amplification is the ratio of flash data actually written to the payload the operating system handed over. The slider uses the simple approximation WA = 1/(1−u) for random write patterns: the fuller the drive, the more valid pages garbage collection has to drag along while cleaning up. Real controllers pick victim blocks more cleverly and separate hot from cold data, so they get off more lightly — but the trend holds.

The FTL is the reason an SSD is not just a large USB stick. It works constantly in the background: wear levelling spreads the ageing, ECC reconstructs individual flipped bits, garbage collection frees blocks up, and TRIM tells the drive which pages the file system no longer needs at all. Without TRIM the controller treats deleted files as valid data and dutifully copies them along during cleanup. On top of that comes a reserve the user never sees — over-provisioning, typically 7 to 28 percent. Together they explain why a drive filled to the brim collapses on writes.

CHAPTER 4SSD versus HDD: two routes to the same bit

A hard disk stores bits as the magnetisation of tiny regions on spinning platters. Before a single byte flows, two things have to happen: the arm swings the head onto the right track, and then the platter keeps turning until the wanted sector passes underneath. Both are mechanical waits — and both can be calculated. Half a revolution at 7200 rpm always takes 4.2 ms, no matter how good the firmware is.

The SSD has nothing to move. It looks the address up in the FTL, reads the flash page and pushes the data out. No travel, no waiting for a revolution.

Figure 4The same read request to both drives, in slow motion
SSD
0,09 ms
HDD
12,08 ms
Factor
138×
SSD accesses meanwhile
0

Seek the trackWait for the sectorTransfer data

The simulation runs in slow motion — one second on screen is four milliseconds inside the drive. Both sides are modelled in comparable class: 7200 rpm, 200 MB/s and a square-root seek model against a SATA SSD at 550 MB/s with 60 µs of NAND read time. With random 4 KiB there are two orders of magnitude between them. With a megabyte in one piece the lead shrinks to less than threefold, because only transfer rate is left to decide — and that is exactly where hard disks stayed competitive. An NVMe SSD would win this field clearly again.

In IOPS, single small accesses come out at roughly 11,000 against 80. Data sheets still like to quote six-figure numbers for SSDs — those apply to deep queues, where the drive spreads dozens of requests across different NAND dies at once. That is precisely what a hard disk cannot do: it has one arm, and the arm is only ever in one place.

Worth remembering

An SSD is not simply a hard disk without platters. It solves the same problem with different physics: electrons instead of magnetisation, mapping instead of head position, and error correction instead of “a bit is a bit”.

Where the hard disk wins has nothing to do with speed: price per terabyte, capacity per enclosure, and wear behaviour that does not depend on how much you write. Where the SSD wins: latency, small accesses, parallelism, silence, shock resistance and power draw.

CHAPTER 5Why flash wears out

Every program-erase cycle stresses the tunnel oxide. Some of the electrons stay stuck inside instead of passing through. The result: the threshold of a programmed cell no longer sits exactly where it should, and the distributions widen. The usable voltage window does not grow along with it — it is set by the physics of the cell, not by the number of bits you would like to put inside.

That is exactly where the trade-off between SLC, MLC, TLC and QLC comes from. Put four bits into a cell and you have to squeeze sixteen levels into the same window: the distance between two levels drops to an eighth of what SLC has available — with practically unchanged noise.

Figure 5Fixed window, more levels, less margin
Wear
Bits per cell
3
Level spacing
750 mV
P/E cycles
0 / 3.000
Raw bit error rate
1,2 · 10⁻⁵

The raw bit error rate here is computed from the overlap of neighbouring Gaussian distributions, not looked up: it shows the order of magnitude, not a specific part. The error correction in modern drives (LDPC) copes with rates up to roughly 10⁻²; once it gets tighter than that, the firmware retires the block and takes a spare.

The cycle counts are typical orders of magnitude for each cell type. Over-provisioning, controller and firmware shift them considerably, and wear levelling makes sure it is never the same blocks that age.

This is why almost every TLC and QLC drive deliberately runs part of its cells in SLC mode: one bit only, huge margins, very fast writes. As long as the data fits into that cache, you see the full data-sheet speed. Once it is full, the controller has to move data over to TLC or QLC while still accepting writes — and that is where the curve collapses during large copies, often from several gigabytes per second to a fraction of it. Not a defect, but the design.

CHAPTER 6What the SSD leaves behind

When you plug in an SSD you see no cells and no NAND blocks, only files. In between works a chain of charge, voltage windows, error correction and address translation. Those layers explain everyday behaviour: SSDs feel instant because the mechanical path is missing. They slow down when full, because write amplification climbs. And they still last, because the controller spreads the wear. Hard disks meanwhile remain what they always were: cheap, large, and not that far behind on long sequential reads.

Sources and further reading

More in the IT Hub

This deep dive belongs to the IT section of MyWorkLog. That is also where the skill tree lives — it turns your training log entries into levels and quizzes you on what you know.