What Disc Assessment Actually Looks Like in Practice

Most people think disc assessment is just running a tool and getting a pass or fail result. It's not. You're evaluating how a storage medium performs under real conditions, and the gap between benchmark numbers and what your application actually experiences is where things fall apart. I spent years dealing with enterprise SAN environments, and the difference between a theoretical assessment and one that tells you anything useful comes down to workload simulation. Here's the basic flow. You define what "useful" means for your scenario first. Then you pick your workload profile. Then you run measurements. Then you interpret results against your actual application requirements, not some generic standard. The order matters because most people skip the first step and end up with data that looks good on paper and completely useless in production.

How Does Disc Assessment Work at the Core Level

At its foundation, disc assessment measures four things: sequential read and write throughput, random read and write IOPS, latency distribution, and sustained performance over time. Most automated tools give you all four numbers in a table and call it a day. The actual assessment requires you to understand which numbers correlate to your workload and which are noise. Sequential throughput matters for backup systems, video editing, and large file transfers. Random IOPS matters for databases, virtualization hosts, and anything doing heavy transaction logging. Latency distribution matters for everything else. A drive might advertise 50,000 IOPS at 99th percentile latency of 2 milliseconds, but if your application can't tolerate more than 500 microseconds, those headline numbers are misleading you. I once worked on a deployment where the initial assessment used a standard fio profile with 4K random reads at queue depth 32, got great numbers, and the system was deployed. Three months later, users were complaining about timeouts during peak hours. The problem wasn't the disc—it was that our workload profile had a 50-50 read-write split in the test, but production was running at 95% writes with occasional large sequential bursts from a reporting batch job. The drive's write caching handled the balanced test fine, but sustained sequential writes during the batch pushed the cache to its limit and latency spiked to over 20 milliseconds. We re-profiled the assessment using a 90/10 write-heavy pattern with periodic large sequential inserts, and the same drive that passed with flying colors now showed a 15-millisecond sustained write latency that made it unsuitable for that workload.

This is the part nobody puts in the product documentation. The assessment isn't wrong. It's just assessing the wrong thing until you calibrate it to your actual traffic pattern.

Get the Full Details

Disc Assessment Sample Report at Tristan Wilkin blog
Disc Assessment Sample Report at Tristan Wilkin blog

The Methodology You Should Actually Follow

Start by documenting your workload characteristics. I mean specifics: average I/O size, read-to-write ratio, queue depth per host, peak concurrent users, any periodic batch operations, and the maximum latency your application can tolerate before users notice. If you don't have historical performance data, capture at least two weeks of production metrics from your existing storage using something like iostat, perf, or your hypervisor's built-in monitoring. Next, select your testing tool. For most scenarios, fio is the standard. It's free, it's flexible, and it gives you enough control to match your workload profile within about 95% accuracy. DD is faster to set up but tells you almost nothing about random I/O performance. Smartmontools gives you drive health information but that's diagnostic, not performance-based. For a proper assessment, you're using fio or a comparable benchmarking framework with a customized job file. Set up your test environment to mirror production as closely as possible. This means testing on the same controller, same cable topology, same drive firmware revision, and ideally the same drive temperature. I've seen assessments where drives tested in a warm server room at 28 degrees Celsius performed noticeably different from identical drives in a cooled test bench, mainly because thermal throttling on some SSDs starts kicking in around those temperatures and the firmware adjusts power limits accordingly.

Run your workload profile for at least 30 minutes for rotational drives and 15 minutes for SSDs. Shorter runs miss the sustained performance degradation that happens as caches fill up or as TRIM operations begin competing with your workload. Record latency percentiles—p50, p95, p99, and p999—not just averages. The average tells you nothing about tail latency, and tail latency is what your users complain about. There's a common pitfall here that I see repeatedly. People test at their average queue depth and call it done. But real workloads aren't steady state. A database server during a commit storm might push queue depth to 64 or 128 while the drive is already near capacity. Run your assessment at multiple queue depths—at QD1, QD8, QD32, and QD128 if the drive supports it—and plot how latency scales. If latency jumps from 1 millisecond to 50 milliseconds when you go from QD8 to QD32, that drive is going to feel terrible under bursty production loads even though its QD1 numbers look healthy.

Interpreting What the Numbers Mean

One throughput number doesn't equal one disk. This is especially true with modern NVMe drives where multiple lanes and queues mean raw bandwidth isn't the bottleneck—latency consistency under load is. A drive that peaks at 3,500 MB/s sequential read might drop to 800 MB/s when mixed random reads hit at the same time, and your assessment needs to capture that interaction. For SSDs, pay attention to SLC cache behavior. Most consumer and enterprise SSDs use a portion of their NAND as a high-speed SLC cache. Write throughput inside the cache is excellent. Once the cache fills during sustained writes, performance drops to the speed of the underlying TLC or QLC NAND, which can be 5 to 10 times slower. A proper disc assessment includes a sustained write test long enough to exhaust the cache and measure the real baseline. Skipping this step is one of the most common reasons assessments look good and production fails. Another counter-intuitive point: higher IOPS doesn't always mean better performance for your workload. A drive with 200,000 random read IOPS at 4K but 15-millisecond latency under load will perform worse for a transactional database than a drive with 80,000 IOPS and 2-millisecond latency. Your application's throughput in terms of completed operations per second depends on latency as much as it depends on raw IOPS, because IOPS is really just operations per second divided by average latency.

The Disk Assessment _ What is Disc – AINZ
The Disk Assessment _ What is Disc – AINZ

The disc assessment process also needs to account for drive health over time. New SSDs often perform differently than drives that have been written to extensively. Some vendors publish a "write endurance" curve showing how performance degrades as the drive fills with data. If your assessment is a one-time snapshot and your drives are three years into their lifecycle, the numbers might not reflect what you'll actually get when you deploy at scale. I've worked with drives that showed excellent read latency when new but degraded to unusable levels after about 60% of their rated terabytes written, and the vendor's datasheet never mentioned that degradation curve.

When Disc Assessment Fails Completely

Benchmarking alone won't tell you everything. There are scenarios where synthetic tests produce meaningless results. Network-attached storage adds variable network latency that no local benchmark can replicate. A drive sitting on a NAS might show perfect numbers on a direct-attach test, but once it's on the network with 512-byte reads from 200 concurrent clients, the aggregate behavior is unpredictable without actual application-level testing. Virtualized environments add another layer of abstraction. Hypervisor scheduling, resource pools, and storage tiering can make a drive's assessed performance irrelevant because your VM isn't getting guaranteed access to the physical hardware. I ran an assessment on a solid-state array that showed 45,000 IOPS, deployed it in a VMware cluster, and the same workload peaked at about 12,000 IOPS because the storage policy and resource allocation didn't match what the physical array could deliver. The disc assessment was correct. The deployment wasn't. If you need higher confidence than benchmarking can provide, the workaround is integration testing. Run your actual application against the storage setup for a defined period under simulated production load. This takes more time—usually 1 to 2 days for a meaningful session compared to 15 to 45 minutes for a synthetic test—but it catches the interaction effects that standalone disc assessment misses entirely.

The assessment framework itself also has blind spots. It doesn't evaluate firmware bugs, drive failures, or controller issues. A drive can score perfectly on all performance metrics and still have a firmware defect that causes intermittent command timeouts under specific conditions. That's why post-assessment monitoring during the first weeks of deployment matters, and why you should track latency and error rates even after the initial test passes. Most importantly, a disc assessment is a point-in-time evaluation of a specific configuration. Change the controller, the firmware, the operating system kernel version, or the workload pattern, and the previous results no longer apply. Re-assess when any of those variables change. That's not a limitation of the method—it's just how storage performance works, and it's the part that trips up people who treat assessment as a one-time checkbox rather than an ongoing practice.

DISC Assessment: A Complete Guide - Mentorink
DISC Assessment: A Complete Guide - Mentorink