RAID, NAS and Storage Array Data Recovery
RAID is not a backup. Member order, geometry and the incident timeline must be preserved before another rebuild or forced start changes the remaining state.
Incident containment
After a second member fails, stop the array and preserve every bay
A degraded array may still be serving stale or inconsistent data; continued writes and automatic replacement policies can consume the remaining evidence.
Power the NAS or array down through a controlled procedure if one remains available, then prevent automatic restart. Do not remove only the disk showing the newest warning. The first failed member, an apparently healthy spare and every active member may all contain different generations needed to explain the incident.
Photograph the front and rear, bay numbers, status display and cabling before removal. Label each device by enclosure, shelf, bay and serial number, never by the order in which it was taken out. Member position is evidence and should not be reconstructed later from memory or drive capacity.
Preserve controller cards, configuration exports, event logs, replacement disks and previous failed members. Record the sequence of alarms, power cuts, disk swaps, rebuilds and firmware changes. The most recent error is not necessarily the original cause, especially when an earlier member dropped out silently.
Keep failed members
A member rejected earlier may contain an older but internally consistent stripe generation. Do not dispose of it after a replacement begins rebuilding.
Keep bay identity
Serial number, slot and enclosure associations support order reconstruction and distinguish disks with identical models. Photograph labels before anti-static packing.
Rebuild risk
A rebuild can overwrite the consistency that still survives
RAID rebuild writes calculated blocks across a member set and assumes that the selected order, source disks and parity history are correct.
When a replacement disk is added, the controller reconstructs what it believes belongs there. If another member has unreadable sectors, stale data or the wrong bay assignment, the process can propagate errors across the new disk. A rebuild is a service-restoration operation, not a read-only diagnostic.
Forcing disks online or clearing foreign configuration can update metadata that identifies membership and generation. Replacing more than one member at once may create a logically plausible but chronologically mixed array. Do not initialise, resynchronise, scrub or expand the source set before every member has been protected.
An interrupted rebuild is documented by percentage, time, source and target members and error messages. The pre-rebuild failed disk and partially written replacement are both retained. Their differing stripe generations may help reconstruct priority data even though neither is a complete current member by itself.
- Do not initialise replacement disks in the original array
- Stop scrub, resync, expansion and forced-online actions
- Retain both previous and partially rebuilt members
- Record controller messages and rebuild chronology
Stale versus missing
A stale member contains older blocks, while a missing member contributes none. Treating the former as current can be more dangerous than modelling the latter as absent.
Partial rebuild
The target may contain new stripes up to one region and blank or old data elsewhere. The boundary and source generation must be inferred and recorded.
Array geometry
Member order, stripe size, parity rotation and offset define the volume
The RAID level alone is insufficient; several layouts can produce a mountable-looking result while joining the wrong blocks.
Recovery identifies active and spare members, logical order, stripe or chunk size, starting offset, parity direction and rotation. Nested RAID, vendor layouts and storage pools add layers. Controller metadata, partition alignment and repeated file-system structures provide complementary evidence.
Parity is tested rather than assumed. For known data stripes, calculated parity should match the candidate parity member and generation. A high mismatch rate can reveal wrong order, stale disks, an incorrect offset or unread regions. Checks are sampled across the address space because one region may look valid by coincidence.
A partition table appearing at sector zero does not prove that the geometry is correct. Many members can contain similar signatures or replicated metadata. Candidate configurations are compared by sustained file-system continuity, parity consistency and meaningful content before one is selected for reconstruction.
Metadata-led evidence
mdadm superblocks, controller headers, ZFS labels or vendor records can reveal UUIDs, roles and event counts. Damaged or rewritten metadata is corroborated rather than trusted alone.
Content-led evidence
File-system structures, virtual-disk headers and known files can test stripe hypotheses. Consistency across distant regions matters more than one recognisable fragment.
Member acquisition
Image each member according to its own physical condition
An array-level problem does not make all disks equally healthy, and the weakest member should not dictate an uncontrolled full-set rebuild.
Each disk is assessed for identity, capacity, interface, error behaviour and physical symptoms. Accessible members are imaged separately with their bay labels retained. Error maps show where unavailable sectors intersect stripe sets and help decide whether another member or parity can supply equivalent data.
A clicking or impact-damaged hard drive may require mechanical diagnosis and, where justified, controlled cleanroom work before imaging. An SSD member follows controller, NAND, TRIM and encryption considerations instead. Cleanroom opening is never applied generically to an array.
Acquisition order can prioritise metadata and business-critical regions when a member is deteriorating. No repair, file-system scan or parity rebuild runs against the original disks. Working images retain identifiers and logs so later virtual reconstruction can distinguish read errors from configuration errors.
- Assess and image every member independently
- Keep bay labels tied to each acquisition
- Record unread regions and device-state changes
- Use cleanroom handling only for diagnosed mechanical damage
Healthy-appearing members
They are still imaged and protected because repeated rebuild reads can expose latent faults. A normal SMART headline does not certify every sector.
Unstable members
Reading strategy adapts to resets, heat, weak sectors or mechanics. The aim is the most valuable recoverable evidence, not an immediate perfect clone at any cost.
Layered reconstruction
Reconstruct the array virtually before opening its storage pool
The member images first produce a read-only virtual block device; partitions, pools and file systems are analysed only after its geometry is supported.
Candidate arrays are assembled from working images with missing members modelled explicitly. Parity can reconstruct missing data only within the fault tolerance of the recorded RAID level: typically one missing contributor for RAID 5 and up to two for RAID 6, provided the required surviving blocks, parity blocks and geometry are coherent. Stripes that exceed that limit or contain stale contributors remain unresolved and are mapped into the virtual result.
Above RAID can sit LVM, Storage Spaces, ZFS, Btrfs, vendor pools, encryption and snapshots. Each layer has identifiers and allocation metadata that must agree. Mounting read-write or importing a pool with force options can replay journals, select a transaction group or update labels.
A virtual mount is performed read-only or on a disposable clone. Candidate layouts remain separate until pool and file-system continuity support one result. The data recovery process keeps this reconstruction repeatable and away from the original member set.
Missing-member modelling
Unavailable stripes are represented honestly. Parity reconstruction is logged, and regions with too many missing or inconsistent contributors remain gaps.
Pool and file system
Volume groups, datasets, shares and snapshots are identified above the array. Repair or journal replay is tested only after a protected baseline exists.
Business triage
Shares, virtual machines, backups and databases need different priorities
An array can contain several workloads whose value, consistency rules and acceptable recovery points are not interchangeable.
File shares may be validated by folder, permissions and representative documents. A virtual-disk recovery case must preserve descriptor and snapshot chains. Backup repositories depend on catalogues and chunk sets, while databases need aligned data and transaction logs.
The client defines critical systems, required dates, acceptable older versions and which services can be rebuilt from elsewhere. This can direct acquisition and reconstruction when unread regions make a complete result impossible. Recovering a secondary archive should not consume the only stable reads before current payroll or production data.
Confidentiality and authorisation are recorded for each workload. A mounted share is not proof that a database is transactionally sound or a backup can restore. Application owners may need to run checks on isolated copies, with external licences and identity dependencies documented separately.
- Rank workloads by business impact and replaceability
- Keep virtual-disk chains and database logs together
- Identify acceptable older recovery points
- Nominate authorised application owners for validation
Data priority
List essential shares, VMs, database names, backup generations and dates. Identify confirmed alternative copies so effort focuses on material with no viable replacement.
Recovery point
Define whether the requirement is the latest possible state or a known consistent earlier point. Stale members and snapshots may support one but not the other.
Outcome evidence
A rebuilt volume still needs parity, file and application validation
The result must connect member acquisition and geometry decisions to usable priority data, not stop at a mounted capacity.
Parity and stripe consistency are sampled across the reconstructed address space. File-system checks then identify allocation damage, stale metadata and unread regions. Representative files are opened or structurally tested across different directories and sizes so validation is not confined to easily accessible early blocks.
Virtual machines, archives and databases receive workload-specific checks on isolated copies. Permissions, timestamps and folder structure are preserved where possible. Recovered data is delivered on healthy storage, not written back to the source array, and the original members remain controlled during client review.
The report distinguishes verified, partial, stale and unavailable material. It records member errors, geometry, parity assumptions and application checks. A green NAS dashboard, complete-looking share tree or headline terabyte count is never substituted for evidence that the required content is usable.
Technical validation
Document member images, unread maps, selected order, stripe and parity behaviour. Alternate layouts or stale-member choices remain traceable.
Business validation
Review agreed shares, dates, database operations or VM content with authorised owners. State what was tested and what still requires restoration work.
Countrywide service
Prepare the bay map, logs and configuration for UK intake
Datastrophe supports RAID and NAS cases across the United Kingdom through arranged media intake. The initial review confirms the full member set, controller context and assigned handling route.
Photograph every enclosure, shelf, bay and cable before changing the system. Export configuration and logs only if this does not start a rebuild or clear foreign metadata. Provide the make, model, controller, RAID level if known, member count, capacities, file-system or pool type and incident chronology.
Package each disk separately in anti-static protection and rigid cushioning, labelled by enclosure and bay. Keep controllers, expansion units, failed disks, spares and partially rebuilt replacements available. Do not stack bare drives or attach adhesive over breather holes and labels.
Use Request a quote to supply the topology and priorities. Pricing information explains why member condition, geometry, pool layers and workload validation affect scope. The destination, transport plan and required enclosure, controller and member set are confirmed for each case.
- Freeze writes, rebuilds, scrubs and forced imports
- Photograph and label every member by original bay
- Retain controllers, failed disks and replacement members
- Confirm complete hardware and dispatch requirements
Topology package
Include bay map, serials, controller and firmware details, configuration exports, event logs and every removed member. Mark uncertainty rather than guessing an order.
Priority package
List shares, VMs, databases, backup dates and acceptable recovery points. Name authorised technical and data reviewers for confidential validation.
FAQ
Frequently asked questions
Should I replace another failed RAID disk and rebuild?
Not before every original and replacement member has been labelled, assessed and imaged. A rebuild writes calculated blocks and assumes the chosen source set is current and readable. Stale data or another weak member can propagate corruption. Preserve the bay map, controller logs, earlier failed disk and any partial rebuild target so candidate generations can be modelled on working copies.
Can RAID member order be determined without the enclosure?
Sometimes, using member metadata, event counts, partition alignment, parity tests and file-system continuity, but the original bay map greatly improves confidence. Do not insert disks into random bays to test an order because controllers may update metadata or start a rebuild. Preserve serial numbers, configuration exports, photographs, controllers and logs; uncertainty in the order should remain explicit.
Does parity guarantee recovery when a RAID disk is missing?
Parity can reconstruct missing data only within the configured RAID level's redundancy: typically one missing contributor in RAID 5 and up to two in RAID 6, when the required surviving blocks, parity and geometry are coherent. More missing or unreadable contributors, stale members and an interrupted rebuild can exceed that protection. Recovery maps errors by stripe, tests parity consistency and records regions that cannot be reconstructed. RAID reduces availability risk; it is not an independent backup.
When does a RAID member need cleanroom handling?
Only when an individual mechanical hard drive has diagnosed internal damage involving heads or platters. The array itself is reconstructed logically from member images. SSD members follow electronic and flash methods, and healthy hard disks do not need opening. Cleanroom work can create a controlled reading opportunity for one damaged member but cannot guarantee that its platter surface or every stripe survives.
Is a mounted NAS share proof of successful recovery?
No. A wrong geometry or stale member can preserve directories while corrupting later blocks, large files, databases or virtual disks. Validation samples parity and file-system consistency, opens priority files and tests workloads such as VMs, archives or databases in isolation. The result should report geometry, member gaps, stale material and application checks rather than relying on mounted capacity alone.
Media
Other expertise
Diagnostic assessment
Unsure about a storage device or fault?
Datastrophe qualifies the risk before any recovery attempt and points you towards the safest next step.