How to check software RAID and drive health on your server
Contents
For administrators of a dedicated server with two or more drives. In half an hour you will check the software RAID and the drives themselves, set up failure alerts and know what to do when a drive fails.
What you will need#
- SSH login as a sudo user (run the commands below as root, after
sudo -i) or Windows Server administrator rights. - Debian 13, Ubuntu 24.04 or 26.04 — package and service names were checked for these releases.
- A fresh backup kept off the server — before you touch partitions or the array.
apt update
apt install smartmontools nvme-cli gdisk dosfstoolsHow the drives are laid out#
With two or more drives, the system is installed by default on all drives with software RAID 1 (a mirror). Typical Linux layout (UEFI): an EFI partition (ESP) on each drive, outside RAID, only one mounted as /boot/efi; /boot in RAID 1, usually /dev/md2; the root / in RAID 1, usually /dev/md3; swap as a separate partition on each drive, outside RAID. On Debian 13, Proxmox VE 9 and Ubuntu 26.04 the EFI partition is mirrored too.
/dev/sda, /dev/sdb and the partition numbers below are examples; NVMe drives are named like /dev/nvme0n1, their partitions like /dev/nvme0n1p2.
Step 1 Array state#
cat /proc/mdstat
lsblk -o NAME,SIZE,TYPE,FSTYPE,MOUNTPOINTS
mdadm --detail /dev/md3In /proc/mdstat, check the square brackets next to each array: [2/2] [UU] — both drives are in place, [2/1] [U_] — the array is degraded and runs on one drive.
In the mdadm --detail output, look at the State line (clean or active is normal, degraded means no redundancy), the Failed Devices counter and each partition’s state at the bottom: active sync, faulty, removed.
Step 2 Drive health#
lsblk -d -o NAME,MODEL,SERIAL,SIZE,ROTA,TRAN
smartctl -H /dev/sda
smartctl -a /dev/sda
smartctl -a /dev/nvme0
nvme smart-log /dev/nvme0smartctl -H shows how the drive rates its own condition. FAILED means replace the drive at once, but PASSED promises nothing: drives often fail with that verdict too. Check the individual indicators.
| Indicator | Warning sign |
|---|---|
SATA: 5 Reallocated_Sector_Ct | RAW_VALUE above zero and growing |
SATA: 197 Current_Pending_Sector, 198 Offline_Uncorrectable | any non-zero value |
SATA SSD: wear indicator, for example Wear_Leveling_Count (name varies by vendor) | VALUE approaches THRESH |
NVMe: Critical Warning | anything other than 0x00 |
NVMe: Available Spare | drops to Available Spare Threshold |
NVMe: Percentage Used | approaches 100% |
NVMe: Media and Data Integrity Errors | above zero |
Step 3 Alerts and periodic array checks#
In /etc/mdadm/mdadm.conf, MAILADDR is set to root by default. Replace it with an address that is actually read, restart monitoring and send a test message:
systemctl restart mdmonitor.service
mdadm --monitor --scan --oneshot --testmdadm sends mail through the local sendmail, so a service that relays mail outside (such as Postfix) is needed. Otherwise alerts stay on the server and you learn of a degraded array only when the second drive fails.
mdmonitor.service emails on failure, mdmonitor-oneshot.timer sends a daily reminder about a degraded array, and mdcheck_start.timer and mdcheck_continue.timer check the arrays monthly, reading the drives in full to find unreadable sectors early. After a check, cat /sys/block/md3/md/mismatch_cnt normally shows 0.
smartd from the smartmontools package watches the drives themselves. In /etc/smartd.conf, replace root after -m in the DEVICESCAN line with your address and run systemctl restart smartmontools.service. To test, temporarily add -M test to the end of the line — a test message is sent after the restart (needs the mail command from bsd-mailx or mailutils).
What to do when the array is degraded#
- Do not reboot the server or “fix” anything at random: the array works on one drive but will not survive a second failure.
- Make a fresh backup off the server — see the 3-2-1 rule.
- Find the failed drive and write down the serial numbers of all drives:
If the failed drive is no longer visible, the healthy drives’ serial numbers help identify it.
mdadm --detail /dev/md3 lsblk -d -o NAME,MODEL,SERIAL,SIZE smartctl -i /dev/sdb nvme list - Contact technical support (available 24/7): phone +38 044 206 08 08, email info@united.net.ua. Data centre staff replace the drive. In the request, give the server address, the serial numbers of the failed drive and all others, and the output of
cat /proc/mdstatandsmartctl -a. Confirm that the data is copied and you understand the risk of losing it, and name a convenient time — replacements are done around the clock.
On servers without hardware RAID, the drive is replaced with the server powered off, so plan for downtime; hot swapping is possible only on servers with a RAID controller. Before the replacement, take any remaining partitions of the failed drive out of the arrays.
Warning. Check the partition name twice: if you mark a partition of the only healthy drive as failed, the array is left without a working copy of the data.
mdadm --manage /dev/md3 --fail /dev/sdb3
mdadm --manage /dev/md3 --remove /dev/sdb3After the replacement: returning the drive to the array#
You rebuild the software RAID yourself. In the example, /dev/sda is the healthy drive with the data and /dev/sdb is the new, empty one.
Warning. Drive names may swap after the replacement. Before every command, check by serial number which drive is which: the new one has no partitions and belongs to no array.
lsblk -o NAME,SIZE,TYPE,FSTYPE,SERIAL,MOUNTPOINTS
cat /proc/mdstat1. Copy the partition table#
Warning. In the
sgdisk --replicatecommand, the drive after the “=” sign is the destination (the new drive; its table will be overwritten), and the last argument is the source (the healthy drive). If you mix them up, the empty table wipes the healthy drive’s table; then restore it at once, without rebooting:sgdisk --load-backup=/root/sda-gpt.bak /dev/sda. Likewise, use--randomize-guidson the new drive only.
sgdisk --backup=/root/sda-gpt.bak /dev/sda
sgdisk --replicate=/dev/sdb /dev/sda
sgdisk --randomize-guids /dev/sdb
sgdisk --print /dev/sdbThe commands assume a GPT table (lsblk -d -o NAME,PTTYPE shows gpt); for MBR (dos), use sfdisk -d /dev/sda | sfdisk /dev/sdb, source on the left.
2. Add the partitions to the arrays#
mdadm --detail shows which partition belongs to which array: if /dev/sda2 is active in /dev/md2, add /dev/sdb2. If the EFI partition is mirrored too (Debian 13, Proxmox VE 9, Ubuntu 26.04), add the new drive’s EFI partition to its array too.
mdadm --manage /dev/md2 --add /dev/sdb2
mdadm --manage /dev/md3 --add /dev/sdb3
watch -n 5 cat /proc/mdstatThe recovery line shows the percentage and estimated time; on large HDDs a resync takes hours. The server keeps working, only slower; avoid rebooting until it finishes.
3. EFI partition and bootloader#
Without a bootloader on the new drive, the server will not boot when the old drive fails. First check the boot mode:
[ -d /sys/firmware/efi ] && echo UEFI || echo BIOSBIOS. Run grub-install /dev/sdb, then dpkg-reconfigure grub-pc and tick both drives so GRUB updates reach each of them.
UEFI. The bootloader lives on the ESP (FAT32). If the ESP is mirrored (lsblk shows a raid1 array under it), the previous step has already returned it. If each drive has its own ESP, the one on the new drive is empty: create a file system and copy the contents from the healthy drive. If /etc/fstab mounts /boot/efi by label (LABEL=), give the new ESP the same label with the -n option of mkfs.vfat.
Warning.
mkfs.vfatdestroys the partition’s contents. Make sure/dev/sdb1is the new drive’s ESP andfindmntshows/boot/efifrom the healthy drive. If it is not mounted, mount the healthy drive’s ESP; if/etc/fstabrefers to it by UUID, correct the UUID (blkidshows it), otherwise the next boot may stop in emergency mode.
findmnt /boot/efi
mkfs.vfat -F 32 /dev/sdb1
mkdir -p /mnt/esp-new
mount /dev/sdb1 /mnt/esp-new
cp -r /boot/efi/. /mnt/esp-new/
umount /mnt/esp-new
efibootmgr -vefibootmgr -v shows the UEFI boot entries; ask support whether the new drive needs an entry and whether the boot order may be changed. Repeat the copy after GRUB updates.
4. Create swap on the new drive#
Swap is not in RAID: create it on the new drive’s partition with the same number as swap on the healthy drive. If /etc/fstab refers to swap by UUID, replace the old drive’s UUID there with the one from blkid before swapon -a.
mkswap /dev/sdb4
blkid -s UUID -o value /dev/sdb4
swapon -a
swapon --showIf the server does not boot#
Ask support for rescue mode: the server is rebooted into a temporary Linux system with SSH access (the password is sent by email, or your public key is used). There you can mount the drives, chroot, repair the bootloader and copy data. To leave it, contact support again to reboot the server from its disk. There is a separate rescue mode for Windows.
Hardware RAID#
Some models have a hardware RAID controller (the configuration card says “hardware RAID”). The system then sees a single logical drive and /proc/mdstat lists no arrays. Check the array and drive state with the controller’s utility, such as storcli or MegaCLI. After a drive replacement, such an array rebuilds itself.
Windows Server#
On Windows Server the mirror is built by Windows itself on dynamic disks. Run PowerShell as administrator to check the drives:
Get-PhysicalDisk | Format-Table DeviceId, FriendlyName, SerialNumber, HealthStatus, OperationalStatusHealthStatus should be Healthy; Warning or Unhealthy is a reason to contact support. The list disk and list volume commands in diskpart show the mirror state: Healthy is normal, Failed Rd means redundancy is lost, Rebuild means a resync is in progress. On failure: first a backup, then a support request with all drives’ serial numbers.
How to check the result#
- In
/proc/mdstat, every array shows[UU]. smartctl -HreportsPASSEDfor each drive, and the indicators from the table are normal.- The test messages from
mdadmandsmartdhave arrived. - After a drive replacement,
swapon --showlists swap on both drives and the server boots without errors. Check with a planned reboot at a convenient time; ask support for the remote console beforehand (the Lite line may not have one) and give your public IP address: the link works only from it, for a limited time.
Common mistakes#
- RAID is treated as a backup. RAID only protects against drive failure: a deleted file, a failed update or ransomware hits both drives at once. Backups are needed separately — see the 3-2-1 rule.
What next#
- Set up backups: restic for Linux, built-in Windows Server tools, databases. The main lines include backup space on separate storage; you set up the copying yourself.