Deleting a file does not always remove its contents right away. A raw disk image may still contain deleted file entries and data in space the filesystem has marked for reuse.
The tools in this guide examine three layers of a disk image: files that still exist, deleted files the filesystem can identify, and raw data left in unallocated space. Each layer requires a different recovery method and a little more validation.
Unallocated space is disk space the filesystem has marked as available for use. It may still contain data from deleted files, which recovery tools can sometimes retrieve if that data has not been overwritten or cleared by the storage device.
Tools at a Glance
GNU tools
| Tool | How it is used |
|---|---|
| GNU ddrescue | Copies a drive into a disk-image file and records unreadable areas |
| GNU Coreutils | Provides sha256sum for checksums, du for directory sizes, and tr for processing text |
| GNU Findutils | Provides find for locating files and directories |
| GNU Grep | Searches configuration files, logs, and recovered text |
| GNU tar | Packages recovered files into a compressed archive |
Other tools
| Tool | How it is used |
|---|---|
| UTM | Runs a Linux virtual machine that can inspect the image |
| LVM2 | Finds logical volumes stored inside Linux partitions |
| util-linux | Provides lsblk, blockdev, findmnt, and mount |
| The Sleuth Kit | Examines filesystem records and recovers deleted data |
| Docker storage documentation | Explains where containers keep persistent application data |
Create and verify the disk image.
A raw disk image is a byte-for-byte copy of a drive. GNU ddrescue creates that image and keeps a map file showing which areas were copied and which could not be read. If the process stops, the map file lets it resume without starting over.
Check the source and destination before running ddrescue. Reversing them could overwrite the drive you are trying to preserve.
After creating the image, calculate a SHA-256 checksum:
sha256sum master.raw
A checksum is a digital fingerprint calculated from a file’s contents. Calculate it again after copying or uploading the image. Matching results confirm that the file did not change during the transfer.
On macOS, use:
shasum -a 256 master.raw
Keep the original drive and master image unchanged. Save recovered files, logs, and reports somewhere else.
Attach the image to a Linux virtual machine.
A Linux virtual machine provides an isolated place to inspect the image. UTM is one option, though other virtualization tools can also attach an image as a read-only data disk.
Once the VM starts, list the disks it can see:
lsblk -o NAME,SIZE,TYPE,FSTYPE,RO,MOUNTPOINTS,MODEL
The output shows the size, filesystem, mount location, and read-only status of each disk. An RO value of 1 means the device is read-only.
Confirm that status directly:
sudo blockdev --getro /dev/sdX
Replace /dev/sdX with the device assigned to the image.
Identify the correct filesystem.
Some Linux systems use Logical Volume Manager, or LVM. LVM creates logical storage devices inside physical partitions, so the main filesystem may be inside a logical volume.
Use these commands to display the LVM structure:
sudo pvs --readonly
sudo vgs --readonly
sudo lvs --readonly -a
If the VM and disk image use similar volume names, check which physical device backs each logical volume:
sudo lvs --readonly -a \
-o lv_path,vg_name,lv_size,devices
Once you find the correct Ext4 filesystem, mount it:
sudo mkdir -p /mnt/image
sudo mount -t ext4 -o ro,noload \
/dev/example-vg/example-lv \
/mnt/image
Ext4 is a common Linux filesystem. Its journal records pending changes so Linux can finish them after an unexpected shutdown. The ro option mounts the filesystem as read-only. The noload option prevents the journal from replaying pending changes.
The Linux Ext4 documentation explains this behavior.
Verify the mount:
findmnt -no SOURCE,TARGET,FSTYPE,OPTIONS /mnt/image
The output should include ro. Some systems display noload as norecovery.
Inspect allocated files.
Start with files that are still allocated. An allocated file still has filesystem records connecting its name and path to the blocks containing its data.
Common places to check include:
/opt
/srv
/home
/usr/local
/var/lib
Use find to locate files, du to check directory sizes, and grep to search configuration and logs.
Include hidden files when searching for source code. A .git directory is part of a Git repository and may contain commits, branches, and files that no longer appear in the working directory.
For Docker-based applications, check the Compose files and these directories:
/var/lib/docker/containers
/var/lib/docker/volumes
/var/lib/containerd
A Docker volume stores data outside a container so it can remain after the container stops or is replaced. Container configuration can point to databases, logs, and other persistent files elsewhere on the disk.
Logs can show whether services started, connected, or failed to write data. Keep their timestamps so you can see when a problem occurred and whether it repeated.
Recover deleted files.
Unmount the filesystem before examining deleted data:
sudo umount /mnt/image
The Sleuth Kit provides several tools for this step:
fsstatdisplays information about the filesystem and its storage blocks.flslists file entries, including deleted entries.tsk_recoverextracts files whose filesystem records still survive.blklsreads unallocated blocks.
Filesystem records, also called metadata, include filenames, timestamps, permissions, and pointers to the blocks that hold a file’s contents.
Inspect the filesystem and list deleted entries:
sudo fsstat /dev/example-vg/example-lv
sudo fls -r -d /dev/example-vg/example-lv \
> deleted-files.txt
Then recover the available files:
mkdir -p recovered-files
sudo tsk_recover \
/dev/example-vg/example-lv \
recovered-files
Use a native Linux filesystem for the initial output when possible. Some shared drives cannot store every Linux filename or permission. You can package the recovered directory and transfer it afterward.
Review the recovery log for write errors. A recovered filename does not guarantee that every part of the file survived because some of its blocks may have been overwritten.
Search unallocated space.
Some deleted data remains after its filename and filesystem records are gone. The Sleuth Kit’s blkls command can read those unallocated blocks.
Because the output may be large, stream it through a search:
sudo blkls /dev/example-vg/example-lv |
LC_ALL=C tr -cs '[:print:]' '\n' |
grep -Eai 'database|application|measurement' \
> unallocated-keyword-hits.txt
Choose keywords related to the files or data you are looking for. A match may come from configuration, logs, program files, or deleted data, so it provides a place to investigate rather than a final answer.
You can also search for binary signatures. A binary signature is a known sequence of bytes that helps identify a file format. Finding one may reveal where a deleted file began.
A matching signature does not confirm that the entire file survived. Check the candidate’s internal structure and use the original application’s inspection tools when available. Save incomplete fragments, but label them clearly.
Package and verify recovered files.
A .tar.gz archive combines a directory into one compressed file while retaining its internal paths and Linux metadata:
sudo tar -czf recovered-files.tar.gz recovered-files
Create a checksum for the archive:
sha256sum recovered-files.tar.gz \
> recovered-exports.sha256
After transferring the files, verify the checksum:
sha256sum -c recovered-exports.sha256
A recovered filename, a file fragment, and a complete file are different results. Check the internal structure of each candidate and use the original application’s inspection tools when they are available.
Keep the disk image unchanged, save the recovery logs, and verify exported files with checksums. Those records give another person enough context to understand what was recovered and how.