At our recent Bareos Expert Circle, Bruno Friess from eXstor GmbH gave a presentation about tape metadata, tape health and what happens inside an LTO environment beyond the backup data itself.
A short look back
Magnetic tape has been used for data storage for decades. The LTO format itself was introduced in 2000 as an open tape standard that could work across products from different vendors. The first LTO-1 cartridge stored 100 GB of uncompressed data. LTO
Today, LTO is already in its 10th generation. Current LTO-10 media offers 30 TB or 40 TB of native capacity depending on the cartridge type. So the technology has changed a lot over the last 25 years but tape is still widely used for backup and long-term storage. LTO
And modern tape stores much more information than just the backup data itself.
A successful backup doesn’t always mean everything was fine
When a backup application reports a successful tape job, we normally assume that everything worked as expected. But a tape drive can correct many problems by itself.
For example, the drive may have trouble positioning the tape correctly. It stops, repositions the head and continues writing. The backup still succeeds so the application may never report an error. If this happens thousands of times, however, you have two problems.
The backup gets slower and the tape or drive may already be showing signs that something is wrong.
So instead of only asking: Did the backup succeed?
it can be useful to ask: What happened while the backup was running?
Your tape knows more than the backup application
An LTO cartridge contains metadata about its usage and condition. The drive keeps its own statistics as well. This can include information about:
- read and write errors
- number of mounts
- data read and written
- positioning problems
- temperature events
- cartridge and drive lifetime
Because the information comes from both the cartridge and the drive, it can also help with troubleshooting. If one cartridge causes problems in different drives, the cartridge becomes suspicious. If several tapes show problems in the same drive, the drive may be the issue. That gives you much more information than a simple successful or failed job status.
Some problems stay hidden
Tape applications receive TapeAlert messages when certain problems become serious enough to report.
But not every corrected error reaches the backup software.
One example from the Expert Circle showed around 34,000 corrected write errors while the backup application itself had not reported a failure.
The tape was still working. But would you want to keep using the same cartridge for important backups for the next few years?
Tape health can also explain poor performance
When a tape backup becomes slow, it is easy to blame the tape technology itself.
But repeated corrections and repositioning can also reduce performance considerably.
The job may still complete successfully while the drive spends a lot of time stopping and repositioning the tape.
Looking at the underlying metadata can therefore help explain why one cartridge or drive performs differently from another.
Collecting the information before it disappears
Some useful information is stored permanently with the cartridge or drive.
Other information exists only for the current tape session.
For example, session data can contain information about compression, media speed and what happened during the last operation. This information needs to be collected when the cartridge is unmounted because it can disappear with the next mount.
That makes automatic collection useful if you want to analyse tape health over time.
ELMM and Bareos
This is one of the ideas behind eXstor Library & Media Manager (ELMM).
ELMM collects information from tape cartridges, drives and individual sessions and stores it in a database. The data can then be visualized, for example in Grafana, to show cartridge health, drive statistics, errors and performance.
eXstor and Bareos are also working on an integration between ELMM and Bareos.
Besides monitoring, ELMM can allow different backup applications to share the same tape infrastructure while keeping their media separate. This can be useful during migrations from another backup solution to Bareos.
Find the problem before the tape fails
The main takeaway is simple.
A backup application tells you whether the job succeeded.
The tape infrastructure can tell you how healthy that successful job actually was.
Corrected errors, positioning problems and lifetime information can give you warning signs before a cartridge or drive finally produces a hard error.
For tape environments that keep important backups for years, seeing these signs earlier can make a real difference.