home | tags


incremental? differential? No… full!

tags: #backup #intro

When speaking of backups, you first hear of tools like tar, and then – as you dig deeper – you’ll hear the words “incremental” and “differential”. This article will try to explain the difference, and also why, today, neither of these is any use.

Generally speaking, both “incremental” and “differential” backups are used in the context of backup systems that retain older versions also, but only apply to systems that are notionally similar to tar. That is, each time you backup, you create a new “tar file” or similar.

In those systems there is the notion of doing a full backup say every Sunday, and either an incremental or a differential backup on other days. The difference between those two is, an incremental backup on Wednesday is whatever changed from Tuesday (i.e. with reference to the last backup) , but a differential backup is always whatever changed from Sunday (i.e., with reference to the last full backup).

Restoring from these backups is typically multi-step: apply the last full backup, then either (a) all the incrementals since or (b) the last differential.

These kinds of backups are very commonly used in enterprises, where backing up to tape is the norm (tapes have much higher reliability in the long term – measured in years or decades – as I understand it) Hence why I used “tar” (derived from “tape archive”, afaik) as the archetypal example.

In all these cases, the actual backups can be offline, and only an index of some kind needed to determine what’s new (or maybe just a reference timestamp).

When your backups are online, the distinction between incremental and differential disappears – typically every backup is a “full” backup, even if the actual data being processed is only what changed since the last backup. rsnapshot is a good example of this kind of backup. It uses hard links to maintain multiple versions of data in less space than all those separate copies would take.

Then there are the most modern backup systems like, say restic and borg, which are more like git in the sense that, not only is every backup a full backup, but you get the benefit of “chunk-based deduplication”. Best way to explain it is if you have a 1 GB file of which only 1 KB has changed since the last backup, restic (and similar tools) will only add 1 KB (plus some overhead) to the backup data. Neither offline backups nor the rsnapshot style backups can do that – they will make a full 1 GB copy again.

(The downside of these highly storage-efficient backup systems is data corruption… you’re totally eliminating redundancy so any data corruption will bite you real hard. Which is why you follow the “3-2-1 rule” for backups (3 copies, 2 different media, 1 offsite)).