Creating backup storage sucks(smarmelling.com) |
Creating backup storage sucks(smarmelling.com) |
Sure it's bothersome to use the sneakernet once a month (manually copy data to USB stick and stash it somewhere), but when that ransomware hits (can happen on personal devices also), it literally can't access the offline copy, and try to encrypt that one also.
Immutable backup solutions are code and thus eventually defeatable from the remote attackers point of view. Offline requires physical access, and maybe a big enough wrench to get the USB stick's disk encryption PIN out of you.
Surrendering just a little bit of absolute control over the physical media can dramatically increase the likelihood that it survives whatever apocalyptic event. I don't know that I've ever heard of tapes being stolen from Iron Mountain. Even if this happened, what are the odds that they would get all the tapes. Your tapes? It probably looks like the warehouse from Raiders of the Lost Ark in there.
For small datasets, a cloud backup (think s3 or azure blob) with credentials that can't be harvested automatically by a malware (eg custom backup script with encrypted credentials - claude will happily write one for you in seconds) is as good as offline. For small datasets (code base, important documents, even photos if you don't go crazy - or perhaps backup a lower resolution as a dooms day last resort thing), this is nearly free.
For larger datasets, you can buy some cheap X11SSH-LN4F or X11SSL-F motherboards on ebay with RAM and CPU for ~$100. These can be remotely switched on and off programmatically with IPMI (same thing, custom scripts with encrypted credentials - claude is your friend). Have your NAS perform an incremental backup once a week or once a month and keep it off the rest of the time (or trigger it from a raspberry pi with own credentials if you don't want to connect IPMI to your LAN or have encrypted IPMI credentials on your NAS). And unless a malware hits right at the time of the backup, it is as good as an offline backup while also being automated. Doesn't protect from a power surge though, which may or may not be a problem depending on where you live.
Also have your backup pull data from your NAS rather than the other way round, and run with different credentials than your NAS, so a malware can't jump from the NAS to the backup (or encrypt the backup). I have seen first hand that if you reuse admin credentials between machines, one machine compromised means all machines compromised within minutes.
Well overdue for a refresh.
I also have a set I curate and put in with the family & household papers. That doesn't have any of my personal crap, just documents, photos, and family history etc.
I have two of these and rotate them weekly and keep one at my work (and I have a self-hosted healthchecks.io instance to remind me if I forget).
An extra protection against the possibility of malicious access such as ransom attacks is a two-step “soft-offline” backup.
The source machines backup to a central place, and the backup machines pull copies from there. The important part is that the source machines can not connect (or at least can not authenticate against) the final backups and vice-versa, so a malicious process/person getting into one can not affect the other and vice versa. Obviously if I don't notice the damage immediately then the most recent backups may be corrupted because damaged data was pulled, but past snapshots (taken after each pull at the backup side) will still be clean.
You have to be very careful about storing credentials to make sure source credentials don't ever end up on the backup machines and backup credentials don't ever end up on the source machines, even well out of the way of normal places like ~/.ssh, because a targetted attack might find them, but it gives almost the assurance of an offline backup while still being fully automatable. My backup site credentials are in my true offline backups, I need to refer to them for maintenance access, and then they get cycled after that use.
Restores can be mediated the same way. Verifying backups can be done by both sides running hashes on the files and posting the list back to the central machine for comparison - anything that differs without having a timestamp after the previous snapshot is likely corruption.
As a side note about corruption: if doing snapshots in the filesystem (“cp -al after rsync” or one of the many similar options) make sure you have more than one snapshot chains (on separate storage if you have resource for that). If you have a file that hasn't changed in years so every one of your snapshots points to that one version, it could only take one random filesystem error to completely lose the file.
Yeah, but an attacker is unlikely to compromise both you and your storage provider at the same point. Many S3-compatible storage providers (e.g. Backblaze and Hetzner) support object lock, where you can lock objects for a certain number of days (and refresh locks if the objects are still used in recent backups). Typically object locks are implemented such that not even you yourself can remove the data from the account settings.
E.g. when I cancelled my Backblaze B2 account, I had to wait until the object locks expired before I could remove the remaining data and delete the account.
Arq on macOS has great support for object locks BTW.
At the end of the day it depends on what is your threat model. Mine is 1) automated malwares and 2) my own fuckups. I am not trying to prevent the NSA from hacking me. Against an automated malware, custom scripts with encrypted credentials that don't show in clear in command lines or environment variables are probably good enough.
Firstly don't try and run a corporate level backup solution or NAS with terabytes of storage and tape at home. You're going to drop dead one day and everyone will hate you for what you leave them to deal with. They will have no idea how to pick through the remains to get important stuff out. It took me a whole 3 days, as someone who knows how this works, to get into a dead relative's NAS appliance and find any mention of any info to get into his EV which had basically bricked itself in the driveway while he was in hospital.
Secondly, delete everything. Not joking. Delete as much as you can. Everything you've watched, delete it. All the old emails, delete them. Tracks you don't like on CDs you've ripped, delete them. Old software ISOs, delete them. Thousands of photos of the same thing taken 1 second apart, delete them. Thousands of PDFs of books you will never read, delete them. I went down from a whopping 8TB to 300Gb doing that and it's still shrinking monthly. Eventually it'll fit on one mediocre computer and that makes life so so so much easier.
Then look at your backup strategy. It'll look simple then. Mine is three external disks.
1. 1TB T7 shield (not encrypted) which lives at my partner's house and gets an rsync once a month.
2. 1TB T7 shield (encrypted) which lives in my bag and gets an rdiff-backup once a week.
3. 2TB Lacie HDD (not encrypted) which lives in my fire safe at home and gets an rdiff-backup once a week.
If I drop dead, my partner can just plug the thing into her computer and get stuff off it.
I'd avoid the cloud if possible as well. One of the worst situations I've seen is someone who confidently pushed their NAS contents to S3 and then when the NAS blew up they had to pay a lot of money in transfer to get it back again. Hundreds of dollars in fact. On top of the price of a new NAS. That might be a last resort option but it should NEVER be the first line backup.
Don't make your life any more complicated than it needs to be. I implore you.
Doesn't that depend entirely on the storage format? For example if you use thunderbird as a client it has database files that you can copy. Since you say that one of your devices is running arch you should check out some of the email options in the repos. There are at least a few of options that can act as a middleman to bridge between your external accounts and your clients.
> So having a plan to roll back to a stable state might be helfpul. I installed and set up Timeshift for this
Snapshots make this trivial. Particularly in the case of btrfs if you configure your bootloader to always use the default subvolume then rollback becomes just `btrfs subvolme set-default`.
However it's better to just delete most stuff so you don't have to back it up. Think my mailbox is about 50Mb.
What's simple because it's "hard" is replacing parts 2 & 3 with a network appliance like TrueNAS running a zfs pool that syncs to backblaze every night. Yeah you have to learn a bit but it won't fall in weird ways like the hard drive part here will just fail to mount one night and not back things up for 3 months until you notice. My 2¢
I use the wonderful https://github.com/garethgeorge/backrest as a web UI around restic.
Every night, I back the data up to a Hetzner storage box https://www.hetzner.com/storage/storage-box/ which is only ~$3.50/mo USD for 1TB of data.
I have three "tiers" of data for myself:
1. Important things I care about: these go into a folder that gets backed up to Hetzner (examples include photos in Immich, code pushed to Forgejo, etc)
2. Smaller files I sort-of care about: these usually just go into Proton Drive (examples include design documents)
3. Large files I don't care about losing: these go into a folder on a hard drive that is not backed up (examples include media)
This took me about a day to set up and I just check in on it every once in a while. Once a year I will do a practice disaster recovery run.
Then you only need to backup the NAS.
git push nasZFS send raw encrypted stream from laptop to a server automatically. Encrypted compressed incremental. No need to verify: if snapshot exists at destination, there is no error. If a problem is encountered, restore from redundancy. Errors can be detected and corrected.
Restic has also been working great. The problem is that, it has no redundancy to correct errors. The assumption is that, server holding repository will have redundancy, but then I can replicate directly to that server. I worry that at some point, there will be an error in repository. I may loose 5 years of snapshots (restic has some functionality to remove involved snapshots and rebuild the index, but may not succeed).
Kopia has error correction, with Reed Solomon.
Borg2 has perpetually remained in beta. Looking forward to test it.
I stay away from sync: problems with permissions, git repo, conflict, incomplete sync status in syncthing etc.
- Live data lives in the cloud. I could self host it, but it would always be inferior to the cloud offerings, and usually more expensive for my ~3TB data.
- I make local backups nightly
- I make remote backups nightly to another cloud.
- Once every week I make a local backup to a device that is offline 6/7 days a week (an old Synology NAS that powers on automatically, and powers off when it's been idle for 30 minutes).
- Every year I curate our photo library and burn a set of M-Disc Blu-Ray copies of all photos created or modified in the past 12 months. Two identical sets, one stored at home, the other stored at a remote location, each clearly labeled "Photo Backup <year>".
- Every year I also update a couple of external HDDs with the entire photo library, again identical copies, contents are verified yearly and updated, and stored again. Disks are also clearly labeled as "Photo backup".
As others have written, curate your data. In my case we have a 2.5TB photo library spanning a couple of decades, but that could easily have been 4TB without curation. I only backup documents in the nightly/weekly backups. (Personal) Documents usually only hold value for a short time, and after that it's mostly sentimental.
Anything media, or anything downloaded from the internet like books, music, etc, regardless of if I purchased it or pirated it, is not getting backed up. If it came from the internet, there's a good chance it can still be found on the internet.
I also (mostly) don't run RAID. RAID is for availability, and since all my important data lives in the cloud, and I will likely survive if one of my backups dies, there's little reason to run raid. The only exception is the share where PhotoSync backs up our photos, which is on a "small" RAID1 volume mainly because it acts as the source of all the other photo backups, so consistency and correctness is important.
Synctrain is a great iOS syncthing client that connects to your swarm but doesn't actually sync any files (unless you ask it to). You can still browse the contents of your shared folders and selectively download individual files.
Ntfs equivalent of Chmod/chown is a "go have a long lunch" type of operation.
That's a core Dropbox-level mistake of putting the backup cart before the main use horse - your files should be sorted according to their primary, not backup use
> Syncthing works very well, its only constraint is that it does need the devices to be often on if you want consistent syncing.
So it's working very poorly, this is a very inconvenient constraint
> phone is a whole other beast. It does not have enough storage to
Indeed
Overall, that fits the title perfectly, and is a bad way to setup your backups. Special needs need special apps.
Just hours ago on another post I had commented
> discard/reject/delete: ≈400GB → 25GB → 3GB (< a decade ago; I remember) (I was shocked to see how little of that really mattered).
This is one of the most important aspect of managing your personal data in any meaningfully sane and sustainable (!) way.
Then being deliberate about one's digital footprint.
I personally have no problem with them going through my stuff. Not the "I have nothing to hide" defence. I just don't give a crap.
With respect to my personal situation, my father took thousands of photos of me, my children etc back as far as 2002. Sent me a selection low res jpegs of them by email. I was thrilled when I got the original RAWs from his NAS in the end. There was so much stuff in there I had forgotten that he hadn't sent.
Everything else will be inaccessible if I'm not around. This also means less faf for them: all the info they need to care about is in that package, they don't need to look at the huge disorganised pile of everything else to find the information that they might need.
I'd say if you want someone to have data or access, you should work that out in time.
But delete everything is good advice. I'd emphasize also keeping some list of services that you may have accounts or logins in and delete and close and erase every single one that you're not using.
Take what you want from it. Discard what you don't.
I think the author is having trouble because he is conflating concepts and roles that should be distinct. Sync, rotating snapshots, and deduplicated backups need to be kept entirely separate if you want any hope of maintaining your sanity.
So he's got sync but he's missing some sort of rotating snapshot system which would solve the stated concern of guarding against syncthing replicating corrupted data. Such automated snapshots can then be used as the source to feed the backup pipeline.
That hard drive doesn't make a good backup because it seems that it is always online. You need an offline backup that you manually plug in to run the job once every so often.
He's also making this more difficult than it needs to be by insisting that the backup drive be compatible with windows. Plug the drives for both snapshots and backups into a linux box, format them with a modern filesystem, and get on with life.
Sync is its own clusterfuck and I have yet to arrive at a satisfactory solution myself despite wasting inordinate amounts of time on it. IMO you either go with a network share or you make due with the "least bad" option of syncthing. Personally I've more or less settled on sshfs at this point not because it's particularly good but because it works well enough and doesn't add any additional complexity.
Personally I use btrfs snapshots on all my devices, those get streamed across the network to a NAS, and the contents of the NAS are periodically (every few months) stuffed into borgbackup on redundant offline devices. Aside from sync the other problem you'll run into if you're a data hoarder is how to split backups across multiple drives once you exceed a few TB. Because external drives only get so large but the NAS will inevitably keep ballooning.
Not if you use git-annex.
There is a straight up data loss bug in 2.4.3 (zeroed out files, totally silent).
https://github.com/openzfs/zfs/issues/18366
This caused all sorts of grief at day job where dedupe was useful for a big cache. Conversely I've run ZFS at home for like 15 years at this point without trouble. But this one is an absolute nightmare.
That gets synced and it's been trouble free so far.
But if you’re on Linux, it’s better to use ZFS or Btrfs and native filesystem snapshots, since they’re atomic.
I've seen tears shed because a wife couldn't get into her dead husband's iPhone to get the photos off from their final holiday together.
Technology is HORRIBLE in this space socially speaking.
https://git-annex.branchable.com/scalability/
> Scaling to hundreds of thousands of files is not a problem, scaling beyond that and git will start to get slow.
So that's probably insufficient for me by at least a couple orders of magnitude. I'm able to maintain my sanity because snapshots simply capture device state, the NAS collects all snapshots while maintaining their independence, and (so far) borg has been sufficiently scalable to deduplicate any collection I've thrown at it.
no information is so important to violate others' right to their own property.
so i contribute as little as possible, in the name of "community", per the guidelines.
seems like, eventually (mostly) no one cares