Persisting state between AWS EC2 spot instances(peteris.rocks) |
Persisting state between AWS EC2 spot instances(peteris.rocks) |
This is leading to rapid progress in clustered/distributed filesystems and it's even built into the Linux kernel now with OrangeFS [1]. There are also commercial companies like Avere [2] who make filers that run on object storage with sophisticated caching to provide a fast networked but durable filesystem.
Kubernetes is also changing the game with container-native storage. This seems to be the most promising model for the future as K8S can take care of orchestrating all the complexities of replicas and stateful containers while storage is just another container-based service using whatever volumes are available to the nodes underneath. Portworx [3] is the great commercial option today with Rook and OpenEBS [4] catching up quickly.
https://aws.amazon.com/about-aws/whats-new/2017/09/amazon-ec...
Twenty years ago, software was hosted on fragile single-node servers with fragile, physical hard disks. Programmers would read and write files directly from and to the disk, and learn the hard way that this left their systems susceptible to corruption in case things crashed in the middle of a write. So behold! People began to use relational databases which offered ACID guarantees and were designed from the ground up to solve that problem.
Now we have a resource (spot instances) whose unreliability is a featured design constraint and OP's advice is to just mount the block storage over the network and everything will be fine?
Here's hoping OP is taking frequent snapshots of their volumes because it sure sounds like data corruption is practically a statistical guarantee if you take OP's advice without considering exactly how state is being saved on that EBS volume.
A spot instance interruption isn't a system crash, it's a shutdown signal. Storing your important spot instance data on EBS is recommended by AWS. If your application can't handle a normal system shutdown without losing data, your application is at fault, not your system setup.
>exactly how state is being saved on that EBS volume
Files are written to a filesystem which is cleanly unmounted at shutdown when interruption happens.
Unless something in the system shutdown fails to give the application what it needs (for instance, time) to shutdown cleanly. Which is entirely possible considering that Amazon is selling you the spot instance on the given assumption that it can give the hardware at any time to somebody who is willing to pay more. Amazon does not guarantee the time needed for a clean shutdown (only that a two-minute warning will be available via their proprietary mechanism, if you architect your application to monitor for it) for a spot instance anywhere in their documentation, and you would be ill-advised to not architect for that.
> Storing your important spot instance data on EBS is recommended by AWS
Because EBS itself is reasonably reliable. If you have configuration data (i.e. in /etc) for a legacy application that isn't managed, it's reasonable to mount that data on EBS since it's rarely written to and writes are generally human-initiated and human-monitored (with operations policy possibly mandating a snapshot even before any changes are made).
That's still very different from daemon writes to /var. Take for instance, the PostgreSQL documentation which warns that snapshots must include WAL logs in order for the snapshot to be recoverable, and that it is quite difficult to restore from a snapshot if you stored your WAL logs on a different mount: https://www.postgresql.org/docs/10/static/backup-file.html
You need to understand precisely how your application is treating your storage and act accordingly. Thinking that all applications interact with storage the same way is dangerous and liable to cause data corruption and loss. That's all.
You're assuming that people are saving their state in databases to begin with. If you're saving state to a database in production, typically you're communicating with that database over a network connection, and not running the database on the same machine as your application. Containerizing databases is a whole separate issue.
OP's specific example is saving /var/opt/gitlab to an EBS volume and expecting to be able to move it from one spot instance to another without corruption. That strikes me as insane.
- Yes, you get a notification, but it's a proprietary notification scheme that your application must be designed to poll for. Why can't Amazon use standard signals like SIGPWR to indicate imminent shutdown?
- Just because it isn't smart for non-spot instances doesn't suddenly make it smart for spot instances ;)
https://aws.amazon.com/about-aws/whats-new/2017/09/amazon-ec...
ec2-spotter classic uses this, but you can also make a pivoting AMI of your favourite Linux distribution.
One thing to watch out for is how to keep the OS automatic kernel updates working. AMIs are rarely updated and you're going to have a "damn vulnerable linux" if you don't get the updates just after booting a new image.
1) https://aws.amazon.com/about-aws/whats-new/2017/09/amazon-ec...
I suppose it's a decent solution if you don't want to deal with prefixes.
* https://github.com/sevagh/goat (my own) * https://github.com/UKHomeOffice/smilodon
This solution looks good, yet only applies to single instance scenarios. I presume this kind of thinking might move forward with EFS + chroot for an actual scalable solution that cannot be ran on Elasticbeanstalk.
http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-inte...
Learn something new everyday. :)
https://aws.amazon.com/blogs/aws/new-ec2-spot-instance-termi...
Personally, because my needs aren't constant. I might need two cores for two months followed by 100 cores for a week.
I would look at providers like OVH and even cheaper (Treudler, TransIP, RamNode, etc.) For example, an SSD with 2 vCPUs, 8GB RAM and 40GB SSD is 13.49$ per month from OVH.
(PS: Don’t use DigitalOcean, they tend to steal your credit if they feel like it. Lost 100 bucks "promotional credit" that way with only a few days notice)
The deal was found on LowEndBox, not sure if it's still available, but there are many other ones.
With a few on-boot scripts to attach-volumes / start-containers, it should be fairly easy to get going as well.
[1] https://engineering.semantics3.com/the-instance-is-dead-long...
Edit: or use AWS EFS
And since it's shared, you don't need to replicate data across multiple nodes... so if 10 compute nodes needs access to the data set, they can all just read it from the same EFS filesystem, no need to download it 10 times to each compute node.
So EFS can still be very cost effective compared to EBS.
A positive thing with EFS is that it can be shared across AZ while EBS needs to be snapshotted and then imported to the other AZ.
Attaching and detaching volumes is a good idea but I wouldn't use that to keep state
You will get a lot of benefit out of it, but may lose in performance, which is fine in 99% of the cases.
Currently they initiate an ACPI shutdown event at the termination time. It's hard to initiate a shutdown in a more standardized manner. An instance shut down via this signal will generally see the init process begin gracefully stopping services, eventually halting on it's own. Typically your init process will get increasingly aggressive with kill signals, as defined by your service definitions, eventually getting to SIGKILL. If your init process fails to get the vcpu halted, after a (undocumented?) period AWS will halt the cpu(s) for you. This is about as graceful a shutdown as you're going to get with 'standard' interfaces.
Termination Notifications go out of their way to give you an extra heads up, in case your application is unlikely to gracefully handle being shut down by the init system. Think DB hosts with a craploads of dirty blocks that take a few minutes to sync to disk at shutdown.
"The EBS root device and attached EBS volumes are saved..."
Some instance types don't support an instance root, but require an EBS root.
Also, NFS has different behavior with respect to buffer caching that needs to be taken into account. It often does not cache as effectively as block storage does.
This makes it easy to declare the volume as part of the deployment and automatically attach storage when the container is run. Mounting on the host isn't very easy (or even possible sometimes), especially with spot/preemptible instances and the increasing abstractions by managed K8S providers. The pricing model might need to be different though if billing on a container-mount level.
Shit happens at scale, it's precisely why ACID guarantees are important. Specifically in GitLab's case, because configuration is stored under /etc/gitlab, relying on EBS snapshots as a safeguard against corruption only works if the snapshot is taken of the entire FS, not just /var/opt/gitlab. If your machine is properly provisioned from an AMI or at least from some kind of configuration management, and you have some kind of reasonably-enforced policy which only permits changes through those management systems, then maybe you can get away with only taking a snapshot of /var/opt/gitlab, but now we're getting into the territory of "I understand how my data is being stored to the EBS volume (in this case, according to documented GitLab instructions) and I am acting accordingly". Then, if the /var/opt/gitlab snapshot ends up being corrupted, the odds of getting an uncorrupted snapshot increase with the more snapshots that you try, and this is probably good-enough in this specific instance because if you needed a better guarantee than that, you'd have a proper HA setup.
Even if you had data journaling, it won't give you consistency between different files. This post used Gitlab as an example, and git will break if some files in its databse are updated, but some not. Git doesn't use fsync to ensure their update order, I don't know if Gitlab enables it or if the performance hit is reasonable.
I can imagine (cough) an application where the application is trying to write some binary blob to disk, doesn't finish before shutdown, and upon reboot, tries to load the binary blob back into memory, fails because the binary blob isn't consistent, doesn't handle the failure well, and refuses to boot.
App's fault? Sure. Does the customer care at 2 am? Nope.
Honestly, it's much safer in that circumstance to have a frequently rebooting instance because it will quickly expose your app's fragility during normal operations instead of that fragility being exposed in a disaster.
I actually happen to agree with you in principle on this, and it's at the root of my current side project.
But sometimes you just don't have the flexibility to fix or replace the app. Ops engineering, like any other kind of engineering, is about dealing with real-world constraints and making the most of the resources you have. Most apps, on some notion of a fragility spectrum, are far closer to fragile than to antifragile, because fragile is the default, and extensive stress-testing to understand and plan for all failure modes before a production deployment isn't typically feasible. At that point, if you can't fix it, you have to work around it.
You will ultimately have many fewer resources available if your strategy is to gloss over failure modes by telling inexperienced engineers to hope they won't happen. It's technical debt and the interest payments are very high.
But ALL cloud providers provide warning before an instance is shutdown. There is absolutely no reason, other than a crash for an instance to have a hard shutdown.
Now I am happy with AWS.
They backtracked on that regarding non-promo credit (referrals etc) and gave a 1-year grace period.
FWIW, I've been very happy with DO, had a couple $5 VPSes there for 3-4 years and they've been remarkably reliable. One host migration, one SLA credit and lengthy failure analysis, and a bunch of notifications ahead of time for maintenance. More than I'd expect for most hosts in the price range.
Not the most powerful for your money, of course, but awesome if you need to run some services with a public IP and consistent uptime.
Actually, they only emailed users to warn them about this ca. 10 days before it was revoked.
I had gotten $100 promotional credit from DO with the GitHub student pack, and planned to use it in my second year of university, as I knew we had to do a practical project there where I’d need it. Well, a few weeks before that project was about to start, I got the email from DO telling me they’d invalidate all my credit next week. In the end, I hosted that project with OVH, and spent over 80€ on it.
But that was extremely annoying, and while I originally wanted to also move servers of a few projects I was hosting to DO, after this I decided not to.
Also, blog post here https://blog.digitalocean.com/details-on-expiring-digitaloce...
This is a question of trust. I have to trust that DO will keep my data safe, that, if the US government would be after my data, DO would prevent them from accessing it. I have to trust that DO won’t access my data.
How am I supposed to trust my, and my user’s personally identifying data, to a company that just like that revokes credit, without warning, and says "well, if you ask nicely, you can get it back"?
This is completely unrealistic. If the [local jurisdiction government] is after your data, they'll have your host, ISP, and anyone else give it to them.
(Inexplicable downtime = your server being imaged.)
Believing anything else, IMO, is purely delusional.
But! Lots of applications aren't built to handle partial writes, which will absolutely occur if apps are hard killed. Any disucssion around this topic should reference Crash-only Software [0][1][2] and Micro Reboots [3]
[0] https://en.wikipedia.org/wiki/Crash-only_software
[1] https://www.usenix.org/conference/hotos-ix/crash-only-softwa...
[2] https://lwn.net/Articles/191059/
[3] https://www.usenix.org/legacy/event/osdi04/tech/full_papers/...
I’m in Germany, my users are in Germany, and if I host with DO in Frankfurt, I have to trust that my data stays in Frankfurt.
I’m sorry, but I can’t trust such a company.