- system dies under memory pressure (regardless of swapping, actually not having swap makes it worse which should be common knowledge by now)
- system dies under disk pressure even if there are tons of free memory (this one is fun to diagnose)
- system can technically not die, but render itself useless (or worse) under memory pressure by the oom killer
- memory compression of any sort is not enabled
- ...
Both Windows and macOS do so much better out of the box for essentially any workload.
We test FreeBSD, Linux, macOS, NetBSD, OpenBSD, and Windows in Zig's CI fleet. Of these, Windows is the only OS that we've had to configure with swap double the size of physical RAM to not hit completely unjustifiable OOMs.
By "unjustifiable", I mean that we're not even close to actually running out of physical memory (let alone swap), but the MM seems to be doing a horrible job of making unused memory actually available to processes.
It's possible there's a relevant configuration knob here that we're just not aware of... but the point is, the default behavior does in fact suck.
It sounds like the application wanted to allocate contiguous regions of memory when none were available. That's a typical indicator of an 'early' OOM condition.
NT is optimized, since day one, to swap.
This is a feature, not a bug.
Working around mm nuttiness is a frequent source of frustration.
Resource control via cgroupsv2 sounds better for both desktop and server use cases, differing by what processes receive minimum resources. But IO isolation remains elusive. Some suggest a database of hardware capabilities, I wonder if a cheap estimator of throughput and latencies could do this dynamically.
Anyway, I'm not convinced user space should care about or rely on the kernel oomkiller. Well before it even considers the situation human interactivity with the system was lost.
for some more user sanity huge ram consumers under memory pressure should be suspended, paged out as much as possible and the user notified (macOS does it right)
I've gone through this exercise in the past on much older kernels which they cover as well and just me personally I ran into less issues by leaving overcommit to 0 and just dropping the overcommit ratio to 0 and setting the oom_score_adj for programs as high as 1000 if I wanted vmscan to leave them alone and of course using the Redhat formulas for setting vm.min_free_kbytes, vm.admin_reserve_kbytes, vm.user_reserve_kbytes. And of course be vigilant in disallowing app owners from using every last bit of memory.
[1] - https://man7.org/linux/man-pages/man5/proc_pid_oom_score_adj...
I agree with the blog post's technical contents, but I feel we came across too strong in the title. For Ubicloud as a managed Postgres provider, we use strict memory overcommit. Our experience with operating Postgres at scale taught us that it's better to enable this than going with the defaults.
However, I can see many other scenarios, where using strict memory overcommit would have unanticipated side-effects. That's why Linux doesn't go with strict memory commit as its default.
For now, we have overcommit_ratio set to a value that is stable from experience, but there really seems to be no silver lining. Go is very happy to allocate a lot of virtual memory, but so are most managed languages. The best solution would probably be to host the backend and the database on separate servers.
There’s no correct-in-general answers to those questions. This is a hard problem due to context dependence; that’s why there are so many knobs.
Firstly, if you take what's written at face value, it seems there's a serious logic error in the OOM killer. If it genuinely counts shared memory against a process, then its logic is wrong because killing that process wouldn't release that much memory, it'd need to kill all the processes sharing that memory to release it. So, maybe it should ignore shared memory in its calculations, or weight them by number of processes sharing it, or whatever.
The other issue is kind of true, but shows a stubbornness from the developers to change how they approach the problem. It's true that if any task could be partially through updating shared memory when it is killed, then all bets are off as to the state of that memory. However, if that's the case, then there should already be some kind of locking mechanisms in place to prevent multiple processes updating the same pages anyway.
Probably the current solution is: something gets locked when modified; every other process would need to spinlock until it's released; locking process is killed; everything else is stuck; another PG thread notices the child has died and kills everything else; on restart the DB has to be recovered.
A different solution could be: give every process its own private part of the shared memory for when it starts a transaction; have an indirection table from page number to shared memory page; for every page that needs to be modified, a new page is allocated from the freed pages list; that allocation is recorded in the private part of the shared memory along with the page number it's replacing; the old page is left untouched and copied to the new page along with any changes; the change list is terminated; then as the last step we update the indirection table for every page that was modified. If the process was killed at any point, we can either roll back the entirety of the transaction (freeing every allocation it made for replacement pages), or if the list was marked as terminated, finish off updating the indirection table with the changes required and marking the original pages as unused. At that point, the lock can be released knowing that shared memory is still entirely consistent.
Some of that is probably happening anyway if postgres supports reading from tables concurrently with an active write transaction on the same table, in which case the logic on how to mark those now freed pages as still in use is required. In that case, each process can also maintain a list of pages it has marked as still being used in case a reading process is killed off.
Took k8s ages to get Swap support.
We lost something when we accepted that Hyperscalers just tell you to use more moemory. It was shitty 5 years ago and today especially after the ram price increases
And now, with PSI + MGLRU, situation is much better, but there are still missing features/subsystems which would be nice to have. For example there's no simple way to lock memory mlockall-style to ensure that rarely used daemon would not face long no-cache-latency upon accessing the first time after long idle time.
The article ignores the proper modern solution to prevent OOM killing of critical processes - OOM Score Adjust.
Tuning CommitLimit manually is an archaic, imprecise, and error-prone way to handle memory limits, only suitable for single-process workloads that can handle ENOMEM properly. It completely ignores dynamic file page cache memory allocation. You still can get OOM if you get unusually high file activity. On the other hand, under low file activity, it wastes memory on the same page cache, because it can't be reclaimed without memory pressure, and memory pressure can't be created because workload hits ENOMEM earlier. Don't use strict overcommit.
Even a revised heuristic that only spots large, individual allocations is not going to do the job.
Oom score adjust also doesn’t do the job: because the only interesting workload is Postgres, if a backend does a page fault that needs memory, who dies? Another sibling Postgres, almost certainly. Then postmaster does crash recovery, which most would rather avoid. High performance databases with distant checkpoints can take a while to come back up.
Unfortunately, many programs commit 2x memory than they actually use. Often I see ~32GB committed and ~16GB resident.
First, Linux's default memory management strategy is bonkers. OOM killing rarely actually works in my experience, at least on desktop. It takes ages to kick in and usually the system just freezes and you have to hard reboot. I've experienced this on every Linux system I've used, even my current one with 128GB of RAM and 64GB of swap, so don't say "it works for me". Windows and Mac do not have this issue at all, so clearly it's possible to do it better.
Has anyone tried using strict overcommit on desktop Linux?
Second, this bug is a great counterpoint to those annoying people who naysay Rust with "but not all bugs are memory safety bugs, what about logic bugs? huh?". Rust code would not have had this bug.
GOMEMLIMIT works very well if you set it to around 90% of available memory as a rough heuristic. You should definitely profile your application to fine tune this number (e.g. if you link with C libraries that hold large memory pools then Go doesn't account for that) but also to identify sources of spikey/leaky allocations. For example, encoding/json is notorious for it's inner sync.Pool hanging on to outsized buffers. There's usually a lot of low hanging fruit.
In my experience Go can be extremely stable in terms of memory footprint at both small (~O(1MiB)) and large (~O(256GiB)) scales, and it takes only a small amount of effort.
As far as GC languages go, it is by far the easiest to work with.
If the database requests more memory, it gets ENOMEM, but if the backend app requests more memory, it does get some more because it can overcommit?
Sounds dangerous, if the go program then writes to the overcommitted memory, you'd still trigger the OOM killer, right?
Whether failed transactions are actually so much more desirable than a OOM-killed process isn't quite obvious, but it might be easier to troubleshoot.
I run Firefox, VSCodium with LSP, Discord, Signal and there's still space left for a game like CS2. I'm not a heavy user by any means.
> I'm not sure they would do much better than crash
I have yet to see a program that silently handles allocation failures and doesn't crash. These days everything is coded to crash if no memory :(
> About once a year a real runaway process (usually a throwaway program I'm working on) gets OOM-killed
In my case it killed system critical processes with no way to recover. With disabled overcommit, it freezes for a while (usually for a minute or two), I close some random program of my choosing and then see in Resource Monitor what's eating my ram.
https://unix.stackexchange.com/questions/797835/disabling-ov...
I dont think it has an option for that.
A memory allocator can implement overcommit, because you can separate reserving virtual memory and having it backed by physical memory into two different system calls. But from the point of view of the kernel, any time it promises to give you physical memory that memory is backed either by RAM or by space reserved in the swap file
I'm not sure what point you're making here. Did you assume we had the page file disabled or something?
> This is a feature, not a bug.
OOMing when there's tens of gigabytes of unused memory is a feature...?
"Are you logged on to DB1?"
"...yes?"
"What did you do, it just died"
This has happened multiple times
Postgres handles allocation failures
It also works for the OOM killer: run a daemon with a child process that holds some fixed amount of memory. Adjust OOM scores of everything else on the system lower than the child. If the parent’s waitpid() returns due to an OOM kill, send an alert/shutdown nonessentials/sync buffers to disk and so on.
But it’s about 1000x slower to call the OS to allocate a new page, than to just use a pointer to preallocated space.
(allocating virtual memory, but not committing is different, and should be handled with MAP_NORESERVE (or similiar)).
The purpose of the system commit limit and commit charge is to track all uses of these resources to ensure they are never overcommitted — that is, that there is never more virtual address space defined than there is space to store its contents, either in RAM or in backing store (on disk).
- Windows Internals, 7th EditionIf no memory is available where a page file would make a difference, this leads to application crashes instead. A crash is (usually) worse than paging.
Certain applications, Photoshop being the historical example, will outright fail to run with no page file present.
Same happens if the page file is full. In that case, why don't those programs use disk directly instead?
No such problem would've ever occured if programs hadn't allocated more than they actually use.
Typically, performance drops enough that the user kills the program or reboots before the page file expands to fill the disk. And other threads here suggest there is something that will prompt users to kill programs in states like this.
> No such problem would've ever occured if programs hadn't allocated more than they actually use.
That's part of the issue, but sometimes things do in fact use too much memory as well as allocate too much.
Another part of the issue is that few programs are built to handle allocation failures.
And then you have a metrics issue. There's not really a good metric to know when you're out of memory, other than performance collapse. If your applications don't use disk, it's not too hard; but when they do use disk, performance will collapse once there's insufficient memory to provide the disk caching needed. In my experience, adding a small swap and monitoring swap i/o can be pretty helpful, and a small swap doesn't tend to allow long thrashing when memory use grows. But that's not universal and everybody loves to hate swap these days.
An application that grows in such a way (besides having backing stores for memory-mapped files, as well) will often perform so poorly that it requires addressing (adding RAM, looking for application faults, etc).
A page file is insurance, one that can last you much longer than available system memory.
They used an int with special meanings for negative/0/positive values. Very common in C, and not at all type safe (all meanings have the same type). In Rust you would use an enum or Result, it would be type safe and the refactoring mistake they made would have been a compile time error.
> On the modern desktop, where programmers don't care about failing malloc(), disabling overcommit is shooting yourself in the foot. As you can observe, the memory allocations start failing long before the memory is exhausted.
If it fails with the default mode the whole process will get killed by the OS. Is that really much better?
Though to be absolutely pedantic, !x is an int for x:int in C, there is no bool coercion involved; an if-statement takes an expression of any scalar value and evals to true on non-zero. Not that that helps to avoid introducing bugs anyway.
In short, Windows partially does the same lazy thing but unlike Linux with its optimistic overcommit, it is stricter about commit/backing-budget reservation.
The Linux Kernel OOM killer kills random things. Userspace OOM killers are meant to improve this, and they work well in a server situation when you already know in advance what is likely to go haywire and what is safe to kill. But they don't work well on desktop (some of them are improving but it doesn't seem to be a priority).
The Windows OOM killer by comparison usually kills something sensible (i.e. the program that is actually using all the memory), and asks the user for permission before killing it (when possible). You do see a lot of memes of situations where it fails.
By default, the Linux kernel kills the largest process in the system (unless OOM adjust was applied).
Don't kill what I'm using.
You don't need it if you have everything allocated upfront. TigerBeetle does this, everybody else can.
Using something like Rust is already a huge win when compared to shipping a browser or running Node.js.
> Your argument falls flat when a page file can be multi-GB and automatically grow
This doesn't solve the original issue and only masks the underlying problem.
You're moving goal posts. No, a page file doesn't solve the problem of a misbehaving application, but it does solve the problem of an app crash because no more VAS allocation can be made.
You should really dive into Windows Internals. Only misinformed gamers turn off page files.
Not in the age of NVMe it doesn't. Swap is fast now. Plus, at least on Linux, you can put zswap in front of the regular swap and introduce an even faster level of memory hierarchy and thereby make page-outs even more profitable.
Your system your rules. But you twiddled a knob to appease your tweak imp, you didn't like the new behaviour, so you called it "fucking stupid". My experience is that the windows vmm is a very high quality component, so simple heuristics tells me PEBCAK.
I can think of two different possible reasons why memory compression might require a page file. Until you understand the technical reasoning, you don't know whether the design is stupid or clever. So taking a strong position is _eliding over a core tenet of wisdom_
I ran out of disk space because Windows decided my page file should suddenly be 64gb. With no space left to save my work, I found out there's no way to shrink the page file without rebooting. Of course the next thing I did was try to cast it off.
Maybe it was my skill issue for not expecting a sudden 64gb file there... or my skill issue for not choosing a nonzero size after?
I see a major footgun in memory compression that Linux doesn't care about - it'll just oomkill your shit, hence TFA. But footguns are not acceptable in mass-market software. Apple also removes them.
It's easy to look at the half-baked ui libraries and fucking Teams and fucking SharePoint and conclude that Microsoft engineers are stupid. But, for your interest: https://blogsystem5.substack.com/p/windows-nt-vs-unix-design