"Stop doing that"
If you care about latency, disable swap. System wide or for the specific the cgroup.
If you care about latency, mlock() your memory, do not disable swap. Swap is good and gives the kernel an equal opportunity to evict data and code pages.
I'd rather have applications be oom_killed than having them swap out, the former is rather obvious and demands action.
I believe that's actually the same reason why Apple stopped using GC in their frameworks in favour of automatic reference counting.
Many make the mistake to think there is only one way to do a GC.
One of the authoritative books on the subject, https://gchandbook.org/contents.html
And a quite well known paper on the matter as well, https://dl.acm.org/doi/10.1145/1035292.1028982
I'm sure an agent can work on this and get some numbers with a day's worth of tokens.
For example, why not swap out not by LRU page but by dense node clusters on the heap graph, maintaining in-memory summaries of inbound and outbound edges for liveness? If you do this, you don't have to swap the cluster in to do a GC involving it.
If the whole cluster becomes unreachable, you wouldn't even have to swap it back in to get rid of it: you'd just drop the swap reference and deem the swap space free.
I don't see anything this deeply integrated happening near-term, but it's fun to think about.
I use Rust where i need low latency.
In general when you tune knobs for GC, you pay for benefits in one area with sacrifices in another. Two big knobs to turn are pause latency and throughput. You probably wouldn’t want to go full “optimize for latency” because you’d end up with poor throughput. Also vice versa. Java’s reputation for poor GC performance is partly due to historical defaults that tune it for throughput.
Go’s GC is already a “concurrent mark-sweep garbage collector” and already has “extremely low mutator pause times, on the order of tens of microseconds”. It sounds like on-the-fly is just a different flavor of what Go already has.
It's a well known algorithm. Folks who do GCs for a living know about it. The folks who work on Go are surely aware of it. I'm assuming that they do not use it for a good reason, hence my question!
Fil-C's GC (Fil's Unbelievable Garbage Collector) uses an alternative on-the-fly algorithm, which I call Phil's Concurrent Marking.
I've documented it here: https://fil-c.org/fugc
Here's the source: https://github.com/pizlonator/fil-c/blob/deluge/libpas/src/l...
Phil's Concurrent Marking differs from DLG in that it only requires a Djikstra barrier and uses a permagrey stack (something that Go used to do).
However, FUGC does clever things for coroutines (as in ucontexts, which Fil-C supports) - they are not permagrey; they only become grey if they execute. That's relevant to Go because Go moved away from permagrey stacks because of coroutine scan overheads, which the FUGC coroutine strategy might avoid.
But even if Go could not go back to permagrey, then the answer would be to use DLG, which would involve using the combined Yuasa+Dijstra barrier, which Go uses today anyway
(though of course a swap is a swap - but you can "trigger" it depending on your memory or file access pattern)
Not really, reference counting can cause a single object deallocation to trigger an arbitrarily long chain of deallocations.
Disabling swap will just moves pressere elsewhere: to code pages. And evicted code page is no better: full stall while kernel loads that page from disk.