Rendered at 11:26:19 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
MaxBarraclough 14 hours ago [-]
It's a non-moving collector. It might be high performance by the standards of C++ garbage collectors, but I doubt its performance could be anywhere near what a decent JVM can manage, especially in the absence of finalizers/destructors.
As Ron Pressler (pron here on HN) has been emphasising recently, [0] the 'sweep' phase of a moving garbage collector is unaffected by the size or number of dead objects in the heap (at least in the typical case, where there are no finalizers). This isn't the case here though (it uses free lists), or in any C++ GC.
The policy of running all destructors on the same thread doesn't really seem like 'high performance' architecture either, even if there are good reasons for it.
Still a neat project though. I rather like this:
> Oilpan uses a Clang plugin that statically verifies, among many other things, that no heap objects are accessed during destruction of an object
I'm not sure I understand this:
> Oilpan is a garbage collector written in C++ for managing C++ memory that can be connected to V8 using cross-component tracing that treats the tangled C++/JavaScript object graph as one heap.
In what sense are they treated as one heap? How can they be, given that V8 uses a moving GC for its JavaScript heap? Does it just mean there's some mechanism for a C++ object to refer to a JavaScript object, and vice versa?
Node 26 added a bunch of profiling stuff that should make this a bit easier to pick apart. I'm being lazy and waiting for perfetto support on windows to get fixed so I haven't bothered yet
whizzter 32 minutes ago [-]
I think it's a good thing to separate requirements/problems/etc.
1: Java has an horrendeous amount of unnecessary garbage due to their erasing generic design (until Valhalla ever arrives), there was some old blog post that chronicled the creation of Dictionary for C#'s System.Collections.Generic (as opposed to the initial Java-like System.Collections) and how it reduced garbage-load by an magnitude.
Java chose compatibility (code continued running as before, with the added compiler checked type safety), so you could use the same Vector, List,etc classes but it also means a ton of boxed Integer,etc objects.
C# went for proper runtime generics, packable struct:s and deprecated their initial collection classes. C++ code with templates can similarly benefit from developer controlled memory usage in places.
You don't need the absolute best GC's when your language doesn't create the same amount of GC-pressure.
2: While the compaction of objects has benefits in compactation and later runtime it does add complications everywhere (you either need to halt while updating moving pointers or have read-barriers).
3: C++ collectors are tricky to get right since compilers don't support it (I'm a bit miffed that they didn't see the GC support through and now removed it), personally I dislike that Boehm has taken so much mindshare since it's lacking in many respects and has given GC's a bad reputation. Reading this article I think it fulfills most parts well (even if it's not 100% faitful to the "abstract" C++ model even if it's probably fine in practice).
4: Reading the article I'm also surprised by the usage of plain pointers, a wrapped pointer type could've provided some concurrent marking support also (assuming optimizing compilers doesn't mess it up), some older GC or other PLang papers measured a 10:1 memory read/write ratio of real world code, so write barriers (esp if only for reference) often aren't that expensive if you're gunning for better latency via concurrent marking.
5: JS <-> C++ is probably the biggest reason they've created their own GC, early JS engines (IE6 iirc) often did reference-counting in C++ and JS GC's and could end up with reference cycles forcing JS developers to avoid certain patterns.
Having unified reachbility in the object graph:s avoids these issues since liveness, in terms of managing moving objects in one heap.
Iirc you created JS-GC:reference objects in V8 interfacing code in the past, didn't check the internals but those objects probably registered themselves with the JS GC,either to be update the references or pin those objects temporarily, you need the interface somewhere, and if the GC's can cooperate at the same time it's an improvement overall.
OskarS 56 minutes ago [-]
It's interesting reading this, wondering how much C++26 static reflection could improve the ergonomics of a system like this. Like, if you tag the class with a specific attribute, can you have have it generate the Trace() function automatically? Can you have it automatically wrap the other GC classes using the Member<> template? I think implementing the Trace() function might be doable (you'd do it in the GarbageCollected<> base class that uses CRTP, right?), but maybe not wrapping the types, I'm not sure if reflection allows you to modify types of fields in that way.
whizzter 27 minutes ago [-]
Automatic trace definetly yes, however since they also removed GC support from the standard it's quite useless (exact object tracing is useless if you cannot accurately scan the stack).
Also co-routines with it's under the hood management of coroutine stack probably would've complicated GC support even more...
d_finch 2 hours ago [-]
Interesting, but I still see `std::shared_ptr` and careful ownership as the C++ way. GC feels like adding another runtime dependency.
whizzter 15 minutes ago [-]
V8 has a C++ GC because if you have shared C++ <-> JS object graphs you either need extra machinery to find cycles in individual code-paths (very error prone) or let C++ data live in JS objects anyhow (initial V8 interop way).
Writing extensions for a GC'd language is painful or you let your interop code live in a GC world, I'm pretty sure you more or less always end up with the latter unless you add a way for the hosted language to support some way to control the host language (like JS code "understanding" std::shared_ptr's but that probably adds other securiy risks in terms of low-level exposure).
It's an engineering tradeoff, if you're 100% C++ you don't need that.
Rhaskins 2 hours ago [-]
Takes me back to trying to wrangle memory in a large C++ app. True high-perf GC would've been a godsend then.
As Ron Pressler (pron here on HN) has been emphasising recently, [0] the 'sweep' phase of a moving garbage collector is unaffected by the size or number of dead objects in the heap (at least in the typical case, where there are no finalizers). This isn't the case here though (it uses free lists), or in any C++ GC.
The policy of running all destructors on the same thread doesn't really seem like 'high performance' architecture either, even if there are good reasons for it.
Still a neat project though. I rather like this:
> Oilpan uses a Clang plugin that statically verifies, among many other things, that no heap objects are accessed during destruction of an object
I'm not sure I understand this:
> Oilpan is a garbage collector written in C++ for managing C++ memory that can be connected to V8 using cross-component tracing that treats the tangled C++/JavaScript object graph as one heap.
In what sense are they treated as one heap? How can they be, given that V8 uses a moving GC for its JavaScript heap? Does it just mean there's some mechanism for a C++ object to refer to a JavaScript object, and vice versa?
[0] https://youtu.be/xr73mR7ii9M?t=1081 Principles of Memory Management in Java, September 2026
1: Java has an horrendeous amount of unnecessary garbage due to their erasing generic design (until Valhalla ever arrives), there was some old blog post that chronicled the creation of Dictionary for C#'s System.Collections.Generic (as opposed to the initial Java-like System.Collections) and how it reduced garbage-load by an magnitude.
Java chose compatibility (code continued running as before, with the added compiler checked type safety), so you could use the same Vector, List,etc classes but it also means a ton of boxed Integer,etc objects.
C# went for proper runtime generics, packable struct:s and deprecated their initial collection classes. C++ code with templates can similarly benefit from developer controlled memory usage in places.
You don't need the absolute best GC's when your language doesn't create the same amount of GC-pressure.
2: While the compaction of objects has benefits in compactation and later runtime it does add complications everywhere (you either need to halt while updating moving pointers or have read-barriers).
3: C++ collectors are tricky to get right since compilers don't support it (I'm a bit miffed that they didn't see the GC support through and now removed it), personally I dislike that Boehm has taken so much mindshare since it's lacking in many respects and has given GC's a bad reputation. Reading this article I think it fulfills most parts well (even if it's not 100% faitful to the "abstract" C++ model even if it's probably fine in practice).
4: Reading the article I'm also surprised by the usage of plain pointers, a wrapped pointer type could've provided some concurrent marking support also (assuming optimizing compilers doesn't mess it up), some older GC or other PLang papers measured a 10:1 memory read/write ratio of real world code, so write barriers (esp if only for reference) often aren't that expensive if you're gunning for better latency via concurrent marking.
5: JS <-> C++ is probably the biggest reason they've created their own GC, early JS engines (IE6 iirc) often did reference-counting in C++ and JS GC's and could end up with reference cycles forcing JS developers to avoid certain patterns.
Having unified reachbility in the object graph:s avoids these issues since liveness, in terms of managing moving objects in one heap.
Iirc you created JS-GC:reference objects in V8 interfacing code in the past, didn't check the internals but those objects probably registered themselves with the JS GC,either to be update the references or pin those objects temporarily, you need the interface somewhere, and if the GC's can cooperate at the same time it's an improvement overall.
Also co-routines with it's under the hood management of coroutine stack probably would've complicated GC support even more...
Writing extensions for a GC'd language is painful or you let your interop code live in a GC world, I'm pretty sure you more or less always end up with the latter unless you add a way for the hosted language to support some way to control the host language (like JS code "understanding" std::shared_ptr's but that probably adds other securiy risks in terms of low-level exposure).
It's an engineering tradeoff, if you're 100% C++ you don't need that.