I am also not entirely sure whether "manorboy" is a good benchmark, but for estimating the function call overhead it should be ok. For the access to variables part, other considerations are far more important IMHO. So I am not sure I would put too much weight on the accessing "k" part of the post.
n3654 contains a criticism of the lambda approach from a language-design perspective. The issue is that it seems not to be a good fit for C. Copying values is always cheap, but this obviously works in toy examples but not necessarily for interesting data structures. In C++ you could then capture a pointer as a value, but then you haven't avoided an indirection either. In general, this is fine in C++ as smart pointer than deal with the memory management, but in C having a captured value in a lambda you can not have explicit access to anymore does not make too much sense.
The "unfortunately, is that unlike C++ there are no templates in C" is also interesting. I fled from C++ back to C exactly because of templates. In a performance context, the fallacy is that you can always create super-optimized code using compile-time techniques that absolutely shines in microbenchmarks (such as this one) but cause a lot of bloat on larger scale (and long compilations times). If you want this, I think you should stick to C++.
(https://news.ycombinator.com/item?id=46243298)
Which rather suggests to me that such a scheme, but generated by the compiler, should have a similar performance to said "Normal Functions" and hence similar to his preferred lambda form.
Since his benchmark environment is so unwieldy, I may have a go at extracting those two code sets to a standalone environment, and measure them so see...
xgcc (GCC) 16.0.0 20260103 (experimental)
1.50 gcc -ftrampoline-impl=stack -Wl,-no-warn-execstack
1.11 gcc -ftrampoline-impl=stack -Wl,-no-warn-execstack -DREFARG
7.21 gcc -ftrampoline-impl=heap
7.34 gcc -ftrampoline-impl=heap -DREFARG
0.93 gcc -DWIDEPTR
1.38 gcc -DWIDEPTR -DREFARG
1.40 gcc -DDIRECT
1.05 gcc -xc++ -std=c++26 -DFUNCREF -DDEDUCING
19.68 gcc -xc++ -std=c++26 -DDEDUCING
20.73 gcc -xc++ -std=c++26
6.31 gcc -xc++ -std=c++26 -DDEDUCING -DREFARG
6.31 gcc -xc++ -std=c++26 -DREFARG
Debian clang version 16.0.6 (15~deb12u1)
21.11 clang -xc++
6.16 clang -xc++ -DREFARG
1.66 clang -fblocks
1.70 clang -fblocks -DREFARGObviously I sympathize as someone who spent the front half of their career writing C and has no interest in writing C++ beyond toy examples. Nevertheless, IMO you're cutting off a choice here, forcing yourself to make potentially sub-optimal choices.
Sometimes the templates would have been better and writing X-macros is just worse, ergonomically at least. Perhaps C++ people use templates more often than they should, but I'm sure "never" is not the correct amount.
Still, if C gets fat pointers that's at least a step in the right direction. So I must wish you the best with that part.
I also recently removed some code written with templates in C++ from one of my projects, because this single file caused compilation times for a ca. 450 file C project to increasse by 50% from 20s to ca. 30s.
AFAIK the C standard doesn't prevent an implementation from using fat pointers, this is one of the reasons why the conversion from pointer to integer only works in one direction. This is actually necessary for segmented memory. The compiler is allowed to optimize based on the assumption, which allocation a pointer comes from (including the allocations boundaries, i.e. size), even if bytewise the pointers would be equal, so you could argue, that C abstract machine already has fat pointers.
I am really not sure if all these observations mentioned in the article are 100% correct, though.
First : Code seems to be compiled with clang. On Linux with gcc the native function one is way faster than the clang one.
Second: The author does run the code on ARM64/MacOS .
At least on my ryzen CPU on Linux with gcc the "normal C code" is way faster than anything else. Not that we do not need to thing about "closure" type functionality, but one should be careful to extrapolate implementations from one compiler on one platform to the rest of the pack.
Regarding N3654 I am not sure how to benchmark it here, since C could potentially use __builtin_call_with_static_chain , but I am not sure how to write the function to use the chain for accessing the variables.
I tried to estimate N3654 it by using "tinygo" which is AFAIK using the usual Calling ABI, but it was a factor of two slower than clang. Even "go" with its very specific ABI is still much slower. I discovered this isn't representative since runtime calling costs had been totally shadowed by costs of allocations.
Even the rust example I am usually using http://www.reddit.com/r/rust/comments/2t80mw/the_man_or_boy_... is much slower than anything else, presumably because of the "Cell" needed
TLDR: This micro benchmark might be misleading
A trick one can do is to let it create the trampoline and then read off the two pointer from the position in the code where it is stored. Not portable and you still have the overhead for creating the trampoline, but you do not need the executable stack anymore.
So, we need to swap to the logarithmic graphs to get a better picture
I wish more people would know about decibels.
> I wish more people would know about decibels.
Huh? Is there any difference? https://en.wikipedia.org/wiki/Decibel:
“The decibel […] expresses the ratio of two values of a power or root-power quantity on a logarithmic scale”
That whole 'zero cost abstraction' idea is not unique to C++ since all the important work to make the 'zero cost thing' happen is performed down in the optimizer passes, and those are identical between C and C++.
- Ada? Many cool features, but that's also the problem - it's damn complicated.
- C++ combines simplicity of Ada with safety of C.
- Rust and Zig are only ~10 years old and haven't really stabilized yet. They also start to suffer from npm-like syndrome, which is much more problematic for a systems language.
- ATS? F#? Not all low-level stuff needs (or can afford) this level of assurance.
- Idris? Much nicer than ATS, but it's even younger than Rust and still a running target (and I'm not sure if zero-runtime support is there yet).
I mean, yes, C is missing tons of potentially useful features (syntactic macros, for one thing), but closures are not one of them.
How so? In C++ a lambda is just a regular type that does not allocate any memory by itself. You have in fact precise control over how/where a lambda is allocated.
Any C implementation of capturing lambdas has the same problem of course, that's why the whole idea doesn't really fit into the C language IMHO.
Basically, with each capturing lambda, a context type _Ctxof(fn) is created that users can then declare on stack or on heap themselves.
Using C++ doesn't mean having to use the whole standard.
Now if C type safety actually was like Modula-2, Object Pascal or Zig, that would not be as bad.
C++ adds more new unsafe concepts on top of C than C ever had though ;) (see std::view, std::range or good old iterator invalidation - at least in C the unsafety is in your face and not hidden under layers of stdlib code)
For example, I'm maintaining some 20 year old C code, which the employer adopted around 10 years ago. It will likely stay in use at least until the current product is replaced, whenever that may be.
C does not have closures. You could simulate closures, but it is neither robust not automatic compared to languages tha truly support them.