Skip to content

MakeCallback is very slow #24

Description

@ronag

Calling into JS from CPP is very slow. Is this due to async hooks? Could we disable that with a global option if the user doesn't need async hooks?

Activity

  1. joyeecheung commented on Nov 21, 2022

    @joyeecheung
    Member

    See past discussions nodejs/node#10101 - I think the conclusion was that it wasn't that slow, but that was measured 6 years ago and before V8 rolled out several architectural redesigns. We have also deprecated domains for a long time. Maybe the situations have changed now. Would be good to have some new measurements to see if it is worth a shot making it some kind of optional thing.

  2. anonrig commented on Nov 21, 2022

    @anonrig
    Member

    Could we disable that with a global option if the user doesn't need async hooks?

    Can you explain this in detail? I'm trying to understand the current implementation, and a quick explanation could/would save a lot of time :)

  3. ronag commented on Nov 22, 2022

    @ronag
    MemberAuthor

    Could we disable that with a global option if the user doesn't need async hooks?

    Can you explain this in detail? I'm trying to understand the current implementation, and a quick explanation could/would save a lot of time :)

    I don't know much about the details. Async hooks are used to propagate async context. It has a constant overhead even if you don't use it.

  4. mcollina commented on Nov 22, 2022

    @mcollina
    SponsorMember

    async_hooks have a tiny overhead if disabled and a significant overhead if enabled.

  5. ronag commented on Nov 22, 2022

    @ronag
    MemberAuthor

    async_hooks have a tiny overhead if disabled and a significant overhead if enabled.

    Is it possible to disable it?

  6. mcollina commented on Nov 22, 2022

    @mcollina
    SponsorMember

    Not really. That's described in the linked issue. Maybe a compile-time flag could do it but we are talking ns of difference.

  7. abenhamdine commented on Dec 1, 2022

    @abenhamdine

    see also previous related work in
    nodejs/node#41331
    nodejs/node#41279
    some big picture datapoints here, related to uWebSockets module : uNetworking/uWebSockets.js#341

  8. benjamingr commented on May 23, 2023

    @benjamingr
    Member

    Interesting stuff (though bad manners) at uNetworking/uWebSockets.js#904 namely that experiment (that does break async_hooks but also improves performance)

  9. H4ad commented on Jul 11, 2023

    @H4ad
    Member

    I could be wrong, but once this nodejs/node#48528 is merged, we can probably rewrite MakeCallback to use AsyncContextFrame, right? Assuming AsyncContextFrame is faster than async_hooks.

  10. benjamingr commented on Jul 12, 2023

    @benjamingr
    Member

    @H4ad that would work for people using AsyncLocalStorage but not for people otherwise using async hooks

  11. Qard commented on Oct 2, 2023

    @Qard
    Member

    I've been wanting to split all the async_hooks stuff in C++ out into a delegate class which we could use a no-op version for until a hook is registered. That would make it easier to at least contain the overhead when it's not in-use.

  12. joyeecheung commented on Oct 19, 2023

    @joyeecheung
    Member

    My understanding (or, from nodejs/node#10101) is that async hooks do not actually contribute much to the overhead when it's not enabled - and it's opt-in already so users would be willingly taking the hit. Most of the overhead comes from the tick callback for process.nextTick. While it also only incurs a lot of overhead when actually used, due to the prevalence of its usage in Node.js internals and the user land, the overhead is also much more prevalent. Doing something about async hooks helps a bit, but it's also only a very small piece of the puzzle because it's actually not used that much in the user land. process.nextTick() is the real beast here that affect almost all use cases.

  13. 15 remaining items

  14. uNetworkingAB commented on Oct 3, 2026

    @uNetworkingAB

    It's still a red flag that node::MakeCallback, which is the central piece of anything that delivers events to Node.js, is more than 2x the cost of the vanilla v8::Function::Call which already do support auto-drain of Promises (as alternative to the nextTick brainfart).

    These kinds of "meh, it's a tiny overhead" is exactly why Bun is in a completely different league performance wise. It's not even possible to make a native addon for Node.js that outperforms Bun, because node::MakeCallback chugs those CPU cycles like they are chocolate milk and stand for 40% of the entire CPU-time.

    It's not just a single "meh", it's a long term trend of a continuous stream of "mehs" that accumulate almost perfectly linearly:

    Image
  15. BridgeAR commented on Oct 4, 2026

    @BridgeAR
    Member

    @uNetworkingAB Node.js 14 is long long long outdated, so for comparison, please use latest versions.
    Next to that, please make concrete suggestions how to improve things using respectful communication / phrasing! Naming something bad is easy, making things better is often hard due to many constraints! There are often many reasons for something that might not immediately be obvious.

  16. nigrosimone commented on Oct 4, 2026

    @nigrosimone

    @uNetworkingAB I agree the gap with a plain v8::Function::Call is still too big: after the open PRs about 100 ns of MakeCallback are Node, a plain call is about 45 ns. The PRs so far take the easy part. To go near the minimum we have an idea, but it needs changes that should be discussed first, so here is a draft to collect feedback early.

    The idea is a fast path for the common case: a top-level callback with no async hook, nobody using executionAsyncResource(), no ALS frame and nothing in the tick queue. There almost all the work of the callback scope changes nothing, so MakeCallback could check this once and do little more than the Function::Call. Three parts:

    1. Async ids. They are pushed and popped on the id stack on every call, also when no hook is enabled. But executionAsyncId(), triggerAsyncId() and process.nextTick() read only the two shared fields, and with no hook only executionAsyncResource() needs the stack. So the scope can swap the two ids in place and skip the stack, with a small chain kept for executionAsyncResource(): who reads the ids pays nothing. It is close to the no-op delegate @Qard proposed above.
    2. Microtask checkpoint. It runs after every top-level callback also when the queue is empty. V8 already returns early in this case since v8/v8@25650aa (V8 15.4), but Node main is on V8 14.6, so this one needs only a cherry-pick.
    3. Environment and context. On every call the Environment is found from the creation context of the function, and its context is entered. An addon that calls the same function many times (uWS calls one handler for every request) could resolve them once, with an opt-in handle. This is new public API.

    It is still in study and nothing here is measured as a whole. When nodejs/node#66326, nodejs/node#66344 and nodejs/node#66395 are in, I will open an issue with the precise plan and the numbers. Feedback on the three points is welcome already now, mostly on the first one, because it changes when the id stack is written.

  17. uNetworkingAB commented on Oct 4, 2026

    @uNetworkingAB

    Node.js 14 is long long long outdated,

    That's irrelevant. That's obviously not the point being made. The point being made is to show objective proof that there has been a long long long running trend of gradually worsening performance (see graph from 10 to 14).

    The point of the post is basically a postmortem. It's a receipt that we aren't just imagining things.

  18. uNetworkingAB commented on Oct 4, 2026

    @uNetworkingAB

    The idea is a fast path for the common case: a top-level callback with no async hook, nobody using executionAsyncResource(), no ALS frame and nothing in the tick queue. There almost all the work of the callback scope changes nothing, so MakeCallback could check this once and do little more than the Function::Call. Three parts:

    Yes. That's how you solve it efficiently. This is how all efficient software is made; in a hierarchy of gradually decreasing probable execution paths. The most likely is the one where nothing needs to be done, it happens 99% of the time. Fixing that path yields massive gains for the most likely case.

  19. uNetworkingAB commented on Oct 4, 2026

    @uNetworkingAB

    Environment and context. On every call the Environment is found from the creation context of the function, and its context is entered. An addon that calls the same function many times (uWS calls one handler for every request) could resolve them once, with an opt-in handle. This is new public API.

    How would this interface look like? Also, probably good to bring in more than 1 library for reference / example. Do you know any other library that could use such an interface?

  20. nigrosimone commented on Oct 4, 2026

    @nigrosimone

    How would this interface look like?

    The smallest shape I measured is one more overload next to the current one:

    NODE_EXTERN v8::MaybeLocal<v8::Value> MakeCallback(Environment* env,
                                                       v8::Local<v8::Object> recv,
                                                       v8::Local<v8::Function> callback,
                                                       int argc,
                                                       v8::Local<v8::Value>* argv,
                                                       async_context asyncContext);

    The addon takes env once with the existing node::GetCurrentEnvironment(context), for example when it is loaded, and the callback must belong to that Environment. In the same runs it goes from 133 to 119 ns per call against the current overload.

    Do you know any other library that could use such an interface?

    Node-API addons do not need a new interface: napi_env already knows its Environment, so napi_make_callback can use it inside Node, and every addon on Node-API or node-addon-api gets it. NAN calls node::MakeCallback() for every Nan::MakeCallback and Nan::AsyncResource::runInAsyncScope, so NAN could switch to the overload where it exists, and every NAN addon gets it too, node-libpq (the native pg driver) for example. Inside Node, AsyncWrap already calls with its own Environment.

  21. nigrosimone commented on Oct 4, 2026

    @nigrosimone

    TL;DR: if everything below lands, on Linux x64 (one core, compare.js, 30 runs), per call:

    • napi_make_callback: from about 290 ns (Node 26.3) to about 170 ns
    • AsyncResource::MakeCallback: from about 270 ns to about 150 ns
    • node::MakeCallback: from 202-208 ns to about 110 ns

    A plain v8::Function::Call is 52 ns, so the part that is Node goes from about 150-240 ns to about 60-120 ns.

    Where we are

    Where we go

    Instead of a new issue I opened the rest as drafts, so the code can be discussed here:

    The final numbers are not yet measured all at once: they add up the parts, each one measured against the build before it. Feedback is welcome on the drafts, mostly on the shape of #66510 and on the new field in #66511.

  22. uNetworkingAB commented on Oct 4, 2026

    @uNetworkingAB

    Sounds like you have excellent control over this and already laid out very competent plans to cut node::MakeCallback overhead by 60% when we've heard for a decade now that "it's already as fast as it can be". It took you what? 3 days? Contrast is night and day.

  23. nigrosimone commented on Oct 4, 2026

    @nigrosimone

    For 10 years there was no AI :D

    Let me tell you what happened behind the scenes, because it can be interesting too. First, many of the solutions were already in this discussion, so I knew where to look. The problem, as you can see from the number of PRs, was more spread out than I expected and it touches different parts of Node.

    I left my PC on for days with frontier models running benchmarks of different possible solutions. Each one needed the machine free of noise, so they ran one after the other, not in parallel. Many ways turned out wrong, others did not convince me, others were too invasive. So my work was more a cut and sew of different good ideas, the ones that gave good results without changing too much, and what is in the PRs I went through line by line, with its benchmark and its tests.

    In practice it was like having an army of developers working 24 hours a day for about a week, while I decided what to try, what to measure and what to keep. For a human this work would not take a week, it is impossible: simply nobody would ever have done it, and maybe this is why it took 10 years. Without AI I would never have done it. It is sad, but it is like that. I am not a fan of AI, but here it did the hardest part.

  24. uNetworkingAB commented on Oct 4, 2026

    @uNetworkingAB

    If it's good, it's good.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    benchmark-neededAdd to issues that does not have a benchmarkhelp wantedExtra attention is needed

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions