Skip to content

Diagnostics Channel - Next steps #180

Description

@mhdawson

The concept of a diagnostic channel was discussed in the Diagnostics summit as well as here: #134

In a nutshell the concept is to provide an API that modules can use to publish events, and an API that consumers (such as APM vendors) can consume. The key differences between the channel provided by these APIs and trace events are:

  • operation is synchronous, the consumer has the ability to take an action before the publish event completes on the emitter side.
  • data is passed (and any other data available indirectly) by the publish and this data can be mutated by the consumer.

This approach requires both implementation of the diagnostic channel as well as updates to modules in order to use the API. An interim step where monkey patching is used to add calls to the API is seen as a good way to bootstrap.

Microsoft already has an experimental implementation which we should evaluate to see what parts (if any) that we can re-use.

An initial cut at next steps include:

  • Create a new repo where we can experiment with a new implementation under the nodejs org - diagnostic-channel
  • review the existing experimental implementation
  • create performance regression suite
  • define the diagnostic publish API
  • define the diagnostic consumer API
  • pull what parts we can from existing implementation into publish API
  • pull what parts we can from existing implementation into consumer API
  • Iterate until we hit a 0.1.0 release that is ready for broader review

Activity

  1. mcollina commented on Apr 3, 2018

    @mcollina
    SponsorMember

    Overhead and performance should be a first level concern about this API. It should result in no-ops if it is not enabled, and get smart if it is enabled. I think we should trade usability for performance whenever possible, as long as it still satisfy our use case.

  2. hybrist commented on Apr 3, 2018

    @hybrist
    Contributor

    @mcollina - [ ] Create performance regression suite?

  3. mcollina commented on Apr 3, 2018

    @mcollina
    SponsorMember

    @jkrems 👍 !

  4. hybrist commented on Apr 3, 2018

    @hybrist
    Contributor

    Added to the top post

  5. Qard commented on Apr 4, 2018

    @Qard
    Member

    I brought it up in the diag meeting just now, but I worked a bit on a proof-of-concept universal diagnostics channel type thing recently, after I got back from the Ottawa summit and then Elastic{ON}. It's a very rough demo and definitely needs more perf focus before it could be actually usable, but you can have a look at it here: https://github.com/Qard/universal-monitoring-bus

    It uses subscription queries to express interest in certain messages in the channel, similar to trace_events, but in a currently more flexible way (sacrificing performance for flexibility while trying to identify interesting filter data). It has no direct graph-forming, instead opting to direct the message bus into aggregators, which can live in-process or out-of-process, and makes no assumptions about transport mechanism. It is simply a callback that receives JSON-serialization-ready objects that you can do whatever you like with.

    It also expresses both spans and metrics updates uniformly in a single message format, with purpose-specific extension, allowing for an easy shared communication channel or aggregator re-use and even correlation between the spans and metrics data.

  6. Qard commented on Aug 5, 2020

    @Qard
    Member

    I'm picking this back up now and working on some ideas of how to reduce/eliminate the runtime overhead. My main concern is that the prototype is basically just an event emitter so it has to lookup the subscribers on every publish. It does have a shouldPublish method to avoid constructing input data if there's nothing listening, but it's basically functionally identical to the listeners method of an event emitter which just results in an additional lookup. And approach I've been looking at is named channel objects that can be acquired once and reused as a way to make the lookup only happen once at the time of trying to acquire the channel.

    I also don't think the module patches and context management aspects should be included. For core adoption, we should stick strictly to the data firehose component and strip that down to the bare essentials for something speedy and highly extendable in userland.

  7. Qard commented on Aug 9, 2020

    @Qard
    Member

    Here's my prototyping progress this week: https://github.com/proxy/gist.github.com/Qard/b21bb8cb1be03274f4a8c585c659adf2

    I tried to follow the original design, for the most part, but added regex subscribers and the concept of channel and subscriber objects to take the lookup cost out-of-band from the publish so it doesn't have to do the lookup every time. The result is publish overhead on-par with a single variable increment, which is pretty good!

    What do people think? Is the simple text name a flexible enough way to search for events in the firehose? Another possibility is object matching, which I did in this prototype several years back. I think it's a bit friendlier and more flexible, but I couldn't figure out a good way to pull the object-style search out-of-band from the messaging like I did for the prototype above. 🤔

  8. mmarchini commented on Aug 12, 2020

    @mmarchini
    Contributor

    Some questions from today's meeting:

    1. Any plans on supporting publishers in C++ (not as a public API necessarily, but as a mechanism to instrument C++ code in core)
    2. there's publish and publishIfActive. IMO it would be better ergonomics for publish to, by default, only publish if active, and then have a forcePublish or something to publish even if inactive.
    3. I wonder if we should have a concept of "parent publisher" to avoid, so that publishing to a child would also publish to the parent. For example, have fs as a parent and fs.readFile as a child publisher, publishing to fs.readFile also publishes to fs (and following 2, it only publishes to active publishers by default, so if fs is active and fs.readFile is not, it would only publish to fs
    4. How does it compare to the existing trace_events module? Could it be a functionality of trace_events instead of a new module?
    5. Based on 4 and assuming we have it separate from trace_events, do we want to have publisher instrumentation on the same places we have trace_events? How to we keep those in sync in that case? How do we decide what to instrument and how do we measure progress of instrumentation in core?
    6. Mentioned in the meeting and on Slack, but just reinforcing I think starting with an userland module published to npm is the way to go to experiment with this until we need something core-specific or until the API is closer to be solidified.
    7. On current implementation subscribers receive the messages syncronously. Should we async subscribers as an option or even by default?
      a) Any chance we can provide a channel that allows publisher and subscribers to be in different isolates (so we could have a worker thread as subscriber, for example)?

    Overall I think it's a great start for the API and it would be good to get input from potential consumers (APMs) and publishers (frameworks) before attempting to land it in core.

  9. Qard commented on Aug 12, 2020

    @Qard
    Member
    1. Not sure about C++ yet. I haven't figured out if there's anywhere specific we'd want to capture C++-side data and, if so, how we could do that in a reasonably fast way.
    2. The reason for publishIfActive(...) is to defer the input data construction to only happen if the channel is active. It's the same as if (channel.shouldPublish()) channel.publish(data) though, so possibly better to tell users to use that. At the moment it only really exists as an alternative API to compare performance and ergonomics. I do think we should settle on one way to do it before landing though.
    3. That complicates the publish logic somewhat and makes it more expensive to decide if there is work to do at runtime. It'd have to walk the tree upward for every publish because the end of the branch might be inactive but other nodes in the graph might be active. It also imposes some structural opinions on the shape a channel can take from the perspective of the subscriber whereas I'd prefer the publisher to be as simple as possible and offload complicated/expensive logic as much as possible to the subscriber and hopefully out-of-band as much as possible. Performance when not listening is a top priority here.
    4. The trace_events API is not consumable within JavaScript so not really useful for production monitoring, it's more of a post-mortem analysis tool. It's also target more specifically at chrome://tracing. It's not really suitable for broadcasting arbitrary/unstructured data.
    5. There may be a few places where we broadcast data for both systems but I don't see there being a lot of overlap.
    6. I disagree about making it a userland module. Most of the data that we (APM vendors) would want to gather from it is from core APIs. Starting as a userland module would just make it even harder to get it into core. (as evidenced by it taking literally half a decade for Node.js to get an async context management API despite APMs asking for it repeatedly for years. 😬 )
    7. At the moment, having it sync is intentional. Timestamps can be added to messages on the consuming end that way. If we move it to async we lose the timing certainly and would likely need to do more within the API--possible, but trying to avoid that unless it's really necessary. As for the worker thread idea, I've thought about it. At the moment I'm looking more at the ergonomics and concepts of the API than the specific implementation details like that.

    There's a lot to consider in how the ecosystem would want to engage with such an API to be able to deliver something that is satisfying, especially given our increasing focus on long-term vision of the project. We need to make sure this is not just relevant now, but will remain relevant and useful for the foreseeable future.

    I have a few months to hack on this so there's plenty of room for iteration. We'll see how successful I am in getting this to a place where people are happy with it and it can deliver on the promise of a high-performance unified data firehose. 😅

  10. Flarna commented on Aug 14, 2020

    @Flarna
    Member

    Sorry for not participating the meetings but the meeting times we have are in general hard to catch for me.

    Regarding (7): I think it's vital that this is sync to allow subscribers to correlate with their current state/context. Just getting the info that e.g. a DB query XXX is sent is not that helpful on its own. APMs correlate events to show e.g. that query X happened during processing request Y. As X and Y in general originate from different modules the events itself can't hold this information.

    I'm not sure if we need the regex subscribers at all. Published data is only of use if subscriber know the structure of the content for each event. Therefore I think subscribers are quite specific and using them for a possibly unbounded amount of publishers sounds not that reasonable. Maybe for something like a log/batching subscriber which just pushes the data as is or stringified further to the next consumer.

  11. RafaelGSS commented on Aug 14, 2020

    @RafaelGSS
    Member

    Let me share my opinions too:

    1. I agree with @Flarna about the regex.
    2. I would like to see this prototype in action taking by example a database_module just to validate and see the gains with it. I would like to help you with this if you agree 😄.

    And I have a question too:

    1. Do you have some idea of what kind of information from node-core we would like to listen/publish? I see this feature as a good idea to all diagnostics on the node, like cpu monitor, memory...
  12. 6 remaining items

  13. Flarna commented on Aug 14, 2020

    @Flarna
    Member

    ok so the main thread publishes data; either via some sensors placed by the APM tool or by modules having this built in. This data is collected in a worker, processed and transmitted.
    Sounds fine so far.
    Thinking on a generic DiagnosticsChannel I still question the async publish to a worker as one tool may run in one worker and another tool in a different worker. So somehow the publish target needs to be part of the subscriber.

  14. mmarchini commented on Aug 14, 2020

    @mmarchini
    Contributor

    fwiw Timestamps can be added to messages on the consuming end that way was enough to convince me that sync in the main thread for the API we expose is the best approach (I was thinking timestamps coming from the publisher, but I guess it makes more sense on the subscriber) :)

  15. Qard commented on Aug 14, 2020

    @Qard
    Member

    Yeah, it's my intuition that it's better to just keep it in the main thread and encourage APMs themselves to dispatch that to their own threaded thing. Timestamps are an easy fix with a wrapper object, but context management gets harder, especially with multi-tenancy.

  16. Qard commented on Aug 23, 2020

    @Qard
    Member

    Pull request opened! 🎉 nodejs/node#34895

  17. github-actions commented on Nov 22, 2020

    @github-actions

    This issue is stale because it has been open many days with no activity. It will be closed soon unless the stale label is removed or a comment is made.

  18. mhdawson commented on Nov 23, 2020

    @mhdawson
    MemberAuthor

    Removing stale as @quad is working on it.

  19. github-actions commented on Feb 23, 2021

    @github-actions

    This issue is stale because it has been open many days with no activity. It will be closed soon unless the stale label is removed or a comment is made.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions