Skip to content

Diagnostics "Best Practices" Guide? #211

Description

@mike-kaufman

Something that would be interesting & valuable to the community is if the WG could come up with a set of diagnostics best practices/techniques for production applications. I'm a little concerned that this is a bit too ambitious, and a little concerned that one person's "best practice" is another's "really dumb thing to do". But I'm willing to throw this out to see what comes out of it. :)

I have a few thoughts about how this could be structured, but would love to hear if anyone else has ideas/suggestions first.

Activity

  1. joyeecheung commented on Jul 10, 2018

    @joyeecheung
    Member

    It would be great if we can come up with even just a guide with links to good resources and advertise it so people know what to read when they run into different types of problems in production. We always receive bug reports with obscure screenshots of resource usages in the core repo, documentations like that would be a good place to redirect to instead of nodejs/help.

  2. mhdawson commented on Jul 10, 2018

    @mhdawson
    Member

    It would be great if we can put this together

  3. mike-kaufman commented on Jul 17, 2018

    @mike-kaufman
    ContributorAuthor

    into different types of problems in production. We always receive bug reports with obscure screenshots of resource usages in the core repo,

    @joyeecheung, can you point me to some examples here?

  4. joyeecheung commented on Jul 18, 2018

    @joyeecheung
    Member

    @mike-kaufman There should be a lot of hits if you search for the issues labeled memory: https://github.com/nodejs/node/issues?utf8=%E2%9C%93&q=is%3Aissue+label%3Amemory+

  5. gireeshpunathil commented on Jul 18, 2018

    @gireeshpunathil
    Member

    I love this idea, support this, and willing to participate / contribute in ways this group needs it and in ways I can.

    While I agree that best practices can be subjective, there are elements of diagnostic steps that are impersonal and to the point - an example would be collecting and analysing heapdumps on memory leak.

    Regarding the structure, I see we can have different flows such as:

    • organized based on symptoms [ memory, exception, performance, hang, crash ... ]
    • organized based on tools [ tracer, profiler, dumper, inspector, debugger, reporter ... ]
    • organized based on subsystems [ installer, net, fs, child_process, natives, uv, ... ]

    I suggest the symptom based categorization as it leads to faster discovery for consumers.

    Rgarding the best practice content, again I see few models:

    • dump of complete diagnostic steps on a failing case, with example code
    • screen-shot assisted illustration of diagnostic methodology
    • plain text elaboration of tools and their usage

    I suggest the first model as it leads to improved education for consumers.

  6. mike-kaufman commented on Jul 18, 2018

    @mike-kaufman
    ContributorAuthor

    @joyeecheung, @gireeshpunathil thanks. Let me strawman an outline and see if this matches what's in your head - feel free to tear it down. :)

    • Production Configuration Best Practices
      • application logging & log management
        • breakouts for different app types
          • server apps
          • serverless
          • desktop
          • IOT
      • APMs
      • Node Internals Tracing
        • configuration
        • interpretation
    • Troubleshooting
      • Profiling Perf Problems
      • Analyzing Memory Leaks
      • Analyzing Core Dumps
  7. gireeshpunathil commented on Jul 19, 2018

    @gireeshpunathil
    Member

    thanks @mike-kaufman . While the troubleshooting part is straightforward for me to relate, the first part (production configuration best practices) looks very wider in scope to me:

    • is it possible to suggest configurations for such wide deployment scenarios?
    • even with say server apps, is it possible to generalize configurations?
    • is it possible to propose configurations independent of execution environment of such app types?
  8. mmarchini commented on Jul 25, 2018

    @mmarchini
    Contributor

    I gave a talk on NodeSummit about this topic, where I showed 6 tools suited for production environments. Even though the topic is subjective, I don’t think there’s much disagreement on which tools or techniques should be used (our current pool of production tools is not that large).

    The tools I showed alongside the examples are available here if anyone is interested. I would like to help to write these guides :)

  9. gireeshpunathil commented on Jul 26, 2018

    @gireeshpunathil
    Member

    I had a discussion with @mike-kaufman and the consensus was to start with a draft and iterate over PRs and refine through collective intelligence.

    So let us start with say Profiling perf problems and then others can follow the structure. I will start looking at dump debugging.

  10. mike-kaufman commented on Jul 26, 2018

    @mike-kaufman
    ContributorAuthor

    @mmarchini - that's a great start :)

    Thanks Gireesh. I'd like to ultimately get content the written to leverage github's auto-html-site feature - i.e., we'd be able to submit markdown updates to this repo, and it will be automatically renedered to html available at http://nodejs.github.io/diagnostics/bestPractices/.

  11. mmarchini commented on Jul 26, 2018

    @mmarchini
    Contributor

    @mike-kaufman auto-generating a website is an awesome idea! But maybe we should try to coordinate with @nodejs/website to have this content available in https://nodejs.org/ as well?

  12. mike-kaufman commented on Jul 26, 2018

    @mike-kaufman
    ContributorAuthor

    @nodejs/website to have this content available in https://nodejs.org/ as well?

    Yes, this would be good. I think as the first step though, we can start getting the content organized, and the github.io "auto-magic-web-site" is a really simple & cheap way for us to get that content rendered & reviewable, w/out the distraction of how we plug into their process.

    @bnb - is there any thinking on how we can plug in content to new website? Ideally, we'd have a bunch of markdown & images here (in diag repo), and this would just get "sucked up" into the website.

  13. fhemberger commented on Jul 27, 2018

    @fhemberger

    Please get in touch with @nodejs/website-redesign, as we are in the middle of planning the content structure for the website relaunch:

    https://github.com/nodejs/website-redesign/issues/

  14. misterdjules commented on Aug 15, 2018

    @misterdjules

    For what it's worth, a long time ago I had written a set of guides to investigate various types of production issues with Node.js. @cjihrig kindly made those guides available publicly at https://github.com/joyent/node-debugging-methodologies. The content is mostly specific to SmartOS but the methodologies/concepts can almost always easily be transferred to other OSes.

  15. mike-kaufman commented on Aug 15, 2018

    @mike-kaufman
    ContributorAuthor

    @misterdjules - thanks, this is great!

  16. 10 remaining items

  17. mhdawson commented on Nov 20, 2018

    @mhdawson
    Member

    @gireeshpunathil I think we should spin up a doodle, with a deadline to complete by end of this week and then choose meeting for next week.

  18. mhdawson commented on Nov 20, 2018

    @mhdawson
    Member

    I suggest opening a new issue for the meeting to make it more visible, and to put the doodle there.

    If we still have a lower number of responses then I think we just need to move forward with whoever responds.

  19. amiller-gh commented on Nov 20, 2018

    @amiller-gh
    Member

    Let me know if I should make a point to come to this 🙂 Otherwise, just sending a friendly reminder that I'd love to see at least one deliverable be a PR adding content to our future website docs, like what Flavio has done with getting started content here: https://github.com/nodejs/website-redesign/pull/105/files

    Excited to see this happen!

  20. goldbergyoni commented on Nov 21, 2018

    @goldbergyoni

    here we work on our 'performance & diagnostics' best practices section:
    goldbergyoni/nodebestpractices#256

    I'll be glad to join our forces

  21. gireeshpunathil commented on Nov 21, 2018

    @gireeshpunathil
    Member

    a separate issue is spawned to track upcoming meetings to discuss this. #254

  22. gireeshpunathil commented on Apr 18, 2020

    @gireeshpunathil
    Member

    removed from wg meeting agenda, as per discussed in the last meeting (rationale: the work is being progressed as part of uesr journey deep dives and subsequent documentation work. If doc work stalls, we could always re-insert this to gain focus)

  23. mmarchini commented on Apr 28, 2020

    @mmarchini
    Contributor

    As a reminder, tomorrow we're meeting (same time as always) to discuss diagnostics on CPU usage.

  24. github-actions commented on Jul 28, 2020

    @github-actions

    This issue is stale because it has been open many days with no activity. It will be closed soon unless the stale label is removed or a comment is made.

  25. github-actions commented on Jul 27, 2022

    @github-actions

    This issue is stale because it has been open many days with no activity. It will be closed soon unless the stale label is removed or a comment is made.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions