Skip to content

reportVersion semantics are not defined #349

Description

@addaleax

The process.report/--experimental-report feature comes with output that contains a version number; however, it is unclear what that version number means and when it is incremented.

For nodejs/node#31386, it would be a good starting point to know if purely additive changes to the JSON format (i.e. only adding previously non-existent keys) should lead to version bumps.

@jasnell also suggested switching to semver, whereas @cjihrig pointed out that in the PR that added versioning, a single-integer versioning scheme was requested. I personally find it hard to make a decision on this question without knowing how consumers are supposed to interact with the version number.

I'm adding this to the WG agenda, I hope that's okay.

Activity

  1. richardlau commented on Jan 17, 2020

    @richardlau
    Member
  2. mhdawson commented on Jan 17, 2020

    @mhdawson
    Member

    For me I think it depends on whether it's ok to ask consumers to check the Node.js version as well. If a breaking change will be signalled by a bump in the Node.js major, then the version number as is might be ok with it being bumped every time something new is added to the report.

    Having said that it may be better to use SemVer and separate out from the Node.js version. For example, a new Node.js major may have no breaking changes in the report so using the Node.js version for when you may need a new version of tools processing reports is sub-optimal.

    The previous comparisons to N-API and ABI using a single number don't necessarily apply in my mind because ABI is always SemVer major and N-API is only SemVer minor so you only need 1 number to represent them.

  3. boneskull commented on Jan 23, 2020

    @boneskull
    Member

    as a tooling author consuming these—and I realize I may be the only one—semver seems like an overcomplication; more difficult to parse and of dubious value. if someone wants to argue for semver, I’m happy to listen about what it’ll enable that I’m failing to consider.

    if the report format changes in any way (new keys, changed, removed, or even moved since the order is deterministic; changed types of values; consider the entire tree) the version should be bumped.

  4. gireeshpunathil commented on Jan 23, 2020

    @gireeshpunathil
    Member

    I agree with @boneskull . Given that report is a diagnostic feature, IMO modification is all about more data, less data or structural changes, and the main party affected is tools. semver means adding more complexity here.

    pinging @june07 (who developed an inspector plugin that understands and renders report https://github.com/june07/NiM ) for an opinion

  5. june07 commented on Jan 23, 2020

    @june07

    @gireeshpunathil Thank you for including me in the discussion.

    While I see the benefit in simplicity vs complexity, I do think that semantic versioning would be the better choice. And given that as software developers we're all very familiar with semver, I'm not even sure I'd categorize using semver as adding to complexity, in fact I think it will help save us from complexity down the road. @boneskull has done a great job with rtk and in fact NiM and related tooling (BrakeCODE) use the rtk library, which demonstrates perfectly the sort of pipelines that do/may exist amongst different software, all which do/may interact with these reports.

    In example, JSON fields are currently added to reports that are ingested by BrakeCODE tooling (see screenshot) to encapsulate processing done by the rtk library. While this did not have to be included directly in the report schema, I think doing so is helpful as it eliminates introducing yet another parent data type. But, having an agreed upon standard (semver) of bumping report versions might make things nicer for other potential consumers of the data further down the pipeline, giving better transparency into data schema. NiM is a perfect example, because it is a consumer of reports from BrakeCODE. While adding keys should not equate to breaking changes, it might be nice to have a way to represent those additions via versioning.

    it would be a good starting point to know if purely additive changes to the JSON format (i.e. only adding previously non-existent keys) should lead to version bumps.

    No with single integer versioning but yes with semver.

    I think using a single integer version scheme might be somewhat limiting, non breaking changes being essentially invisible. @mhdawson brings out a good point regarding the Node.js version, having report versioning tightly coupled with Node.js versioning such that report versions are bumped in sync with Node.js is probably bad since changes in one don't necessarily indicate breaking changes in the other. Semver obviously solves this problem. And while more complicated, ultimately we’re all used to the semver paradigm and the complication really only lies in how semver will map to JSON schema changes vs code changes as we’re all very used to, not not much differently I assume.

    In working together to agree on the answer to...

    it is unclear what that version number means and when it is incremented.

    I feel that semver empowers one to answer that question in the most elegant way, while considering more than breaking changes.

  6. gireeshpunathil commented on Jan 24, 2020

    @gireeshpunathil
    Member

    thanks @june07 for the insights! Feel free to join our one of the upcoming diagnostic WG meetings when we take this up. the next one is going to be on 29th, but unlikely to pick this one up (due to pre-decided agenda, pinging @hekike for a confirmation. )

  7. boneskull commented on Jan 24, 2020

    @boneskull
    Member

    To be clear, my complexity concern is mainly around avoid string-parsing acrobatics or pulling in the userland semver package.

    I see semver is useful for developers who can choose what version of something to consume.

    For example, it enables automated upgrades, assuming the contract is upheld. Or you may know you don’t want to upgrade to a new major release, because it’s likely to contain breaking changes. It works pretty well most of the time!

    But a tool consuming a report file won’t be able to choose the version of said report file; the concept of an “upgrade” doesn’t really apply to the use case.

    A tool will change its behavior based on the version of the report, however. Regardless of what a semver version “means” in a report (e.g., minor version means “new field”), the behavior must still be defined based on information that semver cannot provide.

    Example:

    If we add a report field called foo in a hypothetical version 2.1.0, a tool will read the version of the report file, and if it’s >=2.1.0 the tool must do whatever it’s going to do with foo. If it is <2.1.0 it will not. Note that the semantic version is unable to tell the tool that foo was added. 😉

    If v2.1.1 is a bug fix to field bar, a very similar decision will need to be made—assuming the tool cares about the fix—but this time, based on >=2.1.1 / <2.1.1.

    Likewise with a breaking change in v3.0.0. Note that the semver conventions ~ and ^ are not currently applicable, and will not be until we have a situation in which the same behavior is broken more than once between different major releases.

    OTOH, if you are using incrementing integers for versions, the decision-making is exactly the same as above example... except a tool doesn’t have to parse semver!

    Another problem—this is not technical, but of the “people” variety—is that if a tooling author might see semver and think it’s generally okay to restrict the tool’s operation based on the major release number—because that’s how people use semver! For example, only supporting ^2.0.0, when in fact nobody knows if a v3 report will work or not.

    In summary, SemVer was designed to give people better control of automated software upgrades. I just can’t see how our use case fits with that aim.

  8. jasnell commented on Jan 25, 2020

    @jasnell
    Member

    A simplified two digit semver approach works also here. Just major and minor/patch.

  9. richardlau commented on Jan 29, 2020

    @richardlau
    Member

    Not too keen on a two digit approach. Single digit as @boneskull points out is easier for tools to parse. If we need more digits then semver at least is well defined and has existing parsers that could be reused. If tools were doing their own parsing they could just ignore the patch level (since by definition changes to the patch level should not break them).

  10. boneskull commented on Jan 29, 2020

    @boneskull
    Member

    @jasnell From your point-of-view, what are the advantages of using semver here?

  11. gireeshpunathil commented on Feb 17, 2020

    @gireeshpunathil
    Member

    We deliberated in the last(-to-last) WG meeting on this.

    Report being a data (as opposed to code), patch does not make sense anyways. So the contention is really between x.y semantics versus x semantics.

    I propose a single digit versioning with version bump on every structural changes wherein structural change is anything that:

    • adds a new key
    • removes a key (though chances are less for this to happen)

    It is reasonable to expect report parsers to parse it in JSON-native way, instead of line-parsing, char-parsing etc. One of the main motivation for the report to be in JSON format was easy-parsing for consuming tools. Because of this, section movements / aesthetic modifications do not cause a version bump.

    It is reasonable to expect no data type changes being applied to the fields.

  12. mhdawson commented on Feb 17, 2020

    @mhdawson
    Member

    It is reasonable to expect no data type changes being applied to the fields.

    Can we depend on that? Maybe we should add a point to your list that says:

    • changes the data type for a field
  13. gireeshpunathil commented on Feb 17, 2020

    @gireeshpunathil
    Member

    makes sense @mhdawson . so the modified proposal is:

    a single digit versioning with version bump on every structural changes wherein structural change is anything that:

    • adds a new key
    • removes a key
    • changes the data type for a field
  14. legendecas commented on Feb 17, 2020

    @legendecas
    Member

    It is reasonable to expect no data type changes being applied to the fields.

    Will a format change of string types be counted as a significant change? Like the format of error stacks.

  15. 19 remaining items

  16. gireeshpunathil commented on Aug 23, 2022

    @gireeshpunathil
    Member

    @RafaelGSS - you asked, why can't we attach the containing Node.js version for the report version too? .

    here is the issue with that: if there are more than one structural changes in the report that a Node.js version carry, then we cannot differentiate between those.

  17. No9 commented on Aug 24, 2022

    @No9
    Member

    As agreed in the diagnostics meeting I said I would take a look.

    To me #349 (comment) from Chris is a good summary.
    Specifically the point on there being no schema in the current implementation making the conversation somewhat academic.

    With that I'm +1 on the suggestion from @gireeshpunathil to move this forward with a single number as it reflects the current implementation and can be revisited.

  18. gireeshpunathil commented on Oct 21, 2022

    @gireeshpunathil
    Member

    this is done via nodejs/node#45050 , closing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions