Skip to content
This repository was archived by the owner on Sep 2, 2023. It is now read-only.
This repository was archived by the owner on Sep 2, 2023. It is now read-only.

Tooling: Brainstorming ideas that can lead to efficient loader-oriented designs #203

Description

@SMotaal

Having both ecmascript-modules and @jkrems hackable loader has opened up tremendous scope for experimentation.

Note: This thread does not make claims for or against existing tooling, some of which have stood the test of time, evolved, and are fixtures of the ecosystem. The intent is simply to consider different perspectives being explored in experimental efforts.

As far as things go, the broad range of tooling that applies to loaders basically iterates over productions in each source, irrespective of the specifics of implementation or operations.

Most tools are designed to be used for much more complex applications than merely loading. To that effect, they often avoid the use of new language features that would prevent them from working on older platforms. They can also avoid new features which may have been prematurely associated with inefficiencies in early stages. Some are also built with infrastructures or features that are not ideal or not optimized specifically for loading, like using workers, verbose error checking (ie as a language service)... etc.

I would like to dedicate this thread to brainstorming experimental or just different ideas to implement related patterns for loader-first designs.

Brainstorming: A safe place to discuss ideas and provide constructive feedback

How to contribute

Please avoid emoting that can be confusing (especially if it can construed as passively aggressive)

😄 Indication
👍 To indicate a "Yes" response
👎 To indicate a "No" response
🎉 To indicate a "Aha" moment

Read the Digest

The following is a set of ideas or conclusions curated from the discussions:

Syntax Detection (CJS vs ESM)

  • Safely using RegExp — @SMotaal

    • requires guarding against string hijacking — @devsnek
    • recommend using acorn instead — @devsnek
  • Fallback for ESM without import and export — @targos

    • shouldn't use import(…) to resolve ambiguity — @bmeck
    • can use import.meta — @bmeck
  • Dual parsing a module was deemed inefficient — @MylesBorins

Syntax Identification (CJS vs ESM)

  • Mime type meta data via something like webpackage — @jkrems

  • Magic bytes — @jkrems

Wrapping CJS in an ESM module system

Activity

  1. SMotaal commented on Oct 17, 2018

    @SMotaal
    Author

    ECMAScript modules syntax can arguably be detected using a RegExp which bails on first match.

    Does anyone have ideas for cjs vs esm syntax detection?

  2. changed the title [-]Tooling: Using new language features to design efficient loader extensions[/-] [+]Tooling: Using new language features to design efficient loader-first extensions[/+] on Oct 17, 2018
  3. devsnek commented on Oct 17, 2018

    @devsnek
    Member

    @SMotaal you can't use regexp to parse js grammar (you can always make a pattern of string literals or whatever to confuse the regexp) and the differences between valid cjs and valid esm are ambiguous and can't be reliably detected by just looking at the code.

  4. added
    brainstormingSafe place to discuss ideas and provide constructive feedback
    on Oct 17, 2018
  5. SMotaal commented on Oct 17, 2018

    @SMotaal
    Author

    you can always make a pattern of string literals or whatever to confuse the regexp

    So, can we constructively say that so long as you guard against string hijacking (maybe there is a better term for this), only then can you safely use RegExp?

  6. devsnek commented on Oct 17, 2018

    @devsnek
    Member

    @SMotaal I would just use acorn

  7. SMotaal commented on Oct 17, 2018

    @SMotaal
    Author
  8. SMotaal commented on Oct 17, 2018

    @SMotaal
    Author

    @devsnek humor me in this effort, consider this both an idea-gathering as well as a team-building exercise. Acron is obviously a great solution, but I am trying to create opportunities for people to talk about the aspects that make this and others such great tools. The notion here is that people might just have some evolving ideas that they might want to bounce around. How we connect the dots, like you pointing out the hijacking limitation can potentially inspire untapped solutions to existing problems.

    Sounds fair?

  9. targos commented on Oct 17, 2018

    @targos
    Member

    @SMotaal You could say that a file with import or export syntax is probably an ES Module (the syntax is invalid in Script mode). However, the problem is that files without import and export could be either Script or Module, and depending on how they are written, could have different behaviour in Script vs Module mode.

    For example:

    test = 42;

    In Script mode, this creates the property test on the global object.
    In Module mode, this throws a ReferenceError.

  10. benjamingr commented on Oct 17, 2018

    @benjamingr
    Member

    @targos does the issue get any better if we say that such a loader always imports CJS in strict mode regardless of an explicit "use strict"?

  11. devsnek commented on Oct 17, 2018

    @devsnek
    Member

    are we trying to come up with use cases for loaders or something else?

    if you're using a resolve loader hook you'll always be able to read the contents of whatever you're resolving, at which point you can regex or acorn or whatever it as you see fit.

  12. targos commented on Oct 17, 2018

    @targos
    Member

    I'm having trouble to see the relation between 'cjs vs esm syntax detection" and the OP. Maybe I don't really understand what this thread is about, sorry.

  13. 70 remaining items

  14. SMotaal commented on Oct 25, 2018

    @SMotaal
    Author

    ESX parsing currently scans the full length of source text, but the intent is to actually keep reference of enclosing ranges and not analyze them unless there are no signals of ESM syntax on top-level, then finally scan enclosures to find the first cjs hint or not, this makes it possible to report ESM, CJS, or still ambiguous so use the default based on out-of-band settings... etc.

  15. ljharb commented on Oct 25, 2018

    @ljharb
    SponsorMember

    If we have out of band settings, and that info conflicts with a parsed result, I’d expect it to throw - the two shouldn’t be in disagreement.

  16. SMotaal commented on Oct 26, 2018

    @SMotaal
    Author

    That’s actually a very important aspect, because I in my rushed vision of eliminating parsing errors which are handled normally by the runtime, I have not given thought to certain errors that belong specifically to the intent at hand.

    More of this kind of insights here can go a long way down the road when making decisions. Awesome 🙂

  17. SMotaal commented on Oct 26, 2018

    @SMotaal
    Author
  18. SMotaal commented on Oct 26, 2018

    @SMotaal
    Author

    When considering the case of parsing, I was having trouble mentally placing the metadata communicated between two loaders for instance.

    In this case, it is in-band (imo) but it is not "directly" from source, it is inferred and attributed to the source text, and is triggered (or bypassed) and responds to out-of-band (one-to-many) and out-of-source (one-to-one) aspects or settings.

    Can I propose the following complementary pairs: (examples in brackets)

    1. "out-of-band" — setting that trickles down to one or more resolved specifiers (flag, ext, mime…)
    2. "in-band" — settings determined from resolved source features (pragma, this parse…).
    1. "from-source" — settings declared in the source text (pragma, shebang…)
    2. "out-of-source" — settings inferred or attributed to a source text (some in- and out-of-band)

    Can anyone find a more practical breakdown of such information regarding a source text's journey?

    This is all crude thoughts, it needs magic from the group. I feel that a distinction between what maps to sources and what is specific to a source but not baked right into the body are essential distinctions.

  19. SMotaal commented on Oct 28, 2018

    @SMotaal
    Author

    I finally updated the README and pushed the revisions made last week. Timing is more accurate now. I also converted the rendering pipeline to async APIs. Tokenization APIs remain sync but use generators so they yield and return as needed. I improved the modes for esm, cjs, esx, and added the missing alias es for the regular javascript syntax mode.

    I am really interested to hear some feedback on the three modes (esm, cjs, esx) with various sources, especially if you find a source that breaks or chokes in one of those modes.

  20. devsnek commented on Oct 28, 2018

    @devsnek
    Member

    @SMotaal its cool i guess? i don't really understand why we have an issue open for it though.

  21. SMotaal commented on Oct 28, 2018

    @SMotaal
    Author

    @devsnek This thread is about ideas in general separate from implementation. As we move closer to loaders and defaults, those discussions and demos can be helpful, at the very least, they can serve as a reference for those who need to find more about them.

  22. SMotaal commented on Nov 5, 2018

    @SMotaal
    Author

    @jdalton Can you pitch in on the idea of syntax detection relating to top-level parse. I tried to find a way to model this to the benefit of everyone in the group and was able to show a 200% increase in performance (theoretical) relative to the same method to full ES grammar parsing like ASTs would.

    This was done avoiding the conventional all-or-nothing AST approach, using half-way optimized RegExps addressing usual concerns like hijacking.

    Ideas like dual-parsing (@MylesBorins) and your top-level parse (@bmeck) made me think of a single-parse limited to the minimal subset of both grammars and it was roughly capped at 175% depending on nested complexity but on average better than 150%.

    Since we're trying to find the first clue to determine syntax, the expectation is that such clues will often materialize early on in a text, making it reasonable to bail or delay the rest of the parsing (if at all needed).

    Can we hash out pseudo code for syntax determination based on your initial thoughts on top-level parse?


    About this thread…

    I'm trying to brainstorm ideas parallel to our implementation efforts that make it possible for our broadly diverse members to appreciate the various technical challenges associated with decisions we are making.

    Based on an early digest of this discussion, which I took liberty to summarize at top. I tried to pick ideas which seemed to create rifts in discussions elsewhere, mainly in on the topics of syntax detection and interoperability.

  23. GeoffreyBooth commented on Nov 5, 2018

    @GeoffreyBooth
    Member

    @SMotaal This is impressive . . . just to understand what you’ve done here, is your goal to determine parse goal by analyzing the syntax? A.k.a. a real implementation of the “unambiguous syntax”/grammar that we’ve been discussing?

    If so, and assuming that you find an algorithm that works, have you thought about how to address the related concerns listed in #150 (comment)?

  24. devsnek commented on Nov 5, 2018

    @devsnek
    Member

    to be clear, it's just a lightweight way of parsing js. this doesn't make the ambiguity go away.

  25. ljharb commented on Nov 5, 2018

    @ljharb
    SponsorMember

    Confirmed; there does not exist any approach based on parsing that is unambiguous in all cases, absent a language spec change.

  26. SMotaal commented on Nov 5, 2018

    @SMotaal
    Author

    Yeah, while I would love to be the one that can solve ambiguity of source text and other sources, this is really nothing more than a very modest effort to model different parsing methods separate from the usual tools.

    My gut feeling tells me that while implementing solutions is best served by employing tried and tested tools, coming up with optimal solutions may not always share in those benefits. So in other words, AST's have a way about them that force looking at problems in certain ways, so modeling the problem without is a way to avoid restricting ourselves to the givens of using them.

    So this is far from a solution, just an attempt to provide a way to explore solutions, and the bottom line holds, ambiguity is ultimately a source problem, and if it is, then the only way to resolve it is out of band.

  27. SMotaal commented on Nov 6, 2018

    @SMotaal
    Author

    @devsnek the underlying motivation behind my markup experiment in general is not restricted to JS, in fact, I was interested to find different ways for efficient and responsive multi-syntax parsing without the pitfalls of conventional methods. And on that, I think I am ready to dare make the claim that it can be done with virtually no switching overhead, using less popular features like generators and regexps: html (and script tags)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions