Repository navigation
URL's on incoming requests #71
Description
Activity
I'm highly -1 on using
new URL()forreq.url. It's extremely slow and inefficient due to the host and protocol parsing.I usually use, https://www.npmjs.com/package/request-target. Maybe worth looking into?
https://github.com/nodejs/next-10/blob/master/VALUES_AND_PRIORITIZAION.md might also be worth considering, i.e. "Performance" vs "Web API compatibility".
Reacted by Matteo Collinacould we do something like
req.urlreturns the faster solution and thenreq.url()would return thenew URL()version? This way users have to opt-in to the performance decrease if they want it?Additionally, could we look into improving
host and protocol parsing? Or is has that already been attempted?Reacted by Steven and Owen BuckleyThis is why I was strongly questioning adding that (nodejs/node#35323) in yesterdays meeting.
The perf is a big concern I was not aware of (do we have benchmarks on this, or details on why?), but is also a good reason to push back on the web api compat technical value issue.
do we have benchmarks on this, or details on why and what parts are slow?
From https://www.npmjs.com/package/request-target
$ npm run benchmark legacy url.parse() x 371,681 ops/sec ±0.88% (297996 samples) whatwg new URL() x 58,766 ops/sec ±0.3% (118234 samples) request-target x 552,748 ops/sec ±0.54% (344809 samples)I hadn't clicked your link, but yeah this looks good. I will dig into what it is doing, thanks for the link!
Reacted by Robert NagyWell I very quickly ended on this line, and I am fairly confident we will be able to find a better (and even better perf) way than a big regex.
But this leads me to my next question: do we really want to deviate from the existing URL apis?
Seems like this would add "yet another url" to node core, right? Maybe we could get fast parsing but also expose an api which looks like a
URLinstance, but this gets us into a similar situation of "its a stream but not". Anyway, just thinking out loud on this.It's extremely slow and inefficient due to the host and protocol parsing.
Is the slow performance of
new URL()inherent to the way its spec'ed or just the Node.js implementation?If its the latter, then this would be a good excuse to re-evaluate the implementation to squeeze out better performance.
Overall, I think providing a way to retrieve
req.urlas a URL would be a great addition to Node core because the URL object is the future,url.parse()is a legacy of the past, before there was a standard.Using standard JS
URLobjects allows code to be shared between the browser, Node.js, etc.One of the big foot-guns today is parsing url query string params. Consider this code:
const { parse } = require('url'); const handler = (req, res) => { let { pathname = '/', query = {} } = parse(req.url || '', true); const { user = '' } = query; console.log(user.toLowerCase()); }
This works fine with
https://example.com/api?user=foobut throws when two parameters with the same name are detectedhttps://example.com/api?user=foo&user=barbecauseusersuddenly becomes an array instead of a string.Consider this similar code using URL:
const handler = (req, res) => { const url = new URL(req.url, 'https://example.com'); const user = url.searchParams.get('user') || ''; console.log(user.toLowerCase()); }
This bug doesn't happen with URL because it separates
getAll()vsget(). Now we won't crash in production.My workaround for the time being looks like this:
export function getURL(req) { const proto = req.headers['x-forwarded-proto'] || 'https'; const host = req.headers['x-forwarded-host'] || req.headers.host || 'example.com'; return new URL(req.url || '/', `${proto}://${host}`); }
Reacted by Owen BuckleyIt makes sense to me to offer both (obviously under different names) - ie,
req.urlas the string, andreq.getURL()as the lazily-computed-but-memoized URL object.Reacted by Steven, Ethan Arrowood and Owen BuckleyReacted by StevenIt makes sense to me to offer both
I think before we make any decisions on this, we need to know if we can help fix the perf issues of
URL. If we cannot do that, then I think we should consider alternate api's to expose both. If not, I think one api is best.My workaround for the time being looks like this:
@styfle There are security concerns here, but if we "do it right" I think it is a reasonable approach (if you know those headers can only be set by trusted proxies, you should be fine). The added support for the meta headers would also change this a bit, but not in a bad way.
Also, we would default them to the values computed from
server.address()I think.It makes sense to me to offer both (obviously under different names) - ie,
req.urlas the string, andreq.getURL()as the lazily-computed-but-memoized URL object.Why both when there's
URL::href?Modifying
req.urlto return a URL object instead of a string would be a breaking change and it doesn't seem worth the churn to the ecosystem.10 remaining items
Since the higher level api is intended to usable directly
I think we should probably agree on which APIs have which intended constituencies. @ronag outlined a bit of what the goals of the different
undiciapis are, and the layering which he does enables really low level usage, but also has "direct end user api's" which are high level. The goal was to take a similar approach (see the table in the comment linked above). What we don't yet have is a clear statement on the target usage of the layers.we don't introduce "yet another" url implementation
One of the proposals above was to utilize a specialty parser like
request-targetfor the url. My concern with this is we already have 2 implementations, this would add a third.I think we should probably agree on which APIs have which intended constituencies. @ronag outlined a bit of what the goals of the different undici apis are, and the layering which he does enables really low level usage, but also has "direct end user api's" which are high level. The goal was to take a similar approach (see the table in the comment linked above). What we don't yet have is a clear statement on the target usage of the layers.
And what is wrong with using:
const url = new URL(req.url);
Seems simple enough if a user wants to parse the URL and can also be provided by the frameworks if they must...
One of the proposals above was to utilize a specialty parser like request-target for the url. My concern with this is we already have 2 implementations, this would add a third.
agree with you and I believe this gives an opportunity to let users (or frameworks) chose how they want to parse a URL. As you suggested, we can provide a function to
createServerthat tells the server how to parse the URL. If not provided, just returns a string by default:http.createServer({ allowHTTP1: true, parseUrl(stringUrl) { return new URL(stringUrl); } })
And what is wrong with using:
Nothing wrong, just an extra step and technically it would need to be some combination of this and the following (I have not tested all of this so it might not be completely accurate, and there might be other edge cases I am missing because
URLdoes not parse relative urls):const addr = server.address() const proto = getProto(req) /* as I mentioned, would need the above linked handling for proto */ const url = new URL(req.url, req.headers.host || `${proto}//${addr.address}`);
I think the more this conversation goes on we need more clarity on what the use case are. I think that most cases will be that only the pathname is used and that might lean more toward just telling users to parse it themselves if necessary.
Reacted by Steven and Anton Rybakov - Антон РыбаковI think what'd need to make this work is:
URL.from({ host, path, protocol }). In this way we can skip the non-needed parts of the algorithm that are really expensive.Essentially pushing forward the use of the native
URLobject requires some investigation on how to make its creation more performant.Reacted by Wes Todd, Ethan Arrowood, Jean Burellier and Felix HausOne of the proposals above was to utilize a specialty parser like request-target for the url. My concern with this is we already have 2 implementations, this would add a third.
Agreed, that will likely lead to more confusion.
I think the more this conversation goes on we need more clarity on what the use case are. I think that most cases will be that only the pathname is used and that might lean more toward just telling users to parse it themselves if necessary.
@wesleytodd In addition to pathname, I think query string is quite common too.
I think what'd need to make this work is:
URL.from({ host, path, protocol }). In this way we can skip the non-needed parts of the algorithm that are really expensive.@mcollina I'm not sure that solves the problem of parsing
req.urlinto some object. Would you be able toURL.from(req.url)or evenURL.from(req)?thanks for your message @styfle ! But this group isn't really active anymore (last two years I would say)
I see. I've been bouncing around different issues and all seem to be dead ends.
Reacted by Jean Burellier@mcollina I'm not sure that solves the problem of parsing req.url into some object. Would you be able to URL.from(req.url) or even URL.from(req)?
I think
URL.from(req)would be fantastic. We probably can't shipt that API, but possiblyhttp.getURL(req), which would be ok. Alsoreq.getURL()might work.Reacted by Steven and Anton Rybakov - Антон Рыбаковpossibly
http.getURL(req), which would be ok. Alsoreq.getURL()might work.Either option sounds great!
We could use
original-urlas a starting point (or even vendor it into node core)original-urluses the unsaferequire('url').parsemethod. I don't think anybody should use that implementation as there are multiple security problems with it. The best approach would be to pass vianew URL()or analog.would you like to send a PR for this?
How is
url.parseunsafe? If the incoming request conforms to the generic URI syntax (which obviously should be verified, as you always do for untrusted user input, right?), there's no reason there should be problems.would you like to send a PR for this?
Looks like someone already did: watson/original-url#11
Reacted by Stevenissue related: nodejs/node#51311
Initial context: nodejs/node#12682
I think the new api's should use
URLas the basis forreq.url(or whatever equivalent we expose). As you can see from the thread above, there is some question about how today these are relative urls and the new spec does not support that. My thought, and what I think we need to discuss, is that the "base" should be derived from the incoming request headers and fall back to the server address.If we come to an agreement on this I am happy to write up a more detailed description. Agenda added so we can discuss in the next meeting as well.