Repository navigation
Add eager task creation API to asyncio #97696
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancement
on Sep 30, 2022 - added3.12only security fixesonly security fixesperformancePerformance or resource usagePerformance or resource usage
on Oct 1, 2022 If
condisTrueand some or all ofcoro()completes eagerly, any side-effects of this will be observable in further execution. Iftg.create_task()is used instead, no part ofcoro()will be executed ifcondisTrue.If I didn't already understand the changes to async functions you are proposing I wouldn't understand what this means (or the code fragment above it). What does
condrefer to? I think what you are trying to say here is that, if the body of thewith TaskGroup()block raises before the firstawait, it will cancel all created tasks before they have started executing. Right?Separately -- Can you remind me of the reason to make this a TaskGroup method instead of a new flag on create_task()? And the enqueue() name feels odd, especially since it may immediately execute some of the coroutine, which seems the opposite of putting something in a queue.
To be clear, I have no objection to the feature, and I think making this an opt-in part of task creation avoids worries of changing semantics for something as simple as
await coro().Separately -- Can you remind me of the reason to make this a TaskGroup method instead of a new flag on create_task()? And the enqueue() name feels odd, especially since it may immediately execute some of the coroutine, which seems the opposite of putting something in a queue.
This came up in the conversation we had with @DinoV during the Language Summit. We started out with a flag on
create_task(e.g.eager=True), but we didn't like the fact that when setting this flag, the name of the method may now be misleading (because, in fact, it may not create a task).
enqueuewas the alternative (I don't remember who suggested it and why), but we can definitely bikeshed the naming :)
If we take inspiration from Trio's nurseries, we can usestart_soon.
Other options areexecuteorrun(although these names may be misleading too, since users may assume these APIs never create a task).If I didn't already understand the changes to async functions you are proposing I wouldn't understand what this means (or the code fragment above it). What does cond refer to? I think what you are trying to say here is that, if the body of the with TaskGroup() block raises before the first await, it will cancel all created tasks before they have started executing. Right?
Right. Yeah, this is more contrived than needed, I'll fix up the text.
This came up in the conversation we had with @DinoV during the Language Summit. We started out with a flag on create_task (e.g. eager=True), but we didn't like the fact that when setting this flag, the name of the method may now be misleading (because, in fact, it may not create a task).
I wasn't in this conversation but when writing this issue it occurred to me that even with eager execution we could still return something
Task-like. This would mitigate any issues with knowing the interface of the returned object (e.g. iscancel()available). This would also mean the namecreate_task()would still make sense even with a keyword-argument to alter the behavior. Unless, I'm missing something which doesn't make this viable?Having said this, I think eager execution should be the default behavior users reach for going forward. So it might still be better to have a new method.
Thanks for fixing up the description, it's clear now.
I wasn't in this conversation but when writing this issue it occurred to me that even with eager execution we could still return something
Task-like. This would mitigate any issues with knowing the interface of the returned object (e.g. iscancel()available). This would also mean the namecreate_task()would still make sense even with a keyword-argument to alter the behavior. Unless, I'm missing something which doesn't make this viable?It could return a Future, which has many of the same methods. If the coroutine returns an immediate result that could be the Future's result.
However, if the result is generally not used though the need to return something with appropriate methods would just reduce the performance benefit.
Having said this, I think eager execution should be the default behavior users reach for going forward. So it might still be better to have a new method.
Yeah, that is definitely an advantage of a new method name. I find
enqueue()rather easy to mistype though.Have the performance results been repeated with
TaskGroup.enqueue(), or are the quoted measurements based on the Cinder implementation usinggather()? There could be surprises here.Have the performance results been repeated with TaskGroup.enqueue(), or are the quoted measurements based on the Cinder implementation using gather()? There could be surprises here.
The benchmark showing an 8x improvement (etc.) is with Dino's prototype of
TaskGroup.enqueue()against some revision of 3.11. This is functional, but still needs some work. Notably it does not yet implement management of the currentContextwhen eagerly executing. There will be some extra overhead from that but my hope is it'll be negligible.It could return a Future, which has many of the same methods. If the coroutine returns an immediate result that could be the Future's result.
However, if the result is generally not used though the need to return something with appropriate methods would just reduce the performance benefit.
Dino's prototype currently returns a newly constructed
Futureif the eager execution fully completes, but returns aTaskotherwise. While I can imagine it'll be less common to use the extra fields onTask, it might be annoying to have to think about whether aFutureor aTaskis returned. Again, unless I missed something, I don't think there should be extra overhead from returning some kind of eagerly-completed-Taskvalue vs. a completedFuture.I find enqueue() rather easy to mistype though.
Indeed, I have misspelled it a number of times already. I like
start_soonorstart_immediatewhich to me imply execution start time may not be bound by theasync withscope.For a new API, returning either a
Futureor aTaskis totally fine --Taskis a subclass ofFuture, so we can just document it as returning aFuture.I hadn't realized the prototype doesn't even need C changes -- I had assumed it depended on the changes to vectorcall to pass through a bit indicating it's being awaited immediately. Is that change even needed once we have this? (Probably we resolved that already during discussions at PyCon, it's just so long ago I can't remember anything.)
I think we need to bikeshed some more on the name, but you're right about the impression the name ought to give. Maybe
start_task()?We can just document it as returning a Future.
I think it'd be better to document it as returning a
Task? That way it's clear cancellation is still generally an option.I had assumed it depended on the changes to vectorcall to pass through a bit indicating it's being awaited immediately. Is that change even needed once we have this? (Probably we resolved that already during discussions at PyCon, it's just so long ago I can't remember anything.)
That is an interesting question, and again I don't think I was around if that was discussed in detail. For this feature, that flag is not needed. If we want to add an optimization for the large body of existing code that uses
asyncio.gather(), then it would help significantly. We are probably going to carry this feature forward in Cinder for a while at least as it'll likely take quite some time (years?) before we can fully migrate to the newTaskGroupAPI. I'd definitely like to know what your, or other members of the community feel about adding an optimization to help with the older/existing API.Maybe start_task()?
Fine with me. Maybe we should discuss further once a PR is up.
I think it'd be better to document it as returning a Task? That way it's clear cancellation is still generally an option.
From a static type POV it's always a Future, not always a Task. If you want to cancel it you take your chances -- cancel() will return a bool telling you whether it worked. The docs should just explain the situation without lying.
I will wait for the PR (that you're hopefully working on?). I recommend using
start_task()as the method name until we come up with something better.From a static type POV it's always a Future, not always a Task.
Whoops, I just realized
Futurehas acancel()method. I mistakenly thought that was a feature ofTask. Sorry for the confusion.I will wait for the PR (that you're hopefully working on?)
We should have something up soon. Really wanted to get an issue open first for early feedback on the plan.
51 remaining items
- added a commit that references this issue
on May 9, 2023 - added a commit that references this issue
on May 9, 2023 - added a commit that references this issue
on May 9, 2023 What’s left? Is there a what’s new entry?
What’s left? Is there a what’s new entry?
I think this is done now!
Yes, we added a what's new entry.- moved this from Todo to Done in Release and Deferred blockers 🚫
on May 9, 2023 - added a commit that references this issue
on May 9, 2023 - added a commit that references this issue
on May 9, 2023 - added 3 commits that reference this issue
on May 9, 2023 - added a commit that references this issue
on Feb 26, 2024
Metadata
Metadata
Assignees
Labels
Projects
- StatusShow more project fieldsDone
- StatusShow more project fieldsDone
Feature or enhancement
We propose adding “eager” coroutine execution support to
asyncio.TaskGroupvia a new methodenqueue()[1].TaskGroup.enqueue()would have the same signature asTaskGroup.create_task()but eagerly perform the first step of the passedcoroutine’s execution immediately. If thecoroutinecompletes without yielding, the result ofenqueue()would be an object which behaves like a completedasyncio.Task. Otherwise,enqueue()behaves the same asTaskGroup.create_task(), returning a pendingasyncio.Task.The reason for a new method, rather than changing the implementation of
TaskGroup.create_task()is this new method introduces a small semantic difference. For example in:The exception will cancel everthing scheduled in
tg, but if some or all ofcoro()completes eagerly any side-effects of this will be observable in further execution. Iftg.create_task()is used instead no part ofcoro()will be executed.Pitch
At Instagram we’ve observed ~70% of coroutine instances passed to
asyncio.gather()can run fully synchronously i.e. without performing any I/O which would suspend execution. This typically happens when there is a local cache which can elide actual I/O. We exploit this in Cinder with a modifiedasyncio.gather()that eagerly executescoroutineargs and skips scheduling aasyncio.Taskobject to an event loop if no yield occurs. Overall this optimization saved ~4% CPU on our Django webservers.In a prototype implementation of this proposed feature [2] the overhead when scheduling
TaskGroups with all fully-synchronous coroutines was decreased by ~8x. When scheduling a mixture of synchronous and asynchronouscoroutines, performance is improved by ~1.4x, and when nocoroutines can complete synchronously there is still a small improvement.We anticipate code relying on any semantics which change between
TaskGroup.create_task()andTaskGroup.enqueue()will be rare. So, as the TaskGroup interface is new in 3.11, we hopeenqueue()and its performance benefits can be promoted as the preferred method for scheduling coroutines in 3.12+.Previous discussion
This new API was discussed informally at PyCon 2022, with at least some of this being between @gvanrossum, @DinoV, @markshannon, and /or @jbower-fb.
[1] The name "
enqueue" came out of a discussion between @gvanrossum and @DinoV.[2] Prototype implementation (some features missing, e.g. specifying Context), and benchmark.
Linked PRs