You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This issue is not ready for contributions yet.
The Kestra team still has to review it. Please wait until this message is removed and the issue has been move to Backlog stage before you start working on it. Pull requests opened before that will not be reviewed.
In the meantime, feel free to get started on a different issue by checking out our good first issues list.
Follow-up to #15082. The fix for that issue (#15147, #15153) re-runs non-terminated sibling task runs on replay. It has no special handling for a WorkingDirectory sibling. This was found by reading the code and has not been reproduced yet.
Suspected problem
ExecutionService.replay (core/src/main/java/io/kestra/core/services/ExecutionService.java) restarts every non-terminated task run that is neither the replayed one nor one of its successors. Each one is re-mapped to RESTARTED with a new task run id.
A WorkingDirectory runs its children inside a single worker job. For a WorkingDirectory that is being restarted, removeWorkerTask deletes the children's task runs so they are recreated with it. replay only calls that helper for the replayed task run's ancestors, not for the restarted siblings.
If a WorkingDirectory sibling was RUNNING when the execution ended, this could happen:
the WorkingDirectory task run is restarted;
its SUCCESS children keep their previous results;
its RUNNING child is itself non-terminated, so it is re-mapped to RESTARTED too. The executor dispatches any task run in a created state, so that child may be sent to a worker on its own, outside its WorkingDirectory.
The child could then run twice, or without the working directory files, and the replayed execution could end in an inconsistent state.
Steps to reproduce (to be confirmed)
Create a Dag (or Parallel) with a failing task a and a WorkingDirectoryb that has several children.
Get an execution where a failed, b is still RUNNING with at least one running child, and the execution is terminated. This state is reachable when the executor error path ends the execution (Execution.failedExecutionFromExecutor).
Replay from a.
Observe which task runs are dispatched to workers for the new execution.
Suggested approach
Reproduce first with an executor state-machine test, in the style of DagReplayTest (executor/src/test/java/io/kestra/executor/statemachine). If it confirms the problem, apply removeWorkerTask to the restarted sibling set as well.
Related
restart() copies a RUNNING sibling unchanged and may hang the same way as #15082. That needs its own reproduction.
Important
This issue is not ready for contributions yet.
The Kestra team still has to review it. Please wait until this message is removed and the issue has been move to Backlog stage before you start working on it. Pull requests opened before that will not be reviewed.
In the meantime, feel free to get started on a different issue by checking out our good first issues list.
Follow-up to #15082. The fix for that issue (#15147, #15153) re-runs non-terminated sibling task runs on replay. It has no special handling for a
WorkingDirectorysibling. This was found by reading the code and has not been reproduced yet.Suspected problem
ExecutionService.replay(core/src/main/java/io/kestra/core/services/ExecutionService.java) restarts every non-terminated task run that is neither the replayed one nor one of its successors. Each one is re-mapped toRESTARTEDwith a new task run id.A
WorkingDirectoryruns its children inside a single worker job. For aWorkingDirectorythat is being restarted,removeWorkerTaskdeletes the children's task runs so they are recreated with it.replayonly calls that helper for the replayed task run's ancestors, not for the restarted siblings.If a
WorkingDirectorysibling wasRUNNINGwhen the execution ended, this could happen:WorkingDirectorytask run is restarted;SUCCESSchildren keep their previous results;RUNNINGchild is itself non-terminated, so it is re-mapped toRESTARTEDtoo. The executor dispatches any task run in a created state, so that child may be sent to a worker on its own, outside itsWorkingDirectory.The child could then run twice, or without the working directory files, and the replayed execution could end in an inconsistent state.
Steps to reproduce (to be confirmed)
Dag(orParallel) with a failing taskaand aWorkingDirectorybthat has several children.afailed,bis stillRUNNINGwith at least one running child, and the execution is terminated. This state is reachable when the executor error path ends the execution (Execution.failedExecutionFromExecutor).a.Suggested approach
Reproduce first with an executor state-machine test, in the style of
DagReplayTest(executor/src/test/java/io/kestra/executor/statemachine). If it confirms the problem, applyremoveWorkerTaskto the restarted sibling set as well.Related
restart()copies aRUNNINGsibling unchanged and may hang the same way as #15082. That needs its own reproduction.