Skip to content

Replay restarts children of a running WorkingDirectory sibling outside their WorkingDirectory #20002

Description

@loicmathieu

Important

This issue is not ready for contributions yet.
The Kestra team still has to review it. Please wait until this message is removed and the issue has been move to Backlog stage before you start working on it. Pull requests opened before that will not be reviewed.
In the meantime, feel free to get started on a different issue by checking out our good first issues list.


Follow-up to #15082. The fix for that issue (#15147, #15153) re-runs non-terminated sibling task runs on replay. It has no special handling for a WorkingDirectory sibling. This was found by reading the code and has not been reproduced yet.

Suspected problem

ExecutionService.replay (core/src/main/java/io/kestra/core/services/ExecutionService.java) restarts every non-terminated task run that is neither the replayed one nor one of its successors. Each one is re-mapped to RESTARTED with a new task run id.

A WorkingDirectory runs its children inside a single worker job. For a WorkingDirectory that is being restarted, removeWorkerTask deletes the children's task runs so they are recreated with it. replay only calls that helper for the replayed task run's ancestors, not for the restarted siblings.

If a WorkingDirectory sibling was RUNNING when the execution ended, this could happen:

  • the WorkingDirectory task run is restarted;
  • its SUCCESS children keep their previous results;
  • its RUNNING child is itself non-terminated, so it is re-mapped to RESTARTED too. The executor dispatches any task run in a created state, so that child may be sent to a worker on its own, outside its WorkingDirectory.

The child could then run twice, or without the working directory files, and the replayed execution could end in an inconsistent state.

Steps to reproduce (to be confirmed)

  1. Create a Dag (or Parallel) with a failing task a and a WorkingDirectory b that has several children.
  2. Get an execution where a failed, b is still RUNNING with at least one running child, and the execution is terminated. This state is reachable when the executor error path ends the execution (Execution.failedExecutionFromExecutor).
  3. Replay from a.
  4. Observe which task runs are dispatched to workers for the new execution.

Suggested approach

Reproduce first with an executor state-machine test, in the style of DagReplayTest (executor/src/test/java/io/kestra/executor/statemachine). If it confirms the problem, apply removeWorkerTask to the restarted sibling set as well.

Related

restart() copies a RUNNING sibling unchanged and may hang the same way as #15082. That needs its own reproduction.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/backendNeeds backend code changes

    Type

    Fields

    Stage

    Inbox

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions