<#332 [BUG][Timeline] Timeline is crashing for one...
# flyte-github
a
#332 [BUG][Timeline] Timeline is crashing for one specific workflow, with failed NodeExecution Issue created by anrusina Failing Workflow The Crash for this workflow was fixed. However, the fix was mainly an overcome of corner case than a proper fix. Workflow execution in question: https://demo.nuclyde.io/console/projects/cluster-deploy/domains/development/executions/f76101f1a820a4dda8c4?duration=all Looking closer into it I see that even though node execution is failed, it looks like one of its child is still running, for 7400h+, see screenshot.

Screen Shot 2022-03-17 at 10 02 14 AM▾

Other Workflows with similar behavior: From 10 month ago: run fca7d008dc9ab42f2856 From a week ago: https://demo.nuclyde.io/console/projects/flytesnacks/domains/development/executions/ad4j68tk2wrx7bjw6s9h?duration=all Where wan of a map child tasks are still in the "Running state", while parent is FAILED BE Conversation in Slack: https://unionai.slack.com/archives/C01H0FN1NJX/p1647536794593289 What we currently have The latest solution in this PR Includes check that if task is running longer than 1 week (WEEK_DURATION_SEC) we would assume that it wasn't properly finished. In such case we will mark it with state ABORTED and set its length to
allegedDurationSec
- longest currently available duration happened earlier in the Dag tree. Proper solutions Ideally we should: • Mark status as ABORTED if parent item has FAILED state (which is not that easy to do with current DAG system) • Fix backend to ensure that status is updating properly flyteorg/flyteconsole