[TRLC-DK1] Failure replay and bad episode triage for builders labs (advanced)
TRLC-DK1 discussion on failure replay, bad episode triage, debugging strategy, and how labs decide what to keep, relabel, or discard.
DK1 teams eventually accumulate a pile of questionable runs: some are useful failures, some are broken captures, and some look bad only because the replay tooling is weak.
How are you triaging failed DK1 episodes and using failure replay to decide what to keep, relabel, or discard?
Please share how you replay bad runs, what metadata or signals you inspect first, and when a failure is still useful for training or evaluation.
If you reply, include one exact replay clue that changed your decision about a bad episode.








This thread is most helpful when replies distinguish informative task failure from broken capture or broken labels.
If you have a triage order for logs, video, actions, and calibration state, share it. Searchers want the workflow, not only the conclusion.
A small replay rubric can be more valuable here than a long debugging story.