There should be a stack setting that tells stacks to keep retrying for both proposed and tracked runs until a run is successful. Continuous retries are required to run at scale since stacks fail for reasons that just require a retry (provider download fail, worker pool error, dependent API blip, etc)