Error handling
Temporal automatically retries failed Activities and recovers from infrastructure failures through Durable Execution. But not all failures should be retried. This page covers how to categorize failures, when to mark errors as non-retryable, and how to handle failures that retries cannot resolve.
For background on how Temporal represents and propagates failures, see Application failures.
Categorize failures
When an operation doesn't succeed, the appropriate response depends on whether retrying can resolve it. Durable Execution sorts these outcomes into three categories: transient failures, permanent failures, and negative results.
Transient failures
A transient failure is one where retrying the operation can eventually succeed. Some transient failures are one-time events. For example, a Worker happens to make a network request at the exact moment an administrator replaces a network cable, and the next request succeeds. Others are intermittent. For example, a service that uses rate limiting rejects requests once the threshold is reached, but accepts requests again after the rate limiter resets.
Temporal's default Retry Policy handles transient failures automatically. For
intermittent failures, configure the Retry Policy with an appropriate backoffCoefficient and maximumInterval to
space retries out over a longer period and avoid overwhelming the failing service.
Permanent failures
A permanent failure is one that will recur indefinitely until the cause is fixed. For example, a request that fails due to an invalid email address will continue to fail no matter how many times the operation retries. The only resolution is to correct the email address.
Permanent failures cannot be resolved through retries. They require different input data, a code fix, or some external intervention. Mark these errors as non-retryable so the Activity stops retrying instead of consuming resources on retries that will not succeed.
Failing the Activity fast doesn't have to fail the Workflow. Durable Execution keeps the Workflow's state and progress, so the Workflow can catch the failure, wait for corrected data through a Signal or Update, and then run the Activity again. See the Resumable Activity pattern and Pause a Workflow on failure and resume it after a fix. Bugs in Workflow code work differently: they fail the Workflow Task, not the Workflow Execution, so the Workflow Execution stays open until you deploy a fix.
Negative results
Some outcomes aren't failures. They are your business logic reaching an expected, but negative, result: a customer outside the service area, an order exceeding a credit limit, an expired promotion code, or a card with insufficient funds.
Durable Execution doesn't decide whether a negative result is a dead end or worth waiting on. That depends on your domain, so the decision belongs in your code. You can:
- Return the result as a value and branch on it in the Workflow.
- Raise a non-retryable Application Failure and handle it in the Workflow, for example with compensation.
- Treat it as transient: wait, for example for a Signal that the customer updated their payment details, and then try again.
Mark permanent errors as non-retryable
When your code detects a permanent failure, mark the error as non-retryable to prevent unnecessary retry attempts. For
background on what Application Failures are and how the non_retryable flag works, see
Application Failure.
Use non-retryable errors for situations like:
- Invalid input data: A malformed email address, a negative payment amount, or a missing required field.
- Negative results you handle as failures: A customer outside the service area, an order exceeding credit limits, or an expired promotion code. See Negative results.
- Authorization failures: The caller does not have permission to perform the operation.
- Data validation errors: A referenced record does not exist, or data fails integrity checks.
There are two ways to mark errors as non-retryable:
In the Activity (implementer decides): Set the non_retryable flag when throwing an
Application Failure. This enforces the constraint for all
callers. Use this when the Activity implementer knows that the error can never be resolved through retries.
In the Retry Policy (caller decides): Add the error type to the Retry Policy's list of non-retryable error types. This lets different Workflows make different decisions about the same Activity. Use this when the decision depends on the caller's business logic.
Preserve retryability when wrapping errors
When an Activity returns an error, the SDK checks the outermost error type to determine retryability. If you catch a
non-retryable Application Failure and re-throw it wrapped in a generic language error, the non_retryable flag is lost
and the Activity will be retried.
To add context to an error while preserving its retry behavior, wrap it in another Application Failure with the same
non_retryable flag. Do not wrap Application Failures in generic language errors.
For a detailed explanation of how the SDK-to-server chain works, see The outermost error type determines retryability.
Use non-retryable errors sparingly
In most cases, let the Retry Policy handle retry limits through timeouts
and maximum attempts. Reserve non_retryable for cases where retrying is guaranteed to be futile.
For SDK-specific syntax and code examples, see the error handling guide for your language:
Design Activities for idempotence
Activities may execute more than once due to retries, so design them to be idempotent: producing the same result whether executed once or multiple times.
This is especially important because of an edge case in distributed systems. A Worker can execute an Activity, complete it, and then crash before reporting the result to the Temporal Service. The Activity is retried even though it completed, because the Service has no record of the completion.
Use idempotency keys to prevent duplicate operations. Combine the Workflow Run ID and Activity ID for a value that is consistent across retries but unique across Workflow Executions.
Implement compensation with the Saga pattern
When a multi-step process fails partway through, previous steps may need to be undone. The Saga pattern coordinates a sequence of operations where each step has a compensating action that reverses its effects. If any step fails, the compensating actions for previously completed steps execute in reverse order.
For SDK-specific implementations with working code examples, see: