Back Home

GitHub Repo

LangChain Community Reports Retry Filter Trap: A Single Exception Class May Cause Unexpected Errors to Be Retried

A community reproduction on LangChain 1.4.0 suggests that passing an exception class directly to `retry_on` may cause models and tools to retry errors that do not match the intended filter. The official documentation calls for a tuple of exception classes or a predicate function; a fix and the full scope of impact remain unconfirmed.

Dirck van Baburen · Public domain · Image source
zh-Hant

On October 3, the LangChain community reported an issue with agent retry behavior: configuring retry_on=TimeoutError in ModelRetryMiddleware or ToolRetryMiddleware was intended to retry timeouts only, but could also retry other exceptions. The reporter listed LangChain 1.4.0 as an affected environment and said the issue could also be reproduced on the development branch. At the time of review, the issue remained open and did not list a related fix PR. Community report

In the minimal reproduction, a tool always raises KeyError, with a maximum of two retries configured. The reported result was that the tool ran three times and ultimately returned an error message, instead of raising an exception after the first failure. The author traced the behavior to a shared predicate function that first checks callable(retry_on). Python exception classes are callable, so they are treated as predicate functions; calling one produces a truthy exception object, causing the filter to lose its intended effect. Reproduction and analysis

There is an important boundary here: the official API defines retry_on as either a “tuple of exception classes” or a function that returns a Boolean; it does not promise to accept a single class. Following the documentation, engineers can use retry_on=(TimeoutError,)—note the comma required for a single-element tuple—or pass an explicit isinstance predicate. The documentation also says exceptions that do not match the filter should be raised immediately, without entering the retry-exhaustion handling path. Model retry API

The technical impact concerns how an agent workflow recognizes failures. By default, the tool middleware returns a ToolMessage after retries are exhausted, allowing the model to continue. If the filter is configured in a way that behaves unexpectedly, a programming error that should have stopped execution may instead become an error result in the conversation. Tool retry API For tools with write side effects, repeated execution could also increase the risk of duplicate operations. That is a deployment inference, not an incident confirmed in this report.

Operators should check retry parameters, test whether a non-matching exception runs only once, and verify that errors are returned to the caller as expected. Follow-up questions include whether upstream will add input validation or officially support a single exception class. The current evidence comes from a community case and does not establish that all correctly configured setups are affected.

Sources

  1. ModelRetryMiddleware / ToolRetryMiddleware: retry_on=SomeError retries every exception
  2. ModelRetryMiddleware API reference
  3. ToolRetryMiddleware API reference