Back Home

GitHub Repo

Hermes Agent Community Proposes Context Override Fix After Custom Provider Name Mapping May Break Configuration

A community report on October 6 says that after Hermes Agent connects to a custom model proxy, cached context length and manual overrides may prevent the proxy from starting. A fix for the ineffective override has been proposed, but remains a draft and does not address cache expiration.

Gerald L. Nino · Public domain · Image source
zh-Hant

On October 6, the Hermes Agent community reported two possible issues with context length resolution for custom OpenAI-compatible providers: an old detection result may remain in effect, and a user-configured override may be ignored. The report uses v0.21.5, Docker, and a self-hosted LiteLLM proxy, and includes startup logs and function inspection results. It remains a community report, and the scope of impact has yet to be confirmed upstream. Issue report

The official documentation defines context_length as the combined budget for input and output tokens. Hermes uses it to determine when to compact history and to validate requests. In the resolution order, a manual setting should take precedence over the persistent cache, which is retained across restarts. As a result, even after a fix is deployed on the proxy, the client may still have a different understanding of its capacity. Provider documentation

The reporter described deployments with different capacities behind the same model alias. Startup probing detected 24,000 tokens and cached that value; even after removing the smaller deployment and restarting, Hermes continued using it, causing initialization to fail. When the reporter tried setting model.context_length, custom:litellm was resolved at runtime as custom, and name matching incorrectly treated this as a route change and cleared the override. The reporter said that explicitly setting model.base_url to match the actual endpoint allowed the override to take effect. Reproduction and workaround

Draft PR #133615 addresses only name matching: it treats a named custom provider and its runtime custom identity as the same, while preserving distinctions between differently named providers. The author reports that eight related tests passed, but this validation comes from the author; there is no evidence yet of a merge or official release. The issue of persistent caching without an expiration mechanism is also outside the scope of this fix. Proposed fix

For operations engineers, this case suggests checking backend capacity, the client-side cache, and whether the configuration actually took effect when investigating context errors. Follow-up should track whether the provider identity fix is merged and whether a cache revalidation mechanism is added. The available evidence is not sufficient to conclude that all LiteLLM deployments or custom endpoints are affected.

Sources

  1. Context-length resolution doesn't treat a custom/LiteLLM-style alias as dynamic
  2. fix(agent): keep the context pin when runtime flattens a named custom provider
  3. LLM and Model Providers — Context Length Detection