最新模型
Aleph Alpha Releases Kolibri-1, a 78B MoE Model with Hybrid Attention for Million-Token Context
Kolibri-1 is released with open weights under Apache 2.0. It activates about 3.46B parameters per token and targets German-English bilingual use, document processing, and tool calling. Its million-token window extends beyond the model’s native 256k training length; the company still recommends shorter contexts for complex tasks.

Aleph Alpha released Kolibri-1 on October 3, publishing the weights and configuration files for a mixture-of-experts model with about 78.1B parameters under the Apache 2.0 license. The model focuses on German and English, and offers configurable reasoning modes and tool calling so developers can deploy it themselves for long-document question answering, structured extraction, and agent workflows. Official announcement
The technical focus is managing the compute and long-context costs of a large model. All 50 of Kolibri’s layers use MoE, with 384 routed experts per layer. Each token selects six of them, alongside one shared expert, for about 3.46B active parameters. Its attention combines 40 layers with a 512-token sliding window and 10 layers of global attention. This limits long-sequence costs in most layers, but the full weights still need to fit in memory, so deployment capacity cannot be estimated from the active parameter count alone. Model card
The million-token window also has clear limits: the model was first pretrained with 16k sequences, followed by 64k mid-training and 256k long-context adaptation. Its native training length is 262,144 tokens. The company says that because positional encodings are used only in sliding-window layers, the context can be extended without adjusting positional scaling, and has been validated up to 1,048,576 tokens. However, for complex tasks and services where latency and throughput matter, it still recommends staying at or below 256k. In practice, teams should test window capacity, cross-segment retrieval, and end-to-end task success separately. Model card
Deployment requires the vLLM plugin provided by aleph-alpha-inference, configured with the kolibri1 inference and tool-calling parsers. Reasoning strength is set through the chat template to none, low, medium, or high. Services exceeding 256k must also explicitly override the maximum sequence length. Deployment instructions
The company also says training to refuse when evidence is lacking strengthens document grounding, which may be useful for enterprise RAG. However, performance comparisons at launch are still based on the vendor’s own tests. Chinese is not a target language, so teams working in Traditional Chinese should first validate bilingual documents, refusal quality, tool-argument accuracy, and GPU memory use with long sequences before deciding whether the model suits their existing workloads. Official technical overview