模型與企業部署
Flower Endeavor 1.0 Opens Up Private Enterprise Deployment, but Model Specifications and Independent Evaluations Are Still Missing
Flower Labs has introduced Endeavor 1.0 for reasoning, software development, and long-running agent tasks, available either as a managed service or deployed in customer environments. Its initial results are close to those of some frontier models, but with access currently limited to a preview and the benchmarks based solely on vendor self-evaluation, real-world deployment costs and capabilities cannot yet be verified.

Flower Labs released Endeavor 1.0 on September 1, positioning it as a general-purpose model capable of reasoning, software development, and extended tool use. Rather than making the model weights publicly downloadable, Flower offers it as a managed service or, with licensing and technical support, for deployment on an enterprise’s own infrastructure. During the preview period, applications are limited to a small number of organizations.
The key technical aspect of this arrangement is the ability to run the same model across cloud and private environments. Enterprises can initially access it through an API, then migrate sensitive workloads to internal clusters without rewriting the surrounding agent infrastructure, evaluations, tool interfaces, or data pipelines. Flower says Endeavor manages reasoning depth, context, tool selection, and post-failure inspection and retries at inference time, and that it was post-trained and evaluated using multiple coding and agent harnesses.
The company reports scores of 92.0 on GPQA, 98.2 on HumanEval, 94.1 on IFEval, and 99.9 on AIME 2026. According to its comparison table, Endeavor outperforms GPT-5.6 Sol, Claude Fable 5, Kimi K3, and Nemotron 3 Ultra on HumanEval; on AIME, it ties the first two at 99.9. However, it still trails GPT-5.6 Sol on GPQA and IFEval, while AIME scores approaching 100 are no longer effective at differentiating models. All of these results were produced by Flower, and the test prompts, number of samples, reasoning budgets, and harness configurations have not yet been fully disclosed.
Flower also says the model was developed using signals from internal enterprise tasks in FlowerBench: data and tools remain within participating organizations’ environments, with only sanitized results returned. However, the release materials do not provide Endeavor’s FlowerBench scores, failure categories, or a reproducible protocol.
Engineering teams should not currently equate “privately deployable” with open weights. Flower has not disclosed the parameter count, architecture, context length, weight format, licensing restrictions, recommended hardware, memory footprint, throughput, or pricing. The next things to watch are the formal model card, inference-server compatibility, independent evaluations of long-running tasks, and whether the managed and private versions use the same model, tool semantics, and update cadence.