MBZUAI Launches K2 Horizon 375B to Power AI Agents
Abu Dhabi's MBZUAI has released K2 Horizon 375B A23B, an open-weights mixture-of-experts model designed to reduce hallucinations and excel at complex, real-world agentic workflows.

The Institute of Foundation Models at Abu Dhabi's Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) has launched K2 Horizon 375B A23B. This flagship model leads a new six-model open-weights fleet released under the Apache 2.0 license. Moving away from the dense transformer architecture of its predecessor, the 70-billion-parameter K2 Think V2, the new flagship uses a sparse mixture-of-experts design. It features 375 billion total parameters with 23 billion active per token, alongside a massive 524,000-token context window. The model is available for deployment through partners including AWS, Cerebras, Nebius, and Compass.
K2 Horizon 375B A23B scored 47 on the Artificial Analysis Intelligence Index, representing a 30-point jump over K2 Think V2. In evaluations, the model prioritizes practical agentic workloads over trivia recall, outperforming its closest architectural rival, MiniMax-M3. On the GDPval-AA real-world knowledge work benchmark, K2 Horizon achieved an Elo rating of 1430 compared to MiniMax-M3's 1380. It also dominated the τ³-Banking tool-use simulation, scoring 34.2% against MiniMax-M3's 15.3%. However, it lagged on harder academic reasoning benchmarks, scoring 87.3% on GPQA Diamond and 32.0% on Humanity's Last Exam, compared to MiniMax-M3's 92.9% and 39.0% respectively.
A key feature for practitioners is the model's high rate of abstention on uncertain queries. On the AA-Omniscience benchmark, K2 Horizon attempted only 40% of the questions, helping lower its hallucination rate to 26% and raising its Omniscience Index from -40 to -3, despite a flat raw accuracy of 18%. Beyond the flagship, the K2 Horizon fleet includes a 0.9-billion-parameter model for wearable devices, 3.7-billion and 7-billion versions for mobile phones, and a dense 32-billion alongside a sparse 36-billion-A4B model for on-premise servers. This range allows developers to deploy tailored agentic capabilities across diverse hardware environments.
This is our own summary of reporting by AlphaSignal



