Project MonetRequest demo
Home/Blog/K2 Horizon: Models, Benchmarks, Download & How to Run

AI · Project Monet Briefing

K2 Horizon: IFM’s Fully Open AI Model Fleet Explained

K2 Horizon is IFM’s September 2026 family of six openly released models spanning edge devices to enterprise-scale inference, with unusually broad training artifacts and deployment support.

Published 2026-09-07 · Updated 2026-09-07 · By Project Monet Editorial Team

K2 Horizon model family from 0.9B through 375B with open training and deployment paths

01

What is K2 Horizon?

The Institute of Foundation Models, or IFM, released K2 Horizon on September 3, 2026 as a connected fleet of six foundation models: 0.9B, 3.7B, 7B, 32B, 36B-A4B and 375B-A23B.

IFM positions the family as one deployment spectrum from constrained edge devices through local workstations and enterprise serving. The models share core architectural decisions, training methodology, interfaces, evaluation infrastructure and deployment tooling, while the 0.9B model uses a smaller vocabulary.

02

The six K2 Horizon models

  • 0.9B: the smallest model, designed for highly constrained edge environments.
  • 3.7B: a small dense model aimed at on-device and local use.
  • 7B: a medium dense model and practical local-development starting point.
  • 32B: IFM’s largest dense Horizon model, aimed at stronger local and server workloads.
  • 36B-A4B: a sparse MoVA + MoE model with roughly 4B active parameters per token.
  • 375B-A23B: the flagship sparse model for demanding enterprise-scale workloads.

Parameter labels are not hardware guarantees. Memory use depends on precision, quantization, context length, KV cache, runtime overhead and offloading strategy, so one universal VRAM minimum would be misleading.

03

How open is K2 Horizon?

IFM says it is opening the development lifecycle from pretraining through reasoning and agentic post-training, including final weights, intermediate checkpoints, training code, configurations, logs, evaluation resources, and training data or detailed construction recipes when redistribution is restricted.

The model weights and code are released under Apache 2.0. Datasets retain their applicable licenses, so commercial users should inspect dataset-specific terms instead of assuming Apache 2.0 covers every training artifact.

04

Download, GGUF and local deployment

The models are available through IFM’s Hugging Face organization. Official model cards document Transformers, vLLM and SGLang paths, and IFM now publishes GGUF repositories for multiple Horizon sizes including 3.7B, 7B, 32B and 36B-A4B.

The official 7B GGUF repository is intended for llama.cpp, but its current model card warns that upstream llama.cpp architecture support is still in progress and points to IFM’s fork. That makes runtime compatibility something to verify before treating Ollama or LM Studio support as universal.

05

Context length and reasoning settings

The current 3.7B and 7B cards document a native 524,288-token context window. Their validated serving examples commonly configure shorter contexts, which is a practical reminder that maximum context and sensible local context are not the same thing.

For benchmark-style reasoning, IFM recommends high reasoning effort, temperature 1.0, top_p 0.95 and enough output budget to avoid truncating reasoning. Those are vendor-recommended evaluation settings, not mandatory production defaults.

06

K2 Horizon API availability and pricing

IFM’s current press release says K2 Horizon API access is available through inference partners including Compass, Cerebras and Nebius. The same release does not publish one universal K2 Horizon API price.

07

Benchmarks: useful launch evidence, not independent proof

The current official 7B card reports 70.6 on SWE-bench Verified, 39.1 on Terminal-Bench 2.1 and 59.0 on BrowseComp. The current 3.7B card reports 68.6 on SWE-bench Verified and 25.1 on Terminal-Bench 2.1.

These are IFM-reported results under documented evaluation settings. The 7B card explicitly notes that its BrowseComp protocol can differ from comparison models, so the numbers should not be presented as independent proof of universal superiority.

08

What to verify before adopting K2 Horizon

  • Independent benchmark reproductions are still limited because the release is new.
  • Exact hardware requirements vary by model, precision, quantization and context.
  • Dataset licenses differ even though model weights and code are Apache 2.0.
  • Provider API pricing and model coverage are not universal.
  • GGUF availability does not guarantee every desktop runtime already supports the architecture.

For local deployment, start with the dedicated setup guide below and verify the live model card for the exact checkpoint and runtime you intend to use.

Sources

Primary and supporting sources

Facts were rechecked against the linked sources immediately before publication. Pricing, product availability and rollout status can change.

Project Monet

Useful signals. Clear decisions. Better digital work.

Project Monet turns relevant shifts in AI, creator tools and the web into practical context—and builds focused websites for businesses ready to grow.

Request a free homepage concept