IFM/K2-Horizon-0.9B
Compact dense model for mathematics, code, instruction following, and STEM tasks
1.08B stored parameters including embeddings, with 131K context and native K2-Horizon parsing
Guide
Overview
K2-Horizon-0.9B is a compact dense model in the IFM K2-Horizon family, with 1.08B stored parameters including embeddings and a 131,072-token context window.
K2-Horizon is an open-weight IFM model family built for transparent foundation-model research, staged checkpoint analysis, and practical deployment. The series spans compact dense models such as K2-Horizon-0.9B for local experimentation, dense mid-size and large models for high-quality research workloads, and mixture-of-expert (MoE) models for higher-capacity serving.
Launch Commands
K2-Horizon-0.9B ships YaRN RoPE scaling in its config.json (factor 16 over an 8,192-token original window, rope_theta 1,000,000, beta_fast 128, beta_slow 4), so the full 131,072-token context needs no --hf-overrides. A top-level rope_parameters override replaces that whole block rather than merging into it, so a partial one silently drops rope_theta and the YaRN betas.
Recommended: With YaRN
vllm serve IFM/K2-Horizon-0.9B \
--trust-remote-code \
--chat-template-content-format string \
--dtype bfloat16 \
--max-model-len 131072 \
--gpu-memory-utilization 0.85 \
--reasoning-parser k2_horizon \
--enable-auto-tool-choice \
--tool-call-parser k2_horizon
This configuration enables the full 131,072-token context length and is the recommended setting for K2-Horizon.
Shorter context
To cap memory at the original 8,192-token window, lower --max-model-len. YaRN stays on because it comes from the checkpoint config:
vllm serve IFM/K2-Horizon-0.9B \
--trust-remote-code \
--chat-template-content-format string \
--dtype bfloat16 \
--max-model-len 8192 \
--gpu-memory-utilization 0.85 \
--reasoning-parser k2_horizon \
--enable-auto-tool-choice \
--tool-call-parser k2_horizon
Client usage
from openai import OpenAI
client = OpenAI(
api_key="EMPTY", base_url="http://localhost:8000/v1", timeout=3600
)
resp = client.chat.completions.create(
model="IFM/K2-Horizon-0.9B",
messages=[{"role": "user", "content": "Give me three primes above 100."}],
temperature=1.0, top_p=0.95, max_tokens=2048,
)
print(resp.choices[0].message.content)
Thinking modes
The model supports selectable thinking through chat_template_kwargs, per
request or server-wide via --default-chat-template-kwargs:
{"reasoning_effort": "high"}— full thinking (default).{"reasoning_effort": "medium"}— faster thinking.{"reasoning_effort": "low"}— fastest thinking.