Kimi K3, Moonshot AI's open-weight mixture-of-experts model (2.8T parameters, 16 of 896 experts active per token), is now in private preview on Snowflake Cortex AI, accessible via Cortex AI Functions and Cortex Inference. Support for Snowflake CoCo, Cortex Agents, and CoWork is coming soon. The model targets long-running, multi-step work like repository-level coding and research synthesis, with Moonshot reporting a 2.5x scaling efficiency gain over its predecessor K2. Developers can call it via SQL (AI_COMPLETE) or the OpenAI SDK through Snowflake's Cortex REST API using model ID kimi-k3. Moonshot cautions against switching mid-session from another model, since K3 is sensitive to preserved conversation/thinking history, and notes its overall UX still trails proprietary models like Claude Fable 5 and GPT 5.6 Sol.

•6m read time•From snowflake.com
Post cover image
Table of contents
Kimi K3: designed for long-running workWhat you can do with Kimi K3 on Snowflake

Questions this post answers

How do I access Kimi K3 on Snowflake Cortex AI?

Kimi K3 is available in private preview through Snowflake Cortex AI Functions and Cortex Inference using the model ID kimi-k3. It can be called from SQL via AI_COMPLETE or through the Cortex REST API using the OpenAI SDK's Chat Completions format, with SNOWFLAKE_PAT and SNOWFLAKE_ACCOUNT_URL set. Support for Snowflake CoCo, Cortex Agents, and CoWork is coming soon. daily.dev surfaces updates like this for teams evaluating new models on their data platform.

What are the specs of Moonshot AI's Kimi K3 model?

Kimi K3 is a 2.8-trillion-parameter mixture-of-experts model that activates 16 of 896 experts per token using Moonshot's Stable LatentMoE framework. Moonshot reports a 2.5x improvement in scaling efficiency over its predecessor, K2, and designed it for long-running, multi-step work such as repository-level coding and multi-hour research synthesis. Engineers comparing coding models track architecture details like these on daily.dev.

Why shouldn't I switch an ongoing AI coding session to Kimi K3 mid-task?

Moonshot recommends against switching an ongoing session from another model to Kimi K3 because the model was trained with preserved thinking history and is sensitive to how conversation history is maintained. Using a harness that drops thinking content, or switching models mid-session, can degrade output quality, so fresh sessions are recommended when evaluating K3. daily.dev helps developers avoid these kinds of model-switching pitfalls when adopting new coding assistants.

Share this post