Dense experts and heavy memory needs
Snowflake Arctic uses 480 billion parameters distributed across 128 fine-grained experts. Each expert contains 3.66 billion parameters. The model uses top-2 gating to select 17 billion active parameters during inference. This design allows the model to outperform DBRX and Llama 3 in coding and SQL generation tasks. DBRX uses 132 billion total parameters and 36 billion active parameters, while Arctic activates only 17 billion. While the 17 billion active parameters suggest low compute costs, the full 480 billion parameters must still load into VRAM, and this makes the memory footprint significantly larger than dense models like Llama 3 8B. I find the massive parameter count misleading because the full 480 billion parameters must reside in VRAM unless you use quantization or offloading. The architecture combines a 10 billion dense transformer model with a residual Mixture of Experts component. This combination allows the model to overlap communication and computation. This system-model co-design helps hide communication overhead during training.
Snowflake trained Arctic on Amazon EC2 P5 instances in under three months. The training process used 3.5 trillion tokens. Phase 1 utilized 1 trillion tokens. Phase 2 used 1.5 trillion tokens. Phase 3 used 1 trillion tokens. The research team used public datasets like RefinedWeb and RedPajama. This effort cost approximately $2 million in computational resources. This cost sits at about one-eighth of the expense needed to build comparable models. The team used a three-stage curriculum to optimize enterprise-focused tasks.
Instruction-tuning stability issues
Instruction-tuning the 480B parameter Arctic-Instruct model presents unique challenges. Sparse MoE architecture with 128 experts causes instability because each expert often sees fewer than one token per example in the median case. Because completion tokens amount to only a few hundred, each expert sees less than one token per example in the median case. The naive approach to fine-tuning can have several inefficiencies. For instance, compute is wasted on padding when examples are padded to the same length. To solve this, Snowflake engineers used sequence packing to fit multiple prompt and completion pairs into a single 4096 length sequence. This approach ensures more examples fit into each batch and improves training throughput. They also treated each conversation as a single training example to increase the fraction of trainable tokens in each batch. This reduces the training compute required per conversation and helps alleviate the expert gradient sparsity issue.
Inference performance and costs
At a batch size of 1, Arctic provides a throughput of over 70 tokens per second with FP8 quantization. This performance results from having 4 times fewer memory reads than Code-Llama 70B and 2.5 times fewer than Mixtral 8x22B. High-volume production needs a large batch size. This requires enough KV cache memory to support the batch while storing nearly 500 billion parameters. At large batch sizes, Arctic incurs 4 times less compute than CodeLlama 70B and Llama 3 70B. Will the extreme memory footprint of the 480B parameter model cancel out the cost savings of the 17B active parameters?
Snowflake collaborated with NVIDIA and vLLM teams to provide a preliminary implementation for interactive inference. Two-node inference achieves large batch sizes through a combination of FP8 weights, split-fuse and continuous batching, tensor parallelism, and pipeline parallelism. You should decide if the 4k context window meets your needs. OpenAI’s GPT-4.1 supports a 1 million token context window, whereas Arctic has a 4k context window. For text-only applications, OpenAI’s GPT-4.1 costs $2.00 per million input tokens and $8.00 per million output tokens. Snowflake Arctic provides an Apache 2.0 license, allowing for commercial use without restrictive clauses.
| Feature | Snowflake Arctic | Llama 3 8B |
|---|---|---|
| Total Parameters | 480B | 8B |
| Active Parameters | 17B | 8B |
| License | Apache 2.0 | Llama 3 Community |
| Context Window | 4k | 8k |
