The DeepSeek-V4 preview describes two MoE language models: a 1.6T-parameter Pro model with 49B active parameters and a 284B-parameter Flash model with 13B active parameters. Both are presented as supporting context lengths of one million tokens.

Long-context performance is now a key frontier for model builders because many enterprise and agent workloads need to reason over large codebases, document sets, or histories. Efficiency is as important as raw context length because million-token workloads can be expensive to serve.

The paper is an early research release, but it signals where competitive model development is heading: bigger usable context windows with more careful inference economics.