DeepSeek V4.1 Flash model has recently officially released. It is the smallest model in DeepSeek's new model architecture series and features native multimodal visual understanding capabilities. DeepSeek V4.1 Flash is a 552B-parameter MoE model and significantly reduces the size of the KV Cache. Compared with the previous generation model, demand for HBM has been reduced to 1/4, and demand for SSDs has been reduced to 1/8. In Agent usage scenarios, cache hit costs often account for a relatively high proportion, and the compression of the KV Cache substantially lowers the usage cost of Agent-type tasks.