Instructions to use deepseek-ai/DeepSeek-V4-Flash-0731 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use deepseek-ai/DeepSeek-V4-Flash-0731 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="deepseek-ai/DeepSeek-V4-Flash-0731")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V4-Flash-0731") model = AutoModelForCausalLM.from_pretrained("deepseek-ai/DeepSeek-V4-Flash-0731", device_map="auto") - Inference
- HuggingChat
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use deepseek-ai/DeepSeek-V4-Flash-0731 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "deepseek-ai/DeepSeek-V4-Flash-0731" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "deepseek-ai/DeepSeek-V4-Flash-0731", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/deepseek-ai/DeepSeek-V4-Flash-0731
- SGLang
How to use deepseek-ai/DeepSeek-V4-Flash-0731 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "deepseek-ai/DeepSeek-V4-Flash-0731" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "deepseek-ai/DeepSeek-V4-Flash-0731", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "deepseek-ai/DeepSeek-V4-Flash-0731" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "deepseek-ai/DeepSeek-V4-Flash-0731", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use deepseek-ai/DeepSeek-V4-Flash-0731 with Docker Model Runner:
docker model run hf.co/deepseek-ai/DeepSeek-V4-Flash-0731
multimodal
Thank you for continuously releasing open models that compete directly with frontier AI models. Just wanted to say would be AMAZING if we can support different modalities like image and audio in next DeepSeek models. 🤗
look how @baseten did it here: (not that they gave us a whitepaper or anything, but it can be reverse engineered by GPT 5.6 Sol)
https://huggingface.co/baseten/glm-5-2-projector
https://huggingface.co/baseten/GLM-5.2-Vision-FP8
Thank you for continuously releasing open models that compete directly with frontier AI models. Just wanted to say would be AMAZING if we can support different modalities like image and audio in next DeepSeek models. 🤗
Their chief guy Liang said, 'He only cares about AGI. Multimodal is useful but probably not a necessary part to achieve AGI.'😭
look how @baseten did it here: (not that they gave us a whitepaper or anything, but it can be reverse engineered by GPT 5.6 Sol)
https://huggingface.co/baseten/glm-5-2-projector
https://huggingface.co/baseten/GLM-5.2-Vision-FP8
Those guys reversed vision "only" from Kimi K2.6.
But these guys reversed vision from Kimi K3, plus they have also audio and thermal and infrared, and they can give multimodal even to any API model:
https://huggingface.co/EximiusLabs/fusion-embedding-2-k3-vision
Image input is the last missing piece to make this model perfect. Hopefully deepseek team is cooking this for a future version!
look how @baseten did it here: (not that they gave us a whitepaper or anything, but it can be reverse engineered by GPT 5.6 Sol)
https://huggingface.co/baseten/glm-5-2-projector
https://huggingface.co/baseten/GLM-5.2-Vision-FP8Those guys reversed vision "only" from Kimi K2.6.
But these guys reversed vision from Kimi K3, plus they have also audio and thermal and infrared, and they can give multimodal even to any API model:https://huggingface.co/EximiusLabs/fusion-embedding-2-k3-vision
Well first off I'm not convinced that Kimi K3's ViT is vastly superior to Kimi K2.6's ViT
Secondly I'm not convinced that img2text is a good way to add multimodal capabilities to a model both in quality and In speed, compared to a mmproj
DeepSeek-v4-flash的性能很强,而且速度也很快,非常适合处理各种日常任务,但是缺少了图像输入功能限制了很多使用场景。而且实际上现在的DeepSeek-v4系列模型很喜欢调用视觉工具,然后发现自己无法查看图像。希望可以添加一个图像输入的功能。