multimodal

#13
by MaziyarPanahi - opened

Thank you for continuously releasing open models that compete directly with frontier AI models. Just wanted to say would be AMAZING if we can support different modalities like image and audio in next DeepSeek models. 🤗

look how @baseten did it here: (not that they gave us a whitepaper or anything, but it can be reverse engineered by GPT 5.6 Sol)

https://huggingface.co/baseten/glm-5-2-projector
https://huggingface.co/baseten/GLM-5.2-Vision-FP8

Thank you for continuously releasing open models that compete directly with frontier AI models. Just wanted to say would be AMAZING if we can support different modalities like image and audio in next DeepSeek models. 🤗

Their chief guy Liang said, 'He only cares about AGI. Multimodal is useful but probably not a necessary part to achieve AGI.'😭

look how @baseten did it here: (not that they gave us a whitepaper or anything, but it can be reverse engineered by GPT 5.6 Sol)

https://huggingface.co/baseten/glm-5-2-projector
https://huggingface.co/baseten/GLM-5.2-Vision-FP8

Those guys reversed vision "only" from Kimi K2.6.
But these guys reversed vision from Kimi K3, plus they have also audio and thermal and infrared, and they can give multimodal even to any API model:

https://huggingface.co/EximiusLabs/fusion-embedding-2-k3-vision

Image input is the last missing piece to make this model perfect. Hopefully deepseek team is cooking this for a future version!

look how @baseten did it here: (not that they gave us a whitepaper or anything, but it can be reverse engineered by GPT 5.6 Sol)

https://huggingface.co/baseten/glm-5-2-projector
https://huggingface.co/baseten/GLM-5.2-Vision-FP8

Those guys reversed vision "only" from Kimi K2.6.
But these guys reversed vision from Kimi K3, plus they have also audio and thermal and infrared, and they can give multimodal even to any API model:

https://huggingface.co/EximiusLabs/fusion-embedding-2-k3-vision

Well first off I'm not convinced that Kimi K3's ViT is vastly superior to Kimi K2.6's ViT

Secondly I'm not convinced that img2text is a good way to add multimodal capabilities to a model both in quality and In speed, compared to a mmproj

DeepSeek-v4-flash的性能很强,而且速度也很快,非常适合处理各种日常任务,但是缺少了图像输入功能限制了很多使用场景。而且实际上现在的DeepSeek-v4系列模型很喜欢调用视觉工具,然后发现自己无法查看图像。希望可以添加一个图像输入的功能。

•
This comment has been hidden (marked as Spam)

What the hell is wrong with you with that picture @Sabahudin123 ?

Sign up or log in to comment