Post
4482
๐ผ๏ธ POCKET-Image โ the POCKET series goes visual: character-perfect text in any language, on-device
A new model in VIDRAFT's POCKET family. POCKET put 35B-class models on phones and no-GPU PCs. POCKET-Image carries the same "big capability, small hardware" idea into image generation โ and fixes the one thing nearly every image model gets wrong: text.
Type "์๋ ํ์ธ์" into a typical model and you get "์ใ ๊ธฐ." Hangul alone composes 11,172 syllable blocks; Arabic connects its letters; Thai stacks marks. Diffusion models draw scripts as shapes, so they smear. POCKET-Image renders every glyph exactly โ ํ๊ตญ์ด ยท ไธญๆ ยท ๆฅๆฌ่ช ยท ุงูุนุฑุจูุฉ (RTL) ยท เนเธเธข ยท Latin and more โ onto any scene you describe.
What it is:
โข 100% accurate text, any language โ where global models produce gibberish
โข Any background from a prompt โ text is optional (empty โ a pure image)
โข No GPU, no NPU โ runs on plain CPU + RAM via the POCKET-Core engine
โข Measured footprint: 8.6 GB (RTX 3050/4060) ยท 4.5 GB (offloaded, 6 GB cards) ยท 13.4 GB (MacBook, 16 GB+)
โข Windows ยท macOS ยท Linux ยท fully local, no cloud
Built on the open, commercial-friendly Z-Image (Apache-2.0) foundation.
Honest note: the text is the guaranteed-correct part โ the surrounding scene is ordinary generation, so a busy foreground can crowd the letters. We say so; clean backgrounds stay razor-sharp.
๐จ Studio โ generate right here, any language:
FINAL-Bench/POCKET-Image-Studio
๐งฉ Model card:
FINAL-Bench/POCKET-Image-Zimage
๐ The POCKET collection:
https://huggingface.co/collections/FINAL-Bench/pocket-models
A new model in VIDRAFT's POCKET family. POCKET put 35B-class models on phones and no-GPU PCs. POCKET-Image carries the same "big capability, small hardware" idea into image generation โ and fixes the one thing nearly every image model gets wrong: text.
Type "์๋ ํ์ธ์" into a typical model and you get "์ใ ๊ธฐ." Hangul alone composes 11,172 syllable blocks; Arabic connects its letters; Thai stacks marks. Diffusion models draw scripts as shapes, so they smear. POCKET-Image renders every glyph exactly โ ํ๊ตญ์ด ยท ไธญๆ ยท ๆฅๆฌ่ช ยท ุงูุนุฑุจูุฉ (RTL) ยท เนเธเธข ยท Latin and more โ onto any scene you describe.
What it is:
โข 100% accurate text, any language โ where global models produce gibberish
โข Any background from a prompt โ text is optional (empty โ a pure image)
โข No GPU, no NPU โ runs on plain CPU + RAM via the POCKET-Core engine
โข Measured footprint: 8.6 GB (RTX 3050/4060) ยท 4.5 GB (offloaded, 6 GB cards) ยท 13.4 GB (MacBook, 16 GB+)
โข Windows ยท macOS ยท Linux ยท fully local, no cloud
Built on the open, commercial-friendly Z-Image (Apache-2.0) foundation.
Honest note: the text is the guaranteed-correct part โ the surrounding scene is ordinary generation, so a busy foreground can crowd the letters. We say so; clean backgrounds stay razor-sharp.
๐จ Studio โ generate right here, any language:
FINAL-Bench/POCKET-Image-Studio
๐งฉ Model card:
FINAL-Bench/POCKET-Image-Zimage
๐ The POCKET collection:
https://huggingface.co/collections/FINAL-Bench/pocket-models