Coming Soonvisionmultimodaldocumentsocr
Zentra Vision
A vision-language model that understands images, diagrams, charts, and documents alongside text.
7.1B + ViT params
32,768 tokens context
Transformer + Vision Encoder
Coming Q4 2026
Technical Specifications
| Parameters | 7.1B + ViT |
| Context | 32,768 tokens |
| Architecture | Transformer + ViT |
| Vision | 336×336 patches |
| Layers | 32 |
| Hidden Size | 4,096 |
| Vocab | 128,000 |
| License | Apache 2.0 |
System Requirements
Minimum
RAM16 GB
VRAM8 GB
OSLinux, macOS, Windows
Recommended
RAM32 GB
VRAM16 GB
OSLinux, macOS, Windows
About
Zentra Vision extends the Zentra family with multimodal capabilities. It can process and reason about images alongside text, making it ideal for:
- Document understanding and OCR - Chart and diagram analysis - Visual question answering - UI/UX design feedback and generation - Medical imaging report generation
Get Zentra Vision
Download the GGUF weights for local inference, or try it in the playground.
Open SourceApache 2.0 LicenseHugging Face