Back to Articles

Compact VLMs built for the edge: Tether AI Research introduces VisionPsy-Nano, a family of compact (~460M-parameter) vision-language models (VLMs) purpose-built for on-device and edge deployment. It ships as two variants: VisionPsy-Nano-460M (tuned for quality) and VisionPsy-Nano-460M-Flash (tuned for latency), bringing multimodal understanding that used to require the cloud onto the phone in your pocket.

Best-in-class at ~0.5B: VisionPsy-Nano-460M compares favorably to every other ~0.5B VLM tested, including LFM2.5-VL-450M, SmolVLM2-500M, and its own nanoVLM-460M-8k base, on 16 of 17 benchmarks, with a leading overall normalized score of 62.3 (vs 59.6 / 52.5 / 54.9).

Leads every capability category: VisionPsy-Nano-460M tops all four categories: document understanding & OCR, visual perception, reasoning & knowledge, and instruction following & reliability. Its widest relative margins are in reasoning & knowledge (+7.4% over the next-best ~0.5B model) and visual perception (+4.6%).

Punches above its weight: In the Instruction Following & Reliability category, covering both MM-IFEval (42.3) and POPE (87.9), VisionPsy-Nano-460M outperforms models 1.6x to 2.3x its size: FastVLM-0.5B (759M), Qwen3.5-0.8B (873M), and InternVL3.5-1B (1061M).