
Tether Data has released VisionPsy-Nano, an open-source, 460-million-parameter vision-language model designed to run directly on smartphones. Developed by the company’s AI research arm, QVAC, the model can analyze images, documents, and charts and answer related questions without sending data to cloud servers.
On-device vision-language AI
VisionPsy-Nano is built to perform both input processing and inference on the device itself. By keeping computation local, the model enables users to query visual content on their phones while retaining control of the underlying data. Vision-language models combine visual understanding with natural language capabilities, allowing tasks such as document reading, chart interpretation, and image-based Q&A.
Privacy and latency considerations
Running AI workloads on smartphones can reduce reliance on network connectivity and mitigate data exposure risks associated with cloud-based processing. On-device inference typically offers lower latency and more predictable performance, which can be advantageous for real-time or privacy-sensitive applications.
Model size and deployment
At 460 million parameters, VisionPsy-Nano is comparatively compact next to many cloud-scale models. The smaller footprint is intended to accommodate the compute and memory constraints of mobile hardware, prioritizing efficiency and responsiveness over the largest-possible model size.
Tether’s broader push into AI
Tether Data’s release reflects a growing industry effort to move AI capabilities to the edge. Tether is best known as the issuer of USDT, the largest U.S. dollar-pegged stablecoin by market capitalization, and the company has been expanding into adjacent technology initiatives. QVAC’s work on VisionPsy-Nano underscores that strategy by targeting practical, privacy-preserving AI deployments on consumer devices.