Ultralytics YOLO-VLM
A lightweight YOLO front-end feeding a deeper LLM layer for efficient vision-language pipelines, announced on the Ultralytics roadmap. Build on YOLO26 today and be ready to connect vision to language at launch.
Trusted by the world's leading organizations
Our models' impact
Streamline processes across industries with our cutting-edge vision AI models. Speed, accuracy and ease-of-use powered by Ultralytics.
The evolution of Ultralytics YOLO models
See how Ultralytics YOLO evolved from the practical YOLOv5 workflow to edge-ready YOLO26 inference.
Train on the Best GPUs for Less
26 NVIDIA GPUs starting at $0.24/hr — from Ampere to Blackwell. No markup, no minimums, no commitment.
Explore industry solutions
See how teams apply Ultralytics computer vision across production environments.

Agriculture

Automotive

Healthcare

Logistics

Manufacturing

Retail

Robotics

Agriculture

Automotive

Healthcare

Logistics

Manufacturing

Retail

Robotics

Agriculture

Automotive

Healthcare

Logistics

Manufacturing

Retail

Robotics
Frequently asked questions
YOLO-VLM is an upcoming Ultralytics model announced on the official Ultralytics roadmap as a lightweight YOLO front-end feeding a deeper LLM layer for efficient vision-language pipelines — fast visual perception connected to a language model.
Most vision-language models are large and expensive to run continuously on video. In general, splitting the work — a fast vision front-end feeding a deeper language layer — keeps perception real-time while reserving language compute for where it adds value. That is the pipeline shape the Ultralytics roadmap describes for YOLO-VLM.
In general, vision-language pipelines enable natural-language scene descriptions, visual question answering, open-vocabulary search over video, and alerts that explain what happened rather than only flagging it. YOLO-VLM's announced role is the efficient front-end for pipelines like these.
Ultralytics announces release timing on the official roadmap, which is the authoritative source for what ships when. Follow the roadmap and Ultralytics on GitHub for release announcements.
Start with Ultralytics YOLO26 for the real-time perception layer: detection, segmentation, pose, and tracking. Pipelines built on the Ultralytics workflow today are positioned to add the language layer when YOLO-VLM ships. Train and deploy now on Ultralytics Platform.
Be ready for YOLO-VLM
Build the perception layer with Ultralytics YOLO26 today and connect vision to language at launch.