YOLO Vision 2026:
On the Ultralytics roadmap

Ultralytics YOLO-VLM

A lightweight YOLO front-end feeding a deeper LLM layer for efficient vision-language pipelines, announced on the Ultralytics roadmap. Build on YOLO26 today and be ready to connect vision to language at launch.

Trusted by the world's leading organizations

DuolingoShellSiemensRenaultPhilipsNEURA RoboticsMercado LibreTata SteelFlock SafetyIntelDefense Intelligence AgencyDHL
DuolingoShellSiemensRenaultPhilipsNEURA RoboticsMercado LibreTata SteelFlock SafetyIntelDefense Intelligence AgencyDHL
DuolingoShellSiemensRenaultPhilipsNEURA RoboticsMercado LibreTata SteelFlock SafetyIntelDefense Intelligence AgencyDHL

Our models' impact

Streamline processes across industries with our cutting-edge vision AI models. Speed, accuracy and ease-of-use powered by Ultralytics.

GitHubGitHub stars
Downloads
Ultralytics YOLO usages / day
Open-source contributors

The evolution of Ultralytics YOLO models

See how Ultralytics YOLO evolved from the practical YOLOv5 workflow to edge-ready YOLO26 inference.

Made real-time object detection accessible with a fast, practical PyTorch workflow.

Expanded the unified workflow across detection, segmentation, classification, pose, and OBB.

Improved accuracy, speed, and efficiency while preserving the familiar Ultralytics workflow.

Introduced end-to-end inference and an architecture optimized for efficient edge deployment.

Deploy Anywhere

Export to 20 formats and deploy across edge, cloud, and mobile.

Explore industry solutions

See how teams apply Ultralytics computer vision across production environments.

Frequently asked questions

  • YOLO-VLM is an upcoming Ultralytics model announced on the official Ultralytics roadmap as a lightweight YOLO front-end feeding a deeper LLM layer for efficient vision-language pipelines — fast visual perception connected to a language model.

  • Most vision-language models are large and expensive to run continuously on video. In general, splitting the work — a fast vision front-end feeding a deeper language layer — keeps perception real-time while reserving language compute for where it adds value. That is the pipeline shape the Ultralytics roadmap describes for YOLO-VLM.

  • In general, vision-language pipelines enable natural-language scene descriptions, visual question answering, open-vocabulary search over video, and alerts that explain what happened rather than only flagging it. YOLO-VLM's announced role is the efficient front-end for pipelines like these.

  • Ultralytics announces release timing on the official roadmap, which is the authoritative source for what ships when. Follow the roadmap and Ultralytics on GitHub for release announcements.

  • Start with Ultralytics YOLO26 for the real-time perception layer: detection, segmentation, pose, and tracking. Pipelines built on the Ultralytics workflow today are positioned to add the language layer when YOLO-VLM ships. Train and deploy now on Ultralytics Platform.

Be ready for YOLO-VLM

Build the perception layer with Ultralytics YOLO26 today and connect vision to language at launch.