Free course material
Move from pixels to dependable visual decisions.
Computer Vision & Multimodal AI
A notebook-first path through modern vision, multimodal reasoning, spatial intelligence, embodied systems, and enterprise production concerns.
What this course is about
A practical path from understanding to application.
This field guide connects image and video fundamentals to modern computer vision systems. Each lesson combines technical explanation, a credential-free notebook, experiments, evaluation, failure analysis, and production implications across visual representation learning, multimodal AI, spatial systems, visual agents, and embodied intelligence.
Covered material
What learners will work through.
- Modern CNNs, vision transformers, self-supervised learning, detection, segmentation, retrieval, tracking, and pose
- Vision foundation models, open-vocabulary vision, visual embeddings, and promptable segmentation
- Vision-language models, multimodal reasoning and verification, document intelligence, and multimodal RAG
- Video-language understanding and bounded visual agents
- 3D vision, spatial intelligence, neural rendering, dynamic scenes, and world models
- Embodied vision, spatial memory, evaluation, robustness, privacy, governance, and production operations
Who it is for
Meet learners where they are.
Data scientists, machine learning engineers, AI engineers, researchers, architects, and technical teams building or evaluating image, video, document, multimodal, spatial, or embodied AI systems.
Adapt this course
Bring the material to your organization.
We can adapt the level, examples, exercises, and pace for executives, non-technical groups, technical teams, or mixed audiences. Sessions combine theory with practical and coding work where it fits, led by educators who meet with your group.
Book a custom training conversationRelated One+i service
Intelligent solution design
Connect open learning with the strategy, delivery, and adoption support around it.