Training project
Multimodal vision / generation lab
Related course: Multimodal AI
Problem
Multimodal work requires demonstrating text↔image (and related) pipelines with clear limits and evaluation notes.
Architecture (high level)
Model APIs or open models for text-to-image / image-to-text → embedding comparison → short demo app; aligned with Multimodal AI curriculum topics.
Technologies
- OpenAI
- Hugging Face
- Stable Diffusion
This is a Vector Skill Academy training project completed during the course. It is not confidential client work and does not represent a live production engagement.