Core AI
Multimodal & Perception
22 tools · 3 subcategories
Models that see and hear: vision-language models like GPT-4o and LLaVA, computer vision systems like SAM and Detectron2, and speech recognition engines like Whisper and Deepgram. Pure perception lives here; generative image and audio model weights belong to Foundation Models.










