CODEC-STREAM VIDEO UNDERSTANDING
Vision-Language Model
Glint-VL-Video-8B is the first large video understanding model that uses codec streams as its visual unit, identifying key events, locating time ranges, and extracting evidence for video review, content retrieval, and process analysis.
- Long-Video Understanding
- Temporal Grounding
- Event Retrieval