TwelveLabs Pegasus 1.6: Specs, Context Window, Image Support and Limits
Quick answer: TwelveLabs Pegasus 1.6 is a multimodal video-understanding model released on October 6, 2026. The update improves first-person and egocentric video understanding, entity recognition and metadata extraction, while adding native multi-image analysis. It supports a context window of 261,120 tokens and can analyze long video segments, images and contextual text through the TwelveLabs API.
What is Pegasus 1.6?
Pegasus is TwelveLabs’ generative video-understanding model. Instead of treating video as a collection of unrelated screenshots, it is designed to reason across time, describe events, answer questions about scenes and generate structured information from visual sequences.
Version 1.6 focuses on three practical areas: better understanding of first-person footage, stronger recognition of people and objects, and image support that lets developers analyze still images in the same product family.
Pegasus 1.6 specifications
| Capability | Pegasus 1.6 |
|---|---|
| Context window | 261,120 tokens |
| Video analysis | Up to 2 hours per analyzed portion |
| Large source videos | Can work with longer files when analyzing a selected portion |
| Maximum file size | Up to 10 GB for supported video workflows |
| Images per request | 1–20 |
| Image formats | JPEG, PNG, WebP, GIF, BMP |
| Image size | Up to 20 MB per supported image |
| Core strengths | Video understanding, image analysis, event description, entity extraction |
Better first-person and egocentric video understanding
One of the headline changes is improved performance on egocentric footage: video recorded from the perspective of the person performing an action. Examples include body cameras, smart glasses, action cameras, robotics feeds and training footage captured from a worker’s viewpoint.
First-person video is difficult because the camera moves with the operator, hands can block objects and the visual field changes rapidly. Better egocentric reasoning can make the model more useful for procedure review, training analysis and robotics data.
Improved entity recognition
Pegasus 1.6 is also designed to recognize entities and extract metadata more reliably. For a media archive, that can mean identifying recurring objects, products, people or places and producing more consistent descriptions.
Entity recognition is still probabilistic. High-stakes applications should compare extracted entities with authoritative records rather than automatically treating a model-generated name as confirmed.
Native image analysis
The release adds support for analyzing multiple still images in one request. TwelveLabs currently documents between 1 and 20 images per request, covering common formats such as JPEG, PNG, WebP, GIF and BMP.
This creates useful workflows such as:
- comparing frames from different cameras;
- summarizing product-photo sets;
- checking visual continuity between storyboard images;
- matching still images to video events;
- extracting structured metadata from image batches.
How much video can Pegasus 1.6 analyze?
TwelveLabs documentation supports analyzing up to roughly two hours in a selected video portion. Longer source videos can be handled by choosing the part to analyze rather than expecting the model to reason over an unlimited timeline in one prompt.
For very long media archives, developers should segment files by chapter, scene or time range and store the resulting summaries or embeddings for retrieval.
Why the 261,120-token context window matters
A large context window lets the model combine detailed temporal information with user instructions and supporting text. It is particularly useful when a request requires the model to follow a long sequence and relate an event near the end to something much earlier.
As with text LLMs, more context is not always better. Passing only the relevant time window can reduce cost and improve focus.
Timestamped in-segment events
Current docs describe improved event handling within selected video segments, including timestamp-aware output. That makes Pegasus useful for systems that need answers such as “when did the worker pick up the red tool?” or “which moment contains the product demonstration?”
Timestamp precision should be tested on the type of media you plan to process. Fast sports footage and long static meetings create very different evaluation problems.
Supported image limits
TwelveLabs documents up to 20 MB per image and a maximum of 16,777,216 pixels for supported still-image requests. Extremely large images should be resized before upload to reduce bandwidth and avoid unnecessary preprocessing.
Video resolution and format considerations
Documentation describes supported video dimensions from approximately 360×360 up to 5184×2160, alongside supported aspect-ratio constraints. A practical production pipeline should use FFmpeg or a similar tool to normalize unusual codecs and dimensions before sending media to the API.
Normalization reduces failures and makes latency more predictable, especially when user uploads come from many devices.
Language support
English has the strongest documented support. TwelveLabs also lists partial support across several languages including Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish, Thai and Vietnamese.
For multilingual deployments, test speech, on-screen text and user questions separately. “Language support” does not guarantee identical accuracy across every video type.
Good use cases for Pegasus 1.6
Body-camera and field-service review
Teams can ask questions about procedures, safety steps and equipment interactions in first-person footage.
Media archive search and summarization
Broadcasters and creators can generate summaries, scene descriptions and searchable metadata from video libraries.
Robotics and embodied-AI datasets
Egocentric footage is common in robotics and human-demonstration datasets. Improved first-person understanding can assist with labeling and quality review.
Visual compliance checks
A model can flag candidate moments for human review, although regulated decisions should not rely on automated interpretation alone.
How to integrate Pegasus 1.6
- Create a TwelveLabs account and API key.
- Normalize media into a documented format.
- Upload or reference the video/image inputs.
- Select only the relevant time range for long video.
- Send a focused prompt that specifies the desired output structure.
- Store timestamps and source IDs with generated results.
- Evaluate accuracy against manually labeled examples.
Prompting tips
For video understanding, prompts should be specific about time, entities and format. “Summarize the video” is less reliable than “List the five main actions in chronological order with timestamps and the objects involved.”
If you need structured output, define the fields explicitly. For example: event, start time, end time, person, object and confidence note.
Limitations to plan for
- Entity names can be wrong or ambiguous.
- Small text and distant objects may be missed.
- Fast cuts can complicate timestamp precision.
- Very long files should be segmented strategically.
- Multilingual performance can vary by language and audio quality.
- Visual understanding is not a substitute for human review in safety-critical decisions.
Related AVARIXO coverage
For other multimodal AI launches, see our Mistral Large 4 guide and Reflection Beam overview.
Frequently asked questions
Can Pegasus 1.6 analyze images?
Yes. Version 1.6 adds native multi-image analysis, with current documentation listing 1–20 images per request.
How large is the context window?
261,120 tokens.
Can it analyze first-person video?
Yes. Improved egocentric and first-person understanding is one of the release’s main changes.
How long can a video be?
The model can analyze a selected portion up to roughly two hours, with longer source files handled by selecting the relevant segment.
Sources
Features and technical limits were verified against TwelveLabs’ official Pegasus 1.6 announcement and Pegasus 1.6 documentation.
