Skip to main content
Open a clip in Studio to refine timing, text, subtitles, and finishing touches before you export or share it.

Opening Studio

There are two common ways to open Studio, depending on whether you are editing an existing clip or starting a brand-new project.

Option 1: Edit an existing clip

Open a clip and click Edit Clip. That takes you directly into Studio. You will typically do this from the clip detail page after Overlap has already generated or surfaced a clip for review. A good mental model is:
  • automation and workflows generate or queue the clip
  • the clip detail page gives you a focused view of the selected clip
  • Edit Clip opens the manual editor for that specific output
  • Studio is where you make the final hands-on changes before exporting or sharing
Clip detail page with Edit Clip

Option 2: Start a video from scratch

From the Home page, use New in the Projects section, then choose Edit a video from scratch.
New menu with Edit a video from scratch
That opens a blank project where you can upload a source file and jump straight into editing in Studio.
Blank project upload screen

Understanding the Studio layout

Studio brings the clip preview, playback controls, timeline, and editing tools into one screen.
Overlap Studio editor
The main areas of the editor are:
  • The top bar, where you can go back, see the clip title, and access Export and Share
  • The preview canvas in the center, which shows the current frame of the clip
  • The Reframe control beside the preview for adjusting composition
  • The Thumbnail Editor control beside Reframe, which opens a dedicated thumbnail design surface
  • The playback controls below the preview, including the play button, current time, total duration, speed control, and fit/zoom controls
  • The timeline at the bottom, where Overlap shows each visual or text layer over time
  • The right-side tool rail, which opens different editing panels

Editing with the transcript

Open Transcript in the right-side tool rail to work on the spoken content directly.
Studio transcript editor
In the current Transcript Editor, Overlap shows:
  • speaker-grouped transcript blocks; click a speaker’s name or portrait to identify that speaker across every matching section and every clip made from the same source input, or use the adjacent dropdown to move only the current section to a different speaker track. While either menu is open, the transcript softly previews exactly which sections will be affected. Existing People names and portraits appear automatically when the speaker is already known
  • markers like Video start and Video end so you can see what portion of the source is currently inside the clip
  • a transcript-focused toolbar above the text for quick edit operations
This is the fastest place to make content-aware edits because you can work from the words instead of hunting visually through the timeline.

Highlighting transcript sections

To edit a specific moment, highlight the exact word, phrase, sentence, or paragraph you want to change in the transcript. From there, use transcript actions such as:
  • Cut from Video to remove the selected spoken portion from the clip itself
  • Mute to silence the selected section when you want to keep the visuals but remove the audio
  • Edit Captions to change what appears on screen without treating it as a full video cut
This distinction matters:
  • use Cut from Video when the pacing or content of the clip should change
  • use Mute when the shot should stay but the audio should drop out
  • use Edit Captions when the video timing is fine and only the on-screen text needs adjustment
A practical way to use the transcript editor is:
  1. Read through the clip from top to bottom in Transcript.
  2. Highlight the exact section that feels off.
  3. Decide whether the problem is the spoken content, the audio, or only the captions.
  4. Apply the corresponding action.
  5. Scrub the timeline and preview the result before moving on.

Working with the timeline

The timeline is the fastest way to understand how the clip is assembled over time. In the current studio view, the clip is split into separate tracks such as:
  • Rich Text
  • Subtitles
  • Watermark
  • Video
This layered view helps you see what is happening at each moment in the clip. Use the timeline when you want to:
  • Check when subtitle segments begin and end
  • See how long a text overlay remains on screen
  • Confirm whether the watermark spans the full clip
  • Scrub through the video before exporting
  • Jump back to the saved thumbnail frame from the thumbnail marker
In the normal timeline editor, click a video segment to select it. Use Command+C and Command+V on Mac, or Ctrl+C and Ctrl+V on Windows, to copy and paste the segment after the selected position. Use Command+D or Ctrl+D to duplicate the selected segment in one step. These shortcuts are scoped to the timeline and are not available in Reframe editing mode. The ruler across the top of the timeline shows time markers in seconds, which makes it easier to inspect short-form clips precisely. The saved thumbnail frame appears on the ruler as a small black-and-white marker. Hover the marker to see Thumbnail, or click it to scrub the preview back to that frame.

Designing the clip thumbnail

By default, Overlap uses the first frame as the clip thumbnail. To create a designed cover, scrub to the frame you want and click Thumbnail Editor beside Reframe. The thumbnail editor starts from the clip exactly as it appears at that frame. Inside it, you can:
  • scrub to a different frame
  • add and style text with the Studio text presets
  • upload an image or choose one from the company media library
  • move, resize, replace, or delete thumbnail layers
  • remove clip overlays such as subtitles, titles, watermarks, and b-roll from the thumbnail without removing them from the video
  • adjust exposure, contrast, saturation, vibrance, warmth, and tint
Click Save thumbnail when the design is ready. Overlap composites the selected frame and thumbnail-only layers at the clip’s full output resolution, saves the result as the clip thumbnail, and keeps the editable design so it reopens on the same frame with the same layers and colour settings. Thumbnail edits are independent from the video timeline. Removing or restyling a layer in the thumbnail editor does not change the clip itself, and later Studio saves do not replace a designed thumbnail automatically. You can also open the same editor from a scheduled clip in Social Calendar: open the post and click Design in its cover controls.

Creating a thumbnail in chat

The editing agent can create a polished thumbnail from the current clip, uploaded PNG/JPEG/WebP images, or a mix of both. Ask it to create a thumbnail and describe the headline, subject, and visual direction you want. It samples candidate moments from the clip, uses the company colors, fonts, voice, and logos saved under Brand Kits → My Brand, and returns the generated image in chat. You can also request square or vertical artwork when the destination is not a standard 16:9 thumbnail. Clip sampling compares nine moments in a 3×3 grid. It renders the actual video with the clip’s reframe crop applied and defaults to clean video-only frames. Ask to include the existing title, subtitle, and graphic layers when you want the full clip composition represented in those samples. Continue in the same conversation to revise the result. Feedback such as “make the headline larger,” “use the other speaker,” “remove the logo,” or “go back to the first version” creates a new version without overwriting the earlier image. Generated and revised thumbnails are never saved to the clip automatically. Review any draft, then click Apply to Clip on that thumbnail to save it as the clip thumbnail. You can also explicitly ask the agent to apply a specific version. Each successful generation appears in chat immediately. If visual review calls for another pass, the previous draft stays visible while the next aspect-ratio-matched placeholder renders, and up to three successful attempts are shown together in a wrapping grid. Thumbnail text is kept fully inside safe margins; the composition can zoom out or reflow when needed to prevent clipped letters. The same conversation memory applies to generated files from every editing-agent skill, not just thumbnails. Overlap keeps compact references to earlier images and files with the chat, so follow-up requests such as “move the text up,” “use the latest image,” or “go back to the first version” can resolve the earlier output after a reload. Relevant prior images are shown to the agent again for visual inspection; file metadata and stable references remain available without storing image bytes or private provider responses in the conversation. When the editing agent needs a clip before it can continue, chat shows a compact clip chooser above the message box. Select an existing clip from the Library modal, upload a video to ingest it as a clip, or paste an Overlap clip link. Chat sends the selected clip back to the agent and continues the request; no command or placeholder text needs to be typed manually. Local videos stream directly from the browser into the ingestion pipeline in small, acknowledged chunks; the original video is not first copied into Firebase Storage. After normalization, Overlap saves the video as a regular company clip with its duration, dimensions, thumbnail, transcript, and source metadata. Remote HTTP video links use the same normalization contract. When the agent needs a long-form video, chat presents an inline chooser where you can select a saved source, drag or upload a video file, or enter an HTTP(S) video URL. YouTube watch, share, Shorts, embed, and live URLs are downloaded through the same normalizer and returned to chat as an Overlap source clip. For longer videos, ask chat to find highlights or viral moments. The editing agent reads every window of the normalized transcript, uses your direction to prioritize moments, and checks representative source frames for the finalists before finalizing clean source-time boundaries within exact duration requests such as “45 to 90 seconds.” It then creates the clip records directly instead of sending the source through a separate clipping workflow. The normalized source video is stored once and reused by every clip in the batch. Each result keeps the source’s native aspect ratio and stores its selected source-time ranges in an editable segments array. A result may contain one continuous range or multiple ranges that omit an internal digression or form a coherent montage. Montage ranges are stored in intended playback order, so a strong hook found later in the source can play first when the resulting progression still makes conceptual sense. Overlap does not make a trimmed, reframed, or rendered video copy for each result, and you can continue refining those segments later in Studio. The compact Activity indicator appears only when preparation or transcript scanning takes long enough to be useful, and you can expand it for the current phase. Leaving the conversation does not cancel work still running on the editing server. Completed clip records and their shared normalized source remain available when you reopen the chat. The compact token counter under the conversation reports new, uncached model input plus model output. Cached input is excluded, while reasoning tokens are already included in output and are not added a second time. Hover the counter for the input/output breakdown and an indication when a provider could not return complete usage for an interrupted or direct model call.

Using the right-side tools

The right rail is where you switch between different editing tasks. In the current studio UI, the tool groups are:
  • Transcript
  • Subtitles
  • Media
  • Text
  • AI Tools
  • Transitions
A good way to think about them is:
  • Use Transcript and Subtitles when the spoken words or on-screen captions need refinement
  • Use Media when you want to work with supporting visual assets
  • Use Text when you want to manage overlays such as headline or supporting copy. For title-style text fields, click Generate to create a short AI title. You can optionally prompt the generator with a preferred angle, tone, or emphasis; leave the prompt blank to generate from the clip context and existing title settings
  • Use AI Tools when you want Overlap to apply a broader cleanup or packaging step for you
  • Use Transitions when you want to adjust how the clip moves between visual states

AI-assisted cleanup and finishing

AI Tools collects several of the fastest cleanup and packaging actions in one panel.
Studio AI tools panel
In the current editor, Overlap groups these actions into two sections:
  • Transcript Tools
  • Video Edits
Verified options shown in Transcript Tools include:
  • Filler Words
  • Stutter Words
  • Curse Words
  • Remove Silences
  • Remove Punctuation
  • Keyword Highlights
Verified options shown in Video Edits include:
  • End Card
  • Intro Music
  • Outro Music
  • Add Speaker Cards
This makes Studio useful for both cleanup and packaging. You can remove distracting speech patterns, tighten pacing, highlight important language, and add finishing elements without leaving the editor. The editing agent can also prepare and generate selected motion graphics for a clip. It analyzes the clip title, transcript and word timing, PySceneDetect visual cuts, and sampled semantic-scene frames so the storyboard can distinguish the owner brand from companies or products discussed in the content. Before generation starts, chat presents a compact storyboard row with the scene count, mode mix, brand, review actions, and a fading preview of the first few planned beats. Select the plan to inspect the scrollable brand direction, creative direction, and complete motion sequence. Choose Approve to start one scene batch, Revise to send feedback through the normal conversation and receive a new numbered revision, or Discard to close the proposal without generating anything. The storyboard survives a reconnect, rejects stale decisions, and never exposes its private analysis or moodboard artifacts. The storyboard remains sparse, but it sees the active designer-authored template catalog—including logo capability—before choosing overlay or fullscreen beats and preserves minimum reading time for each mode. Each scene agent must use the selected or best-matching active template first; custom HTML is an escape hatch only when no compatible template serves the required function or an attempted template reports a concrete preview/render failure. While multiple scenes generate, the passive Activity indicator keeps long-running work visible with its real state and available timing without mixing approval actions into that status surface. If a scene calls for a logo, the agent can resolve either a saved My Brand asset or an evidence-grounded referenced brand, prefers a sanitized SVG, previews the exact mark on light and dark backgrounds, and packages it locally before rendering. Preview QA rejects declared images that do not load in the browser, preventing broken-image placeholders from reaching publish. Moodboard references remain inspiration only and never become runtime assets. Final scene packages include the HTML composition, metadata, QA results, previews, and declared local assets so playback and exports do not depend on third-party image URLs. Chat can also generate standalone motion graphics without a clip loaded. Ask for a lower third, logo reveal, headline card, title slide, data graphic, or a full-screen or bounded email/text-message scene. For podcast chapters and conceptual pivots, the template library includes more image-led editorial treatments: topic cutouts, raised folios, framed abstract media, pull-quote collage, stacked posters, paper-peel transitions, loud edge-to-edge title slams, sculptural ribbon folds, dimensional portals, and cinematic contact sheets. Movement-led quote, popup, and comparison families can use original Bauhaus, Art Nouveau, Memphis Milano, Brutalist, Japandi, or Art Deco composition language while still inheriting the project’s palette and typography. Documentary explainers can use original archival cutout collages, landmark chapter openers, event reconstructions, map and route diagrams, evidence boards, and compact sourced annotations. Partial-screen popup blocks, logo slide reveals, balanced comparisons, and kinetic word or language systems cover shorter supporting beats. Their curated inline artwork can adapt to subjects such as ideas, attention, money, conversation, culture, or technology without depending on remote assets. You can also request a reusable art direction such as OpenAI-inspired, Anthropic-inspired, or Comic; themes coordinate palette, typography, corners, borders, shadows, and motion without copying official logos or interface assets. Review the live motion preview directly in chat, iterate on the draft, and save the final version to Brand Kits -> My Brand -> Motion for reuse.

A practical editing flow

If you are editing a clip manually, a good default flow is:
  1. Click Edit Clip to open the clip in Studio.
  2. Review the preview and scrub the timeline so you understand the current cut.
  3. Start in Transcript if the issue is driven by the spoken content.
  4. Highlight sections to Cut from Video, Mute, or Edit Captions as needed.
  5. Check Subtitles, Text, and Rich Text timing so the on-screen messaging still matches the video.
  6. Use AI Tools for broader cleanup passes such as silence removal or filler-word cleanup.
  7. Use Reframe if the subject needs better positioning in the final composition.
  8. Scrub to the frame you want people to see first, open Thumbnail Editor, and save the designed cover.
  9. Finish with Export or Share once the clip looks right.

How this relates to workflows

Workflows are still where you build repeatable automation for generating clips at scale. Studio is where you manually polish a specific output after it has been created. Use workflows when you want the same editing logic to run automatically across incoming content. Use Studio when you want hands-on control over one clip before publishing.