Skip to content

How Editory works

Every video in Editory follows the same six steps, whether you drove it by hand, pressed Auto Generate, or asked an AI client to do it. Knowing the shape is the fastest way to work out where to intervene when something isn’t what you wanted.

flowchart LR
A["1 · Ingest<br/>your story"] --> B["2 · Generate a<br/>script"]
B --> C["3 · Review<br/>& edit"]
C --> D["4 · Record"]
D --> E["5 · Apply<br/>your brand"]
E --> F["6 · Export<br/>for social"]

Paste a URL, upload a PDF or DOCX, pick something from your library — or just describe the story and let Editory write it. That last mode has no article behind it at all: you supply the brief (“announce our partnership with X, point people at their Instagram”) and Editory drafts the facts from what you gave it.

When you paste a link, Editory also pulls the article’s photos. The main photo is selected by default and the rest are one click away.

You choose a persona — News Anchor, Conversational Storyteller, Investigative Reporter, Educational Broadcaster — or take advanced control of duration, speaking style, story focus, depth, tone, pacing and emphasis.

The target length is derived from the runtime you asked for, at roughly 150 words per minute. Ask for 60 seconds and you get about 150 words, not “a short script”.

The editing chat takes plain-English revisions: tone shifts, tightening, restructuring, length changes. Every iteration is kept, so undo always works.

Three routes, and they produce the same kind of input as far as the rest of the pipeline is concerned:

Route What it is
Teleprompter Records in your browser, script scrolling at your pace, with an eye-level alignment guide
Phone as camera Your desktop mints a short-lived link; your phone opens it, gets the script and settings, and records with the better camera
AI voiceover No filming. The script becomes narration, and the video is built from b-roll and graphics

This is where the pipeline does most of its work, and almost none of it needs your attention the second time around.

Stage What it does
Transcription Word-level timing, so captions land on the syllable. Chunk boundaries are extended outward so a cut never clips a word
Shorts planning Picks the breakpoints for short-form clips
B-roll planning Decides where a cutaway belongs and writes the search terms for it
B-roll fetching Pulls from stock providers or your own uploads, scores for relevance, and keeps the credit line
Graphics Quote cards, stat displays, speaker intros, title cards, data tables, flow steps and metric rings
Captions Styled to your preset, plus up to three suggested social captions for the post itself
Music & SFX Generated beds and stings
Assembly Everything becomes one manifest — the single description of the finished video
Render Encoded to H.264 and uploaded

You get the long-form cut plus any short clips, in whichever of 9:16, 16:9 and 1:1 you asked for. From there: download, or publish to a connected account.

The six steps are fixed; how much of each you control is not.

  • Change the words, not the video — reopen the script and edit it, then re-cut.
  • Change the video, not the words — open the timeline editor, or describe the change in Studio chat (“bigger captions, drop the music”).
  • Change it for every future video — set a brand kit, a caption preset and a default for each preset library.
  • Stop making the decisions at all — set up a newsroom series with a source and a schedule, and it runs without you.