Video transcription & analysis API

Send a video. Get back exactly what you asked for.

Upload a video and write your own prompt. Ask for a timed transcript, a summary, translated dialogue, a scene breakdown or clean JSON — the prompt you write decides the result, and nothing is fixed to our format but the video itself.

Your own prompt · millisecond timestamps · speaker attribution · JSON output

1
endpoint for any video task
Yours
the prompt, not ours
4
input types, mixed
JSON
drops into your pipeline
How it works

Three steps, whatever the task.

The same call handles a 30-second clip and a feature-length episode.

Send your video

One request with the file. Attach images or documents alongside it when the task needs context — a reference sheet, a previous script, a style guide.

Write the prompt you need

A transcript with timestamps, a one-paragraph summary, a translation, a list of scenes, a compliance check, or JSON with the exact field names your system expects. It is entirely your prompt.

Get the result back

The response is the answer as text. Ask for JSON and it parses straight into your pipeline; ask for prose and it is ready to hand to a person.

the same endpoint, four different prompts
# transcript
prompt: "Transcribe this video with timestamps in MM:SS.mmm."

# summary
prompt: "Summarise this episode in 150 words for a TV listing."

# structure you define
prompt: "Return JSON. One object per scene with start, end,
          location, characters present and a one-line description."

# dubbing script
prompt: "Segment for dubbing. JSON with start, end, character,
          gender, type, emotion, text_original and text_target in Khmer."

→ one file in, the answer you asked for out
How it behaves

One request, one answer. No memory.

This is not a chat assistant. Every request is independent: you send a video and a prompt, you get one answer back. Nothing carries into the next call, and there is no conversation to continue or come back to.

Independent

Each call stands alone

There is no session, no conversation id and no thread to keep track of. One request in, one answer out — then it is done.

Stateless

Nothing is remembered

A second request knows nothing about the first, and your results are never used to build up context on our side.

Yours

Context is yours to send

If a task needs earlier material — a previous script, a house style, an earlier answer — put it in the prompt. You decide what is in front of the model on every call.

Why that suits a pipeline: with no hidden state, the same input produces the same kind of output, requests can be retried or run in any order, and a failed job never leaves a half-remembered conversation behind.
What you can ask for

Your prompt is the interface.

These are starting points, not product limits — edit them or write your own from scratch.

Timing

Timed transcript

"Transcribe with start and end times to the millisecond."

Captions

Subtitle lines

"Break into caption lines of at most 42 characters with their times."

People

Speaker attribution

"Label who is speaking on every line and estimate their gender."

Direction

Tone and delivery

"Tag each line with emotion and whether it is lip-sync, off-screen or voiceover."

Translation

Translated dialogue

"Translate each line into Khmer so it fits the original duration."

Structure

Scene breakdown

"List every scene with its time range, location and who appears."

Insight

Summary or recap

"Write a 100-word recap and a one-line hook for this episode."

Extraction

Structured JSON

"Return JSON with the exact field names I list below."

Review

Content checks

"Flag any segment containing violence or strong language, with timestamps."

Features

Built for unattended pipelines.

Prompts

Bring your own

Nothing is hard-coded to a single task. Keep one prompt per job — transcripts, recaps, compliance, extraction — and reuse it unchanged.

Timing

Millisecond timestamps

Line-accurate start and end times, so subtitles and audio land on the performance rather than near it.

Segmentation

Cut at natural pauses

Segments break on speaker changes and real pauses, never mid-sentence — which is what makes a script recordable.

Long video

Full-length titles

Feature-length and episodic video is processed as a job and reassembled with one call, with timing preserved across every join.

Output

JSON on request

Name the fields you want and the reply parses directly — no scraping prose, no fragile post-processing.

Traceability

Every call identified

Each response carries an id, so any result in your logs traces back to the exact request that produced it.

Use cases

What teams run through it.

Dubbing & localisation

Timed, translated dialogue scripts ready to record.

Subtitling

Caption tracks from any source language, timed and split correctly.

Series recaps

Episode summaries and hooks for catalogues and listings.

Media monitoring

Search and flag long footage for topics, names or claims.

Archive indexing

Transcribe a back catalogue and make it searchable.

Compliance review

Flag segments that need a human eye, with exact timestamps.

Creative review

Describe cuts, on-screen text and product placement.

Education

Lecture transcripts, chapter markers and study notes.

Pricing

Plans that scale with your library.

Every plan includes the full API, every input type and every output format. Billed on usage.

Starter

For pilots

Enough processing to run real footage through the API and judge the output.

  • Full API access
  • All input types
  • Email support
Create account
Professional

For continuing series

Higher throughput and priority processing for teams shipping to a schedule.

  • Everything in Starter
  • Priority processing
  • Full-episode jobs
  • Priority support
Create account
Enterprise

For catalogues

Volume pricing and the terms a content pipeline needs.

  • Everything in Professional
  • Volume pricing
  • Service agreement
  • Named support contact
Talk to us

Tell us how many hours of footage you process and we will quote within a day.

Questions

Before you start.

Is this a chatbot I can have a conversation with?

No. There is no conversation and no memory. Each request is one prompt in and one answer out, and the next request starts from nothing — that is why results stay repeatable and requests can be retried safely. If you need follow-up behaviour, include the earlier turns inside your prompt.

Do I have to use your prompt?

No. You send whatever instructions you want and the response is that answer. Our recipes are starting points you can copy, edit or ignore entirely.

What can I get back?

Anything you can describe: a timed transcript, caption lines, speaker labels, emotion tags, translated dialogue, a scene list, a summary, a compliance flag, or JSON with field names you choose.

Which languages are supported?

Transcription works across the languages we see in practice, and you can ask for a translation into the language you need. Khmer is the one we are asked for most, and the one our dubbing recipes are tuned for.

Can I send a feature-length video?

Yes. Long video is processed as a job and you fetch one combined result at the end, with timing continuous across the whole title — so it is one request, not a hundred.

How long does processing take?

Short clips finish inside a minute. An episode takes a few minutes. Length of footage drives the time more than file size. Requests are queued and processed in order, so sending many at once does not speed things up.

What happens to my footage?

It is used only to produce the result you asked for and is removed after processing. We do not publish, share or train on your content.

Is there a trial?

Access is by invitation. Ask us for one and we will give you enough processing to run a real title through the API before you commit.

Send us one video.

With the prompt you would really use. You will know within an afternoon whether the output is good enough to build on.