Upload a video and write your own prompt. Ask for a timed transcript, a summary, translated dialogue, a scene breakdown or clean JSON — the prompt you write decides the result, and nothing is fixed to our format but the video itself.
Your own prompt · millisecond timestamps · speaker attribution · JSON output
The same call handles a 30-second clip and a feature-length episode.
One request with the file. Attach images or documents alongside it when the task needs context — a reference sheet, a previous script, a style guide.
A transcript with timestamps, a one-paragraph summary, a translation, a list of scenes, a compliance check, or JSON with the exact field names your system expects. It is entirely your prompt.
The response is the answer as text. Ask for JSON and it parses straight into your pipeline; ask for prose and it is ready to hand to a person.
# transcript prompt: "Transcribe this video with timestamps in MM:SS.mmm." # summary prompt: "Summarise this episode in 150 words for a TV listing." # structure you define prompt: "Return JSON. One object per scene with start, end, location, characters present and a one-line description." # dubbing script prompt: "Segment for dubbing. JSON with start, end, character, gender, type, emotion, text_original and text_target in Khmer." → one file in, the answer you asked for out
This is not a chat assistant. Every request is independent: you send a video and a prompt, you get one answer back. Nothing carries into the next call, and there is no conversation to continue or come back to.
There is no session, no conversation id and no thread to keep track of. One request in, one answer out — then it is done.
A second request knows nothing about the first, and your results are never used to build up context on our side.
If a task needs earlier material — a previous script, a house style, an earlier answer — put it in the prompt. You decide what is in front of the model on every call.
These are starting points, not product limits — edit them or write your own from scratch.
"Transcribe with start and end times to the millisecond."
"Break into caption lines of at most 42 characters with their times."
"Label who is speaking on every line and estimate their gender."
"Tag each line with emotion and whether it is lip-sync, off-screen or voiceover."
"Translate each line into Khmer so it fits the original duration."
"List every scene with its time range, location and who appears."
"Write a 100-word recap and a one-line hook for this episode."
"Return JSON with the exact field names I list below."
"Flag any segment containing violence or strong language, with timestamps."
Nothing is hard-coded to a single task. Keep one prompt per job — transcripts, recaps, compliance, extraction — and reuse it unchanged.
Line-accurate start and end times, so subtitles and audio land on the performance rather than near it.
Segments break on speaker changes and real pauses, never mid-sentence — which is what makes a script recordable.
Feature-length and episodic video is processed as a job and reassembled with one call, with timing preserved across every join.
Name the fields you want and the reply parses directly — no scraping prose, no fragile post-processing.
Each response carries an id, so any result in your logs traces back to the exact request that produced it.
Timed, translated dialogue scripts ready to record.
Caption tracks from any source language, timed and split correctly.
Episode summaries and hooks for catalogues and listings.
Search and flag long footage for topics, names or claims.
Transcribe a back catalogue and make it searchable.
Flag segments that need a human eye, with exact timestamps.
Describe cuts, on-screen text and product placement.
Lecture transcripts, chapter markers and study notes.
Every plan includes the full API, every input type and every output format. Billed on usage.
Enough processing to run real footage through the API and judge the output.
Higher throughput and priority processing for teams shipping to a schedule.
Volume pricing and the terms a content pipeline needs.
Tell us how many hours of footage you process and we will quote within a day.
No. There is no conversation and no memory. Each request is one prompt in and one answer out, and the next request starts from nothing — that is why results stay repeatable and requests can be retried safely. If you need follow-up behaviour, include the earlier turns inside your prompt.
No. You send whatever instructions you want and the response is that answer. Our recipes are starting points you can copy, edit or ignore entirely.
Anything you can describe: a timed transcript, caption lines, speaker labels, emotion tags, translated dialogue, a scene list, a summary, a compliance flag, or JSON with field names you choose.
Transcription works across the languages we see in practice, and you can ask for a translation into the language you need. Khmer is the one we are asked for most, and the one our dubbing recipes are tuned for.
Yes. Long video is processed as a job and you fetch one combined result at the end, with timing continuous across the whole title — so it is one request, not a hundred.
Short clips finish inside a minute. An episode takes a few minutes. Length of footage drives the time more than file size. Requests are queued and processed in order, so sending many at once does not speed things up.
It is used only to produce the result you asked for and is removed after processing. We do not publish, share or train on your content.
Access is by invitation. Ask us for one and we will give you enough processing to run a real title through the API before you commit.
With the prompt you would really use. You will know within an afternoon whether the output is good enough to build on.