Getting access
Access is by invitation. Get in touch and we will send you a code starting with inv_ — then create your account with it, and your API key is ready immediately.
An invitation may be tied to your email address and may carry a request allowance. Once inside, you can create and revoke keys from the account page, and Usage shows how much of your allowance is left.
Processing a series or a whole catalogue? Talk to us about Professional and Enterprise throughput.
Quick start
Send the video and your prompt in one request. The response is the answer as text.
curl -X POST https://169-58-121-148.sslip.io/ask \ -H "x-api-key: $FRAMEWISE_KEY" \ -F "prompt=Transcribe this video with timestamps in MM:SS.mmm and label who is speaking on each line." \ -F "file=@clip.mp4" \ -F "timeoutMs=600000"
{"ok":true,
"reply":"00:00.000 Masked Young Man: …",
"callId":"call-9f2c4d1e-…"}
One-shot, not a chat
This API is not a chat assistant and it does not hold conversations. Every request is independent: one prompt and one video in, one answer out. There is no session, no conversation id, no thread, and no memory of anything you sent before.
Carrying context yourself
Continuity is yours to construct. Anything the model needs in front of it goes inside the prompt on that call — a previous script, a glossary, a house style, an earlier answer:
prompt: "Here is the character glossary from the series bible: Evelyn — female, lead, calm under pressure Alex — male, younger brother, impulsive Now transcribe and segment the attached video for dubbing, using those names and translating into Khmer." + file: episode.mp4
This is how to get the effect of a conversation without one: rebuild the context you want on every request. It also means you can replay a request months later and get a comparable result, which a live conversation cannot do.
Why stateless is the right shape for this
- Repeatable. The same input produces the same kind of output, whatever ran before it.
- Retryable. A failed request can be sent again with no risk of a half-remembered history.
- Order-independent. A hundred titles can be processed in any order, or in parallel across accounts.
- Private by construction. Nothing accumulates in a conversation we would have to keep.
Writing the prompt
The prompt is the whole interface. There is no required template, and the same endpoint behaves completely differently depending on what you ask. Four things make a prompt work well here:
- Say what to produce. "Transcribe with timestamps" and "summarise in 150 words" are different jobs from the same file.
- Name your fields when you want JSON. Listing the exact keys you need is what makes the reply parse without cleanup. When you ask for JSON, reply parses — json.loads(reply) will not raise, and a stray JSON line or code fence the model adds is stripped for you.
- Give the time format. Ask for MM:SS.mmm and the timestamps come back in it.
- State the language. Say which language to transcribe in, and which to translate into, when translation is the point.
Return a JSON array. One object per dialogue line with exactly these fields:
start exact start time, MM:SS.mmm
end exact end time, MM:SS.mmm
character who is speaking
gender male or female
type lip-sync, off-screen or voiceover
emotion the delivery, e.g. neutral, angry, taunting
text_original the line as spoken
text_target the same line translated into Khmer,
written to fit the original duration
Split on natural pauses and speaker turns. Do not merge separate speakers.
Ready-made prompts
Starting points you can send as-is or edit. None of them are required.
Timed transcript
Transcribe this video. Return JSON with one object per line:
{ "start": "MM:SS.mmm", "end": "MM:SS.mmm", "speaker": "…", "text": "…" }
Split at natural pauses and speaker changes.Subtitles
Transcribe this video into subtitle lines. Maximum 42 characters per line, never split mid-sentence. Return JSON with "start", "end" and "text".
Episode recap
Summarise this episode in 150 words for a TV listing, then write a single sentence hook. Plain text, no headings.
Dubbing script
Transcribe and segment this video for dubbing. Return a JSON array with start, end, character, gender, type, emotion, text_original, and text_target translated into Khmer and adapted to fit each line's duration. Keep segments separated by natural pauses and speaker turns.
Scene breakdown
Return JSON, one object per scene, with "start", "end", "location", "characters" (array) and "description" (one line).
Content check
Review this video and return JSON listing any segment containing violence, strong language or unsafe behaviour, with "start", "end" and "reason". Return an empty array if there are none.
Sending media
Send multipart/form-data with one file field per upload — up to eight — plus your prompt. Only video is required; the others are context you can add when the task needs it.
- Video —
mp4,mov,webm - Image —
jpg,png,webp - Documents —
pdf
curl -X POST https://169-58-121-148.sslip.io/ask \ -H "x-api-key: $FRAMEWISE_KEY" \ -F "prompt=Use the style sheet for character names. Transcribe the video." \ -F "file=@episode.mp4" \ -F "file=@style-sheet.png"
Long video
A feature-length title is handled as a job: you start it, send the segments it describes, and fetch one combined result at the end. Timing runs continuously across the joins, so the output reads as one script from the first line to the last.
1. Start the job
curl -X POST https://169-58-121-148.sslip.io/jobs \ -H "x-api-key: $FRAMEWISE_KEY" -H "content-type: application/json" \ -d '{"label":"Ep 87","durationMs":3600000,"chunkSeconds":300}' "plan": [ {"index":0,"id":"job-8a66d77014-c000","startMs":0, "start":"00:00.000","end":"05:00.000"}, {"index":1,"id":"job-8a66d77014-c001","startMs":300000, "start":"05:00.000","end":"10:00.000"}, … ]
Divide the title at those offsets. Each segment keeps its position in the whole film.
2. Send each segment with your prompt
curl -X POST https://169-58-121-148.sslip.io/jobs/<jobId>/chunks \ -H "x-api-key: $FRAMEWISE_KEY" \ -F "index=1" -F "startMs=300000" \ -F "prompt=Transcribe and segment for dubbing. text_target in Khmer." \ -F "timeoutMs=600000" \ -F "file=@part001.mp4;type=video/mp4"
3. Fetch the combined result
curl https://169-58-121-148.sslip.io/jobs/<jobId>/merged \ -H "x-api-key: $FRAMEWISE_KEY" {"complete":true,"chunksDone":12,"chunksPlanned":12,"missing":[], "entries":[{"start":"00:00.000","end":"00:01.000","character":"Evelyn", …]}
- complete is true only when every segment finished.
- missing lists any that did not, so a partial result never passes for a finished one.
- Timing is cumulative: a line at 00:02 in segment two comes back as 05:02 in the title.
- Results are stored, so you can resume without resending what already completed.
Tracking progress
Long media takes a while, so a request can report what it is doing while you wait. Send a progressId with the request and poll it.
{"ok":true,"status":"running",
"steps":[{"name":"processing","detail":"segment 3 of 12"},
{"name":"transcribing","detail":"5381 characters so far"}]}
status is running, done or failed. Most integrations simply wait for the response and use this only for a progress bar.
Rate limits & timing
- One request at a time, per account. Requests are processed one after another, and up to five more wait in the queue while one is running. A request that arrives while the queue is full is refused straight away with
429and aRetry-After— that is the queue saying "come back in a moment", not a problem with your request. - Sending many at once does not speed anything up. The queue is strictly in order, so twelve parallel calls take the same time as twelve sent one after the other — they only risk filling the queue and being refused.
- Short clips finish inside a minute. An episode takes a few minutes. Footage length drives the time more than file size, and video is much slower than a text request.
- Back off on
429: honourRetry-Afterand add a little jitter. Every refusal also carriesretryable, so your loop can decide what to do without hard-coding codes. - Need a series processed faster? Talk to us about Professional and Enterprise throughput. For one very long title, the jobs API handles it in parts.
Errors
Failures return JSON with a stable machine-readable code. Match on the code, not the message — messages are written for humans and may be reworded.
{"ok":false,
"error":{"code":"…","message":"…"},
"callId":"call-…"}
| Status | What it means | What to do |
|---|---|---|
| 400 | The request was not usable — no prompt, bad JSON, or a malformed upload. | Fix it; do not retry unchanged. |
| 401 | The API key is missing or wrong. | Check the x-api-key header. |
| 404 | The job or resource does not exist. | Check the id. |
| 429 | Rate limited, or your queue is full. | Wait and retry with backoff. Retry-After may be set. |
| 500 | An unexpected error on our side. | Retry once; contact support with the callId if it repeats. |
| 504 | Processing did not finish inside timeoutMs. | Raise the timeout and retry. Whatever was produced so far is returned as partial. |
Best practices
- Version your prompts. Keep them beside your code so a change in output can be traced to a change in instructions.
- Send the context you need, every time. Nothing is carried between calls, so anything the model should know goes in the prompt on that request.
- Ask for JSON when a machine reads it. Naming the fields up front removes a whole layer of parsing.
- Use one job per title. It keeps timing continuous and lets you resume only what failed.
- Set a long timeout for video and a short one for text.
- Log the callId with your own identifiers.
Support
Send the callId of the affected request, roughly when it happened, what prompt you used, and what you expected. That is usually enough to resolve it without a reproduction.
Data & privacy
- Footage and prompts are used only to produce the result you asked for, and are removed after processing.
- We do not publish, share, resell or train on your content.
- Results are retained only as long as needed for you to collect them and for the lifetime of a job.
- API keys are per account. Rotate yours if it is ever exposed, and keep it server-side.