Transcription
Cubster's transcription tool turns an audio file into two durable things: the audio itself, stored as a
first-class kind: "audio" asset, and a speaker-labelled transcript (text, segments, per-word timing, the raw provider response) in
the database. The provider — Deepgram Nova-3 today — sits behind an adapter seam, so it can be swapped
without changing the API's contract.
Two ways in — upload-and-transcribe in one call for short audio, or a three-call chunked upload for
anything longer — plus list, get, delete, and retry. Everything is workspace-scoped and kept — a
failed provider call never loses the audio, it just leaves the transcript row
failed
with a retry URL.
Session or key — a key credential must include the transcription service (see Authentication).
multipart/form-data with a file
field (required) plus optional fields:
name
(display name, defaults to the filename),
language
(a BCP-47 tag like en
or en-US,
default en),
and diarize
("true"
or "false",
default true). Responds 201
with the full Transcript JSON on success.
Audio is capped at 4 MB per file on
transcribe, enforced app-side. Netlify's synchronous functions cap request bodies at 6 MB (about 4.5 MB once
you account for binary-to-base64 overhead) and run for at most 10 seconds — 4 MB stays comfortably
inside both limits with room for the multipart envelope and the provider round trip. Record a
shorter clip or lower the bitrate to fit under it, or use chunked upload
below for anything longer.
The 4 MB cap is only half of why transcribe is bounded — the other half is that Netlify's synchronous functions are killed at 10 seconds, and
a 30-minute file takes the provider roughly a minute to transcribe. Chunking the upload alone doesn't
fix that second bound, so long audio needs both a way past the body limit and a way past the
timeout: a three-call chunked upload past the first, and a background job — off the request path
entirely — past the second.
The client splits the finished audio file into fixed-size byte ranges — not audio segments — and sends them in three calls:
-
POST /api/v1/uploads/chunked(Content-Type: application/json— a namespace of its own, deliberately not on/api/v1/uploads, which is the first-party multipart asset upload) withfilename,contentType,totalBytes,partCount, and a lowercase-hexsha256of the whole file. Returns201withuploadIdand anexpiresAtone hour out. -
PUT /api/v1/uploads/chunked/:uploadId/parts/:index— rawapplication/octet-streambytes for that range, in any order. Re-sending an index overwrites it, so retrying a dropped part is just re-sending it. -
POST /api/v1/uploads/chunked/:uploadId/complete— Cubster concatenates every part in order and verifies the reassembled file's SHA-256 against the digest sent in step 1, so the file the provider transcribes is guaranteed byte-identical to the one the client split — one whole file, never a stitched-together approximation, which is why diarization stays global and the timeline has no seams at the part boundaries. Responds202immediately with apendingtranscript — it does not wait for the provider.
From there, poll GET /api/v1/transcripts/:id — the same endpoint below — until status leaves pending. An upload session lives for one hour, and
the whole file is capped at 64 MB total
(roughly three hours at 48 kbps) — well past that expiry or cap and the begin/complete calls start
returning 410/413. POST /api/v1/transcribe above is unchanged and still the right call for anything under 4 MB — chunked upload is strictly
for audio that doesn't fit there.
Every failure mode gets its own status and a specific message. A provider failure (422/500/502/503/504)
still keeps the uploaded audio — the response body carries
{ error, transcript: { id, status: "failed", assetId, retryUrl } }, so you can retry without re-uploading.
| 401 | Missing or invalid credential — no session, no key, or the key was rejected. |
| 400 | Bad language tag or diarize value, e.g. Invalid language 'xx-toolong'; use a BCP-47 tag like en or en-US. |
| 413 | Audio exceeds the 4 MB limit (Netlify synchronous function ceiling). Record a shorter clip or lower the bitrate. |
| 415 | Unsupported audio type '<type>'. Accepted: mp3, mp4/x-m4a/m4a, aac, wav/x-wav/wave, webm, ogg, flac. |
| 422 | Transcription provider rejected the audio: <provider message> — e.g. a file that isn't actually audio. |
| 500 | Not configured (DEEPGRAM_API_KEY unset) or credentials rejected (check DEEPGRAM_API_KEY). |
| 502 | Transcription provider error, or the provider returned a response Cubster couldn't parse. |
| 503 | Transcription provider is rate limiting; retry shortly. |
| 504 | Timed out after 8.5s. The audio was saved — retry with POST <retryUrl>. |
| 403 | The API key doesn't include the 'transcription' service. Body names the missing service and what's enabled — create a key with that service at app.cubster.dev/account/keys. |
GET /api/v1/transcripts
— list, newest first. Returns { transcripts: [...] }, where each item is the list-view shape (below). Query: limit (max 100) and status (pending,
completed, or
failed).
GET /api/v1/transcripts/:id
— detail, including words.
DELETE /api/v1/transcripts/:id
— deletes the transcript row and its audio asset (row + blob) together. Either way, an id from another
workspace — or a non-UUID — resolves to a plain 404,
never a 403.
POST /api/v1/transcripts/:id/retry
— re-runs transcription for a pending
or failed
transcript, reusing the audio already stored — no re-upload. A completed transcript responds 409 { "error": "Transcript is already completed" }.
Public, by opaque key (no listing endpoint — the key is unguessable). Served inline
(Content-Disposition: inline) with HTTP Range support (Accept-Ranges: bytes), so an <audio> element can play and seek it directly — Safari in particular won't play at all without Range support.
A single Range: bytes=a-b request returns 206
with Content-Range;
malformed or unsatisfiable ranges get 416.
Responses are never cached (Cache-Control: private, no-store) — deleting a transcript revokes its audio immediately.
The durable artifact returned by transcribe, transcripts.get, and a successful retry. speaker
is an integer index (0, 1,
…) or null
— never a name; naming speakers is downstream's job. List responses (GET /api/v1/transcripts) omit text/segments/words
and add a preview (first 200 characters of text) instead; detail responses (GET /api/v1/transcripts/:id, and the response from transcribe/retry)
include words.
A transcript's source.kind isn't always "upload" — a transcript can also come from a live realtime voice session
(source.kind: "live"), which has no audio asset at all.
client.transcribe() mirrors uploads.upload() — even on a provider failure it never throws away the file. A failed attempt throws a CubsterError whose transcriptFailure() carries { id, status, assetId, retryUrl } for transcripts.retry. transcripts.get/delete/retry accept an optional { workspace } scope, same as assets.get/update/delete.
cubster transcribe <path> rejects an oversized file locally, before uploading anything. On failure it prints a retry hint with
the transcript id.