CUBSTER Join the waitlist
DOCS / TRANSCRIPTION

Transcription

Cubster's transcription tool turns an audio file into two durable things: the audio itself, stored as a first-class kind: "audio" asset, and a speaker-labelled transcript (text, segments, per-word timing, the raw provider response) in the database. The provider — Deepgram Nova-3 today — sits behind an adapter seam, so it can be swapped without changing the API's contract.

Overview

Two ways in — upload-and-transcribe in one call for short audio, or a three-call chunked upload for anything longer — plus list, get, delete, and retry. Everything is workspace-scoped and kept — a failed provider call never loses the audio, it just leaves the transcript row failed with a retry URL.

POST /api/v1/transcribe

Session or key — a key credential must include the transcription service (see Authentication).

multipart/form-data with a file field (required) plus optional fields: name (display name, defaults to the filename), language (a BCP-47 tag like en or en-US, default en), and diarize ("true" or "false", default true). Responds 201 with the full Transcript JSON on success.

$ curl -X POST https://app.cubster.dev/api/v1/transcribe \
  -H "Authorization: Bearer $CUBSTER_API_KEY" \
  -F "file=@standup.m4a" \
  -F "name=Daily standup" \
  -F "language=en" \
  -F "diarize=true"
The 4 MB cap — and why

Audio is capped at 4 MB per file on transcribe, enforced app-side. Netlify's synchronous functions cap request bodies at 6 MB (about 4.5 MB once you account for binary-to-base64 overhead) and run for at most 10 seconds — 4 MB stays comfortably inside both limits with room for the multipart envelope and the provider round trip. Record a shorter clip or lower the bitrate to fit under it, or use chunked upload below for anything longer.

Long audio

The 4 MB cap is only half of why transcribe is bounded — the other half is that Netlify's synchronous functions are killed at 10 seconds, and a 30-minute file takes the provider roughly a minute to transcribe. Chunking the upload alone doesn't fix that second bound, so long audio needs both a way past the body limit and a way past the timeout: a three-call chunked upload past the first, and a background job — off the request path entirely — past the second.

The client splits the finished audio file into fixed-size byte ranges — not audio segments — and sends them in three calls:

  1. POST /api/v1/uploads/chunked (Content-Type: application/json — a namespace of its own, deliberately not on /api/v1/uploads, which is the first-party multipart asset upload) with filename, contentType, totalBytes, partCount, and a lowercase-hex sha256 of the whole file. Returns 201 with uploadId and an expiresAt one hour out.
  2. PUT /api/v1/uploads/chunked/:uploadId/parts/:index — raw application/octet-stream bytes for that range, in any order. Re-sending an index overwrites it, so retrying a dropped part is just re-sending it.
  3. POST /api/v1/uploads/chunked/:uploadId/complete — Cubster concatenates every part in order and verifies the reassembled file's SHA-256 against the digest sent in step 1, so the file the provider transcribes is guaranteed byte-identical to the one the client split — one whole file, never a stitched-together approximation, which is why diarization stays global and the timeline has no seams at the part boundaries. Responds 202 immediately with a pending transcript — it does not wait for the provider.

From there, poll GET /api/v1/transcripts/:id — the same endpoint below — until status leaves pending. An upload session lives for one hour, and the whole file is capped at 64 MB total (roughly three hours at 48 kbps) — well past that expiry or cap and the begin/complete calls start returning 410/413. POST /api/v1/transcribe above is unchanged and still the right call for anything under 4 MB — chunked upload is strictly for audio that doesn't fit there.

$ curl -X POST https://app.cubster.dev/api/v1/uploads/chunked \
  -H "Authorization: Bearer $CUBSTER_API_KEY" -H "Content-Type: application/json" \
  -d '{"filename":"standup.m4a","contentType":"audio/mp4","totalBytes":12582912,"partCount":3,"sha256":"9f86d081…"}'
$ curl -X PUT https://app.cubster.dev/api/v1/uploads/chunked/$UPLOAD_ID/parts/0 \
  -H "Authorization: Bearer $CUBSTER_API_KEY" -H "Content-Type: application/octet-stream" \
  --data-binary @part-0.bin
$ curl -X POST https://app.cubster.dev/api/v1/uploads/chunked/$UPLOAD_ID/complete \
  -H "Authorization: Bearer $CUBSTER_API_KEY"
Errors

Every failure mode gets its own status and a specific message. A provider failure (422/500/502/503/504) still keeps the uploaded audio — the response body carries { error, transcript: { id, status: "failed", assetId, retryUrl } }, so you can retry without re-uploading.

401 Missing or invalid credential — no session, no key, or the key was rejected.
400 Bad language tag or diarize value, e.g. Invalid language 'xx-toolong'; use a BCP-47 tag like en or en-US.
413 Audio exceeds the 4 MB limit (Netlify synchronous function ceiling). Record a shorter clip or lower the bitrate.
415 Unsupported audio type '<type>'. Accepted: mp3, mp4/x-m4a/m4a, aac, wav/x-wav/wave, webm, ogg, flac.
422 Transcription provider rejected the audio: <provider message> — e.g. a file that isn't actually audio.
500 Not configured (DEEPGRAM_API_KEY unset) or credentials rejected (check DEEPGRAM_API_KEY).
502 Transcription provider error, or the provider returned a response Cubster couldn't parse.
503 Transcription provider is rate limiting; retry shortly.
504 Timed out after 8.5s. The audio was saved — retry with POST <retryUrl>.
403 The API key doesn't include the 'transcription' service. Body names the missing service and what's enabled — create a key with that service at app.cubster.dev/account/keys.
Transcripts: list, get, delete, retry

GET /api/v1/transcripts — list, newest first. Returns { transcripts: [...] }, where each item is the list-view shape (below). Query: limit (max 100) and status (pending, completed, or failed).

GET /api/v1/transcripts/:id — detail, including words. DELETE /api/v1/transcripts/:id — deletes the transcript row and its audio asset (row + blob) together. Either way, an id from another workspace — or a non-UUID — resolves to a plain 404, never a 403.

POST /api/v1/transcripts/:id/retry — re-runs transcription for a pending or failed transcript, reusing the audio already stored — no re-upload. A completed transcript responds 409 { "error": "Transcript is already completed" }.

GET /audio/:key

Public, by opaque key (no listing endpoint — the key is unguessable). Served inline (Content-Disposition: inline) with HTTP Range support (Accept-Ranges: bytes), so an <audio> element can play and seek it directly — Safari in particular won't play at all without Range support. A single Range: bytes=a-b request returns 206 with Content-Range; malformed or unsatisfiable ranges get 416. Responses are never cached (Cache-Control: private, no-store) — deleting a transcript revokes its audio immediately.

Transcript JSON (v1)

The durable artifact returned by transcribe, transcripts.get, and a successful retry. speaker is an integer index (0, 1, …) or null — never a name; naming speakers is downstream's job. List responses (GET /api/v1/transcripts) omit text/segments/words and add a preview (first 200 characters of text) instead; detail responses (GET /api/v1/transcripts/:id, and the response from transcribe/retry) include words.

A transcript's source.kind isn't always "upload" — a transcript can also come from a live realtime voice session (source.kind: "live"), which has no audio asset at all.

transcript.json
{
  "id": "uuid", "workspaceId": "uuid", "assetId": "uuid",
  "displayName": "standup.m4a", "version": 1, "status": "completed", "error": null,
  "language": "en", "diarize": true,
  "source": { "kind": "upload", "filename": "standup.m4a", "contentType": "audio/mp4",
    "size": 812345, "durationSeconds": 93.4, "url": "/audio/<key>.m4a" },
  "provider": { "name": "deepgram", "model": "nova-3", "requestId": "…" },
  "text": "full plain text…", "speakerCount": 2,
  "segments": [{ "speaker": 0, "start": 0.0, "end": 4.2, "text": "…", "confidence": 0.98 }],
  "words": [{ "word": "Hello", "start": 0.0, "end": 0.3, "speaker": 0, "confidence": 0.99 }], // detail only
  "createdAt": "ISO", "updatedAt": "ISO"
}
SDK

client.transcribe() mirrors uploads.upload() — even on a provider failure it never throws away the file. A failed attempt throws a CubsterError whose transcriptFailure() carries { id, status, assetId, retryUrl } for transcripts.retry. transcripts.get/delete/retry accept an optional { workspace } scope, same as assets.get/update/delete.

const transcript = await cubster.transcribe({
  file: new Blob([buffer], { type: 'audio/mp4' }),
  filename: 'standup.m4a',
  name: 'Daily standup',
})
try {
  await cubster.transcribe({ file, filename: 'clip.m4a' })
} catch (err) {
  const failure = err.transcriptFailure() // { id, status, assetId, retryUrl } | null
  if (failure) await cubster.transcripts.retry(failure.id)
}
CLI

cubster transcribe <path> rejects an oversized file locally, before uploading anything. On failure it prints a retry hint with the transcript id.

$ cubster transcribe ./standup.m4a --language en-US
Transcribed "standup.m4a" (t1) — completed, 1:33, 2 speaker(s)
[0] 0:00 Hi there.
[1] 1:33 Bye.
$ cubster transcripts ls
$ cubster transcripts retry t1